System for low-photon-count visual object detection and classification

By using a photon detection system and a low-photon counting classification model, the problems of slow speed, high power consumption, and privacy protection in visual object detection and classification under low light conditions are solved, realizing fast and low-power visual object detection and classification, which is suitable for low-light environments and battery-powered devices.

CN115812164BActive Publication Date: 2025-12-05GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080102578.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-02
Publication Date
2025-12-05
Estimated Expiration
2040-07-02

AI Technical Summary

Technical Problem

Existing visual object detection and classification systems face challenges in achieving fast, low-power, and privacy-preserving operation under low-light conditions, especially in dark rooms and night scenes where they cannot capture enough image data. This results in slow computation speeds, high power consumption, and the inability to perform object recognition without capturing detailed images.

Method used

A photon detection system and a low-photon count classification model are employed. The photon detector outputs a photon signature, and the low-photon count classification model can output a stable classification result after receiving a small number of photons. Parallel processing and sparsification techniques are combined to improve computing speed and reduce power consumption. Machine learning models such as logistic regression and Naive Bayes network are used for classification.

Benefits of technology

It enables fast, low-power visual object detection and classification under low-light conditions, and can output accurate classification in sub-milliseconds. It reduces the dependence on image data, protects privacy, and is suitable for battery-powered devices and time-critical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115812164B_ABST
    Figure CN115812164B_ABST
Patent Text Reader

Abstract

A computing system can be configured for low-photon-count visual object classification. The computing system can include a photon-detection system that includes one or more cells. Each of the one or more cells can include one or more photon detectors. Each of the one or more photon detectors can be configured to output a photon signature in response to a photon incident on the one or more photon detectors. The computing system can include one or more processors and one or more storage devices storing computer-readable data. The data can include a low-photon-count classification model and one or more instructions that, when implemented, cause the one or more processors to perform operations for low-photon-count visual object recognition. The operations can include obtaining a photon signature from the photon-detection system. The operations can include providing the photon signature to the low-photon-count classification model. The operations can include determining, by the low-photon-count classification model, a classification of a visual object placed in a field of view of the photon-detection system based at least in part on the photon signature. The operations can include providing the classification as an output of the low-photon-count classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to object detection and / or classification. More specifically, this disclosure generally relates to systems and methods for low-photon-count visual object detection and / or classification. Background Technology

[0002] Object detection and / or classification refers to using computational systems such as machine learning models to identify the presence of one or more objects in visual data such as images and / or to generate their classification. Summary of the Invention

[0003] Aspects and advantages of embodiments of this disclosure will be set forth in part in the description which follows, or may be learned from the description or by practice of the embodiments.

[0004] One example aspect of this disclosure relates to a computational system configured for classifying low-photon-count visual objects. The computational system may include a photon detection system comprising one or more units. Each of the one or more units may include one or more photon detectors. Each of the one or more photon detectors may be configured to output a photon signature in response to photon incident on the one or more photon detectors. The computational system may include one or more processors and one or more storage devices storing computer-readable data. The data may include a low-photon-count classification model and one or more instructions, when implemented, causing the one or more processors to perform operations for low-photon-count visual object identification. These operations may include obtaining a photon signature from the photon detection system. These operations may include providing the photon signature to the low-photon-count classification model. These operations may include determining, by the low-photon-count classification model, a classification of a visual object placed in the field of view of the photon detection system based at least in part on the photon signature. These operations may include providing the classification as the output of the low-photon-count classification model.

[0005] Another exemplary aspect of this disclosure relates to a computer-implemented method for classifying low-photon-count visual objects. This computer-implemented method may include obtaining a photon signature from a photon detection system by a computing system including one or more computing devices. The computer-implemented method may include providing the photon signature to a low-photon-count classification model by the computing system. The computer-implemented method may include determining a classification of a visual object placed in the field of view of the photon detection system based at least in part on the photon signature by the computing system and the low-photon-count classification model. The computer-implemented method may include providing this classification as the output of the low-photon-count classification model by the computing system.

[0006] Another exemplary aspect of this disclosure relates to a computer-implemented method for training a low-photon counting classification model configured for low-photon counting visual object recognition. This computer-implemented method may include obtaining a training dataset comprising one or more images by a computing system including one or more computing devices. The computer-implemented method may include generating time series of training examples from the one or more images by the computing system. The time series of training examples may include multiple example photon signatures derived from the one or more images. The computer-implemented method may include providing the time series of training examples to the low-photon counting classification model by the computing system. After providing a subset of the time series of training examples to the low-photon counting classification model, the computer-implemented method may include backpropagating a loss from the subset of the time series of training examples by the computing system to train the low-photon counting classification model.

[0007] Other aspects of this disclosure relate to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.

[0008] These and other features, aspects, and advantages of the various embodiments of this disclosure will be better understood by referring to the following description and the appended claims. The accompanying drawings, which are incorporated in and form a part of this specification, illustrate exemplary embodiments of the disclosure and, together with the description, serve to explain the relevant principles. Attached Figure Description

[0009] A detailed discussion of embodiments applicable to those skilled in the art is set forth in the description with reference to the accompanying drawings, wherein:

[0010] Figure 1 A block diagram of an example low-photon-count visual object classification system according to an example implementation of the present disclosure is depicted.

[0011] Figure 2 A block diagram of an example low-photon-count visual object classification system according to an example implementation of the present disclosure is depicted.

[0012] Figure 3 A block diagram of an example low-photon-count visual object classification system according to an example implementation of the present disclosure is depicted.

[0013] Figure 4 A block diagram of an example low-photon-count visual object classification system according to an example implementation of the present disclosure is depicted.

[0014] Figures 5A-5C A block diagram of an example computing system according to an example implementation of the present disclosure is depicted;

[0015] Figure 6 A block diagram depicts an example low-photon counting classification model according to an example implementation of this disclosure;

[0016] Figure 7 A flowchart is depicted illustrating an example computer implementation method for classifying low-photon-count visual objects according to an example implementation of this disclosure; and

[0017] Figure 8 A flowchart is provided illustrating an example computer implementation method for training a low-photon counting classification model configured for low-photon counting visual object recognition, according to an example implementation of this disclosure.

[0018] The repeated reference numerals in multiple figures are intended to identify the same signature in various implementations. Detailed Implementation

[0019] Overview

[0020] In general, this disclosure relates to systems and methods for detecting and / or classifying visual objects with low photon counts (e.g., ultrafast and / or low illumination). According to an example aspect of this disclosure, a photon detection system can be configured to provide input to a low photon count classification model. The photon detection system may include units (e.g., groups) of one or more photon detectors that output a photon signature (e.g., a signal indicating that a photon has been received by the photon detector, such as an electrical signature, such as an electrical spike in a signal) in response to photons incident on the photon detector. The photon signature can be provided to the low photon count classification model (e.g., as a time series, such as at almost the instant the photon signature is generated). The low photon count classification model can be trained based on the photon signature to detect, identify, and / or classify the visual objects that generated the photons. The visual objects can be, for example, text and / or numerical objects, such as text or numerical symbols, and other suitable objects. Visual objects should not be construed as necessarily being three-dimensional objects, although they can be three-dimensional. For example, a visual object can be a visual object placed within the field of view of the photon detection system. Example aspects of this disclosure can advantageously provide aspects for increasing computational speed and / or reducing the number of photons required to provide classification, enabling low-photon-count visual object classification to operate under ultrafast timing (e.g., sub-millisecond timing) and / or low-light conditions. As an example, some example implementations of example aspects of this disclosure can provide accurate classification 100 times faster and / or 10,000 times less illumination compared to some existing systems used for visual object classification. As further described herein, this can contribute to various speed, privacy, power usage, and / or other benefits.

[0021] Object detection and / or classification refers to using a computational system to classify visual objects (e.g., shapes, entities, text, etc.) that exist in computer-readable data (e.g., images, videos, etc.) and / or within the field of view of a visual sensor (e.g., a camera). For example, some existing visual object detection and / or classification systems can feed image data to a classification model (e.g., a machine learning model) that identifies visual objects depicted in image data (e.g., the subject of an image). For example, a classification model can output data indicating that a visual object falls into one of several predetermined categories.

[0022] While this approach may be useful in certain situations, these existing systems can present challenges, especially under suboptimal conditions. As an example, low lighting conditions can prevent some existing systems from capturing enough image data to perform visual object detection and / or classification, particularly over a reasonable duration. As another example, cameras in darkrooms and / or night scenes may not capture enough data, according to some existing systems, to produce interpretable images (e.g., human-interpretable and / or computer-interpretable images) for visual object detection and / or classification. For instance, generating images (e.g., RGB images), such as those suitable for classifying visual objects using some conventional image classification systems, may require capturing data associated with over a billion photons.

[0023] As another example, the computations associated with some existing systems can be very slow. For instance, performing visual object detection and / or classification on image data may require manipulating the image data at various layers, which slows down the processing. While this may not pose a challenge in many cases, some applications may struggle with such large-scale computations. For example, much like in low-light conditions, cameras that are only briefly exposed to the field of view may not capture enough data to produce an interpretable image for visual object detection and / or classification, according to some existing systems. As another example, for highly time-critical applications, obtaining classification as quickly as possible may be desirable and / or necessary. Therefore, improving the computational speed of visual object detection and / or classification may be desired.

[0024] Another example of a drawback of some existing systems is the undesirable high power usage. For instance, some existing systems may perform relatively power-intensive computations (e.g., deep neural network evaluations) and / or require the operation of relatively power-intensive sensors (e.g., cameras). In situations where power savings are desirable (e.g., battery-powered devices such as battery-powered motion sensors, security systems, etc.), these considerations negatively impact the system's power usage. Therefore, reducing power usage may be desirable.

[0025] Another example of some existing systems may include privacy protections. For example, it may be desirable and / or necessary to perform visual object identification on objects without capturing a comprehensive image of the object. As an example, it may be desirable to obtain visual object classification of a face without capturing a detailed image of the face. As another example, it may be desirable to turn off some sensors on the device (e.g., the camera) in the presence of some visual objects while protecting the privacy of the visual objects (e.g., turning off the camera in the presence of an unauthorized user).

[0026] Systems and methods according to exemplary aspects of this disclosure can provide novel ways to address these and other challenges for visual object detection and / or recognition. According to exemplary aspects of this disclosure, a computational system can be configured for classifying visual objects with low photon counts. For example, the computational system can be configured to output a reliable classification after receiving a relatively small number of photons (e.g., less than approximately 10,000 photons, such as less than approximately 1,000 photons). As an example, the classification can be updated incrementally with each incident photon and output once the classification has stabilized. For example, the classification can be considered stable after it has met a stability criterion, such as observing a stability parameter, representing the average fluctuation of the classification, below a threshold when new photons are received. For example, the stability criterion could be that the classification does not change during a period in which a certain threshold number of photons are detected.

[0027] The computing system may include a photon detection system. The photon detection system may include one or more cells. Each cell may be arranged to receive photons from different portions of the field of view of the photon detection system. Each of the one or more cells may include one or more photon detectors. For example, photon detectors may be grouped into cells. For example, photon detectors may be grouped into cells by spatial proximity and / or any other suitable grouping. As an example, photon detectors may be grouped into cell arrays (e.g., square and / or rectangular arrays). As an example, the photon detection system may include a single-photon avalanche diode (SPAD) array. For example, photon detectors may be or may include single-photon avalanche diodes. Additionally and / or optionally, the cells may form a SPAD array.

[0028] In some embodiments, each unit may have an associated signal line. Grouping photon detectors into units with shared signal lines is advantageous for computational speed because the increased number of signal lines can prevent bottlenecks or congestion if multiple photons are incident on different photon detectors simultaneously or nearly simultaneously. Additionally and / or alternatively, in some embodiments, one or more units may share signal lines (e.g., through multiplexing).

[0029] Photon detectors can be configured to output a photon signature in response to photons incident on one or more photon detectors. As an example, a photon signature can include an electrical signature, such as a spike in an electrical signal (e.g., voltage, current, etc.). For example, a photon detector can be a SPAD that outputs an electrical spike in response to an incident photon. Furthermore, a photon signature can be associated with a cell location within one or more cells. As an example, a cell location can be an array location within an array of one or more cells (e.g., a SPAD array) and / or within the cell itself. For example, a cell location can be a serial identifier, (x,y) pairs, etc. Furthermore, for example, a photon signature generated by the detector(s)(s) in any given cell having a cell location can include data such as identifying the cell by recognizing the cell location. For example, a photon signature can include any combination of an electrical signature, a cell location, additional photon data, or other suitable data. In one example embodiment, the photon signature includes only (x,y) pairs. Subsequent processing (e.g., computational systems including low-photon count classification models) can thus locate the photon signature (e.g., an electronic signature) based on cell location, allowing the low-photon count classification model to understand at what cell location the photon that led to the generation of the photon signature was received. Additionally and / or alternatively, in some implementations, the photon signature may include any other suitable information from the photon, such as, for example, the photon's frequency.

[0030] A computing system may include one or more processors and one or more memory devices for storing computer-readable data. According to this disclosure, any suitable type and architecture of processor and / or memory device may be employed, such as CPU, GPU, microcontroller, microprocessor, etc., and / or RAM, hard disk storage, solid-state storage, magnetic tape, compactor, and / or any combination thereof.

[0031] The computing system can store (e.g., in a computer-readable storage device) a low-photon counting classification model. The low-photon counting classification model can be configured to receive photon signatures (e.g., cell locations, such as (x,y) pairs) from a photon detection system. For example, the low-photon counting classification model can receive a series (e.g., a time series) of photon signatures generated at the photon detection system and transmitted to the low-photon counting classification model. Each of these photon signatures can be processed sequentially by the low-photon counting classification model (e.g., in the order of arrival). For example, a series of photon signatures can be processed sequentially to accumulate the results of the series. A classification can be generated from the accumulated results. Therefore, the low-photon counting classification model can provide classifications for visual objects within the field of view of the photon detection system from a set of candidate classifications. For example, this set of candidate classifications can be learned (e.g., by training the low-photon counting classification model) as a set of known and / or related object types.

[0032] Furthermore, in many cases, classification can be stable (and / or provided as the output of a low-photon count classification model) after the photon detection system has received a small number of photons, such as before capturing a complete image of a visual object. For example, classification can be stable after fewer than about 10,000 photons, such as fewer than about 1,000 photons, or even fewer than about 100 photons, have been incident on the photon detection system. Any suitable low-photon count classification model can be used according to the exemplary aspects of this disclosure. For example, in some implementations, low-photon count classification can be or may include one or more logistic regression models of machine learning. It should be understood that while low-photon count classification models can produce classifications based on a relatively small number of incident photons, aspects of this disclosure can be used to provide classifications for any number of photons, including, where appropriate, more than 10,000 photons.

[0033] As an example, a low-photon count classification model may include (e.g., employ) a classification trace vector. For example, the classification trace vector may be or may include a probability vector, such as a logit-likelihood probability vector. The trace vector may be an N-dimensional vector, where N is a positive integer, such as greater than 1, such as a 10-dimensional vector. For example, each dimension of the classification trace vector may correspond to a corresponding class among multiple candidate classes.

[0034] Classification can be based at least in part on classification trace vectors. For example, classification trace vectors can serve as the internal state of a low-photon count classification model that models the likelihood of a visual object within the field of view of a photon detection system belonging to each of a plurality of candidate categories. As an example, classification trace vectors can include photon signatures indicating the logit likelihood of a visual object belonging to one of the plurality of candidate categories. For example, candidate categories can be constructed by providing training data to a low-photon count classification model (e.g., for machine learning).

[0035] In some implementations, the classification trace vector can be updated with each photon signature (e.g., whenever each photon signature is generated and / or received in a low photon count classification model), such that each photon signature contributes to the classification. For example, this can allow the classification to be updated incrementally with each incident photon, making the classification available and / or stable whenever needed. Additionally and / or alternatively, the classification trace vector can be updated after a certain number of photon signatures. As an example, the initial classification trace vector can be zero. The classification vector can be updated repeatedly, for example, for each consecutive photon signature, by adding a corresponding update value to the component of the classification value. The update value can depend on the photon signature, for example, on the cell position associated with the photon signature.

[0036] In some implementations, some or all units may have associated trace vectors (e.g., unit trace vectors). For example, each unit trace vector may be associated with one or more units. As an example, a unit trace vector may correspond to only one unit and / or multiple units, such as multiple units in a group grouped according to units. The unit trace vector can individually measure the contribution of a unit or a group of units to the classification. Unit trace vectors can be combined, for example by summing (e.g., weighted summation and / or unweighted summation), to produce a classification trace vector indicating the classification, which is referred to as the "global trace vector" for illustrative purposes. For example, the global trace vector can be a combination of one or more unit trace vectors, such as a linear combination. Including unit trace vectors can facilitate parallel processing. For example, it is possible to update only the unit trace vectors without updating the global trace vector with each photon (e.g., updating the global trace vector only after multiple photons are received), which can allow parallel processing of two photons incident simultaneously in different units. For example, the components of the unit trace vector may initially be zero and / or equal (e.g., ambiguous), and these components can be updated whenever a photon is detected in one of the associated units. When a cell trace vector is associated with multiple cells, updating the cell trace vector can typically depend on which cell among those cells a photon is detected.

[0037] In some implementations, the low-photon count classification model may include a trace model. The trace model maps photon signatures to classification trace vectors. For example, the trace model may include one or more layers that map photon signatures to classification trace vectors. The layers (multiple) may be or may include a network of one or more nodes, weights, and / or biases. For example, in some embodiments, the layers (multiple) may form a single-layer trace model, such as a trace model with only a single layer of weights and / or biases. As another example, these layers may form a deep model with multiple layers. In some embodiments, one or more layers may be linear layers, such as layers that scale the input data to form a linear trace model without considering an activation function. For example, each output of a linear layer may be a weighted sum of the inputs to that layer, defined by a corresponding set of weights, optionally with corresponding biases added to the weighted sum. In some implementations, a single-layer linear trace model may be advantageous for computational speed while still achieving the desired classification performance. The weights, biases, and / or other parameters of the trace model (e.g., a linear trace model) may be learned as part of training the low-photon count classification model. In some embodiments, each cell trace vector may have a separate trace model (e.g., with a unique set of weights and / or biases) and / or a shared trace model (e.g., having the same composition for some or all cell trace vectors).

[0038] For each cell, accumulating data in the classification trace vector can be computationally more efficient than recording how many photons have been incident on that cell (e.g., where accumulation is performed for each cell), and a linear layer is applied to the recorded data specifying the number of accumulated photons for each cell. For example, this approach might require multiplying the weights by the number of recorded photons detected in each cell whenever a classification update is performed.

[0039] Additionally and / or alternatively, in some implementations, the low photon count classification model may include an embedding space model. The embedding space model maps photon signatures to (e.g., updates values ​​to) an embedding space map. The embedding space model may employ an embedding space map within an embedding space. The embedding space map may initially be zero at all points in the embedding space. For example, the embedding space may define a dimensional space, such as an M-dimensional space, where M is a positive integer, such as greater than one (e.g., a 30-dimensional space). Similar to the classification trace vector, the embedding space map may be updated for each photon signature or alternatively after multiple photon signatures. This can be done by adding the output of a linear layer for each photon signature to an existing embedding space map.

[0040] For example, an embedding space model may include one or more layers that map photon signatures to update the values ​​of the embedding space map. The layers (multiple) may be or may include a network of one or more nodes, weights, and / or biases. For example, in some embodiments, the layers (multiple) may form a single-layer trace model, such as an embedding space model with only single-layer weights and / or biases. As another example, these layers may form a deep model with multiple layers. In some embodiments, the layers may be linear layers, such as layers that scale the input data to form a linear embedding space model without considering the activation function. In some implementations, a single-layer linear embedding space model can be advantageous for computational speed while still achieving the desired classification performance. The weights, biases, and / or other parameters of the embedding space model (e.g., a linear embedding space model) may be learned as part of training a low-photon-count classification model.

[0041] Additionally and / or alternatively, the low-photon count classification model may include a feature cross model configured to obtain multiple feature crosses of the embedding space map. As an example, the feature cross model may be configured to perform an outer product (e.g., with itself) of the embedding space map to produce multiple feature crosses, such as multiple quadratic feature crosses. For example, if the embedding space map is defined in an M-dimensional embedding space, the feature cross model may produce an M×M feature cross matrix. In some cases, symmetry in the feature cross matrix may allow the removal of approximately half of the entries from the M×M feature cross matrix. In some implementations, this is advantageous for reducing memory usage, increasing processing speed, reducing computational resources, and / or other benefits. For example, when the classification trace vector and / or embedding space map are progressively updated with each photon signature, it may only be necessary to recompile half of the M×M entries in the feature cross matrix. Furthermore, in some cases, sparsity implemented during training of the low-photon count classification model can further reduce the number of entries in the feature cross matrix that must be recomputed by the feature cross model. As an example, sparsification can remove some links in a feature cross model, so that each unit, photon detector, etc., contributes only to a subset of the embedding space dimension and / or feature cross.

[0042] Additionally and / or alternatively, the low-photon count classification model may include a feature cross-trace model. The feature cross-trace model can be configured to map multiple feature crosses (e.g., feature cross matrices) to multiple feature cross-trace vectors. For example, the feature cross-trace model may include one or more layers that map feature crosses to feature cross-trace vectors. The layers (multiple) may be or may include a network of one or more nodes, weights, and / or biases. For example, in some embodiments, the layers (multiple) may form a single-layer feature cross-trace model, such as a feature cross-trace model with only a single layer of weights and / or biases. As another example, these layers may form a deep model with multiple layers. In some embodiments, the layers may be linear layers, such as layers that scale the input data to form a linear feature cross-trace model without considering the activation function. In some implementations, a single-layer linear feature cross-trace model can be advantageous for computational speed while still achieving the desired classification performance. The weights, biases, and / or other parameters of the feature cross-trace model (e.g., a linear feature cross-trace model) may be learned as part of training the low-photon count classification model. In some implementations, the embedding space map may simply be the output of the embedding space model, and the accumulation of data can be performed downstream, for example, by setting the feature cross vectors (multiple) to zero initially and updating them based on the embedding space map (e.g., for each photon signature).

[0043] Additionally and / or alternatively, in some implementations, an aggregated trace vector can be determined from multiple feature-intersecting trace vectors and / or classification trace vectors (e.g., from a linear trace model). For example, multiple feature-intersecting trace vectors and / or trace vectors can be combined (e.g., summed) to produce an aggregated trace vector. The aggregated trace vector can be associated with a classification. For example, the maximum value in the dimension of the aggregated trace vector can correspond to a classification. Additionally and / or alternatively, in some implementations, the aggregated trace vector can be mapped to a probability vector, e.g., via a softmax function. For example, the instructions can further include mapping the aggregated trace vector to a probability vector, e.g., via a softmax function. The probability vector can include multiple probabilities associated with each of a plurality of candidate categories. As an example, a low-photon count classification model can be configured to estimate the probability that a visual object belongs to one of a plurality of candidate categories (e.g., Rogerette likelihood).

[0044] In some embodiments, any model described herein may include statistical networks (e.g., Bayesian networks, such as Naive Bayes networks), in addition to layers that substitute for weights, biases, etc. In some implementations, statistical networks can be learned through training.

[0045] In some embodiments, a low-photon count classification model may include an accumulator. The accumulator may be configured to accumulate classification trace vectors over multiple photon signatures after receiving photon signatures, and output a classification when the classification trace vectors have stabilized over the multiple photon signatures. For example, the accumulator may track the internal state of the low-photon count classification model (e.g., a normalized form of the vectors, such as a classification trace vector, an aggregated trace vector, etc.) such that predictions can be suppressed until the classification, trace vectors, etc., have stabilized. Including an accumulator can be beneficial because predictions can be provided after stabilization (e.g., as a representation of accuracy) and / or before waiting for enough photons to fully understand the visual object. As an example, the accumulator may track one or more differences (e.g., in each dimension) between the current value of the classification trace vector and previous values ​​of the classification trace vector. These differences can be tracked until they stabilize, and then a prediction can be output. In some embodiments, the accumulator may be omitted and / or predictions may be continuously available and / or provided after a certain number of photons, regardless of stability. Examples of stability include consistent differences (e.g., stable changes in each photon signature toward the same classification) and / or consistent trace vector values ​​(e.g., little or no change in each photon signature, such as changes less than a threshold).

[0046] In some implementations, at least one FIFO buffer can be coupled to the photon detection system. The FIFO buffer can be configured to store photon signatures before they are input to the low-photon count classification model. For example, the FIFO buffer can save photon signatures in the order they are generated (e.g., along with other information such as the time associated with the photon signature) and output the photon signatures when the low-photon count classification model is ready for computation. For example, some or all of the low-photon count classification model processing may become a bottleneck, preventing two photon signatures from being processed simultaneously. Therefore, to prevent the loss of simultaneously and / or nearly simultaneously incident photons, the FIFO buffer can save photon signatures until the model is ready to accept another photon signature. In some implementations, the FIFO buffer can be a fragmented FIFO buffer. For example, a fragmented FIFO buffer can split an entire data entity into fragments or partial pieces of the data entity. In some implementations, such as if the low-photon count classification model includes multiple cell trace vectors associated with various sets of one or more cells, a FIFO buffer can be provided for each set of one or more cells.

[0047] The computing system may store one or more instructions that, when implemented, cause one or more processors to perform operations for low-photon-count visual object recognition. For example, these operations may include a computer-implemented method for low-photon-count visual object recognition according to exemplary aspects of this disclosure.

[0048] The computer-implemented method may include obtaining photonic signatures from a photonic detection system (e.g., via a computing system comprising one or more computing devices). As an example, photonic signatures (e.g., including electrical signals such as electrical spikes) may be generated by photonic detectors and / or units of the photonic detection system. For example, a photonic detector may generate an electrical signature and / or associated data (e.g., the energy of the photon, etc.) from photons incident on it.

[0049] Additionally and / or alternatively, according to exemplary aspects of this disclosure, a computer-implemented method may include providing (e.g., via a computing system) a photon signature to a low-photon count classification model. For example, the photon signature may be transferred from a photon detection system, such as from a photon detector within the system, to the low-photon count classification model. In some implementations, the photon signature may be transferred to a FIFO buffer, such as a sliced ​​FIFO buffer, before being transferred to the low-photon count classification model. In some embodiments, the photon signature may be transmitted via one or more signal lines. For example, the signal line may be unique for each unit of the photon detection system and / or shared among some or all units.

[0050] Additionally and / or alternatively, the computer-implemented method may include determining (e.g., by a computational system) the classification of a visual object placed within the field of view (e.g., located within the field of view) of a photon detection system, at least in part based on photon signatures, using a low-photon-count classification model. As an example, the low-photon-count classification model may include a classification trace vector. The classification may be at least in part based on the classification vector. For example, the classification trace vector may store classification states on multiple photon signatures (e.g., the probability that a visual object belongs to each of multiple candidate categories). The computer-implemented method may include updating the classification trace vector at least in part based on the photon signatures. For example, the photon signature at a cell location may contribute to the likelihood or probability that a visual object in the field of view of the photon detection system belongs to one of the multiple candidate categories.

[0051] For example, in some implementations, a computer-implemented method may include providing photon signatures as input (e.g., via a computational system) to an embedding space model (e.g., a linear embedding space model). For instance, a low photon count classification model may include an embedding space model. The embedding space model may be configured to map photon signatures (or data derived from the cell locations of photon signatures) to form an embedding space mapping (e.g., updating values ​​used for updating) in an embedding space (e.g., an M-dimensional embedding space). This computer-implemented method may include receiving an embedding space mapping as output from a linear embedding space model in response to providing a photon signature to the linear embedding space model.

[0052] In some implementations, the computer-implemented method may include providing (e.g., via a computational system) an embedding space map as input to a feature cross model. For example, a low-photon count classification model may include a feature cross model. The feature cross model may be configured to obtain multiple feature crosses of the embedding space map. As an example, the feature cross model may be configured to perform an outer product of the embedding space map (e.g., with itself) to produce multiple feature crosses, such as multiple quadratic feature crosses. The computer-implemented method may include receiving multiple feature crosses as the output of the feature cross model in response to providing the embedding space map to the feature cross model.

[0053] In some implementations, the computer-implemented method may include providing (e.g., via a computational system) multiple feature intersections as input to a linear feature intersection trace model. The feature intersection trace model may be configured to map multiple feature intersections (e.g., feature intersection matrices) to multiple feature intersection trace vectors. The computer-implemented method may include receiving multiple feature intersection trace vectors as output of the linear feature intersection trace model in response to providing the multiple feature intersections to the linear feature intersection trace model.

[0054] In some implementations, computer-implemented methods may include determining (e.g., via a computational system) an aggregated trace vector as a combination of a classification trace vector and multiple feature-intersecting trace vectors. As an example, multiple feature-intersecting trace vectors may be added to the classification trace vector (e.g., from a trace model) to produce the aggregated trace vector. Classification may be based at least in part on the aggregated trace vector. For example, the aggregated trace vector may indicate the likelihood (e.g., Rogert likelihood) of a visual object belonging to each of multiple candidate categories.

[0055] Additionally and / or alternatively, a computer-implemented method may include providing a classification as the output of a low-photon count classification model. As an example, a classification can be provided once it has stabilized (e.g., a normalized form of a classification trace vector). As another example, a classification may be provided continuously (e.g., with each update) and / or upon request from a system, such as a system that performs additional actions based on the classification. Additional actions may include, for example, providing data related to the classification and / or visual objects to a user and / or an additional system, performing control actions based on the classification, such as waking up and / or disabling devices, sensors, etc., providing textual data depicting the visual objects based on the classification (e.g., optical character recognition, handwriting recognition, etc.), or any other suitable additional action.

[0056] According to an example aspect of this disclosure, a computing system may additionally and / or alternatively store one or more instructions that, when implemented, cause one or more processors to perform operations for training a low-photon-count visual object recognition model configured for low-photon-count visual object recognition. For example, according to an example aspect of this disclosure, these operations may include a computer-implemented method for training the low-photon-count visual object recognition model configured for low-photon-count visual object recognition. According to an example aspect of this disclosure, other suitable methods for training the low-photon-count visual object recognition model configured for low-photon-count visual object recognition may be employed.

[0057] The computer-implemented method may include obtaining (e.g., via a computing system comprising one or more computing devices) a training dataset comprising one or more images. For example, the training dataset may include images stored in computer-readable data and / or any suitable format, including RGB, CYMK, or any other suitable format such as PNG, JPEG, BMP, GIF, and / or SVG. Many suitable datasets exist with millions (if not billions) of labeled and / or tagged images.

[0058] Computer-implemented methods may include generating (e.g., via a computing system) time series of training examples from one or more images. For example, the time series of training examples may include time series of photon signatures derived from one or more images. Additionally and / or optionally, other data, such as photon intensity, energy, etc., may be derived from one or more images. For example, the time series of training examples may reflect how the photons that initially generated each image were received (e.g., at the camera that previously captured the image). This approach can be advantageous because, although, according to an exemplary aspect of this disclosure, the time series of photon signatures can be provided directly as training data, generating photon signatures from image data (e.g., existing image data) can allow for comprehensive training on many large existing image datasets.

[0059] The computer-implemented method may include providing time series of training examples from a computing system to a low-photon-count classification model. For example, the time series of training examples may be provided to the model in a manner that simulates how training examples are received from a photon detection system (e.g., simulates how photon signatures are received via signal lines, etc.) and / or in other ways that allow the model to accurately learn how to classify visual objects in a low-photon-count environment.

[0060] The computer-implemented method may include backpropagating (e.g., via a computational system) the loss from the time series of the training examples after providing a subset of time series of training examples to the low-photon counting classification model to train the model. As an example, the loss may be the difference between the expected classification (e.g., labels from images in the training dataset) and the classification output by the low-photon counting classification model in response to the time series of training examples. As an example, the loss may define a gradient that can be used to train the model via backpropagation (e.g., via gradient descent, stochastic gradient descent, etc.). For example, the parameters of a low-photon counting classification model (e.g., weights, biases, activations, etc.) can be adjusted based on this loss to provide a more accurate classification of the training data. As an example, training a low-photon counting visual object classification model may include training one or more sub-models that define the low-photon counting classification model, such as a trace model having one or more layers (e.g., linear layers) and / or configured to map photon signatures to classification trace vectors, an embedding space model having one or more layers (e.g., linear layers) and / or configured to map photon signatures to embedding space mappings, a feature cross model configured to obtain multiple feature crosses of the embedding space mapping, and / or a feature cross trace model having one or more layers (e.g., linear layers) and / or configured to map multiple feature crosses to multiple feature cross trace vectors. In some cases, this step may be performed multiple times for different corresponding subsets of the time series of the training examples.

[0061] In some implementations, sparsity can be performed as part of the training process. Sparsity can include removing some parameters (e.g., linkages, biases, etc.) that have a contribution less than a certain minimum (e.g., magnitude, value, etc.) such that these parameters do not significantly contribute to the accuracy of the low-photon count classification model. In some cases, sparsity can beneficially reduce the amount of computation and / or other computational resources that must be consumed in evaluating a low-photon count classification model, contributing to speed and / or power efficiency benefits. For example, in some implementations, sparsity can reduce the number of rows and / or columns that must be evaluated in the feature space matrix. As an example, regularization such as L1 regularization can be applied to sparsify low-photon count classification models.

[0062] In some implementations, some or all of the computations described herein (e.g., low-photon count classification models, some or all of the instructions, etc.) can be implemented using clockless logic. As an example, one or more processors may include clockless processors configured to execute clockless logic (e.g., clockless instructions). For instance, clockless logic can perform computations without relying on clock signals for synchronization, but can instead be executed in response to signal stimuli (e.g., photon signatures). Clockless logic can allow for slightly faster computations (e.g., by not having to wait for a clock edge) and / or improved power savings, since driving a clock can sometimes require considerable power.

[0063] Additionally and / or alternatively, in some implementations, the system and method can be implemented in a pipelined manner. For example, after receiving a single photon signature, matrix rows can be added to it to update the embedding space map. Feature crosses can then be updated from the updated embedding space map, and so on. In some cases, subsequent photons can only be processed after the computation of previous photons has been completed, which may introduce unnecessarily long detector dead time. However, in these cases, the computation can be broken down into pipeline stages. As an example, each pipeline stage can use a latch. The pipeline stage can be configured such that the update from the first photon to the feature cross can occur simultaneously with the update from subsequent photons arriving shortly thereafter to the embedding space map.

[0064] The systems and methods according to exemplary aspects of this disclosure can provide numerous technical effects and benefits, including improvements to computer-related technologies. As an example, the systems and methods according to exemplary aspects of this disclosure can provide classification of visual objects requiring fewer photons. For instance, by analyzing received photons relative to time and / or location (e.g., as opposed to still images), the systems and methods according to exemplary aspects of this disclosure can provide accurate and / or reliable classification for scenarios such as those with reduced photon counts, such as low-light conditions, scenarios where the photon detection system is only briefly exposed to the target, scenarios where it may be desirable to protect the privacy of the target, and / or other low-photon count scenarios.

[0065] As another example, the systems and methods according to exemplary aspects of this disclosure can improve the speed of computer-related techniques for visual object classification. For instance, once classification has stabilized, but before accumulating sufficient photons for the synthesized image, accurate classification can be achieved in a fraction of the time required by some conventional classification methods, including sub-millisecond classification.

[0066] As another example, systems and methods according to exemplary aspects of this disclosure can provide reduced power usage for classification tasks. For instance, photon detection systems and / or low-photon-count classification models according to exemplary aspects of this disclosure, particularly if they include a single-layer linear sub-model, can provide reduced power usage compared to some existing systems (e.g., cameras, full machine learning models, etc.). Furthermore, systems and methods according to exemplary aspects of this disclosure can provide power by “gating” more power-consuming and / or more complex (e.g., the most expensive to operate, especially power-intensive) systems, allowing the system to be activated in response to a specific classification from the low-photon-count classification system and method according to exemplary aspects of this disclosure.

[0067] Exemplary embodiments of the present disclosure will now be discussed in more detail with reference to the accompanying drawings. Repeated reference numerals in the various figures are intended to identify the same features in various implementations. As used herein, “about” in conjunction with numerical values ​​means within 20% of said value.

[0068] Figure 1A block diagram depicts an example low-photon-count visual object classification system 100 according to an example implementation of this disclosure. System 100 may include a photon detection system 110. Photon detection system 110 may include one or more units 112. Each of the one or more units 112 may include one or more photon detectors 114. For example, photon detectors 114 may be grouped into units 112. For example, photon detectors 114 may be grouped into units 112 by spatial proximity and / or any other suitable grouping. As an example, photon detectors 114 may be grouped into an array of units 112 (e.g., a square and / or rectangular array). As an example, photon detection system 110 may include a single-photon avalanche diode (SPAD) array. For example, photon detectors 114 may be or may include single-photon avalanche diodes. Additionally and / or alternatively, units 112 may form a SPAD array.

[0069] In some embodiments, each unit 112 may have an associated signal line 105. Grouping photon detectors into units 112 with shared signal lines 105 is beneficial for computational speed because the increased number of signal lines 105 can prevent bottlenecks or congestion of the signal lines 105 if multiple photons are incident on different photon detectors simultaneously or nearly simultaneously. Additionally and / or alternatively, in some embodiments, one or more units 112 may share the signal line 105 (e.g., through multiplexing).

[0070] Photon detector 114 can be configured to output a photon signature in response to photons incident on one or more photon detectors 114. As an example, the photon signature can include an electrical signature, such as a spike in an electrical signal (e.g., voltage, current, etc.). For example, photon detector 114 can be a SPAD that outputs an electrical spike in response to incident photons. Furthermore, the photon signature can be associated with a cell location within one or more cells 112. As an example, a cell location can be an array location within an array of one or more cells 112 (e.g., a SPAD array) and / or within the cells 112 themselves. For example, a cell location can be a serial identifier, (x,y) pairs, etc. For example, the photon signature can include any combination of an electrical signature, cell location, additional photon data, or other suitable data. In one example embodiment, the photon signature only includes (x,y) pairs. Subsequent processing (e.g., a computational system including a low photon count classification model 120) can therefore locate the photon signature (e.g., an electrical signature) based on the cell location, such that the low photon count classification model 120 can receive information indicating the cell location where the photon was received. Additionally and / or alternatively, a photonic signature may include any other suitable information from the photon.

[0071] A computing system may include one or more processors and one or more memory devices for storing computer-readable data. According to this disclosure, any suitable type, architecture, etc., of processor and / or memory device may be employed, such as CPU, GPU, microcontroller, microprocessor, etc., and / or RAM, hard disk storage, solid-state storage, magnetic tape, compactor, and / or any combination thereof.

[0072] The computing system can store (e.g., in a computer-readable storage device) a low-photon counting classification model 120. The low-photon counting classification model 120 can be configured to receive photon signatures (e.g., cell locations, such as (x,y) pairs) from the photon detection system 110. For example, the low-photon counting classification model 120 can receive a series (e.g., time series) of photon signatures generated at the photon detection system 110 and transmitted to the low-photon counting classification model 120. The low-photon counting classification model 120 can process each of the series of photon signatures sequentially (e.g., in the order of arrival). Thus, the low-photon counting classification model 120 can provide classifications for visual objects in the field of view of the photon detection system 110 from a set of candidate classifications (e.g., based on cumulative information from multiple photons and / or photon signatures). For example, this set of candidate classifications can be learned (e.g., by training the low-photon counting classification model 120) as a set of known and / or related object types.

[0073] Furthermore, in many cases, classification can be stable (and / or provided) after a small number of photons, such as before capturing a complete image of a visual object. For example, classification can be stable after fewer than about 10,000 photons, such as fewer than about 1,000 photons, or even fewer than about 100 photons, incident on the photon detection system 110. According to the example aspects of this disclosure, any suitable low-photon count classification model 120 can be used. For example, in some implementations, low-photon count classification can be or can include one or more logistic regression models of machine learning. It should be understood that while the low-photon count classification model 120 can produce classification based on a relatively small number of incident photons, aspects of this disclosure can be used with any number of photons provided, including, where appropriate, more than 10,000 photons.

[0074] The low-photon counting classification model 120 may include a classification trace vector 122. For example, the classification trace vector 122 may be or may include a probability vector, such as a Rogers likelihood probability vector. The classification trace vector 122 may be an N-dimensional vector, such as a 10-dimensional vector. For example, each dimension of the classification trace vector 122 may correspond to one of a plurality of candidate categories.

[0075] The classification can be based at least in part on the classification trace vector 122. For example, the classification trace vector 122 can serve as the internal state of the low-photon count classification model 120, which models the likelihood that a visual object in the field of view of the photon detection system 110 belongs to each of a plurality of candidate categories. As an example, the classification trace vector 122 may include a photon signature indicating the logit likelihood that a visual object belongs to one of the plurality of candidate categories. For example, the candidate categories can be established by providing training data to the low-photon count classification model 120 (e.g., machine learning).

[0076] In some implementations, each component of the classification trace vector 122 may initially be zero, and / or the classification trace vector 122 may be updated with each photon signature, such that each photon signature contributes to classification. For example, this could allow classification to be updated incrementally with each incident photon, making classification available and / or stable whenever needed. Additionally and / or alternatively, the classification trace vector 122 may be updated after a certain number of photon signatures.

[0077] In some embodiments, the low photon count classification model 120 may include an accumulator 124. The accumulator 124 may be configured to accumulate a classification trace vector 122 over multiple photon signatures after receiving a photon signature, and output a classification when the classification trace vector 122 (e.g., its normalized form) has stabilized over the multiple photon signatures. For example, the accumulator 124 may track the internal state of the low photon count classification model 120 (e.g., classification trace vector 122, aggregation trace vector 428...). Figure 4 (e.g.,) allows predictions to be suppressed until stability is achieved. Including accumulator 124 may be beneficial because predictions can be provided after stability is achieved for classification and / or the classification trace vector 122 (e.g., as a representation of accuracy) and / or before waiting for enough photons to fully understand the visual object. As an example, accumulator 124 can track one or more differences (e.g., in each dimension) between the current value of classification trace vector 122 and previous values ​​of classification trace vector 122. These differences can be tracked until they stabilize, and then a prediction can be output. In some embodiments, accumulator 124 can be omitted, and / or predictions can be continuously available, and / or predictions can be provided after a certain number of photons, regardless of stability. Examples of stability include consistent differences (e.g., stable changes in each photon signature toward the same classification) and / or consistent trace vector 122 values ​​(e.g., minimal changes in each photon signature, such as changes less than a threshold).

[0078] The low-photon-count visual object classification system 100 may additionally include a trace model (not shown). The trace model maps photon signatures to classification trace vectors 122. For example, the trace model may include one or more layers that map photon signatures to classification trace vectors 122. The layers (multiple) may be or may include a network of one or more nodes, weights, and / or biases. For example, in some embodiments, the layers (multiple) may form a single-layer trace model, such as a trace model with only a single layer of weights and / or biases. As another example, these layers may form a deep model with multiple layers. In some embodiments, the layers may be linear layers, such as layers that scale the input data to form a linear trace model without considering an activation function. In some implementations, a single-layer linear trace model may be advantageous for computational speed while still achieving the desired classification performance. The weights, biases, and / or other parameters of the trace model (e.g., a linear trace model) may be learned as part of training the low-photon-count classification model 120.

[0079] Figure 2 A block diagram depicts an example low-photon counting visual object classification system 200 according to an example implementation of this disclosure. The low-photon counting classification system 200 may include... Figure 1 Components of the low-photon-count visual object classification system 100, such as, for example, photon detection system 110 and / or low-photon-count classification model 120.

[0080] The low-photon-count visual object classification system 200 may additionally include a FIFO buffer 202. The FIFO buffer 202 may be coupled to a photon detection system. The FIFO buffer 202 may be configured to store photon signatures before they are input to the low-photon-count classification model 120. For example, the FIFO buffer 202 may save photon signatures in the order they are generated (e.g., along with other information such as the time associated with the photon signature) and output the photon signatures when the low-photon-count classification model 120 is ready for computation. For example, some or all of the low-photon-count classification model processing may become a bottleneck, preventing two photon signatures from being processed simultaneously. Therefore, to prevent the loss of photons that are simultaneously and / or nearly simultaneously incident, the FIFO buffer 202 may save photon signatures until the low-photon-count classification model 120 is ready to accept another photon signature. In some implementations, the FIFO buffer 202 may be a fragmented FIFO buffer. For example, a fragmented FIFO buffer 202 may split an entire data entity into fragments or partial segments of the data entity.

[0081] Figure 3 A block diagram depicts an example low-photon counting visual object classification system 300 according to an example implementation of this disclosure. The low-photon counting classification system 300 may include... Figure 1 Components of the low photon count visual object classification system 100, such as, for example, photon detection system 110.

[0082] The low-photon-count visual object classification system 300 (e.g., low-photon-count classification model 320) may additionally include unit trace vectors 322 associated with some or all of the units 112. For example, each unit trace vector 322 may be associated with a set of one or more units 112. As an example, a unit trace vector 322 may correspond to a set of one or more units 112. For example, in some implementations, each unit trace vector 322 may be fed from a separate signal line 105. The unit trace vectors 322 may individually measure the contribution of a unit 112 or a group of units 112 to the classification. In some embodiments, each unit trace vector 322 may have a separate trace model (e.g., with a unique set of weights and / or biases) and / or a shared trace model (e.g., having the same composition for some or all of the unit trace vectors 322).

[0083] The unit trace vector 322 can be combined, for example, by summing (e.g., weighted summing and / or unweighted summing) to produce a classification trace vector 122 indicating classification, which, for illustrative purposes, is referred to as the "global trace vector 122". For example, the global trace vector 122 can be a combination of one or more unit trace vectors 322, such as a linear combination. Including the unit trace vector 322 facilitates parallel processing. For example, for each photon, only the unit trace vector 322 can be updated without updating the global trace vector 122, and / or the global trace vector 122 can be updated only after the photon detection system 110 and / or the low photon count classification model 320 has received multiple photons. This allows for parallel processing of two photons simultaneously incident in different units 112.

[0084] Figure 4 A block diagram depicts an example low-photon counting visual object classification system 400 according to an example implementation of this disclosure. The low-photon counting classification system 400 may include... Figure 1 Components of the low photon count visual object classification system 100, such as, for example, photon detection system 110.

[0085] Additionally and / or alternatively, in some implementations, the low photon count classification model 120 may include an embedding space model. The embedding space model maps photon signatures to an embedding space mapping 422 within an embedding space 422. The embedding space model includes the embedding space mapping 422 within the embedding space. All components of the embedding space mapping 422 may initially be zero. For example, the embedding space may define a dimensional space, such as an M-dimensional space (e.g., a 30-dimensional space). For example, the embedding space model may include one or more layers that map photon signatures to the embedding space mapping 422 (e.g., update their values). The layers (multiple) may be or may include a network of one or more nodes, weights, and / or biases. For example, in some embodiments, the layers (multiple) may form a single-layer trace model, such as an embedding space model with only a single layer of weights and / or biases. As another example, these layers may form a deep model with multiple layers. In some embodiments, the layers may be linear layers, such as layers that scale the input data to form a linear embedding space model without considering the activation function. In some implementations, a single-layer linear embedding space model may be advantageous for computational speed while still achieving the desired classification performance. The weights, biases, and / or other parameters of the embedding space model (e.g., a linear embedding space model) can be learned as part of training the low-photon counting classification model 420. For example, in some implementations, the embedding space model can generate an updated value for the embedding space map 422 when a photon signature arrives, and then update the embedding space map 422 by adding the updated value to the corresponding component of the embedding space map 422.

[0086] Additionally and / or alternatively, the low-photon count classification model 120 may include a feature cross model configured to obtain multiple feature crosses 424 of an embedding space map 422. As an example, the feature cross model may be configured to perform an outer product (e.g., with itself) of the embedding space map 422 to produce multiple feature crosses 424, such as multiple quadratic feature crosses 424. For example, if the embedding space map 422 is defined in an M-dimensional embedding space 422, the feature cross model may produce an M×M matrix of feature crosses 424. In some cases, symmetries in the feature cross matrix 424 may allow the removal of approximately half of the entries from the M×M feature cross matrix 424. In some implementations, this is advantageous for reducing memory usage, increasing processing speed, reducing computational resources, and / or other benefits. For example, when the classification trace vector 122 and / or the embedding space map 422 are progressively updated with each photon signature, only half of the M×M entries in the feature cross matrix 424 must be recomputed. Furthermore, in some cases, sparsification implemented during the training of the low-photon count classification model 420 can further reduce the number of entries in the feature cross matrix 424 that must be recomputed by the feature cross model. As an example, sparsification can remove some links in the feature cross model, such that each unit 112, photon detector 114, etc., contributes only to a subset of the embedding space dimension 422 and / or the feature cross 424.

[0087] Additionally and / or alternatively, the low-photon count classification model 120 may include a feature cross-trace model. The feature cross-trace model may be configured to map multiple feature crosses 424 (e.g., feature cross matrices 424) to multiple feature cross-trace vectors 426. For example, the feature cross-trace model may include one or more layers that map feature crosses 424 to feature cross-trace vectors 426. The layers (multiple) may be or may include a network of one or more nodes, weights, and / or biases. For example, in some embodiments, the layers (multiple) may form a single-layer feature cross-trace model, such as a feature cross-trace model with only single-layer weights and / or biases. As another example, these layers may form a deep model with multiple layers. In some embodiments, the layers may be linear layers, such as layers that scale the input data to form a linear feature cross-trace model without considering the activation function. In some implementations, a single-layer linear feature cross-trace model may be advantageous for computational speed while still achieving the desired classification performance. The weights, biases, and / or other parameters of the feature cross-trace model (e.g., a linear feature cross-trace model) may be learned as part of training the low-photon count classification model 120.

[0088] Additionally and / or alternatively, in some implementations, the aggregate trace vector 428 can be determined from multiple feature-intersecting trace vectors 426 and / or classification trace vectors 122 (e.g., from a linear trace model). For example, multiple feature-intersecting trace vectors 426 and / or classification trace vectors 122 can be combined (e.g., summed) to produce the aggregate trace vector 428. The aggregate trace vector 428 can be associated with a classification. For example, the maximum value in the dimension of the aggregate trace vector 428 can correspond to a classification. As an example, the classification could be the category corresponding to the largest component of the aggregate trace vector 428. Additionally and / or alternatively, in some implementations, the aggregate trace vector 428 can be mapped to a probability vector, such as via a softmax function. For example, the probability vector can include multiple probabilities associated with each of a plurality of candidate categories. As an example, a low-photon count classification model 420 can be configured to estimate the probability that a visual object belongs to one of the plurality of candidate categories (e.g., Rogert likelihood).

[0089] In some embodiments, any model used to generate trace vector 122, embedding space mapping 422, feature cross 424, and / or feature cross trace vector 426, etc., may include a statistical network (e.g., a Bayesian network, such as a Naive Bayes network), in addition to and / or layers that substitute weights, biases, etc. In some implementations, the statistical network can be learned through training.

[0090] Example devices and systems

[0091] Figures 5A-5C A block diagram of an example computing system according to an example implementation of this disclosure is depicted. For example, Figure 5A A block diagram of an example computing system 500 performing low-photon-count object classification according to an example embodiment of the present disclosure is depicted. System 500 includes a user computing device 502, a server computing system 530, and a training computing system 550 communicatively coupled via a network 580.

[0092] User computing device 502 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0093] User computing device 502 includes one or more processors 512 and memory 514. The one or more processors 512 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 514 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 514 can store data 516 and instructions 518 executed by the processor 512 to cause the user computing device 502 to perform operations.

[0094] In some implementations, the user computing device 502 may store or include one or more low-photon counting classification models 520. For example, the low-photon counting classification model 520 may be, or may otherwise include, various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including nonlinear and / or linear models. Neural networks may include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Reference Figure 1-4 520 is an exemplary low-photon counting classification model discussed in section 6.

[0095] In some implementations, one or more low-photon counting classification models 520 may be received from a server computing system 530 via a network 580, stored in a user computing device memory 514, and then used or otherwise implemented by one or more processors 512. In some implementations, the user computing device 502 may implement multiple parallel instances of a single low-photon counting classification model 520 (e.g., performing parallel low-photon counting classification across multiple instances of the low-photon counting classification model 520).

[0096] More specifically, the low photon counting classification model 520 can be configured to analyze photon counts from a photon detection system (e.g., Figure 1-4 The low-photon counting classification model 520 receives photon signatures (e.g., cell locations, such as (x, y) pairs). For example, when a series of photon signatures (e.g., time series) are generated at the photon detection system and transmitted to the low-photon counting classification model 520, the low-photon counting classification model 520 can receive them. The low-photon counting classification model 520 can process each of the series of photon signatures sequentially (e.g., in the order of arrival). Therefore, the low-photon counting classification model 520 can provide classifications for visual objects in the field of view of the photon detection system from a set of candidate classifications. For example, this set of candidate classifications can be learned (e.g., by training the low-photon counting classification model 520) as a set of known and / or related object types.

[0097] Additionally or alternatively, one or more low-photon count classification models 540 may be included in, or otherwise stored and implemented by, a server computing system 530 that communicates with the user computing device 502 according to a client-server relationship. For example, the low-photon count classification model 540 may be implemented by the server computing system 540 as part of a network service (e.g., a low-photon count classification service). Thus, one or more models 520 may be stored and implemented at the user computing device 502, and / or one or more models 540 may be stored and implemented at the server computing system 530.

[0098] User computing device 502 may also include one or more user input components 522 for receiving user input. For example, user input component 522 may be a touch-sensitive component (e.g., a touch-sensitive display or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). Touch-sensitive components can be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other components that a user can use to provide user input.

[0099] Server computing system 530 includes one or more processors 532 and memory 534. The one or more processors 532 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 534 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 534 can store data 536 and instructions 538 executed by the processor 532 to cause the server computing system 530 to perform operations.

[0100] In some implementations, the server computing system 530 includes one or more server computing devices, or is implemented by one or more server computing devices. When the server computing system 530 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0101] As described above, the server computing system 530 may store or otherwise include one or more machine learning low-photon count classification models 540. For example, model 540 may be, or may otherwise include, various machine learning models. Example machine learning models include neural networks or other multi-layer nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. (Reference) Figure 1-4 Example model 540 was discussed in section 6.

[0102] User computing device 502 and / or server computing system 530 can train models 520 and / or 540 via interaction with training computing system 550, which is communicatively coupled through network 580. Training computing system 550 may be separate from server computing system 530, or it may be part of server computing system 530.

[0103] The training computing system 550 includes one or more processors 552 and memory 554. The one or more processors 552 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. The memory 554 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 554 can store data 556 and instructions 558 executed by the processor 552 to cause the training computing system 550 to perform operations. In some implementations, the training computing system 550 includes one or more server computing devices, or is otherwise implemented by one or more server computing devices.

[0104] The training computing system 550 may include a model trainer 560 that uses various training or learning techniques, such as, for example, error backpropagation, to train machine learning models 520 and / or 540 stored at user computing device 502 and / or server computing system 530. For example, a loss function can be backpropagated through the models(s)(s)(s)(s)(s)(s)(s)(s))(s)(s)(s)(s)) to update one or more parameters of the models(s) ...)(s)(s)(s)(s)))(s)(s)(s)(s)))(s)(s)(s)(s))

[0105] In some implementations, backpropagation of the execution error may include truncated backpropagation over time. The model trainer 560 can perform various generalization techniques (e.g., weight decay, exit, etc.) to improve the generalization ability of the trained model.

[0106] Specifically, model trainer 560 can train low photon count classification models 520 and / or 540 based on a set of training data 562. Training data 562 may include, for example, time series of example photon signatures. In some embodiments, example photon signatures may be generated as is, such as from an example photon detection system. In some embodiments, according to exemplary aspects of this disclosure, example photon signatures can be generated from example image data (e.g., a training dataset comprising one or more images, such as one or more labeled images). For example, in some embodiments, example photon signatures can be generated to reproduce how a photon detection system perceives image data in real life. This can be advantageous because it allows training data to be generated from image data, especially from large existing image data training datasets.

[0107] In some implementations, if the user has provided consent, training examples can be provided by the user computing device 502. Therefore, in such an implementation, the training computing system 550 can train the model 520 provided to the user computing device 502 based on user-specific data received from the user computing device 502. In some cases, this process can be referred to as model personalization.

[0108] Model trainer 560 includes computer logic for providing the required functionality. Model trainer 560 can be implemented using hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, model trainer 560 includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, model trainer 560 includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium, such as RAM, a hard disk, or optical or magnetic media.

[0109] Network 580 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL) can be used to carry communication on Network 580 via any type of wired and / or wireless connection.

[0110] The machine learning models described in this specification can be used in a variety of tasks, applications, and / or use cases.

[0111] In some implementations, the input to the machine learning model(s) of this disclosure can be image data. The machine learning model(s) can process the image data to generate output. As an example, the machine learning model(s) can process image data to generate image recognition output (e.g., image data identification, latent embedding of image data, encoded representation of image data, hashing of image data, etc.). As another example, the machine learning model(s) can process image data to generate image segmentation output. As another example, the machine learning model(s) can process image data to generate image classification output. As another example, the machine learning model(s) can process image data to generate image data modification output (e.g., image data alteration, etc.). As another example, the machine learning model(s) can process image data to generate encoded image data output (e.g., encoded and / or compressed representation of image data, etc.). As another example, the machine learning model(s) can process image data to generate enhanced image data output. As another example, the machine learning model(s) can process image data to generate predictive output.

[0112] In some implementations, the input to the machine learning model(s) of this disclosure may be one or more photon signatures from text images or natural language data. The machine learning model(s) may process the photon signatures to generate outputs. For example, the machine learning model(s) may process natural language data to generate language-encoded outputs. As another example, the machine learning model(s) may process text or natural language data to generate latent text embedding outputs. As another example, the machine learning model(s) may process photon signatures to generate translated outputs. As another example, the machine learning model(s) may process photon signatures to generate classification outputs. As another example, the machine learning model(s) may process photon signatures to generate text segmentation outputs. As another example, the machine learning model(s) may process photon signatures to generate semantic intent outputs. As another example, the machine learning model(s) may process photon signatures to generate enhanced text or natural language outputs (e.g., text or natural language data of higher quality than the input text or natural language). As another example, the machine learning model(s) may process photon signatures to generate predictive outputs.

[0113] Figure 5AThe illustration shows an example computing system that can be used to implement this disclosure. Other computing systems may also be used. For example, in some implementations, the user computing device 502 may include a model trainer 560 and a training dataset 562. In such an implementation, the model 520 can be trained and used locally on the user computing device 502. In some such implementations, the user computing device 502 may implement the model trainer 560 to personalize the model 520 based on user-specific data.

[0114] Figure 5B A block diagram depicts an example computing device 50 implemented according to an example embodiment of the present disclosure. The computing device 50 may be a user computing device or a server computing device.

[0115] Computing device 50 includes multiple applications (e.g., applications 5 to N). Each application contains its own machine learning library and (multiple) machine learning models. For example, each application may include a machine learning model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0116] like Figure 5B As illustrated, each application can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In other implementations, the API used by each application is application-specific.

[0117] Figure 5C A block diagram depicts an example computing device 50 implemented according to an example embodiment of the present disclosure. The computing device 50 may be a user computing device or a server computing device.

[0118] Computing device 50 includes multiple applications (e.g., applications 5 to N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application may communicate with the central intelligence layer (and the multiple models stored therein) using an API (e.g., a common API across all applications).

[0119] The central intelligence layer includes many machine learning models. For example, such as... Figure 5CAs illustrated, a corresponding machine learning model (e.g., a model) can be provided for each application and managed by a central intelligent layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligent layer can provide a single model (e.g., a single model) for all applications. In some implementations, the central intelligent layer is included in the operating system of computing device 50, or otherwise implemented by the operating system of computing device 50.

[0120] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data warehouse for computing device 50. For example... Figure 5C As illustrated, the central device data layer can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0121] Example model layout

[0122] Figure 6 A block diagram depicts an example low-photon counting classification model according to an example implementation of this disclosure. In some implementations, the low-photon counting classification model 600 is trained to receive a set of input data 604 describing a time series of photon signatures, and as a result of receiving the input data 604, provides output data 606 representing low-photon counting classification data (e.g., trace vectors and / or determined classifications). Therefore, in some implementations, the low-photon counting classification model 600 may include a machine learning low-photon counting classification model 602 operable to perform low-photon counting classification as described herein. For example, the machine learning low-photon counting classification model may include a machine learning logistic regression model.

[0123] Example Method

[0124] Figure 7 A flowchart is depicted of an example computer implementation of a method 700 for classifying low-photon-count visual objects according to an exemplary implementation of this disclosure. For example, a computing system may store one or more instructions that, when implemented, cause one or more processors to perform operations for low-photon-count visual object recognition according to the example computer implementation of method 700. Although for illustrative and discussion purposes... Figure 7 The steps are depicted in a specific order, but the method of this disclosure is not limited to the specific order or arrangement shown in the illustrations. Without departing from the scope of this disclosure, the various steps of method 700 may be omitted, rearranged, combined, and / or adapted in various ways.

[0125] In some implementations, multiple instances of method 700 may be executed simultaneously. For example, a photon detection system executes step 702 while a computing system (e.g., including a low photon count classification model according to an example aspect of this disclosure) is executing one or more of steps 704-708 for previously received photons.

[0126] The computer-implemented method 700 may include, at 702, obtaining a photon signature from a photon detection system (e.g., via a computing system including one or more computing devices). As an example, the photon signature (e.g., including an electrical signal, such as an electrical spike) may be generated by a photon detector and / or unit of the photon detection system. For example, the photon detector may generate an electrical signature and / or associated data (e.g., the energy of the photon, etc.) from photons incident on it.

[0127] Additionally and / or alternatively, according to an example aspect of this disclosure, a computer-implemented method 700 may include, at 704, providing a photon signature (e.g., via a computing system) to a low-photon count classification model. For example, the photon signature may be transferred from a photon detection system, such as from a photon detector in the photon detection system, to the low-photon count classification model. In some implementations, the photon signature may be transferred to a FIFO buffer, such as a sliced ​​FIFO buffer, before being transferred to the low-photon count classification model. In some embodiments, the photon signature may be transmitted via one or more signal lines. For example, the signal line may be unique for each unit of the photon detection system and / or shared among some or all units.

[0128] Additionally and / or alternatively, the computer-implemented method 700 may, at 706, determine (e.g., by a computational system) the classification of a visual object placed in the field of view of a photon detection system, at least in part based on photon signatures, using a low-photon count classification model. As an example, the low-photon count classification model may include a classification trace vector. The classification may be based at least in part on the classification trace vector. Updates to the classification trace vector may be generated for each photon signature. For example, the classification trace vector may store a classification state accumulated across multiple photon signatures (e.g., the probability that a visual object belongs to each of a plurality of candidate categories). The computer-implemented method may include updating the classification trace vector at least in part based on photon signatures. For example, a photon signature at a cell location may contribute to the likelihood or probability that a visual object in the field of view of the photon detection system belongs to one of a plurality of candidate categories.

[0129] For example, in some implementations, a computer-implemented method may include providing (e.g., via a computational system) a photon signature to an embedding space model (e.g., a linear embedding space model) as input. For instance, a low photon count classification model may include an embedding space model. The embedding space model may be configured to map the photon signature and / or classification trace vector to an embedding space mapping (e.g., updating values) in an embedding space (e.g., an M-dimensional embedding space). The computer-implemented method may include receiving an embedding space mapping (e.g., updating it) as output from the linear embedding space model in response to providing the photon signature to the linear embedding space model.

[0130] In some implementations, the computer-implemented method may include providing (e.g., via a computational system) an embedding space map as input to a feature cross model. For example, a low-photon count classification model may include a feature cross model. The feature cross model may be configured to obtain multiple feature crosses of the embedding space map. As an example, the feature cross model may be configured to perform an outer product of the embedding space map (e.g., with itself) to produce multiple feature crosses, such as multiple quadratic feature crosses. The computer-implemented method may include receiving multiple feature crosses as the output of the feature cross model in response to providing the embedding space map to the feature cross model.

[0131] In some implementations, the computer-implemented method may include providing (e.g., via a computational system) multiple feature intersections as input to a linear feature intersection trace model. The feature intersection trace model may be configured to map multiple feature intersections (e.g., feature intersection matrices) to multiple feature intersection trace vectors. The computer-implemented method may include receiving multiple feature intersection trace vectors as output of the linear feature intersection trace model in response to providing the multiple feature intersections to the linear feature intersection trace model.

[0132] In some implementations, computer-implemented methods may include determining (e.g., via a computational system) an aggregated trace vector as a combination of a classification trace vector and multiple feature-intersecting trace vectors. As an example, multiple feature-intersecting trace vectors may be added to the classification trace vector (e.g., from a trace model) to produce the aggregated trace vector. Classification may be based at least in part on the aggregated trace vector. For example, the aggregated trace vector may indicate the likelihood (e.g., Rogert likelihood) of a visual object belonging to each of multiple candidate categories.

[0133] Additionally and / or alternatively, the computer-implemented method 700 may include, at 708, providing the classification as the output of a low-photon count classification model. As an example, the classification can be provided once the classification (e.g., a classification trace vector) has stabilized. As another example, the classification may be provided continuously (e.g., with each update) and / or upon request from a system, such as a system that performs additional actions based on the classification. Additional actions may include, for example, providing data related to the classification and / or visual object to a user and / or an additional system, performing control actions based on the classification, such as waking up and / or disabling devices, sensors, etc., providing textual data depicting the visual object (e.g., optical character recognition, handwriting recognition, etc.) based on the classification, or any other suitable additional action.

[0134] Figure 8 A flowchart depicts an example computer implementation method for training a low-photon count classification model configured for low-photon count visual object recognition, according to an exemplary implementation of the present disclosure. For example, a computing system may store one or more instructions that, when implemented, cause one or more processors to perform operations according to an exemplary computer implementation method 800 to train a low-photon count classification model configured for low-photon count visual object recognition, according to an exemplary implementation of the present disclosure. Although for illustrative and discussion purposes, Figure 8 The steps are depicted in a specific order, but the method of this disclosure is not limited to the specific illustrated order or arrangement. Without departing from the scope of this disclosure, the various steps of method 800 may be omitted, rearranged, combined, and / or adapted in various ways.

[0135] The computer-implemented method 800 may include, at 802, obtaining (e.g., via a computing system including one or more computing devices) a training dataset comprising one or more images. For example, the training dataset may include images stored in computer-readable data and / or any suitable format, including RGB format, CYMK format, or any other suitable format such as, for example, PNG, JPEG, BMP, GIF, and / or SVG. Many suitable datasets exist with millions (if not billions) of labeled and / or tagged images.

[0136] The computer-implemented method 800 may include, at 804, generating (e.g., via a computing system) time series of training examples from one or more images. For example, the time series of training examples may include time series of photon signatures derived from one or more images. Additionally and / or optionally, other data, such as photon intensity, energy, etc., may be derived from one or more images. For example, the time series of training examples may reflect how the photons initially generated for each image were received (e.g., at the camera where the image was previously captured). This approach may be advantageous because, although, according to an exemplary aspect of this disclosure, the time series of photon signatures can be provided directly as training data, generating photon signatures from image data (e.g., existing image data) can allow for comprehensive training on many large existing image datasets.

[0137] The computer-implemented method 800 may include, in 806, providing a time series of training examples to the low-photon-count classification model by the computing system. For example, the time series of training examples may be provided to the model in a manner that simulates how training examples are received from a photon detection system (e.g., simulating how photon signatures are received via a signal line, etc.), and / or in other ways that allow the model to accurately learn how to classify visual objects in a low-photon-count environment.

[0138] The computer-implemented method 800 may include, at 808, backpropagating (e.g., via a computational system) the loss from the time series of the training examples to train the low-photon counting classification model after providing a subset of the time series of training examples to the low-photon counting classification model. As an example, the loss may be the difference between the expected classification (e.g., labels from images in the training dataset) and the classification output by the low-photon counting classification model in response to the time series of training examples. As an example, the loss may define a gradient that can be used to train the model via backpropagation (e.g., via gradient descent, stochastic gradient descent, etc.). For example, the parameters of the low-photon counting classification model (e.g., weights, biases, activations, etc.) may be adjusted based on the loss to provide a more accurate classification for the training data. As an example, training a low-photon-count visual object classification model may include training one or more sub-models that define the low-photon-count classification model, such as, for example, a trace model having one or more layers (e.g., linear layers) and / or configured to map photon signatures to classification trace vectors, an embedding space model having one or more layers (e.g., linear layers) and / or configured to map photon signatures to embedding space mappings, a feature cross model configured to obtain multiple feature crosses of the embedding space mapping, and / or a feature cross trace model having one or more layers (e.g., linear layers) and / or configured to map multiple feature crosses to multiple feature cross trace vectors. In some cases, step 808 may be performed multiple times, such as each time using a different subset of the time series of training examples.

[0139] In some implementations, sparsity can be performed as part of the training process (e.g., after step 808 and / or at any point in method 800, such as between multiple executions of step 808 and / or after all instances of step 808). Sparsity may include removing some parameters (e.g., linkages, biases) that have a smaller contribution (e.g., magnitude, value, etc.) than a certain minimum contribution, such that these parameters do not significantly contribute to the accuracy of the low-photon count classification model. In some cases, sparsity can beneficially reduce the amount of computation that must be performed and / or other computational resources that must be consumed in evaluating a low-photon count classification model, which contributes to speed and / or power usage benefits. For example, in some implementations, sparsity can reduce the number of rows and / or columns that must be evaluated in the feature space matrix. As an example, regularization such as L1 regularization can be applied to sparsify a low-photon count classification model.

[0140] Additional disclosure

[0141] The technologies discussed here involve servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and received from these systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functions between and within components. For example, the processes discussed here can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0142] While this subject matter has been described in detail with respect to various specific example embodiments, each example is provided by way of explanation and not as a limitation of this disclosure. Those skilled in the art will readily recognize changes, modifications, and equivalents to these embodiments upon understanding the foregoing. Therefore, this disclosure does not exclude such modifications, modifications, and / or additions to the subject matter, which will be apparent to those skilled in the art. For example, features illustrated or described as part of one embodiment may be used with another embodiment to produce yet another embodiment. Therefore, this disclosure is intended to cover such changes, modifications, and equivalents.

Claims

1. A computing system configured for low-photon-count visual object classification, the computing system comprising: a photon detection system comprising one or more units, each of the one or more units comprising one or more photon detectors, each of the one or more photon detectors configured to output a photon signature in response to a photon incident on the one or more photon detectors; one or more processors; and one or more memory devices storing computer-readable data, the data comprising: a low-photon-count classification model; and one or more instructions, when implemented, cause the one or more processors to perform operations for low-photon-count visual object recognition, the operations comprising: obtaining a photon signature from the photon detection system, wherein the photon signature comprises a signal indicative of a photon having been received by one or more photon detectors; providing the photon signature to the low-photon-count classification model; determining, by the low-photon-count classification model, a classification of a visual object placed in a field of view of the photon detection system based at least in part on the photon signature; and providing the classification as an output of the low-photon-count classification model, wherein the low-photon-count classification model comprises a classification trace vector, the classification trace vector modeling a likelihood of the visual object belonging to each of a plurality of classes, and wherein determining the classification comprises updating the classification trace vector based at least in part on the photon signature, and wherein the low-photon-count classification model comprises an accumulator, wherein the accumulator is configured to accumulate the classification trace vector over a plurality of photon signatures after receiving the photon signature, and to output the classification when the classification trace vector has stabilized over the plurality of photon signatures. the classification trace vector comprises a logit likelihood of the visual object belonging to one of a plurality of candidate classes.

2. The computing system of claim 1, wherein, 3. The computing system of claim 1 or 2, wherein the low-photon-count classification model comprises one or more unit trace vectors related to the one or more units, wherein the classification trace vector comprises a combination of the one or more unit trace vectors. the low-photon-count classification model comprises a linear trace model comprising one or more linear layers, the linear trace model configured to map the photon signature to the classification trace vector, wherein each linear layer outputs a weighted sum of the layer’s input, optionally with a respective bias term added in the weighted sum.

4. The computing system of claim 1 or 2, wherein, the low-photon-count classification model further comprises:

5. The computing system of claim 1 or 2, wherein, a linear embedding space model comprising one or more linear layers, the linear embedding space model configured to map the photon signature to an embedding space mapping; a feature cross model configured to obtain a plurality of feature crosses of the embedding space mapping; and a linear feature cross trace model comprising one or more linear layers, the linear feature cross trace model configured to map the plurality of feature crosses to a plurality of feature cross trace vectors; wherein the instructions further comprise: ​ providing the photon signature to the linear embedding space model as input, and receiving, as output from the linear embedding space model, the embedding space map in response to providing the photon signature to the linear embedding space model; providing the embedding space map to the feature cross model as input, and receiving, as output from the feature cross model, the plurality of feature crosses in response to providing the embedding space map to the feature cross model; providing the plurality of feature crosses to the linear feature cross trace model as input, and receiving, as output from the linear feature cross trace model, the plurality of feature cross trace vectors in response to providing the plurality of feature crosses to the linear feature cross trace model; and determining an aggregate trace vector as a combination of the classification trace vector and the plurality of feature cross trace vectors, wherein the classification is based at least in part on the aggregate trace vector.

6. The computing system of claim 5, wherein the instructions further comprise mapping the aggregate trace vector to a probability vector by a softmax function.

7. The computing system of any one of claims 1, 2, and 6, wherein, The low-photon-count classification model comprises a machine-learned logistic regression model.

8. The computing system of any one of claims 1, 2, and 6, wherein the photon signature comprises an electrical signature and a cell location within the one or more cells.

9. The computing system of any one of claims 1, 2, and 6, wherein, The photon detection system comprises an array of single-photon avalanche diodes, and wherein the one or more photon detectors comprise one or more single-photon avalanche diodes.

10. The computing system of any one of claims 1, 2, and 6, wherein, The photon detection system comprises a rectangular array of the one or more cells.

11. The computing system of any one of claims 1, 2, and 6, further comprising a FIFO buffer coupled to the photonic detection system, wherein, The FIFO buffer is configured to store the photon signature prior to the photon signature being input to the low-photon-count classification model.

12. The computing system of claim 11, wherein, The FIFO buffer comprises a sliced FIFO buffer.

13. The computing system of any one of claims 1, 2, 6, and 12, wherein the one or more processors comprise clockless logic, and wherein the one or more instructions are implemented by the clockless logic.

14. A computer-implemented method of low-photon-count visual object classification, the computer-implemented method comprising: obtaining, by a computing system comprising one or more computing devices, a photon signature from a photon detection system, wherein the photon signature comprises a signal indicative of a photon having been received by one or more photon detectors; providing, by the computing system, the photon signature to a low-photon-count classification model; determining, by the computing system and the low-photon-count classification model, a classification of a visual object placed in a field of view of the photon detection system based at least in part on the photon signature; and providing, by the computing system, the classification as output of the low-photon-count classification model, wherein the low-photon-count classification model comprises a classification trace vector that models a likelihood of the visual object belonging to each of a plurality of classes, and wherein determining the classification comprises updating the classification trace vector based at least in part on the photon signature, and wherein the low-photon-count classification model comprises an accumulator, wherein the accumulator is configured to accumulate the classification trace vector over a plurality of photon signatures after receiving the photon signature, and output the classification when the classification trace vector has stabilized over the plurality of photon signatures.

15. The computer-implemented method of claim 14, further comprising: providing, by the computing system, the photon signature to a linear embedding space model as input, and in response to providing the photon signature to the linear embedding space model, receiving an embedding space map as output from the linear embedding space model; providing, by the computing system, the embedding space map to a feature cross model as input, and in response to providing the embedding space map to the feature cross model, receiving a plurality of feature crosses as output from the feature cross model; providing, by the computing system, the plurality of feature crosses to a linear feature cross trace model as input, and in response to providing the plurality of feature crosses to the linear feature cross trace model, receiving a plurality of feature cross trace vectors as output from the linear feature cross trace model; and determining, by the computing system, an aggregate trace vector as a combination of the classification trace vector and the plurality of feature cross trace vectors, wherein the classification is based at least in part on the aggregate trace vector.