Efficient Image Sensor

JP2025502456A5Pending Publication Date: 2026-01-20VOXELSENSORS SRL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024543313
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-01-27
Filing Date
2023-01-19
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Conventional image sensors face inefficiencies in power consumption due to frame-based processing of similar frames, unnecessary processing of both detected and undetected photons, and the need for components like analog-to-digital converters and random access memory, leading to power inefficiencies and increased system size.

Method used

An image sensor design that integrates photodetectors directly with processing means, utilizing a neural network for parallelized computation, eliminating the need for traditional processors, ADCs, and RAM, and minimizing data transmission distances through short, direct connections to neurons.

Benefits of technology

The solution enables faster and more power-efficient processing, reduces power consumption, and eliminates the need for intermediate components, resulting in a compact and efficient optical sensing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to an image sensor (1) for efficient light sensing. The sensor (1) comprises a number of pixel sensors (3), each pixel sensor (3) having a photodetector (4), each pixel adapted to output a digital signal. The sensor (1) further comprises processing means, e.g. a neural network (5) having a number of neurons (6). Each photodetector (4) is located in the vicinity of said processing means. Furthermore, each photodetector (4) is connected to the processing means.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image sensor and a light sensing system comprising such an image sensor, in particular to an image sensor for efficient light sensing and a light sensing system comprising such an image sensor. [Background technology]

[0002] When performing image processing using a conventional or AI-based image signal processing unit, power consumption is one of the most important parameters. Typically, the photodetectors of the unit detect impinging photons in a frame-based manner. The higher the frame rate, the higher the overall power consumption. This is a disadvantage when the photons impinging on the set of photodetectors remain the same from frame to frame and across many consecutive frames. In this case, a higher frame rate will only increase the power consumption as the unit processes similar or the same frames repeatedly.

[0003] Another drawback is that it is power inefficient because all photodetector outputs must be processed, regardless of whether a photon is detected or not. For example, the photodetector output must be processed both for the logic "1" (e.g. a photon is detected, e.g. image information is available) and for the logic "0" (e.g. no photon, e.g. image information is not available). Thus, the absence of detection must also be processed, leading to unnecessary power consumption.

[0004] Another drawback is that a data bus is required to transmit data from the photodetectors to the processor. Furthermore, a quantizer with a certain resolution is required at the output of each photodetector. The higher the resolution of the quantizer, the higher the power consumption per quantization. Also, an analog-to-digital converter and a random access memory are required. Both of these require space, and the power dissipation in each component makes the system inefficient.

[0005] Therefore, there is a need for power efficient image sensors and light sensing systems, and signal processing units.

[0006] The present invention aims to solve at least some of the above problems. Summary of the Invention

[0007] It is an object of embodiments of the present invention to provide an efficient image sensor and an optical sensing system comprising such an image sensor.The above object is achieved by an apparatus and a method according to the present invention.

[0008] In a first aspect, the present invention relates to an image sensor for efficient light sensing, said image sensor comprising: a plurality of pixel sensors, each pixel sensor including a photodetector adapted to output a logic signal; - a processing means, Characterized in that each photodetector is located in the vicinity of said processing means and each photodetector is connected, preferably directly, to said processing means.

[0009] An advantage of embodiments of the present invention is that the processing of the sensing information is not performed by a conventional processor with serialized data and instructions, but by a processing structure with massively parallel input connections to each of the pixel sensors, for example implemented using a neural network. For readability, the remainder of this specification uses a neural network as the main implementation example, but the present invention is also applicable to the generalization of highly parallelized computation with a dedicated connection to each pixel sensor in the array. For example, the processing means is at least one neural network with a plurality of neurons, with each photodetector located in the vicinity of at least one corresponding neuron, and each photodetector connected, preferably directly, to at least one corresponding neuron.

[0010] An advantage of embodiments of the present invention is that they can process sensing information much faster and with much less power than conventional processors.

[0011] An advantage of embodiments of the present invention is that the need for a conventional processor is eliminated.

[0012] An advantage of embodiments of the present invention is that they eliminate the need for an analog-to-digital converter (ADC).An advantage of embodiments of the present invention is that they eliminate the need for a traditional random access memory (RAM).

[0013] An advantage of embodiments of the present invention is that they provide a power efficient optical sensing system.

[0014] It is an advantage of embodiments of the present invention that limitations of optical signal processors, such as limitations of frame-based imaging (eg, similar or the same frames have to be processed repeatedly, which is not power efficient), are overcome.

[0015] An advantage of embodiments of the present invention is that a power efficient optical sensing system is obtained by using the network to process optical data only when the detector detects an impinging photon.

[0016] An advantage of embodiments of the present invention is that a compact system is obtained, eliminating the need for long and / or wide buses for data transmission or very high speed serialized links between each detector and the corresponding neuron in the network.

[0017] An advantage of embodiments of the present invention is that intermediate components between the detectors and the neurons are avoided, thus improving the power efficiency of the system. It is an advantage that a power efficient and simple connection is obtained.

[0018] An advantage of embodiments of the present invention is that the outputs of the detectors are processed simultaneously by corresponding neurons, resulting in efficient and fast data processing using a network.

[0019] Preferred embodiments of the first aspect of the invention have one or suitable combinations of two or more of the following features.

[0020] The photodetector is a single photon detector, preferably a single photon avalanche diode (SPAD). An advantage of an embodiment of the present invention is that the SPAD acts as a quantizer, which eliminates the need for an additional quantizer. An advantage of an embodiment of the present invention is that the intensity data (as well as the intensity change data) of the impinging photon at the diode is encoded in the time domain at the output of the diode.

[0021] The distance between each detector and its corresponding neuron is preferably at most 100 micrometers, more preferably at most 50 micrometers, even more preferably at most 25 micrometers, and most preferably at most 10 micrometers. An advantage of embodiments of the present invention is that short and nearly direct connections are obtained between each detector and its corresponding neuron.

[0022] The pixel sensors may be stacked on the neurons. An advantage of an embodiment of the present invention is that a shortest possible connection between the pixel sensors and the neurons is obtained. An advantage of an embodiment of the present invention is that a need for a long bus for data transmission between the pixel sensors and the neurons is eliminated. An advantage of an embodiment of the present invention is that although the connections between the pixel sensors and the neurons are fixed, the weights in each neuron are configurable and reconfigurable.

[0023] Each neuron may be connected to at least one neighboring neuron, and each neuron may generate an output based on the output of the detector as well as on the input signal of at least one neighboring neuron. An advantage of an embodiment of the present invention is that a more accurate detection is obtained. An advantage of an embodiment of the present invention is that false positive detection and noise are filtered and reduced.

[0024] Each neuron is preferably capable of generating an output based on the current input to the neuron as well as the current state of the neuron. An advantage of embodiments of the present invention is that future possible detections are predicted. An advantage of embodiments of the present invention is that the network exhibits memory-like behavior.

[0025] The network is preferably adapted to respond to a pulse train. An advantage of an embodiment of the invention is that the single detector output (a train of pulses or spikes) is compatible with the input of the neural network. An advantage of an embodiment of the invention is that the output of the detector does not need to be transformed to be acceptable to the neural network.

[0026] The neural network is advantageously a spiking neural network. An advantage of embodiments of the present invention is that an efficient neural network is used to obtain a power efficient optical sensing system. An advantage of embodiments of the present invention is that a faster optical sensing speed is obtained using a spiking neural network. An advantage of embodiments of the present invention is that said network is compatible with the output of the detector.

[0027] In a second aspect, the present invention relates to an optical sensing system comprising an image sensor according to the first aspect and an optical system capable of generating an image of a scene on said image sensor.

[0028] Preferred embodiments of the second aspect of the invention have one or suitable combinations of two or more of the following features. The system further comprises at least one light source, said light source being adapted to illuminate dots on the scene, said light source having means adapted to scan, preferably continuously, said dots on the scene. The network is trained to extract features of a set of input signals to the network. The features are the mass, density, motion and / or proximity of objects in the scene.

[0029] In a third aspect, the present invention relates to a method for extracting features of a scene using an optical sensing system according to the second aspect, said method comprising: - an adaptation step of said neural network to receive a train of pulses; - adjusting the activation level of neurons based on said pulse train; - adapting said neurons to generate an output signal when said activation level exceeds a threshold; - a propagation step of propagating said output signal to at least one neuron of a subsequent layer of said network; - an extraction step for extracting features of the output signal of said network; has.

[0030] Preferred embodiments of the third aspect of the invention have one or suitable combinations of two or more of the following features. The method further comprises a step of determining a depth profile of the field of view, the determining step comprising: a) projecting at least one light pattern by said light source onto a field of view, said projection being performed in a time window of less than 10 μsec, preferably less than 1 μsec; b) imaging the projection pattern by the image sensor and optical system during at least one observation window within the time window, in synchronization with the projection of the projection pattern, the pixel sensors being in a false state when no light is detected by the corresponding photodetector and in a true state when light is detected by the corresponding photodetector, thereby obtaining a first binary matrix of pixels representing the field of view; c) a separation step of separating the projection pattern from ambient light noise on the first binary matrix of pixels by considering only pixels that are in the true state and have at least one neighboring pixel also in the true state on the first matrix of pixels obtained by one observation or a combination of at least two observations within the time window, thereby obtaining a second binary matrix of pixels representing the projection pattern; d) calculating a depth profile corresponding to the projection pattern based on triangulation between the light source position, the image sensor position, and the second binary matrix of pixels (triangulation can also be performed between two or more instances using relative positions of the image sensor and the corresponding second binary matrix); e) scanning said projection pattern by repeating steps a to d over all said fields of view to determine a depth profile of all said fields of view; having Each separated element of the projection pattern lies on at least two consecutive pixels in the binary representation.

[0031] said method comprising the steps of: an encoding step of encoding the intensity data and the intensity change data of the incident photons in the time domain of each detector; a decoding step of decoding the intensity data and the intensity variation data by a neural network; It further has: The method further comprises a training step of training the network to extract features of a set of input signals to the neural network.

[0032] The above and other characteristics, features and advantages of the present invention will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the invention, the description being given for the purposes of example only, without limiting the scope of the invention. [Brief description of the drawings]

[0033] The present disclosure is further explained by the following description and accompanying drawings. [Figure 1] FIG. 1 is a top view of an image sensor (1) including a neural network (5) according to an embodiment of the present invention. [Diagram 2] 2 is a cross-sectional view of an image sensor (1) and a neural network (5) according to an embodiment of the present invention, where each detector (4) is in the vicinity of at least one corresponding neuron (6). For neural networks with a certain depth, there may be additional neurons from subsequent layers in the network. Thus, the network (5) is not limited to only neurons corresponding to or interfacing with the image sensor (1). [Diagram 3] FIG. 3 is a cross-sectional view of an image sensor (1) and neural network (5) according to an embodiment of the present invention, where the detectors (4) of the pixel sensors (3) are stacked on the neurons (6), and where there is a physical connection between the pixel sensor layer and the neuron layer, including an electrical connection between each pixel sensor (3) and its corresponding interface neuron (6). The physical connection can be obtained by wafer-to-wafer bonding, whereby the electrical connection is made using Cu-Cu connections. Local pixel circuitry for controlling and / or biasing the detectors and pre-processing or processing the detector signals may be present in the pixel sensor layer. To reduce the manufacturing complexity of the pixel sensor layer, such circuitry may be partially or completely implemented in the neural network layer (5). For example, all circuitry using MOS transistors may be implemented in the neural network layer, avoiding the processing steps required for the manufacturing of MOS transistors. [Figure 4]FIG. 4 relates to an embodiment of the present invention and shows a top view of an image sensor (1) and neural network (5) where the detectors (4) of the pixel sensors (3) are stacked on top of the neurons (6). [Diagram 5] FIG. 5 illustrates an embodiment of the present invention, a cross-sectional view of an image sensor (1) and neural network (5) stacked together with at least one contact pad (8) for each of the detectors (4) and neurons (6). [Figure 6] FIG. 6 shows multiple detectors (4) feeding neurons (6) in a neural network (5) according to an embodiment of the present invention. [Figure 7] FIG. 7 illustrates an implementation of an image sensor (1) using a spiking neural network according to an embodiment of the present invention. [Figure 8] FIG. 8 illustrates the encoding and decoding of intensity of photons impinging on the pixel sensor (3) according to an embodiment of the present invention. [Figure 9] FIG. 9 illustrates the reduction of thermal and environmental noise in a detector output (11) using a neural network (5) according to an embodiment of the present invention. [Figure 10] FIG. 10 illustrates the operation of neurons in a pulse-coupled neural network (27) according to an embodiment of the present invention. [Figure 11] FIG. 11 shows three successive input signals from a pixel sensor array for an embodiment of the present invention, where (a,b,c) are detections for an ideal case with no environmental and thermal noise, (d,e,f) are detections for a realistic case with environmental and thermal noise, and (g,h,i) are the filtered outputs. [Figure 12] FIG. 12 relates to an embodiment of the present invention and shows in (a) a cross-sectional view of an image sensor (34) having a plurality of pixel sensors (35), a processing unit (38), and a memory element (37), and in (b) a top view of the plurality of pixel sensors (35) in a matrix configuration. [Figure 13]FIG. 13 relates to an embodiment of the present invention and shows a matrix configuration of the plurality of pixel sensors (35) and their detections (39), where in one pixel sensor (35'), two detections (39', 39'') occur in two successive time steps (T1, T2). [Figure 14] FIG. 14 illustrates an embodiment of the present invention showing a matrix arrangement of the plurality of pixel sensors (35) and their detections (39), where two detections (39', 39'') occur at two immediately adjacent pixel sensors (35', 35'') at two successive time steps (T1, T2) and a future detection (39''') at a future time instance (T3) is extrapolated. [Figure 15] FIG. 15 illustrates an embodiment of the present invention showing a matrix arrangement of the plurality of pixel sensors (35) and their detections (39), where two detections (39', 39'') occur at two non-adjacent pixel sensors (35', 35''') at two successive time steps (T1, T2) and a future detection (39''') at a future time instance (T3) is extrapolated. [Figure 16] Fig. 16 shows an example of a triangular object (43) moving in a scene (44) between first and second successive time instances (T1, T2) according to an embodiment of the present invention. Reference signs in the claims should not be construed as limiting the scope thereof. In different drawings, the same reference signs refer to the same or similar elements. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0034] The present invention relates to an image sensor for efficient image sensing.

[0035] The present invention will be described with respect to particular embodiments and with reference to certain drawings but the invention is not limited thereto but only by the claims. The drawings described are schematic and non-limiting. In the drawings, the size of some of the elements may be exaggerated and not drawn to scale for illustrative purposes. The dimensions and relative dimensions do not correspond to actual implementations of the invention.

[0036] The terms first, second, etc. in this specification and claims are used to distinguish between similar elements and are not necessarily used to describe an order in time, space, sequence, or otherwise. The terms so used are interchangeable under appropriate circumstances, with the understanding that the embodiments of the invention described herein are capable of operating in sequences other than those described or illustrated herein.

[0037] Moreover, terms such as top, under, and the like are used in this specification and claims for descriptive purposes and not necessarily to describe relative positions, and it is to be understood that the terms so used are interchangeable under appropriate circumstances, and that the embodiments of the invention described herein are capable of operation in orientations other than those described or illustrated herein.

[0038] In this specification, numerous specific details are described. However, it will be understood that embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this specification.

[0039] References throughout this specification to "one embodiment" or "one embodiment" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. Thus, the appearances of the phrases "in one embodiment" or "in one embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment, although they may. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this disclosure, in one or more embodiments.

[0040] Similarly, in describing exemplary embodiments of the invention, it should be understood that various features of the invention may be grouped together in a single embodiment, figure, or description for the purpose of streamlining the disclosure and aiding in understanding one or more various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single previously disclosed embodiment. Thus, the claims following the detailed description are expressly incorporated into this specification, with each claim standing on its own as a separate embodiment of the invention.

[0041] Furthermore, although some embodiments described herein include some features and not others included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the present invention and form different embodiments, as would be understood by one of ordinary skill in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0042] Unless otherwise defined, all terms used in disclosing the present invention, including technical and scientific terms, have the meanings commonly understood by one of ordinary skill in the art to which this invention belongs. As a further guide, definitions of terms are included to better understand the teachings of the present invention.

[0043] As used herein, the following terms have the following meanings.

[0044] As used herein, "A," "an," and "the" refer to both the singular and the plural, unless the context clearly indicates otherwise. By way of example, "a contaminant" refers to one or more contaminants.

[0045] The recitation of numerical ranges by endpoints includes not only the recited endpoints but also all values ​​and fractions subsumed within that range.

[0046] In a first aspect, the present invention relates to an image sensor for efficient image sensing as claimed in claim 1.

[0047] The image sensor comprises a number of pixel sensors. Each pixel sensor comprises a photodetector adapted to output a logic signal (e.g. a sequence of "1"s and "0"s, e.g. "1" being detection and "0" being absence of detection). The image sensor further comprises processing means. The image sensor is characterized in that each photodetector is located in close proximity to the processing means, e.g. in close proximity so that the power consumption due to the distance between each photodetector and the processing means is reduced. This is advantageous for obtaining a compact sensor and at the same time a sensor with minimal power losses, e.g. due to long connections between different components (e.g. between each of the photodetectors and the processing means), or long and / or wide buses for data transmission, or very fast serialized links. Such a distributed computing architecture is also much faster.

[0048] The processing means is connected to each detector in parallel. For example, the output of each photodetector is received by (or fed to) or connected to) the processing means in parallel, for example the processing means has multiple input ports for receiving the outputs of the photodetectors. This is advantageous in that the data at the output of the pixel sensors (i.e. the output of the photodetectors) is processed by the processing means simultaneously (i.e. processing the data in parallel). For example the processing means has multiple processing elements, each processing element being connected to one photodetector in a parallel manner. For example, for 1,000,000 pixel sensors, if 1 bit of data is available at the output of each pixel sensor (i.e. the output of each photodetector), then 1,000,000 bits of data are input to the processing means and processed simultaneously at one instance. And then, at another instance, another 1,000,000 bits of data are input to the processing means and processed. This is advantageous in that all these data are input to the processing means and processed efficiently and at high speed.

[0049] The processing means is preferably at least one neural network. However, the invention is not limited to the use of neural networks. Various features of the invention described throughout this specification with reference to neural networks may also be implemented in processing units of a parallel and distributed computing architecture such as those described above.

[0050] The neural network comprises a number of neurons. The image sensor is characterized in that each photodetector is located in the vicinity of at least one corresponding neuron. For example, they are located in close proximity so that the power consumption due to the distance between each photodetector and the neuron is reduced. For example, each photodetector has one corresponding neuron. For example, the distance between each photodetector and the neuron is the shortest distance and is as close as possible. This is advantageous for obtaining a compact sensor and at the same time a sensor with minimal power loss. The power loss may be due, for example, to long connections, long and / or wide buses for data transmission, or very high speed serialized links between different components (for example, between the photodetectors and the neurons).

[0051] The neural network processes the sensing information from multiple photodetectors. For example, the neural network replaces a conventional processor. This is advantageous in that it not only eliminates the need for a conventional processor, but also the need for an analog-to-digital converter (ADC) required to convert the analog data generated by the photodetectors into the digital data required by the processor. This is also advantageous in that it eliminates the need for conventional random access memory (RAM), since the neurons in a neural network can store history, at least temporarily. For example, the state of a neuron is influenced by its previous state. The elimination of all these components (conventional processor, conventional RAM, and ADC) and their connections and signal conversions between them increases the efficiency of the image sensor.

[0052] In addition, neural networks are particularly advantageous in processing sensing information faster than conventional processors and significantly reducing power consumption. For example, conventional processors must process all input sensing data, such as logic "1" as well as logic "0", whereas neural networks can process only logic "1" data and limit further processing of "0" data. For example, conventional processors consume power when repeatedly processing logic "0" data, especially at high frame rates, whereas the present invention does not need to process "0" data because no information is required for further processing.

[0053] Each photodetector is connected to at least one corresponding neuron. For example, the output of each photodetector is received (or fed or connected) by at least one corresponding neuron. This is advantageous in that the data at the output of the pixel sensors (i.e. the output of the photodetectors) are processed simultaneously by the corresponding neurons (i.e. the data are processed in parallel). For example, for 1,000,000 pixel sensors, if 1 bit of data is available at the output of each pixel sensor (i.e. the output of each photodetector), then 1,000,000 bits of data are input to the neural network and processed simultaneously at one instance. At another instance, another 1,000,000 bits of data are input to the neural network and processed. This is advantageous in that all these data are input to the neural network and processed efficiently and at high speed.

[0054] In such a sensor, it is not trivial to eliminate the traditional processor and other necessary components (e.g., ADC and traditional RAM). Combining an image sensor with a neural network can result in a mismatch between the output of the image sensor and the input to the neural network. This does not allow the detector to be placed in close proximity to the neural network. The invention is made possible by having a photodetector adapted to output a logic signal, which can generate an input (i.e., a pulse train) that is accepted by the neural network. In a preferred embodiment, the photodetector is a single photon detector.

[0055] In a preferred embodiment, there may be a synchronization element, for example a flip-flop, between each photodetector and its corresponding neuron (or between each photodetector and the processing means), which allows converting asynchronous pulses or spikes (i.e. coming from the photodetector) into synchronous pulses or spikes (i.e. entering the neural network). This allows the invention to be implemented with different types of neural networks that may require synchronous operation or synchronous input data.

[0056] In a preferred embodiment, such a network mimics a biological neural network, such as the human brain, e.g., such that information (e.g., visual input in the form of digital images or video sequences) is processed and analyzed in a manner similar to the human brain.

[0057] In a preferred embodiment, each pixel detector is capable of generating a stream of pulses, each pulse being the result of the detection of an impinging photon. The time interval between pulses is related to the intensity of the detected optical signal. The higher the intensity in a certain time window, the shorter the average time between pulses. Of course, as known to those skilled in the art, the actual time between pulses follows a Poisson distribution, with the average time being defined by the intensity of the optical signal.

[0058] In a preferred embodiment, the image sensor preferably operates between 100 nm and 10 micrometers, more preferably between 100 nm and 1 micrometer. For example, each photodetector can detect photons impinging on it within a detection window that is within the range of 100 nm to 10 micrometers, preferably within the range of 100 nm to 1 micrometer, and more preferably within the optical window defined by the wavelength of the light source used. For example, when building a detection system that uses a light source with a center wavelength of 850 nm, the optical window for detection is also centered on the same wavelength.

[0059] In a preferred embodiment, the single photon detector may be a single photon avalanche diode (SPAD). Such a diode acts as a photon quantizer, eliminating the need for an additional quantizer in the sensor. Moreover, such a diode allows the intensity and intensity change information of an impinging photon to be encoded in the time domain information of the output of the diode. This is due to the nature of the output of the diode. The output is a stream of pulses, each pulse corresponding to a photon detected on the diode. By setting observation windows that are equally spaced from each other, the average time delay or time difference between two output pulses (i.e. of the detector) depends on the intensity (e.g. optical power) of the impinging photon. For example, for an average of three detected photons in one observation window, the average time difference between the pulses of the detector is ΔT1. On the other hand, for an average of two detected photons in one observation window, the average time difference between the pulses of the detector is ΔT2, with ΔT2>ΔT1. Thus, the intensity of the incident photon is higher at lower ΔT and lower at higher ΔT. Thus, the intensity information (and intensity changes) of the impinging photons at the diodes is encoded in the time domain of the diode output. This may be decoded using a conventional processor, e.g. the processing means, or it may be decoded using a neural network. The neural network only needs to decode the time domain information to obtain the intensity information of the impinging photons. In contrast, if an ADC and a processor are used, the ADC needs to convert all analog data of each frame of all diodes in the sensor to the digital domain, and then the processor needs to process the data, e.g. count and accumulate the photons, as well as record the intensity of the photons on the diodes. This is a power consuming and time consuming process, especially at high frame rates, and especially when the number of photodetectors is large. Due to its efficient and fast nature, neural networks are highly advantageous (compared to conventional processors) in decoding the time domain information.

[0060] In a preferred embodiment, each photodetector may preferably be directly connected to at least one corresponding neuron of the neural network (or the processing means). For example, there are no intermediate components between the photodetectors and the neural network (or the processing means). For example, each photodetector is directly connected to a corresponding neuron (or the processing means). This is advantageous in obtaining a power-efficient sensor, since intermediate components such as a quantizer with high power consumption are avoided.

[0061] In a preferred embodiment, the distance between each detector and its corresponding neuron (or said processing means) may be at most 100 micrometers, preferably at most 50 micrometers, more preferably at most 25 micrometers, most preferably at most 10 micrometers. This is advantageous for obtaining a short and almost direct connection between each detector and its corresponding neuron (or said processing means). For example, the metal pads provided on each of the detector and the neuron (or said processing means) are simply placed on top of each other (for example, are complementary to each other, for example, are permanently joined), for example, metallically connected and almost directly connected. For example, the detector and the corresponding neuron (or said processing means) are separated by the thickness of the pads. For example, there is no wire connection. In a less preferred embodiment, for example, by having a short connection between the detector and its corresponding neuron (or said processing means), the distance may be at most 1000 micrometers, preferably at most 500 micrometers.

[0062] In a preferred embodiment, the distance between each detector and its nearest neighbor may be at most 20 micrometers, preferably at most 10 micrometers, more preferably at most 5 micrometers, and most preferably at most 1 micrometer. The distance between each neuron and its nearest neighbor is similarly adapted.

[0063] In a preferred embodiment, the ratio between the distance between two neurons and the distance between the detector and the corresponding neuron is between 1:1000 and 1:5, preferably between 1:100 and 1:10. This is equally applicable if the processing means is not a neural network.

[0064] In a preferred embodiment, each detector is connected to (corresponds to) the closest neuron, however, the corresponding neuron of each detector may not be the closest, but may be, for example, the second closest, or the third closest.

[0065] In a preferred embodiment, a plurality of pixel sensors may be stacked on a plurality of neurons (or on a plurality of said processing elements of said processing means). For example, pixel sensors are stacked such that each detector is stacked on a corresponding neuron, and the distance between said detector and said neuron is minimized. Stacking may refer to arranging pixel sensors to be located in close proximity to said neuron, e.g., on top of each other, e.g., facing each other. This allows the output of each detector to be easily available for input to the corresponding neuron. For example, there is a minimum length connector between said detector and said neuron, or preferably no connector is required. This removes extra losses that may appear in a bus connection between said detector and said neuron.

[0066] In a preferred embodiment, the connection between the pixel sensor and the neuron (or the processing element of the processing means) may be fixed. For example, it is not easy to reconnect a pixel sensor to another neuron after stacking is completed. However, the weights of each neuron in the neural network are adjustable (configurable and reconfigurable). This allows a user to reconfigure the neural network without disconnecting the pixel sensor from the corresponding neuron. Similarly, the processing elements are reconfigurable.

[0067] In a preferred embodiment, there may be a misalignment between the pixel sensor and the neuron (or the processing means). The misalignment is preferably at most 10 micrometers, more preferably at most 5 micrometers, even more preferably at most 2 micrometers, most preferably at most 1 micrometer or less. Since an image sensor of the present invention may have 1,000,000 pixel sensors, the misalignment can easily be pushed down to less than 1 micrometer. This is done, for example, by aligning the top left corner and the bottom right corner of the image sensor using alignment markers, resulting in a very high alignment accuracy.

[0068] In a preferred embodiment, the detector may comprise a semiconductor, such as silicon, germanium, GaAs, InGaAs, InGaAsP, InAlGaAs, InAlGaP, InAsSbP, etc. For example, the detector may comprise different layers of different semiconductors. Also, the corresponding neuron may comprise a semiconductor, such as silicon or germanium. For example, the neural network may comprise different layers of different materials. The detector and the corresponding neuron may be connected by two contact pads, such as, for example, a conductive metal. For example, the conductive metal may be copper, gold, silver, nickel, aluminum, titanium, or preferably a combination thereof.

[0069] In a preferred embodiment, the detectors and the corresponding neurons may be on the same chip or on two different chips, with contact pads e.g. on both chips and connected, e.g. by stacking the two chips on top of each other.

[0070] In a preferred embodiment, each detector may be connected to a corresponding neuron (or said processing means) via a conductive metal connection, for example a metal pad. For example, said pad is arranged on each pixel sensor (i.e. each detector) and a complementary pad is arranged on each corresponding neuron. For example, there are 1,000,000 pads on 1,000,000 pixel sensors and 1,000,000 pads on the 1,000,000 neurons corresponding to said pixel sensors. The location of each detector in close proximity to its corresponding neuron is advantageous because it avoids additional losses due to extra wire connections. For example, additional losses due to the increased distance between each detector and its corresponding neuron would be multiplied by 1,000,000 (i.e. the same losses are accumulated for each pixel sensor), which is avoided in the present invention.

[0071] In a preferred embodiment, the neuron may be capable of firing (i.e., generating or generating an output signal) based on the output of the detector. For example, if a detector detects a photon impinging thereon, the output of the detector may be, for example, a logic "1" or, for example, a pulse, and the neuron may fire, for example, based on at least one pulse of the detector. On the other hand, if there is no at least one pulse, the neuron may not fire. This is advantageous in that data representing no detection does not need to be processed (as would be the case in conventional processors, which also process logic "0" data). This improves power efficiency and processing speed, and reduces both memory and computational burden.

[0072] In a preferred embodiment, each neuron is connected to at least one neighboring neuron. Different neurons may be connected (e.g., via a link). For example, two neighboring neurons may be located closer to each other than the other neuron, i.e., very close to each other. Each neuron may be able to fire based on the output of the detector (i.e., the corresponding detector), as well as an input signal from the at least one neighboring neuron. For example, in case of a false positive detection where the output of the detector is a logic "1" or produces a pulse, the input signal from the at least one neighboring neuron may serve as a confirmation signal to determine whether this is a false positive detection or a true positive detection. This allows noise to be reduced and filtered.

[0073] For example, the definition of a nearby neuron is not necessarily literal: a nearby neuron may be located in the neighborhood of another neuron, not necessarily the closest neuron.

[0074] In a preferred embodiment, a neuron may also be connected to and receive signals from non-adjacent neurons, for example to determine whether a minimum number of detections in the image sensor has been achieved before determining whether a detection is a false positive detection or a true positive detection.

[0075] In a preferred embodiment, each detector may be arranged in a reverse bias configuration. The detectors are capable of detecting a single photon impinging thereon. The detectors have an output for outputting an electrical detection signal upon detection of a photon. For example, the detection signal may be represented by a signal having a logic "1" (e.g., detection), and no detection signal may be represented by a signal having a logic "0" (e.g., no detection). Alternatively, the detection signal may be represented by or be the result of a pulse signal, such as a transition from logic "0" to logic "1" and then back from logic "1" to logic "0", and no detection may be represented by (or be the result of) the absence of such a pulse signal.

[0076] In a preferred embodiment, the network may have at least one layer of neurons. Preferably, the network has an input layer of neurons, an output layer of neurons, and multiple intermediate layers (e.g., at least one layer, preferably multiple layers) between the input layer and the output layer for processing data.

[0077] In a preferred embodiment, each neuron in one layer may be connected to at least one neuron, preferably multiple neurons, in the next layer. For example, each detector is connected to at least one neuron, preferably multiple neurons, in the input layer. For example, each neuron in the input layer is connected to at least one neuron, preferably multiple neurons, in the second layer.

[0078] In a preferred embodiment, there may be 1,000,000 pixel sensors and one neuron (or one processing element) corresponding to each of the pixel sensors. Thus, in the input layer of the network, there may be 1,000,000 neurons connected by 1,000,000 connections to the 1,000,000 pixel sensors, each neuron receiving the output of its corresponding detector. There are also 1,000,000 neurons in the output layer of the network, and at least 1,000,000 neurons in each intermediate layer between the input layer and the output layer of the network.

[0079] In a preferred embodiment, the neural network may be trained to process the data of the detector outputs. For example, the network may be trained to decode time domain information of the detector outputs. For example, the network may be trained to decode intensity and / or intensity change information of impinging photons at the detectors using the time domain information of the detector outputs. Processing the output of each detector using a neural network is much faster and requires much less power than processing the output using a conventional processor.

[0080] In a preferred embodiment, the network may be trained to extract features of a set of input signals to the network, e.g., at a given instance, e.g., a combined set of input signals. The features may be mass, density, motion, and / or proximity of objects in a scene, e.g., a scene to be analyzed and / or imaged (e.g., proximity of one object to another or from the detector). Other examples of features may be intensity values, contours, ultrafast events, fluorescence lifetime imaging, etc. For example, a detector monitors at least one object, preferably multiple objects, in a scene, and photons that strike the detector are transformed by the detector into a pulse train that is input to the network. The network then analyzes the input and extracts the features, e.g., by applying a transformation to the input. Thus, the network functions as an analysis module. Nevertheless, the neural network is capable of reconstructing an image of the scene. For example, an image representation module takes the output of the neural network and reconstructs an image of the scene.

[0081] In a preferred embodiment, the neural network may be adapted to respond to a pulse train (i.e., to receive a pulse train input). This is particularly advantageous in processing sensing information at a speed faster than that of a conventional processor and with much lower power consumption. For example, the output of the detector does not need to be converted to be accepted by the network. The neural network may be, for example, a spiking neural network or a pulse-coupled neural network. In a preferred embodiment, the spiking neural network may be a spiking convolutional neural network (SCNN) or a spiking recurrent neural network (SRNN).

[0082] In a preferred embodiment, the neural network comprises a spiking neural network. The inputs provided at the input layer of said spiking neural network are typically trains of pulses or spikes. For example, each neuron in the input layer receives a series of spikes as input (e.g., from at least one corresponding detector output). These input spikes are transmitted across different layers. The output of the output layer is also typically a train of pulses or spikes.

[0083] In a preferred embodiment, there may be at least one of three types of connections between the neurons in the spiking neural network: The connections may be feedforward connections (i.e., connections from the input through a hidden layer to the output layer), or alternatively, negative lateral connections or reciprocal connections.

[0084] In a preferred embodiment, each neuron in a spiking neural network may operate as an integrate-and-fire neuron. The neurons of the network may have a threshold at which they fire. For example, the neurons may fire only at the moment the threshold is crossed. When the potential of a neuron reaches the threshold, the neuron fires, generating a signal that is transmitted to other neurons. In response to this signal, the potential of other neurons increases or decreases. Thus, a spike increases or decreases the value of the potential of the neuron until the potential of the neuron decays or fires. After firing, the potential of the neuron is reset to a low value. This is advantageous in reducing power consumption, since a signal is generated only if there is a sufficiently strong change in the incident photons (e.g., determined by a predefined number of photons impinging on the pixel sensor). This reduces the power consumption of the network and suppresses the activity of some of it. For example, a neuron will not fire if only one photon is incident on the corresponding detector, and a neuron will not fire if, for example, only one photon or a small number of photons (less than the predefined number of photons required to overcome the threshold) are incident on a group of corresponding detectors. The network also helps reduce and / or eliminate the addition of zeros when computing the weighted sum of the inputs to each neuron. This reduction in data traffic translates into reduced power consumption and increased processing speed.

[0085] In a preferred embodiment, the neuron is only allowed to fire if the time since its last output is greater than the refractory period, which is advantageous to avoid the neuron firing too frequently if it is overstimulated.

[0086] In a preferred embodiment, each neuron in a spiking neural network operates as a leaky-integrate-and-fire neuron. Neurons in the network may have a threshold at which they fire, similar to integrate-and-fire neurons. However, neurons may have leaky mechanisms. For example, the potential of a neuron decreases over time. This is advantageous to only fire a neuron when there is a significant and sustained increase in potential, e.g., an increase due to noise decays over time.

[0087] In a preferred embodiment, the neural network may be a pulse-coupled neural network, in which each neuron is connected to and receives input signals from the detector of the corresponding pixel sensor as well as from neighbouring neurons. Each neuron of the pulse-coupled neural network has a linking compartment and a feeding compartment. The linking compartment receives and sums the input signals from the neighbouring neurons. The feeding compartment receives and sums the input signals from the neighbouring neurons and from the detector of the corresponding pixel sensor. The outputs of both compartments are multiplied and fed into a threshold compartment. If the input to the threshold compartment is equal to or greater than a threshold, the neuron fires. It is particularly advantageous that each of the linking compartment, the feeding compartment and the threshold compartment has a feedback mechanism (i.e. the output depends on the previous output). This is advantageous for keeping a history of the previous states of the neuron and its neighbouring neurons. Due to the feedback mechanism, the threshold is dynamic. This is advantageous for obtaining sensors that are robust to noise and can react to slight changes in the intensity of the input pattern, etc.

[0088] In a preferred embodiment, single-photon avalanche diodes are advantageously compatible with neural networks that respond to pulse trains, such as spiking neural networks and pulse-coupled neural networks, for example, the encoded intensity and intensity change information at the output of the diode is suitable as input to the neural network, which can decode the intensity information.

[0089] In a preferred embodiment, each neuron in a neural network is capable of generating an output based on the current input to the neuron and the current state of the neuron. The current state of each neuron corresponds to the previous input to the neuron and the previous state of the neuron. They are also influenced by inputs from other neurons and possibly the states of nearby neurons. Thus, each neuron can store (e.g., at least temporarily) and compare data of different detection instances to conclude the trajectory of the moving object therefrom. For example, knowing the trajectory of a moving object moving in front of the multiple pixel sensors and detected by some of the detectors in instances T1, T2, and T3, a neuron can easily determine a possible future detection in the next instance T4. For example, a neuron can predict which detector (e.g., among nearby detectors) is likely to detect an impinging photon in the next instance. Spiking neural networks are particularly advantageous for achieving this, since the output of each neuron is determined by a combination of the current state of the neuron (determined by the previous input to the neuron) and the current input (e.g., the input from the detector in this case). This allows the network and each neuron to keep history and exhibit memory-like behavior, allowing detections across different instances to be compared, helping to determine the next output for each neuron. This also allows the neural network to predict when noisy detections are likely to occur, making it easier to reduce noise and filter out noisy detections (e.g. false positive detections).

[0090] In a preferred embodiment, the image sensor may comprise at least one type of neural network, or alternatively, the image sensor may comprise different neural networks of the same or different types.

[0091] In a preferred embodiment, the pixel sensors and corresponding neurons may be arranged in multiple rows and / or multiple columns. For example, the pixel sensors and corresponding neurons may be arranged in a matrix. The corresponding neurons (i.e., in the input layer of the neural network) are preferably configured in the same or similar configuration as the detectors of the pixel sensors. The intermediate and output layers of the neural network may be configured differently. Such an arrangement allows detection of a larger area and each neuron can have neighboring neurons.

[0092] In a preferred embodiment, the image sensor may have 100 or more pixel sensors, preferably 1,000 or more pixel sensors, more preferably 10,000 or more pixel sensors, even more preferably 100,000 or more pixel sensors, and most preferably 1,000,000 or more pixel sensors. For example, the image sensor may be arranged in a matrix, and the image sensor may have 1,000 pixel sensor rows and 1,000 pixel sensor columns.

[0093] In a preferred embodiment, the image sensor may be movable, for example the orientation of the image sensor can be changed in either the x, y or z direction, for example it is possible to obtain a 3D image with just one image sensor.

[0094] In a preferred embodiment, the image sensor can be used for 3D imaging applications. For example, the image sensor can be used to visualize objects in three dimensions. Alternatively, the sensor may be capable of analysing a scene without necessarily generating an image of the scene, for example by extracting features of objects in the scene.

[0095] In a second aspect, the present invention relates to an optical sensing system according to claim 8, comprising an image sensor according to the first aspect. The optical sensing system further comprises an optical system capable of generating an image of a scene on the image sensor. For example, the means is connected to a neural network, for example to an output of the neural network. For example, an image of a scene may be reproduced by the optical system, for example the positions of points of objects in the scene may be reproduced by the optical system. The optical system may use the neural network as an array of integrators to reproduce the image information. For example, the optical input signals of each pixel sensor creating a logical output signal by each pixel sensor may be accumulated and aggregated so that an image can be generated.

[0096] In a preferred embodiment, the system may further comprise at least one light source. For example, the light source is adapted to a wavelength detectable by the pixel sensor, for example between 100 nanometers and 10 micrometers, preferably between 100 nanometers and 1 micrometer. The light source is advantageous in enabling triangulation. For example, using a system with a light source illuminated on a scene (for example, the environment) and two of the image sensors, for example arranged in different orientations with respect to each other, where the two image sensors have a shared view of the scene, it is possible to convert the xy-time data of the two sensors into xyz-time data by triangulation. For example, by adapting the light source to illuminate dots on the scene in an illumination trace, the light source has means adapted to scan the dots on the scene (preferably continuously), the image sensor monitors the dots and outputs the position of at least one object point in the scene (for example a point on the surface of the object) along the trace in multiple instances, and the xy-time data of the two image sensors can be converted into xyz-time data by triangulation, for example using the neural network. The light source may, for example, act as a reference point and synchronize the position of the point of the object on the first image sensor with the position of the point of the object on the second image sensor, creating a z-dimension. The light source may illuminate and scan the scene to be imaged in a Lissajous manner or pattern. This illumination trajectory is advantageous in that after a few illumination cycles a significant part of the image is already illuminated, allowing efficient and fast image detection. Other illumination patterns are also conceivable.

[0097] In a preferred embodiment, the system may further comprise multiple image sensors and / or multiple light sources. This may be advantageous when creating a 3D image. For example, each image sensor may be oriented differently such that a 3D perception of the imaged scene may be captured, for example by triangulating the output of the two image sensors.

[0098] In a preferred embodiment, the system may comprise image representation means, for example a screen or other image representation device, to reproduce the positions of the object points in the scene.

[0099] In a third aspect, the present invention relates to a method for extracting features of a scene using an optical sensing system according to the second aspect, said method comprising an adapting step of adapting said neural network to receive a pulse train, for example adapting each neuron to receive a pulse train generated by a corresponding optical detector.

[0100] The method further includes adjusting an activation level or potential of a neuron based on the pulse train, e.g., for each pulse received by a neuron, an activation level of the potential of the neuron is increased, the activation level representing the intensity of a photon impinging on a photodetector corresponding to the neuron.

[0101] The method further comprises an adapting step of adapting a neuron to fire (i.e. generate an output signal) when the activation level exceeds a threshold. The method further comprises a propagating step of propagating the output signal to at least one neuron in a subsequent layer of the neural network. The method may further comprise a propagating step of propagating the output signal to at least one adjacent neuron.

[0102] The method further comprises the step of extracting features of an output signal of the network, for example extracting features of an output signal of an output layer of the network.

[0103] In a preferred embodiment, the method may further comprise the step of determining a depth profile of a field of view (e.g. the scene), the determining step comprising: a) projecting at least one light pattern onto a field of view by an illuminator (e.g., said light source), said projection being performed in a time window of less than 10 μsec, preferably less than 1 μsec; b) imaging the projected pattern by a camera sensor (e.g., the image sensor) and an optical system synchronously with the projection of the pattern during at least one observation window within the time window, the camera sensor having a matrix of pixels (e.g., the pixel sensor) each including a photodetector, the pixels being in a false state when no light is detected by a corresponding photodetector and in a true state when light is detected by a corresponding photodetector, thereby obtaining a first binary matrix of pixels representing a field of view; c) a separation step of separating the projection pattern from ambient light noise on a binary matrix of pixels obtained by one observation or a combination of at least two observations within the time window, by considering only pixels that are in a true state and have at least one neighboring pixel in a true state, thereby obtaining a binary matrix of pixels representing the projection pattern; d) calculating a depth profile corresponding to the projected pattern based on a triangulation between the illuminator position, the camera position, and a second binary matrix of pixels; e) scanning the projection pattern by repeating steps a to d for all fields of view to determine a depth profile for all fields of view; having Each separated element of the pattern is located on at least two consecutive pixels in the binary representation.

[0104] In a preferred embodiment, the method may further comprise training the network to extract features of a set of input signals to the network. For example, for optical input signals to the network, the signals are pulse trains. The features may include mass, density, motion and / or proximity of at least one object in a scene. The method may further comprise applying a transform to the input signals.

[0105] In a preferred embodiment, the method may further comprise encoding the intensity and intensity change of the impinging photons in the time domain of the detector output, for example using a single photon avalanche diode to encode the intensity and intensity change of the impinging photons. The method may further comprise training the neural network to decode the intensity and intensity change data. The method may further comprise decoding the intensity and intensity change data.

[0106] The features of the third aspect (method) are as described in the first aspect (sensor) and the second aspect (system).

[0107] In a fourth aspect, the present invention relates to an image sensor for efficient image sensing.

[0108] The sensor comprises a plurality of pixel sensors, each pixel sensor comprising a photodetector, and the sensor further comprises a processing unit capable of identifying the detection of each pixel sensor. Preferably, the processing unit is replaced by the neural network of the first aspect of the invention (or the processing element of the processing means of the first aspect of the invention), wherein each pixel sensor and photodetector is connected to a corresponding neuron as described in the first aspect.

[0109] The sensor is characterized in that the unit is able to identify detections that are spatially connected (or semi-connected) over successive time steps or that show a particular pattern. For example, the detections of each pixel sensor are at least two detections over at least two successive time steps, the at least two detections corresponding to at least one pixel sensor or at least one pixel sensor and one pixel sensor immediately adjacent to the pixel sensor. For example, each pixel sensor has multiple immediate neighboring pixel sensors. For example, the two detections are spatially connected over successive time steps, e.g. detected at one detection point or adjacent detection points over successive time steps.

[0110] For example, in a first time step, the detection point may shift from a pixel sensor to a first immediately adjacent pixel sensor, and in a subsequent second time step, the detection point may shift to a second immediately adjacent pixel sensor of the first immediately adjacent pixel sensor, or the detection point may remain at the same detection point without shifting to another detection point.

[0111] This is advantageous in determining consistent detections, e.g., detections where the detections are not random, e.g., have a certain pattern. The pixel sensor sampling is fast enough so that consistent detections of an object defined by a point are detected by one pixel sensor over at least two consecutive time steps, i.e., where the object has not moved between the two time steps, or is detected by a nearby pixel sensor. For example, the dot detection does not skip one pixel sensor as it moves between a first time instance and a second time instance. This is further advantageous in allowing filtering out random detections, e.g., environmental noise and thermal noise, that are likely to be false positive detections.

[0112] The neural network is particularly advantageous for replacing the processing unit, for example to filter noisy detections, because the neural network is a fast and power efficient way to process information, for example to filter noisy detections, and furthermore, the network can be trained to find patterns of a more complex nature than can be described by more traditional rule-based processing.

[0113] In a preferred embodiment, the detection further corresponds to at least one non-neighbor pixel sensor for the pixel sensor, and there is at most one pixel sensor between each pixel sensor and the non-neighbor pixel sensor, and preferably there are at most two pixel sensors. This is a design choice based on the sampling rate of the pixel sensors. For example, if the sampling rate is not high enough, a fast moving object may skip one or two pixel sensors. In this case, the detection between the pixel sensor and the non-neighbor pixel sensor may be interpolated.

[0114] In a preferred embodiment, the sampling rate of each pixel sensor is at least 1 MHz, preferably at least 10 MHz, more preferably at least 100 MHz. This is advantageous because objects do not move very fast compared to the sampling rate of the detector. For example, the sensor has a time resolution of at least 1 microsecond, preferably at least 100 nanoseconds, more preferably at least 10 nanoseconds.

[0115] In a preferred embodiment, the unit is capable of calculating a trajectory along the pixel sensor. The trajectory describes the movement of a dot in a scene. The dot can be generalized to an object moving in a scene, the object being defined by a number of dots. For example, a trajectory of points defining the moving object is calculated, for example between a first and a second consecutive time instance, such as points having a spatially consistent trajectory within the time step. The trajectory allows filtering out points that deviate from the trajectory. The trajectory may be a line trajectory, but also a circular trajectory, a semicircular trajectory or any other trajectory.

[0116] In a preferred embodiment, the unit is capable of extrapolating future detections based on the trajectory, for example determining the most likely points to occur based on the trajectory. For example, it may be assumed that the object or pattern moves along a certain trajectory. The extrapolation requires the detection of at least two points and a trajectory defined by the two points, but the extrapolation is more accurate if based on more than two points, for example five points over five time steps, for example ten points over ten time steps. For example, an object moving right along the x-axis between a first instance T1 and a tenth instance T10 is likely to continue moving in the same direction in an eleventh instance T11 and further time instances. Preferably, the unit is capable of verifying the extrapolated detections, for example after each given time step, for example every two or three time steps, for example to check if the trajectory of the object has changed. The extrapolation is advantageous in that it allows for example to obtain faster sensors by allowing to calculate or predict future possible detections before they are actually detected. This also reduces power consumption since fewer detections are required to detect the movement of objects in the scene.

[0117] In a preferred embodiment, the unit is capable of grouping points that move along similar or identical trajectories, e.g. points that define an object. For example, the processing unit can define a contour based on said grouped points. This allows filtering out points that do not define said object or that fall outside said contour, e.g. false positive detection.

[0118] In a preferred embodiment, the sensor further comprises at least one memory element, which is capable of storing the detection values. For example, the detection values ​​are stored at least temporarily in the memory element and then read out by the unit for processing, for example for filtering false positive detection values, for calculating the trajectory, or for extrapolating the future detection values. The memory element may be replaced by the neural network. For example, the neural network may replace the functions of both the processing unit and the memory element, and may for example store data so that noisy detections are filtered. For example, one neuron of the network corresponds to one pixel sensor.

[0119] In a preferred embodiment, different regions of the neural network, e.g. different groups of neurons, are capable of processing data of pixel sensors corresponding to said neurons, e.g. to identify trajectories of motion detection along a region of interest.

[0120] In a fifth aspect, the present invention relates to a method for optical sensing, the method comprising determining a first detection at a first detection point in a first time step. The method further comprises determining a second detection at a second detection point in a second time step consecutive to the first time step. The method further comprises identifying the first and second detection points that are spatially connected (or semi-connected) over the first and second time steps. For example, the first and second detection points are the same point or immediately adjacent points.

[0121] In a preferred embodiment, the method further comprises filtering detections that are not spatially connected over the first and second time steps, for example when the first and second detection points are not the same point or immediately adjacent points.

[0122] In a preferred embodiment, the method further comprises the step of calculating a trajectory of the first and second detections, e.g. a trajectory describing a movement of said detections from one detection point to another caused by a dot moving in a scene, e.g. an object defined by a number of dots moving in a scene.

[0123] In a preferred embodiment, the method further comprises filtering detections that deviate from the trajectory, which is advantageous in filtering noisy detections, e.g. caused by thermal and environmental noise. Preferably, the method may further comprise extrapolating future detections based on the trajectory.

[0124] Features of the fifth aspect (method) may be correspondingly described in the fourth aspect (sensor). Furthermore, any features of the fifth and fourth aspects may be correspondingly described in the first and second aspects (sensor). For example, the neural network of the first aspect is particularly advantageous for the sensor of the fourth aspect. Thus, all features of the neural network of the first aspect may be combined with the sensor of the fourth aspect.

[0125] In a sixth aspect the present invention relates to the use of a sensor according to the first aspect and / or an apparatus according to the second aspect and / or a method according to the third aspect and / or a sensor according to the fourth aspect and / or a method according to the fifth aspect for optical sensing, preferably efficient optical sensing.

[0126] Further features and advantages of embodiments of the present invention will now be described with reference to the drawings, in which it should be noted that the present invention is not limited to the specific embodiments shown in these drawings or described in the examples, but only by the claims.

[0127] FIG. 1 is a top view of an image sensor (1). The image sensor (1) includes a number of pixel sensors (3), each of which includes a single-photon detector (4). The image sensor (1) further includes a neural network (5), which includes a number of neurons (6). Each detector (4) receives a photon (10) impinging thereon and generates a corresponding signal based thereon. For example, an electrical pulse is generated by each detector (4) after detecting one photon (10). Then, each detector (4) is connected to a corresponding neuron. For example, a detector (4') is connected to a neuron (6'), or a detector (4'') is connected to a neuron (6''). The detectors (e.g., 4' and 4'') and corresponding neurons (e.g., 6' and 6'') are arranged in corresponding positions such that connections are neat and minimal. For example, detector (4') and neuron (6') are connected by connection (18'), and similarly detector (4'') and neuron (6'') are connected by connection (18'').

[0128] 2 is a cross-sectional view of the image sensor (1) and the neural network (5). Each detector (4) is adjacent to and connected to one corresponding neuron (6). For example, detector (4') is adjacent to neuron (6') and connected to neuron (6') via connection (18'), and similarly detector (4'') is adjacent to neuron (6'') and connected to neuron (6'') via connection (18''). Each detector (4) and its corresponding neuron (6) are separated from each other by a distance (7).

[0129] 3 is a cross-sectional view of an image sensor (1) and a neural network (5) in which the detectors (4) of a pixel sensor (3) are stacked on top of the neurons (6). For example, the distance (7) is zero or very close to zero since the connections (20', 20'') are eliminated. For example, the distance (7) is equal to the thickness of a contact pad (not shown).

[0130] 4 shows a top view of an image sensor (1) and a neural network (5), where the detectors (4) of a pixel sensor (3) are stacked on top of the neurons (6), i.e., each detector (4) is stacked on top of a corresponding neuron (6), and such stacking is carefully done to minimize alignment errors.

[0131] FIG. 5 shows a cross-sectional view of the image sensor (1) and the neural network (5) with the respective contact pads (8) of the detectors (4) and neurons (6) stacked on top of each other. The contact pads (8) are made of a conductive metallic material, while the neurons and the detectors are made of a semiconductor material (9). The distance (7) between each detector (4) and the corresponding neuron (6) may be only the thickness of the contact pad (8). The pads may be suitable for stacking on top of each other and may for example have a complementary shape. The pads may also be permanently stacked on top of each other to prevent any further movement of the image sensor (1) with respect to the neural network (5).

[0132] FIG. 6 shows multiple detectors (4) that feed into neurons (6) of a neural network (5). Photons (10) that strike each detector (4) result in a detector output (11) that is input to each corresponding neuron. In this case, each detector has one corresponding neuron (6). The neural network can be divided into three layers: an input layer (14), a hidden layer (15), and an output layer (16). The input layer (14) has neurons (6) that receive the detector output (11) as input. The hidden layer (15) has neurons that process the input (11) to the input layer (14). The output layer (16) has neurons that output the final output of the neural network (5). Each neuron is connected to its neighbors. For example, neurons (6) and (6') are connected via an input signal (12). The input signals (12) of neighboring neurons allow neurons to communicate with neighboring neurons to achieve more accurate detection, e.g., reduce the number of false positive detections and filter noise. Each neuron in one layer is connected to at least one neuron in the next layer, and preferably to two or more neurons in the next layer. The number of neurons, connections, and layers shown in FIG. 6 are merely illustrative and may be different in practice.

[0133] FIG. 7 shows an implementation of a sensor (1) using a spiking neural network. For example, if a detector (4) of the sensor (1) detects an object or detection pattern moving along a trajectory (17) from time T=T1 through T=T2 to T=T3 as shown in the figure, the neural network (5) can predict the next detection point where an object is present based on the trajectory (17). This makes it easier to detect the object. This may be done by a neuron getting input from nearby neurons, which act as feedback information. For example, if a neuron knows (through nearby neurons) that a neuron 10 neurons away is firing, followed by another neuron 9 neurons away, which is then firing, and so on, and so on, which is 8 neurons away, it can predict when it will fire. The neural network can also reduce noise since it can predict that such detections are likely to occur. Furthermore, in addition to predicting the next detection point, the trajectory can identify consistent movement of the object, which can filter out noisy detections.

[0134] FIG. 8 shows the encoding and decoding of the intensity of photons incident on the pixel sensor (3). The encoding of the intensity is performed using a SPAD and the decoding is performed using a neural network (5). By setting observation windows (19) that are equally spaced from one another, the probability of detecting noise (20) in the windows (19) is reduced, while the probability of detecting an impinging photon remains the same. Two examples are shown in FIG. 8. In the first example, FIG. 8(a) shows a train of impinging photons on the SPAD. The detector (4) generates a train of pulses shown in FIG. 8(c), with the average time difference between the pulses being ΔT1, where ΔT1 corresponds to the intensity of the incident photon. In the second example, FIG. 8(b) shows another train of impinging photons, and the detector (4) generates a train of pulses shown in FIG. 8(d). The average time difference between the pulses is ΔT2, where ΔT2 corresponds to the intensity of the incident photon. As shown in FIG. 8(c,d), the detected intensity in FIG. 8(b) is lower than in FIG. 8(a), so ΔT2>ΔT1. A neural network could reconstruct the intensity level based on the time difference between the pulses, as shown in Figure 8(e,f).

[0135] FIG. 9 shows the reduction of thermal and environmental noise at the detector output (11) using a neural network (5). Two examples are shown in FIG. 9(a,b). In FIG. 9(a), a detector (4) detects an incident photon (10). The neuron (6) corresponding to the detector (4) decides whether this detection is true or false, based on the history of the neuron (6) and on the input signals (12) from the neighboring neurons (6', 6''). In this case, it is a false detection. On the other hand, FIG. 9(b) shows a true detection, where the neighboring neurons (6', 6'') (and their detectors (4', 4'') provide an indication to the neuron (6) that this is a true detection, and therefore a pulse is generated by the neuron (6). FIG. 9 simplifies the operation, and in a real network, a pulse is not necessarily generated at the detector output (13) after each incident photon to the detector (4). In a real network, a neuron (e.g., (6)) generates a pulse only when a detection threshold is reached for a certain period of time.

[0136] FIG. 10 illustrates the operation of a neuron in a pulse-coupled neural network (27). A receptive field compartment (21) receives input signals to the neuron. These signals may be input signals from neighboring neurons (12, 12', 12''). Alternatively, these signals may be a stimulating input signal (11) to the neuron from another layer (e.g., a previous layer) or an output signal from a detector if said neuron is in the input layer (14) of the neural network (5). Within the receptive field compartment (21), the compartment that receives the neighboring neuron input signal (12, 12', 2'') is called the linking compartment (22), and the compartment that receives input signals from both the neighboring neurons and the stimulating input signal (11) is called the feeding compartment (23). In the linking section (22), the neighboring neuron input signals (12, 12', 12'') are summed, and similarly in the feeding section (23), the neighboring neuron input signals (12, 12', 12'') and the stimulus input signal (11) are summed. The output signals of the linking section (24) and the feeding section (25) are multiplied using a multiplier (28), and then the output (29) of said multiplier (28) is fed to the threshold section (26). Each of the linking section (22), the feeding section (23) and the threshold section (26) has a feedback mechanism (not shown), where each section retains its previous state but has a decay coefficient. Thus, each section has a memory of its previous state, which decays over time. If the input to the threshold section exceeds the threshold, the neuron fires, i.e., the output of the threshold section (13) is generated. However, the threshold changes over time due to the feedback mechanism. After a neuron fires, the threshold increases significantly and then decays over time until the neuron fires again.

[0137] FIG. 11 shows three consecutive input signals from an array of pixel sensors in (a,b,c), showing an object or pattern moving between a first position (30), a second position (31) and a third position (32). This assumes that the dot does not move more than one pixel in one time step. FIG. 11 (a,b,c) is an ideal case detection without considering the thermal noise environment. FIG. 11 (d,e,f) shows three realistic consecutive input signals from an array of pixel sensors, taking into account the environmental noise, e.g. (33), and the thermal noise. FIG. 11 shows in (g,h,i) the filtering output of each input signal by a processing means, preferably a neural network. The processing means, preferably a neural network, can filter the environmental noise, e.g. (33), and the thermal noise, by checking the consistency of the object's motion, e.g. an object moving in a consistent manner in a certain direction. Furthermore, optionally, the neural network can predict the object's motion at a later time instance based on the consistent motion.

[0138] FIG. 12(a) is a cross-sectional view of an image sensor (34) having a plurality of pixel sensors (35), each pixel sensor (35) having a photodetector (36). The sensor (34) further comprises at least one memory element (37). Each photodetector (36) is connected to the memory element (37) for storing the detection output (39', 39'') of each said photodetector (36) at different time instances. The sensor (34) further comprises a processing unit (38). The processing unit (38) is able to retrieve the data stored in the memory element (37). This unit (38) may also perform other operations such as extrapolation, calculation of the detection trajectory, etc., as mentioned above.

[0139] FIG. 12(b) is a top view of the sensor 34, in which the pixel sensors 35 are arranged in a matrix configuration, e.g., having a plurality of pixel sensor rows and columns. Each detector 36 detects a photon 42 impinging thereon and generates a corresponding signal based thereon. For example, an electrical pulse is generated by each detector 36 after detecting one photon 42. Each pixel sensor 35' has a plurality of immediate neighbor pixel sensors, e.g., 35'' and a non-immediate neighbor pixel sensor, e.g., 35'''.

[0140] FIG. 13 shows a matrix of pixel sensors (35) and their detections (39), where two detections (39', 39'') occur in two consecutive time steps (T1, T2) in one pixel sensor (35'). The detections (39', 39'') are considered to be consistent detections because they occurred in two consecutive time steps (T1, T2) in the same pixel sensor (35'). Therefore, unlike (41), they are not considered to be noisy detections.

[0141] FIG. 14 shows a matrix of pixel sensors (35) and their detections (39), where two detections (39', 39'') occur at two immediately adjacent pixel sensors (35', 35'') at two successive time steps (T1, T2). Since the detections occur at two immediately adjacent pixel sensors (35', 35'') within two successive time steps (T1, T2), they are considered to be consistent detections, unlike the noisy (i.e. false positive) detection of (41). Based on the two detections (39', 39''), a trajectory (40) of detections may be calculated by the unit (38). Based on the trajectory (40), the noisy detections (41) may be filtered. Furthermore, based on the trajectory (40), future detections (39''') at future time instances (T3) may be extrapolated. This allows for prediction of detections before they occur, such as the trajectory of a moving object. The object moving in front of the sensor 34 is in this case a simple point moving along a linear trajectory 40. In reality, however, said object is more complex, as shown in the following figure.

[0142] FIG. 15 shows a matrix of pixel sensors (35) and their detections (39), where two detections (39', 39'') occur at two non-neighboring pixel sensors (35', 35''') in two consecutive time steps (T1, T2). Based on design choices, e.g., based on the sampling rate of each pixel sensor, a detection that skips one pixel sensor is considered consistent, e.g., when the object motion is faster than the sampling rate of the pixel sensor. This is an alternative to the operation of FIG. 3. Based on the two detections (39', 39''), a trajectory (40) can be calculated to filter out noisy detections (41) and extrapolate future detections (39''').

[0143] FIG. 16 shows an example of a triangular object (43) moving in a scene (44) between first and second successive time instances (T1, T2). For simplicity, it is assumed that the object (43) is defined by three detection points (45', 45'', 45'''), and that each of the three detection points (45', 45'', 45''') has moved by one pixel between the first and second time instances (T1, T2). This allows filtering out other noisy detections around the triangle, such as (46', 46'', 46'''). Other arrangements for achieving the objectives of the method and apparatus embodying the present invention will be apparent to those skilled in the art.

[0144] The description in this specification describes the details of a specific embodiment of the present invention. However, no matter how detailed the above description is in text, it will be clear that the present invention can be applied in many ways. It should be noted that the use of a specific term in describing a specific feature or aspect of the present invention should not be interpreted as meaning that the term in this specification is redefined to be limited to the specific feature or aspect of the present invention with which the term is associated. [Explanation of symbols]

[0145] 1 Image sensor 3 pixel sensor 4. Single-photon detector 5. Neural Networks 6 Neurons 7 distance 8 Contact Pads 9. Silicon 10 incident photons 11 Detector output or neuron input 12 Input signals of neighboring neurons 13 Neuron Output 14 Input Layer 15 Middle Class 16 Output layer 17 Trajectory of moving objects 18 Connection between detectors and neurons 19 Observation window 20. Noise 21 Receptive Field 22 Linking Section 23 Feeding Area 24 Linking partition output 25 Feeding plot output 26 Threshold Zone 27 Neurons in a Pulse-Coupled Neural Network 28 Multiplier 29 Multiplier Output 30 1st position 31 Second position 32 Third position 33 Noise T1 First time step T2 Second time step T3 Future time step 34 Image Sensor 35 Multiple pixel sensor 35' pixel sensor 35'' Near-pixel sensor 35'' non-adjacent pixel sensor 36 Photodetector 37 Memory elements 38 Processing Unit 39 Detection 39' First detection 39'' Second Detection 39'' Future detection 40 Trajectory 41 Noisy Detection 42 photons 43 Triangular Objects 44 Scenes 45', 45'', 45'' 3 detection points 46', 46'', 46''' 3 noisy detection points

Claims

1. An image sensor (1) for efficient light sensing, comprising: a plurality of pixel sensors (3), each pixel sensor (3) including a photodetector (4) and adapted to output a logic signal; a processing means, which is at least one neural network (5) having a plurality of neurons (6); and Each photodetector (4) is located in the vicinity of at least one corresponding neuron (6), and each photodetector (4) is connected in parallel, preferably directly, to said at least one corresponding neuron (6); the photodetector (4) is a single-photon detector, preferably a single-photon avalanche diode; Each neuron (6) receives a pulse train, Image sensor (1).

2. 2. An image sensor (1) according to claim 1, Each neuron (6) is connected to at least one neighboring neuron (6'), and each neuron (6) is capable of generating an output based not only on the output (11) of the photodetector (4) but also on the input signal (12) of at least one neighboring neuron (6'). Image sensor (1).

3. 2. An image sensor (1) according to claim 1, Each neuron (6) is capable of generating an output (13) based on the current inputs (11, 12) to said neuron (6) as well as the current state of said neuron (6). Image sensor (1).

4. 2. An image sensor (1) according to claim 1, The neural network (5) is a spiking neural network. Image sensor (1).

5. 2. An image sensor (1) according to claim 1, the distance (7) between each detector (4) and said processing means is at most 100 micrometers, preferably at most 50 micrometers, more preferably at most 25 micrometers, most preferably at most 10 micrometers; Image sensor (1).

6. 2. An image sensor (1) according to claim 1, the plurality of pixel sensors (3) are stacked on the processing means; Image sensor (1).

7. A light sensing system comprising an image sensor (1) according to any one of claims 1 to 6 and an optical system capable of generating an image of a scene on said image sensor.

8. 8. The optical sensing system according to claim 7, the system (1) further comprises at least one light source, the light source adapted to illuminate dots on a scene, the light source having means adapted to scan, preferably continuously, the dots on the scene; Optical sensing system.

9. 8. The optical sensing system according to claim 7, The network (5) is trained to extract features of a set of input signals to the network (5). Optical sensing system.

10. 8. A method for extracting scene features using the optical sensing system of claim 7, comprising: an adaptation step of adapting said neural network (5) to receive a pulse train; an adjusting step of adjusting the activation level of said neuron (6) after said neuron (6) receives each pulse; an adaptation step of adapting said neuron (6) to generate an output signal when said activation level exceeds a threshold; a propagation step of propagating said output signal to at least one neuron in a subsequent layer of said network (5); an extraction step for extracting features of the output signal of said network (5); A method having the following.

11. 11. The method of claim 10, further comprising determining a depth profile of the field of view, said determining comprising: a) projecting at least one light pattern onto the field of view by the light source, the projection being performed in a time window of less than 10 μsec, preferably less than 1 μsec; b) imaging the projection pattern by the image sensor and optical system during at least one observation window within the time window in synchronization with the projection of the projection pattern, the pixel sensors being in a false state when no light is detected by the corresponding photodetector and in a true state when light is detected by the corresponding photodetector, thereby obtaining a first binary matrix of pixels representing the field of view; c) a separation step of separating the projection pattern from ambient light noise on a binary matrix of pixels obtained from one observation or a combination of at least two of the observations within the time window by considering only pixels that are in the true state and have at least one neighboring pixel also in the true state, thereby obtaining a second binary matrix of pixels representing the projection pattern; d) calculating a depth profile corresponding to the projection pattern based on a triangulation between the light source position, the image sensor position, and the second binary matrix of pixels; e) scanning the projection pattern by repeating steps a to d for all of the fields of view to determine a depth profile for all of the fields of view; and each separated element of the projection pattern extends over at least two consecutive pixels in binary representation; method.

12. 11. The method of claim 10, further comprising: an encoding step of encoding intensity data and intensity variation data of impinging photons in the time domain of each detector (4); a decoding step of decoding the intensity data and the intensity variation data by the neural network (5); further comprising method.