Imaging sensor device using array of single photon avalanche diode photodetectors
By adopting a stacked pixel array layer and processing layer structure in the imaging sensor device, using SPAD and reconfigurable processing cores, the flexibility and efficiency problems of the imaging sensor device in the prior art in photon detection and image data processing are solved, and efficient autonomous processing capabilities and the execution of complex algorithms are achieved.
Patent Information
- Application Number
- CN202280102531.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-13
- Publication Date
- 2025-07-18
AI Technical Summary
Existing imaging sensor devices lack flexibility and efficiency in photon detection and image data processing, making it difficult to achieve a wide range of configurable operating modes.
The imaging sensor device is adopted in a stacked arrangement, which includes a pixel array layer and a processing layer. The pixel layer is composed of SPAD and a preprocessing circuit. The processing layer is composed of multiple processing cores. The processing cores can communicate in two directions and can be reconfigured to provide autonomous processing capabilities.
It realizes the flexibility and efficiency of the imaging sensor device, can perform custom algorithm execution and data processing, supports complex operations such as LSTM calculation, and improves the flexibility and efficiency of image data processing.
Smart Images

Figure CN120345263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an imaging sensor device, and more particularly to an imaging sensor device that uses photon detection using single photon avalanche diodes and has improved and configurable flexibility in image data processing. Background Art
[0002] Generally, an imaging sensor device includes a two-dimensional array of light detectors. The light detectors can be configured to detect one or more incident photons and provide corresponding photon detection signals to generate image information of the received light distribution. One type of light detector includes a single photon avalanche diode (SPAD), which is configured to detect light by generating electron-hole pairs and multiplying them through an electric field when performing photon detection, and the electric field generates a detectable electron avalanche.
[0003] This so-called Geiger mode photodetection cell is typically fabricated on / in a silicon substrate having a p-n junction that is electrically biased beyond its breakdown voltage such that each electron-hole pair can trigger an avalanche multiplication process to form a photon detection signal as an electrical pulse signal.
[0004] After identifying this avalanche, the avalanche is quenched by reducing the electric field that accelerates the generated electrons to stop the avalanche process. Then, the electric field is increased again to prepare the photodetection cell for the next photon detection.
[0005] Document US 9,210,350 discloses an imaging system that includes a pixel array. The pixel array includes a plurality of pixels, where each pixel of the plurality of pixels includes a single photon avalanche diode (SPAD) coupled to detect photons in response to incident light. A readout circuitry includes a plurality of photon counters, where each photon counter of the plurality of photon counters is coupled to a corresponding pixel of the plurality of pixels to count the number of photons detected by the corresponding pixel of the plurality of pixels. Each photon counter of the plurality of photon counters is coupled to stop counting photons that reach a threshold photon count for the corresponding pixel of the plurality of pixels, and where each photon counter of the plurality of photon counters is coupled to continue counting photons that do not reach the threshold photon count for the corresponding pixel of the plurality of pixels. A control circuitry is coupled to the pixel array to control the operation of the pixel array and includes an exposure time counter coupled to count the exposure time elapsed before each pixel of the plurality of pixels detects a threshold photon count. The corresponding exposure time counts and photon counts for each pixel of the plurality of pixels of the pixel array are combined.
[0006] The document A.C. Ulku, C. Bruschini, I.M. Antolovic et al., “A 512×512 SPAD image sensor with integrated gating for widefield FLIM”, IEEE Journal of Selected Topics in Quantum Electronics, Vol. 25, No. 1, pp. 1 - 12, 2019 discloses an image sensor with 512×512 photon counting pixels, each pixel including a single photon avalanche diode (SPAD), a 1-bit memory, and a gating mechanism capable of turning the SPAD on and off. The sensor is designed to achieve a high frame rate.
[0007] The literature "Megapixel time-gated SPAD image sensor for 2D and 3D imaging applications" by K. Morimoto, A. Ardelean, M.-L. Wu, etc. in Optica, Vol. 7, No. 4, pp. 346-354, 2020 discloses a 1-megapixel single-photon avalanche diode camera, which is characterized by a 3.8-nanosecond time gating and a frame rate of 24 kfps.
[0008] The literature "A CMOS SPAD imager with collision detection and 128 dynamically reallocating TDCs for single-photon counting and 3D time-of-flight imaging" by C. Zhang, S. Lindner, I.M. Antolovic, M. Wolf and E. Charbon in Sensors, Vol. 18, No. 11, 2018 discloses a single-photon avalanche diode (SPAD) sensor, which has a pixel-by-pixel time-to-digital converter (TDC) architecture to achieve a high photon flux. A 32×32 pixel SPAD sensor fabricated using 180 nm CMOS image sensor technology is disclosed, in which a dynamically reallocated TDC is implemented to achieve the same photon flux as that of the pixel-by-pixel TDC. Every 4 TDCs are shared by 32 pixels through a collision detection bus.
[0009] The document "A high-PDE, backside-illuminated SPAD in 65 / 40-nm 3D IC CMOS pixel with cascaded passive quenching and active recharge" by S. Lindner, S. Pellegrini, Y. Henrion, B. Rae, M. Wolf, and E. Charbon, published in IEEE Electron Device Letters, Vol. 38, No. 11, pp. 1547-1550, 2017, discloses a detector pixel based on a single-photon avalanche diode (SPAD) fabricated with backside-illuminated (BSI) 3D IC technology. The chip stack includes an image-sensing layer fabricated with 65-nm image sensor technology and a data-processing layer fabricated with 40-nm CMOS. Using simple CMOS-compatible technology, the pixel can perform passive quenching and active recharge at voltages far higher than those applied by a single transistor while ensuring that no device exceeds the reliability limits of gate-source (VGS), gate-drain (VGD), and drain-source (VDS).
[0010] The object of the present invention is to improve the architecture of an imaging sensor device having a stack of an image-sensing layer and a data-processing layer, thereby providing a wide range of configurable operating modes for efficient processing. Summary of the Invention
[0011] This object has been achieved by an imaging sensor device according to claim 1. Further embodiments are indicated in the dependent subclaims.
[0012] According to a first aspect, there is provided an imaging sensor device in a stacked arrangement, the imaging sensor device comprising:
[0013] - a pixel array layer including a plurality of pixel segments, each pixel segment having a plurality of pixels for photon detection, each pixel providing a digital pixel output;
[0014] - a processing layer including a plurality of processing cores, each processing core associated with one of the plurality of pixel segments to receive the pixel outputs of the pixels of the corresponding pixel segment, wherein the processing cores each communicate bidirectionally with one or more adjacent processing cores;
[0015] Each of the processing cores is configured to receive pixel outputs of pixels of an associated pixel segment and to distribute processing of the pixel outputs between the processing core and at least one of the adjacent processing cores.
[0016] The above imaging sensor device is a reconfigurable and scalable computational imaging sensor having fully autonomous processing capabilities provided by processing cores. The flexibility of the architecture stems from the ability to run custom programs / algorithms in each processing core and also from the reconfigurable hardware at the pixel interface that can be customized by software at runtime.
[0017] In addition, the pixels of the pixel array layer include detector diodes, in particular SPADs; and preprocessing circuitry coupled to the detector diodes to provide corresponding pixel outputs as signals, where specifically, the signals can be provided to the processing layer, for example, through vias in the pixel array layer. The pixel array layer is provided on a semiconductor (e.g., silicon) substrate and is processed by semiconductor processing techniques to produce the structure of the pixels and the preprocessing circuitry.
[0018] The processing layer can also be provided on a semiconductor (e.g., silicon) substrate and is processed by semiconductor processing techniques to produce the structure of the pixels and the preprocessing circuitry.
[0019] According to one embodiment, each processing core can include a front-end block that can include combinatorial logic and / or at least one look-up table, the combinatorial logic and / or the at least one look-up table being freely configurable to provide masking and / or logical operations on the pixel outputs for preprocessing and / or combining the pixel outputs.
[0020] It can be provided that the timing block is configured to receive the preprocessed pixel outputs and perform various timing functions known in the art such as pulse width and / or phase shift measurements.
[0021] In addition, a processing block and a control block can be provided in each processing core of the processing cores, the processing block including a set of general-purpose registers, an arithmetic and logic unit (ALU), and RAM, the control block being configured to provide flexible data processing capabilities, where the control block is configured to control the data processing operations of the processing block.
[0022] Specifically, a control block of at least one processing core associated with a pixel segment may be configured to split one or more processing tasks to be performed on the pixels of the one pixel segment into processing portions, where at least two of the processing portions may be processed simultaneously in parallel, and where the at least two processable processing portions are executed in the processing core and at least one adjacent processing core as follows: directly controlling, in the processing block of the processing core, the processing of the corresponding processing portion, and instructing the at least one adjacent processing core to execute the other processing portions in the corresponding processing portion respectively.
[0023] According to one embodiment, the processing task may be an LSTM calculation performed on an image detected by the imaging sensor device, where the LSTM calculation includes matrix operations, addition operations, sigmoid operations, and hyperbolic tangent (tangens hyperbolicus) operations, and where the control block of each processing core in the processing core may be respectively configured to
[0024] - split the execution of the matrix operation and the addition operation to be performed on the pixels of the one associated pixel segment into multiple processing portions,
[0025] - execute at least one of the multiple processing portions in the corresponding control block, and
[0026] - transfer at least one of the multiple processing portions to at least one adjacent processing core among the adjacent processing cores and instruct the at least one adjacent processing core to execute the corresponding at least one of the multiple processing portions. Description of the Drawings
[0027] The embodiments are described in more detail in conjunction with the accompanying drawings, in which:
[0028] Figure 1 A top view of the imaging sensor device is shown.
[0029] Figure 2 A cross-sectional view through the imaging sensor device is shown.
[0030] Figure 3 A circuit diagram of the pixel circuitry of the pixel array layer is shown.
[0031] Figure 4 The pixel layout is schematically shown.
[0032] Figure 5 A close-up top view of the edge of the processing layer substrate is shown.
[0033] Figure 6 A block diagram of the processing core is shown.
[0034] Figure 7 Shows a schematic diagram of a reconfigurable front - end circuit block.
[0035] Figure 8 Shows a group arrangement of pixel outputs associated with a LUT.
[0036] Figure 9 Shows a schematic diagram of a timing block implemented in each processing core.
[0037] Figure 10 Shows a processing block implemented in each processing core.
[0038] Figure 11 Shows the architecture of an imaging sensor device, which includes a plurality of processing cores arranged in a grid.
[0039] Figure 12 Presents a scheduling diagram of equations implementing LSTM inference operations.
[0040] Figure 13 Shows an exemplary arrangement of 5 processing cores with one main processing core and four slave processing cores. Detailed Description
[0041] Figure 1 Shows a top view of the imaging sensor device 1, and Figure 2 Shows a cross - sectional view through the imaging sensor device 1, which has a stack of a pixel array layer 2 and a processing layer 3. Both layers 2 and 3 can be fabricated on a silicon substrate, for example, using CMOS technology and / or FinFET 3D technology.
[0042] The pixel array layer 2 can be an exemplary 12×24 array (or any other size) of SPAD pixels 21 (SPAD: Single Photon Avalanche Diode), and the pixels are grouped into pixel segments 22 consisting of 4×4 SPAD pixels. Each pixel output is electrically connected to the processing layer 3 through a Through - Substrate Via (TSV) 23.
[0043] The processing layer 3 has a 3×6 array of independent processing cores 31, and each processing core is connected to the SPAD pixels 21 of a corresponding pixel segment 22. The processing cores 31 can share information and exchange data with their directly adjacent processing cores 31, and can be synchronized with each other by using internal and external handshake signals. The processing layer 3 is designed for 3D integration with large TSV landing pads 32, and the structure of the large TSV landing pads is similar to the structure of traditional flip - chip ball - bond pads.
[0044] The processing layer 3 contains digital processing electronics and processes the raw pixel array data of the pixel array layer 2 (including the pixel output of each SPAD pixel 21) as a front end. This stacked architecture offers the possibility of using the processing layer 3 as a general-purpose readout IC, which is coupled to a custom detector technology that is not limited to SPAD-based pixel arrays.
[0045] Figure 3 The circuit diagram of the electronic circuit system of the pixel array layer 2 as an exemplary pixel schematic is shown. The detector diode D1 (such as an SPAD) is in series with the cascode transistor T1 and the reset transistor T2 between the high voltage potential VHV and the ground potential GND. The cascode transistor T1 is used to extend the bias voltage range of the pixel, while the reset transistor T2 implements clock-driven active recharge controlled by the provided RST signal.
[0046] During avalanche, the voltage at the node A between the transistors T1 and T2 will rise, and depending on the state of the gate transistor T3 coupled to the node a as a transfer gate, it can act on the node B and the gate of the transistor T5. In the case where the level of the node B is high, the transistor T5 will drive the gate of the super-large (thick oxide) transistor T8, and the super-large transistor discharges the large parasitic capacitance of the TSV23. The RST signal will also drive the transistor T4 coupled to the node B to reset the node B, the transistor T6 serially coupled to the transistor T4, and the transistor T7 serially coupled to the transistor T8 to charge the capacitance of the TSV 23. The voltage level conversion can be achieved by setting the power supply voltage VDDBOT to, for example, 0.8V plus the threshold voltage of the transistor T8.
[0047] Figure 4 The pixel layout is schematically shown. The detector diode 21a is located at the center of the pixel area. The octagonal shape at the lower left is the metal contact of the TSV 23. All pixel circuit systems are located in the rectangular section 21b at the bottom and on the left. Due to the small area and pitch requirements, all transistors can be formed as thick oxide NMOS to circumvent the minimum pitch rules between different types of transistors. The pixel areas of adjacent pixels are indicated by dashed lines.
[0048] The processing layer 3 can accommodate the TSV landing pads 32 to the underlying processing core 31 through a set of ESD protection diodes (not shown), and the landing pads can be formed by rectangular aluminum contacts in groups of 16. The TSV landing pads 32 are used to receive the pixel output.
[0049] Figure 5 A close-up top view of the edge of the processing layer substrate is shown, in which the TSV landing pads 32 can be seen next to the normal-sized bond pads 33 for the external connection of the imaging sensor device 1.
[0050] Figure 6 A block diagram showing a schematic diagram of a processing core 31. Each processing core 31 is connected to 16 pixels on a pixel array layer 2 arranged in a 4×4 pixel pattern. Pixel output signals (pixel outputs) from the pixel circuitry on the pixel array layer 2 are connected to a reconfigurable front-end block 41, which includes combinational logic and a look-up table (LUT).
[0051] The front-end circuit block 41 can preprocess the pixel output data before transferring it to the timing block 42 and / or the processing block 43. The timing block 42 is a dedicated circuit that can perform various timing functions, such as pulse width or phase shift measurement.
[0052] The processing block 43 is a collection of a general-purpose register 431, an arithmetic and logic unit (ALU) 432, and a RAM 433. The overall operation of the processing core 31 is coordinated by a control block 44, in which an instruction ROM 441 and an instruction decoder 442 can be accommodated, and the instruction ROM and the instruction decoder define the operation of the processing block 43.
[0053] The control block 44 can receive inputs from adjacent processing cores 31 and, in turn, can provide software control outputs for synchronization. Similarly, the timing block 42 can receive inputs from adjacent processing cores 31 and can provide fast-propagating signals through a dedicated channel, e.g., outside the device through other processing cores 31.
[0054] The front-end block 41 and the timing block 42 are synchronized by implementing a dedicated signal path (1 bit line), through which the output of the front-end module or the timing module can be routed to the front-end block and / or the timing block. Thus, the timing blocks 42 of all adjacent processing cores 31 can be triggered by the main processing core detecting the same photon. In this case, the output of the front-end circuit block 41 of the main processing core 31 will be routed to all four adjacent processing cores 31 through a multiplexer, and the four adjacent processing cores can process the output as if it came from their own front-end circuit block 41 (slave processing cores 31).
[0055] When the range of the timing block 42 needs to exceed 16 bits, the most significant bits of the timing block 42 can be routed to the timing blocks 42 of adjacent processing cores 31, and these bits are used as a counter clock at the timing block, which actually uses two separate 16-bit counters as a single 32-bit counter.
[0056] The control block 44 generates a synchronization signal as a simple 1-bit signal from an adjacent processing core 31, and the signal can be checked using conditional instructions in the code. For example, if the adjacent processing core 31 has sent a signal, jump to line XX. There are also dedicated instructions for emitting a signal to a specific adjacent processing core 31, such as a strobe signal for the adjacent processing core 31.
[0057] Figure 7 A schematic diagram of the reconfigurable front-end circuit block 41 is shown. According to Figure 8 the diagram shown in, sixteen pixel outputs PXL[0…15] are connected in groups of 4 to a group of 4 LUTs 411 (LUT0 to LUT3). The aim is to create the possibility of binning 4×4 pixels into groups of 2×2. Each LUT can be programmed by the control block 44 in 4 instruction cycles to implement any logic function of the following type:
[0058] Q3Q2Q1Q0 = a[0]P3P2P1P0 + a[1]P3P2P1P0 +... + a
[15] P3P2P1P0
[0059] where P and Q are 4-bit LUT inputs and outputs respectively, and a[] is an array of 16 values consisting of 0s and 1s.
[0060] The outputs of the first layer of LUTs LUT0, LUT1, LUT2, LUT3 are connected to the second layer of a circuit containing an adder 412 and a LUT4 413. The adder adds the 16 bits of the output Q of the LUT 411 to form a single 5-bit number, which essentially counts the number of 1s. LUT4 is another larger version of 4, with 8-bit inputs and 1-bit output. Contrary to the adder, only the most significant two bits of the output from the previous LUT 411 are connected to the LUT4.
[0061] LUT4 can also be configured by the control block 44 in 16 instruction cycles to implement any function of the following type:
[0062] Q0 = a[0]P7P6P5P4P3P2P1P0
[0063] + a[1]P7P6P5P4P3P2P1P0
[0064] +...
[0065] + a
[255] P7P6P5P4P3P2P1P0
[0066] where P is an 8-bit input assembled from 2 bits of the output of LUT 411, Q is a 1-bit output, and a is an array of 256 values consisting of 0s and 1s.
[0067] Sixteen pixel outputs are connected to two OR trees 414 and can be individually masked using a set of AND gates 415. This secondary path is designed with a separate set of constraints to allow for fast signal propagation and can be used as an input to timing block 42 or other adjacent processing cores 31.
[0068] Various signals from the front-end circuit block 41 (such as, but not limited to, LUT and adder outputs and raw pixel values) are connected to output DMUX 416, which, like all other circuits, can be controlled by control block 44. This allows software to flexibly select various preprocessed versions of the inputs without having to reconfigure the front-end circuit block 41.
[0069] Figure 9 A schematic diagram of the timing block 42 implemented in each processing core 31 is shown. A 16-bit counter 421 acts as the central element of the timing block 42, and its value is used as the timing block output. A set of multiplexers 422 is used to select which sources serve as the counter clock signal C and the enable signal EN, and there is a wide range of choices available for both cases. The inputs to the multiplexers 422 can be Figure 7 fast output signals. The counter clock signal and the enable signal can be selected in various ways from the front-end block 41.
[0070] The enable signal EN can directly originate from the front-end circuit block output or pass through an SR latch 423, which can combine two separate input signals from the front-end output. Additionally, pulses generated by adjacent processing cores 31 or the control block 44 can be used as the counter clock signal C or the enable signal EN.
[0071] The local oscillator 424 can be formed by a ring consisting of multiple (e.g., seven) NAND gates and can be used to generate a higher-frequency clock reference for the counter.
[0072] The timing block 42 generates two flags that can be used by the control block 44 for conditional instructions: the counter overflow CO and the latch set LS. The counter overflow CO is set when an overflow is detected in the counter 421 and is essentially the 17th counter bit latched. The latch set LS is the state of the input SR latch and can be used to detect the arrival of an input.
[0073] The control block 44 is able to reconfigure the functionality of the timing block 42 by setting all the multiplexers 422 and resetting the counter 421 and the two flags CO, LS. A full reconfiguration requires two instruction cycles, but for most cases, a single cycle will be sufficient.
[0074] As Figure 10As shown, processing block 43 includes a plurality of general registers 431, a byte selector block 432, an ALU 433, and a RAM 434. The inputs of processing block 43 are connected to the front-end block 41, the timing block 42, and adjacent processing cores through a set of multiplexers managed by control block 44. The output is the RAM memory itself, which can be read out from processing core 31 through an external system circuitry or a set of registers connected to adjacent processing core 31.
[0075] The general registers 431 can be loaded with data from input ALU 433 or RAM 434. The load signals of general registers 431 are independently driven by control block 44 to enable the same data to be written to multiple locations simultaneously when needed.
[0076] The byte selector block 432 is a dedicated circuit for shifting a specific byte out of or extracting a specific byte from a data word presented at its input. The byte selector block can be used alone with dedicated instructions or can be used in combination with other operations. The following table summarizes the manipulations that the byte selector block 432 can perform:
[0077]
[0078] ALU 433 is a combinational circuit block with three inputs and a single output of the same bit size. The ALU control signal selects one of its 25 possible operations for calculating the output. The following operations can be performed: NOT O = NOT Ain, AND O = Ain AND Bin, OR O = Ain OR Bin, XOR O = Ain XOR Bin, NEG O = -Cin, ADD O = Ain + Cin, SUB O = Ain - Cin, MUL O = Ain x Bin, MAC O = Ain x Bin + Cin, CMP Ain < Bin, RL O = Ain[30:0] & Ain
[31] , RR O = Ain[0] & Ain[31:1], SL O = Ain << 1, SR O = Ain >> 1, MAX O = max(Ain, Bin), and MIN O = min(Ain, Bin).
[0079] In addition to the integer output, ALU 433 also generates two flags for conditional jumps or instruction calls: zero and carry. These flags are set or cleared according to the result of an arithmetic operation and remain in the same state until another operation acts on them. As an exception, no logical operation affects the carry flag.
[0080] The inputs Ain, Bin, and Cin to the ALU can be provided from multiple sources, which can be general-purpose registers 432, explicit RAM addresses, pointers to RAM addresses, or hard-coded values in the instruction code. Similarly, the operation result can be written to a register or an explicit or pointed-to RAM location.
[0081] The memory can be a dual-port RAM block. The read address port and the write address port are independently controlled to allow the data source flexibility described above. Additionally, the RAM can be accessed externally by overriding all connections, which is a feature for debugging and extracting the processor output.
[0082] The control block 44 includes an instruction memory 441 formed as a ROM, an instruction decoder circuit 442, and a finite state machine 443. 1-bit synchronization signals from each adjacent processing core 31 can be used for conditional instructions during runtime. Similarly, 1-bit output signals are connected to each adjacent processing core 31 and can be gated or set by software. Additionally, two 1-bit inputs and one 1-bit output can operate in the same manner as the connections presented by the adjacent processing core 31 and are designed for synchronization with modules external to the device.
[0083] The instruction memory 441 can be a 256×24-bit dual-port RAM block. Compared with the scratchpad RAM from the processing block 43, the processing core 31 can only access the read port of the instruction memory 441, and thus the read port functions like a ROM. The write port can be connected to the device bus and is only used during setup or in special cases where the program execution is paused and the instruction memory 441 can be rewritten during runtime.
[0084] The processing core 31 can be configured to follow the fetch-decode-execute sequence, which exactly occupies 3 clock cycles for each instruction. During the fetch phase, the instruction pointed to by the program counter register (PC) is read from the ROM 441 and passed to the instruction decoder 442, which is a combinational circuit that drives all the processing core control signals. During the decode phase, in addition to setting the control signals, the required data can be fetched from the RAM by directly driving the RAM address bus or using the general-purpose register as a pointer. At the end of the last phase, the operation result is written to the requested destination and the PC is incremented.
[0085] The length of each instruction can be 24 bits and starts with a variable-length opcode, followed by a payload.
[0086] The architecture of the imaging sensor device includes multiple processing cores 31, which are arranged in a grid (such as a 6x3 grid in this example), as Figure 11As shown. Each processing core 31 is connected to its four directly adjacent processing cores 31 via a bidirectional data bus 33 and a bidirectional timing signal line 34. In the case of the absence of one or more neighbors (for edge cores), the corresponding signals can be connected to a register 35.
[0087] Programming and reading out of the device can be performed via a conventional AXI bus, where the output of each processing core 31 is mapped to a memory location, which can include an instruction ROM, RAM, and any accompanying registers (when needed).
[0088] Two input signals are allocated to all processing cores 31 and can arrive at each processing core 31 in parallel and can be used for synchronization.
[0089] Programming of each processing core 31 can be performed by uploading a program to a separate instruction memory (ROM) 441 via the AXI bus. To avoid any unexpected behavior, during this program, the target processing core 31 should remain in the reset state. However, when the processing core 31 can be reprogrammed while running, there can be two exceptions: determining that the execution of a certain part of the program will not occur before the reprogramming is completed, or the processing core 31 remains frozen waiting for an external stimulus using a special-purpose WAIT instruction.
[0090] Instruction types can be divided into 5 categories: logical instructions, arithmetic instructions, manipulation instructions, flow instructions, and special instructions. Depending on the source of the operands and the destination of the result, most instructions have multiple variants.
[0091] Logical instructions can have one or two operands sourced from registers or RAM locations. The result can be written to any register or RAM location, including the location acting as the source. This type of operation does not support explicit operands. After execution, all ALU carry flags will be cleared, and regardless of the result, these instructions replace dedicated clear flag commands. The ALU zero flag functions normally. The four logical operations supported by the architecture are: NOT, AND, OR, and XOR.
[0092] Arithmetic instructions can be executed by the ALU: sign inversion, addition, subtraction, multiplication, MAC, MAX, MIN, and value comparison. Similarly, for logical instructions, they can act on data from general registers or RAM, but can also use three operands (MAC instruction) or explicit values (ADD and SUB instructions).
[0093] The manipulation instructions act on a single operand and are used to apply rotations and shifts or to select specific bytes. The category also contains RAM STORE and FETCH instructions that transfer data from or to the general-purpose register 432 to or from RAM, with the latter option supporting multiple destinations simultaneously. Additionally, a LOAD instruction is provided to write an explicit value into any general-purpose register 432.
[0094] The flow instructions affect the execution of the program by changing the PC register. The JUMP instruction can be used to jump unconditionally or based on the status of the available flags to any address in the instruction memory. The CALL and RET instructions can be used to execute subroutines. The former acts exactly like the JUMP instruction and its variants but pushes the PC value onto the stack so that when RET is called, program execution can resume from that point.
[0095] The highly customized architecture requires a special set of instructions that are not typically encountered in other CPUs. Communication with the adjacent processing cores 31 and external circuits can be accomplished through the SAVEN, GETN, PUTN, and TELL instructions. The first two instructions are used to sample the adjacent data buses and transfer the values to the general-purpose register 432. These two instructions support multiple sources and destinations simultaneously. The PUTN instruction latches the value from the general-purpose register onto one or more adjacent data buses. Finally, TELL is used to strobe or set the synchronization signals of the adjacent control unit or external IO pads.
[0096] The dedicated GETC and GETP instructions can be used to read data from the timing module or the front end, and these instructions support simultaneous byte selection and multiple destinations.
[0097] All combinational logic paths have their own configuration instructions that start with the front-end multiplexers (SETFM, SETTM) and LUT functions (SETLUT) and end with the fast-path OR tree (SETOR) and the timing module (SETTIME).
[0098] Finally, a special WAIT instruction can be provided to facilitate the simultaneous synchronization of the cores under asynchronous external trigger conditions. When this instruction is running, the control block 44 is frozen in the execution phase until the specified condition is met, and after the specified condition is met, it resumes operation immediately in the next clock cycle. The condition is verified by monitoring the adjacent processing cores 31 and external synchronization signals and can only be met when pulses from all the requested sources are detected (regardless of the order).
[0099] As a possible application, an LSTM LiDAR sensor is described. Long Short-Term Memory (LSTM) is a special type of artificial neural network that contains feedback connections, which allow the processing of data sequences such as audio signals or video signals. Recently, research has focused on extending the use of LSTM to LiDAR applications, in which the data stream generated by a Time-of-Flight (ToF) image sensor is processed by this network to determine the depth map of the detected scene. The unique characteristics of the above architecture allow the implementation of LSTM because the processing cores 31 can share information between them, which allows for a high degree of parallelization, and because the reconfigurable front-end block 41 can implement preprocessing techniques such as coincidence detection without speed or processing penalties.
[0100] As an example, an imaging sensor is proposed that acts as a single-point ToF detector for an X-Y scanning setup and implements LSTM cells of size 20. The front-end block 41 is configured to trigger the timing block 42 with the first detected input pulse within the exposure window (any 4×4 pixel segment or only one pixel segment, or when at least multiple pixels from a 4×4 pixel segment are excited).
[0101] The following equations describe the LSTM at time step t:
[0102] g f = σ(W f × x[t] || h[t-1] + b f )
[0103] g i1 = σ(W i1 × x[t] || h[t-1] + b i1 )
[0104] 8 i2 = tanh(W i2 × x[t] || h[t-1] + b i2 )
[0105] g o = σ(W o × x[t] || h[t-1] + b o )
[0106] c[t] = g f · C[t-1] + g i1 · g i2
[0107] h[t] = g o · tanh(c[t])
[0108] where g f 、g i1 、gi2 and g o are 20×1 arrays of values of the forget gate, input gate, and output gate, h and c are 20×1 vectors representing the hidden state and cell state saved from one LSTM iteration to the next, x is the input value given by the timing block, W f 、W i1 、W i2 and W o are 20×21 weight matrices, and b f 、b i1 、b i2 and b o are 20×1 bias arrays. All W and b values are constant and are determined during LSTM training before running. σ and tanh are the sigmoid function and hyperbolic tangent function, and respectively, ||, ×, and. represent concatenation, matrix multiplication, and Hadamard multiplication.
[0109] Figure 12 Presents the scheduling diagram of the above equations. Step S0 is simple because the concatenation operation can be replaced by a memory write. Steps S1 and S2 are the most resource-intensive, requiring a total of 420 fixed-point MACs. Respectively, steps S4, S5, and S7 require only 40 fixed-point multiplications and 20 fixed-point add / multiplications. The non-linear tanh and σ functions can be implemented as LUTs, i.e., simple memory read operations. To improve the execution speed, steps S1 and S2 are allocated to 4 separate slave cores, while the remaining steps are assigned to a single master core.
[0110] The total available RAM in each core is 512 bytes, and thus, all weight and bias coefficients must be stored as 8-bit signed fixed-point numbers with 3 fractional bits, 4 coefficients per RAM word. g f 、g i1 、g i2 and g o 、h, and c values have a 16-bit signed fixed-point representation with 3 fractional bits and are stored in pairs at each RAM location. The LUTs for the non-linear activation functions have the same data format as the data format of the previous variables.
[0111] Figure 13 Shows the arrangement of 5 processing cores 31 for LSTM implementation. The main processing core 31 is surrounded by slave processing cores 31 in four basic directions to allow the fastest possible data transfer. After the main processing core 31 completes the exposure period, it transfers the x[t] variable to the four slave processing cores and waits until all matrix multiplications and additions are completed. Then, the main processing core 31 will read the results and perform all the remaining operations.
[0112] Figure 13Shows a possible way of arranging a 5-core cluster to form a large-format image. In this case, due to geometric constraints, the processing cores 31 of the clusters at the array edges have different arrangements. Two cores cannot be used and are represented by black squares. It must be noted that the current setup can be extended such that the main processing cores 31 also perform calculations on the timestamps from their corresponding slave processing cores 31, which essentially results in a consistent LSTM imager.
[0113] The processing cores 31 can also interoperate in the form of a master-master configuration. For example: when an image is to be compressed by applying a function, the function includes a lookup value in a lookup table that uses the current pixel value or the TDC timestamp. Each processing core 31 will require a copy of this table, but in most cases, the table will be too large to fit into each processing core 31. Therefore, instead, the table can be broken into several parts and distributed among all the processing cores 31. In this way, each processing core 31 can process its own input, but if the input is outside its range, the processing core will send the input to an adjacent processing core 31 that stores the corresponding part of the table.
[0114] Although any processing core can send data only to an adjacent processing core, the data can be further transferred from the adjacent processing core to its adjacent processing core. In this way, information can be shared between any two cores, but since there is no direct physical connection between the two cores, it can be shared indirectly and more slowly.
Claims
1. An imaging sensor device arranged in a stacked configuration, the imaging sensor device comprising: - A pixel array layer comprising a plurality of pixel segments, each pixel segment having a plurality of pixels for photon detection, each pixel providing a digital pixel output; - A processing layer comprising a plurality of processing cores, each processing core being associated with one of the plurality of pixel segments to receive the pixel outputs of the pixels of the corresponding pixel segment, wherein the processing cores each communicate bidirectionally with one or more adjacent processing cores; Wherein each of the processing cores is configured to receive the pixel outputs of the pixels of the associated pixel segment and distribute the processing of the pixel outputs between the processing core and at least one of the adjacent processing cores.
2. The imaging sensor device according to claim 1, wherein the pixels of the pixel array layer comprise detector diodes, in particular SPADs; and a preprocessing circuitry coupled to the detector diodes to provide the corresponding pixel outputs as signals, wherein specifically, the signals are provided to the processing layer through vias in the pixel array layer.
3. The imaging sensor device according to claim 1 or 2, wherein each processing core comprises a front-end block comprising combinatorial logic and / or at least one look-up table, the combinatorial logic and / or the at least one look-up table being freely configurable to provide masking and / or logical operations on the pixel outputs to preprocess the pixel outputs.
4. The imaging sensor device according to claim 3, wherein a timing block is configured to receive the preprocessed pixel outputs and perform various timing functions such as pulse width and / or phase shift measurements.
5. The imaging sensor device according to any one of claims 1 to 4, wherein a processing block is provided, the processing block comprising a set of general-purpose registers, an arithmetic and logic unit (ALU) and a RAM, and wherein a control block is provided to control the operation of the processing block.
6. The imaging sensor device according to any one of claims 5, wherein the control block of at least one processing core associated with a corresponding pixel segment is configured to split one or more processing tasks to be performed on the pixels of the one pixel segment into processing parts, wherein at least two of the processing parts can be processed simultaneously in parallel, and wherein the at least two processing parts are at least partially executed in parallel in the processing core and at least one adjacent processing core by: directly controlling the processing of the corresponding processing part in the processing block of the processing core and instructing the at least one adjacent processing core to execute the other processing parts of the corresponding processing part respectively.
7. The imaging sensor device according to any one of claims 6, wherein the processing tasks include performing matrix operations and / or addition operations on the image detected by the imaging sensor device, and specifically comprise LSTM calculations, the LSTM calculations including matrix operations, addition operations, sigmoid operations and hyperbolic tangent operations, wherein the control block of each processing core in the processing cores can each be configured to: - Split the execution of the matrix operation and the addition operation to be performed on the pixels of the one associated pixel segment into a plurality of processing parts, wherein at least two of the processing parts can be processed simultaneously in parallel; - Execute at least one of the processing parts in a corresponding control block, and - Transmit at least one of the plurality of processing parts to at least one adjacent processing core adjacent to the corresponding processing core in the adjacent processing cores, and instruct at least one corresponding adjacent processing core to execute the corresponding at least one of the plurality of processing parts.
8. The imaging sensor device according to any one of claims 1 to 7, wherein the processing cores associated with the pixel segments located at the edges of the pixel array each communicate bidirectionally with one or more registers.
9. The imaging sensor device according to any one of claims 1 to 8, wherein the processing cores each communicate bidirectionally through a bidirectional data bus and a bidirectional timing signal line.
10. The imaging sensor device according to any one of claims 1 to 9, wherein the processing layer and the pixel array layer are formed on separate substrates, and the separate substrates are stacked to form the imaging sensor device.
11. The imaging sensor device according to any one of claims 1 to 10, wherein processing tasks are divided into processing parts, and at least one of the processing cores is configured to distribute the processing parts among the at least one processing core and at least one adjacent processing core adjacent to the at least one processing core in the adjacent processing cores.
Citation Information
Patent Citations
Low power imaging system with single photon avalanche diode photon counters and ghost image reduction
US9210350B2
Cited By
Lightweight readout circuit for large-array SPAD real-time data compression
CN121691956A