Image sensor with in-sensor computing functionality and image processing method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2026-04-22
- Publication Date
- 2026-08-07
AI Technical Summary
研究表明,在传统QVGA规格传感器中,原始数据的处理和ADC环节可消耗高达总功耗的96%,而核心的像素感光电路仅占约4%,形成了严峻的“转换壁垒”
[0017] The beneficial technical effects of this application are as follows: by employing a pixel-level analog-to-digital converter based on voltage-to-frequency conversion and a reconfigurable counter, combined with global pipelined broadcast control, ultra-high energy efficiency fully parallel inductive convolution computation is achieved. This not only significantly reduces the overall power consumption and area of quantization and computation, supporting high frame rate global shutter imaging, but also forms a complete front-end processing link from optical signal to compressed feature map by integrating activation and pooling functions in the readout link. Furthermore, it supports dual-mode operation of raw data and feature map output, enhancing the system's flexibility and practicality.
Smart Images

Figure CN122534334A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of integrated circuit design, image sensor technology and artificial intelligence edge computing, and more particularly to a complementary metal-oxide-semiconductor (CMOS) image sensor with in-sensor computing (ISC) functionality and an image processing method. Background Technology
[0002] With the widespread penetration and deployment of artificial intelligence (AI) technology in edge vision applications such as unmanned systems, intelligent monitoring, the Internet of Things, and mobile devices, increasingly stringent requirements are being placed on front-end visual perception systems: they not only need to capture high-quality image information, but also need to achieve high frame rates and low latency real-time intelligent processing (such as target detection, recognition, and tracking) with extremely low power consumption and limited hardware resources. Traditional visual processing paradigms follow a separation process of "perception-transmission-computation": image sensors capture raw light signals and convert them into analog electrical signals, which are then quantized by on-chip column-level or array-level analog-to-digital converters (ADCs) to generate digital image data. This data is then transmitted via a high-speed interface to a back-end independent processor (such as a CPU, GPU, or dedicated AI accelerator) for complex algorithmic processing. This classic architecture exposes several fundamental bottlenecks when pursuing high-resolution, high-frame-rate imaging requirements: The "conversion wall" problem: High-resolution, high-frame-rate imaging means that massive amounts of analog signals need to be digitized quickly and accurately. Traditional CMOS image sensors typically use column-level shared ADCs or more peripheral ADC arrays. As speed and accuracy requirements increase, the area and power consumption of the ADC circuit increase dramatically and non-linearly, becoming the main contributor to system power consumption. Studies have shown that in traditional QVGA sensors, raw data processing and the ADC stage can consume up to 96% of the total power consumption, while the core pixel photosensitive circuit accounts for only about 4%, forming a severe "conversion wall." The "memory wall" and "IO wall" problems: The quantized raw image data (RAW data) is enormous and needs to be read out and stored at high speed, or transmitted to an external processor through limited I / O pin bandwidth. This not only consumes a large amount of dynamic power but also constitutes a bottleneck for system latency and data throughput. More importantly, for many intelligent vision tasks (such as classification and detection), the final decision relies only on high-order feature information in the image. The large amount of redundant details contained in the raw image (such as background and texture details) actually causes a huge waste of bandwidth and energy during transmission, forming a "perception wall." To address these challenges, emerging paradigms such as Near-Sensor Computing (NSC) and In-Sensor Computing (ISC) have emerged. NSC places the processing unit close to the sensor, reducing data loss during long-distance transmission. ISC goes a step further, aiming to embed some fundamental but critical computational tasks (such as feature extraction, filtering, and preliminary classification) directly into the sensor pixels or adjacent readout circuitry, performing data compression and preprocessing in the analog domain or early digital domain, thereby significantly reducing the amount of data that needs to be transmitted and processed at the source.
[0003] Despite the promising future of the ISC concept, existing technologies still have significant limitations: Existing technology one (fully parallel PWM pixel-level in-sensor computation): This method uses pulse-width modulation (PWM) pixels to directly convert light intensity to time. However, this approach has a long exposure time, limiting the frame rate; slight changes in photodiode voltage during the comparison phase result in poor linearity; dynamic range is limited; if applied to fully parallel computation, a shared digital-to-analog converter (DAC) at the array level is required, leading to poor ramp-up speed and consistency, and high fixed pattern noise (FPN); adding a buffer to each pixel to improve performance would severely sacrifice pixel area and power consumption. Existing technology two (column-parallel ADC and proximity sensing computation): This method uses a column-level single-slope ADC for quantization and performs proximity sensing computation at the edge of the sensor array. Because this approach uses a column-level analog bus, it is prone to introducing inter-column mismatch and column fixed pattern noise, resulting in a long signal settling time and a low frame rate; furthermore, column-level operation typically only supports rolling shutters, which can cause distortions such as "jelly effect" when shooting high-speed scenes. Furthermore, its computing units are located on the periphery of the array, failing to fully utilize data locality. The amount of data moved within the chip remains large, and there is still room for improvement in the reconfigurability and energy efficiency of the computing.
[0004] Therefore, there is an urgent need for a new CMOS image sensor architecture that can deeply integrate photosensitive, quantization and preliminary intelligent processing, and simultaneously achieve high frame rate, global shutter, high energy efficiency and good computational flexibility. Summary of the Invention
[0005] The purpose of this application is to provide an image sensor and image processing method with in-sensor computing capabilities, overcoming the technical bottlenecks of existing CMOS image sensors such as the "conversion wall," "memory wall," and "IO wall" in high-speed and high-precision imaging. It achieves global shutter and fully parallel operation to support high frame rate imaging, and significantly reduces data volume and transmission power consumption from the source by directly performing preprocessing calculations such as convolution within pixels, thereby providing a high-efficiency and low-latency integrated solution for edge artificial intelligence vision applications.
[0006] To achieve the above objectives, the image sensor with in-sensory computation function provided in this application includes: a pixel array comprising multiple pixel units arranged in an array; a global control unit; wherein each pixel unit includes: a sensing and conversion module for generating a pixel current based on incident light; and a processing module connected to the sensing and conversion module; the global control unit is configured to: provide a global exposure control signal to all pixel units, enabling the sensing and conversion module to synchronously sense light and generate pixel current; and generate multiple shift clock signals in a predetermined order; the processing module is configured to: convert the pixel current into pulse signals; and, under the control of the multiple shift clock signals, sequentially receive pulse signals from pixel units associated at different positions within a convolution kernel window, perform weighted accumulation of the received pulse signals according to pre-stored weight values, and output the convolution operation result corresponding to the pixel unit.
[0007] In some embodiments of this application, the processing module includes: an analog-to-digital conversion unit for converting the pixel current into the pulse signal; and a calculation unit connected to the analog-to-digital conversion unit for performing the weighted accumulation operation.
[0008] In some embodiments of this application, the analog-to-digital conversion unit is a voltage-to-frequency converter, and the oscillation frequency of the voltage-to-frequency converter is controlled by the pixel current.
[0009] In some embodiments of this application, the voltage-frequency converter includes a ring oscillator, wherein the power supply terminal of at least one inverter in the ring oscillator is powered by the pixel current.
[0010] In some embodiments of this application, the computing unit includes a reconfigurable counter; the reconfigurable counter is configured to perform enable control, counting direction control and / or counting weight control on the input pulse signal according to a pre-stored weight value.
[0011] In some embodiments of this application, the reconfigurable counter performs at least one of the following based on a weight configuration signal: according to a first configuration bit, turning on or off the path of the pulse signal to the counter; according to a second configuration bit, guiding the pulse signal to different input terminals of the counter to achieve different effective counting weights; and according to a third configuration bit, controlling the counter to increment or decrement the count.
[0012] In some embodiments of this application, the sensing and conversion module includes: a photoelectric conversion circuit for generating a pixel voltage under the global exposure control signal; and a voltage-to-current conversion circuit connected to the photoelectric conversion circuit for linearly converting the pixel voltage into the pixel current.
[0013] In some embodiments of this application, the voltage-to-current conversion circuit includes a cascaded current mirror composed of transistors with different threshold voltages.
[0014] In some embodiments of this application, an output processing unit is also included, connected to the pixel array, for performing nonlinear activation processing and pooling downsampling processing on the convolution operation result.
[0015] In some embodiments of this application, the output processing unit includes a cascaded linear rectifier circuit and a max-pooling circuit.
[0016] In another embodiment of this application, an image processing method applicable to the image sensor with in-sensory computing function is also provided. The method includes: applying a global exposure control signal to all pixel units through a global control unit to control all pixel units to synchronously perform photoelectric conversion to obtain the pixel voltage corresponding to each pixel unit; converting the pixel voltage through a voltage-current conversion circuit and a voltage-frequency converter of each pixel unit to obtain a pulse signal corresponding to the light intensity of each pixel unit; controlling the pulse signal of each pixel unit to be broadcast sequentially to different associated pixel units in a convolution kernel window in multiple consecutive periods by generating multiple shift clock signals with a predetermined order through the global control unit; weighting the pulse signal received by broadcasting according to a pre-stored weight value in each period, and accumulating the weighting result through a reconfigurable counter to obtain the convolution result corresponding to each pixel unit after all periods are completed; and constructing and outputting a feature map based on the convolution result of each pixel unit.
[0017] The beneficial technical effects of this application are as follows: by employing a pixel-level analog-to-digital converter based on voltage-to-frequency conversion and a reconfigurable counter, combined with global pipelined broadcast control, ultra-high energy efficiency fully parallel inductive convolution computation is achieved. This not only significantly reduces the overall power consumption and area of quantization and computation, supporting high frame rate global shutter imaging, but also forms a complete front-end processing link from optical signal to compressed feature map by integrating activation and pooling functions in the readout link. Furthermore, it supports dual-mode operation of raw data and feature map output, enhancing the system's flexibility and practicality. Attached Figure Description
[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of a current-mode pixel structure provided in an embodiment of this application; Figure 2 This is a schematic diagram of the three-phase state of oscillation provided in an embodiment of this application; Figure 3This is a schematic diagram of the overall structural logic provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a reconfigurable counter provided in an embodiment of this application; Figure 5 This is a schematic diagram of convolutional logic provided in an embodiment of this application; Figure 6 This is a schematic diagram of the reading logic of a pixel counter provided in an embodiment of this application; Figure 7 This is a schematic diagram of the overall array structure provided in an embodiment of this application; Figure 8 This is a schematic flowchart of an image processing method provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] Specific embodiments of this application are disclosed in detail with reference to the following description and accompanying drawings, indicating how the principles of this application can be adopted. It should be understood that the embodiments of this application are not limited in scope. Within the spirit and scope of the appended claims, embodiments of this application include many changes, modifications, and equivalents.
[0021] Features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, combined with features in other embodiments, or substituted for features in other embodiments.
[0022] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, whole, step, or component, but does not exclude the presence or addition of one or more other features, wholes, steps, or components.
[0023] The image sensor with in-sensory computing function provided in this application includes: A pixel array contains multiple pixel units arranged in an array. Global control unit; Each pixel unit includes: a sensing and conversion module, used to generate a pixel current based on the incident light; The processing module is connected to the sensing and conversion module; The global control unit is configured to: provide a global exposure control signal to all pixel units, enabling the sensing and conversion module to synchronously sense light and generate pixel current; and generate a plurality of shift clock signals in a predetermined order. The processing module is configured to: convert the pixel current into a pulse signal; and, under the control of the plurality of shift clock signals, sequentially receive pulse signals from pixel units associated at different positions within a convolution kernel window, perform weighted summation on the received pulse signals according to pre-stored weight values, and output the convolution operation result corresponding to the pixel unit.
[0024] Specifically, in practical applications, the core architecture of the image sensor with in-sensory computing capabilities provided in this application can be referenced. Figure 1 , Figure 3 and Figure 7 As shown, the image sensor mainly includes a pixel array, a global control unit, and an output processing unit.
[0025] The pixel array consists of 128 rows × 128 columns of pixel units arranged in a two-dimensional matrix (the specific size can be adjusted according to the design). Each pixel unit is the basic unit for performing photosensitivity, quantization, and calculation.
[0026] The global control unit is the core timing engine that coordinates the synchronous operation of the entire array. It primarily generates two types of key signals: first, a global exposure control signal (global shutter signal), used to simultaneously initiate the photoelectric conversion process of all pixel units in the array, eliminating motion distortion (like the jelly effect) caused by the rolling shutter; second, multiple (e.g., nine) shift clock signals (Φ0 to Φ8) with a predetermined sequence. These shift clock signals have different phases and are effective sequentially according to a specific timing sequence, used to control the broadcasting of pulse signals between pixel units and the convolution calculation process.
[0027] Each pixel unit includes at least a sensing and conversion module and a processing module.
[0028] Sensing and Conversion Module: Responsible for photoelectric conversion and initial signal processing. It includes a photodiode (PD), a transfer transistor (TX), a reset transistor (RST), and an improved current output stage, forming a structure similar to a 4T active pixel sensor (4T-APS). Specifically, its current output stage employs a cascaded structure consisting of a high threshold voltage (HVT) PMOS transistor (MP_H) and a low threshold voltage (LVT) PMOS transistor (MP_L) connected in a common-gate configuration (e.g., Figure 1 (As shown). This structure replaces the source follower (SF) in a traditional 4T pixel. By precisely designing transistor parameters (such as aspect ratio), MP_H operates at the edge of the saturation region while MP_L operates at the edge of the linear region, thereby enabling the pixel current (I) output from this cascaded structure to...pixel The voltage of the floating diffusion node (VFD) after the photodiode is reset exhibits an approximately linear relationship with that of the photodiode, thus realizing I pixel Voltage-to-current conversion to VFD.
[0029] In practical work, I pixel The approximate higher-order functional relationship with VFD (ignoring channel length modulation effects) can be referenced from the following model: ; ; The pixel current is: ; In the above two equations, This refers to the threshold voltage of a low-threshold MOSFET. Where is the thermal voltage, and n is a non-ideal constant, typically between 1 and 2, depending on the device characteristics. These are constants related to device geometry and material properties. For quantized pixel current, To correspond to the intrinsic parameters and aspect ratio of the MOSFET, , , These represent the gate-source voltage, threshold voltage, and drain-source voltage of a high-threshold MOSFET. Let V be the threshold voltage difference between the two MOSFETs. A linear approximation can be made for the pixel current biased at the boundary between the subthreshold and linear regions, assuming that V... GS,H Approaching V TH,H At that time, I pixel It is a linear function of VFD.
[0030] Processing Module: Connected to the output of the sensing and conversion module. It primarily includes an analog-to-digital converter (ADC), specifically a voltage-to-frequency converter (VFC). The core of this VFC is a three-stage ring oscillator (e.g., Figure 2 , Figure 3 As shown). The aforementioned pixel current (I) pixel The current is directed to the ring oscillator, serving as the primary current source for the charging and discharging of the interstage capacitor (CL) of its internal inverters. According to the current-to-frequency conversion principle, the frequency (f) of the pulse frequency modulation (PFM) signal generated by the oscillator... osc ) and I pixel Proportional. The three-phase oscillation signal (with a phase difference of about 120°) output by the ring oscillator is processed by subsequent shaping circuits (such as shaping inverters) and combinational logic circuits (such as three-input XOR gates) to eliminate jitter and generate a stable single-channel PFM signal (P0).
[0031] The calculation model for the oscillation frequency is as follows: ; In the above formula, I is the oscillation frequency of the ring oscillator. pixel Let I be the pixel current, N represent the use of an N-stage ring oscillator, CL be the N-stage parasitic node load capacitance, and VP be the amplitude of the resulting oscillation signal. We consider the pixel current I... pixel The parasitic capacitance CL in the cascaded node of the N-stage ring oscillator is fully charged, and the charging swing is a relatively fixed VP (related to the ratio of the output impedance of the voltage-to-current conversion module looking upward to the input impedance of the ring oscillator looking downward). Therefore, a current-controlled frequency unit similar to a current-starved ring oscillator is realized, realizing the conversion of current to the frequency domain.
[0032] The processing module also includes a computing unit, the core of which is a reconfigurable counter. This counter has a bit width of up to 9 bits (the highest bit can be used as the sign bit). Its "reconfigurable" characteristic is reflected in its ability to dynamically configure its counting behavior based on locally stored weight values (e.g., a signed weight with a value range of -2, -1, 0, +1, +2). Configuration is achieved through weight control signals, for example: Weight enable signal: Controls whether the PFM signal (P0) can be input to the counter. When the weight is 0, this path is disconnected.
[0033] Weighted amplitude signal: controls the position of the PFM signal connected to the counter. For example, connecting it to the least significant bit will count each pulse as 1, or connecting it to the second least significant bit (through logical equivalence) will count each pulse as 2.
[0034] Weight symbol signal: controls the counter to increment (corresponding to positive weight) or decrement (corresponding to negative weight).
[0035] With the above configuration, the weighted operation of the input pulses is directly realized during the counting process.
[0036] In some embodiments of this application, the processing module includes: an analog-to-digital conversion unit for converting the pixel current into the pulse signal; and a calculation unit connected to the analog-to-digital conversion unit for performing the weighted accumulation operation. The analog-to-digital conversion unit is a voltage-to-frequency converter, and the oscillation frequency of the voltage-to-frequency converter is controlled by the pixel current. The voltage-to-frequency converter includes a ring oscillator, and the power supply terminal of at least one inverter in the ring oscillator is powered by the pixel current.
[0037] In conjunction with the foregoing embodiments, the overall global collaborative workflow of the image sensor with in-sensory computing function provided in this application is as follows: Global Exposure and Quantization Stage: Under the global exposure control signal, all pixel units are synchronously exposed to light and reset, subsequently transferring charge to generate VFD. The VFD of each pixel is linearly converted to I by its voltage-to-current conversion circuit. pixel I pixel The local ring oscillator is driven to generate a PFM pulse train. Within a preset fixed time window (integration time), the reconfigurable counter of each pixel (usually configured to count positively with a weight of 1) counts its own PFM pulses, and the resulting count value (D_i) is the digitized result of the original light intensity of that pixel.
[0038] Fully parallel pipelined convolution stage: The global control unit sequentially activates nine shift clock signals Φ0 to Φ8. Within each signal cycle, the entire array performs a uniform operation: the PFM pulse signal currently generated by each pixel is routed to a reconfigurable counter of a specific "associated pixel" defined by the signal of that cycle for weighted accumulation. For example, cycle Φ0 might cause each pixel's pulse to act on its own counter, cycle Φ1 might cause it to act on the counter of its right-hand neighboring pixel, and so on. Figure 5 As shown, after nine complete cycles, for any 3x3 local window in the image, the counter of the central pixel has sequentially accumulated the pulse count values contributed by all nine pixels within that window, modulated by their own weights. Finally, the value in each pixel's counter is the result of one 3x3 convolution operation (feature map).
[0039] The output processing unit is connected to the readout circuit of the pixel array (e.g., Figure 6 , Figure 7 (As shown). It performs subsequent processing on the read convolution results in the digital domain, mainly including: Linear Rectified Lumen (ReLU) processing: Each data point read (e.g., a 9-bit signed number) is evaluated; if it is negative, the output is zero; if it is positive or zero, the original value is output. This implements the functionality of the ReLU activation function in a neural network.
[0040] Max pooling: Typically, a 2x2 window is used to downsample the data after ReLU processing. Four values are compared within each window and the maximum value is output, thus compressing the data volume to 1 / 4.
[0041] In some embodiments of this application, the image sensor has configurable data output modes, including a first mode that outputs the quantized values of the original image and a second mode that outputs the result of the convolution operation.
[0042] The sensor in this embodiment supports switching of operating modes. It can be configured to: in the first mode, directly output the raw digital values (D_i) of each pixel obtained in the quantization stage; in the second mode, output the feature map obtained after the above convolution calculation and subsequent processing (ReLU, pooling).
[0043] In other embodiments of this application, the computing unit includes a reconfigurable counter; the reconfigurable counter is configured, according to pre-stored weight values, to perform enable control, counting direction control, and / or counting weight control on the input pulse signal. Further, the reconfigurable counter performs at least one of the following based on a weight configuration signal: according to a first configuration bit, turning on or off the path of the pulse signal to the counter; according to a second configuration bit, guiding the pulse signal to different input terminals of the counter to achieve different effective counting weights; and according to a third configuration bit, controlling the counter to increment or decrement counting.
[0044] In some embodiments of this application, the sensing and conversion module includes: a photoelectric conversion circuit for generating a pixel voltage under the global exposure control signal; and a voltage-to-current conversion circuit connected to the photoelectric conversion circuit for linearly converting the pixel voltage into the pixel current. Further, the voltage-to-current conversion circuit includes a cascaded current mirror composed of transistors with different threshold voltages.
[0045] For details, please refer to Figure 4 As shown, the core of the aforementioned reconfigurable counter is a dynamic counter based on true single-phase clock logic. For example... Figure 4 As shown, it consists of a series of cascaded TSPCD flip-flops (DFFs) forming an n-bit binary counter (e.g., n=9). The highest bit (MSB, the 9th bit) Q[8] is designed as the sign bit. During the reset (RST) phase, the counter is preset to a specific initial value. According to the invention, it can be preset to "1xxxxxxxx" (i.e., the sign bit is '1' and the data bits are any value, usually set to '0') during reset, which is equivalent to a negative value of a signed binary number (e.g., -256), providing a suitable starting point for subsequent weighted accumulation (especially decrementing operations under negative weights). The TSPC structure is suitable for processing high-frequency PFM pulse signals (P0) from a ring oscillator due to its high speed and low power consumption characteristics.
[0046] The "reconfigurable" feature of the counter is achieved by additional weight control logic, which receives a 3-bit weight configuration signal (W[0], W[1], W[2]) from the local storage unit and generates real-time control over the counter's behavior. For example... Figure 4As shown, the logic mainly includes the following three functional modules: Enable control module: controlled by the weight configuration signal W[0]. Its implementation can be an AND gate or a transmission gate, connected in series on the path from the PFM pulse signal (P0) to the counter clock input terminal (CLK). When W[0]=1, the gate is open, the P0 pulse can pass normally, and the counter is enabled; when W[0]=0, the gate is closed, the P0 pulse is blocked, and the counter does not work. This directly corresponds to the case where the convolution kernel weight is "0", and the pixel does not participate in the accumulation. Weight amplitude control module: controlled by the weight configuration signal W[1]. Its core is a multiplexer (MUX). The two data input terminals of the MUX receive respectively: Path A: The original PFM pulse signal (P0).
[0047] Path B: The signal after weighted logic processing. This processing can be achieved by directly connecting P0 to the second lowest bit of the counter (the clock terminal of Q[1]) or by using equivalent logic to make each pulse cause a +2 / -2 count change, or by using a simple frequency divider or pulse widening circuit so that an input pulse can equivalently trigger two counting operations.
[0048] When W[1]=0, MUX selects path A, and each valid P0 pulse causes the counter to change by ±1 (the specific ± is determined by the direction control), corresponding to a weight absolute value of 1.
[0049] When W[1]=1, MUX selects path B, and each valid P0 pulse causes the counter to change by ±2, corresponding to a weight absolute value of 2.
[0050] Weight sign / direction control module: controlled by weight configuration signal W[2]. This signal is directly connected to the up / down mode control terminal (UP / DOWN) of the n-bit TSPC counter. When W[2]=1, the counter is set to the up (UP) mode, corresponding to positive weight. When W[2]=0, the counter is set to the down (DOWN) mode, corresponding to negative weight.
[0051] Each pixel unit contains a tiny weight storage unit, such as a 3-bit static latch or a small SRAM unit. This storage unit is used to store the specific weight value (Wij) assigned to that pixel location in the convolution kernel. During system initialization or kernel replacement, the global control unit writes the weight data of the convolution kernel matrix into the weight storage units of each pixel in parallel or serially via row and column address gating. During the convolution calculation phase, the stored weight values are read out and used as (W[0],W[1],W[2]) signals to control the reconfigurable counter in real time.
[0052] Please refer to this again. Figure 4 and Figure 5 As shown, within each of the nine cycles (Φ0-Φ8) of the convolution calculation: 1. Based on the current shift clock signal Φx, the PFM pulse signal P0 of each pixel is routed to the counter input node of a target pixel (itself or a neighboring pixel).
[0053] 2. The target pixel's weight control logic reads its locally stored weight value (W[2:0]).
[0054] 3. If W[0]=1, the pulse passes through the enable control gate.
[0055] 4. The pulse is processed by the amplitude control module into counting events with different weights (+1 / +2 or -1 / -2) according to the value of W[1].
[0056] 5. The counter responds to the counting event and updates its count value according to the direction (UP / DOWN) set by W[2].
[0057] 6. After nine cycles of accumulation, the final value in each pixel counter completes the 3x3 window convolution operation centered on itself, that is: ; In the above formula, For the calculated 3 3. Results of digital convolution before normalization The weighted result is implemented using a programmable counter. The oscillation frequency is the digital result obtained in a general pulse counter, therefore, in the digital domain (both weights and inputs are digital values), for 3... The method uses 3 convolutional kernels to solve the problem, thus realizing the computational operation of a convolutional neural network in the entire digital domain.
[0058] With the above structure, this application successfully simplifies the multiplication and accumulation (MAC) operation into a configurable counting process based on pulse counting and direction control, eliminating the need for traditional digital multipliers. This greatly reduces the hardware complexity and power consumption of the intra-pixel computing unit and is a key innovative circuit for realizing efficient and flexible intra-pixel computing.
[0059] In some embodiments of this application, an output processing unit is further included, connected to the pixel array, for performing nonlinear activation processing and pooling downsampling processing on the convolution operation result. The output processing unit includes a cascaded linear rectifier circuit and a max-pooling circuit.
[0060] Specifically, in practical applications, the cascaded linear rectifier (ReLU) circuit can be implemented using digital comparison logic. The circuit monitors the sign bit (most significant bit, MSB) of the input data. If the MSB is '1' (indicating a negative number), the output is forced to all zeros; if the MSB is '0' (indicating a positive number or zero), the original data is output. This circuit can be deployed in parallel on multiple readout channels to improve throughput. For 2x2 pooling, the max pooling circuit can employ a two-stage pipelined structure. The first stage contains two digital comparators that compare the data of two adjacent pixels in the same row and output their respective maximum values. The second stage contains one digital comparator that compares the output results from the first stage from two consecutive rows and outputs the final maximum value. This structure can continuously process scanned readout data in a pipelined manner.
[0061] Figure 6 This is a schematic diagram of the linear rectification and max-pooling circuit in the output processing unit provided in an embodiment of this application. (Refer to...) Figure 6 In this embodiment, the output processing unit is connected to the column readout bus of the pixel array and is used to perform nonlinear activation and pooling downsampling processing on the convolution operation results output by the pixel array. It includes a cascaded linear rectifier circuit and a max-pooling circuit. Specifically, the linear rectifier circuit is implemented using digital comparison logic. Its input receives the convolution operation result data (e.g., a 9-bit signed binary number) from the pixel array. ReLU activation is achieved by monitoring the sign bit (most significant bit) of the input data: when the sign bit is "1", it indicates that the data is negative, and the circuit output is forced to all zeros; when the sign bit is "0", it indicates that the data is positive or zero, and the circuit outputs the original data. This linear rectifier circuit can be deployed in parallel on multiple readout channels to match the readout throughput of the pixel array. The max-pooling circuit is connected to the output of the linear rectifier circuit and is used to perform 2×2 window max-pooling downsampling on the data after ReLU processing. This max pooling circuit employs a two-stage pipelined structure: the first stage contains two digital comparators that compare the data of two adjacent pixels in the same row and output the maximum value within their respective windows; the second stage contains one digital comparator that compares the outputs from the first stage from two consecutive rows and outputs the final maximum value. Through this pipelined structure, the max pooling circuit can continuously process the scanned data stream, compressing the data volume to one-quarter of the original convolution result.
[0062] Through the cascaded linear rectification and max pooling circuits described above, the output processing unit can directly perform activation and pooling operations commonly used in convolutional neural networks in the digital domain, further compressing the amount of output data and reducing the data handling and computational overhead of the back-end processor.
[0063] Figure 7This is a schematic diagram of the overall array structure provided in an embodiment of this application. (Refer to...) Figure 7 The image sensor in this embodiment includes a pixel array, a global control unit, and an output processing unit. The pixel array consists of multiple pixel units arranged in a two-dimensional matrix, for example, 128 rows × 128 columns. Each pixel unit integrates a sensing and conversion module, as well as a processing module including an analog-to-digital conversion unit and a computing unit. The global control unit is located on the periphery of the pixel array and is connected to each pixel unit via a global exposure control signal line and a multi-phase shift clock signal line. It synchronously provides global exposure control signals to all pixel units to control the exposure timing of the global shutter, and simultaneously generates multiple shift clock signals (e.g., Φ0 to Φ8) in a predetermined order to control the pipeline broadcasting process within the pixel array. The output processing unit is connected to the column readout bus of the pixel array and includes a cascaded linear rectifier circuit and a max-pooling circuit. It sequentially performs nonlinear activation and pooling downsampling processing on the convolution operation results output by the pixel array, thereby finally outputting compressed feature map data. Through this overall array architecture, the image sensor can complete fully parallel processing of photoelectric conversion, analog-to-digital conversion and convolution calculation at the pixel level, and realize post-processing output of feature maps at the array level, effectively reducing the data output bandwidth and the computing power requirements of subsequent processing units.
[0064] Please refer to Figure 8 As shown, in another embodiment of this application, an image processing method suitable for the image sensor with in-sensory computing function is also provided, the method comprising: S801 applies a global exposure control signal to all pixel units through a global control unit, and controls all pixel units to synchronously perform photoelectric conversion to obtain the pixel voltage corresponding to each pixel unit; S802 converts the pixel voltage through the voltage-current conversion circuit and voltage-frequency converter of each pixel unit to obtain a pulse signal corresponding to the light intensity of each pixel unit. The S803 generates multiple shift clock signals in a predetermined order through a global control unit to control the pulse signal of each pixel unit to be broadcast sequentially to different associated pixel units in a convolution kernel window within multiple consecutive cycles. In each cycle, S804 weights the pulse signal received via broadcast according to the pre-stored weight value for each pixel unit, and accumulates the weighting result through a reconfigurable counter to obtain the convolution result corresponding to each pixel unit after all cycles are completed. S805 constructs and outputs feature maps based on the convolution results of each pixel unit.
[0065] Specifically, in practical work, the above image processing method can be automatically executed by the image sensor through the following steps: 1. Global synchronous exposure and photoelectric conversion: The global control unit sends a global exposure control signal, and the photodiodes in all pixel units start integrating photogenerated charge synchronously.
[0066] 2. Signal Conversion and Quantization: After exposure, each pixel converts its integrated charge into a pixel voltage (VFD), which is then converted into an Ipixel through a voltage-to-current conversion circuit. The Ipixel drives a ring oscillator to generate PFM pulses. Within a fixed time interval, the counter of each pixel (initialized to a specific state, such as the sign bit being 1) counts its own PFM pulses, completing the analog-to-digital conversion and obtaining the original pixel digital value.
[0067] 3. Weight Configuration: Assign the weight values of the required convolutional kernel (e.g., 3x3 size) to each pixel unit. The weights of the same convolutional kernel are distributed in the array according to their position relative to the center.
[0068] 4. Pipeline-style convolution calculation: a) The global control unit starts the first shift clock cycle Φ0. In this cycle, each pixel unit inputs the PFM pulses it continuously generates into its own reconfigurable counter, and performs weighted accumulation according to its stored weights (corresponding to the center weights of the convolution kernel).
[0069] b) After the Φ0 cycle ends, the Φ1 cycle begins. In this cycle, each pixel unit broadcasts its own PFM pulse to a specific neighboring pixel (e.g., the right pixel). Each pixel unit then receives the PFM pulses broadcast from its neighbor in a specific direction (e.g., the left pixel) and inputs them into its own counter, which is then weighted and accumulated according to its stored weights (corresponding to the weights in that neighboring direction in the convolution kernel).
[0070] c) Similarly, the global control unit activates cycles Φ2 to Φ8 in a predefined order, with each cycle completing the pulse broadcast and weighted accumulation in a specific direction (or position).
[0071] 5. Feature Map Generation and Readout: After all nine cycles are completed, the value in the counter of each pixel unit is its convolution result. These results are read out sequentially through row and column address decoding.
[0072] 6. Post-processing: The read data stream passes through the output processing unit and undergoes ReLU and 2x2 max pooling operations in sequence.
[0073] 7. Output: Output the final processed feature map data. If the system requires the original image, the original digital values of each pixel can be read directly after step 2.
[0074] The beneficial technical effects of this application are as follows: by employing a pixel-level analog-to-digital converter based on voltage-to-frequency conversion and a reconfigurable counter, combined with global pipelined broadcast control, ultra-high energy efficiency fully parallel inductive convolution computation is achieved. This not only significantly reduces the overall power consumption and area of quantization and computation, supporting high frame rate global shutter imaging, but also forms a complete front-end processing link from optical signal to compressed feature map by integrating activation and pooling functions in the readout link. Furthermore, it supports dual-mode operation of raw data and feature map output, enhancing the system's flexibility and practicality.
[0075] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The terms "upper," "lower," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this application and for simplification, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Unless otherwise expressly specified and limited, the terms "installation," "connection," and "linkage" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0076] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the embodiments in this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0077] This application uses specific embodiments to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An image sensor with in-sensory calculation function, characterized in that, include: A pixel array contains multiple pixel units arranged in an array. Global control unit; Each of the pixel units includes: The sensing and conversion module is used to generate pixel current based on incident light; The processing module is connected to the sensing and conversion module; The global control unit is configured to: provide a global exposure control signal to all pixel units, enabling the sensing and conversion module to synchronously sense light and generate pixel current; and generate a plurality of shift clock signals in a predetermined order. The processing module is configured to: convert the pixel current into a pulse signal; and, under the control of the plurality of shift clock signals, sequentially receive pulse signals from pixel units associated at different positions within a convolution kernel window, perform weighted summation on the received pulse signals according to pre-stored weight values, and output the convolution operation result corresponding to the pixel unit.
2. The image sensor according to claim 1, characterized in that, The processing module includes: An analog-to-digital converter unit is used to convert the pixel current into the pulse signal; A calculation unit, connected to the analog-to-digital conversion unit, is used to perform the weighted accumulation operation.
3. The image sensor according to claim 2, characterized in that, The analog-to-digital conversion unit is a voltage-to-frequency converter, and the oscillation frequency of the voltage-to-frequency converter is controlled by the pixel current.
4. The image sensor according to claim 3, characterized in that, The voltage-frequency converter includes a ring oscillator, wherein the power supply terminal of at least one inverter in the ring oscillator is powered by the pixel current.
5. The image sensor according to claim 2, characterized in that, The computing unit includes a reconfigurable counter; The reconfigurable counter is configured, based on pre-stored weight values, to perform enable control, counting direction control, and / or counting weight control on the input pulse signal.
6. The image sensor according to claim 5, characterized in that, The reconfigurable counter performs at least one of the following based on the weight configuration signal: Based on the first configuration bit, the path of the pulse signal to the counter is turned on or off; According to the second configuration bit, the pulse signal is guided to different input terminals of the counter to achieve different effective counting weights; The third configuration bit controls the counter to increment or decrement.
7. The image sensor according to claim 1, characterized in that, The sensing and conversion module includes: A photoelectric conversion circuit is used to generate pixel voltage under the global exposure control signal; A voltage-to-current conversion circuit, connected to the photoelectric conversion circuit, is used to linearly convert the pixel voltage into the pixel current.
8. The image sensor according to claim 7, characterized in that, The voltage-to-current conversion circuit includes a cascaded current mirror composed of transistors with different threshold voltages.
9. The image sensor according to claim 1, characterized in that, It also includes an output processing unit connected to the pixel array, used to perform nonlinear activation processing and pooling downsampling processing on the convolution operation results.
10. The image sensor according to claim 9, characterized in that, The output processing unit includes a cascaded linear rectifier circuit and a max-pooling circuit.
11. An image processing method applicable to an image sensor with in-sensory computing function as described in any one of claims 1 to 10, characterized in that, The method includes: By applying a global exposure control signal to all pixel units through the global control unit, all pixel units are controlled to synchronously perform photoelectric conversion to obtain the pixel voltage corresponding to each pixel unit; The pixel voltage is converted into a pulse signal corresponding to the light intensity of each pixel unit through a voltage-to-current conversion circuit and a voltage-to-frequency converter. The global control unit generates multiple shift clock signals in a predetermined order to control the pulse signal of each pixel unit to be broadcast sequentially to different associated pixel units in a convolution kernel window in multiple consecutive cycles. Within each cycle, each pixel unit weights the pulse signal received via broadcast according to the pre-stored weight values, and accumulates the weighting results through a reconfigurable counter to obtain the convolution result corresponding to each pixel unit after all cycles are completed. The feature map is constructed and output based on the convolution results of each pixel unit.