Heterogeneous high-speed parallel reading circuit for storage and calculation array
By designing a heterogeneous high-speed parallel readout circuit for memory arrays, the existing ADC arrays solve the problems of data readout consistency, area, and power consumption when taking into account data depth and conversion rate, and realize efficient data quantization and activation processing, supporting the high-speed parallel and ultra-low power consumption characteristics of brain-like computing hardware.
Patent Information
- Application Number
- CN202510556647.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-06-24
AI Technical Summary
When existing large-scale ADC arrays take into account data depth and conversion rate, there are problems such as data read consistency, area, and power consumption, which is difficult to effectively support the high-speed parallel and ultra-low power consumption characteristics of brain-like computing hardware.
A heterogeneous high-speed parallel readout circuit for memory arrays is designed. Using the multi-order conductance characteristics of memory arrays, only one integrator, comparator and digital signal processing module is required for each row. It draws on the working principle of dual-slant ADC and realizes the merge and multiplexing of modules through the time dimension quantization principle, supporting data processing in two modes of artificial neural network and pulse neural network.
It realizes efficient data quantization and activation processing, supports ANN and SNN neuron behavior, reduces power consumption and area occupation, improves data read consistency and conversion efficiency, and is suitable for heterogeneous high-speed parallel data processing of memory arrays.
Smart Images

Figure CN120199305A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microelectronic integrated circuits, and particularly relates to a heterogeneous high-speed parallel readout circuit for a memory-computation array. Background Art
[0002] Brain-inspired computing is a new computing theory based on the integration of biological brain science and computer science. Currently, generalized brain-inspired computing chips integrate spiking neural networks (SNNs) simulated by brain inspiration and traditional artificial neural networks (ANNs). In order to achieve the high-speed parallel and ultra-low-power computing characteristics of the biological brain, a memory-computation integrated architecture is proposed. By reducing data movement to reduce power consumption, it can achieve high energy efficiency, which is an effective way to implement brain-inspired computing hardware.
[0003] A truly memory-computation integrated architecture relies on new memory devices (such as memristors and floating-gate FLASH). Its characteristic is that it can store multi-level information in the form of conductance and thus can perform a large number of parallel data operations. Therefore, it is necessary to design a high-speed parallel readout circuit for a memory-computation array architecture based on new memory devices, supporting both ANN and SNN data processing modes. Conventionally, a successive approximation register analog-to-digital converter (SAR ADC) array is used to read out the memory-computation current and then combined with digital modules for data operation. Because such a large-scale ADC array needs to consider both data depth and conversion rate at the same time, there are problems such as data readout consistency, area, and power consumption. Therefore, there is an urgent need for a heterogeneous high-speed parallel readout circuit for a memory-computation array. Summary of the Invention
[0004] The purpose of the present invention is to provide a heterogeneous high-speed parallel readout circuit for a memory-computation array, which can perform data quantization and activation processing on the multiply-accumulate operation current output in parallel by a memory-computation array based on new memory devices, and support both ANN and SNN neuron behaviors, so as to solve the problems of data readout consistency, area, power consumption, etc. that exist when a large-scale ADC array needs to consider both data depth and conversion rate at the same time.
[0005] To solve the above technical problems, the present invention provides a heterogeneous high-speed parallel readout circuit for a memory-computation array, The memory-computation array is a crossbar switch matrix integrating i rows and j columns of memory device units for performing convolution operations, including word lines, bit lines, and source lines; the memory device units are located at the intersection of each word line and bit line, store information in the form of conductance, output multiplication current through the source lines, and the source lines of multiple memory device units are combined to achieve multiply-accumulate operations; The heterogeneous high-speed parallel readout circuit performs voltage clamping on the source line and reads the multiply-accumulate computing current, converts the multiply-accumulate computing current into the time domain for quantization, and at the same time selects to accumulate the integration voltage in multiple time steps, so as to enter the leaky integrate-and-fire (LIF) neuron functional state, and realizes heterogeneous high-speed parallel data readout according to subsequent selections. The heterogeneous high-speed parallel readout circuit includes modules used for data processing in two modes: artificial neural network and spiking neural network. Each row of the heterogeneous high-speed parallel readout circuit integrates an integrator, a comparator, a digital latch unit, an activation operation module, a pulse emission module, and a time-step latch unit; among them, Each row of the heterogeneous high-speed parallel readout circuit shares the integrator and the comparator, and the entire heterogeneous high-speed parallel readout circuit shares a set of synchronous clock management module, a counter, and an external bias.
[0006] In one implementation, in the artificial neural network mode, data processing is completed by the combination of an analog-to-digital conversion module and an activation operation module. The analog-to-digital conversion module includes a shared integrator, a comparator, and a counter; Among them, the negative terminal of the integrator is connected to the source line of each row of the computing-in-memory array to perform voltage clamping and read the computing current. The output terminal of the integrator is connected to the negative terminal of the comparator, and the positive terminal of the comparator is grounded. The control clock is generated by the synchronous clock management module. The counter is responsible for time-domain quantization, starts counting when entering the quantization stage, stops counting when the comparator signal goes high, and the counter outputs a digital code to the activation calculation module for non-linear calculation, and finally outputs to the shift multiplexer.
[0007] In one implementation, in the artificial neural network mode, the working process of the heterogeneous high-speed parallel readout circuit includes: Step 11: The control clock is sent out by the synchronous clock management module, and the integrators of all rows are started; each row contributes different currents through the storage device units storing different weights to form the multiply-accumulate computing current on the source line. During the fixed integration time, the integration voltages are respectively obtained at the output terminals of the integrators. Step 12: The counter starts counting, the word lines of each row are turned off, and the integration voltage generated by the corresponding integrator starts to drop until it is lower than 0V, and the output terminal of the comparator generates a high signal. Each row latches the current count value according to the high signal. Step 13: The digital latch unit of each row inputs the latched value into the activation operation module for non-linear calculation, and outputs in parallel to the shift multiplexer.
[0008] In one implementation, in the spiking neural network mode, data processing is completed by the combination of a shared integrator and comparator, a pulse emission module, and a time-step latch unit; The negative terminal of the integrator is connected to the source line of each row of the memory - computing array for voltage clamping and memory - computing current reading. The output terminal of the integrator is connected to the negative terminal of the comparator. The positive terminal of the comparator is connected to the threshold voltage Vth, and they are commonly connected for each row. The synchronization clock management module generates time steps, and an input and integration are performed once per clock cycle within a time step. Integration is performed when the clock is high, and input data is switched when the clock is low. Each time the comparator signal is pulled low within a time step, the pulse emission module generates a pulse signal. The pulse signal lasts for one clock cycle and is latched by the time - step latching unit. Until the time step is completed, the data of the time - step latching unit is output to the shift multiplexer.
[0009] In one implementation, under the said spiking neural network, the working process of the heterogeneous high - speed parallel readout circuit includes: Step 21: The synchronization clock management module issues a control clock, and the integrators of all rows are started; integration is performed when each clock level is high. Each row contributes different currents through the memory device units storing different weights to form a multiply - accumulate memory - computing current on the source line, which is accumulated and converted into an integration voltage at the output terminal of the integrator. Integration pauses when each clock level is low, and the input is switched. The speed of each integration is related to the input. Step 22: The integration voltage generated by each row's integrator is connected to the negative terminal of the comparator and compared with the threshold voltage Vth at the positive terminal of the comparator. Until it exceeds the threshold voltage Vth, a pull - down signal is generated at the output terminal of the comparator, and the pulse emission module generates a pulse signal. Step 23: After the pulse signal appears, the integrator is reset. Step 24: Until the end of the time step, the time - step latching unit outputs the latched data.
[0010] In one implementation, the memory device unit is a non - volatile memory unit, and the non - volatile memory unit is one of floating - gate FLASH, RRAM, and CBRAM, which realizes multi - level conductance storage.
[0011] In one implementation, the parameters of the traditional artificial neural network and the spiking neural network are stored in the non - volatile memory units in the memory - computing array. The input data type of the memory - computing array is voltage, and the output data type is current. Convolution operations are realized in the analog domain through Kirchhoff's laws.
[0012] A heterogeneous high-speed parallel readout circuit for a memory and computing array provided by the present invention draws on the working principle of a dual-slope ADC and utilizes the characteristics of the multi-order conductance of the memory and computing array. Only one integrator, one comparator, and some digital signal processing modules are required for each row, and the structure is simple and suitable for array implementation. It avoids the problem that a large area needs to be occupied by a CDAC (capacitive digital-to-analog converter) in the SAR ADC structure, and the layout can achieve a relatively low height to match the memory and computing array. Through the time dimension quantization principle, module merging and multiplexing are realized, achieving data processing in two modes and increasing the flexibility of the network architecture while maintaining simplicity. In addition, all comparators adopt a unified bias and share a set of counters, with high data conversion consistency, and can obtain the quantization result of the memory and computing current or the result of the LIF rule operation at high speed in parallel. Description of the Drawings
[0013] Figure 1 It is a schematic diagram of the circuit flow of the n-bit memory and computing current readout activation processing circuit in the ANN mode of the present invention.
[0014] Figure 2 It is a schematic diagram of the principle of the n-bit memory and computing current readout activation processing circuit in the ANN mode of the present invention.
[0015] Figure 3 It is a schematic diagram of the circuit flow of the LIF function circuit with 3 bits as one time step in the SNN mode of the present invention.
[0016] Figure 4 It is a schematic diagram of the principle of the LIF function circuit with 3 bits as one time step in the SNN mode of the present invention.
[0017] Figure 5 It is an overall architecture diagram of the heterogeneous high-speed parallel readout circuit for the memory and computing array of the present invention. Detailed Embodiments
[0018] The following further elaborates on a heterogeneous high-speed parallel readout circuit for a memory and computing array proposed by the present invention in conjunction with the drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the drawings are all in a very simplified form and use non-precise scales, only for conveniently and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0019] To achieve the above object, first, the characteristics of the hardware architecture of the memory and computing array are described.
[0020] Both traditional artificial neural networks (ANNs) and spiking neural networks (SNNs) involve a large number of convolutional operations, which are completed by the memory-compute array. Each memory-compute unit stores weights through conductance programming, and the weight deployment of the entire memory-compute array is completed through a programmer and a driving bias. Each row of the memory-compute array can perform convolutional operations in parallel, and the operation data of this part is output in the form of analog current.
[0021] The memory-compute array is a crossbar (cross-switch matrix) array integrating new storage device units with i rows and j columns, which includes: Word Lines (WL): Used to select a specific row, usually a voltage is applied during programming or reading operations; Bit Lines (BL): Used to transmit input signals or read current signals, usually related to input data; Source Lines (SL): Used to provide a reference voltage or establish a current path; New storage device units: Located at the intersection of each WL and BL, store information in the form of conductance, output multiplication current through the source line, and the source lines of multiple new storage device units are combined to achieve MAC operations (Multiply-Accumulate).
[0022] As Figure 5 shown, the heterogeneous high-speed parallel readout circuit includes modules used for data processing in both artificial neural network and spiking neural network modes. Each row of the heterogeneous high-speed parallel readout circuit integrates an integrator, a comparator, a digital latch unit, an activation operation module, a pulse emission module, and a time step latch unit; among them, each row of the heterogeneous high-speed parallel readout circuit includes all modules used for data processing in both modes, which are uniformly operated by a synchronous clock management module and share a set of external bias voltages and counters.
[0023] The heterogeneous high-speed parallel readout circuit performs voltage clamping on the SL and reads the memory-compute current. The data processing modes for the two networks, ANN and SNN, are described separately below.
[0024] There are various non-linear activation functions introduced in traditional artificial neural networks (ANNs), such as Sigmoid, ReLU, Tanh, etc. Therefore, the data processing in this mode is completed by the combination of an analog-to-digital conversion module and an activation operation module. The analog-to-digital conversion module includes a shared integrator, a comparator, and a counter, and the activation calculation module is implemented by digital logic.
[0025] Among them, the negative terminal of the shared integrator is connected to the SL of each row of the memory-compute array for voltage clamping and reading the memory-compute current; The output terminal of the integrator is connected to the negative terminal of the comparator, and the positive terminal of the comparator is grounded; The control clock is generated by the synchronous clock management module; The counter is responsible for time-domain quantization. It starts counting when entering the quantization stage and stops counting when the comparator signal goes high. The counter outputs digital codes to the activation calculation module for non-linear calculation, and finally outputs to the shift multiplexer.
[0026] The specific working process includes the following steps: Step 1: The control clock is sent by the synchronous clock management module, and all row integrators are started. Each row contributes different currents through a new type of memory device storing different weight values to form a multiply-accumulate storage and calculation current in the SL. During the fixed integration time, the integration voltages are obtained at the output ends of the integrators respectively; Step 2: The counter starts counting. Each row of WL is turned off, and the integration voltage generated by the corresponding integrator starts to drop until it is lower than 0V. The comparator output generates a high signal, and each row latches the current count value according to the high signal; Step 3: The digital latch unit of each row inputs the latched value into the activation operation module for non-linear calculation and outputs in parallel to the shift multiplexer.
[0027] The spiking neural network (SNN) inspired by the brain usually uses LIF neurons for non-linear operations. The operation process is the accumulation of signals, the discharge after reaching the threshold, and the emission of pulse signals. This mode is measured by time steps, so there are multiple signal inputs and multiple integrations. The whole process of data processing is completed by the combination of a shared integrator and comparator, a pulse emission module, and a time-step latch unit.
[0028] Among them, the negative terminal of the shared integrator is connected to the SL of each row of the storage and calculation array for voltage clamping and reading the storage and calculation current; the output terminal of the integrator is connected to the negative terminal of the comparator, and the positive terminal of the comparator is connected to the threshold voltage Vth and is shared by each row; the synchronous clock management module generates time steps, and one input and integration are performed in each clock cycle within the time step; integration is performed when the clock level is high, and the input data is switched when the clock level is low; each time the comparator signal is pulled low within the time step, the pulse emission module generates a pulse signal; the pulse signal lasts for one clock cycle and is latched by the time-step latch unit; until the time step is completed, the data of the time-step latch unit is output to the shift multiplexer.
[0029] The specific working process includes the following steps: Step 1: The control clock is sent by the synchronous clock management module, and all row integrators are started. Integration is performed when each clock level is high. Each row contributes different currents through a new type of memory device storing different weight values to form a multiply-accumulate storage and calculation current in the SL, and it is accumulated and converted into an integration voltage at the output terminal of the integrator. Integration is paused when each clock level is low, and the input is switched. The speed of each integration is related to the input; Step 2: The integration voltage generated by each row integrator is connected to the negative terminal of the comparator and compared with the threshold voltage Vth at the positive terminal of the comparator. Until the threshold voltage Vth is exceeded, a pull-down signal is generated at the output terminal of the comparator, and the pulse emission module generates a pulse signal. Step 3: After the pulse signal appears, reset the integrator. Step 4: Until the end of the time step, the time step latch unit outputs the latched data. Embodiment
[0030] As Figure 1 shown, the embodiment of the present invention provides a single-row data processing flow of a heterogeneous high-speed parallel readout circuit for a memory and computing array in the ANN mode, and the depth of the memory and computing current data is n-bit.
[0031] In a specific implementation, the negative terminal of the integrator is connected to the SL output of the memory and computing array, the positive terminal is grounded, the voltage at the SL terminal is clamped to the ground through a capacitive feedback path, and the memory and computing current is integrated at the output terminal to generate an integration voltage. Each new memory device unit stores different weight values. BL is the input data path, and the input voltage and the conductance G are multiplied to form the memory and computing current. The memory and computing currents generated by multiplexed multiplication converge at SL to form a multiply-accumulate memory and computing current, thereby realizing vector convolution operation. The synchronous clock management module generates a fixed integration time T0. The multiply-accumulate current generated in one row generates an integration voltage at the capacitor and the integration output terminal. The integration voltage gradually rises with the integration time, and the rising speed is linearly related to the magnitude of the multiply-accumulate memory and computing current. After the fixed integration time T0 ends, the integration voltage V O is reached. Control WL to turn off row readout, start the counter to count, the standard cells in the memory and computing array perform a fixed-size current outflow, the integration voltage starts to drop, and the time to drop to 0V is linearly related to the magnitude of V O .
[0032] When V O drops below 0V, the output signal of the comparator is pulled high, and the counter ends counting.
[0033] The digital code output by the counter is the quantization result of the memory and computing current. The digital code is output to the activation operation module to complete the non-linear operation, and finally output to the shift multiplexer.
[0034] As Figure 2 shown, the specific working principle of the circuit in the embodiment of the present invention for reading and quantifying the memory and computing current by nbit in the ANN mode is as follows: Looking at the memory and computing current generated by the row of the memory and computing array from the SL terminal, there is a maximum value and a minimum value. At the fixed integration duration T0, the corresponding integration voltage values reach the maximum and the minimum. When WL is turned off, only the standard cells contribute to the fixed current path, and the integrated voltage generated at T0 decreases at the same slope. The minimum integrated voltage corresponds to the shortest decrease time T, and the maximum integrated voltage corresponds to the maximum decrease time T. min These are used as quantization boundaries, that is, the states where the digital code corresponds to the outputs of "0" and "255". max As the quantization boundaries, namely the states where the digital code corresponds to the outputs of "0" and "255". The counter counts during the process of the integrator's integrated voltage decrease and stops counting when the integrated voltage drops below 0V. The corresponding quantization result is latched and output to the activation operation module for non-linear calculation. During the fixed integration time T0, the integration capacitor C is a constant value, and the integrated voltage generated at the output of the integrator and the row multiply-accumulate storage current I are linearly correlated. MAC Linearly correlated: After entering the quantization stage, the time for the integrated voltage to drop to 0V is linearly correlated with the magnitude of the integrated voltage generated during the fixed integration T0. Referring to Figure 2 , the quantization values corresponding to V1 and V2 are respectively: Where is the synchronous clock period, and count1 and count2 are the latched count codes.
[0035] In this embodiment, the equivalent input resistance of the storage and computing array is utilized to be programmable as the input for integrator quantization. By referring to the dual-slope ADC principle, the storage current is converted into the time domain for quantization operation. The circuit structure is simple and supports multiplexing, which is suitable for the large-scale parallel data output processing of the storage and computing array. Embodiment
[0036] As Figure 3 shown, the embodiment of the present invention provides a heterogeneous high-speed parallel readout circuit for a storage and computing array for single-row data LIF function processing flow in the SNN mode, with a single time step having a 3-bit width.
[0037] In specific implementation, the negative terminal of the integrator is connected to the SL output of the storage and computing array, and the positive terminal is grounded. The voltage at the SL terminal is clamped to the ground through a capacitive feedback path, and an integrated voltage is generated at the output. Each new memory device unit stores different weight values, and BL is the input data path. The clock signal is generated by the synchronous clock management module. Within one time step, under each synchronous clock period, integration is performed when the clock level is high, the input data is switched when the clock level is low, and at the same time, the integrated voltage is compared with the threshold voltage V th . If the comparator pulls low, it indicates that the integrated voltage exceeds the threshold voltage V th, the pulse emission module generates a pulse signal and controls the integrator to reset. If the comparator remains high, the input data is switched and the next integration is continued until 8 clock cycles are completed; When the integrated voltage exceeding the threshold voltage V is not detected th , the integrated voltage maintains the value of the previous time period, and the speed change of each integration is related to the input data; After the counter finishes counting a time step period, the pulse information data latched by the time step latch unit is output to the shift multiplexer.
[0038] As Figure 4 shown, the specific working principle of the circuit in the LIF operation in the SNN mode in the embodiment of the present invention is as follows: Referring to Figure 4 , the function of the LIF neuron has three states, namely the accumulation of the membrane potential, the threshold comparison pulse emission, and the reset. The accumulation of the membrane potential can be represented by the integrator forming an integrated voltage. The threshold comparison can be completed by the comparator. The pulse emission is realized by the pulse emission module, and finally the reset is realized by controlling the integrator reset switch.
[0039] The memory and computing array can be regarded as a synaptic array connected by LIF neurons in this mode, and the membrane potential accumulation and pulse emission are carried out in the way of memory and computing current transmission.
[0040] The time step depth needs to be specified in this mode, usually defined by the algorithm. The time step width in this example is 3it, that is, the number of integrations per single time step is 8 times.
[0041] For a determined array conductance, threshold voltage, and integration capacitance within one time step, the slope of each integration is directly related to the input data.
[0042] In this process, the integrated membrane voltage accumulated at the output end of the integrator in the nth integration period is equal to the previous integrated voltage plus the integrated voltage generated this time : is compared with V input to the positive terminal of the comparator th . Here, V th is the threshold voltage, and the pulse emission condition is obtained as: When the output signal of the comparator is pulled low, the pulse emission module will control the integrator to reset. After the reset, the output end of the integrator will be set to 0V, and the next integration will start from 0V.
[0043] The time-step latch unit records the pulse emission situation of each clock cycle in a time step and outputs after the time step ends.
[0044] In this embodiment, a memory-computation array is used as a synaptic array, and the time-step concept is adopted for integration, threshold comparison, and reset, realizing the function of a LIF neuron, multiplexing the integrator and comparator modules, simplifying the data readout processing circuit in two network modes based on a new memory device array, and realizing heterogeneous and efficient data processing.
[0045] The above description is only a description of the preferred embodiments of the present invention and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art of the present invention based on the above disclosure shall fall within the protection scope of the claims.
Claims
1. A heterogeneous high-speed parallel readout circuit for a storage and computing array, characterized in that: The storage and calculation array is a cross switch matrix integrating i rows and j columns of storage device units, which is used to complete the convolution operation, including word lines, bit lines and source lines; the storage device unit is located at each intersection of the word line and the bit line, stores information in the form of conductance, outputs multiplication current through the source line, and the source lines of multiple storage device units are combined to realize the multiplication and accumulation operation; The heterogeneous high-speed parallel readout circuit clamps the voltage of the source line and reads the storage current, converts the multiplication-accumulation storage current into the time domain for quantization, and simultaneously selects to accumulate the integral voltage in multiple time steps, thereby entering the leaky integrate-and-fire LIF neuron functional state, and realizing heterogeneous high-speed parallel data readout according to subsequent selections; The heterogeneous high-speed parallel readout circuit includes modules used for data processing in two modes: artificial neural network and pulse neural network. Each row of the heterogeneous high-speed parallel readout circuit integrates an integrator, a comparator, a digital latch unit, an activation operation module, a pulse emission module, and a time step latch unit; wherein, Each row of heterogeneous high-speed parallel readout circuits shares an integrator and a comparator, and the entire heterogeneous high-speed parallel readout circuit shares a set of synchronous clock management modules, counters and external biases.
2. The heterogeneous high-speed parallel readout circuit for a storage and computing array according to claim 1, characterized in that: In the artificial neural network mode, data processing is completed by a combination of an analog-to-digital conversion module and an activation operation module, and the analog-to-digital conversion module includes a shared integrator, a comparator, and a counter; The negative end of the integrator is connected to the source line of each row of the storage and calculation array to perform voltage clamping and storage and calculation current reading; The output terminal of the integrator is connected to the negative terminal of the comparator, and the positive terminal of the comparator is grounded; The control clock is generated by the synchronous clock management module; The counter is responsible for time domain quantization. It starts counting when entering the quantization stage and stops counting when the comparator signal is pulled high. The counter outputs the digital code to the activation calculation module for nonlinear calculation and finally outputs it to the shift multiplexer.
3. The heterogeneous high-speed parallel readout circuit for a storage and computing array according to claim 2, characterized in that: In the artificial neural network mode, the workflow of the heterogeneous high-speed parallel readout circuit includes: Step 11: The synchronous clock management module sends a control clock, and the integrators of all rows are started; each row contributes different currents to form a multiplication and accumulation storage current on the source line through storage device units storing different weights, and within a fixed integration time, the output ends of the integrators respectively obtain integrated voltages; Step 12: The counter starts counting, each row of word lines is closed, and the integral voltage generated by the corresponding integrator starts to decrease until it is lower than 0V. The comparator output generates a pull-up signal, and each row latches the current count value according to the pull-up signal; Step 13: The digital latch unit of each row inputs the latch value into the activation operation module for nonlinear calculation and outputs it in parallel to the shift multiplexer.
4. The heterogeneous high-speed parallel readout circuit for a storage and computing array according to claim 1, characterized in that: Under the spiking neural network, data processing is completed by a combination of a shared integrator and comparator, a pulse transmitting module, and a time step latch unit; The negative end of the integrator is connected to the source line of each row of the storage and calculation array to perform voltage clamping and storage and calculation current reading; The output terminal of the integrator is connected to the negative terminal of the comparator, the positive terminal of the comparator is connected to the threshold voltage Vth, and each row is connected in common; The time step is generated by the synchronous clock management module, and input and integration are performed once in each clock cycle within the time step; Integration occurs when the clock is high, and input data switching occurs when the clock is low; Each time the comparator signal is pulled low within a time step, the pulse emission module generates a pulse signal; The pulse signal maintains one clock cycle and is latched by the time step latch unit; Until the time step is completed, the data of the time step latch unit is output to the shift multiplexer.
5. The heterogeneous high-speed parallel readout circuit for a storage and computing array according to claim 4, characterized in that: Under the spiking neural network, the workflow of the heterogeneous high-speed parallel readout circuit includes: Step 21: The synchronous clock management module sends a control clock, and the integrators of all rows are started; when each clock level is high, integration is performed, and each row contributes different currents through storage device units storing different weights to form multiplication and accumulation storage currents on the source line, which are accumulated and converted into integrated voltages at the output end of the integrator. When each clock level is low, integration is paused, and the input is switched. The speed of each integration is related to the input; Step 22: The integrated voltage generated by each row integrator is connected to the negative end of the comparator and compared with the threshold voltage Vth of the positive end of the comparator until it exceeds the threshold voltage Vth, the comparator output generates a pull-down signal, and the pulse transmission module generates a pulse signal; Step 23: Reset the integrator after the pulse signal appears; Step 24: Until the time step ends, the time step latch unit outputs the latched data.
6. The heterogeneous high-speed parallel readout circuit for a storage and computing array according to claim 1, characterized in that: The storage device unit is a non-volatile storage unit, which is a floating gate type FLASH, RRAM, or CBRAM, and realizes multi-level conductance storage.
7. The heterogeneous high-speed parallel readout circuit for a storage and computing array according to claim 6, characterized in that: The parameters of the traditional artificial neural network and the pulse neural network are stored by non-volatile storage units in a storage and computing array. The input data type of the storage and computing array is voltage, and the output data type is current. Convolution operations are implemented in the analog domain through Kirchhoff's law.
Citation Information
Cited By
Neuron simulation system and method based on LTspice and fused with time
CN120745720A
Storage and calculation integrated device, storage and calculation method, processing device, tile module and accelerator
CN120874919A