Heterogeneous high-speed parallel reading circuit optimization structure for storage and calculation array
By introducing positive and negative weight determination units and up-down counters into the readout circuit of the arithmetic array, the structure is optimized to eliminate the consistency problem of the arrayed circuit, and the problem of complex structure, slow speed and low accuracy of the readout circuit of the arithmetic array is solved, and high-precision, high-speed and consistent data reading is achieved.
Patent Information
- Application Number
- CN202510556650.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-06-27
AI Technical Summary
The readout circuit structure of the computing array is complex, the area is very different from the computing array, the layout is difficult, the power consumption is high and the speed is slow, and the accuracy is difficult to improve due to the analog operation noise, resulting in the effect of the integrated computing technology in improving energy consumption, area and speed.
A heterogeneous high-speed parallel readout circuit optimization structure is designed, a positive and negative weight determination unit is introduced, and the subtraction operation of the analog domain is realized through positive and reverse current superposition. The up-down counter and TDC are used for structure optimization, eliminating the consistency problem of array circuits, and improving the readout speed through coarse and fine quantization.
It realizes high accuracy, high speed and consistency improvement of data reading of memory arrays, saves computing resources, expands the data quantization range and accuracy, and improves the readout speed of the circuit.
Smart Images

Figure CN120220768A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microelectronic integrated circuits, and particularly relates to an optimized structure of a heterogeneous high-speed parallel readout circuit for a memory-computation array. Background Art
[0002] The memory-computation integrated technology stores multi-bit information based on the conductance programmability feature of non-volatile devices, and performs large-scale multiply-accumulate operations by expanding the crossbar (cross-switch matrix) array structure. This array is called a memory-computation array.
[0003] The computing mode of the memory-computation array conforms to the convolution algorithm, which is the most commonly used and resource-consuming algorithm in neural networks. By means of neuromorphic, analog / digital hybrid computing, the power consumption loss and time delay caused by the back-and-forth data transfer are avoided, which is an effective technical way to implement a more powerful edge intelligent chip.
[0004] Due to the high-density integration of non-volatile devices in the memory-computation array, the memory-computation array often occupies a very small area and power consumption. However, the driving and readout circuits for the memory-computation array face problems such as complex circuit structures, difficult layout with a large disparity in area from the memory-computation array, high power consumption, slow speed, and difficulty in improving the accuracy due to the influence of analog operation noise. As a result, the improvement of the memory-computation integrated technology in terms of energy consumption, area, and speed is not obvious. Therefore, it is very necessary to design a peripheral circuit adapted to the high-density integrated memory-computation array. Summary of the Invention
[0005] The purpose of the present invention is to provide an optimized structure of a heterogeneous high-speed parallel readout circuit for a memory-computation array, so as to solve the problem of actual data readout errors caused by the process mismatch of the readout circuit array, improve data accuracy, speed, and consistency, and achieve efficient positive and negative weight operations and the inhibitory forgetting process of LIF neurons.
[0006] To solve the above technical problems, the present invention provides an optimized structure of a heterogeneous high-speed parallel readout circuit for a memory-computation array, including: The memory-computation array is a cross-switch matrix integrating i rows and j columns of storage device units, including word lines, bit lines, and source lines; the storage device units are located at the intersection of each word line and bit line, store information in the form of conductance, output multiplication current through the source lines, and the source lines of multiple storage device units are combined to achieve multiply-accumulate operations; The optimized structure of the heterogeneous high-speed parallel readout circuit includes modules used for data processing in both the artificial neural network and spiking neural network modes. Each row of the heterogeneous high-speed parallel readout circuit integrates an integrator, a comparator, a digital latch unit, an activation operation module, a pulse emission module, a time step latch, a data merging unit, and a time-to-digital converter; wherein, The optimized structure of the per-row heterogeneous high-speed parallel readout circuit shares an integrator and a comparator, and the entire optimized structure of the heterogeneous high-speed parallel readout circuit shares a set of synchronous clock management modules, up-down counters, and external biases; The optimized structure of the heterogeneous high-speed parallel readout circuit introduces the characteristic of membrane potential leakage generated over time in the LIF neuron function, and determines the reset state according to the membrane potential overflow threshold intensity, matching the working mechanism of the brain-inspired computing neuron; A bypass reference resistance state std cell is provided in the last column of each row in the memory and computing array for integration and quantization in a fixed time period, and then integrated and quantized with the result of the memory and computing array. An up-down counter is introduced for differential input counting, and the counting result is the difference between the standard quantization results; the optimized structure of the heterogeneous high-speed parallel readout circuit performs coarse and fine quantization through a time-to-digital converter, adopts the rule of outputting the first-converted first, and introduces asynchronous arbitration output.
[0007] In one implementation, a positive and negative weight determination unit is introduced at the intersection of each word line and bit line. The bit line input potential is input as Vsl±△V through the positive and negative weight determination unit. If the positive weight is input as Vsl - △V, the stored computing current contributed flows out from the source line and the integration voltage rises; if the negative weight is input as Vsl+△V, the stored computing current contributed flows into the source line and the integration voltage drops; the positive and negative weight data operations are realized in the analog domain through the superposition of the positive and negative currents on the source line.
[0008] In one implementation, in the artificial neural network mode, data processing is completed by the combination of an analog-to-digital conversion module, an activation operation module, a data merging unit, a digital latch unit, and a time-to-digital converter. The analog-to-digital conversion module includes a shared integrator, a comparator, and an up-down counter; Step 1: According to the stored weights in the memory and computing array and the positive and negative determination unit, after inputting data through the bit line, there are two directions of the stored computing current generated on the source line. One is flowing out from the source line, and the other is flowing into the source line. A bypass reference resistance state std cell contributes an offset to keep the input range of the stored computing current within the range of flowing out from the source line; Step 2: Turn off the word line, and the synchronous clock management module only performs integration and data quantization in a fixed clock period through the bypass reference resistance state std cell as a reference standard; Step 3: The up-down counter is started. During standard quantization, it performs Down counting. After standard quantization is completed, the word line is turned on, and the stored-computation result quantity is input for integration for a standard time. This integration result includes the contribution of the bypass reference resistance state std cell and the true stored-computation result. After the integration ends, it performs Up counting. According to the positive or negative of the true stored-computation result, the time digital converter is started. After quantization is completed, the values of the up-down counter and the time digital converter are input to the data merging unit for digital merging, and finally latched into the digital latch unit. Step 4: The digital latch unit inputs the latched value into the activation operation module for non-linear calculation, and finally outputs it to the asynchronous axon module.
[0009] In one implementation manner, under the said spiking neural network, data processing is completed by the combination of a pulse emission module, a time step latch, a shared integrator, a comparator, and an up-down counter. Step 1: When the membrane potential accumulates, the bypass reference resistance state std cell is in the off state. Since the storage array stores weights and the positive / negative determination unit, after data is input through the bit line, the membrane potential accumulation voltage generated will rise or fall. A decrease is the inhibitory behavior of neuron learning. Step 2: When the membrane potential accumulates, the bypass reference resistance state std cell is in the off state, while the bypass reference resistance state std cell is turned on during the membrane potential leakage process, and the storage array is turned off through the word line. Step 3: The integration voltage generated by each row of integrators is connected to the negative terminal of the comparator and compared with the threshold voltage V th at the positive terminal of the comparator until it exceeds the threshold voltage V th . A pull-down signal is generated at the output terminal of the comparator, the pulse emission module generates a pulse signal, and the overflow threshold voltage V overflow is retained. Step 4: After the pulse signal appears, the integrator is reset, and the storage array is turned off through the word line. Only the bypass reference resistance state std cell is used for integration leakage. The difference between the membrane potential threshold voltage V th and the initial voltage V ini is fixed. Therefore, the integration leakage time period through the bypass reference resistance state std cell is also fixed. After the time period arrives, the bypass reference resistance state std cell is turned off, and the reset ends. The membrane potential stays at the overflow threshold voltage V overflow . Step 5: The pulse state during the reset process is not counted into the time step. After each reset, the integration step is repeated until the time step ends. The time step latch outputs the latched data to the asynchronous axon module.
[0010] An optimized structure of a heterogeneous high-speed parallel readout circuit for a memory-computation array provided by the present invention adds positive and negative determination units in the memory-computation array, thereby introducing positive and negative integration directions, realizing subtraction operations in the analog domain, saving computing resources, and expanding the data quantization range and accuracy. Based on the heterogeneous high-speed parallel readout circuit for the memory-computation array, a structure optimization is carried out by using an up-down counter and a TDC, eliminating the consistency problems caused by the process in the parallel array circuit, and improving the readout speed of the circuit through a coarse and fine quantization method. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a schematic diagram of the process of the 4-bit memory-computation current readout activation processing circuit in the ANN mode in the embodiment of the present invention.
[0012] Figure 2 It is a schematic diagram of the principle of the 4-bit memory-computation current readout activation processing circuit in the ANN mode in the embodiment of the present invention.
[0013] Figure 3 It is a schematic diagram of the process of the LIF functional circuit with 3 bits as one time step in the SNN mode in the embodiment of the present invention.
[0014] Figure 4 It is a schematic diagram of the principle of the LIF functional circuit with 3 bits as one time step in the SNN mode in the embodiment of the present invention.
[0015] Figure 5 It is a schematic diagram of the overall architecture of the optimized memory-computation array and heterogeneous high-speed parallel readout circuit in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The following further details an optimized structure of a heterogeneous high-speed parallel readout circuit for a memory-computation array proposed by the present invention in conjunction with the accompanying drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the drawings are all in a very simplified form and use non-precise scales, only for conveniently and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0017] The present invention provides an optimized structure of a heterogeneous high-speed parallel readout circuit for a memory-computation array. By introducing a bypass reference resistance state std unit on the memory-computation array and combining it with an up-down counter for differential data input, the data offset caused by the process mismatch of the readout circuit array is eliminated.
[0018] The memory-computation array generally uses positive and negative weight storage, which has the advantage of expanding the memory-computation input range, thereby improving the quantization accuracy and accuracy. The present invention designs a readout circuit for the positive and negative weight memory-computation unit, realizes subtraction operations in the analog domain, saves computing resources, and thus realizes efficient positive and negative weight data operations.
[0019] Since the circuit structure adopts the principle of time-domain quantization, the higher the precision of the data to be processed, the slower the speed. For example, the quantization of 8-bit data requires 256 Tclk to complete. The present invention introduces a TDC (Time-to-Digital Converter) in the circuit structure for coarse and fine quantization, greatly shortening the quantization time, flexibly improving the data precision and processing speed. At the same time, the rule of outputting the data that is converted first is adopted, and an asynchronous arbitration output is introduced to improve the fluidity of large-scale parallel data.
[0020] The present invention also introduces the characteristic of membrane potential leakage generated over time in the LIF neuron function, and determines the reset state according to the threshold intensity of membrane potential overflow, which better matches the working mechanism of brain-like computing neurons.
[0021] The optimized architecture of the heterogeneous high-speed parallel data readout circuit is described below: The crossbar memory and computing array structure integrating the new memory device units of i rows and j columns includes: Word Lines (WL): Used to select a specific row, usually applying a voltage during programming or reading operations; Bit Lines (BL): Used to transmit input signals or read current signals, usually related to input data; Source Lines (SL): Used to provide a reference voltage or establish a current path; New memory device units: Located at the intersection of each WL and BL, storing information in the form of conductance, outputting multiplication current through the source line, and the source lines of multiple new memory device units are combined to achieve MAC (Multiply Accumulate) operations.
[0022] Since conductance cannot represent positive or negative, a positive and negative weight determination unit is introduced at each intersection, with the symbol ±. The BL input potential is input as Vsl±△V through the positive and negative weight determination unit. For example, for positive weight input Vsl-△V, the stored computing current contributed flows out from the SL, and the integral voltage rises; for negative weight input Vsl+△V, the stored computing current contributed flows into the SL, and the integral voltage drops; the positive and negative weights of data are calculated in the analog domain through the superposition of positive and negative currents on the SL.
[0023] There is a bypass reference resistance state std unit at the last column of each row in the memory and computing array, which is used for the fixed current path during the integrator quantization process and the integral leakage in the LIF function.
[0024] The optimized structure of the heterogeneous high-speed parallel data readout circuit includes: an integrator, a comparator, an up-down counter, an activation operation module, a synchronous clock management module, a data merging unit, a pulse emission module, a digital latch unit, a time step latch, an external bias voltage, and a TDC. By converting the input data into time-domain information, data processing of the ANN / SNN neural network type can be completed.
[0025] The data is input into the memory and computing array. According to the conductance stored in the unit and the positive and negative judgment unit, a positive and negative weight superposition vector multiply-accumulate memory and computing current is formed. For the specific working process of the circuit in the optimized ANN (traditional artificial neural network) mode, it includes the following steps: Step 1: According to the memory weights in the memory and computing array and the positive and negative judgment unit, after inputting data through the BL, there are two directions of the memory and computing current generated on the SL. One is flowing out of the SL, and the other is flowing into the SL. By bypassing the reference resistance state std unit to contribute an offset, the input range of the memory and computing current is kept flowing out of the SL. Step 2: Turn off the WL. The synchronous clock management module only performs integration and data quantization for a fixed clock cycle through the bypass reference resistance state std unit as a reference standard. Step 3: The up-down counter starts and performs Down counting during standard quantization. After completing the standard quantization, turn on the WL and input the memory and computing result quantity for integration for the standard time. This integration result includes both the contribution of the bypass reference resistance state std unit and the real memory and computing result. After the integration ends, perform Up counting, select to start the TDC according to the positive and negative of the real memory and computing result, input the values of the up-down counter and the TDC into the data merging unit for digital merging after quantization is completed, and finally latch them into the digital latch unit. Step 4: The digital latch unit inputs the latched value into the activation operation module for non-linear calculation, and finally outputs it to the asynchronous axon module.
[0026] For the specific working process of the circuit in the optimized SNN (spiking neural network) mode, it includes the following steps: Step 1: The initial voltage of the membrane potential is V ini , when the membrane potential accumulates, the bypass reference resistance state std unit is in the off state. Due to the memory weights in the memory and computing array and the positive and negative judgment unit, after inputting data through the BL, the accumulated voltage of the membrane potential will rise or fall. The fall is the inhibitory behavior of neuron learning. Step 2: When the membrane potential accumulates, the bypass reference resistance state std unit is in the off state, while the bypass reference resistance state std unit is turned on during the membrane potential leakage process, and the memory and computing array is turned off through the WL. Step 3: The integration voltage generated by each row integrator is connected to the negative terminal of the comparator and compared with the threshold voltage V th at the positive terminal of the comparator until it exceeds the threshold voltage V th . A pull-down signal is generated at the output terminal of the comparator, and the pulse emission module generates a pulse signal, and the overflow threshold voltage V overflow is retained; Step 4: After the pulse signal appears, the integrator is reset, and the memory and computing array is turned off through WL. Only the bypass reference resistive state std cell is used for integration leakage. Since V th -V ini is fixed, the integration leakage time period through the bypass reference resistive state std cell is also fixed. After the time period arrives, the bypass reference resistive state std cell is turned off, and the reset ends. The membrane potential stays at the overflow threshold voltage V overflow ; Step 5: The pulse state during the reset process is not counted into the time step. After each reset, the integration step is repeated until the time step ends. The time step latch outputs the latched data to the asynchonous axon module.
[0027] As Figure 1 shown, the embodiment of the present invention provides a heterogeneous high-speed parallel readout circuit for a memory and computing array, and a single-row data processing flow in the ANN mode. The memory and computing current data depth is 4-bit.
[0028] In a specific implementation, the negative terminal of the integrator is connected to the SL output of the memory and computing array, and the positive terminal is grounded. The voltage at the SL terminal is clamped to the ground through a capacitive feedback path; First, all array units are turned off. The V ref voltage of the bypass reference resistive state std cell is switched for fixed-period input standard integration and quantization. This process only needs to be performed once after the entire circuit is started, and is used to calibrate the conversion accuracy error of the analog circuit. During the standard quantization process, the up-down counter starts reverse counting; When performing analog convolution operations, the bypass reference resistive state std cell and the array units are turned on. BL is the input data path. The input voltage Vbl and the conductance G are multiplied to form the memory and computing current. The memory and computing currents generated by multiplexed multiplication converge at SL to form the multiply-accumulate memory and computing current; The multiply-accumulate current generates an integration voltage at the output terminal of the capacitor and the integrator. The integration voltage gradually rises with the integration time, and the rising speed is linearly related to the magnitude of the multiply-accumulate memory and computing current. The integration period remains the same as that in the standard quantization process; After the period ends, WL is controlled to turn off the array units, and the up-down counter switches to forward counting. Depending on the specific weight storage situation of the array units, the memory and computing current is positive or negative, and the quantization result is higher or lower than the standard quantization value.
[0029] When the integrated voltage drops below 0V, the comparator signal is pulled high. At the same time, the TDC is started, and start_dealy begins to generate delay clocks. At the standard data quantization point, the TDC generates a stop_dealy delay clock at the fourth rising edge of the clk during data quantization. As Figure 2 shown, the data quantization time T1 is subjected to 2-bit fine quantization. The TDC data and the data latched by the up-down counter are input to the data merging unit for data merging to form a 4-bit quantization result, which is output to the activation operation module for non-linear calculation. During the fixed integration time T0, the integration capacitor C is a constant value. The integrated voltage generated at the output end of the integrator and the row multiply-accumulate storage current I MAC are linearly correlated:
[0030] There are two cases during the merging process. When the MAC storage current is negative, the integrated voltage will be less than the standard integration. As Figure 2 shown by the data quantization time T1, at this time the TDC fine quantization is: the integrated voltage drops to a value between 0V and the standard quantization, plus the up-down count value as the quantization result; when the MAC storage current is positive, the integrated voltage will be greater than the standard integration. As Figure 2 shown by the data quantization time T2, at this time the TDC fine quantization is: the up-down counter counts from the rising edge of the clock cycle before the integrated voltage drops to 0V to the value when the integrated voltage drops to 0V, plus the up-down count value as the quantization result.
[0031] In this embodiment, the differential processing of the stored and calculated data is realized by using the calculation principle of the storage and calculation array and the standard quantization of the bypass reference resistance state std unit, eliminating the data offset that may be caused by the process and consistency. The TDC is used to accelerate the data conversion rate, increasing the feasibility and practicality for the engineering implementation of this readout circuit.
[0032] As Figure 3 shown, the embodiment of the present invention provides a heterogeneous high-speed parallel readout circuit for a storage and calculation array, and the single-row data LIF function processing flow in the SNN mode, with a single time step of 3-bit depth.
[0033] In a specific implementation, the negative terminal of the integrator is connected to the SL output of the storage and calculation array, and the positive terminal is connected to V ini , and the voltage at the SL terminal is clamped to the initial membrane potential voltage V ini through a capacitive feedback path. The storage current is formed by multiplying the input voltage Vbl and the conductance G, and an integrated voltage is generated at the output end. A clock signal is generated by the synchronous clock management module. Within one time step, in each synchronous clock cycle, when the clock voltage is high, the bypass reference resistive state std cell is turned off and the array cells are turned on. The input data is multiplied and added with the weight parameters to form a computing-in-memory current for voltage integration. When the clock voltage is low, the bypass reference resistive state std cell is turned on and the array cells are turned off. The input data is switched and the integration voltage leaks. At the same time, the integration voltage is compared with V th If the comparator pulls low, it indicates that the integration voltage exceeds V th A pulse signal is generated by the pulse emission module and the integrator is controlled to reset. If the comparator remains high, the input data is switched and the next integration is continued until 8 clock cycles are completed. After the counter finishes counting one time step cycle, the pulse information data latched by the time step latch is output.
[0034] Refer to Figure 4 Since positive and negative values are introduced by the input computing-in-memory current, the initial integration voltage needs to be raised to V ini When the multiplication and addition result of the input data and the weight parameter is negative, the integration voltage will decrease. At this time, it is the inhibitory state of the neuron membrane potential. After each integration, the bypass reference resistive state std cell leaks the integration voltage through a fixed current path. This is the process of SNN neuron memory decay. When the positive integration reaches V th , the membrane potential value accumulated by the neuron often exceeds V th . At this time, the integration reset is also achieved through the fixed current I std path of the bypass reference resistive state std cell, but the discharge time is a fixed T rst , thus realizing the voltage discharge reset of the V th -V ini difference, retaining the membrane potential overflow value. The pulse data in the fixed reset period is not included in the time step.
[0035]
[0036] In this embodiment, the computing-in-memory array calculation principle and the fixed leakage of the bypass reference resistive state std cell are utilized. The concept of time step is adopted to realize functions such as the accumulation, inhibition, leakage, emission, and retention reset of the quasi-biological neuron membrane potential. The circuit structure optimizes the data readout processing method in two network modes based on the new memory device array, realizing heterogeneous, efficient, and accurate data output.
[0037] The above description is only a description of the preferred embodiments of the present invention, and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art of the present invention based on the above disclosure shall fall within the protection scope of the claims.
Claims
1. An optimized structure of a heterogeneous high-speed parallel readout circuit for a storage and computing array, characterized in that: include: The storage and computing array is a crossbar switch matrix integrating i rows and j columns of storage device units, including word lines, bit lines and source lines; The memory device unit is located at each intersection of a word line and a bit line, stores information in the form of conductance, outputs a multiplication current through a source line, and the source lines of multiple memory device units are combined to realize a multiplication and accumulation operation; The optimized structure of the heterogeneous high-speed parallel readout circuit includes modules used for data processing in two modes: artificial neural network and pulse neural network. Each row of the heterogeneous high-speed parallel readout circuit integrates an integrator, a comparator, a digital latch unit, an activation operation module, a pulse emission module, a time step latch, a data merging unit, and a time digital converter; wherein, Each row of the heterogeneous high-speed parallel readout circuit optimization structure shares an integrator and a comparator, and the entire heterogeneous high-speed parallel readout circuit optimization structure shares a set of synchronous clock management modules, up-down counters and external biases; The optimized structure of the heterogeneous high-speed parallel readout circuit introduces the membrane potential leakage characteristics generated over time in the LIF neuron function, and determines the reset state according to the membrane potential overflow threshold strength, matching the working mechanism of brain-like computing neurons; The last column of each row in the storage and calculation array is provided with a bypass reference resistance state std unit, which performs integration and quantization in a fixed time period, and the storage and calculation array result is added for integration and quantization, and an up-down counter is introduced to perform differential input counting, and the counting result is the difference between the standard quantization results; the heterogeneous high-speed parallel readout circuit optimization structure performs coarse and fine quantization through a time digital converter, adopts the rule of first conversion first output, and introduces asynchronous arbitration output.
2. The optimized structure of heterogeneous high-speed parallel readout circuit for storage and computing array according to claim 1, characterized in that: A positive and negative weight determination unit is introduced at each intersection of a word line and a bit line. The bit line input potential is input as Vsl±△V through the positive and negative weight determination unit. If the positive weight input is Vsl-△V, the contributed storage current flows out from the source line and the integrated voltage rises; if the negative weight input is Vsl+△V, the contributed storage current flows into the source line and the integrated voltage decreases. The positive and negative weight data operations are realized in the analog domain by superimposing the positive and reverse currents on the source line.
3. The optimized structure of heterogeneous high-speed parallel readout circuit for storage and computing array according to claim 2, characterized in that: In the artificial neural network mode, data processing is completed by a combination of an analog-to-digital conversion module, an activation operation module, a data merging unit, a digital latch unit, and a time-to-digital converter. The analog-to-digital conversion module includes a shared integrator and comparator, and an up-down counter; Step 1: According to the storage weights and positive and negative judgment units of the storage calculation array, after inputting data through the bit line, there are two directions of the storage calculation current generated on the source line, one is flowing out from the source line, and the other is flowing in from the source line. By contributing an offset through the bypass reference resistance state std unit, the storage calculation current input range is kept in the direction of flowing out from the source line; Step 2: Turn off the word line, and the synchronous clock management module only performs integration and data quantization of a fixed clock cycle by bypassing the reference resistance state std unit as a reference standard; Step 3: The up-down counter is started, and Down counting is performed during standard quantization. After standard quantization is completed, the word line is opened, and the storage result is input to perform integration of the standard time. This integration result includes the contribution of the bypass reference resistance state std unit and the actual storage result; After the integration is completed, the Up count is performed, and the time digital converter is started according to the positive or negative value of the actual stored calculation result. After the quantization is completed, the values of the up-down counter and the time digital converter are input into the data merging unit for digital merging, and finally latched into the digital latch unit; Step 4: The digital latch unit inputs the latch value into the activation operation module for nonlinear calculation, and finally outputs it to the asynchronous axon module.
4. The optimized structure of heterogeneous high-speed parallel readout circuit for storage and computing array according to claim 2, characterized in that: Under the spiking neural network, data processing is completed by a combination of a pulse emission module, a time step latch, a shared integrator, a comparator, and an up-down counter; Step 1: When the membrane potential accumulates, the bypass reference resistance state std unit is in the off state. Since the storage array stores weights and positive and negative judgment units, after the data is input through the bit line, the generated membrane potential accumulation voltage will rise or fall. The fall is the neuron learning inhibition behavior; Step 2: When the membrane potential is accumulated, the bypass reference resistance state std unit is in the closed state, while during the membrane potential leakage process, the bypass reference resistance state std unit is opened, and the storage and calculation array is closed through the word line; Step 3: The integrated voltage generated by each row integrator is connected to the negative terminal of the comparator and the threshold voltage V th The comparison is performed until the threshold voltage V th , the comparator output generates a pull-down signal, the pulse emission module generates a pulse signal, and retains the overflow threshold voltage V overflow ; Step 4: After the pulse signal appears, the integrator is reset, and the storage array is closed through the word line, and only the integral leakage is performed through the bypass reference resistance state std unit, and the membrane potential threshold voltage V th With the initial voltage V ini The difference is fixed, so the integral leakage time period through the bypass reference resistance state std unit is also fixed. After the time period is reached, the bypass reference resistance state std unit is turned off, the reset ends, and the membrane potential stays at the overflow threshold voltage V overflow ; Step 5: The pulse state during the reset process is not included in the time step. The integration step is repeated after each reset until the time step ends. The time step latch outputs the latched data to the asynchronous axon module.