A highly flexible storage computing array

By changing the word line direction in the 1T1R array perpendicular to the input, combining the row and column selector and linear voltage regulator, the problems of high power consumption and calculation delay of traditional in-memory computing arrays are solved, and efficient vector matrix multiplication calculation is realized.

CN115831170BActive Publication Date: 2025-07-18PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211396106.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-09
Publication Date
2025-07-18
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

Traditional non-volatile memory arrays have problems such as high power consumption, calculation delay, and physical matrix scale do not match the actual calculation needs when performing in-memory computing, especially when input voltages using digital-to-analog converters or buffers.

Method used

The 1T1R array is adopted, and the word line direction is perpendicular to the input direction. The array row and column are selectively activated through row multiplexer and column multiplexer. Combined with a linear regulator, it provides multi-valued voltage, reduces the use of digital-to-analog converters, and realizes flexible sub-array calculations.

Benefits of technology

Reduces power consumption, reduces computing delay, improves computing efficiency, and allows flexible selection of sub-arrays for calculation in large-scale arrays to avoid waste of power consumption of unused units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115831170B_ABST
    Figure CN115831170B_ABST
Patent Text Reader

Abstract

The present invention proposes a highly flexible storage computing array, belonging to the technical field of semiconductor non-volatile memories and in-memory computing. In the 1T1R array of the present invention, the word line direction is perpendicular to the input direction. The word line driver is used to control the word line input power supply voltage or ground voltage of the 1T1R array to turn on or off a column of word lines. The input unit is provided with an input register, a voltage multiplexer, and a row multiplexer. The voltage multiplexer selects one of the multiple voltages generated by the linear regulator as the input according to the value of the input register. The row multiplexer is connected to the source line of the 1T1R array. The bit lines of the 1T1R array are connected to the clamping circuit and the analog-to-digital converter through the column multiplexer in the output unit. By using the highly flexible storage computing array provided by the present invention, unnecessary power consumption can be saved, multi-value input can be achieved without an analog-to-digital converter, the computing speed is improved, and the number of times the array is turned on and the resulting power consumption are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of non-volatile memory and compute-in-memory in semiconductors and CMOS very large scale integrated circuits (VLSIs), and particularly relates to an array structure for performing vector matrix multiplication using a non-volatile memory array. Background Art

[0002] With the development of artificial intelligence and deep learning technologies, artificial neural networks have been widely applied in fields such as natural language processing, image recognition, autonomous driving, and graph neural networks. However, the gradually increasing network scale has led to a large amount of energy consumption in the transfer of data between memory and traditional computing devices such as CPUs and GPUs, which is known as the von Neumann bottleneck. The most dominant calculation in artificial neural network algorithms is vector matrix multiplication. Compute-in-memory based on non-volatile memory stores weights in non-volatile memory cells and performs analog vector matrix multiplication in the array, avoiding frequent data transfer between memory and computing units, and is considered a promising approach to solve the von Neumann bottleneck.

[0003] Figure 1 It is a schematic diagram for performing vector matrix multiplication based on a non-volatile device array. After writing the weights into non-volatile memory devices such as RRAM, PCRAM, MRAM, etc., the weights are stored in the conductance values of the devices. The devices are organized in an array form. A voltage is input from one end as the input of the vector matrix multiplication. In the array, calculations are performed through Ohm's law and Kirchhoff's law, and the current obtained at the other end of the array is the summation result of the vector matrix multiplication. The device cells in the array can use 1R devices or 1T1R devices. The input can be a multi-valued voltage input through a digital-to-analog converter (DAC) or a binary voltage input through a buffer. The summation result is usually read out using an analog-to-digital converter (ADC). Due to the mismatch in length between the area of the analog-to-digital converter and the area of the array cells, a multiplexer (MUX) is usually used to allow multiple columns in the array to share one analog-to-digital converter.

[0004] Since 1T1R avoids the problem of write crosstalk, 1T1R devices are usually used in larger arrays. In the agreed naming method, the line connecting the gate of the transistor is the word line (WL), the line connecting the source of the transistor is the source line (SL), and the line connecting one end of the device is the bit line (BL). The traditional 1T1R array structure used for vector-matrix multiplication is as follows: Figure 2 (a) or (b). Figure 2 In (a), all transistors are turned on by the word line, voltage is input from the source line, and current and are read from the bit line. Figure 2 In (b), the same read voltage is input to the source line, and the word line is used to control the opening and closing of a row to represent the input of "1" or "0", and the current and are still read from the bit line. The common feature of the above two array structures and other commonly used array structures is that the word line WL is parallel to the input direction, which is a design used in traditional memory array structures.

[0005] However, there are two problems with the traditional array structure in realizing in-memory computing: 1. If a digital-to-analog converter is used to input multi-value voltages, there is a mismatch between the area of the digital-to-analog converter and the area of the array unit in terms of layout height, and a digital-to-analog converter must be used for each row, which brings about the problem of high power consumption. If a buffer is used to input binary voltages, high-precision inputs can only be represented by multiple input pulse sequences, which increases the delay of the calculation, and requires multiple array calculations to be turned on for multi-bit inputs, which increases the number of times the array is turned on and the number of times the analog-to-digital converter works, and also increases power consumption. 2. The traditional design in which the word line is parallel to the input direction cannot turn off the current of the devices on the unused columns when the physical matrix scale is larger than the matrix scale required for actual calculations, which leads to a waste of power consumption. Summary of the invention

[0006] In view of the above problems, the present invention provides a highly flexible storage computing array.

[0007] The technical solution provided by the present invention is as follows:

[0008] A highly flexible in-memory computing array, characterized by comprising a 1T1R array. Several rows of the 1T1R array are divided into a row segment, and several columns are divided into a column segment. Each row segment corresponds to an input unit, and each column segment corresponds to an output unit. Each 1T1R unit consists of a MOS transistor and a non-volatile memory device. The gate of the MOS transistor is connected to the word line, the source of the MOS transistor is connected to the source line, the drain of the MOS transistor is connected to one end of the non-volatile memory device, one end of the non-volatile memory device is connected to the drain of the MOS transistor, and the other end is connected to the bit line. One source line connects the sources of all the MOS transistors in a row of the array, parallel to the input direction; one bit line connects the non-volatile devices of all the units in a column of the array, perpendicular to the input direction; one word line connects the gates of all the MOS transistors in a column of the array, perpendicular to the input direction. The peripheral circuit of the 1T1R array includes a word line driver, an input unit, an output unit, a linear voltage regulator, and a control module, which are used to select a corresponding row and a corresponding column, float the inputs of the unselected rows, and input a ground voltage to the word lines corresponding to the unselected columns to turn off the transistors. If the 1T1R array is divided into m row segments, corresponding to m input units in the layout design, and at the same time the 1T1R array is divided into n column segments, corresponding to n output units in the layout design, each row segment contains B rows of memory units, and each column segment contains C columns of memory units, then a physical array is divided into B*C logical arrays. The matrix size of each logical array is m rows * n columns. Each time a vector-matrix multiplication is performed, one of the logical arrays is selected for calculation, and no power consumption is generated on the unselected memory units. In the logical array, any row and any column are selected to form a logical sub-array for calculation.

[0009] Furthermore, the word line driver is used to control the word line of the 1T1R array to input a power supply voltage or a ground voltage to turn on or off a column of word lines.

[0010] Furthermore, the input unit includes an input register, a voltage multiplexer, and a row multiplexer. The input register consists of (a + 1) D flip-flops. Under the control of a clock signal, the (a + 1)-bit scan chain input is sequentially scanned into the input register of the upper input unit from the bottommost input unit; the first a outputs of the input register are the decoding signals of the voltage multiplexer, connected to the a-A decoder in the voltage multiplexer, and the (a + 1)-th output is the enable signal of the row multiplexer, connected to the AND gate of the row multiplexer. The voltage multiplexer contains an a-A decoder and A transmission gates; the a-A decoder converts the a decoding signals output by the input register into an A-bit one-hot code output, that is, only one bit of the A-bit output is high level, and the rest of the bits are low level, to turn on one of the transmission gates. The quantitative relationship is A = 2 a; The voltage multiplexer selects one of the A voltages output by the LDO and sends it to the row multiplexer. The row multiplexer includes a b-B decoder, B AND gates, and B transmission gates; the b-B decoder converts the b-bit row decoding signal into a B-bit one-hot code output, that is, only one of the B-bit outputs is high level, and the rest are low level, and is connected to the B two-input AND gates. The quantitative relationship is B = 2 b ; The other input of the B two-input AND gates is the row multiplexer enable signal output by the input register. When the enable signal is high level, the output of the AND gate is the same as the output of the b-B decoder, and one of the transmission gates is selected to be opened, and one of the B row source lines is selected to be connected to the output of the voltage multiplexer, and the rest of the source lines are in a floating state. When the enable signal is low level, the outputs of all AND gates are low, all transmission gates are closed, and all B row source lines are in a floating state.

[0011] Further, the output unit structure includes a column multiplexer, a clamping circuit, and an analog-to-digital converter. The column multiplexer includes a c-C decoder and C transmission gates. The c-C decoder converts the c-bit column decoding signal into a C-bit one-hot code output, that is, only one of the C-bit outputs is high level, and the rest are low level. The quantitative relationship is C = 2 c ; The output of the decoder selects to open one of the transmission gates, and one of the C column bit lines is selected to be connected to the clamping circuit, and the rest of the bit lines are in a floating state; the clamping circuit uses an operational amplifier OP and a feedback resistor R f to clamp the selected bit line at the reference potential. At this time, the output voltage of the operational amplifier is equal to the result of the vector matrix multiplication calculation. The output voltage is stored on C1 by the sample-and-hold circuit composed of transistor N1 and capacitor C1, and finally input to the analog-to-digital converter to be converted into an x-bit digital signal. x is the design accuracy of the analog-to-digital converter, and the analog-to-digital converter adopts a Flash-ADC or SAR-DAC structure.

[0012] Further, the input unit uses a linear voltage regulator to uniformly generate multiple input voltages required for the calculation. The voltage multiplexer is used to input the voltage to the row multiplexer. The row multiplexer is connected to the source lines of the 1T1R array, and the bit lines of the 1T1R array are connected to an analog-to-digital converter through the column multiplexer.

[0013] The present invention mainly includes three characteristics compared with the traditional array structure.

[0014] The first feature is to change the word line direction from parallel to the input to perpendicular to the input. This change brings two benefits: one is that by keeping the unselected row inputs floating, a row can be shielded so that it does not affect the calculation of the rest of the array. The other is that by applying a low voltage to the word line input, the transistors can be turned off to achieve the effect of shielding a column so that it does not affect the calculation of the rest of the array. Therefore, the proposed design can shield any number of rows and columns in the array, and any sub-array in a large array can be selected for calculation without current flowing through the shielded devices, resulting in increased power consumption.

[0015] The second feature is the addition of a row multiplexer in the input unit. The row multiplexer selects one row from multiple rows of the array to be connected to the voltage multiplexer. The reason for this design is that even though using a voltage multiplexer instead of a digital-to-analog converter can reduce the layout height mismatch problem, multiple row heights are still required to match one voltage multiplexer.

[0016] The third feature is to replace the digital-to-analog converter or buffer in the input unit with a shared linear voltage regulator. The linear voltage regulator uniformly generates multiple input voltages required for calculation and inputs them to all input units simultaneously. In the input unit, one of the voltages is selected through the voltage multiplexer and connected to the array. This reduces the layout mismatch problem caused by using a digital-to-analog converter for each row input and the high power consumption of the digital-to-analog converter. At the same time, the benefits brought by multi-valued voltage input can be utilized, and there is no need to achieve multi-bit input by turning on the array multiple times in the way of a buffer inputting binary voltages, reducing the calculation delay and the number of operations of the analog-to-digital converter. Description of the Drawings

[0017] Figure 1 Schematic diagram for matrix multiplication based on a non-volatile device array;

[0018] Figure 2 Schematic diagrams of two traditional 1T1R array structures for vector-matrix multiplication;

[0019] Figure 3 Schematic diagram of the array structure and peripheral circuit proposed by the present invention;

[0020] Figure 4 Circuit diagram of the input unit proposed by the present invention;

[0021] Figure 5 Circuit diagram of the output unit used by the present invention;

[0022] Figure 6 Schematic diagram of the working state of the array structure selection logic array proposed by the present invention;

[0023] Figure 7 Schematic diagram of the working state of the array structure selection logic sub-array proposed by the present invention. Detailed implementation mode

[0024] The present invention will be described in detail below with reference to the accompanying drawings and specific examples.

[0025] Refer to Figure 3 , the high-flexibility in-memory computing array and its peripheral circuit design of the present invention include a 1T1R array, word line driver, input unit, output unit and control module. The 1T1R array is composed of 1T1R units, and each 1T1R unit is composed of a MOS transistor and a non-volatile memory device. The gate of the MOS transistor is connected to the word line, the source of the MOS transistor is connected to the source line, the drain of the MOS transistor is connected to one end of the non-volatile memory device, one end of the non-volatile memory device is connected to the drain of the MOS transistor, and the other end is connected to the bit line. One source line connects the sources of the MOS transistors of all the units in one row of the array and is parallel to the input direction; one bit line connects the non-volatile devices of all the units in one column of the array and is perpendicular to the input direction; one word line connects the gates of the MOS transistors of all the units in one column of the array and is perpendicular to the input direction. A word line driver is provided to control the input power supply voltage or ground voltage of the word lines of the 1T1R array to turn on or off a column of word lines. In the layout, several rows of the 1T1R array are divided into a row segment, and several columns are divided into a column segment. Each row segment corresponds to an input unit, and each column segment corresponds to an output unit. As shown in the layout on the right, the 1T1R array is divided into m row segments, corresponding to m input units in the layout design. At the same time, the 1T1R array is divided into n column segments, corresponding to n output units in the layout design.

[0026] Use the input unit to replace the DAC to achieve the multi-value input function. The input unit circuit is as Figure 4 shown, and internally includes an input register, a voltage multiplexer and a row multiplexer. The input register is composed of (a + 1) D flip-flops. Under the control of the clock signal, the (a + 1)-bit scan chain input is sequentially scanned into the input register of the upper input unit from the bottommost input unit. The first a outputs of the input register are the decoding signals of the voltage multiplexer and are connected to the a-A decoder in the voltage multiplexer. The (a + 1)-th output is the enable signal of the row multiplexer and is connected to the AND gate of the row multiplexer. The voltage multiplexer contains an a-A decoder and A transmission gates. The a-A decoder converts the a decoding signals output by the input register into A-bit one-hot code output, that is, only one bit of the A-bit output is high level, and the rest of the bits are low level, to turn on one of the transmission gates. The quantitative relationship is A = 2 a。The voltage multiplexer selects one of the A voltages output by the LDO and sends it to the row multiplexer. The row multiplexer includes a b-B decoder, B AND gates, and B transmission gates. The b-B decoder converts the b-bit row decoding signal into a B-bit one-hot code output, that is, only one of the B-bit outputs is high level, and the rest are low level, and is connected to B two-input AND gates. The quantitative relationship is B = 2 b 。The other input of the B two-input AND gates is the row multiplexer enable signal output by the input register. When the enable signal is high level, the output of the AND gate is the same as the output of the b-B decoder, and one of the transmission gates is selected to be opened, and one of the B row source lines is selected to be connected to the output of the voltage multiplexer, and the rest of the source lines are in a floating state. When the enable signal is low level, the outputs of all AND gates are low, all transmission gates are closed, and all B row source lines are in a floating state.

[0027] The output unit circuit is as Figure 5 shown. The output unit structure includes a column multiplexer, a clamping circuit, and an analog-to-digital converter. The column multiplexer includes a c-C decoder and C transmission gates. The c-C decoder converts the c-bit column decoding signal into a C-bit one-hot code output, that is, only one of the C-bit outputs is high level, and the rest are low level. The quantitative relationship is C = 2 c 。The output of the decoder selects one of the transmission gates to be opened, and one of the C column bit lines is selected to be connected to the clamping circuit, and the rest of the bit lines are in a floating state. The clamping circuit uses an operational amplifier OP and a feedback resistor R f to clamp the selected bit line at the reference potential. At this time, the output voltage of the operational amplifier is equal to the result of the vector matrix multiplication calculation. This output voltage is stored on C1 by the sampling and holding circuit composed of transistor N1 and capacitor C1, and finally input into the analog-to-digital converter to be converted into an x-bit digital signal. x is the design accuracy of the analog-to-digital converter. The analog-to-digital converter can use a Flash-ADC or SAR-DAC structure.

[0028] One row multiplexer only selects one corresponding row of source lines, and one column multiplexer only selects one corresponding column of bit lines. The source lines in the unselected rows and the bit lines in the unselected columns are all kept floating. The row multiplexer adds the function of floating all corresponding source lines compared with the column multiplexer to achieve a flexible sub-array selection function. Suppose a physical array contains m row segments and n column segments, each row segment contains B rows of storage units, and each column segment contains C columns of storage units. Then a physical array can be divided into B*C logical arrays, and the matrix size of each logical array is m rows * n columns. Each time the vector matrix multiplication is performed, one of the logical arrays can be selected for calculation, and no power consumption is generated on the unselected storage units. In the logical array, any row and any column can also be selected to form a logical sub-array for calculation, and no power consumption is generated on all unselected storage units.

[0029] Figure 6 This is a schematic diagram of the working state of the logic array selected for the array structure proposed by the present invention. Figure 6 Taking the 1T1R physical array with a size of 6 rows * 6 columns as an example, the physical array is divided into three row segments, each row segment includes two rows of devices, the physical array is divided into three column segments, and each column segment includes two columns of devices. Each row segment corresponds to an input unit, and each column segment corresponds to an output unit. In a vector matrix multiplication calculation, the row multiplexer in an input unit can only select one row in one row segment, and the column multiplexer of an output unit can only select one column in one column segment. The devices on all the selected rows and columns form the array used for this vector matrix multiplication calculation, which is called the logical array (Logical Array). In contrast, the entire array is called the physical array (Physical Array). The physical array is divided into 4 logical arrays, and each logical array includes 3 rows * 3 columns of storage units. Each vector matrix multiplication calculation can only select one of the logical arrays for calculation. As Figure 6 shown in, when the row decoding signal is 0, all row decoders select rows 1, 3, and 5 to form an array, and when the column decoding signal is 0, all column decoders select columns 1, 3, and 5 to form an array. The selected 3 rows * 3 columns of devices form the logical array for this calculation, as shown in red in the figure. When the column decoders select columns 1, 3, and 5, the word line driver needs to input the word line voltages of columns 1, 3, and 5 into the power supply voltage (VDD) at the same time, and input the remaining word lines into the ground voltage (GND). The unselected cells can be divided into two categories: ① cells in unselected columns; ② cells in unselected rows on selected columns. It can be analyzed that ① in unselected columns, since the gates of all transistors are grounded, the path between the source line and the bit line of the 1T1R device is closed, and no current flows through. ② In unselected rows on selected columns, the transistor gates are connected to the power supply voltage, the transistors are turned on, the bit lines are clamped to the ground level, and the source lines are floating, so no current flows through either. This shows that each time a vector matrix multiplication is executed, one of the logical arrays can be selected for calculation, and no power consumption is generated on the unselected storage cells. In the input module, the input data is stored in the input register of each input unit through the scan chain, and a voltage multiplexer decoding signal is generated to select one of the multiple input voltages generated by the linear regulator and send it to the row multiplexer to replace the function of the digital-to-analog converter. In the output module, the bit line current of the selected column is converted into the voltage on the capacitor through the clamping circuit and finally read out as a digital signal by the ADC.

[0030] Figure 7 This is a schematic diagram of selecting a part of the logical sub-array (Logical Sub Array) for calculation in the logical array. In the logical array, any part of rows and columns can be selected to form a logical sub-array for vector matrix multiplication calculation. As Figure 7Taking the example shown in [figure], the row corresponding to the row multiplexer 3 is selected to be turned off, and the column corresponding to the column multiplexer 3 is selected to be turned off. The rows and columns corresponding to the row multiplexers 1 and 2 and the column multiplexers 1 and 2 form a logical sub-array with a size of 2 rows * 2 columns. The method is to set the row multiplexer enable signal in the input unit 3 low through the scan chain to turn off all AND gates, floating all the source lines corresponding to the input unit 3. At the same time, the word line driver sets the word lines of all the memory cells corresponding to the output unit 3 to the ground level to turn off all the columns corresponding to the output unit 3. As shown in the figure, the unselected cells can be divided into two categories: ① cells in the unselected column; ② cells in the unselected row on the selected column. It can be analyzed that: ① In the unselected column, since the gates of all transistors are grounded, the path between the source line and the bit line of the 1T1R device is turned off and no current flows through. ② In the unselected row on the selected column, the gate of the transistor is connected to the power supply voltage, the transistor is turned on, the bit line is clamped to the ground level, and the source line is floating, so no current flows through either. This shows that in the logic array, any part of the rows and columns can be arbitrarily selected to form a logical sub-array for vector matrix multiplication calculation, and no power consumption is generated on the unselected memory cells. For the devices in the selected column and row, the calculation method is the same as Figure 6 the same.

[0031] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.

Claims

1. A highly flexible in-memory computing array, characterized in that, It includes a 1T1R array. Several rows of the 1T1R array are divided into a row segment, and several columns are divided into a column segment. Each row segment corresponds to an input unit, and each column segment corresponds to an output unit. Each 1T1R unit consists of a MOS transistor and a non-volatile memory device. The gate of the MOS transistor is connected to the word line, the source of the MOS transistor is connected to the source line, the drain of the MOS transistor is connected to one end of the non-volatile memory device, one end of the non-volatile memory device is connected to the drain of the MOS transistor, and the other end is connected to the bit line. One source line connects the sources of all the MOS transistors in one row of the array and is parallel to the input direction; One bit line connects the non-volatile devices in one column of the array and is perpendicular to the input direction; One word line connects the gates of all the MOS transistors in one column of the array and is perpendicular to the input direction. The peripheral circuit of the 1T1R array includes a word line driver, an input unit, an output unit, a linear voltage regulator, and a control module, which are used to select a corresponding row and a corresponding column, float the inputs of the unselected rows, and input a ground voltage to the word lines corresponding to the unselected columns to turn off the transistors. If the 1T1R array is divided into m row segments, it corresponds to m input units in the layout design. At the same time, the 1T1R array is divided into n column segments, which corresponds to n output units in the layout design. Each row segment contains B rows of memory cells, and each column segment contains C columns of memory cells. Then a physical array is divided into B*C logical arrays. The matrix size of each logical array is m rows * n columns. Each time a vector-matrix multiplication is performed, one of the logical arrays is selected for calculation, and no power consumption is generated on the unselected memory cells. In the logical array, any row and any column are selected to form a logical sub-array for calculation.

2. The high-flexibility in-memory computing array according to claim 1, wherein The word line driver is used to control the word lines of the 1T1R array to input a power supply voltage or a ground voltage to turn on or off a column of word lines.

3. The highly flexible in-memory computing array according to claim 1, wherein The input unit includes an input register, a voltage multiplexer, and a row multiplexer. The input register consists of (a + 1) D flip-flops. Under the control of a clock signal, the (a + 1)-bit scan chain input is sequentially scanned into the input register of the upper input unit from the bottommost input unit. The first a outputs of the input register are decoding signals for the voltage multiplexer, which are connected to the a-A decoder in the voltage multiplexer. The (a + 1)-th output is the enable signal for the row multiplexer, which is connected to the AND gate in the row multiplexer. The voltage multiplexer contains an a-A decoder and A transmission gates. The a-A decoder converts the a decoding signals output by the input register into an A-bit one-hot code output, that is, only one bit of the A-bit output is high level, and the rest are low level, to open one of the transmission gates. The quantitative relationship is A = 2 a ; The voltage multiplexer selects one of the A voltages output by the LDO and sends it to the row multiplexer. The row multiplexer contains a b-B decoder, B AND gates, and B transmission gates. The b-B decoder converts the b-bit row decoding signal into a B-bit one-hot code output, that is, only one bit of the B-bit output is high level, and the rest are low level, and is connected to B two-input AND gates. The quantitative relationship is B = 2 b ; The other input of the B two-input AND gates is the enable signal for the row multiplexer output by the input register. When the enable signal is high level, the output of the AND gate is the same as the output of the b-B decoder, and one of the transmission gates is selected to be opened, and one of the B row source lines is selected to be connected to the output of the voltage multiplexer, and the rest of the source lines are in a floating state. When the enable signal is low level, the outputs of all AND gates are low, all transmission gates are closed, and all B row source lines are in a floating state.

4. The highly flexible in-memory computing array according to claim 1, wherein The output unit structure includes a column multiplexer, a clamping circuit, and an analog-to-digital converter. The column multiplexer includes a c-C decoder and C transmission gates. The c-C decoder converts c column decoding signals into C one-hot code outputs, that is, only one of the C outputs is high level, and the rest are low level. The quantitative relationship is C = 2 c ; The output of the decoder selects to open one of the transmission gates to select one of the C column bit lines to connect to the clamping circuit, and the rest of the bit lines are in a floating state; The clamping circuit uses an operational amplifier OP and a feedback resistor R f to clamp the selected bit line at the reference potential. At this time, the output voltage of the operational amplifier is equal to the result of the vector matrix multiplication calculation. The output voltage is stored on the capacitor C1 by the sample-and-hold circuit composed of the transistor N1 and the capacitor C1, and finally input to the analog-to-digital converter to be converted into an x-bit digital signal, where x is the design accuracy of the analog-to-digital converter.

5. The highly flexible in-memory computing array according to claim 4, wherein The analog-to-digital converter adopts a Flash-ADC or SAR-DAC structure.

6. The highly flexible in-memory computing array according to claim 3, wherein The input unit uses a linear voltage regulator to uniformly generate various input voltages required for calculation. The voltage multiplexer is used to input the voltage to the row multiplexer. The row multiplexer is connected to the source line of the 1T1R array. The bit line of the 1T1R array is connected to an analog-to-digital converter through the column multiplexer.

Citation Information

Patent Citations

  • Computing array based on 1T1R device, operation circuit of computing array based on IT1R device and operation method thereof

    CN108111162A

  • In-memory computing circuit suitable for full-connection binary neural network

    CN110414677A