FeFET-based MLC memory, MAC circuit, CIM circuit and chip
By designing a FeFET-based MLC memory, each cell circuit consists of FeFET transistor, PMOS transistor and NMOS transistor. The voltage level of the storage node Vs is used to characterize the 2-bit data, and the 2-bit multiplication operation is implemented through the encoding of WL1, WL2, and INR, which solves the control complexity and area overhead problems of FeFET-type MLC memory, and realizes efficient data storage and logical operations.
Patent Information
- Application Number
- CN202510663976.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-29
AI Technical Summary
The existing FeFET-type MLC memory has complex control and quantization logic when implementing logical operations, and the peripheral circuit area overhead is large, and the different signal quantization methods lead to different signal quantization methods in memory mode and calculation mode.
A FeFET-based MLC memory is adopted. Each cell circuit is composed of 2 FeFET transistors F1, F2, 1 PMOS transistor M1, 1 NMOS transistor M2, and 1 sampling capacitor C0. The 2-bit data is characterized by four different voltage levels of the storage node Vs. The data pre-stored in the storage node Vs is used as the weight of 2-bit in the calculation mode. WL1, WL2, and INR are jointly encoded as 2-bit inputs to realize the multiplication operation of 2-bit.
It realizes efficient storage of 2-bit data in the unit circuit and performs 2-bit multiplication operations, which reduces reading power consumption, improves calculation accuracy and area utilization, and reduces the area overhead of peripheral circuits.
Smart Images

Figure CN120564801A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of memory, and in particular relates to an MLC memory based on FeFET and its corresponding MAC circuit, CIM circuit and chip. Background Art
[0002] In recent years, the continuous development of technologies such as big data and artificial intelligence has led to an increasing demand for massive data storage and processing. On the one hand, traditional computers are based on the von Neumann architecture, which uses a design that separates computing and storage. When performing calculations, data in the memory needs to be transferred to the computing unit to perform the calculation, and then the calculation results are written back to the memory. The movement of data between the memory and the computing unit leads to increased energy consumption and latency issues. On the other hand, with the expiration of Moore's Law, the storage density of traditional memories based on complementary metal-oxide-semiconductor transistor (CMOS) technology has slowed down.
[0003] The Compute-in-Memory (CIM) architecture overcomes the challenges of the traditional von Neumann architecture by embedding computing units within memory, integrating data storage and computation into a single entity and addressing the power consumption and latency issues associated with data migration. Numerous CIM circuits based on static random access memory (SRAM) have been proposed for analog CIM circuits. However, the SRAM latch structure necessitates complex circuit architectures to implement multi-bit CIM functionality. This resulting large area overhead, coupled with SRAM's inherent static power consumption, has limited its development.
[0004] At the same time, various new non-volatile memories (NVMs) such as resistive random access memory (ReRAM), spin transfer torque based magnetic random access memory (STT-MRAM), and ferroelectric field effect transistors (FeFETs) have attracted attention for their high-density data storage capabilities and extremely low static power consumption. Among them, FeFETs, as voltage-controlled devices, have a structure developed from traditional MOSFETs and have good CMOS process compatibility. In addition, they have the characteristics of fast switching speed, high switching ratio, and low leakage current. They can realize multi-level cell (MLC) storage in CMOS processes and are expected to efficiently perform multiply and accumulate (MAC) operations in deep neural network (DNN) reasoning.
[0005] FeFET can adjust the polarization intensity of HfO2 ferroelectric layer by different gate (G) programming voltage. Fe ) is adjusted to achieve the transistor threshold voltage V th The positive polarization voltage can make the V th Lowering the negative polarization voltage can make the V th Increased. Taking advantage of this circuit characteristic, technicians hope to use FeFET to replace some MOSFETs to develop CIM circuits. Compared with mature MOSFET MLC memories, single-function FeFET MLC memories have problems such as complex peripheral circuits and current-type storage state reading. By taking advantage of the fact that a single FeFET device can store higher-bit information, multiplication and addition operations can be performed based on the MLC read and write functions, and higher-bit operations can be achieved within a single unit, which has a significant advantage over the MOSFET structure. However, the existing FeFET in-memory computing units have different signal quantization methods in memory mode and computing mode, which increases the area overhead of the peripheral circuit. Summary of the Invention
[0006] In order to solve the problems of complex control and quantization logic and large peripheral circuit area overhead when using FeFET-type MLC memory to implement logical operations in the existing technology, the present invention provides a FeFET-based MLC memory and its corresponding MAC circuit, CIM circuit and chip.
[0007] The technical solution provided by the present invention is:
[0008] A FeFET-based MLC memory with data storage and logic operation functions. Each unit circuit consists of two FeFET transistors F1 and F2, a PMOS transistor M1, an NMOS transistor M2, and a sampling capacitor C0. The gates of F1 and F2 are connected to input ports WL1 and WL2, respectively; the gate of M2 is connected to the enable port EN; the drain of F1 is connected to the control port BL; the source of F2 and the drain of M1 are grounded. The sources of F1 and M1 are connected to the drains of F2 and M2, forming the storage node Vs; the gate of M2 is connected to the input port INR; the source of M2 is connected to one end of C0, and the connection point serves as the output port Vout; the other end of C0 is grounded.
[0009] In the MLC storage mode, the value of the stored 2-bit data is represented by four different voltage levels of the storage node Vs.
[0010] In calculation mode, the data pre-stored in the storage node Vs is used as a 2-bit weight, and WL1, WL2, and INR are jointly encoded as a 2-bit input. The voltage of Vout is quantized to obtain the product of the weight and the input.
[0011] As a further improvement of the present invention, in storage mode, the operating logic of the FeFET-based MLC memory to implement data writing is:
[0012] First, set the INR terminal to low level, and set the BL and EN terminals to 0V; then apply different programming voltages to WL1 and WL2 according to the preset coding rules; then after the specified writing time t w After that, the data writing is completed.
[0013] Among them, the preset encoding rules are:
[0014] (1) When data "00" needs to be written, WL1 is set to -V0 and WL2 is set to V1+3ΔV;
[0015] (2) When data "01" needs to be written, WL1 is set to V0 and WL2 is set to V1+2ΔV;
[0016] (3) When data "10" needs to be written, WL1 is set to V0+ΔV and WL2 is set to V1+ΔV;
[0017] (4) When data "11" needs to be written, WL1 is set to V0+2ΔV and WL2 is set to V1;
[0018] Wherein, V0 and V1 are references for the programming voltages of WL1 and WL2 respectively, and ΔV is the gradient change value of each programming voltage.
[0019] As a further improvement of the present invention, in the MLC storage mode, the operation logic of the FeFET-based MLC memory to implement data reading is:
[0020] First set BL to VDD and EN to the linear voltage V M ; Then set WL1 and WL2 to read voltage V R , INR is set to read voltage V RB ; then after the specified reading time t R Then, the corresponding stored data is read according to the voltage value of Vout:
[0021] In the data reading phase, Vout has four levels: 0V, V cap , 2*V cap , 3*V cap , V cap The above four level states respectively indicate that the storage data read from the storage node is “00”, “01”, “10”, and “11”.
[0022] As a further improvement of the present invention, in calculation mode, the operation logic for performing multiplication operation is as follows:
[0023] (1) Weight pre-storage
[0024] In MLC storage mode, a 2-bit weight value is pre-stored in the storage node.
[0025] (2) Input and calculation
[0026] Set BL to VDD and EN to the linear voltage V M Switch to calculation mode; then adjust the voltage values of WL1, WL2, and INR to complete the input of the corresponding input number. The encoding rules of the input number are:
[0027] When WL1=V n0 , when WL2=0V, INR=Random, the characterization input is “00”;
[0028] When WL1=V p1 , WL2=V p , INR=V R1 When , the characterization input is "01";
[0029] When WL1=V p2 , WL2=V p , INR=V R2When , the characterization input is "10";
[0030] When WL1=V p3 , WL2=V p , INR=V R3 When , the characterization input is "11";
[0031] Among them, V n0 <0 and V p1 >V p2 >V p3 >0;0 <V R1 <V R2 <V R3 ; Random means that INR can be set to any voltage value.
[0032] (3) Operation output
[0033] The input and weight are multiplied in the circuit and after the specified sampling time t sampling Then, the corresponding calculation result is read according to the voltage value of Vout:
[0034] In calculation mode, Vout has 7 levels: 0V, V cap , 2*V cap , 3*V cap , 4*V cap , 6*V cap 、9*V cap ; The above 7 level states respectively represent the multiplication results: "0000", "0001", "0010", "0011", "0100", "0110", and "1001".
[0035] The present invention also includes a MAC circuit composed of multiple unit circuits from the aforementioned FeFET-based MLC memory arranged in rows. In the MAC circuit, the sources of M2 in all unit circuits are connected to the same computation bit line SL, and all unit circuits share the same sampling capacitor C0. One end of C0 is connected to SL, and the other end is grounded.
[0036] After the MAC circuit completes the corresponding multiplication operation in each row of unit circuits, it uses the bit line voltage V SL Represents the result of the multiplication and accumulation operations of all rows.
[0037] As a further improvement of the present invention, the process of the MAC circuit performing the multiplication-accumulation operation is as follows:
[0038] (1) Weight pre-storage
[0039] In the MLC storage mode, the weight value of the corresponding multiplication operation part is pre-stored in the storage node of the unit circuit of the selected corresponding row.
[0040] (2) Input and calculation
[0041] Set BL to VDD and EN to V M2 Switch to calculation mode; then adjust the voltage values of WL1, WL2 and INR in the unit circuit of the corresponding row to complete the input of the input number of the multiplication operation part.
[0042] (3) Operation output
[0043] Each unit circuit completes the multiplication operation of its corresponding input and weight, and after the specified sampling time, according to the bit line voltage V SL Read out the corresponding calculation results, where V SL Relative gradient voltage V cap The multiple of is the result of the multiplication and accumulation operation.
[0044] The present invention also includes a CIM circuit, which includes a storage and calculation array and its peripheral circuits.
[0045] The memory and computation array consists of multiple MAC circuits, as described above, arranged in columns. In the memory and computation array, the gates of F1 and F2 in each unit circuit in the same row are connected to the same set of input word lines, denoted as WLI and WL2. The gate of M1 in each unit circuit in the same row is connected to the same input word line, denoted as INR. The drain of F1 in each unit circuit in the same column is connected to the same control bit line, denoted as BL. The gate of M2 in each unit circuit in the same column is connected to the same control bit line, denoted as EN. The computation bit line SL in the same column is used to input the computation results of each column.
[0046] The peripheral circuitry includes wordline drivers, bitline drivers, a switch array, a sense amplifier array, an ADC, and a shift adder. The wordline driver controls the wordline voltages of the input wordlines WLI, WL2, and INR for each row in the memory and calculation array. The bitline driver switches the voltage levels of the control bitlines BL and EN for each column in the memory and calculation array. The calculation bitlines for each column in the memory and calculation array are connected to the sense amplifiers in the corresponding columns of the sense amplifier array via switches in the corresponding columns of the switch array. The output of each column's sense amplifier is connected to the input of the ADC, which in turn is connected to the input of the shift adder.
[0047] In the memory array, each column calculates the voltage signal V output by the bit line to represent the calculation result. SL It is first quantized by a sensitive amplifier and then converted into a digital quantity of the calculation result by ADC processing.
[0048] As a further improvement of the present invention, each unit circuit in the CIM circuit's memory array is used to implement 2-bit data storage and multiplication operations between a 2-bit weight and a 2-bit input. Each column of unit circuits constitutes a MAC circuit and is used to perform multiplication-accumulation operations between a 2-bit weight and a 2-bit input. Multiple columns of MAC circuits are used to perform multiplication-accumulation operations between a 2-bit input and a 2n-bit weight, where n>1.
[0049] The operation logic of the multiplication and accumulation operation between the 2-bit input and the 2n-bit weight is as follows:
[0050] (1) The 2n-bit weight is decomposed into multiple 2-bit weights with weights according to every two bits; and then pre-stored into each unit circuit in different columns.
[0051] (2) Perform multiplication and accumulation operations between each 2-bit input and the 2-bit weights of various weights in different columns of the storage array.
[0052] (3) Each switch in the switch array is closed in sequence; then the numerical value of the multiplication and accumulation result of the corresponding column is generated through the sensitive amplifier and ADC, and finally the digital value of the multiplication and accumulation result of each column is shifted and added according to the preset weight through the shift adder to obtain the final operation result.
[0053] As a further improvement of the present invention, the MAC circuits of multiple columns in the CIM circuit are further configured to perform a multiplication-accumulation operation between a 2-bit weight and a 2n-bit input; n>1; and the operation logic is as follows:
[0054] (1) Decompose the 2n-bit input into U 2-bit inputs with input weights at every two bits;
[0055] (2) In each cycle, a column is selected to perform a multiplication-accumulation operation between one of the 2-bit inputs and the 2-bit weight;
[0056] (4) After each cycle, the corresponding switch of the switch array is closed; the numerical value of the corresponding column multiplication and accumulation result is generated through the sense amplifier and the ADC; and the digital value of the multiplication and accumulation result of the current cycle and the multiplication and accumulation result of the previous cycle are shifted and added according to the preset input weights through the shift adder;
[0057] (4) After U cycles, the shift adder outputs the final operation result.
[0058] The present invention also includes a chip, which is packaged by the aforementioned CIM circuit.
[0059] The technical solution provided by the present invention has the following beneficial effects:
[0060] The present invention designs a new type of MLC memory, which has data storage and logic operation functions. In terms of data storage function, the scheme adopts the coupling method of F1, F2, and M1 to store data, especially using M1 to ensure the accurate storage of stored information. The node voltage signal V is used in the reading phase. S Representing the storage result, the storage state of the cell can be accurately detected within one operation, shortening the clock cycle length and reducing read power consumption.
[0061] In terms of logical operation function, the circuit performs input (IN) in the in-memory calculation process through synchronous input of WL1, WL2, and INR, thereby realizing the input of 2-bit signal and multiplication of 2-bit weight (W). Based on the same principle, the multiplication of higher-bit weight (W) and 2-bit input can be further realized through performance expansion. Among them, PMOS can achieve precise regulation of Vs voltage through the control of INR, and the working scheme can be well promoted in FeFETs with different characteristics. Good linearity and discrimination of different IN and W multiplication results are achieved, which is conducive to multi-row parallel operations in array design.
[0062] This solution uses the quantization method of capacitor voltage to realize the reading process in MLC mode and the accumulation process in calculation mode at the same time. Compared with the threshold voltage V th The measured reading scheme and quantization circuit achieve good reusability and lower area overhead. Different from the existing technical features, the circuit design of the present invention realizes the gate reading voltage V of FeFET when realizing MLC mode reading and calculation mode input. R The gate voltage corresponding to different inputs (IN) is much smaller than the MLC mode storage programming voltage, which also achieves the stability of the cell storage state and the reusability of the cell during in-memory computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 This is a circuit diagram of a basic unit circuit in the FeFET-based MLC memory provided in Example 1 of the present invention.
[0064] Figure 2 This is a circuit diagram of the MAC circuit provided in Example 2 of the present invention.
[0065] Figure 3 This is a circuit diagram of the CIM circuit provided in Example 3 of the present invention.
[0066] Figure 4 This is the circuit diagram of the traditional MLC memory that does not use PMOS transistors in the control group of the test experiment.
[0067] Figure 5This is a comparison diagram of storage node voltage changes of the MLC memory of the present invention and the MLC memory of the control group under the same control signal in the test experiment.
[0068] Figure 6 This is a signal flow diagram of the MLC memory of the present invention during the data reading and writing and data reading stages in the test experiment.
[0069] Figure 7 This is a signal flow diagram when the CIM circuit of the present invention performs MAC operation in the test experiment. DETAILED DESCRIPTION
[0070] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0071] Example 1
[0072] FeFET-based non-volatile compute-in-memory (NV-CIM) has demonstrated significant advantages in non-volatile data storage and low-power applications, demonstrating its potential as both MLC memory and CIM units. Compared to other novel non-volatile memory MLC units and MAC arithmetic units, it offers advantages such as fewer control transistors, smaller area consumption, and lower power consumption.
[0073] However, when such devices are used as MLC memories, the threshold voltage V th The detection method used to determine the storage state within the cell cannot guarantee accurate detection of the specific storage state of the cell in a single operation, and requires a long clock cycle. When used as a component of a MAC unit, most structures can only handle MAC operations under 1-bit input (IN) and 1-bit weight (Weight, W), and can only be used to build low-precision simple neural networks. In addition, the difference between the input (IN) voltage and the storage programming voltage used in traditional solutions is small, which can easily change the storage state of the FeFET, affecting the reusability of the weight information within the cell and potentially causing a decrease in calculation accuracy.
[0074] Based on this, this embodiment provides a FeFET-based MLC memory, which is a completely new NV-CIM circuit structure that combines data storage and logic operation functions. Unlike traditional NV-CIM circuits, the solution provided in this embodiment can store 2-bit data within a single unit circuit, thus providing MLC storage functionality. It can also perform 2-bit × 2-bit multiplication within a single unit circuit, significantly improving its advantages in area overhead, computational power consumption, and accuracy as a CIM circuit with multi-bit multiplication functionality.
[0075] Specifically, if Figure 1 As shown, each unit circuit in the FeFET-based MLC memory provided in this embodiment is composed of two FeFET transistors F1 and F2, one PMOS transistor M1, one NMOS transistor M2, and one sampling capacitor C0. The gates of F1 and F2 are connected to input ports WL1 and WL2, respectively; the gate of M2 is connected to the enable port EN; the drain of F1 is connected to the control port BL; the source of F2 and the drain of M1 are grounded. The sources of F1 and M1 are connected to the drains of F2 and M2, forming a storage node Vs; the gate of M2 is connected to the input port INR; the source of M2 is connected to one end of C0, and the connection point serves as the output port Vout; the other end of C0 is grounded.
[0076] In MLC storage mode, the unit circuit represents the value of the stored 2-bit data using the four different voltage levels that the storage node Vs can exhibit. In calculation mode, the data pre-stored in the storage node Vs is used as a 2-bit weight, and WL1, WL2, and INR are jointly encoded as a 2-bit input. The voltage at Vout is quantized to obtain the product of the weight and the input, thus completing a 2-bit × 2-bit multiplication operation.
[0077] In the FeFET-based MLC memory provided in this embodiment, Figure 1 The new unit circuit shown adopts a new working principle to realize data storage and logic operation. Conventional circuits usually use the circuit structure of F1 and F2 in this embodiment to realize the coupling of binary signals, that is, to realize the storage function of single bit. Although the solution of this embodiment also utilizes the MLC characteristics of FeFET, it does not use different V th The data is stored in a different way according to the state of the data, but the different on-resistance R of two FeFETs (F1, F2) and a PMOS (M1) is used. on The coupling of states represents a storage state.
[0078] Specifically, in the storage mode, for F1 and F2 in this embodiment, applying voltage signals with different pulse heights and the same pulse width to the FeFET gate will change the storage state in the FeFET. The change in storage state is essentially a change in the threshold voltage V of the FeFET. th Or the on-resistance R on Among them, the positive voltage makes the FeFET V th Lower, negative voltage makes FeFET V th Increase. Therefore, by collaboratively regulating the gate voltages of F1 and F2, the voltage level of Vs between the two can be effectively changed. On this basis, the solution of this embodiment continues to introduce a PMOS tube M1. PMOS acts as an equivalent resistor in the on state and as an open circuit in the off state. By utilizing this characteristic, in the circuit of this embodiment, the gate voltages (WL1, WL2, INR) of F1, F2 and M1 are jointly regulated to couple out a four-state signal with high linearity at the storage node Vs, thereby realizing the MLC storage function. On this basis, this embodiment continues to utilize an NMOS tube M2 and a sampling capacitor in the circuit to realize data reading and writing in the MLC mode, and outputting the logical operation results in the operation mode.
[0079] In addition, it should be emphasized that the presence of the PMOS tube M1 in this embodiment can not only regulate the linearity of the output analog signal, but also has the function of ensuring that the source, drain and substrate voltages of the FeFET are 0V during the information storage stage, thereby ensuring accurate regulation of the FeFET state.
[0080] In order to make the performance and advantages of the circuit provided by this embodiment more prominent, the circuit having MLC storage and logic operation functions is described in detail below;
[0081] (1) MLC storage mode
[0082] The circuit provided in this embodiment belongs to an NV memory, and the storage function includes two basic steps: data writing and data reading.
[0083] 1.1 Data Writing
[0084] In the MLC storage mode, the unit circuit in the FeFET-based MLC memory of this embodiment implements the data writing operation logic as follows:
[0085] First, set the INR terminal to a low level (denoted as V WB) and set the BL and EN terminals to 0V; at this time, M1 is in the on state, making the source (S) and drain (D) voltages of F1 and F2 0V, ensuring the precise control of the polarization strength of the ferroelectric layer. Then, according to the preset coding rules, different programming voltages are applied to WL1 and WL2; after the specified writing time t w After that, the data writing is completed.
[0086] For example, when data “00” needs to be stored, a negative voltage V is applied to the gate of F1 at the WL1 terminal for a time period of t0. n01 , a positive voltage V is applied to the gate of F2 at the WL2 terminal for a time of t0. p02 ; At this time, the threshold voltage information stored in F1 is V th01 , the threshold voltage information stored in F2 is V th02 .
[0087] When data "01" needs to be stored, a positive voltage V is applied to the F1 gate at the WL1 terminal for a time of t0. p11 , apply a positive voltage V to the F2 gate at time t0 at WL2 p12 ; At this time, the threshold voltage information stored in F1 is V th11 , the threshold voltage information stored in F2 is V th12 .
[0088] When data "10" needs to be stored, a positive voltage V is applied to the F1 gate at the WL1 terminal for a time of t0. p21 , apply a positive voltage V to the F2 gate at time t0 at WL2 p22 ; At this time, the threshold voltage information stored in F1 is V th21 , the threshold voltage information stored in F2 is V th22 .
[0089] When data "11" needs to be stored, a positive voltage V is applied to the F1 gate at the WL1 terminal for a time of t0. p31 , apply a positive voltage V to the F2 gate at time t0 at WL2 p32 ; At this time, the threshold voltage information stored in F1 is V th31 , the threshold voltage information stored in F2 is V th32 .
[0090] In order to ensure that the voltage level of Vs increases linearly with the values "00", "01", "10" and "11" it needs to represent, the programming voltage applied to the WL1 terminal in this embodiment needs to meet the following requirements:
[0091] V n01 <0 <V p11 <V p21 <Vp31 , and V p11 、V p21 、V p31 The values between them increase in a gradient
[0092] That is, the larger the 2-bit number stored, the larger the programming voltage applied by WL1. Under the same read voltage, the threshold voltage V th The smaller the on-resistance R onF1 The smaller it is, the greater the voltage at point Vs.
[0093] Correspondingly, the programming voltage of WL2 needs to meet the following requirements:
[0094] V p02 >V p12 >V p22 >V p32 , and V p02 、V p12 、V p22 、V p32 The values between them decrease in a gradient
[0095] That is, the larger the 2-bit number stored, the smaller the programming voltage applied by WL2. Under the same read voltage, the threshold voltage V th The larger the on-resistance R onF2 The larger the V S The greater the point voltage.
[0096] Based on this principle, in the data writing phase, the coding rules used by the coding voltages WL1 and WL2 input to the circuit in this embodiment are shown in the following table:
[0097] Table 1: Signal encoding rules for FeFET-based MLC memory in MLC storage mode
[0098]
[0099] The specific contents are:
[0100] (1) When data "00" needs to be written, WL1 is set to -V0 and WL2 is set to V1+3ΔV;
[0101] (2) When data "01" needs to be written, WL1 is set to V0 and WL2 is set to V1+2ΔV;
[0102] (3) When data "10" needs to be written, WL1 is set to V0+ΔV and WL2 is set to V1+ΔV;
[0103] (4) When data "11" needs to be written, WL1 is set to V0+2ΔV and WL2 is set to V1;
[0104] Wherein, V0 and V1 are references for the programming voltages of WL1 and WL2 respectively, and ΔV is the gradient change value of each programming voltage.
[0105] 1.2 Data Reading
[0106] In the MLC storage mode, the unit circuit in the FeFET-based MLC memory of this embodiment implements the data reading operation logic as follows:
[0107] First set BL to VDD and EN to the linear voltage V M ; Then set WL1 and WL2 to read voltage V R , INR is set to read voltage V RB .
[0108] Among them, the linear voltage V M It is a voltage value related to device parameters and can make the storage node voltage Vs output linear current after passing through M2. R Refers to a voltage that can make F1 and F2 in the on state; V RB It refers to a voltage that can make M1 in the on state. When M1 is at this gate voltage, the Vs voltage can be adjusted so that the current corresponding to different storage states output by M2 is linear.
[0109] At this time, F1 and F2 are in different threshold voltage states, that is, the on-resistance R onF1 ≠R onF2 , the different conduction states of F1, F2 and M1 produce different voltage divisions, so the intermediate node V S There are different voltages in different storage states; after the voltage is applied to the drain of M2, a current is generated between the shared capacitor C0, thereby charging the capacitor. R After that, the corresponding stored data can be read according to the voltage value of Vout.
[0110] Assume that when the original storage data in the storage node Vs is "00", the Vs voltage is 0V; when the original storage data is "01", the Vs voltage is V S0 +V offset ; When the original storage data is "10", the Vs voltage is 2*V S0 +V offset ; When the original storage data is "11", the Vs voltage is 3*V S0 +V offset .
[0111] Then, when a discharge path is formed through M2 and C0 is charged, the voltage of the storage node Vs corresponding to different storage states "00", "01", "10", and "11" is different, so the output current on the discharge path generated by M2 is also different, and roughly changes in a gradient close to 0A, I0, 2*I0, and 3*I0. When the control reading time is t R When the charge transfer amount of different storage states on C0 is different, the voltage of the sampling capacitor (corresponding to Vout) is 0V, V c0 , 2*V c0 , 3*V c0 Therefore, the values of different stored data represented by the original storage node Vs in the unit circuit can be read by quantizing the voltage of Vout.
[0112] In summary, in the data reading phase, the output of Vout has four level states: 0V, V cap , 2*V cap , 3*V cap , V cap Indicates the gradient voltage of the preset read signal. The larger the Vout, the larger the stored data. The above four level states respectively indicate that the stored data read from the storage node is "00", "01", "10", and "11".
[0113] (2) Logical operation mode
[0114] As mentioned above, the unit circuit in the FeFET-based MLC memory provided in this embodiment has a 2-bit×2-bit logical operation function. Specifically, when performing a logical operation task, the operands involved in the multiplication operation are composed of two parts: input (IN) and weight (W). Among them, the 2-bit weight (W) value is pre-stored in the unit circuit; the 2-bit input (IN) is obtained by jointly encoding the three signals WL1, WL2, and INR, and the four groups of different voltage value combinations jointly encoded are input into the unit circuit to complete the operation. The final multiplication result is represented by the voltage value of Vout.
[0115] Specifically, the operation logic of the unit circuit in the FeFET-based MLC memory to perform multiplication is as follows:
[0116] (1) Weight pre-storage
[0117] In MLC storage mode, a 2-bit weight value is pre-stored in the storage node.
[0118] (2) Input and calculation
[0119] Set BL to VDD and EN to the linear voltage VM , switch to calculation mode; then adjust the voltage values of WL1, WL2 and INR to complete the input of the corresponding input number.
[0120] (1) When the input (IN) is "00", WL1 applies a negative voltage V to the gate of F1. n0 , WL2 applies 0V to the gate of F1, and INR applies 0V to the gate of M1. At this time, the current flowing through F1 and F2 is close to 0A, V S The voltage is close to 0V.
[0121] Under this condition, when the weight (W) pre-stored in the corresponding unit circuit is any value, the discharge current of M2 is close to 0A.
[0122] (2) When the input (IN) is "01", WL1 applies a positive voltage V to the gate of F1. p1 , WL2 applies a positive voltage V to the F1 gate p , INR applies V to the gate of M1 R1 .
[0123] At this time, if the corresponding unit has different weight values "00", "01", "10", and "11", then V S The voltage at point is 0V, V S0 +V offset , 2*V S0 +V offset , 3*V S0 +V offset By applying voltage to the gate of M2 through EN, M2 operates in the linear region, and Vs voltage is applied to the drain of M2. The discharge current of M2 is close to 0A, I0, 2*I0, and 3*I0.
[0124] (3) When the input (IN) is "10", WL1 applies a positive voltage V to the gate of F1. p2 , WL2 applies a positive voltage V to the F1 gate p , INR applies V to the gate of M1 R2 .
[0125] At this time, if the corresponding unit has different weight values "00", "01", "10", and "11", then V S The voltage at the point is 0V, 2*V S0 +V offset , 4*V S0 +V offset , 6*V S0 +V offset By applying voltage to the gate of M2 through EN, M2 operates in the linear region, and Vs voltage is applied to the drain of M2. The discharge current of M2 is close to 0A, 2*I0, 4*I0, and 6*I0.
[0126] (4) When the input (IN) is "11", WL1 applies a positive voltage V to the gate of F1. p3 , WL2 applies a positive voltage V to the F1 gate p , INR applies V to the gate of M1 R3 .
[0127] At this time, if the corresponding unit has different weight values "00", "01", "10", and "11", V S The voltage at the point is 0V, 3*V S0 +V offset , 6*V S0 +V offset 、9*V S0 +V offset By applying voltage to the gate of M2 through EN, M2 operates in the linear region, and Vs voltage is applied to the drain of M2. The discharge current of M2 is close to 0A, 3*I0, 6*I0, and 9*I0.
[0128] (3) Operation output
[0129] The input and weight are multiplied in the circuit and after the specified sampling time t sampling After that, the corresponding calculation result is read out according to the voltage value of Vout. Specifically, in the logic operation mode, the sampling capacitor C0 is charged by the discharge current generated by M2. When the sampling time t sampling Because different multiplication results satisfy the linear current size, the capacitor voltage Vout sampled at the same charging time will also increase linearly, that is, Vout has 7 levels, which can be represented as 0V, V cap , 2*V cap , 3*V cap , 4*V cap , 6*V cap 、9*V cap ; The above 7 level states respectively represent the multiplication results: "0000", "0001", "0010", "0011", "0100", "0110", and "1001".
[0130] To summarize the above process, the logic operation truth table of the circuit in a straight line 2-bit × 2-bit operation is as follows:
[0131] Table 2: Truth table of FeFET-based MLC memory in logic operation mode
[0132]
[0133]
[0134] Among them, V n0 <0 and V p1 >V p2 >V p3 >0;0 <V R1 <V R2 <V R3 ; Random means that INR can be set to any voltage value.
[0135] Example 2
[0136] Based on the solution of embodiment 1, this embodiment provides a MAC circuit, such as Figure 2 As shown, the MAC circuit is composed of multiple unit circuits arranged in rows, as in the FeFET-based MLC memory of Example 1. The sources of M2 in all unit circuits are connected to the same computation bit line SL, and all unit circuits share the same sampling capacitor C0. One end of C0 is connected to SL, and the other end is grounded.
[0137] Combining this circuit design, it can be found that each unit circuit in the MAC circuit can complete an independent 2-bit×2-bit multiplication operation. The result of each operation is reflected in the magnitude of the discharge current output through the source of M2. The discharge current generated by all unit circuit operations will eventually be output to the same calculation bit line SL, and the sampling capacitor C0 shared by the column will be charged. According to the charge effect of the capacitor, it can be seen that the final capacitance of the sampling capacitor C0 and its voltage will change linearly with the multiplication and accumulation result of the unit circuit. Therefore, in the MAC circuit provided by this embodiment, after the unit circuits in each row complete the corresponding multiplication operation, the bit line voltage V SL Represents the result of the multiplication and accumulation operations of all rows.
[0138] Accordingly, the process of performing the multiplication-accumulation operation by the MAC circuit provided in this embodiment is as follows:
[0139] (1) Weight pre-storage
[0140] In the MLC storage mode, the weight value of the corresponding multiplication operation part is pre-stored in the storage node of the unit circuit of the selected corresponding row.
[0141] (2) Input and calculation
[0142] Set BL to VDD and EN to V M2 Switch to calculation mode; then adjust the voltage values of WL1, WL2 and INR in the unit circuit of the corresponding row to complete the input of the input number of the multiplication operation part.
[0143] (3) Operation output
[0144] Each unit circuit completes the multiplication operation of its corresponding input and weight, and after the specified sampling time, according to the bit line voltage V SL Read out the corresponding calculation results, where V SL Relative gradient voltage V cap The multiple of is the result of the multiplication and accumulation operation.
[0145] For example, assuming a multiplication-accumulation operation is: 01×01+11×10, the 01×01 operation can be performed in the unit circuit of the first row, and the 11×10 operation can be performed in the unit circuit of the second row. At this time, after the operation is completed, combined with the previous content, it can be seen that the discharge currents of the two unit circuits to the calculation bit line SL are I0 and 6*I0 respectively. After the same sampling time, the unit circuit of the first row can make V SL From 0V to V cap , the unit circuit in the second row can make V SL From 0V to 6V cap ; After the two phases are superimposed, V SL Finally it rises to 7V cap , the corresponding binary number is "0111". That is, the MAC circuit completes the multiplication and accumulation operation of 01×01+11×10=0111.
[0146] Example 3
[0147] Based on the embodiments 1 and 2, this embodiment further provides a CIM circuit, such as Figure 3 As shown, it includes a storage and computing array and its peripheral circuits.
[0148] The memory and computation array is composed of multiple MAC circuits arranged in columns, as in Example 2. In the memory and computation array, the gates of F1 and F2 in each unit circuit in the same row are connected to the same set of input word lines, denoted as WLI and WL2. The gate of M1 in each unit circuit in the same row is connected to the same input word line, denoted as INR. The drain of F1 in each unit circuit in the same column is connected to the same control bit line, denoted as BL. The gate of M2 in each unit circuit in the same column is connected to the same control bit line, denoted as EN. The computation bit line SL in the same column is used to input the computation results of each column.
[0149] The peripheral circuit includes a word line driver (WL1 / WL2 / INR Driver), a bit line driver (CaculateDriver), a switch array, a sense amplifier array (Sense Amplifier), an ADC, and a shift adder (Shift-Adder). The word line driver is used to regulate the word line voltage of the input word lines WLI, WL2, and INR of each row in the storage and calculation array; the bit line driver is used to switch the level state of the control bit lines BL and EN of each column in the storage and calculation array. The calculation bit lines of each column in the storage and calculation array are connected to the sense amplifiers of the corresponding columns in the sense amplifier array through the switches on the corresponding columns of the switch array. The output of the sense amplifier of each column is connected to the input of the ADC, and the output of the ADC is connected to the input of the shift adder. In the storage and calculation array of this embodiment, the voltage signal V representing the calculation result output by the calculation bit line of each column is connected to the input of the ADC. SL It is first quantized by a sensitive amplifier and then converted into a digital quantity of the calculation result by ADC processing.
[0150] exist Figure 3 In the CIM circuit's memory-computation array, each unit circuit implements 2-bit data storage and multiplication operations between a 2-bit weight and a 2-bit input. The unit circuits in each column form a MAC circuit, which performs multiplication-accumulation operations between a 2-bit weight and a 2-bit input.
[0151] In particular, in a CIM circuit including multiple columns of MAC circuits, it can also implement 2-bit×2n-bit (n>1) multiplication or multiply-accumulate operations. In practical applications, this embodiment also provides two different operation modes for performing such operations:
[0152] (1) Single-cycle operation mode
[0153] In this mode, the MAC circuit performs operations as multiplication-accumulation operations between 2-bit inputs and 2n-bit weights. The specific operation logic is as follows:
[0154] (1) The 2n-bit weight is decomposed into multiple 2-bit weights with weights according to every two bits; and then pre-stored into each unit circuit in different columns.
[0155] (2) Perform multiplication and accumulation operations between each 2-bit input and the 2-bit weights of various weights in different columns of the storage array.
[0156] (3) Each switch in the switch array is closed in sequence; then the numerical value of the multiplication and accumulation result of the corresponding column is generated through the sensitive amplifier and ADC, and finally the digital value of the multiplication and accumulation result of each column is shifted and added according to the preset weight through the shift adder to obtain the final operation result.
[0157] (2) Multi-cycle operation mode
[0158] In this mode, the MAC circuit performs operations as multiplication-accumulation operations between 2-bit weights and 2n-bit inputs. The specific operation logic is as follows:
[0159] (1) Decompose the 2n-bit input into U 2-bit inputs with input weights at every two bits;
[0160] (2) In each cycle, a column is selected to perform a multiplication-accumulation operation between one of the 2-bit inputs and the 2-bit weight;
[0161] (4) After each cycle, the corresponding switch of the switch array is closed; the numerical value of the corresponding column multiplication and accumulation result is generated through the sense amplifier and the ADC; and the digital value of the multiplication and accumulation result of the current cycle and the multiplication and accumulation result of the previous cycle are shifted and added according to the preset input weights through the shift adder;
[0162] (4) After U cycles, the shift adder outputs the final operation result.
[0163] In summary, the FeFET-based MLC memory provided in this embodiment can perform 2-bit × 2-bit multiplication operations within a single memory cell. The expanded MAC circuit can perform multi-row parallel MAC operations by sharing SL output lines and column-shared accumulation capacitor structures with cells in the same column. The further expanded CIM circuit can expand to 2-bit × 2n-bit MAC operations in an n-column memory array by changing the maximum number of bits supported by the shift-adder. This new circuit design also demonstrates significant advantages in area overhead, computing power consumption, and computational accuracy.
[0164] In practical applications, the MLC memory, MAC circuit, and CIM circuit provided in the embodiments of the present invention can be implemented as circuit modules or chip products. For example, a CIM chip can be provided, which is packaged with the aforementioned CIM circuit. This circuit then has both MLC storage functionality and multiplication and multiply-accumulate operations between multi-bit numbers.
[0165] Performance Testing
[0166] In order to verify the circuit performance of the FeFET-based MLC memory and its corresponding MAC circuit, CIM circuit, and other solutions provided by the present invention, technicians simulated and tested them in relevant solutions.
[0167] 1. Node voltage linearity
[0168] This experiment simulates and tests the influence of the PMOS transistor in the novel MLC memory provided by the present invention on the linearity of the node voltage Vs of the storage node. Figure 1 The solution of the present invention is as shown in Figure 4 The performance of the control group solution without PMOS is compared, and the level state of Vs is obtained when the same gate voltage is applied to F1 and F2.
[0169] According to the experimental results, the Figure 5 The comparison chart shown in the figure. Analyzing the data in the figure, we can find that: when writing the same data, Figure 5 Part (a) shows that the 2FeFET and 1PMOS scheme provided by the present invention can couple different level states that change linearly at the storage node. Figure 5 Part (b) shows that the various level states coupled out by the control group scheme are obviously nonlinear.
[0170] 2. Basic functional testing
[0171] (2.1) MLC storage function
[0172] This experiment Figure 1 The data storage function of the MLC memory shown in the MLC storage mode is tested, and the tested signal flow diagram is as follows Figure 6 shown.
[0173] analyze Figure 6 The data can be found:
[0174] The voltage division of F1, F2, and M1 is completed in a very short time. Compared with the existing FeFET structure, the read delay is greatly reduced and can be at the same level as SRAM. The read delay is significantly better than other non-volatile memories. Due to the non-volatile nature of the circuit, in the hold phase that does not require processing after saving the data, all port levels are set to 0V, and there is almost no static power consumption. Compared with the existing FeFET structure, the circuit ensures that the Vs voltage is 0 during the storage phase, achieving precise control of the FeFET threshold voltage. The threshold voltage adjustment error is small, further ensuring the linearity of the result during the operation. In addition, considering that the precise control of the threshold voltage of other FeFETs requires ensuring that the S / D terminals are both 0, the operating logic of the solution of the present invention is simpler.
[0175] (2.2) MAC operation function
[0176] This experiment Figure 3 The function of the CIM circuit shown in the figure is tested in the logic operation mode. The tested signal flow diagram is shown in the figure below. Figure 7 shown.
[0177] analyze Figure 7 The data can be found:
[0178] The computational efficiency of the solution of the present invention is higher, and only the reading time t R Multiplication operations can be performed in the same amount of time, with a short time cycle. Existing FeFET multiplication and accumulation circuits are mostly current-type, while the calculation results of the present invention are voltage-type, which is easier to quantize. Therefore, the present invention can reduce the area and power consumption of quantization-related circuits. In the operational logic of the present invention, different MAC result values consume almost the same amount of time. The voltage corresponding to each value of the Vs point is much smaller than VDD, and the SL current is the current in the M2 linear region, which is smaller than the saturation region current of the existing circuit. This results in low power consumption, smaller required capacitance value, and smaller area overhead.
[0179] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A FeFET-based MLC memory, wherein the unit circuit used has data storage and multiplication functions, characterized by: Each unit circuit consists of two FeFET transistors F1 and F2, a PMOS transistor M1, an NMOS transistor M2, and a sampling capacitor C0. The gates of F1 and F2 are connected to input ports WL1 and WL2, respectively; the gate of M2 is connected to the enable port EN; the drain of F1 is connected to the control port BL; the source of F2 and the drain of M1 are grounded; the sources of F1 and M1 are connected to the drains of F2 and M2, forming a storage node Vs; the gate of M2 is connected to the input port INR; the source of M2 is connected to one end of C0, and the connection point serves as the output port Vout; the other end of C0 is grounded. In MLC storage mode, the value of the stored 2-bit number is represented by four different voltage levels of the storage node Vs; In calculation mode, the data pre-stored in the storage node Vs is used as a 2-bit weight, and WL1, WL2, and INR are jointly encoded as a 2-bit input. The voltage of Vout is quantized to obtain the product of the weight and the input.
2. The FeFET-based MLC memory according to claim 1, wherein: In storage mode, the operation logic for writing data is as follows: First, set the INR terminal to low level and the BL and EN terminals to 0V; then apply different programming voltages to WL1 and WL2 according to the following coding rules; after the specified writing time t w After that, the data is written: When data "00" needs to be written, WL1 is set to -V0 and WL2 is set to V1+3ΔV; When data "01" needs to be written, WL1 is set to V0 and WL2 is set to V1+2ΔV; When data "10" needs to be written, WL1 is set to V0+ΔV and WL2 is set to V1+ΔV; When data "11" needs to be written, WL1 is set to V0+2ΔV and WL2 is set to V1; Wherein, V0 and V1 are references for the programming voltages of WL1 and WL2 respectively, and ΔV is the gradient change value of each programming voltage.
3. The FeFET-based MLC memory according to claim 1, wherein: In MLC storage mode, the operation logic for data reading is: First set BL to VDD and EN to the linear voltage V M ; Then set WL1 and WL2 to read voltage V R , INR is set to read voltage V RB ; then after the specified reading time t R Then, the corresponding stored data is read according to the voltage value of Vout: When Vout is 0V, V cap , 2*V cap , 3*V cap When , it means the storage data of the storage node is "00", "01", "10", "11" respectively; V cap Indicates the gradient voltage of the preset read signal.
4. The FeFET-based MLC memory according to claim 1, wherein: In calculation mode, the operation logic for performing multiplication operations is as follows: (1) Weight pre-storage In MLC storage mode, the weight value is pre-stored in the storage node; (2) Input and calculation Set BL to VDD and EN to the linear voltage V M Switch to calculation mode; then adjust the voltage values of WL1, WL2, and INR to complete the input of the corresponding input number. The encoding rules of the input number are: When WL1=V n0 , when WL2=0V, INR=Random, the characterization input is "00"; When WL1=V p1 , WL2=V p , INR=V R1 When , the character input is "01"; When WL1=V p2 , WL2=V p , INR=V R2 When , the characterization input is "10"; When WL1=V p3 , WL2=V p , INR=V R3 When , the character input is "11"; Among them, V n0 <0 and V p1 >V p2 >V p3 >0;0 <V R1 <V R2 <V R3 ; (3) Operation output The input and weight are multiplied in the circuit and after the specified sampling time t sampling Then, the corresponding calculation result is read according to the voltage value of Vout: When Vout is 0V, V cap , 2*V cap , 3*V cap , 4*V cap , 6*V cap 、9*V cap , indicating that the product results are: "0000", "0001", "0010", "0011", "0100", "0110", "1001".
5. A MAC circuit, characterized in that: It is composed of a plurality of unit circuits in the FeFET-based MLC memory according to any one of claims 1 to 4 arranged in rows; in the MAC circuit, the source of M2 in all unit circuits is connected to the same calculation bit line SL, and each unit circuit shares the same sampling capacitor C0; one end of C0 is connected to SL, and the other end is grounded; After the MAC circuit completes the corresponding multiplication operation in each row of unit circuits, the SL bit line voltage V SL Represents the result of the multiplication and accumulation operations of all rows.
6. The MAC circuit according to claim 5, wherein: The process of performing multiplication and accumulation operations is as follows: (1) Weight pre-storage In the MLC storage mode, the weight value of the corresponding multiplication operation part is pre-stored in the storage node of the unit circuit of the selected corresponding row; (2) Input and calculation Set BL to VDD and EN to V M2 Switch to calculation mode; then adjust the voltage values of WL1, WL2 and INR in the unit circuit of the corresponding row to complete the input of the input number of the multiplication operation part; (3) Operation output Each unit circuit completes the multiplication operation of its corresponding input and weight, and after the specified sampling time, according to the bit line voltage V SL Read out the corresponding calculation results, where V SL Relative gradient voltage V cap The multiple of is the result of the multiplication and accumulation operation.
7. A CIM circuit, characterized in that: It includes a storage and computing array and its peripheral circuits; The memory and calculation array is formed by arranging a plurality of MAC circuits according to claim 5 in columns; in the memory and calculation array, the gates of F1 and F2 in each unit circuit in the same row are connected to the same set of input word lines, denoted as WLI and WL2; the gate of M1 in each unit circuit in the same row is connected to the same input word line, denoted as INR; the drain of F1 in each unit circuit in the same column is connected to the same control bit line, denoted as BL; the gate of M2 in each unit circuit in the same column is connected to the same control bit line, denoted as EN; the calculation bit line SL in the same column is used to input the calculation results of each column; The peripheral circuit includes a word line driver, a bit line driver, a switch array, a sense amplifier array, an ADC, and a shift adder; the word line driver is used to regulate the word line voltages of the input word lines WLI, WL2, and INR of each row in the storage and calculation array; the bit line driver is used to switch the level states of the control bit lines BL and EN of each column in the storage and calculation array; the calculation bit lines of each column in the storage and calculation array are connected to the sense amplifiers of the corresponding columns in the sense amplifier array through switches on the corresponding columns of the switch array; the output of the sense amplifier of each column is connected to the input of the ADC, and the output of the ADC is connected to the input of the shift adder; In the storage array, each column calculates the voltage signal V output by the bit line to represent the calculation result. SL It is first quantized by a sensitive amplifier and then converted into a digital quantity of the calculation result by ADC processing.
8. The CIM circuit according to claim 7, wherein: Each unit circuit in the memory-calculation array is used to implement 2-bit data storage and multiplication operations between 2-bit weights and 2-bit inputs; each unit circuit in each column constitutes a MAC circuit and is used to perform multiplication and accumulation operations between 2-bit weights and 2-bit inputs; multiple columns of MAC circuits are used to perform multiplication and accumulation operations between 2-bit inputs and 2n-bit weights; n>1; The operation logic of the multiplication and accumulation operation between the 2-bit input and the 2n-bit weight is as follows: Decompose the 2n-bit weight into multiple 2-bit weights with weights according to every two bits; then pre-store them into each unit circuit in different columns; Perform multiplication and accumulation operations between each 2-bit input and the 2-bit weights of various weights in different columns of the storage array; Each switch in the switch array is closed in sequence; the numerical value of the corresponding column multiplication and accumulation result is then generated through the sensitive amplifier and ADC. Finally, the digital value of the multiplication and accumulation result of each column is shifted and added according to the preset weight through the shift adder to obtain the final calculation result.
9. The CIM circuit according to claim 8, wherein: The multi-column MAC circuit is also used to perform the multiplication and accumulation function between the 2-bit weight and the 2n-bit input; n>1; its operation logic is as follows: Decompose the 2n-bit input into U 2-bit inputs with input weights for every two bits; In each cycle, a column is selected to perform a multiplication-accumulation operation between one of the 2-bit inputs and the 2-bit weight. After each cycle, the corresponding switch of the switch array is closed; the digital value of the multiplication and accumulation result of the corresponding column is generated through the sense amplifier and ADC; and the digital value of the multiplication and accumulation result of the current cycle and the multiplication and accumulation result of the previous cycle are shifted and added according to the preset input weights through the shift adder; After U cycles, the shift adder outputs the final operation result.
10. A chip, characterized in that: It is encapsulated by the CIM circuit according to any one of claims 7 to 9.