Method for realizing full adder based on 2T-2C FRAM memory operation
By using an in-memory computing unit based on 2T-2C FRAM in the von Neumann architecture, Majority operations are implemented to support the full adder function, and the problems of limited computing speed and high power consumption in traditional architectures are solved, and efficient in-memory computing is achieved.
Patent Information
- Application Number
- CN202510231265.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-24
AI Technical Summary
The traditional von Neumann architecture has problems with "memory wall" and "power wall" when dealing with big data and high computing power requirements, resulting in limited computing speed and high power consumption.
The in-memory computing unit based on 2T-2C FRAM is adopted to realize the function of a full adder through Majority operations, and data is calculated and processed in memory, reducing the need for data transmission.
It realizes efficient fully added logic operations in FRAM in-memory computing units, reduces power consumption, improves calculation speed, and fills the gap in logical calculation methods in FRAM in-memory computing.
Smart Images

Figure CN120196587A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of integrated circuits and relates to a method for implementing a full adder based on 2T-2C FRAM in-memory computing Background Art
[0002] In recent years, with the continuous development of artificial intelligence, large models and big data algorithms have emerged in an endless stream, and the requirements for chip computing power have also been continuously increasing. The traditional von Neumann architecture can no longer meet the increasingly huge data computing requirements. In the von Neumann architecture, data is stored in a memory. When computing is required, the processor needs to call data from the memory, and the memory and the processor are connected by a data bus for data transmission. Such a structure has significant disadvantages. First, the computing speed of the processor is faster than the access speed of the memory, so the computing power of the chip will be limited by the bandwidth, resulting in the actual computing power of the processor being much lower than the theoretical computing power, making it difficult to meet the requirements of fast computing and accurate response of intelligent chips. This problem is called the "memory wall" problem. By increasing the bandwidth of the bus and the clock frequency, the data transmission speed can be increased, thereby improving the performance of the processor to a certain extent. However, this will lead to high power consumption and integration costs, and its scalability is also severely limited. Second, in the von Neumann architecture, the memory and the processor are separated, and the data to be processed will be frequently transmitted between the processor and the memory, which will generate huge transmission power consumption (transmission power consumption accounts for 70% of the overall power consumption), which is also called the "power wall" problem. For example, a research report by NVIDIA points out that the data transmission power consumption required for floating-point operations is about 200 times that of data processing power consumption. The above "memory wall" and "power wall" problems are collectively referred to as the bottlenecks of the von Neumann architecture
[0003] In order to break through the bottlenecks of the von Neumann architecture, two new architectures, namely the near-memory computing architecture and the in-memory computing architecture, have been proposed. The near-memory computing architecture increases the data bandwidth by means of high-speed interfaces, three-dimensional stacking, and increasing on-chip caches, etc., and at the same time reduces the distance between the processor and the memory to reduce power consumption. However, the near-memory computing architecture essentially still belongs to the von Neumann architecture and can only alleviate the "memory wall" and "power wall" bottlenecks of the von Neumann architecture by increasing the bandwidth and reducing the transmission distance between the memory and the processor, and cannot fundamentally solve the bottlenecks of the von Neumann architecture. Therefore, a completely new in-memory computing architecture has been proposed. The in-memory computing architecture uses the memory to calculate and process data, without the need for data to be repeatedly called between the processor and the memory, realizing the integration of storage and computing, and is expected to break through the "memory wall" and "power wall" bottlenecks of the von Neumann architecture. Since in-memory computing is expected to significantly improve the computing speed and reduce the computing power consumption, this technology has broad application prospects in intelligent chips
[0004] So far, the industry has developed various computing-in-memory architectures based on static random access memory (SRAM), dynamic random access memory (DRAM), flash memory, resistive random access memory (ReRAM), phase change memory (PCM), ferroelectric field-effect transistor (FeFET), magnetic random access memory (MRAM), etc. However, they still face various problems and challenges on the road to industrialization. SRAM has the advantages of mature technology and advanced process nodes, but it is a volatile memory, and power failure will cause data loss. The computing-in-memory unit of SRAM also occupies a large area, which is not conducive to highly integrated in-memory computing chips with high computing performance. DRAM also has a mature process, and the area of the computing-in-memory unit of DRAM is small, but like SRAM, it is a volatile memory and cannot save data when power is off. And because DRAM uses capacitors to store data, it needs to be refreshed regularly and there is a leakage phenomenon, making it difficult to achieve high-precision in-memory computing. DRAM is widely used in the near-memory computing architecture of three-dimensional stacking. ReRAM is non-volatile, can save data when power is off, and can implement large-scale cross-point arrays. It is one of the potential chips for the industrialization of computing-in-memory chips in the future; however, the current process of ReRAM is not yet mature. ReRAM requires a large programming voltage, so it is difficult to be manufactured using advanced nodes. The multi-bit in-memory computing accuracy of ReRAM in in-memory computing is poor (generally less than 8 bits), and its robustness is also poor. Phase change memory PCM also belongs to non-volatile memory and can implement large-scale cross arrays, but the read / write power consumption of PCM is large, the read / write speed is slow, and its durability is poor. FeFET is a non-volatile memory and can implement cross-point arrays, but the current process is not yet mature, and its data retention characteristics are poor, and its read / write tolerance is also poor. MRAM is a non-volatile memory, with the advantages of high durability, high speed, low power consumption, etc. The process of MRAM is relatively mature and has good scalability. However, the ratio of the high-resistance state to the low-resistance state of MRAM is low (about 250%), and its reliability in multi-bit in-memory computing is low. Flash is a non-volatile memory, with a mature process and low cost, and has achieved mass production of in-memory computing chips; however, Flash still needs to be further improved in terms of miniaturization, and the programming time of Flash is long.
[0005] Compared with DRAM and SRAM, ferroelectric random access memory (FRAM), as a non-volatile memory, has the ability to store data when power is off, which is beneficial for low-power design. Moreover, there is no leakage problem in ferroelectric capacitors. Compared with DRAM, FRAM has better reliability in in-memory computing. Compared with other non-volatile memories used for in-memory computing (ReRAM, PCM, MRAM, Flash, FeFET), FRAM memory has lower read and write power consumption than MRAM, Flash, and PCM, faster read and write speeds than Flash and PCM, and more read and write cycles than Flash, ReRAM, and PCM. In addition, FRAM based on hafnium oxide thin film also has the advantages of high compatibility with CMOS process and strong radiation resistance. However, since logical computing needs to consider low carry and high carry, there is currently no logical computing method based on FRAM in-memory computing. Therefore, on the basis of previous work, we propose a method for implementing a full adder based on a 2T-2C FRAM in-memory computing unit, filling the gap in logical computing methods in FRAM in-memory computing and promoting the application of FRAM in-memory computing units in neural networks and embodied intelligent chips. Summary of the Invention
[0006] In the present invention, a method for implementing a full adder through in-memory operation based on 2T-2C FRAM is designed and proposed for the first time. Through the Majority operation that fully conforms to the 2T-2C unit, a method for implementing a full adder in the FRAM in-memory computing unit is given. This method can minimize the invocation of FRAM units and has the simplest timing. The functions and timing of the FRAM in-memory computing full adder are simulated and verified.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] The circuit is composed of an in-memory computing circuit composed of 2T-2C ferroelectric storage units ( Figure 1 ), and by simultaneously activating multiple word lines to read the data in the ferroelectric storage units, logical operations can be performed in the ferroelectric storage units. The 4T-2C unit can store the non-logic of the data through the word line WLN, thereby realizing non-logical operations ( Figure 2 ). The two ends of the latch-type sense amplifier are respectively connected to BL and BLN. After comparing the voltage magnitudes on BL and BLN, the end with the larger voltage is pulled up to VDD, and the end with the smaller voltage is pulled to zero level to complete the calculation
[0009] In order to design the full adder algorithm that best conforms to the 2T-2C FRAM in-memory computing unit, we first define the Majority operation:
[0010] Majority(x1, x2.....x n ) = xi (x1 - x n Among them, the number of x i values is the largest) (1)
[0011] For the 2T - 2C FRAM in - memory computing circuit, the Majority operation fully conforms to its circuit structure. Since the 2T - 2C FRAM in - memory computing unit uses complementary ferroelectric capacitors for comparison and then amplifies them in the sense amplifier, when multiple word lines are activated for calculation, if there are more data "1"s stored in the computing capacitors connected to BL, there will be more data "0"s stored in the comparison capacitors connected to BLN. After calculation and amplification through the sense amplifier, the potential on BL is pulled up to VDD, the potential on BLN is pulled down to "0", and the output calculation result is "1". If there are more data "0"s stored in the computing capacitors (A1, B1, C1) connected to BL, then there will be more data "1"s stored in the comparison capacitors (A2, B2, C2) connected to BLN. After calculation and amplification through the sense amplifier, the potential on BL is pulled down to 0, the potential on BLN is pulled up to VDD, and the output calculation result is "0". When the 2T - 2C FRAM in - memory computing unit performs calculations, the calculation result depends on the value of the data with the majority, so the 2T - 2C FRAM in - memory computing unit is completely consistent with the Majority operation.
[0012] Next, the majority algorithm is used to perform full - adder operations. The truth table of the full - adder is as Figure 3 shown, where Xi - 1 represents the low - order carry, Ai represents the augend, Bi represents the addend, Si represents the sum, and Xi represents the carry. Converting the full - add result into a Majority operation, the logic of the full - adder can be expressed by the following logical expressions:
[0013] Xi = Majority(Ai, Bi, Xi - 1) (2)
[0014] Si = Majority(Ai, Bi, Xi - 1, Xi’, Xi’) (3)
[0015] The Majority operation defines that the calculation result is the number that accounts for the majority within the parentheses. For example, when Ai = 1, Bi = 0, Xi - 1 = 0, through the Majority operation, Xi = 0 and Si = 1 can be obtained, and the result is consistent with the truth table. Using the ferroelectric 2T - 2C in - memory computing unit, the logical full - add function can be realized. In addition, through serial addition, the calculation result S of A + B is stored in the FRAM array, and S can be called to participate in subsequent calculations when needed. The FRAM full - adder we proposed can realize the full - add calculation of consecutive multiple numbers such as A + B + C+....
[0016] Advantages of the present invention:
[0017] For the first time, a method for implementing a full adder based on a FRAM in-memory computing unit is designed and proposed, which can perform logical full addition operations on input data. The previously proposed FRAM in-memory computing unit can only perform Boolean logic operations. Based on this, the present invention proposes a method for implementing full addition logic operations in a FRAM in-memory computing unit, improving the functions of the FRAM in-memory computing unit. The FRAM in-memory computing circuit we proposed has the advantages of high reliability, high tolerance, and low power consumption, and is expected to be applied to brain-like chips, embodied intelligent chips, and intelligent chips in the military, aerospace and other fields. Brief Description of the Drawings
[0018] Figure 1 It is a schematic diagram of the basic structure of an in-memory computing unit based on 2T-2C FRAM.
[0019] Figure 2 It is the truth table of the input and output of the full adder.
[0020] Figure 3 It is a schematic diagram of the structure and operation for implementing in-memory full addition calculation in FRAM.
[0021] Figure 4 It is a schematic diagram of the 4T-2C unit structure for implementing non-logical operations.
[0022] Figure 5 It is a schematic diagram of the simulation circuit for performing two-bit full addition.
[0023] Figure 6 It is the full addition simulation waveform diagram of "11" + "11".
[0024] Figure 7 It is the full addition simulation waveform diagram of "10" + "11".
[0025] Figure 8 It is the full addition simulation waveform diagram of "10" + "01".
[0026] Figure 9 It is the full addition simulation waveform diagram of "10" + "00".
[0027] Figure 10 It is the full addition simulation waveform diagram of "01" + "11".
[0028] Figure 11 It is the full addition simulation waveform diagram of "01" + "01". Detailed Implementation Manner
[0029] When adding multiple numbers, n ferroelectric memory cells and corresponding sense amplifiers are connected to each bit line. Taking the full addition of two two-digit numbers (each bit is added separately) as an example, its circuit structure is asFigure 4 As shown, 13 2T-2C ferroelectric memory cells and 2 4T-2C ferroelectric cells are connected in parallel to a pair of complementary bit lines BL and BLN, and both ends of the sense amplifier are also connected to BL and BLN respectively.
[0030] A schematic diagram of the calculation process is as Figure 5 shown. When calculating Ai + Bi, first select the word lines corresponding to Ai, Bi, and Xi-1 to calculate Xi and Xi'. For subsequent calculations, the results of Xi and Xi' need to be stored in two different 2T-2C ferroelectric memory cells respectively. The calculation result of Xi' needs to be written into the 4T-2C non-logical operation unit through the WLN word line of the 4T-2C unit we proposed before. Next, simultaneously select the word lines of Ai, Bi, Xi-1, Xi', and Xi' to calculate Si, and write the calculation result of Si into the row storing the calculation result. Then, perform the full addition operation for the next bit.
[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the beneficial effects of the present invention will be further elaborated through six specific implementation schemes as follows:
[0032] Example 1:
[0033] Add two numbers in the FRAM in-memory computing unit, where each number is two digits. Take the full addition logic operation of "11" + "11" as an example. The circuit calls 13 2T-2C ferroelectric memory cells, 2 4T-2C ferroelectric cells, and a latch-type sense amplifier, as Figure 4 shown.
[0034] The specific operation process is as follows: Write the input data "11" and "11" into the 2T-2C ferroelectric memory cells controlled by WL1 - WL8. Due to the destructive readout characteristic of the ferroelectric capacitor, each input data needs to be written into two different 2T-2C cells, so 8 2T-2C ferroelectric memory cells are required. Then write the low-bit carry signal (the carry signal of the lowest bit is "0") into the 2T-2C ferroelectric memory cells controlled by WL9 and WL10. Then calculate in sequence to obtain X0 = Majority(A0, B0, 0), X0', S0 = Majority(A0, B0, 0, X0', X0'), X1 = Majority(A1, B1, X0)(S2), X1', S1 = Majority(A1, B1, X0, X1', X1'). Finally, perform read operations on the 2T-2C ferroelectric memory cells controlled by WL11, WL12, and WL13 in sequence to read out the calculation result "110". The simulation waveform diagram is as Figure 6 shown.
[0035] Example 2:
[0036] Add two numbers in the FRAM in-memory computing unit. Each number is two digits. Taking the full adder logic operation of "10" + "11" as an example. The circuit calls 13 2T-2C ferroelectric memory cells, 2 4T-2C ferroelectric cells and a latch-type sensitive amplifier, as Figure 4 shown.
[0037] The specific operation process is as follows: Write the input data "10" and "11" into the 2T-2C ferroelectric memory cells controlled by WL1-WL8. Due to the destructive readout characteristic of the ferroelectric capacitor, each input data needs to be written into two different 2T-2C cells, so 8 2T-2C ferroelectric memory cells are required. Then write the low-order carry signal (the carry signal of the lowest bit is "0") into the 2T-2C ferroelectric memory cells controlled by WL9 and WL10. Then calculate in sequence to get X0 = Majority(A0, B0, 0), X0', S0 = Majority(A0, B0, 0, X0', X0'), X1 = Majority(A1, B1, X0)(S2), X1 , , S1 = Majority(A1, B1, X0, X1', X1'), and finally perform read operations on the 2T-2C ferroelectric memory cells controlled by WL11, WL12, and WL13 in sequence to read out the calculation result "101". The simulation waveform diagram is as Figure 7 shown.
[0038] Embodiment 3:
[0039] Add two numbers in the FRAM in-memory computing unit. Each number is two digits. Taking the full adder logic operation of "10" + "01" as an example. The circuit calls 13 2T-2C ferroelectric memory cells, 2 4T-2C ferroelectric cells and a latch-type sensitive amplifier, as Figure 4 shown.
[0040] The specific operation process is as follows: Write the input data "10" and "01" into the 2T-2C ferroelectric memory cells controlled by WL1-WL8. Due to the destructive readout characteristic of the ferroelectric capacitor, each input data needs to be written into two different 2T-2C cells, so 8 2T-2C ferroelectric memory cells are required. Then write the low-order carry signal (the carry signal of the lowest bit is "0") into the 2T-2C ferroelectric memory cells controlled by WL9 and WL10. Then calculate in sequence to get X0 = Majority(A0, B0, 0), X0', S0 = Majority(A0, B0, 0, X0', X0'), X1 = Majority(A1, B1, X0)(S2), X1', S1 = Majority(A1, B1, X0, X1', X1'), and finally perform read operations on the 2T-2C ferroelectric memory cells controlled by WL11, WL12, and WL13 in sequence to read out the calculation result "011". The simulated waveform diagram is as Figure 8 shown.
[0041] Example 4:
[0042] Add two numbers in the FRAM in-memory computing unit. Each number is two digits. Take the full adder logic operation of "10" + "00" as an example. The circuit calls 13 2T-2C ferroelectric memory cells, 2 4T-2C ferroelectric cells and a latch-type sensitive amplifier, as Figure 4 shown.
[0043] The specific operation process is as follows: Write the input data "10" and "00" into the 2T-2C ferroelectric memory cells controlled by WL1-WL8. Due to the destructive readout characteristic of the ferroelectric capacitor, each input data needs to be written into two different 2T-2C cells, so 8 2T-2C ferroelectric memory cells are required. Then write the low-order carry signal (the carry signal of the lowest bit is 0) into the 2T-2C ferroelectric memory cells controlled by WL9 and WL10. Then calculate in sequence to get X0 = Majority(A0, B0, 0), X0', S0 = Majority(A0, B0, 0, X0', X0'), X1 = Majority(A1, B1, X0)(S2), X1', S1 = Majority(A1, B1, X0, X1', X1'), and finally perform read operations on the 2T-2C ferroelectric memory cells controlled by WL11, WL12, and WL13 in sequence to read out the calculation result "010". The simulated waveform diagram is as Figure 9 shown.
[0044] Example 5:
[0045] Add two numbers in the FRAM in-memory computing unit. Each number is two digits. Taking the full adder logic operation of "01" + "11" as an example. The circuit calls 13 2T-2C ferroelectric memory cells, 2 4T-2C ferroelectric cells and a latch-type sensitive amplifier, as Figure 4 shown.
[0046] The specific operation process is as follows: Write the input data "01" and "11" into the 2T-2C ferroelectric memory cells controlled by WL1-WL8. Due to the destructive readout characteristic of the ferroelectric capacitor, each input data needs to be written into two different 2T-2C cells, so 8 2T-2C ferroelectric memory cells are required. Then write the low-order carry signal (the carry signal of the lowest bit is 0) into the 2T-2C ferroelectric memory cells controlled by WL9 and WL10. Then calculate in sequence to get X0 = Majority(A0, B0, 0), X0', S0 = Majority(A0, B0, 0, X0', X0'), X1 = Majority(A1, B1, X0)(S2), X1', S1 = Majority(A1, B1, X0, X1', X1'). Finally, perform read operations on the 2T-2C ferroelectric memory cells controlled by WL11, WL12, and WL13 in sequence to read out the calculation result "100". The simulation waveform diagram is as Figure 10 shown.
[0047] Example 6:
[0048] Add two numbers in the FRAM in-memory computing unit. Each number is two digits. Taking the full adder logic operation of "01" + "01" as an example. The circuit calls 13 2T-2C ferroelectric memory cells, 2 4T-2C ferroelectric cells and a latch-type sensitive amplifier, as Figure 4 shown.
[0049] The specific operation process is as follows: Write the input data "01" and "01" into the 2T-2C ferroelectric memory cells controlled by WL1-WL8. Due to the destructive readout characteristic of the ferroelectric capacitor, each input data needs to be written into two different 2T-2C cells, so 8 2T-2C ferroelectric memory cells are required. Then write the low-order carry signal (the carry signal of the lowest bit is 0) into the 2T-2C ferroelectric memory cells controlled by WL9 and WL10. Then calculate in sequence to get X0 = Majority(A0, B0, 0), X0', S0 = Majority(A0, B0, 0, X0', X0'), X1 = Majority(A1, B1, X0)(S2), X1', S1 = Majority(A1, B1, X0, X1', X1'). Finally, perform a read operation on the 2T-2C ferroelectric memory cells controlled by WL11, WL12, and WL13 in sequence to read out the calculation result "010". The simulation waveform diagram is as Figure 11 shown.
[0050] The embodiments described above only represent the implementation manners of the present invention, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A method for implementing a full adder based on 2T-2C FRAM in-memory operation, characterized in that: Full-addition calculation is performed in a ferroelectric memory cell, and the basic components of the circuit include a 2T-2C ferroelectric memory cell, a 4T-2C ferroelectric cell and a sensitive amplifier.
2. The method for implementing a full adder based on 2T-2C FRAM in-memory operation according to claim 1, characterized in that: When performing full-addition calculation, each bit line contains n ferroelectric memory cells and corresponding sensitive amplifiers. When writing data, multiple word lines (WL) can be opened at one time for writing; when reading data, only one word line can be opened at one time for reading data; when performing calculation, an odd number of ferroelectric memory cells must be turned on at the same time for calculation.
3. The method for implementing a full adder based on 2T-2C FRAM in-memory operation according to claim 2, characterized in that: When performing full addition calculation, three word lines will be activated simultaneously for calculation when the carry signal is calculated; and five word lines will be activated simultaneously for calculation when the result bit signal is calculated.
4. The method for implementing a full adder based on 2T-2C FRAM in-memory operation according to claim 2, characterized in that: When performing n-bit full addition calculation, according to C0(C0 , ), S0, C1(C1 , ), S1...Cn-1(Cn-1 , ), Sn-1. The value of Cn-1 is also the value of Sn.
5. The method for implementing a full adder based on 2T-2C FRAM in-memory operation according to claim 2, characterized in that: Input data (addends, augends), intermediate values (carry signals) and results are all stored in 2T-2C or 4T-2C ferroelectric memory cells without any cache.
6. The method for implementing a full adder based on 2T-2C FRAM in-memory operation according to claim 2, characterized in that: The calculation result S can be stored in the FRAM array, and subsequent calculations can directly call the existing calculation result S, which can realize the full addition calculation of multiple consecutive numbers such as A+B+C...