Method for implementing in-memory computing with a small-signal ferroelectric capacitor based on non-destructive readout

Through the in-memory calculation method based on small signal ferroelectric capacitors, the problems of high hardware cost and high energy consumption in the existing technology are solved, and signed MAC operation and CAM search with low hardware cost and high energy efficiency are realized, which improves the energy efficiency and linearity of the computing system.

CN115565574BActive Publication Date: 2025-07-18PEKING UNIV

Patent Information

Application Number
CN202211206511.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2025-07-18
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

In the prior art, artificial neural networks and content addressable memory based on nonvolatile memory have problems such as high hardware cost, high energy consumption and high DC power consumption, making it difficult to realize signed MAC operation and CAM search with low hardware cost and high energy efficiency.

Method used

A small signal ferroelectric capacitor based on non-destructive read is used to realize in-memory calculation by applying small signal voltage pulses, and an in-memory calculation unit is formed using a single ferroelectric capacitor FeCap to realize the functions of the CAM unit and signed local multiplication operation. A linear distance measurement and signed MAC operation are achieved by combining ferroelectric capacitor cross arrays and operational amplifiers.

Benefits of technology

It reduces hardware cost and energy consumption, achieves high computing energy efficiency, avoids DC power consumption and "IR drop" problems, improves the calculation linearity and detection range, and has the advantages of machine learning tasks at the edge end and neural network hardware acceleration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115565574B_ABST
    Figure CN115565574B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for realizing in-memory computing based on a small-signal ferroelectric capacitor with non-destructive readout, belonging to the field of novel computing architectures. The present invention realizes the linearly inseparable comparison operation of a content-addressable memory (CAM) unit and the local multiplication operation of a signed synaptic unit on a ferroelectric capacitor, reducing the hardware cost. Compared with the method of distance measurement based on the traditional CAM architecture, it has higher computational linearity and detection range, as well as lower search power consumption; at the same time, compared with the resistive synaptic array, it has no DC power consumption and no potential DC conduction path, avoiding the "IR drop" problem of large-scale arrays, and has lower computational power consumption, showing significant advantages in edge machine learning tasks based on feature retrieval and neural network hardware acceleration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of novel computing architectures, and particularly to a method for realizing in-memory computing based on a small-signal ferroelectric capacitor with non-destructive readout. Background Art

[0002] With the rapid development of artificial intelligence and big data technologies, the bottlenecks of the traditional von Neumann computing architecture with separate memory and computing have become increasingly prominent. The transmission of data between storage units and computing units will cause a large amount of delay and energy consumption waste, and the computing system has higher requirements for speed and energy efficiency. Inspired by the human brain's operation mode, researchers have proposed in-memory computing architectures, constructing a distributed computing network that integrates memory and computing and is highly parallel. While improving the processing efficiency of complex data, it can avoid the delay and energy consumption problems caused by data transfer in the traditional von Neumann computing architecture. In a classic artificial neural network, the input feature vector and the weight matrix perform a vector-matrix multiplication to generate an output vector, and then the output activation vector is obtained through an activation function. Among them, the multiply-accumulate (MAC) operation occupies most of the computing resources, and the artificial neural network shows excellent performance in many fields such as pattern recognition, natural language processing, and automatic control. In addition, the content-addressable memory (CAM) is a new in-memory computing paradigm that can complete the matching operation of the input vector (query) and all stored vectors (entry) within one search cycle, and perform feature retrieval based on distance measurement according to the degree of mismatch, which is extremely attractive in processing efficient machine learning models such as in-memory computing based on feature retrieval.

[0003] The in-memory computing architecture ultimately needs to be fully hardwareized to completely break away from the limitation of the "memory wall" bottleneck. Artificial neural networks are mainly used to accelerate MAC operations, and CAMs are mainly used to accelerate search operations. The artificial neural networks and CAMs built based on traditional CMOS circuits require a large hardware overhead and bring high energy consumption. Researchers have constructed artificial neural networks and CAM units with synapses as the core based on various non-volatile memory devices, such as resistive random access memory (RRAM), phase change memory (PCM), ferroelectric field-effect transistor (FeFET), etc., which have reduced hardware overhead and improved the energy efficiency of computing.

[0004] However, currently, artificial neural networks based on non-volatile memories mainly represent different weight magnitudes by modulating the resistance states of storage elements, and then apply an input voltage for MAC operations. When multiple rows are simultaneously enabled, such a resistive-based synaptic array usually has a relatively high DC current, and a pair of complementary resistive storage elements need to be used as a signed synaptic weight unit. By performing differential processing on the result of matrix-vector multiplication, signed MAC operations are achieved, resulting in problems of high hardware cost and large energy consumption. In addition, the current CAM design can achieve a relatively linear distance metric by directly detecting the match line (ML) current linearly related to the number of mismatched cells, but it will inevitably form a DC path and generate a large DC power consumption. When using the method of pre-charging the ML and then searching, it is necessary to detect the discharge speed of the ML voltage related to the degree of mismatch through complex timing control and detection circuits for distance measurement. Although it avoids a large DC power consumption, it has poor computational linearity and a limited detection range. In summary, a hardware implementation of an energy-efficient signed MAC operation and CAM with both low hardware cost and no DC power consumption is of great significance. Summary of the Invention

[0005] In view of the above problems existing in the prior art, the present invention proposes a method for in-memory computing implemented by a small-signal ferroelectric capacitor based on non-destructive readout, which can realize CAM search and MAC operations, and has extremely low hardware cost and high computing energy efficiency.

[0006] The technical solution of the present invention is as follows:

[0007] A method for in-memory computing implemented by a small-signal ferroelectric capacitor based on non-destructive readout, characterized in that the in-memory computing unit is composed of a single ferroelectric capacitor FeCap, and the in-memory computing array is composed of a ferroelectric capacitor FeCap cross array. One end SL of the FeCap is used for programming and input, and the other end is connected to the ground potential. This in-memory computing unit can realize the functions of a CAM unit and signed local multiplication operations, and its operations are as follows:

[0008] 1) In the stage of programming the FeCap to store an entry, a set / reset write pulse is applied to the SL to program the initial polarization state of the FeCap into positive / negative remanent polarization, representing stored entry 1 and entry 0 respectively. In the search stage, when biased at V R1 / V R0Apply small-signal input pulses to the SL, representing search queries 1 and 0 respectively. Only when the input query matches the stored entry, the FeCap is in a low capacitance state (LCS); otherwise, the FeCap is in a high capacitance state (HCS). During the application of small-signal search pulses, FeCaps in different capacitance states generate different amounts of charge, representing the CAM matching result of an in-memory computing unit.

[0009] 2) During the stage of programming the weight values stored in the FeCap, apply a programming voltage pulse higher than the ferroelectric coercive voltage to the SL terminal to change the polarization state of the ferroelectric and thus program the stored weight. It has different small-signal capacitance values under a DC bias. When performing a local multiplication operation, apply small-signal input pulses to the SL terminal biased at V +1 / V -1 , representing +1 or -1 as the input respectively. FeCaps in different capacitance states generate different amounts of charge, representing the result of signed local multiplication of an in-memory computing unit.

[0010] This invention is different from the destructive readout of traditional ferroelectric random access memory (FeRAM), that is, reading information by applying a large voltage according to the change in ferroelectric polarization. In this invention, the small-signal capacitance (C FE ) of the FeCap is read under a certain DC voltage bias. Since no penetrating ferroelectric domains are formed during the application of this DC voltage bias, and the DC voltage bias is removed after reading the information, the FeCap will return to its original polarization state, realizing non-destructive readout of the stored information.

[0011] Furthermore, the in-memory computing array is composed of a ferroelectric capacitor FeCap cross array. Each row of the cross array shares a charge integration circuit composed of an operational amplifier and an integration capacitor (C ref ), and each column of the cross array shares one SL. This in-memory computing array can implement the search of the CAM array and the MAC operation, and its operations are as follows:

[0012] 1) When performing a CAM array search, apply small-signal voltages ΔV SL corresponding to the input query to all SLs simultaneously. The output terminal of the ferroelectric capacitor cross array is ML. The ML of each row will obtain an output voltage ΔV ML as shown in formula (1) according to the matching situation between the input query and the stored entry in that row, where N represents the total number of in-memory computing units in a row, and n represents the number of non-matching in-memory computing units in that row.

[0013]

[0014] As can be seen from Equation (1), the voltage of ML is linearly related to the number n of cells that do not match this row. When the CAM is used for distance measurement, n represents the Hamming distance between the input vector and the stored vector. The in-memory computing architecture proposed by the present invention can achieve linear distance measurement; a reference row is further set, and the input query and the stored entry of this row are completely matched. The final output result is the voltage of ML minus the voltage of the reference row, as shown in Equation (2), enabling more detectable mismatch numbers to be allowed within a limited voltage range and achieving a larger range of distance measurement.

[0015]

[0016] 2) When performing the MAC operation, small signal voltages ΔV corresponding to the input vector are simultaneously applied to all SLs SL , and the output end of the ferroelectric capacitor cross array is the result V of the signed MAC operation Out , which is the accumulation of the charges obtained by local multiplication of all synaptic units in each row through the charge integration circuit, as shown in Equation (3). A signed vector matrix multiplication operation is realized in the ferroelectric capacitor cross array;

[0017]

[0018] The in-memory computing design of the small signal ferroelectric capacitor based on non-destructive readout proposed by the present invention, where the ferroelectric material can be HfO2 doped with Zr (HZO), HfO2 doped with Al (HfAlO), or other types of multi-domain ferroelectric materials.

[0019] The technical effects of the present invention are as follows:

[0020] 1. The present invention's CAM and signed MAC operation schemes based on ferroelectric capacitors adopt a new non-destructive readout scheme, that is, under a non-destructive DC voltage bias, the search and MAC operations are completed by applying small signal voltage pulses. Compared with the traditional FeRAM's destructive readout method of reading the polarization change amount, there is no need to perform a rewrite operation after each read operation, greatly reducing the energy consumption and operation complexity.

[0021] 2. The in-memory computing unit of the present invention is based on a single FeCap, and utilizes the non-monotonic characteristics of the small signal ferroelectric capacitor. Compared with the traditional method that requires two complementary branch paths to implement the CAM unit and differential operation of two complementary storage elements to achieve signed weights, it reduces the hardware cost and improves the computing energy efficiency.

[0022] 3. In the present invention, by accumulating all CAM matching results in a row during the search process, a matching result linearly related to the number of non-matching units in that row is obtained, enabling linear distance measurement. Moreover, the search method based on small-signal voltage pulses has no DC power consumption, and has significant advantages compared with the method of directly detecting the matching current or detecting the ML discharge speed based on the traditional CAM architecture for distance measurement.

[0023] 4. The present invention uses small-signal pulses as input and utilizes the charge transfer principle to achieve the accumulation operation. Compared with the resistive synaptic array, it has no DC power consumption and has no potential direct current path, avoiding the "IR drop" problem of large-scale arrays. In addition, it has the potential for three-dimensional integration on top of the front-end CMOS circuit, which can further improve the synaptic storage density of the artificial neural network.

[0024] 5. The hafnium oxide-based FeCap adopted in the present invention has high CMOS process compatibility and good scaling ability, and has high durability, retention characteristics and consistency, enabling it to have the ability to implement large-scale artificial neural networks and higher reliability.

[0025] 6. The present invention has the potential for three-dimensional integration on top of the front-end CMOS circuit, which can further improve the storage density of the CAM. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 FIG. is a schematic diagram of the CAM search function implemented based on small-signal ferroelectric capacitors in the present invention;

[0027] Figure 2 FIG. is a schematic diagram of the MAC operation implemented based on small-signal ferroelectric capacitors in the present invention;

[0028] In the figure: 1 - Schematic diagram of the implementation principle of the binary signed synaptic weight unit based on FeCap; 2 - Schematic diagram of the implementation principle of the multi-valued signed synaptic weight unit based on FeCap

[0029] Figure 3 FIG. is a schematic structural diagram of the CAM search based on the ferroelectric capacitor cross array in the present invention;

[0030] Figure 4 FIG. is a schematic structural diagram of the MAC operation based on the ferroelectric capacitor cross array in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The following further clearly and completely elaborates the present invention with reference to the accompanying drawings and through specific embodiments.

[0032] The principle of the CAM search implemented based on small-signal ferroelectric capacitors in the present invention is as Figure 1As shown, one end of the FeCap serves as the SL end of the in-memory computing unit, and the other end is connected to the ground potential. The FeCap is programmed to +P by applying a set / reset write pulse to the SL end. r / -P r which represent stored entry 1 and entry 0 respectively, and they exhibit small-signal capacitances of LCS1 / HCS0 and HCS1 / LCS0 respectively under the DC biases of V R1 and V R0 ; during the search operation, by applying a positive small-signal pulse to the SL end biased at V R1 / V R0 to represent input query 1 and query 0, if the input query matches the stored entry, the FeCap is in LCS under this search bias and less charge will be generated during the application of the small-signal pulse, while if the input query does not match the stored entry, more charge corresponding to HCS will be generated during the application of the small-signal pulse, thus realizing the linearly inseparable comparison operation of the CAM based on a single FeCap.

[0033] The principle of realizing the MAC operation based on the small-signal ferroelectric capacitor in the present invention is as Figure 2 shown. One end of the FeCap serves as the SL end of the in-memory computing unit for programming and input, and the other end is connected to the ground potential. When the synapse has binary weights (+1, -1), the FeCap is programmed to positive and negative remanent polarizations (+P r / -P r ) by applying a set / reset write pulse to the SL end, which represent weights +1 and -1 respectively, and they exhibit small-signal capacitances of C +1 / C -1 and C +1H / C -1L and C +1L / C -1H respectively under the DC biases of V +1 / V -1 ; during input, by applying a positive small-signal pulse to the SL end biased at V +1 / V -1 to represent input +1 and -1, when x·w = +1, the FeCap is in C +1H or C -1H , and more charge will be generated during the application of the small-signal pulse. When x·w = -1, the FeCap is in C +1L or C -1L, less charge will be generated, and the calculation results are shown in the truth table. Further, using the ferroelectric multi-domain theory, the ferroelectric polarization state can be continuously modulated by applying different programming voltage pulses, and then multi-value weights can be stored. In this embodiment, taking the 2-bit synaptic weight as an example, after applying a reset pulse at the SL end, programming voltage pulses with different amplitudes are applied to program the FeCap into four polarization states, representing stored weights +2, +1, -1, and -2 respectively; during input, according to the four calculation results of x·w, the FeCap will be in small-signal capacitance states of different sizes and obtain corresponding charge responses under the input pulse, and the calculation results are shown in the truth table.

[0034] Figure 3 Each row of the ferroelectric capacitor cross array of the present invention is connected to the positive input terminal of the operational amplifier, and the integrating capacitor C ref is connected across the positive input terminal and the output terminal of the operational amplifier, and the negative input terminal of the operational amplifier is grounded. Using the "virtual short" characteristic of the operational amplifier, each row of the cross array is equipotential with the ground potential; each column of the ferroelectric capacitor cross array shares the SL, that is, the entire CAM search operation is completed in parallel. During the search, search pulses corresponding to the input query vector are simultaneously applied to all the SLs. Each CAM unit generates different amounts of charge according to the matching situation, and the total charge of each row is integrated through C ref Finally, the matching result of this row is obtained at the output terminal ML. A CAM array of the same length is set as a reference row. The input query vector of this row is completely matched with the stored entry vector, then the total accumulated charge of this row is the total charge generated by N FeCaps in the LCS under the small-signal search pulse. The final output result is the voltage of ML minus the voltage of the reference row. For each unmatched unit in a row of the CAM array, there will be a FeCap in the HCS in this row, and the finally obtained output voltage is proportional to the number of unmatched units. Based on the present invention, linear distance measurement can be realized, and its detection margin depends on the difference between the HCS and the LCS.

[0035] Figure 4 Each row of the ferroelectric capacitor cross array based on the present invention is connected to the positive input terminal of the operational amplifier, and the integrating capacitor C ref is connected across the positive input terminal and the output terminal of the operational amplifier, and the negative input terminal of the operational amplifier is grounded. During input, small-signal pulses corresponding to the input feature vector are simultaneously applied to all the SLs. Each in-memory computing unit generates different amounts of charge according to the local multiplication result, and the total charge of each row is integrated through C ref Finally, the accumulated result of the local multiplication of this row is obtained at the output terminal, completing the signed MAC operation.

[0036] This embodiment fully and elaborately expounds the in-memory computing design of a small-signal ferroelectric capacitor based on non-destructive readout. It realizes the CAM comparison operation and the local multiplication operation of signed synapses on a single FeCap, and also realizes the CAM search for linear distance measurement and the signed MAC operation. Compared with the method of distance measurement based on the traditional CAM architecture, it has higher computational linearity and detection range, and reduces the search power consumption. At the same time, compared with the resistive synaptic array, it has no DC power consumption and "IR drop" problems, and has lower computational power consumption.

[0037] Finally, it should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art can understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention shall be subject to the scope defined by the claims.

Claims

1. A method for realizing in-memory computing with a small-signal ferroelectric capacitor based on non-destructive readout, characterized in that, The in-memory computing unit is composed of a single ferroelectric capacitor FeCap, and the in-memory computing array is composed of a cross array of ferroelectric capacitors FeCap. One end SL of the ferroelectric capacitor FeCap is used for programming and input, and the other end is connected to the ground potential. This in-memory computing unit realizes the function of the CAM unit and signed local multiplication operation, and its operation is as follows: 1) During the programming of the FeCap storage entry, a set / reset write pulse is applied to the SL to program the initial polarization state of the FeCap into positive / negative remanent polarization, representing storage entry 1 and entry 0 respectively. During the search phase, a small-signal input pulse is applied to the SL biased at V R1 / V R0 representing search query 1 and query 0 respectively. Only when the input query matches the stored entry, the FeCap is in a low capacitance state; otherwise, the FeCap is in a high capacitance state. During the application of the small-signal search pulse, FeCaps in different capacitance states generate different amounts of charge, representing the CAM matching result of an in-memory computing unit; 2) During the stage of programming the weight values stored in FeCap, a programming voltage pulse higher than the ferroelectric coercive voltage is applied to the SL terminal to change the polarization state of the ferroelectric and thus program the stored weights. It has different small-signal capacitance magnitudes under an input DC bias. When performing a local multiplication operation, a small-signal input pulse is applied to the SL terminal biased at V +1 / V -1 , which respectively represents an input of +1 or -1. FeCaps in different capacitance states will generate different amounts of charge, representing the result of signed local multiplication of a memory computing unit.

2. The method for realizing in-memory computing based on a small-signal ferroelectric capacitor with non-destructive readout according to claim 1, wherein Each row of the ferroelectric capacitor FeCap cross array shares a charge integration circuit composed of an operational amplifier and an integration capacitor, and each column of this cross array shares an SL to realize the search of the CAM array and the MAC operation. Its operation is as follows: 1) By accumulating all the CAM matching results in a row during the search process, a matching result linearly related to the number of non-matching units in this row is obtained, and the search operation of the CAM array, that is, the distance metric, is completed; 2) By accumulating the results of signed local multiplication in each row, a signed MAC operation is realized.

3. The method for realizing in-memory computing with a small-signal ferroelectric capacitor based on non-destructive readout according to claim 1, wherein The ferroelectric material of the ferroelectric capacitor FeCap is HfO2 doped with Zr, HfO2 doped with Al, or other types of multi-domain ferroelectric materials.

Citation Information

Patent Citations

  • Delay modulation method based on ferroelectric transistor

    CN113903378A

  • Method for realizing ternary content addressable memory (TCAM) based on ferroelectric tunneling field effect transistor (FeTFET)

    CN114743578A

Cited By

  • Storage and calculation integrated logic unit and calculation architecture of state definition logic

    CN121935206A