High-throughput parallel ferroelectric storage system, storage device and storage equipment

By introducing charge replication cancellation circuit and integration circuit in high-throughput parallel ferroelectric memory computing system, the problem of the accumulation of output results of the memory cell in the low-capacitance state affecting the calculation accuracy, achieving higher calculation accuracy and system robustness.

CN120183458AInactive Publication Date: 2025-06-20XIDIAN UNIV HANGZHOU RES INST +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510670324.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In a high-throughput parallel ferroelectric memory computing system, the accumulation of output results of the memory cells on the same column in the low-capacitance state will produce the same effect as that of the high-capacitance state, affecting the accuracy of the calculation results.

Method used

The charge replication cancellation circuit is used to connect it to multiple memory units to generate cancellation charges and cancel it with the accumulated charge generated by the multiple column memory units in a low capacitance state. At the same time, the calculation results of the one column memory unit are added and output.

Benefits of technology

The charge replication cancellation circuit eliminates the impact of accumulated charge generated by the calculation unit in the low capacitance state on the calculation results, which improves the calculation accuracy and solves the problem of low calculation accuracy in the ferroelectric memory system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183458A_ABST
    Figure CN120183458A_ABST
Patent Text Reader

Abstract

The invention discloses a high-throughput parallel ferroelectric storage calculation system, a storage calculation device and storage calculation equipment, the high-throughput parallel ferroelectric storage calculation system comprises a plurality of storage calculation units, the plurality of storage calculation units are arranged in a row and column structure, and the plurality of storage calculation units are used for storing output data of external equipment; the charge replication counteracting circuit is connected with the plurality of storage and calculation units and is used for generating counteracting charges and counteracting the counteracting charges with accumulated charges generated by the plurality of storage and calculation units in a low-capacitance state; and the plurality of integrating circuits are connected with the plurality of storage and calculation units, and each integrating circuit is used for adding and outputting the calculation results of one column of storage and calculation units. The invention aims to solve the problems that in a ferroelectric storage and calculation system, the accumulation of output results of storage and calculation units on the same column in a low-capacitance state can generate the same effect as that in a high-capacitance state, so that the calculation results of the storage and calculation units are influenced, and the calculation precision is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of in-memory computing, and particularly to a high-throughput parallel ferroelectric memory-computation system, a memory-computation device, and a memory-computation equipment. Background Art

[0002] With the continuous enhancement of the capabilities of artificial intelligence models, the demand for hardware systems that can support large-scale, high-performance, and efficient computing has also increased. Due to the separation of storage and computing in the traditional von Neumann architecture, the so-called memory wall and power wall problems, namely the so-called von Neumann bottleneck, severely limit the development of artificial intelligence. The emergence of in-memory computing represents a major innovation in the computing architecture. Different from the traditional von Neumann architecture centered on computing, in-memory computing is centered on data storage. By integrating storage and computing, specific computing tasks can be completed internally, effectively reducing the frequent data interaction between the processor and memory.

[0003] Currently, the analog in-memory computing technology based on resistive random-access memory has received extensive attention. However, problems such as high static power consumption, voltage drop, and sneak path leakage have hindered its application in large-scale high-performance systems. Compared with the analog in-memory computing scheme based on resistive random-access memory, the scheme based on ferroelectric capacitors can naturally overcome these disadvantages, and with its advantages such as high integration density and easy backend integration, it has become one of the strong competitors for high-performance in-memory computing systems. However, due to the limitation of the switching ratio, the parallelism and accuracy of the ferroelectric memory-computation array are not high, and the computing efficiency of the inference system is low.

[0004] When storing data in a ferroelectric array, for example, when mapping neural network weights to a ferroelectric array, the cells representing logic 1 are programmed to a high capacitance state, and the cells representing logic 0 are programmed to a low capacitance state. The ratio of their capacitance values is the switching ratio. Since the low capacitance of an actual ferroelectric capacitor cannot reach 0, in actual calculations, the accumulation of the output results of the cells in the low capacitance state on the same column will produce the same effect as the high capacitance state, directly affecting the calculation results of the ferroelectric array and reducing the calculation accuracy. Summary of the Invention

[0005] In view of the above deficiencies of the prior art, the purpose of the present invention is to provide a high-throughput parallel ferroelectric memory-computation system, a memory-computation device, and a memory-computation equipment to solve the problem that in a high-throughput parallel ferroelectric memory-computation system, the accumulation of the output results of the memory-computation units in the low capacitance state on the same column will produce the same effect as the high capacitance state, thereby affecting the calculation results of the memory-computation units and reducing the calculation accuracy.

[0006] The technical solution of the present invention is as follows: A high-throughput parallel ferroelectric memory-computation system includes: Multiple memory - computing units, where multiple of the memory - computing units are arranged in a row - column structure, and the multiple memory - computing units are used to store the output data of an external device; A charge replication cancellation circuit, connected to the multiple memory - computing units, where the charge replication cancellation circuit is used to generate cancellation charges and cancel the cancellation charges against the cumulative charges generated by multiple columns of the memory - computing units in a low - capacitance state; Multiple integration circuits, connected to the multiple memory - computing units, where each integration circuit is used to add up the calculation results of a column of the memory - computing units and then output the result.

[0007] Optionally, the high - throughput parallel ferroelectric memory - computing system further includes: A word line, connected to the control terminals of the memory - computing units in each row; A bit line, connected to the input terminals of the memory - computing units in each column; A plate line, connected to the output terminals of the memory - computing units in each column; A first driver, connected to the word line, where the first driver is used to output a first driving signal to the word line to drive the multiple memory - computing units to work; A second driver, connected to the bit line and the plate line, where the second driver is used to output a second driving signal to the bit line and the plate line to drive the multiple memory - computing units to perform data reading and writing.

[0008] Optionally, the memory - computing unit includes a first transistor and a first capacitor. The gate of the first transistor is connected to the word line, the source of the first transistor is connected to the bit line, the drain of the first transistor is connected to the first end of the first capacitor, and the second end of the first capacitor is connected to the plate line.

[0009] Optionally, the charge replication cancellation circuit includes: A charge cancellation reference column, which is used to generate cancellation reference charges when the multiple memory - computing units are working; An offset charge cancellation circuit, connected to the charge cancellation reference column, where the offset charge cancellation circuit is used to generate corresponding cancellation charges according to the cancellation reference charges generated by the charge cancellation reference column and cancel the cancellation charges against the cumulative charges generated by multiple columns of the memory - computing units in a low - capacitance state.

[0010] Optionally, the charge cancellation reference column includes multiple reference units. The multiple reference units in the charge cancellation reference column are connected to the multiple memory - computing units in a column of the memory - computing units one - to - one. The reference unit includes: a second transistor and a second capacitor. The gate of the second transistor is connected to the word line, the source of the second transistor is connected to the bit line, the drain of the second transistor is connected to the first end of the second capacitor, and the second end of the second capacitor is connected to the plate line.

[0011] Optionally, the offset charge cancellation circuit includes a first operational amplifier, a plurality of first reference capacitors, a second reference capacitor, and a third switch. The positive input terminal of the first operational amplifier is grounded. The negative input terminal of the first operational amplifier and the second terminal of the second reference capacitor are connected to the output terminal of the charge cancellation reference column. The output terminal of the first operational amplifier, the first terminal of the second reference capacitor, and the first terminals of the plurality of first reference capacitors are connected to the first terminal of the third switch. The second terminal of the third switch is grounded. The second terminals of the plurality of first reference capacitors are connected to the output terminals of multiple columns of the memory and computing units one-to-one.

[0012] Optionally, the integration circuit includes a second operational amplifier, a first switch, and a third capacitor. The first terminal of the first switch and the positive input terminal of the second operational amplifier are grounded. The second terminal of the first switch, the output terminal of the second operational amplifier, and the first terminal of the third capacitor are connected to each other. The negative input terminal of the second operational amplifier and the second terminal of the third capacitor are connected to the output terminal of one column of the memory and computing units.

[0013] Optionally, the integration circuit further includes: a second switch, the first terminal of the second switch is connected to the output terminal of the second operational amplifier, and the second terminal of the second switch is connected to the first terminal of the third capacitor.

[0014] The present invention also provides a memory and computing device, including the high-throughput parallel ferroelectric memory and computing system as described above.

[0015] The present invention also provides a memory and computing equipment, including the memory and computing device as described above.

[0016] The technical solution of the present invention constitutes a high-throughput parallel ferroelectric computing-in-memory system through multiple computing-in-memory units, a charge replication cancellation circuit, and multiple integration circuits. Among them, the multiple computing-in-memory units are arranged in a row-column structure, and the multiple computing-in-memory units can store the output data of external devices; the charge replication cancellation circuit is connected to the multiple computing-in-memory units, and the charge replication cancellation circuit can generate cancellation charges when the multiple computing-in-memory units are working, and make the cancellation charges cancel the cumulative charges generated by the multiple columns of computing-in-memory units in the low-capacitance state; the multiple integration circuits are connected to the multiple computing-in-memory units, and each integration circuit can add up the calculation results of a column of computing-in-memory units and then output; thus, the present invention can generate cancellation charges equal to the cumulative charges generated by the computing-in-memory units in the low-capacitance state when the multiple computing-in-memory units are working through the charge replication cancellation circuit, so that the cancellation charges can cancel the cumulative charges generated by the computing-in-memory units, thereby eliminating the influence of the cumulative charges generated by the computing-in-memory units in the low-capacitance state during the actual working process on the calculation results; it solves the problem that in the ferroelectric computing-in-memory system, the accumulation of the output results of the computing-in-memory units on the same column in the low-capacitance state will produce the same effect as that in the high-capacitance state, thus affecting the calculation results of the computing-in-memory units and reducing the calculation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0018] Figure 1 It is a schematic circuit structure diagram of an embodiment of the high-throughput parallel ferroelectric computing-in-memory system of the present invention.

[0019] Figure 2 It is a schematic circuit structure diagram of another embodiment of the high-throughput parallel ferroelectric computing-in-memory system of the present invention.

[0020] Figure 3 It is a schematic diagram of the simulation result of the traditional array scheme.

[0021] Figure 4 It is a schematic diagram of the simulation result of the high-throughput parallel ferroelectric computing-in-memory system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the purpose, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0023] In the embodiments and the claims, unless otherwise specifically defined in the context, the articles "a", "an", "the" and "said" may also include the plural forms. If descriptions such as "first", "second", etc. are involved in the embodiments of the present invention, such descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second" may explicitly or implicitly include at least one such feature.

[0024] It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.

[0025] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood as having a meaning consistent with the meaning in the context of the prior art and will not be interpreted with an idealized or overly formal meaning unless specifically defined as herein.

[0026] In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0027] With the continuous enhancement of the capabilities of artificial intelligence models, the demand for hardware systems that can support large-scale, high-performance, and efficient computing has also increased. Due to the separation of storage and computing in the traditional von Neumann architecture, the so-called memory wall and power wall problems, namely the so-called von Neumann bottleneck, have severely restricted the development of artificial intelligence. The emergence of in-memory computing represents a major innovation in the computing architecture. Different from the traditional von Neumann architecture centered on computing, in-memory computing is centered on data storage. By integrating storage and computing, specific computing tasks can be completed internally, effectively reducing the frequent data interaction between the processor and memory.

[0028] At present, the analog in-memory computing technology based on resistive random access memory has received extensive attention. However, problems such as high static power consumption, voltage drop, and sneak path leakage have hindered its application in large-scale high-performance systems. Compared with the analog in-memory computing scheme based on resistive random access memory, the scheme based on ferroelectric capacitors can naturally overcome these disadvantages, and with its advantages such as high integration density and easy back-end integration, it has become one of the strong competitors for high-performance in-memory computing systems. However, due to the limitation of the switching ratio, the parallelism and accuracy of ferroelectric memory-compute arrays are not high, and the computing efficiency of the inference system is low.

[0029] When storing data in the ferroelectric array, for example, when mapping neural network weights to the ferroelectric array, the cells representing logic 1 are programmed to a high capacitance state, while the cells representing logic 0 are programmed to a low capacitance state, and the ratio of their capacitance values is the switching ratio. Since the low capacitance of the actual ferroelectric capacitor cannot reach 0, in actual calculations, the accumulation of the output results of the cells in the low capacitance state in the same column will produce the same effect as the high capacitance state, directly affecting the calculation results of the ferroelectric array and reducing the calculation accuracy. This effect becomes particularly obvious when the array parallelism is high due to the increase in the number of cells in the low capacitance state, limiting the high-throughput parallel calculation of the ferroelectric memory-compute array.

[0030] To solve the above problems, the present invention proposes a high-throughput parallel ferroelectric memory-compute system.

[0031] Refer to Figure 1 , in one embodiment, the high-throughput parallel ferroelectric memory-compute system includes: A plurality of memory-compute units 10, the plurality of memory-compute units 10 are arranged in a row-column structure, and the plurality of memory-compute units 10 are used to store the output data of external devices; A charge replication cancellation circuit 20, connected to the plurality of memory-compute units 10, the charge replication cancellation circuit 20 is used to generate cancellation charges and cancel the cancellation charges with the cumulative charges generated by the plurality of columns of memory-compute units 10 in the low capacitance state; A plurality of integration circuits 30, connected to the plurality of memory-compute units 10, and each integration circuit 30 is used to add up the calculation results of one column of memory-compute units 10 and then output.

[0032] In this embodiment, the memory - computing unit 10 can be composed of a combination of a transistor and a capacitor. The transistor can be used as a switching element to control the voltage application to the capacitor and the data reading process. By controlling the on - off state of the transistor, the operation of the memory - computing unit 10 can be precisely controlled to achieve selective writing and reading of data. Specifically, an external electric field can be applied through an electrode to control the polarization state of the capacitor, which is also the channel for current input and output, realizing data writing, reading, and computing operations. In the row - column structure of the memory - computing unit 10, the electrodes can be word lines and bit lines, etc. Arranging multiple memory - computing units 10 in a row - column structure to form a matrix structure can effectively utilize the chip area, making the arrangement of the high - throughput parallel ferroelectric memory - computing system more compact, thereby increasing the storage density. And through the row - column structure, a specific memory cell can be accessed by using row selection and column selection methods. This method simplifies the complexity of address decoding and access control and improves the access efficiency. The row - column structure also has advantages such as reducing power consumption, increasing the speed of data reading and writing, enhancing scalability, and increasing anti - interference ability. The memory - computing unit 10 can store the output data of external devices, which can be the storage of binary data, the storage of neural network weight values, or the storage of other analog signals.

[0033] The charge replication cancellation circuit 20 can generate cancellation charges to cancel the accumulated charges generated by the memory - computing unit 10 in the low - capacitance state when the memory - computing unit 10 is working. It can be understood that the memory - computing unit 10 in the low - capacitance state generates a kind of polarized charge. Taking negative charge as an example, in this embodiment, the charge replication cancellation circuit 20 can generate an equal amount of positive charge to cancel the negative charge generated by the memory - computing unit 10 in the low - capacitance state. In this way, the influence of the charge generated by the memory - computing unit 10 in the low - capacitance state during the actual working process on the calculation result can be eliminated. Specifically, the charge replication cancellation circuit 20 can adopt a design opposite to that of the memory - computing unit 10. For example, a capacitor is set in the charge replication cancellation circuit 20 to be connected to the capacitor in the memory - computing unit 10. When the connection end between the capacitor in the memory - computing unit 10 and the capacitor in the charge replication cancellation circuit 20 outputs negative charge, the capacitor in the charge replication cancellation circuit 20 is controlled to output an equal amount of positive charge, thus realizing charge cancellation.

[0034] The integration circuit 30 can adopt an operational amplifier and a capacitor to form a trans - impedance integrator. The polarization state of each memory - computing unit 10 is encoded as the weight value required for calculation. The input voltage is applied to the memory - computing unit 10 through the word line. The output currents of all the memory - computing units 10 on the same column (bit line) will be naturally superimposed on the plate line to form a total current. The integration circuit 30 can convert the total current into a voltage signal through the operational amplifier and the capacitor for output, so as to facilitate subsequent circuit processing, such as analog - to - digital conversion, further digital computing, etc.

[0035] The technical solution of the present invention constitutes a high-throughput parallel ferroelectric memory and computing system through multiple memory and computing units 10, a charge replication cancellation circuit 20, and multiple integration circuits 30. Among them, the multiple memory and computing units 10 are arranged in a row-column structure, and the multiple memory and computing units 10 can store the output data of external devices; the charge replication cancellation circuit 20 is connected to the multiple memory and computing units 10, and the charge replication cancellation circuit 20 can generate cancellation charges when the multiple memory and computing units 10 are working, and make the cancellation charges cancel the accumulated charges generated by the multiple columns of memory and computing units 10 in the low-capacitance state; the multiple integration circuits 30 are connected to the multiple memory and computing units 10, and each integration circuit 30 can add up the calculation results of a column of memory and computing units 10 and then output; thus, the present invention can generate cancellation charges equal to the accumulated charges generated by the memory and computing units 10 in the low-capacitance state when the charge replication cancellation circuit 20 is working, so that the cancellation charges can cancel the accumulated charges generated by the memory and computing units 10, thereby eliminating the influence of the accumulated charges generated by the memory and computing units 10 in the low-capacitance state during the actual working process on the calculation results; it solves the problem that in the ferroelectric memory and computing system, the accumulation of the output results of the memory and computing units 10 in the same column in the low-capacitance state will produce the same effect as that in the high-capacitance state, thereby affecting the calculation results of the memory and computing units 10 and reducing the calculation accuracy.

[0036] Referring to Figure 2 , in one embodiment, the high-throughput parallel ferroelectric memory and computing system further includes: Word lines, connected to the controlled ends of each row of the memory and computing units 10; Bit lines, connected to the input ends of each column of the memory and computing units 10; Plate lines, connected to the output ends of each column of the memory and computing units 10; A first driver, connected to the word lines, and the first driver is used to output a first driving signal to the word lines to drive the multiple memory and computing units 10 to work; A second driver, connected to the bit lines and the plate lines, and the second driver is used to output a second driving signal to the bit lines and the plate lines to drive the multiple memory and computing units 10 to perform data reading and writing.

[0037] In this embodiment, the word line, bit line, and plate line are key electrode structures for realizing data storage and reading. The word line can be used to select a specific storage row. By applying a voltage, i.e., a driving signal, to the word line through a first driver, all the memory and computing units 10 in that row can be activated, enabling them to perform read or write operations. Specifically, access control to the memory and computing units 10 can be achieved by controlling the switching element through the word line, such as turning on and off the transistor in the memory and computing unit 10. When reading data, the activation of the word line enables the memory and computing units 10 in the corresponding row to be connected to the bit line, thereby transmitting the stored data to the bit line. When writing data, the activation of the word line allows programming of the memory and computing units 10 in that row. The bit line is used to transmit the data of the memory and computing units 10, and the memory and computing units 10 calculate the data transmitted by the bit line and store the calculation result in the capacitor of the memory and computing units 10. In a read operation, the bit line receives the signal from the memory and computing units 10 in the activated row, reflecting the stored data state. In a write operation, the bit line provides the data voltage signal to be written, allowing the memory and computing units 10 to update their states. The plate line is generally used to provide a reference voltage or electric field to help control the state of the memory and computing units 10. It can be used to enhance the stability and performance of the storage unit. In a high-throughput parallel ferroelectric memory and computing system, the plate line can be used to apply a constant electric field to help the memory and computing units 10 maintain their polarization states during write and read processes. Also, while providing charge cancellation, the plate line can also perform an output function, outputting a voltage. By controlling the voltage of the plate line, the read and write characteristics of the storage unit can be optimized. Additionally, the word line can be connected to a decoder, which can select a specific output line for connection to other circuits, such as the memory and computing units 10. Thus, by combining the decoder and the driver, the decoder selects a specific memory and computing unit 10 according to the input address signal, and the driver is responsible for transmitting the control signal to the selected memory and computing unit 10 for read and write operations. In this embodiment, the first driver and the second driver can output different voltage signals, i.e., the first driving signal and the second driving signal, to the word line, bit line, and plate line, thereby controlling whether the memory and computing units 10 are activated and controlling the read and write working states of the memory and computing units 10. The specific voltage signals output by the first driver and the second driver can be selected according to the actual situation and user requirements.

[0038] Referring to Figure 1 and Figure 2 , in one embodiment, the memory and computing unit 10 includes a first transistor Q1 and a first capacitor C1. The gate of the first transistor Q1 is connected to the word line, the source of the first transistor Q1 is connected to the bit line, the drain of the first transistor Q1 is connected to the first end of the first capacitor C1, and the second end of the first capacitor C1 is connected to the plate line.

[0039] In this embodiment, the storage and calculation unit 10 can be composed of a first transistor Q1 and a first capacitor C1. The first transistor Q1 can be an NMOS tube. In this way, when a write operation is required for a certain storage and calculation unit 10, the corresponding word line is activated (usually at a high level), which can make all the NMOS tubes in the row turned on, allowing the signal on the bit line to enter the storage unit. For the unselected storage and calculation units 10, the corresponding word lines remain at a low level. This puts the NMOS tubes of these units in a cut-off state, thereby isolating the connection between these units and the bit lines. Therefore, the use of NMOS tubes in this embodiment can shield the unselected units to avoid write interference affecting the robustness of the high-throughput parallel ferroelectric storage and calculation system.

[0040] In one embodiment, the charge replication cancellation circuit 20 includes: A charge cancellation reference column, wherein the charge cancellation reference column is used to generate a cancellation reference charge when the plurality of storage and calculation units 10 are working; The offset charge cancellation circuit is connected to the charge cancellation reference column. The offset charge cancellation circuit is used to generate corresponding cancellation charges based on the cancellation reference charges generated by the charge cancellation reference column, and to cancel the cancellation charges with the accumulated charges generated by multiple columns of the storage and computing units 10 in a low capacitance state.

[0041] In this embodiment, the charge replication cancellation circuit 20 can be composed of a charge cancellation reference column and an offset charge cancellation circuit. The charge cancellation reference column can specifically adopt the same transistor and capacitor design as the storage and calculation unit 10, and is used to generate a cancellation reference charge with the same charge amount as the storage and calculation unit 10. The offset charge cancellation circuit can adopt a design opposite to the storage and calculation unit 10. For example, a capacitor is set in the offset charge cancellation circuit to be connected to the capacitor in the storage and calculation unit 10. When the connection end between the capacitor of the storage and calculation unit 10 and the capacitor in the offset charge cancellation circuit is in the stage of outputting negative charge, the capacitor in the offset charge cancellation circuit is controlled to output an equal amount of positive charge. In this way, the offset charge cancellation circuit in this embodiment can generate corresponding cancellation charges according to the cancellation reference charge generated by the charge cancellation reference column, and cancel the cancellation charge with the accumulated charge generated by the multiple columns of storage and calculation units 10 in the low capacitance state. In this way, the accumulated charge generated by the storage and calculation unit 10 in the low capacitance state will not affect the calculation result of the storage and calculation unit 10, thereby improving the accuracy of the calculation.

[0042] Reference Figure 1 and Figure 2In one embodiment, the charge cancellation reference column includes a plurality of reference units, and the plurality of reference units in the charge cancellation reference column are connected one-to-one with the plurality of storage and calculation units 10 in a column of the storage and calculation units 10; the reference unit includes: a second transistor Q2 and a second capacitor C2, the gate of the second transistor Q2 is connected to the word line, the source of the second transistor Q2 is connected to the bit line, the drain of the second transistor Q2 is connected to the first end of the second capacitor C2, and the second end of the second capacitor C2 is connected to the plate line.

[0043] In this embodiment, a reference unit in the charge cancellation reference column can adopt the same circuit structure as the storage and calculation unit 10, specifically composed of a second transistor Q2 and a second capacitor C2; in this way, the charge cancellation reference column can generate a cancellation reference charge with the same charge amount as the accumulated charge generated by the storage and calculation unit 10 in a low-capacitance state, and then the subsequent offset charge cancellation circuit generates a corresponding cancellation charge based on the cancellation reference charge generated by the charge cancellation reference column. The specific principle can refer to the storage and calculation unit 10, and will not be repeated here.

[0044] Furthermore, in one embodiment, the charge cancellation reference column can be set in the middle of multiple storage and computing units 10; considering the impact of global chip errors, placing the charge cancellation reference column in the middle of multiple storage and computing units 10 can make the distance between it and each storage and computing unit 10 relatively balanced, thereby reducing the signal delay or attenuation difference caused by different physical distances, and reducing the impact on system performance; it should be noted that the specific position of the charge cancellation reference column in this embodiment is not limited to being set in the middle of multiple storage and computing units 10. For example, when the number of columns of storage and computing units 10 is an odd number, such as 3 columns of storage and computing units 10, it can be set in the front column or the back column of the second column of storage and computing units 10 to ensure that the distance between the charge cancellation reference column and each storage and computing unit 10 is relatively balanced. For details, please refer to Figure 2 .

[0045] Reference Figure 1 and Figure 2 In one embodiment, the offset charge cancellation circuit includes a first operational amplifier OP1, a plurality of first reference capacitors C11, a second reference capacitor C12 and a third switch S3, the positive input terminal of the first operational amplifier OP1 is grounded, the negative input terminal of the first operational amplifier OP1 and the second end of the second reference capacitor C12 are connected to the output terminal of the charge cancellation reference column, the output terminal of the first operational amplifier OP1, the first end of the second reference capacitor C12 and the first ends of the plurality of first reference capacitors C11 are connected to the first end of the third switch S3, the second end of the third switch S3 is grounded, and the second ends of the plurality of first reference capacitors C11 are connected one-to-one to the output terminals of the plurality of columns of the storage and calculation units 10.

[0046] In this embodiment, the first operational amplifier OP1 and the second reference capacitor C12 can form an integrating circuit 30 for the charge cancellation reference column. After the cancellation reference charge is generated on the plate line of the charge cancellation reference column, the first operational amplifier OP1 and the second reference capacitor C12 can read out the cancellation reference charge generated by the charge cancellation reference column and copy it into each first reference capacitor C11. In this way, an equal amount of cancellation charge will also be generated on the plate line of each column of memory and computing units 10 connected to the first reference capacitor C11, which cancels out the original accumulated charge on the plate line of the memory and computing units 10 during actual operation, thereby reducing the influence of the accumulated charge generated by the memory and computing units 10 in the low-capacitance state on the calculation result. Controlling the third switch S3 to be in the closed state during the write operation stage can prevent multiple first reference capacitors C11 and the second reference capacitor C12 from being charged; controlling the third switch S3 to be in the open state during the read operation stage can charge multiple first reference capacitors C11 and the second reference capacitor C12 to prevent short circuits.

[0047] Refer to Figure 1 and Figure 2 In one embodiment, as shown in FIGS. 5 and 6, the integrating circuit 30 includes a second operational amplifier OP2, a first switch S1, and a third capacitor C3. The first end of the first switch S1 and the positive input terminal of the second operational amplifier OP2 are grounded. The second end of the first switch S1, the output terminal of the second operational amplifier OP2, and the first end of the third capacitor C3 are connected to each other. The negative input terminal of the second operational amplifier OP2 and the second end of the third capacitor C3 are connected to the output terminal of a column of the memory and computing units 10.

[0048] In this embodiment, the second operational amplifier OP2 and the third capacitor C3 can form a transimpedance integrator. The input voltage is applied to the memory and computing units 10 through the word line. The output currents of all the memory and computing units 10 on the same column will naturally superimpose on the bit line to form a total current. The integrating circuit 30 can then convert the total current into a voltage signal through the second operational amplifier OP2 and the third capacitor C3 for subsequent circuit processing; the input node of the integrating circuit 30 is the bit line connected to the inverting input terminal of the second operational amplifier OP2, which can ensure that the currents of all the memory and computing units 10 on a column directly flow into the integrating circuit 30; the feedback path is composed of the third capacitor C3, which determines the integration time constant and signal gain. Additionally, by controlling the conduction or cutoff of the first switch S1, the working state of the integrating circuit 30 can also be controlled.

[0049] Refer to Figure 1 and Figure 2 In one embodiment, the integrating circuit 30 further includes: The second switch S2, the first end of the second switch S2 is connected to the output end of the second operational amplifier OP2, and the second end of the second switch S2 is connected to the first end of the third capacitor C3.

[0050] As can be seen from the content of the above embodiments, multiple second operational amplifiers OP2 are used in the present invention, which will result in relatively high power consumption. To alleviate this problem, in this embodiment, a switch is added between the output end of the second operational amplifier OP2 and the third capacitor C3. After each column integration stage is completed, the second switch S2 is controlled by the controller to be turned on to maintain the voltage output, and the second operational amplifier OP2 is turned off to reduce unnecessary power consumption. In this way, the second operational amplifier OP2 only needs to work during the integration stage, effectively alleviating the problem of increased power consumption caused by the large number of second operational amplifiers OP2.

[0051] To better elaborate the technical concept of the present invention, in combination with the above embodiments and Figures 1 to 4 the working principle of the high-throughput parallel ferroelectric memory and computing system proposed by the present invention is described as follows: The charge replication cancellation circuit 20 is composed of a charge cancellation reference column and an offset charge cancellation circuit, and is used to cancel the accumulated charge generated by the memory and computing unit 10 in the low-capacitance state. During the write operation, for the unselected memory and computing units 10, their corresponding word lines remain at a low level. This makes the NMOS transistors of these memory and computing units 10 in the cut-off state, thus isolating the connection between these units and the bit lines. That is, the NMOS transistors are used to shield the unselected units to avoid write interference affecting the system robustness. In the array structure composed of multiple memory and computing units 10, the calculation results of each column of memory and computing units 10 are added by the integration circuit 30 and output in the form of voltage. The output formula is as follows: ; Among them, represents the output voltage of the integration circuit 30, represents the input voltage of the word line, represents the capacitance value of the first capacitor C1, represents the capacitance value of the third capacitor C3, i represents the number of rows in the row-column structure, j represents the number of columns in the row-column structure. Due to the limited switching ratio, the memory and computing units 10 in the low-capacitance state will affect the output voltage of the integration circuit 30. Therefore, by transforming in the formula, the following formula can be obtained: ; Among them, N represents the number of activated rows in a column, that is, the number of rows with the word line input of 1, and N = H + L, where L represents the number of units with the input of 1 and the stored data of 0, and H represents the number of units with the input of 1 and the stored data of 1. The voltage for charging the first capacitor C1 during the charging phase represents the capacitance value of the third capacitor C3 represents the capacitance value of the first capacitor C1 in the low-capacitance state represents the capacitance value of the first capacitor C1 in the high-capacitance state. Due to the non-ideality of the device is not equal to 0, and the sum of multiple may be equal to , resulting in an error in the calculation result. Therefore, reducing the charge accumulation caused by the memory and computing unit 10 in the low-capacitance state can effectively improve the accuracy of the array. The most natural way to reduce the influence of unnecessary charge accumulation caused by the memory and computing unit 10 in the low-capacitance state is to subtract the sum of the currents of the memory and computing unit 10 in the low-capacitance state from the resulting charge. Since N represents the number of activated rows in a column, we can subtract from each unit with input 1 during the calculation process. The result can obtain the following equation: ; Finally, when quantifying in the digital-to-analog converter, using ( − ) / as the unit voltage, the correct calculation result can be obtained.

[0052] The circuit operation of the charge replication cancellation scheme can be divided into two stages. In stage 1, that is, during the write operation phase, based on the input data, the word line is activated through the driver, and the first capacitor C1 is charged with the voltage , and the current flow direction is from the word line to the source of the first transistor Q1 in the memory and computing unit 10; the memory and computing unit will calculate the written data and store the calculation result in the capacitor of the memory and computing unit. In stage 2, that is, during the read operation phase, all bit line voltages are pulled to the low level through the driver. Due to capacitive coupling, a large amount of negative charge will be generated on each plate line. For the charge cancellation reference column, the amount of charge Qref = N· · is generated on the plate line to cancel the reference charge, where is the charging voltage of the first capacitor C1; then these charges are read out and copied by the first operational amplifier OP1 and the second reference capacitor C12 and output to the first reference capacitor C11 connected to each column of the memory and computing unit 10. At this time, due to capacitive coupling, an equal amount of positive charges will also be generated on the plate lines of each column of the memory and computing unit 10. In this process, the direction of current flow is that the first operational amplifier OP1 charges the first reference capacitor C11 on the plate line, while the first capacitor C1 on the plate line discharges the first transistor Q1. And in the integrating circuit, the third capacitor C3 is charged, a current flowing towards the plate line is generated at the first end of the third capacitor C3, and the positive voltage at the second end of the third capacitor C3 is used as the output. Therefore, the first capacitor C1 on the same plate line is in a discharging state, and the first reference capacitor C11 is in a charging state. Then, the negative charges generated at the second end of the first capacitor C1 on the same plate line will be partially offset by the positive charges generated at the second end of the first reference capacitor C11, thereby reducing the influence of the charges generated by the memory and computing unit 10 in the low-capacitance state.

[0053] Under the conditions that the switching ratio is 10 and the device-to-device variation rate is 10%, simulations are respectively carried out on the traditional crossbar array scheme and the scheme with the charge replication and cancellation circuit 20 introduced. Figure 3 and Figure 4 The dots in represent the simulation results, and the dashed lines represent the ideal results. Figure 3 is the simulation result diagram of the traditional array scheme. When there are a large number of memory and computing units 10 in the high-capacitance state with an input of 1, the calculation results will deviate significantly from the ideal values. Figure 4 is the simulation result diagram of the scheme of the present invention after introducing the charge replication and cancellation circuit 20. From the simulation result diagram, it can be known that the high-throughput parallel ferroelectric memory and computing system of the technical scheme of the present invention improves the accuracy of array calculation and is beneficial to improving the parallelism of the array.

[0054] The present invention effectively reduces the negative impact of the memory and computing unit 10 in the low-capacitance state on the output voltage by introducing a charge cancellation reference column and an offset charge cancellation circuit, thereby improving the accuracy and robustness of the high-throughput parallel ferroelectric memory and computing system. By effectively reducing the influence of the memory and computing unit 10 in the low-capacitance state on the output voltage, this scheme significantly improves the calculation accuracy of the high-throughput parallel ferroelectric memory and computing system under high parallelism conditions and is beneficial to increasing the parallelism of the array of the memory and computing units 10 in the high-throughput parallel ferroelectric memory and computing system.

[0055] The present invention also proposes a memory and computing device.

[0056] In one embodiment, the memory - computing device includes the high - throughput parallel ferroelectric memory - computing system as described above. It can be understood that since the above - mentioned high - throughput parallel ferroelectric memory - computing system is used in the memory - computing device of the present invention, the embodiments of the memory - computing device of the present invention include all the technical solutions of all the embodiments of the above - mentioned high - throughput parallel ferroelectric memory - computing system, and the achieved technical effects are also exactly the same, so they will not be elaborated here. The specific memory - computing device can be a non - volatile memory or a neural network accelerator, etc.

[0057] The present invention also provides a memory - computing device.

[0058] In one embodiment, the memory - computing device includes the memory - computing device as described above. It can be understood that since the above - mentioned memory - computing device is used in the memory - computing device of the present invention, the embodiments of the memory - computing device of the present invention include all the technical solutions of all the embodiments of the above - mentioned memory - computing device, and the achieved technical effects are also exactly the same, so they will not be elaborated here. The specific memory - computing device can be an artificial intelligence machine, a computer, a smart sensor, etc.

[0059] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A high-throughput parallel ferroelectric memory and computing system, characterized in that, Comprising: A plurality of memory - in - computing units, the plurality of memory - in - computing units being arranged in a row - column structure, and the plurality of memory - in - computing units being used to store the output data of an external device; A charge replication cancellation circuit, connected to the plurality of memory - in - computing units, the charge replication cancellation circuit being used to generate cancellation charges and cancel the cancellation charges against the cumulative charges generated by the plurality of columns of memory - in - computing units in a low - capacitance state; A plurality of integration circuits, connected to the plurality of memory - in - computing units, each integration circuit being used to add up the calculation results of a column of memory - in - computing units and then output.

2. The high-throughput parallel ferroelectric memory and computing system according to claim 1, characterized in that, The high - throughput parallel ferroelectric memory - in - computing system further includes: Word lines, connected to the control terminals of the memory - in - computing units in each row; Bit lines, connected to the input terminals of the memory - in - computing units in each column; Plate lines, connected to the output terminals of the memory - in - computing units in each column; A first driver, connected to the word lines, the first driver being used to output a first driving signal to the word lines to drive the plurality of memory - in - computing units to work; A second driver, connected to the bit lines and the plate lines, the second driver being used to output a second driving signal to the bit lines and the plate lines to drive the plurality of memory - in - computing units to perform data reading and writing.

3. The high-throughput parallel ferroelectric memory and computing system according to claim 2, characterized in that, The memory - in - computing unit includes a first transistor and a first capacitor. The gate of the first transistor is connected to the word line, the source of the first transistor is connected to the bit line, the drain of the first transistor is connected to the first end of the first capacitor, and the second end of the first capacitor is connected to the plate line.

4. The high-throughput parallel ferroelectric memory and computing system according to claim 2, characterized in that, The charge replication cancellation circuit includes: A charge cancellation reference column, the charge cancellation reference column being used to generate cancellation reference charges when the plurality of memory - in - computing units are working; An offset charge cancellation circuit, connected to the charge cancellation reference column, the offset charge cancellation circuit being used to generate corresponding cancellation charges according to the cancellation reference charges generated by the charge cancellation reference column and cancel the cancellation charges against the cumulative charges generated by the plurality of columns of memory - in - computing units in a low - capacitance state.

5. The high-throughput parallel ferroelectric memory and computing system according to claim 4, characterized in that, The charge cancellation reference column includes a plurality of reference units. The plurality of reference units in the charge cancellation reference column are connected to the plurality of memory - in - computing units in a column of memory - in - computing units one - to - one. The reference unit includes: a second transistor and a second capacitor. The gate of the second transistor is connected to the word line, the source of the second transistor is connected to the bit line, the drain of the second transistor is connected to the first end of the second capacitor, and the second end of the second capacitor is connected to the plate line.

6. The high-throughput parallel ferroelectric memory and computing system according to claim 4, characterized in that, The offset charge cancellation circuit includes a first operational amplifier, a plurality of first reference capacitors, a second reference capacitor, and a third switch. The positive input terminal of the first operational amplifier is grounded. The negative input terminal of the first operational amplifier and the second end of the second reference capacitor are connected to the output terminal of the charge cancellation reference column. The output terminal of the first operational amplifier, the first end of the second reference capacitor, and the first ends of the plurality of first reference capacitors are connected to the first end of the third switch. The second end of the third switch is grounded. The second ends of the plurality of first reference capacitors are connected to the output terminals of the plurality of columns of memory - in - computing units one - to - one.

7. The high-throughput parallel ferroelectric memory and computing system according to claim 1, characterized in that, The integrating circuit includes a second operational amplifier, a first switch, and a third capacitor. A first end of the first switch and a positive input terminal of the second operational amplifier are grounded. A second end of the first switch, an output terminal of the second operational amplifier, and a first end of the third capacitor are connected to each other. A negative input terminal of the second operational amplifier and a second end of the third capacitor are connected to an output terminal of a column of the memory and computing units.

8. The high-throughput parallel ferroelectric memory and computing system according to claim 7, characterized in that, The integrating circuit further includes: A second switch, a first end of the second switch is connected to the output terminal of the second operational amplifier, and a second end of the second switch is connected to the first end of the third capacitor.

9. A memory and computing device, characterized in that, It includes the high-throughput parallel ferroelectric memory and computing system according to any one of claims 1-8.

10. A memory and computing equipment, characterized in that, It includes the memory and computing device according to claim 9.

Citation Information

Patent Citations

  • Technology and device for canceling memory cell variations

    CN110277118A