Memory device and operation method thereof

By designing multiple memory cells to output cell currents in 3D memory and summing them, the MAC computing volume is increased without increasing the circuit area, and the problem of insufficient MAC computing volume in 3D memory is solved, and an efficient AI accelerator is realized.

CN114842893BActive Publication Date: 2025-05-16MACRONIX INTERNATIONAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111125191.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-15
Filing Date
2021-09-24
Publication Date
2025-05-16
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

In 3D memory, how to increase the calculation amount of multiplication accumulation calculation (MAC) without taking up additional circuit area is suitable for artificial intelligence (AI) accelerators.

Method used

By designing multiple memory cells in a memory device, using the weights of these memory cells to output the cell current, and sum it up through the region bit lines and signal lines, the overall signal line current is finally converted into an output, and MAC operation is realized.

Benefits of technology

Without increasing the circuit area, the amount of MAC operations is increased, and an efficient AI accelerator is realized, with the advantages of low circuit cost but high-speed computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842893B_ABST
    Figure CN114842893B_ABST
Patent Text Reader

Abstract

The present disclosure provides a memory device and an operation method thereof. The operation method comprises: when performing a multiplication-accumulation operation, inputting a plurality of inputs to a plurality of memory cells of the memory device through a plurality of first signal lines; according to a plurality of weights of the memory cells, the memory cells output a plurality of cell currents on a plurality of regional bit lines; summing up the cell currents on each of the regional bit lines into a plurality of signal line currents; summing up the signal line currents into an overall signal line current; and converting the overall signal line current into an output, wherein the output represents a multiplication-accumulation operation result of the inputs and the weights.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a memory device and an operating method thereof. Background Art

[0002] Artificial intelligence (AI) is becoming increasingly important. The Multiply Accumulate (MAC) operation is the core operation of AI.

[0003] Traditionally, to complete a MAC operation, data must be retrieved from a memory through an arithmetic logic unit (ALU), a floating-point operator, etc., for operation, which requires a large amount of data to be moved, so the operation speed is slow.

[0004] Computing-in-Memory (CIM) memories have been developed to achieve fast MAC performance, making them suitable for implementing AI accelerators.

[0005] Currently, memory devices have been developed towards 3D stacking to increase memory density. In terms of 3D structure, in addition to 3D NAND flash memory and 3D NOR flash memory, 3D AND flash memory has also been developed.

[0006] How to increase the MAC computing power in 3D memory without occupying additional circuit area is one of the directions the industry is working on.

[0007] Public Content

[0008] According to an embodiment of the present disclosure, a method for operating a memory device is provided, the method comprising: when performing a multiply accumulate (MAC) operation, inputting a plurality of inputs to a plurality of memory cells of the memory device through a plurality of first signal lines; outputting a plurality of cell currents on a plurality of regional bit lines according to a plurality of weights of the memory cells; summing up the cell currents on each of the regional bit lines into a plurality of signal line currents; summing up the signal line currents into an overall signal line current; and converting the overall signal line current into an output, wherein the output represents a result of a multiply accumulate operation of the inputs and the weights.

[0009] According to another embodiment of the present disclosure, a memory device is proposed, including: a memory array, including a plurality of memory cells, these memory cells store a plurality of weights, these memory cells are coupled to a plurality of first signal lines and a plurality of regional bit lines; at least one first regional signal line decoder, coupled to the memory array and at least one first overall signal line; and at least one conversion unit, coupled to the at least one first regional signal line decoder and the at least one first overall signal line.

[0010] In order to better understand the above and other aspects of the present invention, embodiments are given below and described in detail with reference to the accompanying drawings: BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 FIG. 4 is a circuit diagram of a memory device according to an embodiment of the present disclosure.

[0012] Figure 2 A schematic diagram showing a memory device performing a MAC operation according to an embodiment of the present disclosure.

[0013] Figure 3 A flow chart of a memory operation method according to an embodiment of the present disclosure is shown.

[0014] 4A to 4D A diagram showing device performance characteristics according to an embodiment of the present disclosure.

[0015] Description of Reference Numerals

[0016] 100: memory device

[0017] 110: Memory Array

[0018] D_LBL(1)~ D_LBL(M): Local bit line decoder

[0019] D_LSL(1)~ D_LSL(M): Regional source line decoder

[0020] ADC(1)~ADC(M): conversion unit

[0021] BLT(1)~BLT(Q): bit line transistor

[0022] SLT(1)~SLT(Q): Source line transistor

[0023] MC(i, j, k): memory cell

[0024] WL(1)~WL(N): word line

[0025] LBL: Local Bit Line

[0026] LSL: Local Source Line

[0027] 310~350: Steps DETAILED DESCRIPTION

[0028] The technical terms in this specification refer to the customary terms in the technical field. If some terms are explained or defined in this specification, the interpretation of these terms shall be based on the explanation or definition in this specification. Each embodiment of the present disclosure has one or more technical features. Under the premise of possible implementation, ordinary technicians in this technical field may selectively implement some or all of the technical features in any embodiment, or selectively combine some or all of the technical features in these embodiments.

[0029] Please refer to Figure 1 , which illustrates a circuit diagram of a memory device according to an embodiment of the present disclosure. Figure 1 As shown, a memory device 100 of an embodiment of the present disclosure includes: a memory array 110, a plurality of regional bit line decoders D_LBL(1)~D_LBL(M) (M is a positive integer), a plurality of regional source line decoders D_LSL(1)~D_LSL(M), and a plurality of conversion units ADC(1)~ADC(M). Here, the conversion unit is an analog-to-digital conversion unit for illustration, but it should be understood that the present disclosure is not limited thereto. The memory device 100 is, for example but not limited to, a 3D (three-dimensional) AND type memory device, and the memory array 110 is a 3D AND type memory array.

[0030] Each of the regional bit line decoders D_LBL(1) to D_LBL(M) includes a plurality of bit line transistors BLT(1) to BLT(Q) (Q is a positive integer). Similarly, each of the regional source line decoders D_LSL(1) to D_LSL(M) includes a plurality of source line transistors SLT(1) to SLT(Q).

[0031] The memory array 110 includes a plurality of memory cells MC (i, j, k) arranged in an array. The memory cells MC (i, j, k) are coupled to a plurality of word lines WL (1) to WL (N) (N is a positive integer), a plurality of regional source lines LSL, and a plurality of regional bit lines LBL. i = 1 to N, j = 1 to M, k = 1 to Q. i, j, and k are positive integers.

[0032] Taking the bit line transistor BLT(1) as an example, the bit line transistor BLT(1) has: a first terminal (such as a source) coupled to the local bit line LBL, a second terminal (such as a drain) coupled to the input terminal of the conversion unit and an overall bit line (not shown), and a control terminal (such as a gate) receiving a control signal (not shown). The bit line transistors BLT(2) to BLT(Q) have a similar coupling relationship.

[0033] Similarly, taking source line transistor SLT(1) as an example, source line transistor SLT(1) has: a first terminal (such as source) coupled to local source line LSL, a second terminal (such as drain) coupled to an overall source line (not shown), and a control terminal (such as gate) receiving a control signal (not shown). Source line transistors SLT(2)~SLT(Q) have similar coupling relationships.

[0034] When performing a multiply accumulate (MAC) operation, the word lines WL(1) to WL(N) receive word line voltages VWL(1) to VWL(N), where the word line voltages VWL(1) to VWL(N) are high level voltages or low level voltages. When performing a MAC operation, these word line voltages VWL(1) to VWL(N) are inputs.

[0035] These memory cells can be programmed to logic 1 or logic 0, that is, in one embodiment of the present disclosure, these memory cells are single-level cells (SLC), but the present disclosure is not limited thereto. In other possible embodiments of the present disclosure, these memory cells may be multi-level cells (MLC), which is also within the spirit of the present disclosure. When the memory cell is programmed to logic 1 and a high-level voltage is applied to the associated word line, the memory cell will output a cell current; when the memory cell is programmed to logic 1 and a low-level voltage is applied to the associated word line, the memory cell will not output a cell current; and when the memory cell is programmed to logic 0, regardless of whether a high-level voltage or a low-level voltage is applied to the associated word line, the memory cell will not output a cell current. The cell current Icell(i, j, k) output by the memory cell MC(i, j, k) can be expressed as Icell(i, j, k)=VWL(i)*w(i, j, k), wherein w(i, j, k) represents the weight value stored in the memory cell MC(i, j, k), that is, the transconductance of the memory cell MC(i, j, k).

[0036] Therefore, for the same regional bit line LBL, the bit line current (signal line current) flowing from the regional bit line LBL to the conversion unit ADC(j) is the sum of the cell currents of the N memory cells on the regional bit line LBL.

[0037] Each regional bit line decoder D_LBL(1)~D_LBL(M) adds up the bit line currents (signal line currents) on these regional bit lines LBL into an overall bit line current (also called overall signal line current). Therefore, it can be deduced that the overall bit line current = .

[0038] The conversion units ADC(1)~ADC(M) receive the individual overall bit line currents of the local bit line decoders D_LBL(1)~D_LBL(M) and convert them into outputs (digital codes) to obtain outputs OUT(1)~OUT(M). For example but not limited to, when the conversion units ADC(1)~ADC(M) have an 8-bit resolution, the input current can be converted into 8-bit outputs OUT(1)~OUT(M). Therefore, the output OUT(j) can be expressed as: , where IN(i) represents input data to the word line WL(i) of the memory array 110. When the input data IN(i) is logic high, the word line voltage VWL(i) is a high level voltage; and when the input data IN(i) is logic low, the word line voltage VWL(i) is a low level voltage.

[0039] That is, the output OUT(j) of the conversion unit ADC(j) is related to the MAC operation result of the storage weights of the memory cells coupled to the same conversion unit ADC(j) and the related word line voltages (input data).

[0040] Please refer to Figure 2 , which shows a schematic diagram of a memory device performing a MAC operation according to an embodiment of the present disclosure. Figure 2 As shown, during the MAC operation, the bit line transistors BLT(1) to BLT(3) and the source line transistors SLT(1) to SLT(3) are turned on, and the overall bit line voltage applied to the overall bit line GBLj is 1.8 V, while the overall source line voltage applied to the overall source line GSLj is 0 V. The high level voltage of the word line voltage VWL(1) to VWL(4) is 2.8 V, while the low level voltage is 0 V.

[0041] Therefore, in Figure 2 In the example, the overall bit line current = .

[0042] Depend on Figure 2 It can be seen that the total current added by the three bit line transistors BLT(1) to BLT(3) can represent multi-level weights 0, 1, 2 and 3, that is, 0 represents 2-level 00, 1 represents 2-level 01, 2 represents 2-level 10 and 3 represents 2-level 11. Each memory cell stores a single-level weight 1 or 0.

[0043] Furthermore, when one wants to represent x-order weights, the number of local bit lines coupled to the same conversion unit is: Q = 2 x -1. For example, if you want to represent a 4-order weight, the number of regional bit lines coupled to the same conversion unit is: Q = 2 4 -1=15.

[0044] That is to say, in an embodiment of the present disclosure, even if a single-level storage unit is used, multi-level weight calculations can still be performed. Therefore, the embodiment of the present disclosure has the advantage of a simple architecture but can perform complex MAC calculations.

[0045] Figure 3 A flow chart of a memory operation method according to an embodiment of the present disclosure is shown. In step 310, when performing a multiply accumulate (MAC) operation, multiple inputs are input to multiple memory cells of the memory device through multiple first signal lines. In step 320, according to multiple weights of these memory cells, these memory cells output multiple cell currents on multiple regional bit lines. In step 330, these cell currents on each of these regional bit lines are summed into multiple signal line currents. In step 340, these signal line currents are summed into an overall signal line current. In step 350, the overall signal line current is converted into an output, wherein the output represents a result of a multiply accumulate operation of these inputs and these weights.

[0046] 4A to 4D A device performance characteristic diagram according to an embodiment of the present disclosure is shown. Figure 4A As shown, in one embodiment of the present disclosure, if the difference between the on current (Ion) and the off current (Ioff) of the memory cell can be made larger (for example, (Ion / Ioff)>10 4 ), which allows for more parallel summation and reduces background leakage. In addition, the threshold voltage can be gradually raised by the Increment Step Programming Pulse (ISPP). When the word line voltage is fixed at 2.8V, the cell current can be gradually modified to be smaller.

[0047] Figure 4B Displays the unit current that is tunable and tight. Figure 4BAs shown, the cell current Icell can be modified to different ranges, for example but not limited to, the range of the cell current Icell can be from 150nA to 1.5μA. Moreover, the distribution of the cell current Icell is more compact, and the standard deviation (Standard Deviation, mathematical symbol σ (sigma)) of the cell current Icell can be less than 2% (σ<2%).

[0048] Figure 4C It is shown that according to an embodiment of the present disclosure, a 3D AND memory device can be read-disturb free, for example, when the word line voltage is about +7 V to +8 V. For a memory device with CIM function, the read bias voltage is about 2.8 V, which can further reduce the read disturbance.

[0049] Figure 4D It is shown that in one embodiment of the present disclosure, the memory device has a small random telegraph noise (RTN). When the cell current Icell is 150nA, the random telegraph noise is only + / -0.02μA, which is about 1.9% of the average value.

[0050] Furthermore, in an embodiment of the present disclosure, the operating voltage VCC (eg, 3.3V) can be stepped down to generate a word line voltage (eg, but not limited to, 2.8V), so no additional charge pump is required, which has the advantage of cost saving.

[0051] In an embodiment of the present disclosure, a 3D memory device may provide parallel N*M MAC operations to provide a high operation bandwidth.

[0052] In addition, in an embodiment of the present disclosure, if more word line voltages and more ADC outputs can be provided, the amount of computation can be greatly increased. For example, if 1000 word line voltages (N=1000) and 1000 ADC outputs (M=1000) can be provided (approximately equal to an 8Mb memory tile), up to 1M MAC computation can be calculated in a very short read time (e.g., 150ns, 8-bit ADC output reading is approximately 150ns), which is equivalent to 6.7TOPS MAC computing power, where TOPS is the abbreviation of Tera Operations Per Second, and 1 TOPS represents one trillion (10^12) operations per second.

[0053] Furthermore, the disclosed embodiment can perform ultra-high-speed MAC operations while occupying a very small memory circuit area. Therefore, the disclosed embodiment has the advantages of low circuit cost but high-speed operation.

[0054] In summary, although the present invention has been disclosed in the above embodiments, it is not intended to limit the present invention. A person skilled in the art of the present invention may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the scope of the attached claims.

Claims

1. A method for operating a memory device, characterized in that: The method of operation includes: When performing a multiply-accumulate (MAC) operation, inputting a plurality of inputs to a plurality of memory cells of the memory device through a plurality of word lines, the plurality of memory cells being controlled by the plurality of word lines; According to the multiple weights of the memory cells, the memory cells output multiple cell currents to multiple regional bit lines; summing up the cell currents on the bit lines of the respective regions into a plurality of signal line currents; summing the signal line currents into an overall signal line current; and Converting the entire signal line current into an output through a conversion unit, wherein the output represents a result of a multiplication and accumulation operation of the inputs and the weights; Wherein, when performing an x-order weight operation of a weight unit composed of at least one memory cell among the plurality of memory cells, at least one memory cell among the weight unit is coupled to a single word line among the plurality of word lines, and the weight unit receives a single input among the plurality of inputs via the single word line among the plurality of word lines, the number of regional bit lines coupled to the same conversion unit is: Q=2 x -1, where x and Q are both positive integers, and Q is the number of regional bit lines coupled to the same conversion unit; The memory cells are single-level memory cells; and When performing the product-add operation, a plurality of first transistors of at least one first region signal line decoder are turned on.

2. A memory device, characterized in that: include: A memory array includes a plurality of memory cells, the memory cells storing a plurality of weights, the memory cells being coupled to a plurality of regional source lines, a plurality of word lines and a plurality of regional bit lines, the plurality of memory cells being controlled by the plurality of word lines; at least one first regional signal line decoder coupled to the memory array and at least one first global signal line; as well as At least one conversion unit is coupled to the at least one first regional signal line decoder and the at least one first overall signal line, Wherein, when performing an x-order weight operation of a weight unit composed of at least one memory cell among the plurality of memory cells, at least one memory cell among the weight unit is coupled to a single word line among the plurality of word lines, and the weight unit receives a single input among the plurality of inputs via the single word line among the plurality of word lines, the number of regional bit lines coupled to the same conversion unit is: Q=2 x -1, where x and Q are both positive integers, and Q is the number of regional bit lines coupled to the same conversion unit; The memory cells are single-level memory cells; and When performing a multiplication-accumulation operation, a plurality of first transistors of the at least one first region signal line decoder are turned on.

3. The memory device according to claim 2, wherein: When performing a multiply-accumulate (MAC) operation, a plurality of inputs are input to the memory cells through the word lines; According to the weights of the memory cells, the memory cells output a plurality of cell currents to the regional bit lines; The cell currents are summed up on the bit lines of each of the regions into a plurality of signal line currents and input to the at least one first region signal line decoder; The at least one first region signal line decoder sums the signal line currents into an overall signal line current; as well as At least one first conversion unit converts the overall signal line current output by the at least one first regional signal line decoder to obtain an output, wherein the output represents a multiplication and accumulation operation result of the inputs and the weights.

4. The memory device according to claim 2, wherein: The memory device is a three-dimensional AND memory device.

5. The memory device according to claim 2, wherein: The at least one first integral signal line includes a plurality of first integral signal lines, the at least one conversion unit includes a plurality of conversion units, and each of the first integral signal lines is coupled to each of the conversion units.

6. The memory device according to claim 3, wherein: The first transistor has: a first terminal coupled to one of the regional bit lines, a second terminal coupled to an input terminal of the at least one conversion unit and the at least one first overall signal line, and a control terminal receiving a control signal.

7. The memory device according to claim 2, wherein: The at least one conversion unit is an analog-to-digital conversion unit (ADC).

8. The memory device according to claim 2, wherein: It also includes at least one second regional signal line decoder, which includes multiple second transistors. The second transistor has: a first end coupled to one of multiple third signal lines, a second end coupled to at least one second overall signal line, and a control end receiving a control signal.

Citation Information

Patent Citations

  • Decoders for analog neural memory in deep learning artificial neural network

    TW201941209A

  • Neural Network Classifier Using Array Of Four-Gate Non-volatile Memory Cells

    US20190237142A1