Memory device and method of operating the same

By performing multiplication operations in the memory array and combining analog, digital and hybrid accumulation circuits, MAC operations are optimized, solving the efficiency and accuracy of multi-bit input and weight value operations in the AI architecture, and achieving efficient in-memory operations.

CN114153421BActive Publication Date: 2025-07-18MACRONIX INTERNATIONAL CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110829792.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-14
Filing Date
2021-07-22
Publication Date
2025-07-18
Estimated Expiration
2041-07-22

AI Technical Summary

Technical Problem

The existing AI architecture is prone to encounter problems such as output bottlenecks and low efficiency when performing multi-bit input and multi-bit weight values, especially in memory operations, which are difficult to take into account both operation speed and accuracy.

Method used

The multiplication circuit in the memory array is used for multiplication operations, combined with analog, digital and hybrid accumulation circuits, the accumulative method is determined by the decision unit, the MAC operation is optimized by the analog-to-digital conversion unit and the comparator, and the data mapping and counting process are optimized by combining the grouping circuit and the counting unit.

Benefits of technology

It significantly reduces the number of memory cells and computing costs, improves the speed and accuracy of MAC operations, reduces the impact of error bits, and improves the overall computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114153421B_ABST
    Figure CN114153421B_ABST
Patent Text Reader

Abstract

The present disclosure provides a memory device and an operation method thereof. The memory device includes: a memory array including a plurality of memory cells for storing a plurality of weight values in these memory cells of the memory array; a multiplication circuit that multiplies a plurality of input data by these weight values to obtain a plurality of multiplication results, wherein when performing the multiplication, these memory cells generate a plurality of memory cell currents; a digital accumulation circuit that performs a digital accumulation on these multiplication results; an analog accumulation circuit that performs an analog accumulation on these memory cell currents to generate a first multiply-accumulate (MAC) operation result; and a determination unit that determines to perform the analog accumulation, the digital accumulation, or a hybrid accumulation, wherein when performing the hybrid accumulation, it is determined whether to trigger the digital accumulation circuit according to the first multiply-accumulate operation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a memory device with in-memory computing (IMC) and an operation method thereof. Background Art

[0002] Artificial intelligence (AI) has become a highly effective solution in many fields. The key operation of AI is to perform multiply-and-accumulation (MAC) operations on a large amount of input data (such as input feature maps) and weight values.

[0003] However, in the current AI architecture, it is prone to encounter input / output bottlenecks (IO bottlenecks) and inefficient MAC operation flows.

[0004] To achieve high accuracy, MAC operations with multiple-bit inputs and multiple-bit weight values can be performed. However, the input / output bottleneck becomes more serious, and the efficiency will be lower.

[0005] In-memory computing (IMC) can be used to accelerate MAC operations because IMC can reduce the complex arithmetic logic units (ALUs) required in the central processing architecture and provide high parallelism for in-memory MAC operations.

[0006] In the case of non-volatile memory-based IMC (NVM-based IMC), its advantages are, for example, non-volatile storage, reduced data transfer, etc.

[0007] When performing IMC, if both "operation speed" and "operation accuracy" can be taken into account, it will be beneficial to the performance of IMC.

[0008] Disclosure

[0009] According to an example of the present invention, a memory device is provided, including: a memory array including a plurality of memory cells, which can be used to store a plurality of weight values in these memory cells of the memory array; a multiplication circuit coupled to the memory array, the multiplication circuit multiplies a plurality of input data by these weight values to obtain a plurality of multiplication results, wherein when performing the multiplication, these memory cells generate a plurality of memory cell currents; a digital accumulation circuit coupled to the multiplication circuit, which performs a digital accumulation on these multiplication results; an analog accumulation circuit coupled to the memory array, which performs an analog accumulation on these memory cell currents to generate a first multiply-accumulate (MAC) operation result; and a decision unit coupled to the digital accumulation circuit and the analog accumulation circuit, which decides to perform the analog accumulation, the digital accumulation or a hybrid accumulation, wherein when performing the hybrid accumulation, it is determined whether to trigger the digital accumulation circuit according to the first multiply-accumulate operation result.

[0010] According to another example of the present invention, a method for operating a memory device is provided, including: storing a plurality of weight values in a plurality of memory cells of a memory array of the memory device; performing bit multiplication on a plurality of input data and these weight values to obtain a plurality of multiplication results, wherein when performing the multiplication, these memory cells generate a plurality of memory cell currents; and deciding to perform an analog accumulation, a digital accumulation or a hybrid accumulation, wherein when performing the analog accumulation, these memory cell currents are subjected to the analog accumulation to generate a first multiply-accumulate (MAC) operation result; when performing the digital accumulation, these multiplication results are subjected to the digital accumulation to generate a second multiply-accumulate operation result; and when performing the hybrid accumulation, it is determined whether to trigger the digital accumulation according to the first multiply-accumulate operation result.

[0011] In order to have a better understanding of the above and other aspects of the present invention, specific embodiments are given below and described in detail in conjunction with the accompanying drawings as follows: Description of the Drawings

[0012] Figure 1 It is a functional block diagram of a memory device with in-memory computing function according to an embodiment of the present invention.

[0013] Figure 2 It is a schematic diagram of data mapping according to an embodiment of the present invention.

[0014] Figures 3A to 3C They are several examples of data mapping according to an embodiment of the present invention.

[0015] Figure 4 They are schematic diagrams of two exemplary multiplication operations of an embodiment of the present invention.

[0016] Figure 5A And Figure 5B It is a schematic diagram of the grouping operation (majority decision operation) and counting according to an embodiment of the present invention.

[0017] Figure 6 It is to compare several MAC operation processes according to an embodiment of the present invention.

[0018] Figure 7A It is a flowchart of programming a fixed memory page in an embodiment of the present invention, Figure 7B It is a flowchart of adjusting the read voltage in an embodiment of the present invention.

[0019] Figure 8 It is a MAC operation process according to an embodiment of the present invention.

[0020] Description of Reference Numerals

[0021] 100: Memory device

[0022] 110: Memory array

[0023] 120: Multiplication circuit

[0024] 130: Input / output circuit

[0025] 140: Grouping circuit

[0026] 150: Counting unit

[0027] 111: Memory cell

[0028] 121: Unit multiplication unit

[0029] 121A: Input latch

[0030] 121B: Sense amplifier

[0031] 121C: Output latch

[0032] 121D: Common data latch

[0033] 141: Grouping unit

[0034] 135: Digital accumulator circuit

[0035] 160: Analog accumulator circuit

[0036] 170: Decision unit

[0037] 161: Analog-to-digital conversion unit

[0038] 163: Comparator

[0039] 301A, 303A, 301B, 303B, 311A, 313A, 311B, 313B: bits

[0040] 302, 312, 314: weight values

[0041] 405: latch

[0042] 410: bit line switch

[0043] 710 - 750: Steps 810 - 860: Steps Detailed implementation manners

[0044] The technical terms in this specification refer to the customary terms in this technical field. If this specification explains or defines some terms, the explanations of these terms shall prevail according to the explanations or definitions in this specification. Each embodiment of the present disclosure has one or more technical features. On the premise of possible implementation, those skilled in the art in this technical field can selectively implement some or all of the technical features in any embodiment, or selectively combine some or all of the technical features in these embodiments.

[0045] Please refer to Figure 1 , which is a functional block diagram of a memory device 100 with an In-Memory-Computing (IMC) function according to an embodiment of the present invention. The memory device 100 with an in-memory computing function includes: a memory array 110, a multiplication circuit 120, an input / output circuit 130, a digital accumulation circuit 135, an analog accumulation circuit 160, a decision unit 170, and a comparator 163. The digital accumulation circuit 135 includes a grouping circuit 140 and a counting unit 150. The analog accumulation circuit 160 includes an analog-to-digital conversion unit 161. Among them, the memory array 110, the multiplication circuit 120, the analog accumulation circuit 160, and the analog-to-digital conversion unit 161 are analog, while the digital accumulation circuit 135, the grouping circuit 140, and the counting unit 150 are digital.

[0046] The memory array 110 includes a plurality of memory cells 111. In an embodiment of the present invention, the memory cell 111 is, for example but not limited to, a non-volatile memory cell. When performing a MAC operation, the memory cell 111 can be used to store weight values.

[0047] The multiplication circuit 120 is coupled to the memory array 110. The multiplication circuit 120 includes a plurality of unit multiplication units 121. Each unit multiplication unit 121 includes: an input latch 121A, a sense amplifier (SA) 121B, an output latch 121C, and a common data latch (CDL) 121D. The input latch 121A is coupled to the memory array 110. The sense amplifier 121B is coupled to the input latch 121A. The output latch 121C is coupled to the sense amplifier 121B. The common data latch 121D is coupled to the output latch 121C.

[0048] The input / output circuit 130 is coupled to the multiplication circuit 120, the grouping circuit 140, and the counting unit 150 to receive input data and output the output data obtained by the memory device 100.

[0049] The digital accumulation circuit 135 is used for digital accumulation, and its details will be described below.

[0050] The analog accumulation circuit 160 is used for analog accumulation, and its details will be described below.

[0051] The decision unit 170 determines whether the memory device 100 performs analog accumulation, digital accumulation, or hybrid accumulation. The decision unit 170 can respectively output enable signals EN1 and EN2 to the analog accumulation circuit 160 and the digital accumulation circuit 135 to determine whether to activate the analog accumulation circuit 160 or the digital accumulation circuit 135.

[0052] "Analog accumulation" means activating the analog accumulation circuit 160 but not activating the digital accumulation circuit 135. "Digital accumulation" means activating the digital accumulation circuit 135 but not activating the analog accumulation circuit 160. "Hybrid accumulation" means activating the digital accumulation circuit 135 and the analog accumulation circuit 160.

[0053] The analog-to-digital conversion unit 161 is coupled to these memory cells 111 of the memory array 110. The cell currents of these memory cells 111 can be accumulated and input to the analog-to-digital conversion unit 161 to be converted into the first MAC operation result OUT1.

[0054] Comparator 163 is coupled to the analog-to-digital conversion unit 161 for comparing the first MAC operation result OUT1 with a trigger reference value. When "hybrid accumulation" is selected, when the first MAC operation result OUT1 is lower than the trigger reference value, the comparator 163 does not output a trigger signal TS to the digital accumulation circuit 135 (i.e., the digital accumulation circuit 135 will not be triggered); and when the first MAC operation result OUT1 is higher than the trigger reference value, the comparator 163 outputs the trigger signal TS to the digital accumulation circuit 135 to trigger the digital accumulation circuit 135 to perform digital accumulation. When "analog accumulation" is selected, the trigger signal TS output by the comparator 163 will be ignored by the digital accumulation circuit 135.

[0055] In an embodiment of the present invention, "analog accumulation" can be used to quickly filter out useless data to increase the operation speed of the MAC. And "digital accumulation" can accumulate the unfiltered data to increase the accuracy of the MAC. And "hybrid accumulation" can reduce the variation influence due to the use of low-resolution quantization operations. In addition, it can also avoid accumulating useless data and can maintain the resolution. That is, "hybrid accumulation" takes into account the advantages of "analog accumulation" and "hybrid accumulation" while reducing their disadvantages.

[0056] The grouping circuit 140 is coupled to the multiplication circuit 120. The grouping circuit 140 includes a plurality of grouping units 141. These grouping units 141 perform a grouping operation on the plurality of multiplication results of these unit multiplication units 121 to obtain a plurality of grouping results. In a possible embodiment of the present invention, the grouping operation can be implemented, for example, by a majority technique, such as a majority function technique. The grouping circuit 140 is implemented by a majority grouping circuit according to the majority function technique, and the grouping unit 141 is implemented by a distributed majority grouping unit, but the present invention is not limited thereto. The grouping technique can be implemented by other similar techniques. In an embodiment of the present invention, the grouping circuit 140 can be selectively provided.

[0057] The counting unit 150 is coupled to the grouping circuit 140 or the multiplication circuit 120. In an embodiment of the present invention, the counting unit 150 is configured to perform bitwise counting or bitwise accumulation on the multiplication result of the multiplication circuit 120 to generate a second MAC operation result OUT2 (when the memory device 100 does not include the grouping circuit 140). Alternatively, the counting unit 150 is configured to perform bitwise counting or bitwise accumulation on the grouping result (e.g., the majority decision result) of the grouping circuit 140 to generate a second MAC operation result OUT2 (when the memory device 100 includes the grouping circuit 140). In an embodiment of the present invention, the counting unit 150 can be implemented by a known counting circuit, such as, but not limited to, a ripple counter. In the description of the present invention, counting and accumulation basically have the same meaning, and a counter and an accumulator basically have the same meaning.

[0058] Specifically, when the decision unit 170 determines that the memory device 100 performs analog accumulation, the first MAC operation result OUT1 is used as the MAC operation result. When the decision unit 170 determines that the memory device 100 performs digital accumulation, the second MAC operation result OUT2 is used as the MAC operation result. When the decision unit 170 determines that the memory device 100 performs hybrid accumulation, before the digital accumulation circuit 135 is triggered by the trigger signal TS, the first MAC operation result OUT1 is used as the MAC operation result; and after the digital accumulation circuit 135 is triggered by the trigger signal TS, the second MAC operation result OUT2 is used as the MAC operation result.

[0059] Please now refer to Figure 2 , which shows a schematic diagram of data mapping according to an embodiment of the present invention. As Figure 2 shown, taking the example that each input data (or each weight value) has N dimensions (N is a positive integer) and is 8 bits (but it should be known that the present invention is not limited thereto).

[0060] The following takes the data mapping of the input data as an example for illustration, but it should be known that the present invention is not limited thereto. The following description also applies to the data mapping of the weight value.

[0061] When the input data is represented in 8-bit binary, the input data (or weight value) is divided into a most significant bit (MSB) vector and a least significant bit (LSB) vector. The most significant bit vector of the 8-bit input data (or weight value) includes 4 bits B7 to B4, and the least significant bit vector includes 4 bits B3 to B0.

[0062] The most significant bit (MSB) vector and the least significant bit (LSB) vector of the input data are represented in unary coding (i.e., value format). For example, bit B7 of the MSB vector of the input data can be represented as B70 to B77, bit B6 of the MSB vector of the input data can be represented as B60 to B63, bit B5 of the MSB vector of the input data can be represented as B50 to B51, and bit B4 of the MSB vector of the input data is represented as B4 as well.

[0063] The bits of the MSB vector of the input data represented in unary coding (value format) and the bits of the LSB vector of the input data are repeated multiple times to form an unfolding dot product (unFDP) form. For example, the bits of the MSB of the input data are repeated (24 - 1) times, and similarly, the bits of the LSB of the input data are repeated (24 - 1) times. In this way, the input data can be represented in an unfolding dot product form.

[0064] A multiplication operation is performed on the input data (in unfolding dot product form) and the weight value to obtain a multiplication operation result.

[0065] For the sake of understanding, an example is given below, but it should be noted that it is not used to limit the present invention.

[0066] Now, please refer to Figure 3A , which shows an example of one-dimensional data mapping according to an embodiment of the present invention. As Figure 3A shown, the input data = (IN1, IN2) = (2, 1), and the weight value = (We1, We2) = (1, 2). The MSB and LSB of the input data are represented in binary form. Therefore, IN1 = 10, and IN2 = 01. Similarly, the bits of the MSB and LSB of the weight value are represented in binary form. Therefore, We1 = 01, and We2 = 10.

[0067] The MSB and LSB of the input data, and the MSB and LSB of the weight value are encoded in unary coding (value format). That is, the MSB of the input data is encoded as 110, the LSB of the input data is encoded as 001. Similarly, the MSB of the weight value is encoded as 001, and the LSB of the weight value is encoded as 110.

[0068] After that, each bit of the MSB (110) of the input data encoded in unary code is repeated multiple times with each bit of the LSB (001) of the input data encoded in unary code to become the form of unfolding dot product (unFDP). For example, each bit of the MSB (110) of the input data is repeated 3 times, so the product expansion of the MSB of the input data is 111111000. Each bit of the LSB (001) of the input data is repeated 3 times, so the product expansion of the LSB of the input data is 000000111.

[0069] Perform a MAC operation on the input data (product expansion) and the weight value to obtain the MAC operation result. The MAC operation result is: 1*0 = 0, 1*0 = 0, 1*1 = 1, 1*0 = 0, 1*0 = 0, 1*1 = 1, 0*0 = 0, 0*0 = 0, 0*1 = 0, 0*1 = 0, 0*1 = 0, 0*0 = 0, 0*1 = 0, 0*1 = 0, 0*0 = 0, 1*1 = 1, 1*1 = 1, 1*0 = 0. Adding these values together, we can get: 0+0+1+0+0+1+0+0+0+0+0+0+0+0+0+1+1+0 = 4.

[0070] As can be seen from the above, if the input data is i bits and the weight value is j bits (both i and j are positive integers), then the number of memory units used is: (2i - 1)*(2j - 1).

[0071] Now please refer to Figure 3B , which shows another possible example of data mapping according to an embodiment of the present invention. In Figure 3B , the input data is (IN1) = (2), and the weight value is (We1) = (1). The input data and the weight value are 4 bits.

[0072] When the input data is represented in binary format, IN1 = 0010. Similarly, when the weight value is represented in binary format, We1 = 0001.

[0073] Encode the input data and the weight value into unary code (numerical form). For example, the highest bit "0" of the input data is encoded as "00000000", and the lowest bit "0" of the input data is encoded as "0", and so on. Similarly, the highest bit "0" of the weight value is encoded as "00000000", and the lowest bit "1" of the weight value is encoded as "1".

[0074] Each bit of the input data encoded in unary code is copied multiple times to become the product expansion. For example, the highest bit 301A of the input data encoded in unary code is copied 15 times to become bit 303A; and the lowest bit 301B of the input data encoded in unary code is copied 15 times to become bit 303B.

[0075] The weighted value 302 encoded into unary code is also copied 15 times to represent a product expansion.

[0076] A multiplication operation is performed on the input data represented as a product expansion and the weighted value represented as a product expansion to generate a MAC operation result. Specifically, bit 303A of the input data is multiplied by the weighted value 302; bit 303B of the input data is multiplied by the weighted value 302, and so on. Summing up the multiplication values can generate the MAC operation result ("2").

[0077] Now please refer to Figure 3C , which shows another possible example of data mapping according to an embodiment of the present invention. In Figure 3C , the input data is (IN1) = (1), and the weighted value is (We1) = (5). The input data and the weighted value are 4 bits.

[0078] When the input data is represented in binary format, IN1 = 0001. Similarly, when the weighted value is represented in binary format, We1 = 0101.

[0079] The input data and the weighted value are encoded into unary code (numerical form).

[0080] Each bit of the input data encoded into unary code is copied multiple times to become a product expansion. In Figure 3C , when copying each bit of the input data and the weighted value, a bit "0" is added. For example, the highest bit 311A of the input data encoded into unary code is copied 15 times and a bit "0" is added to become bit 313A; and the lowest bit 311B of the input data encoded into unary code is copied 15 times and a bit "0" is added to become bit 313B. Thus, the input data is represented as a product expansion.

[0081] Similarly, the weighted value 312 encoded into unary code is also copied 15 times, and an additional bit "0" is added to each weighted value 314. Thus, the weighted value is represented as a product expansion.

[0082] A multiplication operation is performed on the input data represented as a product expansion and the weighted value represented as a product expansion to generate a MAC operation result. Specifically, bit 313A of the input data is multiplied by the weighted value 314; bit 313B of the input data is multiplied by the weighted value 314, and so on. Summing up the multiplication values can generate the MAC operation result ("5").

[0083] In the prior art, for an 8-bit input data and an 8-bit weighted value to perform a MAC operation, if the direct MAC algorithm is used, the number of memory units used is 255 * 255 * 512 = 33,292,822.

[0084] Conversely, as described above, in the embodiments of the present invention, when performing a MAC operation on 8-bit input data and 8-bit weight values, the number of memory cells used is 15 * 15 * 512 * 2 = 115,200 * 2 = 230,400. Therefore, the number of memory cells used in the embodiments of the present invention during the MAC operation is approximately 0.7% of that of the conventional technology.

[0085] In the embodiments of the present invention, by using the unFDP-based data mapping, the number of memory cells used during the operation can be reduced. Therefore, the operation cost can be reduced, and the error correction code (ECC) cost can also be reduced. In addition, the fail-bit effect can be tolerated.

[0086] Please refer again to Figure 1 . In the embodiments of the present invention, during the multiplication operation, the weight values (transduction values) are stored in these memory cells 111 of the memory array 110, and the input data (voltage) is read by the input / output circuit 130 and transmitted to the common data latch 121D. The common data latch 121D transmits the input data to the input latch 121A.

[0087] To better understand the multiplication operation of the embodiments of the present invention, please now refer to Figure 4 , which shows a schematic diagram of an exemplary multiplication operation of the embodiments of the present invention. Figure 4 Applied to the memory device to support the selected bit-line read function. Figure 4 In, the input latch 121A includes a latch (first latch) 405 and a bit-line switch 410.

[0088] As Figure 4 shown, the weight values are represented in one-hot encoding (numerical form) (such as Figure 2 ). Therefore, the most significant bit of the weight value is stored in 8 memory cells 111, the second most significant bit of the weight value is stored in 4 memory cells 111, the third most significant bit of the weight value is stored in 2 memory cells 111, and the least significant bit of the weight value is stored in 1 memory cell 111.

[0089] Similarly, the input data is represented in one-hot encoding (numerical form) (such as Figure 2 ), so the most significant bit of the input data is stored in 8 common data latches 121D, the second most significant bit of the input data is stored in 4 common data latches 121D, the third most significant bit of the input data is stored in 2 common data latches 121D, and the least significant bit of the input data is stored in 1 common data latch 121D. The input data is sent from the common data latch 121D to the latch 405.

[0090] In Figure 4 Figure 4 , these multiple bit-line switches 410 are coupled between the memory cell 111 and the sense amplifier 121B. The bit-line switch 410 is controlled by the latch 405. For example, when the latch 405 outputs a bit 1, the bit-line switch 410 is turned on, and when the latch 405 outputs a bit 0, the bit-line switch 410 is turned off.

[0091] In addition, when the weight value in the memory cell 111 is a bit 1 and the bit-line switch 410 is turned on (the input data is a bit 1), the sense amplifier 121B will sense the memory cell current to generate a multiplication result of "1". When the weight value in the memory cell 111 is a bit 0 and the bit-line switch 410 is turned on (the input data is a bit 1), the sense amplifier 121B does not sense the memory cell current. When the weight value in the memory cell 111 is a bit 1 and the bit-line switch 410 is turned off (the input data is a bit 0), the sense amplifier 121B does not sense the memory cell current to generate a multiplication result of "0". When the weight value in the memory cell 111 is a bit 0 and the bit-line switch 410 is turned off (the input data is a bit 0), the sense amplifier 121B does not sense the memory cell current.

[0092] That is, via Figure 4 Figure 4 's layout, when the input data is a bit 1 and the weight value is a bit 1, the sense amplifier 121B senses the memory cell current to generate a multiplication result of "1". For other cases, the sense amplifier 121B does not sense the memory cell current to generate a multiplication result of "0".

[0093] The memory cell currents IMC generated by these memory cells 111 are jointly input to the analog-to-digital conversion unit 161.

[0094] As for the relationship between the input data, the weight value, the digital multiplication result, and the analog memory cell current IMC, it is shown in the following table:

[0095] Input data Weight value Result of mathematical multiplication IMC 0 0 (HVT) 0 0 0 +1 (LVT) 0 0 1 0 (HVT) 0 IHVT 1 +1 (LVT) 1 ILVT

[0096] In the above table, HVT and LVT respectively represent high-threshold memory cells and low-threshold memory cells. And IHVT and ILVT respectively represent the analog memory cell currents IMC generated by the high-threshold memory cell and the low-threshold memory cell (the weight values are 0 (HTV) and +1 (LTV) respectively) when the input data is logic 1.

[0097] In an embodiment of the present invention, when performing a multiplication operation, a selected bit line read (SBL-read) instruction can be reused. Therefore, the embodiment of the present invention can reduce the variation influence caused by single-bit representation.

[0098] Now, please refer to Figure 5A , which shows a schematic diagram of a grouping operation (majority decision operation) and bitwise counting according to an embodiment of the present invention. As Figure 5A shown, reference symbol GM1 represents the first multiplication result obtained after performing a bitwise multiplication on the first MSB vector of the input data and the weight value; reference symbol GM2 represents the second multiplication result obtained after performing a bitwise multiplication on the second MSB vector of the input data and the weight value; reference symbol GM3 represents the third multiplication result obtained after performing a bitwise multiplication on the third MSB vector of the input data and the weight value; reference symbol GL represents the fourth multiplication result obtained after performing a bitwise multiplication on the LSB of the input data and the weight value. After the grouping operation (majority decision operation), the grouping result of the first multiplication result GM1 is the first grouping result CB1 (whose cumulative weight is 22); the grouping result of the second multiplication result GM2 is the second grouping result CB2 (whose cumulative weight is 22); the grouping result of the third multiplication result GM3 is the third grouping result CB3 (whose cumulative weight is 22); and the grouping result of the fourth multiplication result GL is the fourth grouping result CB4 (whose cumulative weight is 20).

[0099] Figure 5B shows Figure 3C the cumulative example of Figure 3C and Figure 5B . Please refer to Figure 5B . As Figure 3C shown, bit 313B of the input data ( Figure 3C ) is multiplied by the weight value 314. The first four bits ("0000") of the multiplication result generated by multiplying bit 313B of the input data ( Figure 3C ) by the weight value 314 are grouped as the first multiplication result "GM1". Similarly, the fifth to eighth bits ("0000") of the multiplication result generated by multiplying bit 3-13B of the input data ( Figure 3C ) by the weight value 314 are grouped as the second multiplication result "GM2". The ninth to twelfth bits ("1111") of the multiplication result generated by multiplying bit 313B of the input data ( Figure 3CThe thirteenth to sixteenth bits (“0010”) of the multiplication result generated by multiplying bit 313B of

[0100] After the grouping operation (majority decision operation), the first grouping result CB1 is “0” (its cumulative weight is 22); the second grouping result CB2 is “0” (its cumulative weight is 22); the third grouping result CB3 is “1” (its cumulative weight is 22). When counting, these grouping results CB1 to CB4 are multiplied by their respective cumulative weights and accumulated to generate the MAC operation result. For example, as Figure 5B shown, the MAC operation result (the second MAC operation result OUT2) is CB1 * 22 + CB2 * 22 + CB3 * 22 + CB4 * 20 = 0 * 22 + 0 * 22 + 1 * 22 + 1 * 20 = 00000000000000000000000000000101 = 5.

[0101] In an embodiment of the present invention, the grouping principle (majority decision principle) can be as follows:

[0102] Group bit Clustering result (majority decision result) 1111 (Condition A) 1 1110 (Condition B) 1 1100 (Condition C) 1 or 0 1000 (Condition D) 0 0000 (Condition E) 0

[0103] In the above table, for situation A, since all groups are correct (“1111” has no error bits), therefore, the majority decision result is 1. For situation E, since all groups are correct (“0000” has no error bits), therefore, the majority decision result is 0.

[0104] For situation B, since there is 1 error bit in the group (“0” in “1110” is incorrect), through majority decision, “1110” can be determined as “1”. For situation D, since there is 1 error bit in the group (“1” in “0001” is incorrect), through majority decision, “0001” can be determined as “0”.

[0105] For situation C, there are 2 error bits in the group (“00” in “1100” is incorrect, or “11” in “1100” is incorrect), through majority decision, “1100” can be determined as “1” or “0”.

[0106] Therefore, in the embodiment of the present invention, through the grouping (majority decision) function, the number of error bits can be reduced.

[0107] The grouping result of the grouping circuit 140 is input to the counting unit 150 for bit counting.

[0108] When counting, the counting result of the multiplication operation result of the MSB vector is accumulated with the counting result of the multiplication operation result of the LSB vector. For Figure 5AIn the case of, two accumulators are used. The first accumulator is to be assigned a higher accumulation weight value (e.g., 22). The first accumulator is to accumulate: (1) "the grouped (majority decision) result obtained by grouping (majority decision) the multiplication operation result GM1: 1 bit" plus "the grouped (majority decision) result obtained by grouping (majority decision) the multiplication operation result GM2: 1 bit" plus "the grouped (majority decision) result obtained by grouping (majority decision) the multiplication operation result GM3: 1 bit". The count result obtained by the first accumulator is then multiplied by the higher accumulation weight value (e.g., 22). The second accumulator is to be assigned a lower accumulation weight value (e.g., 20). The second accumulator directly counts the multiplication operation result GL (multiple bits). Adding the two accumulated results weighted by the accumulation weights gives the MAC result. For example, the grouped result obtained by grouping the multiplication operation result GM1 is 1 (1 bit), the grouped result obtained by grouping the multiplication operation result GM2 is 0 (1 bit), and the grouped result obtained by grouping the multiplication operation result GM3 is 1 (1 bit). The count result (1 + 0 + 1) obtained by the first accumulator is multiplied by 22, which is 2 * 22 = 8. The multiplication operation result GL is 4 (3 bits), which can be directly counted. Adding the two accumulated results weighted by the accumulation weights gives the MAC result of 8 + 4 = 12.

[0109] As can be seen from the above, in the embodiments of the present invention, when performing counting or accumulation, since the input data has been expanded into the unFDP form, the data stored in the common data latch can be grouped (i.e., divided into the MSB vector and the LSB vector), and the error bits in the MSB vector / LSB vector can be reduced through the grouping mechanism (majority decision mechanism).

[0110] In addition, in the embodiments of the present invention, even when using a traditional accumulator (counter), the counting / accumulation time can still be reduced because the embodiments of the present invention use digital counting instructions (error bit counting) and different accumulation weights are given to the accumulation results of different vectors (MSB vector and LSB vector). For example, the accumulation operation time can be reduced to about 40%.

[0111] Figure 6 Shows the MAC operation flow of an embodiment of the present invention. In Figure 6 wherein, DMAC represents the first digital accumulation (but without performing a grouping operation, that is, the memory device 100 does not include a grouping circuit 140), mDMAC represents the second digital accumulation (performing a grouping operation, that is, the memory device 100 includes a grouping circuit 140), AMAC represents "analog accumulation", and HMAC represents "hybrid accumulation".

[0112] In the first digital accumulation operation process of the embodiment of the present invention, input data is transmitted to the memory device. Bit line setting and word line setting are performed at the same time. After the bit line setting is completed, sensing is performed. Digital accumulation operation is performed. And the digital accumulation operation result is returned. The above operation is repeated until all input data have been calculated.

[0113] In the second digital accumulation operation flow of the embodiment of the present invention, the operation speed of digital accumulation can be accelerated by grouping operation.

[0114] In the analog accumulation operation process of the embodiment of the present invention, when sensing is performed, ADC conversion and comparison operations can be completed simultaneously, so the MAC operation can be further improved.

[0115] In the hybrid accumulation operation flow of the embodiment of the present invention, since both analog accumulation and digital accumulation are required, the operation speed of the hybrid accumulation is slower than that of the analog accumulation but faster than that of the digital accumulation. However, the accuracy of the hybrid accumulation is almost equal to that of the digital accumulation and higher than that of the analog accumulation.

[0116] Depend on Figure 6 It can be seen that the MAC operation of the embodiment of the present invention can be divided into two sub-operation types. The first sub-operation type is a multiplication operation, which multiplies the input data by a weight value, and is performed according to the selected bit line read instruction. The second sub-operation type is accumulation (data counting), especially fail bit counting. In other possible embodiments of the present invention, more counting units can be added to speed up the counting / accumulation operation.

[0117] The operation time of digital accumulation mainly depends on the accumulation speed of the counting unit 150, because the counting unit 150 is bit-by-bit calculation. The quantization accuracy of the analog-to-digital conversion unit 161 mainly depends on the variation tolerance of the memory unit. Therefore, digital accumulation has high accuracy but low accumulation speed compared to analog accumulation.

[0118] In addition, in the embodiment of the present invention, the read voltage can also be adjusted. Figure 7A A flowchart showing a method of programming a fixed memory page in an embodiment of the present invention is shown. Figure 7B A flow chart showing adjustment of a read voltage in an embodiment of the present invention is shown.

[0119] exist Figure 7A In step 710, a known input data is programmed into a fixed memory page. For example but not limited to, the bit ratio of the known input data is: 75% is bit 0, and 25% is bit 1.

[0120] existFigure 7B In step 720, the fixed memory page is read, and the ADC is started. In step 730, it is judged whether the output value of the ADC is close to the reference test value (if the ADC is 8-bit, the reference test value is 127, but the present invention is not limited thereto). If step 730 is NO, the process proceeds to step 740. If step 730 is YES, the process proceeds to step 750.

[0121] In step 740, if the output value of the ADC is less than the reference test value, the read voltage is increased; and if the output value of the ADC is greater than the reference test value, the read voltage is decreased. After step 740 ends, the process returns to step 720.

[0122] In step 750, the current read voltage is recorded for use in subsequent read operations.

[0123] As is known, the read voltage will affect the ADC output value and the reading of bit 1. Therefore, in the embodiments of the present invention, the read voltage can be periodically corrected according to operating conditions (such as but not limited to, programming cycles, temperature, or read interference, etc.) to maintain high accuracy and reliability.

[0124] Figure 8 Shows a MAC operation flow according to an embodiment of the present invention. In step 810, a plurality of weight values are stored in a plurality of memory cells of a memory array of the memory device. In step 820, bit multiplications are performed on a plurality of input data and these weight values to obtain a plurality of multiplication results, wherein when performing the multiplications, these memory cells generate a plurality of memory cell currents. In step 830, it is determined to perform an analog accumulation, a digital accumulation, or a hybrid accumulation. In step 840, when performing the analog accumulation, the memory cell currents are analog-accumulated to obtain a first multiply-accumulate (MAC) operation result. In step 850, when performing the digital accumulation, the multiplication results are digitally accumulated to obtain a second multiply-accumulate (MAC) operation result. In step 860, when performing the hybrid accumulation, it is determined whether to trigger the digital accumulation according to the first multiply-accumulate operation result.

[0125] Embodiments of the present invention can be applied to NAND flash memories, or memory devices sensitive to retention and thermal variations, such as but not limited to, NOR flash memories, phase change (PCM) flash memories, magnetic random access memories (magnetic RAM), or resistive RAMs.

[0126] Embodiments of the present invention can be applied to 3D memories and 2D memories, such as but not limited to, 2D / 3D NAND flash memories, 2D / 3D NOR flash memories, 2D / 3D phase change (PCM) flash memories, 2D / 3D magnetic random access memories (magnetic RAM), or 2D / 3D resistive RAM.

[0127] Although in the above embodiments, the input data and / or weight values are divided into an MSB vector and an LSB vector (two vectors), the present invention is not limited thereto. In other possible embodiments of the present invention, the input data and / or weight values can also be divided into more vectors, which is also within the spirit of this disclosure.

[0128] Embodiments of the present disclosure can not only apply the majority decision clustering technique, but also apply other clustering techniques to accelerate accumulation.

[0129] Embodiments of the present disclosure can be applied to, for example but not limited to, AI technologies such as face recognition.

[0130] In embodiments of the present disclosure, the analog-to-digital conversion unit 161 can be a current-mode analog-to-digital conversion unit, or a voltage-mode analog-to-digital conversion unit, or a hybrid-mode analog-to-digital conversion unit.

[0131] Embodiments of the present disclosure can be applied not only to serial MAC operations, but also to parallel MAC operations.

[0132] In summary, although the present disclosure has been disclosed above with embodiments, it is not intended to limit the present invention. Those skilled in the art of the present disclosure can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be defined by the scope of the appended claims.

Claims

1. A memory device, comprising: A memory array including a plurality of memory cells for storing a plurality of weight values within these memory cells of the memory array; A multiplication circuit coupled to the memory array, the multiplication circuit multiplying a plurality of input data by these weight values to obtain a plurality of multiplication results, wherein when performing the multiplication, these memory cells generate a plurality of memory cell currents; A digital accumulation circuit coupled to the multiplication circuit for performing a digital accumulation on these multiplication results; An analog accumulation circuit coupled to the memory array for performing an analog accumulation on these memory cell currents to generate a first multiplication accumulation operation result; And A decision unit coupled to the digital accumulation circuit and the analog accumulation circuit for determining to perform the analog accumulation, the digital accumulation or a hybrid accumulation, wherein, when performing the hybrid accumulation, it is determined whether to trigger the digital accumulation circuit according to the first multiplication accumulation operation result; when performing the hybrid accumulation, when the first multiplication accumulation operation result is lower than a trigger reference value, the digital accumulation circuit is not triggered; when the first multiplication accumulation operation result is higher than the trigger reference value, the digital accumulation circuit is triggered to perform digital accumulation; when the decision unit determines that the memory device performs hybrid accumulation, before the digital accumulation circuit is triggered, the first multiplication accumulation operation result is used as the multiplication accumulation operation result; and after the digital accumulation circuit is triggered, the second multiplication accumulation operation result is used as the multiplication accumulation operation result.

2. The memory device according to claim 1, wherein when performing the analog accumulation, the decision unit can activate the analog accumulation circuit but cannot activate the digital accumulation circuit; when performing the digital accumulation, the decision unit can activate the digital accumulation circuit but cannot activate the analog accumulation circuit; And when performing the hybrid accumulation, the decision unit can activate the digital accumulation circuit and the analog accumulation circuit.

3. The memory device according to claim 1, further comprising: A comparator coupled to the analog accumulation circuit and the digital accumulation circuit for comparing the first multiplication accumulation operation result with a trigger reference value to output a trigger signal to the digital accumulation circuit to trigger the digital accumulation circuit to perform the digital accumulation, wherein, the analog accumulation circuit includes an analog-to-digital conversion unit coupled to the memory array, and the memory cell currents of these memory cells are accumulated and then input to the analog-to-digital conversion unit to be converted into the first multiplication accumulation operation result.

4. The memory device according to claim 1, wherein, The digital accumulation circuit includes: A counting unit coupled to the multiplication circuit for performing bit counting on these multiplication results to obtain an operation result of a second multiplication accumulation.

5. The memory device according to claim 4, further comprising a grouping circuit coupled to the multiplication circuit and the counting unit, the grouping circuit performs a grouping operation on these multiplication results of the multiplication circuit to obtain a plurality of grouping results, and inputs these grouping results to the counting unit.

6. The memory device according to claim 1, each of these input data or each of these weight values has a plurality of bits divided into a plurality of bit vectors; each bit of these bit vectors is converted from a binary form to a unary encoding representation; each bit of these bit vectors in the unary encoding representation is repeated multiple times to form a product expansion; and the multiplication circuit performs a multiplication operation on these input data of the product expansion and these weight values of the product expansion to obtain these multiplication operation results.

7. The memory device according to claim 5, wherein when performing a grouping operation on these multiplication results, the grouping circuit performs a grouping operation on a plurality of multiplication results of these vectors respectively to obtain these grouping results; when performing bit counting, different cumulative weight values are given to these grouping results and then accumulated to obtain the second multiplication accumulation operation result; and the grouping circuit is a majority decision circuit including a plurality of majority decision units.

8. An operation method of a memory device, including: storing a plurality of weight values in a plurality of memory cells of a memory array of the memory device; performing a bit multiplication on a plurality of input data and these weight values to obtain a plurality of multiplication results, wherein when performing multiplication, these memory cells generate a plurality of memory cell currents; and deciding to perform an analog accumulation, a digital accumulation or a hybrid accumulation, wherein when performing the analog accumulation, performing the analog accumulation on these memory cell currents to generate a first multiplication accumulation operation result; when performing the digital accumulation, performing the digital accumulation on these multiplication results to generate a second multiplication accumulation operation result; and when performing the hybrid accumulation, deciding whether to trigger the digital accumulation according to the first multiplication accumulation operation result; when performing the hybrid accumulation, when the first multiplication accumulation operation result is lower than the trigger reference value, the digital accumulation circuit of the memory device is not triggered; when the first multiplication accumulation operation result is higher than the trigger reference value, the digital accumulation circuit is triggered to perform digital accumulation; when deciding to perform hybrid accumulation, before the digital accumulation circuit is triggered, using the first multiplication accumulation operation result as the multiplication accumulation operation result; and, after the digital accumulation circuit is triggered, using the second multiplication accumulation operation result as the multiplication accumulation operation result.

9. The method of operating a memory device according to claim 8, wherein, accumulating these memory cell currents of these memory cells and then performing an analog-to-digital conversion to obtain the first multiplication accumulation operation result; and comparing the first multiplication accumulation operation result with a trigger reference value to decide whether to trigger the digital accumulation.

10. The operation method of the memory device according to claim 8, wherein each of these input data or each of these weight values is divided into a plurality of bit vectors; each bit of these bit vectors is converted from a binary form to a unary encoding representation; each bit of these bit vectors in the unary encoding representation is repeated multiple times to form a product expansion; and performing a multiplication operation on these input data of the product expansion and these weight values of the product expansion to obtain these multiplication operation results; When performing bit accumulation, different accumulation weight values are given to these clustering results and then accumulated to obtain the result of the second multiplication accumulation operation; and perform a clustering operation on these multiplication results and a majority decision operation on these multiplication results.

Citation Information

Patent Citations

  • Mixed signal computer architecture

    US20200134268A1

  • Digital to analog converters and memory devices and related methods

    US9892782B1

  • Memory-integrated neural network

    WO2020159800A1