Memory device and method of operating the same
By implementing MAC operations in the memory array, multiplication circuit and counting units are used for multiplication and bit counting, the output bottlenecks and inefficiency problems existing in the existing AI architecture in the MAC operations are solved, and efficient MAC operations are achieved.
Patent Information
- Application Number
- CN202110792246.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-01
- Filing Date
- 2021-07-13
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-07-13
AI Technical Summary
Existing AI architectures are prone to encounter problems such as output input bottlenecks and inefficiency when performing MAC operations, especially when dealing with multi-bit inputs and multi-bit weight values, which are less efficient.
In-Memory-Computing (IMC) technology is used to couple multiply and bit count multiple input data and weight values by coupling a multiplication circuit and counting unit in the memory array to realize MAC operation.
Reduces complex arithmetic logic units required under a central processing architecture, provides high parallelism, improves the efficiency of MAC operations, and reduces the number of memory cells and power consumption.
Smart Images

Figure CN114153419B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a memory device with in-memory computing (IMC) and an operation method thereof. Background Art
[0002] Artificial intelligence (AI) has become a highly effective solution in many fields. The key operation of AI is to perform multiply-and-accumulation (MAC) operations on a large amount of input data, such as input feature maps, and weight values.
[0003] However, in a natural AI architecture, it is easy to encounter input / output bottlenecks (IO bottlenecks) and inefficient MAC operation flows.
[0004] To achieve high accuracy, MAC operations with multiple-bit inputs and multiple-bit weight values can be performed. However, the input / output bottleneck becomes more serious, and the efficiency will be lower.
[0005] In-memory computing (IMC) can be used to accelerate MAC operations because IMC can reduce the complex arithmetic logic units (ALUs) required in a central processing architecture and provide high parallelism for in-memory MAC operations.
[0006] In the case of non-volatile memory-based IMC (NVM-based IMC), its advantages are, for example, non-volatile storage and reduced data migration. However, the challenges of non-volatile memory-based IMC are the error-bit effect of the most significant bit (MSB), indistinguishable current summation results, and the need for a large number of ADC / DACs, which increases power consumption and chip area.
[0007] Disclosure
[0008] According to an example of the present invention, a memory device is provided, including: a memory array including a plurality of memory cells, which can be used to store a plurality of weight values in these memory cells of the memory array; a multiplication circuit coupled to the memory array, the multiplication circuit multiplies a plurality of input data by these weight values to obtain a plurality of multiplication results; and a counting unit coupled to the multiplication circuit, which performs bit counting on these multiplication results to obtain a multiply-accumulate (MAC) operation result.
[0009] According to another example of the present invention, a method for operating a memory device is provided, including: storing a plurality of weight values in a plurality of memory cells of a memory array of the memory device; performing bit multiplication on a plurality of input data and these weight values to obtain a plurality of multiplication results; and performing bit counting on these multiplication results to obtain a multiply-accumulate (MAC) operation result.
[0010] In order to have a better understanding of the above and other aspects of the present invention, specific embodiments are given below and described in detail in conjunction with the accompanying drawings as follows: Description of the Drawings
[0011] Figure 1 A functional block diagram of a memory device with in-memory computing function according to an embodiment of the present invention is shown.
[0012] Figure 2 A schematic diagram of data mapping according to an embodiment of the present invention is shown.
[0013] Figures 3A to 3C Several examples of data mapping according to an embodiment of the present invention are shown.
[0014] Figure 4A And Figure 4B Schematic diagrams of two exemplary multiplication operations according to an embodiment of the present invention are shown.
[0015] Figure 5A And Figure 5B A schematic diagram of a grouping operation (majority decision operation) and counting according to an embodiment of the present invention is shown.
[0016] Figure 6 A MAC operation flow comparing an embodiment of the present invention with the prior art is shown.
[0017] Figure 7A A flowchart showing the programming of a fixed memory page in an embodiment of the present invention is shown, Figure 7B A flowchart showing the adjustment of the read voltage in an embodiment of the present invention is shown.
[0018] Figure 8 A MAC operation flow according to an embodiment of the present invention is shown.
[0019] Description of Reference Numerals
[0020] 100: Memory device
[0021] 110: Memory array
[0022] 120: Multiplication circuit
[0023] 130: Input / output circuit
[0024] 140: Grouping circuit
[0025] 150: Counting unit
[0026] 111: Memory cell
[0027] 121: Unit multiplication unit
[0028] 121A: Input latch
[0029] 121B: Sense amplifier
[0030] 121C: Output latch
[0031] 121D: Common data latch
[0032] 141: Grouping unit
[0033] 301A, 303A, 301B, 303B, 311A, 313A, 311B, 313B: Bits
[0034] 302, 312, 314: Weight values
[0035] 405, 415: Latches
[0036] 410: Bit line switch
[0037] 420: AND gate
[0038] 710 - 750: Steps
[0039] 810 - 860: Steps Detailed implementation manners
[0040] The technical terms in this specification refer to the customary terms in this technical field. If this specification explains or defines some terms, the explanations of these terms shall prevail according to the explanations or definitions in this specification. Each embodiment of the present invention has one or more technical features. On the premise of possible implementation, those skilled in the art can selectively implement some or all of the technical features in any embodiment, or selectively combine some or all of the technical features in these embodiments.
[0041] Please refer toFigure 1 , which shows a functional block diagram of a memory device 100 with in-memory computing (IMC) function according to an embodiment of the present invention. The memory device 100 with in-memory computing function includes: a memory array 110, a multiplication circuit 120, an input / output circuit 130, a grouping circuit 140, and a counting unit 150. Among them, the memory array 110 and the multiplication circuit 120 are analog, while the grouping circuit 140 and the counting unit 150 are digital.
[0042] The memory array 110 includes a plurality of memory cells 111. In an embodiment of the present invention, the memory cell 111 is, for example but not limited to, a non-volatile memory cell. When performing a MAC operation, the memory cell 111 can be used to store weight values.
[0043] The multiplication circuit 120 is coupled to the memory array 110. The multiplication circuit 120 includes a plurality of unit multiplication units 121. Each unit multiplication unit 121 includes: an input latch 121A, a sense amplifier (SA) 121B, an output latch 121C, and a common data latch (CDL) 121D. The input latch 121A is coupled to the memory array 110. The sense amplifier 121B is coupled to the input latch 121A. The output latch 121C is coupled to the sense amplifier 121B. The common data latch 121D is coupled to the output latch 121C.
[0044] The input / output circuit 130 is coupled to the multiplication circuit 120, the grouping circuit 140, and the counting unit 150, for receiving input data and outputting the output data obtained by the memory device 100.
[0045] The grouping circuit 140 is coupled to the multiplication circuit 120. The grouping circuit 140 includes a plurality of grouping units 141. These grouping units 141 perform a grouping operation on the multiple multiplication results of these unit multiplication units 121 to obtain multiple grouping results. In a possible embodiment of the present invention, the grouping operation can be implemented, for example, by a majority technique, such as a majority function technique. The grouping circuit 140 is implemented by a majority grouping circuit according to the majority function technique, and the grouping unit 141 is implemented by a distributed majority grouping unit, but the present invention is not limited thereto. The grouping technique can be implemented by other similar techniques. In an embodiment of the present invention, the grouping circuit 140 can be selectively provided.
[0046] The counting unit 150 is coupled to the grouping circuit 140 or the multiplication circuit 120. In an embodiment of the present invention, the counting unit 150 is used to perform bitwise counting or bitwise accumulation on the multiplication result of the multiplication circuit 120 to generate a MAC operation result (when the memory device 100 does not include the grouping circuit 140). Alternatively, the counting unit 150 is used to perform bitwise counting or bitwise accumulation on the grouping result (e.g., the majority decision result) of the grouping circuit 140 to generate a MAC operation result (when the memory device 100 includes the grouping circuit 140). In an embodiment of the present invention, the counting unit 150 can be implemented by a known counting circuit, such as, but not limited to, a ripple counter. In the description of the present invention, counting and accumulation basically have the same meaning, and a counter and an accumulator basically have the same meaning.
[0047] Please refer now to Figure 2 , which shows a schematic diagram of data mapping according to an embodiment of the present invention. As Figure 2 shown, take the example that each input data (or each weight value) has 8 bits with N dimensions (N is a positive integer) (it should be noted that the present invention is not limited thereto).
[0048] The following takes the data mapping of the input data as an example for illustration, but it should be noted that the present invention is not limited thereto. The following description also applies to the data mapping of the weight value.
[0049] When the input data is represented in 8-bit binary, the input data (or weight value) is divided into a most significant bit (MSB) vector and a least significant bit (LSB) vector. The MSB vector of the 8-bit input data (or weight value) includes 4 bits B7 to B4, and the LSB vector includes 4 bits B3 to B0.
[0050] Each bit of the MSB vector and the LSB vector of the input data is represented in unary coding (i.e., value format). For example, bit B7 of the MSB vector of the input data can be represented as B70 to B77, bit B6 of the MSB vector of the input data can be represented as B60 to B63, bit B5 of the MSB vector of the input data can be represented as B50 to B51, and bit B4 of the MSB vector of the input data is represented as B4 as well.
[0051] The bits of the MSB vector of the input data represented in one-hot encoding (numerical form) are repeated multiple times with the bits of the LSB vector of the input data to form an unfolding dot product (unFDP). For example, the bits of the MSB of the input data are repeated (2 4 - 1) times. Similarly, the bits of the LSB of the input data are repeated (2 4 - 1) times. In this way, the input data can be represented in the form of an unfolding dot product.
[0052] A multiplication operation is performed on the input data (unfolding dot product) and the weight value to obtain a multiplication operation result.
[0053] For ease of understanding, an example is given below, but it should be understood that it is not used to limit the present invention.
[0054] Now please refer to Figure 3A , which shows an example of one-dimensional data mapping according to an embodiment of the present invention. As Figure 3A shown, the input data = (IN1, IN2) = (2, 1), and the weight value = (We1, We2) = (1, 2). The MSB and LSB of the input data are represented in binary form. Therefore, IN1 = 10, and IN2 = 01. Similarly, the bits of the MSB and LSB of the weight value are represented in binary form. Therefore, We1 = 01, and We2 = 10.
[0055] The MSB and LSB of the input data, and the MSB and LSB of the weight value are encoded in one-hot encoding (numerical form). That is, the MSB of the input data is encoded as 110, the LSB of the input data is encoded as 001. Similarly, the MSB of the weight value is encoded as 001, and the LSB of the weight value is encoded as 110.
[0056] After that, the bits of the MSB (110) of the input data encoded in one-hot encoding are repeated multiple times with the bits of the LSB (001) of the input data encoded in one-hot encoding to form an unfolding dot product (unFDP). For example, the bits of the MSB (110) of the input data are repeated 3 times, so the unfolding dot product of the MSB of the input data is 111111000. The bits of the LSB (001) of the input data are repeated 3 times, so the unfolding dot product of the LSB of the input data is 000000111.
[0057] Perform a MAC operation on the input data (dot product expansion) and the weight values to obtain the MAC operation result. The MAC operation results are: 1*0 = 0, 1*0 = 0, 1*1 = 1, 1*0 = 0, 1*0 = 0, 1*1 = 1, 0*0 = 0, 0*0 = 0, 0*1 = 0, 0*1 = 0, 0*1 = 0, 0*0 = 0, 0*1 = 0, 0*1 = 0, 0*0 = 0, 1*1 = 1, 1*1 = 1, 1*0 = 0. Summing these values gives: 0+0+1+0+0+1+0+0+0+0+0+0+0+0+0+1+1+0 = 4.
[0058] As can be seen from the above, if the input data is i bits and the weight value is j bits (both i and j are positive integers), the number of memory cells used is: (2 i -1)*(2 j -1).
[0059] Now, please refer to Figure 3B , which shows another possible example of data mapping according to an embodiment of the present invention. In Figure 3B , the input data is (IN1) = (2), and the weight value is (We1) = (1). The input data and the weight value are 4 bits.
[0060] When the input data is represented in binary format, IN1 = 0010. Similarly, when the weight value is represented in binary format, We1 = 0001.
[0061] Encode the input data and the weight value into unary encoding (numerical form). For example, the highest bit "0" of the input data is encoded as "00000000", and the lowest bit "0" of the input data is encoded as "0", and so on. Similarly, the highest bit "0" of the weight value is encoded as "00000000", and the lowest bit "1" of the weight value is encoded as "1".
[0062] The bits of the input data encoded into unary encoding are replicated multiple times to become the dot product expansion. For example, the highest bit 301A of the input data encoded into unary encoding is replicated 15 times to become bit 303A; and the lowest bit 301B of the input data encoded into unary encoding is replicated 15 times to become bit 303B.
[0063] The weight value 302 encoded into unary encoding is also replicated 15 times to be represented as the dot product expansion.
[0064] Perform a multiplication operation on the input data represented as the product expansion and the weight value represented as the dot product expansion to produce the MAC operation result. Specifically, bit 303A of the input data is multiplied by the weight value 302; bit 303B of the input data is multiplied by the weight value 302, and so on. Summing the multiplication values can produce the MAC operation result ("2").
[0065] Now, please refer to Figure 3C , which shows another possible example of data mapping according to an embodiment of the present invention. In Figure 3C , the input data is (IN1) = (1), and the weight value is (We1) = (5). The input data and the weight value are 4 bits.
[0066] When the input data is represented in binary format, IN1 = 0001. Similarly, when the weight value is represented in binary format, We1 = 0101.
[0067] Encode the input data and the weight value into unary encoding (numeric form).
[0068] The bits of the input data encoded into unary encoding are replicated multiple times to form a dot product expansion. In Figure 3C , when replicating the bits of the input data and the bits of the weight value, the bit "0" is added. For example, the highest bit 311A of the input data encoded into unary encoding is replicated 15 times and the bit "0" is added to become bit 313A; and, the lowest bit 311B of the input data encoded into unary encoding is replicated 15 times and the bit "0" is added to become bit 313B. Thus, the input data is represented as a dot product expansion.
[0069] Similarly, the weight value 312 encoded into unary encoding is also replicated 15 times, and an additional bit "0" is added to each weight value 314. Thus, the weight value is represented as a dot product expansion.
[0070] Perform a multiplication operation on the input data represented as a dot product expansion and the weight value represented as a dot product expansion to produce a MAC operation result. Specifically, bit 313A of the input data is multiplied by weight value 314; bit 313B of the input data is multiplied by weight value 314, and so on. Summing up the multiplication values can produce a MAC operation result ("5").
[0071] In the conventional technology, when performing a MAC operation on 8-bit input data and 8-bit weight values, if the direct MAC algorithm is used, the number of memory units used is 255 * 255 * 512 = 33,292,822.
[0072] On the contrary, as described above, in the embodiment of the present invention, when performing a MAC operation on 8-bit input data and 8-bit weight values, the number of memory units used is 15 * 15 * 512 * 2 = 115,200 * 2 = 230,400. Therefore, the number of memory units used in the embodiment of the present invention during the MAC operation is approximately 0.7% of that in the conventional technology.
[0073] In an embodiment of the present invention, by using the unFDP-based data mapping, the number of memory cells used in the operation can be reduced, so that the operation cost can be reduced, and the error correction code (ECC) cost can be reduced. In addition, the fail-bit effect can also be tolerated.
[0074] Please refer again to Figure 1 . In an embodiment of the present invention, when performing a multiplication operation, the weight values (transduction values) are stored in these memory cells 111 of the memory array 110, and the input data (voltage) is read by the input / output circuit 130 and transmitted to the common data latch 121D. The common data latch 121D transmits the input data to the input latch 121A.
[0075] To better understand the multiplication operation of the embodiment of the present invention, please now refer to Figure 4A and Figure 4B , which show two schematic diagrams of exemplary multiplication operations of the embodiment of the present invention. Figure 4A Applied to a memory device that supports the selected bit-line read function, Figure 4B Applied to a memory device that does not support the selected bit-line read function. Figure 4A In Figure 4B , the input latch 121A includes a latch (first latch) 405 and a bit-line switch 410; and,
[0076] As Figure 4A shown, the weight values are represented in one-hot encoding (numerical form) (as Figure 2 ). Therefore, the most significant bit of the weight value is stored in 8 memory cells 111, the second most significant bit of the weight value is stored in 4 memory cells 111, the third most significant bit of the weight value is stored in 2 memory cells 111, and the least significant bit of the weight value is stored in 1 memory cell 111.
[0077] Similarly, the input data is represented in one-hot encoding (numerical form) (as Figure 2 ), so that the most significant bit of the input data is stored in 8 common data latches 121D, the second most significant bit of the input data is stored in 4 common data latches 121D, the third most significant bit of the input data is stored in 2 common data latches 121D, and the least significant bit of the input data is stored in 1 common data latch 121D. The input data is sent from the common data latch 121D to the latch 405.
[0078] In Figure 4AAmong them, these multiple bit-line switches 410 are coupled between the memory cell 111 and the sense amplifier 121B. The bit-line switch 410 is controlled by the latch 405. For example, when the latch 405 outputs a bit 1, the bit-line switch 410 is turned on, and when the latch 405 outputs a bit 0, the bit-line switch 410 is turned off.
[0079] In addition, when the weight value in the memory cell 111 is a bit 1 and the bit-line switch 410 is turned on (the input data is a bit 1), the sense amplifier 121B will sense the memory cell current to generate a multiplication result of "1". When the weight value in the memory cell 111 is a bit 0 and the bit-line switch 410 is turned on (the input data is a bit 1), the sense amplifier 121B does not sense the memory cell current. When the weight value in the memory cell 111 is a bit 1 and the bit-line switch 410 is turned off (the input data is a bit 0), the sense amplifier 121B does not sense the memory cell current to generate a multiplication result of "0". When the weight value in the memory cell 111 is a bit 0 and the bit-line switch 410 is turned off (the input data is a bit 0), the sense amplifier 121B does not sense the memory cell current.
[0080] That is to say, via Figure 4A the layout, when the input data is a bit 1 and the weight value is a bit 1, the sense amplifier 121B senses the memory cell current to generate a multiplication result of "1". In other cases, the sense amplifier 121B does not sense the memory cell current to generate a multiplication result of "0".
[0081] In Figure 4B , the input data is sent from the common data latch 121D to the latch 415. One end of the AND gate 420 receives the sensing result (i.e., the weight value) of the sense amplifier 121B, and the other end receives the output bit of the latch 415 (i.e., the input data). When the weight value stored in the memory cell 111 is a bit 1, the sensing result of the sense amplifier 121B is logic high (sensing the memory cell current); when the weight value stored in the memory cell 111 is a bit 0, the sensing result of the sense amplifier 121B is logic low (not sensing the memory cell current).
[0082] When the latch 415 outputs a bit 1 (i.e., the input data is a bit 1) and the sensing result of the sense amplifier 121B is logic high (i.e., the weight value is a bit 1), the AND gate 420 outputs a bit 1 to generate a multiplication result of "1", and sends it to the grouping circuit 140 or the counting unit 150. In other cases, the AND gate 420 outputs a bit 0 to generate a multiplication result of "0", and sends it to the grouping circuit 140 or the counting unit 150.
[0083] Figure 4B The embodiments of can be applied not only to non-volatile memories but also to volatile memories.
[0084] In an embodiment of the present invention, when performing a multiplication operation, the selected bit line read (SBL-read) instruction can be reused. Therefore, the embodiment of the present invention can reduce the variation influence caused by single-bit representation.
[0085] Please now refer to Figure 5A , which shows a schematic diagram of a grouping operation (majority decision operation) and bitwise counting according to an embodiment of the present invention. As Figure 5A shown, the reference symbol GM1 represents the first multiplication result obtained after performing bitwise multiplication on the first MSB vector of the input data and the weight value; the reference symbol GM2 represents the second multiplication result obtained after performing bitwise multiplication on the second MSB vector of the input data and the weight value; the reference symbol GM3 represents the third multiplication result obtained after performing bitwise multiplication on the third MSB vector of the input data and the weight value; the reference symbol GL represents the fourth multiplication result obtained after performing bitwise multiplication on the LSB of the input data and the weight value. After the grouping operation (majority decision operation), the grouping result of the first multiplication result GM1 is the first grouping result CB1 (whose cumulative weight is 2 2 ); the grouping result of the second multiplication result GM2 is the second grouping result CB2 (whose cumulative weight is 2 2 ); the grouping result of the third multiplication result GM3 is the third grouping result CB3 (whose cumulative weight is 2 2 ); and the grouping result of the fourth multiplication result GL is the fourth grouping result CB4 (whose cumulative weight is 2 0 ).
[0086] Figure 5B Shows Figure 3C the cumulative example. Please refer to Figure 3C and Figure 5B . As Figure 5B shown, the bit 313B of the input data ( Figure 3C ) is multiplied by the weight value 314. The first four bits ("0000") of the multiplication result generated by multiplying the bit 313B of the input data ( Figure 3C ) by the weight value 314 are grouped as the first multiplication result "GM1". Similarly, the fifth to eighth bits ("0000") of the multiplication result generated by multiplying the bit 313B of the input data ( Figure 3C ) by the weight value 314 are grouped as the second multiplication result "GM2". From the input data ( Figure 3CThe ninth to twelfth bits ("1111") of the multiplication result generated by multiplying bit 313B of ( Figure 3C ) by the weight value 314 are grouped as the third multiplication result "GM3". From the input data (
[0087] The thirteenth to sixteenth bits ("0010") of the multiplication result generated by multiplying bit 313B of ) by the weight value 314 are directly counted. 2 ) The first grouped result CB1 is "0" (its cumulative weight is 2 2 ) The second grouped result CB2 is "0" (its cumulative weight is 2 2 ) The third grouped result CB3 is "1" (its cumulative weight is 2 Figure 5B As shown, for example, the MAC operation result is CB1 * 2 2 + CB2 * 2 2 + CB3 * 2 2 + CB4 * 2 0 = 0 * 2 2 + 0 * 2 2 + 1 * 2 2 + 1 * 2 0 = 00000000000000000000000000000101 = 5.
[0088] In an embodiment of the present invention, the grouping principle (majority decision principle) can be as follows:
[0089] Group bit Grouping result (majority decision result) 1111 (Condition A) 1 1110 (Condition B) 1 1100 (Condition C) 1 or 0 1000 (Condition D) 0 0000 (Condition E) 0
[0090] In the above table, for case A, since all groups are correct ("1111" has no error bits), the majority decision result is 1. For case E, since all groups are correct ("0000" has no error bits), the majority decision result is 0.
[0091] For case B, since there is 1 wrong bit in the group ("0" in "1110" is wrong), through majority decision, "1110" can be determined as "1". For case D, since there is 1 wrong bit in the group ("1" in "0001" is wrong), through majority decision, "0001" can be determined as "0".
[0092] For case C, there are 2 wrong bits in the group ("00" in "1100" is wrong, or "11" in "1100" is wrong), through majority decision, "1100" can be determined as "1" or "0".
[0093] Therefore, in the embodiments of the present invention, the number of error bits can be reduced through the clustering (majority decision) function.
[0094] The clustering result of the clustering circuit 140 is input to the counting unit 150 for bit counting.
[0095] When counting, the counting result of the multiplication operation result of the MSB vector is accumulated with the counting result of the multiplication operation result of the LSB vector. For Figure 5A example, two accumulators are used. The first accumulator is assigned a higher accumulation weight value (for example, 2 2 ). The first accumulator accumulates: (1) "the clustering (majority decision) result obtained by clustering (majority decision) the multiplication operation result GM1: 1 bit" plus "the clustering (majority decision) result of clustering (majority decision) the multiplication operation result GM2: 1 bit" plus "the clustering (majority decision) result of clustering (majority decision) the multiplication operation result GM3: 1 bit". The counting result obtained by the first accumulator is then multiplied by the higher accumulation weight value (for example, 2 2 ). The second accumulator is assigned a lower accumulation weight value (for example, 2 0 ). The second accumulator directly counts the multiplication operation result GL (multiple bits). Adding the two accumulation results weighted by the accumulation weight gives the MAC result. For example, the clustering result obtained by clustering the multiplication operation result GM1 is 1 (1 bit), the clustering result of clustering the multiplication operation result GM2 is 0 (1 bit), and the clustering result of clustering the multiplication operation result GM3 is 1 (1 bit). The counting result (1 + 0 + 1) obtained by the first accumulator is multiplied by 2 2 , which is equal to 2 * 2 2 = 8. For the multiplication operation result GL of 4 (3 bits), it can be directly counted. Adding the two accumulation results weighted by the accumulation weight gives the MAC result of 8 + 4 = 12.
[0096] As can be seen from the above, in the embodiments of the present invention, when counting or accumulating, since the input data has been expanded into the unFDP form, the data stored in the common data latch can be clustered (i.e., divided into the MSB vector and the LSB vector), and the number of error bits in the MSB vector / LSB vector can be reduced through the clustering mechanism (majority decision mechanism).
[0097] In addition, in the embodiments of the present invention, even when using a conventional accumulator (counter), the counting / accumulation time can still be reduced because the embodiments of the present invention use digital counting instructions (error bit counting) and assign different accumulation weights to the accumulation results of different vectors (MSB vector and LSB vector). For example, the accumulation operation time can be reduced to about 40%.
[0098] Figure 6 Show the MC operation processes comparing an embodiment of the present invention with the prior art. For the MAC operation processes of the embodiment of the present invention and the prior art, the input data is transmitted to the memory device. At the same time, bit line setting and word line setting are performed. After the bit line setting is completed, sensing is performed. An accumulation operation is carried out. And the result of the accumulation operation is transmitted back. The above operations are repeated until all the input data is processed.
[0099] From Figure 6 it can be seen that the MAC operation of the embodiment of the present invention can be divided into two sub-operation types. The first sub-operation type is the multiplication operation, which multiplies the input data by the weight value and is performed according to the selected bit line read instruction. The second sub-operation type is accumulation (data counting), especially error bit counting. In other possible embodiments of the present invention, more counting units can be added to accelerate the counting / accumulation operation.
[0100] Compared with the prior art, in the embodiments of the present invention, the accumulation operation is faster, so the MAC operation can be accelerated.
[0101] In addition, in the embodiments of the present invention, the read voltage can also be adjusted. Figure 7A Show the flowchart of programming a fixed memory page in the embodiment of the present invention, Figure 7B Show the flowchart of adjusting the read voltage in the embodiment of the present invention.
[0102] In Figure 7A , in step 710, a known input data is programmed into a fixed memory page, where the bit ratio of the known input data is: 50% is bit 0 and 50% is bit 1.
[0103] In Figure 7B , in step 720, the fixed memory page is read and the ratio of bit 1 is counted. In step 730, it is determined whether the ratio of bit 1 is close to 50%. If step 730 is no, the process proceeds to step 740. If step 730 is yes, the process proceeds to step 750.
[0104] In step 740, if the ratio of bit 1 is less than 50%, increase the read voltage; and if the ratio of bit 1 is greater than 50%, decrease the read voltage. After step 740 ends, the process returns to step 720.
[0105] In step 750, record the current read voltage for use in subsequent read operations.
[0106] As is known, the read voltage will affect the reading of bit 1. Therefore, in the embodiments of the present invention, the read voltage can be periodically corrected according to operating conditions (such as but not limited to, programming cycles, temperature, or read interference, etc.) to maintain high accuracy and reliability.
[0107] Figure 8 Show the MAC operation flow according to an embodiment of the present invention. As Figure 8 shown, in step 810, periodically check the read voltage. If the read voltage needs to be adjusted, it can be adjusted according to the Figure 7B flow.
[0108] In step 820, store the input data in the common data latch 121D.
[0109] In step 830, transfer the input data from the common data latch 121D to the input latch 121A.
[0110] In step 840, perform a multiplication operation under the condition of supporting or not supporting the selected bit line read instruction.
[0111] In step 850, perform accumulation.
[0112] In step 860, output the MAC operation result (for example, output through the input / output circuit 30).
[0113] The embodiments of the present invention can be applied to NAND flash memory, or memory devices sensitive to retention and thermal changes, such as but not limited to, NOR flash memory, phase change (PCM) flash memory, magnetic random access memory (magnetic RAM), or resistive RAM.
[0114] The embodiments of the present invention can be applied to 3D memory and 2D memory, such as but not limited to, 2D / 3D NAND flash memory, 2D / 3D NOR flash memory, 2D / 3D phase change (PCM) flash memory, 2D / 3D magnetic random access memory (magnetic RAM), or 2D / 3D resistive RAM.
[0115] Although in the above embodiments, the input data and / or weight values are divided into an MSB vector and an LSB vector (two vectors), the present invention is not limited thereto. In other possible embodiments of the present invention, the input data and / or weight values may also be divided into more vectors, which is also within the protection scope of the present invention.
[0116] The embodiments of the present invention can not only apply the majority decision clustering technology, but also apply other clustering technologies to accelerate accumulation.
[0117] The embodiments of the present invention can be applied to, for example but not limited to, AI technologies such as face recognition.
[0118] In summary, although the present invention has been disclosed above in embodiments, it is not intended to limit the present invention. Those skilled in the art in the technical field to which the present invention pertains can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the appended claims.
Claims
1. A memory device, characterized in that, Comprising: A memory array including a plurality of memory cells, which can be used to store a plurality of weight values in these memory cells of the memory array; A multiplication circuit coupled to the memory array, the multiplication circuit multiplying a plurality of input data and these weight values to obtain a plurality of multiplication results; And A counting unit coupled to the multiplication circuit, which performs bit counting on these multiplication results to obtain a multiply-accumulate operation (MAC) operation result; Wherein, a plurality of bits of these input data or these weight values are divided into a plurality of bit vectors; each bit of these bit vectors is converted from a binary form to a unary code representation; each bit of these bit vectors represented in unary code is repeated multiple times to become a dot product expansion; the multiplication circuit multiplies these input data of the dot product expansion and these weight values of the dot product expansion to obtain these multiplication operation results.
2. The memory device according to claim 1, characterized in that, The multiplication circuit includes a plurality of unit multiplication units, and each unit multiplication unit includes: An input latch coupled to the memory array, A sense amplifier coupled to the input latch, An output latch coupled to the sense amplifier, and A common data latch coupled to the output latch, Wherein, The common data latch transmits the input data to the input latch.
3. The memory device according to claim 2, characterized in that, The unit multiplication unit generates these multiplication results and inputs them to the counting unit.
4. The memory device according to claim 2, characterized in that, It further includes a grouping circuit coupled to the multiplication circuit and the counting unit, the grouping circuit performs a grouping operation on these multiplication results of the multiplication circuit to obtain a plurality of grouping results, and inputs these grouping results to the counting unit, wherein, the unit multiplication unit generates the multiplication results and inputs them to the grouping circuit.
5. The memory device according to claim 4, characterized in that, It further includes: An input / output circuit coupled to the multiplication circuit and the counting unit, which is used to receive these input data and output the multiply-accumulate operation result obtained by the memory device; Wherein, the grouping circuit includes a plurality of grouping units, and these grouping units perform a grouping operation on these multiplication results to obtain these grouping results; The memory array and the multiplication circuit are analog, while the grouping circuit and the counting unit are digital.
6. The memory device according to claim 2, characterized in that, Each input latch further includes a first latch and a bit line switch, these first latches receive the input data of the common data latch, these bit line switches are coupled between these memory cells and these sense amplifiers, these bit line switches are controlled by the input data stored in these first latches to control whether the weight values in these memory cells are transmitted to these sense amplifiers, and, these sense amplifiers generate these multiplication results by sensing the input of these bit line switches.
7. The memory device according to claim 2, characterized in that, Each input latch further includes a second latch and a logic gate, these second latches receive the input data of the common data latch, these sense amplifiers sense the weight values stored in these memory cells, and these logic gates generate these multiplication results according to the input data of these second latches and the weight values transmitted through these sense amplifiers.
8. The memory device according to claim 5, characterized in that, When performing a grouping operation on these multiplication results, the grouping circuit performs a grouping operation on multiple multiplication results of these vectors respectively to obtain these grouping results; When performing bit counting, different cumulative weight values are given to these grouping results and then accumulated to obtain the multiplication accumulation operation result; and The grouping circuit is a majority decision circuit and includes a plurality of majority decision units.
9. A method for operating a memory device, characterized in that, Including: Storing a plurality of weight values in a plurality of memory cells of a memory array of the memory device; Performing bit multiplication on a plurality of input data and these weight values to obtain a plurality of multiplication results; and Performing bit counting on these multiplication results to obtain a multiplication accumulation (MAC) operation result, dividing these input data or these weight values into a plurality of bit vectors; converting each bit of these bit vectors from a binary form to a unary code representation; repeating each bit of these bit vectors represented in the unary code form multiple times to form a dot product expansion; performing a multiplication operation on these input data of the dot product expansion and these weight values of the dot product expansion to obtain these multiplication operation results.
10. The method for operating a memory device according to claim 9, characterized in that, When performing bit accumulation, different cumulative weight values are given to these grouping results and then accumulated to obtain the multiplication accumulation operation result; and Performing a grouping operation on these multiplication results is performing a majority decision operation on these multiplication results.
Citation Information
Patent Citations
Neuron circuit
CN111210015A