integrated circuit

In-memory and near-memory operations are realized through memory arrays and page buffers in integrated circuits, which solves the performance bottleneck caused by the separation of data storage units and processing units in the Van Newman-type architecture, and improves the efficiency of processing huge amounts of data.

CN112750487BActive Publication Date: 2025-08-26MACRONIX INTERNATIONAL CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201911067337.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-29
Filing Date
2019-11-04
Publication Date
2025-08-26
Estimated Expiration
2039-11-04

AI Technical Summary

Technical Problem

In the Van Newman-type architecture, the data round-trip time and energy consumption caused by the separation of the data storage unit and the data processing unit, especially when processing huge amounts of data, there is a performance bottleneck.

Method used

An integrated circuit is designed that includes a memory array and a page buffer that can write weights in memory mode and perform the sum function of the product terms through in-memory and near-memory operations in operation mode, reducing the round trip between cells.

Benefits of technology

Significantly improve the instruction cycle, avoid performance bottlenecks, and is suitable for computing huge amounts of data, especially in artificial intelligence and machine learning systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112750487B_ABST
    Figure CN112750487B_ABST
Patent Text Reader

Abstract

The present invention discloses an integrated circuit comprising a memory array, a plurality of word lines, a plurality of bit lines, and a page buffer. The memory array comprises a plurality of memory cells, each of which is configured to be written with a weight. The plurality of word lines are respectively connected to a column of memory cells among the plurality of memory cells. The plurality of bit lines are respectively connected to a column of memory cells connected in series with each other among the plurality of memory cells. A plurality of the plurality of bit lines in a block of the memory array or a plurality of the plurality of word lines in a plurality of blocks of the memory array are configured to receive a plurality of input voltages, and the memory cells receiving the plurality of input voltages are configured to multiply the write weight with the received input voltage. The page buffer is coupled to the memory array and is configured to sense a plurality of products of the weight and the input voltage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an integrated circuit and a calculation method thereof, and in particular to a memory circuit. Background Art

[0002] In calculators designed using the Von Neumann architecture, the data storage unit and the data processing unit are separated. Data must be transferred between the data storage unit and the data processing unit via input / output (I / O) ports and buses, which consumes time and energy. Furthermore, when processing massive amounts of data, the data transfer between these units creates a bottleneck in processing performance. In recent years, with the rise of artificial intelligence (AI) technology, the amount of data that calculators must process has increased significantly, making this performance bottleneck increasingly severe. Summary of the Invention

[0003] The present invention provides an integrated circuit which can operate in a memory mode and a calculation mode.

[0004] The integrated circuit of the present invention includes: a memory array, including multiple memory cells, each configured to have weights written thereto; multiple word lines and multiple bit lines, wherein the multiple word lines are respectively connected to a column of memory cells among the multiple memory cells, and the multiple bit lines are respectively connected to a column of memory cells connected in series with each other among the multiple memory cells, multiple of the multiple bit lines in a block of the memory array or multiple of the multiple word lines in multiple blocks of the memory array are configured to receive multiple input voltages, and multiple of the multiple memory cells receiving the multiple input voltages are configured to multiply multiple of the multiple weights written thereto by the multiple input voltages received; and a page buffer coupled to the memory array and configured to sense multiple products of the multiple weights and the multiple input voltages.

[0005] In some embodiments, the plurality of the plurality of bit lines in the block receive the plurality of input voltages, and one of the plurality of word lines in the block is configured to receive a read voltage while the others of the plurality of word lines in the block are configured to receive a pass voltage.

[0006] In some embodiments, memory cells corresponding to the ones of the bit lines and the one of the word lines are configured to multiply the ones of the stored weights with the received input voltages and generate the products.

[0007] In some embodiments, the integrated circuit further includes a counter, wherein the counter is coupled to the page buffer and configured to sum the plurality of products.

[0008] In some embodiments, at least two of the plurality of input voltages are different from each other.

[0009] In some embodiments, the plurality of input voltages are identical to one another.

[0010] In some embodiments, the page buffer includes a first cache and a second cache. The first cache is configured to receive a plurality of first logic signals converted by multiplying the plurality of weights by the plurality of input voltages and is pre-written with a plurality of second logic signals converted by a plurality of additional input voltages. The second cache is configured to multiply the plurality of first logic signals by the plurality of second logic signals and accumulate the plurality of products of the plurality of first logic signals and the plurality of second logic signals.

[0011] In some embodiments, at least two of the plurality of additional input voltages are different from one another and are converted to different logic signals.

[0012] In some embodiments, the plurality of word lines in the plurality of blocks are configured to receive the plurality of input voltages, the word line of one of the plurality of blocks is electrically isolated from the word line of another of the plurality of blocks, the plurality of bit lines are respectively shared by the plurality of blocks of the memory array, and one of the plurality of bit lines is configured to receive a read voltage, while the other plurality of bit lines are configured to receive a pass voltage.

[0013] In some embodiments, memory cells corresponding to the plurality of the plurality of word lines and the one of the plurality of bit lines are configured to multiply the plurality of stored weights with the received plurality of input voltages and generate the plurality of products.

[0014] In some embodiments, the plurality of products are summed across the one of the plurality of bit lines.

[0015] In some embodiments, memory cells corresponding to the plurality of the plurality of word lines and the one of the plurality of bit lines have a starting voltage greater than or equal to 0V.

[0016] In some embodiments, the memory array is a NAND flash memory array, and the plurality of memory cells are a plurality of flash memory cells.

[0017] In some embodiments, the number of the page buffers is plural, and a block of the memory array has a plurality of sub-blocks, each of the sub-blocks being coupled to one of the plurality of page buffers.

[0018] The operation method of the integrated circuit of the present invention includes: performing at least one programming operation to write multiple weights into the multiple memory cells respectively; applying multiple input voltages to multiple of the multiple bit lines in a block of the memory array or multiple of the multiple word lines in multiple blocks of the memory array, wherein the memory cells receiving the multiple input voltages are configured to multiply multiple of the stored multiple weights with the received multiple input voltages to obtain multiple products; and summing the multiple products via the page buffer or via one of the multiple bit lines.

[0019] In some embodiments, the step of applying the plurality of input voltages and the step of summing the plurality of products form a loop, and the operation method of the integrated circuit includes performing the loop a plurality of times.

[0020] In some embodiments, the step of applying the multiple input voltages in one of the multiple cycles precedes the step of applying the multiple input voltages in a subsequent one of the multiple cycles.

[0021] In some embodiments, the step of applying the plurality of input voltages in one of the plurality of cycles overlaps in time with the step of summing the plurality of products in a previous one of the plurality of cycles.

[0022] In some embodiments, the plurality of input voltages are applied to the plurality of the plurality of bit lines in the block, and the page buffer is configured to sum the plurality of products.

[0023] In some embodiments, the plurality of input voltages are applied to the plurality of the plurality of blocks of the plurality of word lines, and the plurality of products are summed via the one of the plurality of bit lines.

[0024] Based on the above, the integrated circuit of the present invention can operate in both memory mode and computation mode. The integrated circuit includes a memory array, such as a NAND flash memory array. The integrated circuit can execute a sum-of-products function and can be used in learning programs for artificial intelligence applications, neuromorphic computing systems, and machine learning systems. In memory mode, weights are written to memory cells in the memory array. In computation mode, the stored weights are multiplied by an input voltage transmitted to the memory cells via a bit line or word line, and the product of the weights and the input voltage is accumulated. Compared to the van Neumann architecture, which performs computations in a data processing unit (such as a central processing unit) separate from a data storage unit (such as a memory integrated circuit), the integrated circuit of the present invention can operate in both memory mode and computation mode. Therefore, data no longer needs to be transferred back and forth between the data processing unit and the data storage unit, and instruction cycle time can be significantly improved. In particular, the page buffer used to write weights to the memory cells and receive the product of the weights and the input voltage is coupled to the memory array via a large number of highly parallel bit lines, resulting in a very high bandwidth. Therefore, integrated circuits can be used for computing massive amounts of data and may not encounter performance bottlenecks like those in the Van Neumann architecture.

[0025] In order to make the above features and advantages of the present invention more clearly understood, embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1A is a schematic diagram of an integrated circuit according to some embodiments of the present invention.

[0027] Figure 1B yes Figure 1A A flow chart of an operation method of an integrated circuit is exemplarily depicted.

[0028] Figure 2 is a schematic diagram of an integrated circuit according to some embodiments of the present invention.

[0029] Figure 3 is a schematic diagram of an integrated circuit according to some embodiments of the present invention.

[0030] Figure 4 is a schematic diagram of an integrated circuit according to some embodiments of the present invention.

[0031]

Explanation of symbols

[0032] 10, 10a, 10b, 20: Integrated circuits

[0033] 100, 100', 200: Memory array

[0034] BL: bit line

[0035] BK1, BK2: Block

[0036] BS: Inter-subblock bus system

[0037] CA1: First cache

[0038] CA2: Second cache

[0039] CT: Counter

[0040] GSL: Ground Select Line

[0041] GST: Ground Select Transistor

[0042] MC: Memory Cell

[0043] PB, PB': Page Buffer

[0044] S100, S102, S1021, S1022, S102 n ,S104,S1041,S1042,S104 n :step

[0045] SL: Source line

[0046] SSL: string select line

[0047] SST: string select transistor

[0048] TL: Subblock

[0049] W i , W1, W2: weights

[0050] WL, WL1, WL2, WL3, WLn: word lines

[0051] X, X i , X1, X2: input voltage DETAILED DESCRIPTION

[0052] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0053] Figure 1A is a schematic diagram of an integrated circuit 10 according to some embodiments of the present invention. Figure 1B yes Figure 1A A flow chart of an operation method of the integrated circuit 10 is exemplarily shown.

[0054] Please refer to Figure 1A, the integrated circuit 10 may be a memory circuit, such as a non-volatile memory circuit. In some embodiments, the integrated circuit 10 is a NAND flash memory circuit and may be used in applications such as neuromorphic computing systems, machine learning systems, and artificial intelligence that include performing multiply-and-accumulate (MAC) operations. The MAC operation can be represented by a sum-of-products function, as shown in Equation (1):

[0055]

[0056] In formula (1), the accumulated product terms are respectively the input value X i With weight W i The weight of the accumulated multiple product terms W i The values ​​can be different from each other. The weights can be specified as a set of constants, and the sum of multiple product terms changes as the input values ​​change. Furthermore, when the algorithm executes a learning procedure, the weights of multiple learning procedures can be different from each other, and learning is performed from the sum of multiple product terms. For example, the weights can be obtained through remote training performed on a computer and downloaded to the integrated circuit 10. These weights can be downloaded and updated in the integrated circuit 10 as the remote training mode changes.

[0057] The integrated circuit 10 includes a memory array 100. The memory array 100 has a plurality of memory cells MC. In some embodiments, the memory array 100 is a three-dimensional memory array. Figure 1A As shown, the memory cells MC of each block are configured to have a plurality of columns (or strings) and a plurality of rows (or pages). In an embodiment where the integrated circuit 10 is a NAND flash memory circuit, the memory cells MC may be floating gate transistors, semiconductor-oxide-nitride-oxide-semiconductor (SONOS) transistors or the like. The memory cells MC of each column (or string) are connected in series and are connected between a bit line BL and a source line SL. In some embodiments, the memory cells MC of the majority of columns (or strings) share a source line SL. On the other hand, a plurality of word lines WL (such as Figure 1AAs shown, for example, one of word lines WL1, WL2, WL3, ..., and WLn is connected to each row (or page) of memory cells MC. In some embodiments, the memory array 100 further includes a string select transistor SST and a ground select transistor GST. In these embodiments, each column (or string) of memory cells MC is connected between the string select transistor SST and the ground select transistor GST. The plurality of string select transistors SST can be respectively connected to one of the plurality of bit lines BL, while the plurality of ground select transistors GST can be connected to the source line SL. In addition, a string select line SSL connects the string select transistors SST in a row, and a ground select line GSL connects the ground select transistors GST in a row.

[0058] The integrated circuit 10 can operate in a memory mode and an operation mode. In the memory mode, data can be written to or read from the memory cells MC using programming operations, erase operations, and read operations. Peripheral circuits coupled to the memory array 100 can support the aforementioned programming operations, erase operations, and read operations. For example, the peripheral circuits may include a decoder (not shown), a page buffer PB, and the like. During a programming operation, a word line WL and some bit lines BL are selected, and data is written to the memory cells MC corresponding to the selected word line WL and bit lines BL via the page buffer PB and the selected word line WL. On the other hand, during a read operation, data is read from the memory cells MC corresponding to the selected word line WL and bit lines BL via the page buffer PB and the selected bit line BL. In some embodiments, each programming operation writes data to a page of memory cells MC, and each read operation reads data from a page of memory cells MC. In an embodiment where the integrated circuit 10 is configured to perform the sum of product terms function (as shown in equation (1)), the weight W is set by multiple times of the above-mentioned programmed operations. i (For example, including Figure 1A The weights W1 and W2 shown are written into the plurality of memory cells MC. i Determine the conductance or transconductance of these memory cells MC. In some embodiments, the memory cells MC are programmed using a binary mode, and the weight W i In an alternative embodiment, the weights W i The signal is stored as a multi-bit level or an analog code. For example, the multi-bit level may be N levels, where N is a positive integer greater than 2.

[0059] In the operation mode of the integrated circuit 10, the weight W stored in the memory cell MC is i With input voltage X i Multiply and accumulate multiple weights W i Corresponding to input voltage X i In some embodiments, a plurality of bit lines BL of a block of the memory array 100 are configured to receive an input voltage X i (like Figure 1A As shown, for example, it includes input voltage X1 and input voltage X2. In some embodiments, the multiple input voltages X1 and X2 received by these bit lines BL are i has a specific distribution (pattern), and these input voltages X i For example, by applying multiple input voltages X in a dual-bit mode i , and wherein the input voltage X1 is a high logic level "1", and the input voltage X2 is a low logic level "0". Alternatively, a plurality of input voltages X i The voltage applied may be a multi-bit level (e.g., N levels, where N is a positive integer greater than 2) or an analog code. One of the multiple word lines WL of a block of the memory array 100 is selected to receive a read voltage, while the other word lines WL of the block of the memory array 100 receive a pass voltage. In some embodiments, a page of memory cells MC connected to the selected word line WL receives the read voltage and is turned on. In addition, when the bit line BL inputs the voltage X i When inputted to these conductive memory cells MC, the weights W stored in these conductive memory cells MC are i The corresponding input voltage X i Multiply. At input voltage X i In the embodiment of the invention, the weight W stored in the memory cell MC is transmitted to the memory cell MC via the bit line BL. i It can be regarded as the conductance of the memory cell, and the weight W i With input voltage X i The product is output as current. i With input voltage X i The multiplication occurs in the memory array 100. This multiplication operation can be regarded as an in-memory computing.

[0060] In some embodiments, the multiple weights W i Corresponding to input voltage X iThe multiple products of the current signals are output to a page buffer PB coupled to the memory array 100 via the bit lines BL. A sense amplifier (not shown) in the page buffer PB can be configured to sense the current signals outputted therefrom. In addition, a counter CT coupled to the page buffer PB can be configured to sum the current signals outputted therefrom (i.e., the multiple weights W i Corresponding to input voltage X i multiple products of ). Although Figure 1A The page buffer PB and counter CT are shown as separate components, but they may alternatively be integrated into a single component. The page buffer PB and counter CT are disposed in an area surrounding the memory array 100 and are immediately adjacent to the memory array 100. Therefore, the addition operation performed using the page buffer PB and counter CT can be considered a near-memory operation.

[0061] So far, the memory operation (multiple weights W i Corresponding to input voltage X i multiplication) and near memory operations (multiplying multiple weights W i Corresponding to input voltage X i The sum of the products is performed by adding multiple products (as shown in formula (1)). Compared with the Van Neumann architecture, which performs operations in a data processing unit (such as a central processing unit) separated from a data storage unit (such as a memory integrated circuit), the integrated circuit 10 of the present invention can operate in both memory mode and operation mode. Therefore, data no longer needs to be transferred between the data processing unit and the data storage unit, and the instruction cycle can be significantly improved. In particular, the weight W is used to perform the operation in the data processing unit (such as a central processing unit) separated from the data storage unit (such as a memory integrated circuit). i Write memory cell MC and receive weight W i With input voltage X i The page buffer PB, which contains the product of the bit lines BL and the page buffer PB, is coupled to the memory array 100 via a large number of highly parallel bit lines BL. Therefore, the page buffer PB has a relatively high bandwidth. Therefore, the integrated circuit 10 can be applied to large-scale data operations without encountering performance bottlenecks such as those encountered in the Van Neumann architecture. In some embodiments, the page buffer PB can have a bandwidth greater than or equal to 32 kB.

[0062] Please refer to Figure 1A and Figure 1B The calculation method of the integrated circuit 10 may include the following steps. In step S100, the weight W is set by performing the above-mentioned programmed operation multiple times. i Writing to a plurality of memory cells MC.

[0063] In step S102, multiple input voltages X i is applied to a page of memory cells MC connected to a word line WL (eg, word line WL1). i With input voltage X i Multiply in memory cell MC, and weight W i With input voltage X i The product of is outputted in the form of a current signal via the bit line BL. In addition, the page buffer PB is configured to sense the current signals outputted. In step S104, the current signals outputted are summed up by a component such as a counter CT. Steps S102 and S104 may constitute a single loop for executing the sum-of-products function for the memory cells MC of a single page. Subsequently, other loops are performed to execute the sum-of-products function for the memory cells MC of other pages. For example, other loops include a loop containing steps S1021 and S1041, a loop containing steps S1022 and S1042... and a loop containing steps S102 n Same as step S104 n In two consecutive cycles of executing the product term sum function for the memory cells MC of the adjacent pages, the input voltage X i One of the steps (e.g., step S1021) of applying the current to the memory cells MC of the adjacent page occurs after the other (e.g., step S102) and may at least partially overlap with the step of summing the current signal in the earlier cycle (e.g., step S104). Based on this pipeline timing flow design, some steps overlap in time, thereby further improving the instruction cycle time of the integrated circuit 10.

[0064] Figure 2 is a schematic diagram of an integrated circuit 10a according to some embodiments of the present invention. Figure 2 The integrated circuit 10a and its operation method described are similar to those of reference Figure 1A 、 Figure 1B The integrated circuit 10 and its operation method are described below. Only the differences between the two are described below, and the same or similar parts are not repeated.

[0065] Please refer to Figure 2In some embodiments, in the operation mode, the multiple bit lines BL of a block of the memory array 100 receive the same input voltage X. In other words, in these embodiments, the multiple input voltages X received by these bit lines BL do not have a specific distribution (pattern). For example, in the dual-bit mode, all the bit lines BL can be configured to receive the input voltage X of the low logic level "1". In this way, the multiple weights W stored in the multiple memory cells MC are i The same input voltage X is multiplied, and the resulting multiple products are converted into logic signals (e.g., 1 and 0) in the form of current signals via an amplifying sensor (not shown) and input to the page buffer PB'. In some embodiments, the page buffer PB' includes a first cache CA1 and a second cache CA2. The first cache CA1 is configured to receive and temporarily store the aforementioned logic signal (hereinafter referred to as the first logic signal) and is pre-written with the multiple input voltages X. i The other logic signals converted from these signals (hereinafter referred to as second logic signals) are: i In other words, multiple input voltages X i For example, in the dual-bit mode, multiple input voltages X i One of the input voltages X can be converted into a high logic level signal "1", and the multiple input voltages X i The other of the two signals may be converted to a low logic level signal "0." Subsequently, a counter (not shown) within the second cache CA2 is configured to perform a multiplication-accumulation operation on the first logic signal and the second logic signal. In other words, the second cache CA2 is configured to multiply the first logic signal and the second logic signal and sum the resulting products. Thus, the sum-of-products function has been executed through multiplication and addition operations, and these multiplication and addition operations can be considered near-memory operations.

[0066] Figure 3 is a schematic diagram of an integrated circuit 10b according to some embodiments of the present invention. Figure 3 The integrated circuit 10b and its operation method described are similar to those of reference Figure 1A 、 Figure 1B The integrated circuit 10 and its operation method are described below. Only the differences between the two are described below, and the same or similar parts are not repeated.

[0067] Please refer to Figure 3 In some embodiments, a block of the memory array 100' of the integrated circuit 10b is divided into a plurality of tiles. For example, Figure 3As shown, a block of the memory array 100' is divided into four sub-blocks TL. Each of the sub-blocks TL includes a portion of the memory array 100', and the sub-blocks TL are physically separated from each other. Figure 3 Only the bit lines BL and word lines WL of each sub-block TL are shown, and other components of each sub-block TL (such as the Figure 1A Memory cells MC, string select transistors SST, ground select transistors GST, string select lines SSL, and ground select lines GSL are shown. Multiple sub-blocks TL are arranged along a plurality of columns and rows. In some embodiments, an inter-sub-block bus system BS is coupled to the multiple sub-blocks TL and extends between the multiple sub-blocks TL. Furthermore, the inter-sub-block bus system BS may be further coupled to a sequencing controller (not shown). Furthermore, each sub-block TL is coupled to peripheral circuits including a page buffer PB and a counter CT. In some embodiments, peripheral circuits coupled to adjacent sub-blocks TL in the same column face each other, and peripheral circuits coupled to adjacent sub-blocks TL in the same row are located on the same side of the sub-blocks TL. However, those skilled in the art may adjust the number of sub-blocks TL and the configuration of the sub-blocks TL and peripheral circuits based on design requirements, and the present invention is not limited thereto. Furthermore, in some embodiments, each sub-block TL is coupled to a row decoder and a column decoder (neither shown). By dividing the memory array 100 ′ into a plurality of sub-blocks TL, the resistance-capacitance delay (RC delay) effect of the integrated circuit 10 b can be reduced, and the instruction cycle of the integrated circuit 10 b can be further improved.

[0068] Figure 4 is a schematic diagram of an integrated circuit 20 according to some embodiments of the present invention. Figure 4 The integrated circuit 20 and its operation method described are similar to those of reference Figure 1A 、 Figure 1B The integrated circuit 10 and its operation method are described below. Only the differences between the two are described below, and the same or similar parts are not repeated.

[0069] Figure 4 The diagram shows a plurality of blocks of the memory array 200 of the integrated circuit 20, for example, including block BK1 and block BK2. Each block of the memory array 200 is similar to Figure 1AThe illustrated block of memory array 100 has a plurality of columns (or strings) and a plurality of rows (or pages) of memory cells MC. One of a plurality of word lines WL connects each column of memory cells MC, while each column (or string) of memory cells MC is connected between a bit line BL and a source line SL. In some embodiments, within the same block, a plurality of columns (or strings) of memory cells MC share the same source line SL. Furthermore, the word lines WL of one block (e.g., block BK1) and the word lines WL of another block (e.g., block BK2) are disconnected (or electrically isolated), while the bit lines BL of different blocks (e.g., blocks BK1 and BK2) are connected to each other. In other words, the multiple blocks each have independent word lines WL and a shared bit line BL. In some embodiments, the source lines SL of different blocks may be coupled to each other. In an alternative embodiment, the source lines SL of one block (eg, block BK1 ) and the source lines SL of another block (eg, block BK2 ) are not connected to each other (or are electrically isolated).

[0070] When the integrated circuit 20 operates in the memory mode, by referring to Figure 1A The multiple programmed operations described above are used to convert multiple weights W i On the other hand, when the integrated circuit 20 operates in the operation mode, a plurality of word lines WL of different blocks and a bit line BL shared by different blocks are selected, and the selected word line WL receives the input voltage X i In some embodiments, these input voltages X i has a specific distribution (pattern), and these input voltages X i For example, in the dual-bit mode, multiple input voltages X i One of them is a high logic level "1", and multiple input voltages X i The other one is a low logic level "0". In addition, the selected bit line BL receives a read voltage, while the other bit lines BL receive a pass voltage (for example, 0V). The weight W stored in the memory cell MC corresponding to the selected word line WL and bit line i In these memory cells MC, the input voltage X i Multiply. The input voltage X is multiplied by the word line WL. i In the embodiment of transferring to the memory cell MC, the weight W stored in the memory cell MC i It can be regarded as the transconductance of the memory cell MC. i Corresponding to input voltage X iThe multiple products of W are outputted as current signals via the selected bit line BL. Since each bit line BL is shared by different blocks of the memory array 200, the output current signals from different blocks are accumulated at the selected bit line BL. In some embodiments, the multiple weights W are sensed by a page buffer PB coupled to the memory array 200. i Corresponding to input voltage X i The sum of multiple products of .

[0071] Based on Figure 4 In the configuration shown, the multiplication operation is performed in the memory cell MC, while the addition operation is performed via the bit line BL shared by different blocks. Therefore, both the multiplication and addition operations can be regarded as in-memory operations.

[0072] In reference Figure 4 In the illustrated embodiment, over-erasing of the memory cells MC is avoided before programming them. Specifically, in embodiments where the memory cells MC are N-type transistors, the threshold voltage of the erased memory cells is greater than or equal to 0V. Consequently, in the operating mode, the memory cells MC corresponding to unselected word lines WL receive a pass voltage, such as 0V, and are completely turned off. Consequently, the output current signal is contributed solely by the memory cells MC corresponding to the selected word line WL and bit line BL, thereby improving the reliability of the integrated circuit 20.

[0073] In summary, the integrated circuit of the present invention can operate in both memory mode and computation mode. The integrated circuit includes a memory array, such as a NAND flash memory array. The integrated circuit can execute a sum-of-products function and can be used in learning programs for artificial intelligence applications, neuromorphic computing systems, and machine learning systems. In memory mode, weights are written to memory cells in the memory array. In computation mode, the stored weights are multiplied by an input voltage transmitted to the memory cells via a bit line or word line, and the product of the weights and the input voltage is accumulated. Compared to the van Neumann architecture, which performs computations in a data processing unit (such as a central processing unit) separate from the data storage unit (such as a memory integrated circuit), the integrated circuit of the present invention can operate in both memory mode and computation mode. Therefore, data no longer needs to be transferred back and forth between the data processing unit and the data storage unit, and instruction cycle time can be significantly improved. In particular, the page buffer used to write weights to the memory cells and receive the product of the weights and the input voltage is coupled to the memory array via a large number of highly parallel bit lines, resulting in a very high bandwidth. Therefore, integrated circuits can be used for computing massive amounts of data and may not encounter performance bottlenecks like those in the Van Neumann architecture.

[0074] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An integrated circuit comprising: a memory array comprising a plurality of memory cells, each configured to be written with a weight; a plurality of word lines and a plurality of bit lines, wherein the plurality of word lines are respectively connected to a column of memory cells among the plurality of memory cells, and the plurality of bit lines are respectively connected to a column of memory cells connected in series, wherein a plurality of the plurality of bit lines in a block of the memory array or a plurality of the plurality of word lines in a plurality of blocks of the memory array are configured to receive a plurality of input voltages, and the plurality of the plurality of memory cells receiving the plurality of input voltages are configured to multiply a plurality of the plurality of written weights by the received plurality of input voltages; as well as a page buffer coupled to the memory array and configured to sense a plurality of products of the plurality of the plurality of weights and the plurality of input voltages; The plurality of input voltages are identical to one another; the page buffer includes a first cache and a second cache, the first cache being configured to receive a plurality of first logic signals converted by multiplying the plurality of weights by the plurality of input voltages and being pre-written with a plurality of second logic signals converted by a plurality of additional input voltages; and the second cache being configured to multiply the plurality of first logic signals by the plurality of second logic signals and accumulate the plurality of products of the plurality of first logic signals and the plurality of second logic signals.

2. The integrated circuit of claim 1 , wherein the plurality of the plurality of bit lines in the block receive the plurality of input voltages, and one of the plurality of word lines in the block is configured to receive a read voltage, while the others of the plurality of word lines in the block are configured to receive a pass voltage.

3. The integrated circuit of claim 2 , wherein the memory cells corresponding to the plurality of the plurality of bit lines and the one of the plurality of word lines are configured to multiply the plurality of the stored plurality of weights with the received plurality of input voltages and generate the plurality of products. 4 . The integrated circuit of claim 3 , further comprising a counter, wherein the counter is coupled to the page buffer and configured to sum the plurality of products.

5. The integrated circuit of claim 1, wherein at least two of the plurality of additional input voltages are different from one another and are converted to different logic signals.

6. The integrated circuit of claim 1 , wherein the plurality of word lines in the plurality of blocks are configured to receive the plurality of input voltages, the word lines of one of the plurality of blocks are electrically isolated from the word lines of another of the plurality of blocks, the plurality of bit lines are respectively shared by the plurality of blocks of the memory array, and one of the plurality of bit lines is configured to receive a read voltage, while the other of the plurality of bit lines are configured to receive a pass voltage.

7. The integrated circuit of claim 6, wherein the memory cells corresponding to the plurality of the plurality of word lines and the one of the plurality of bit lines are configured to multiply the plurality of the stored weights by the received plurality of input voltages and generate the plurality of products.

8. The integrated circuit of claim 7, wherein the plurality of products are summed across the one of the plurality of bit lines.

9. The integrated circuit of claim 7, wherein memory cells corresponding to the plurality of the plurality of word lines and the one of the plurality of bit lines have a starting voltage greater than or equal to 0V.

10. The integrated circuit of claim 1, wherein the memory array is a NAND flash memory array and the plurality of memory cells are a plurality of flash memory cells.

11. The integrated circuit according to claim 1, wherein the number of the page buffers is plural, and a block of the memory array has a plurality of sub-blocks, each of the sub-blocks being coupled to one of the plurality of page buffers.

Citation Information

Patent Citations

  • Semiconductor memory device

    CN107346666A

  • Device for neuromorphic computer system and manufacturing method thereof

    CN110163351A

  • Non-volatile (NV) memory (NVM) matrix circuits employing NVM matrix circuits for performing matrix computations

    US20190019538A1