Design for efficient near-memory computing and in-memory digital computing.
The in-memory computing circuit with tristate multipliers and shared accumulator addresses space and performance issues in dCIM systems by enabling concurrent MAC operations and faster weight downloads, reducing layout area and operation time.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- MACRONIX INTERNATIONAL CO LTD
- Filing Date
- 2024-10-25
- Publication Date
- 2026-07-23
AI Technical Summary
Conventional digital computing-in-memory (dCIM) systems face issues with space-intensive adder trees and slow performance due to the need to pause multiplication operations for weight downloads, which occupy a large layout area and require undesirably long time to update memory content.
Implement an in-memory computing circuit with tristate multipliers and a shared accumulator circuit, allowing subgroups to perform MAC operations concurrently and reducing the number of adder trees, enabling faster weight downloads and pipelined adder trees to reduce overall operation time.
The solution reduces layout area and improves performance by allowing simultaneous multiplication operations across subgroups, eliminating the need to pause MAC operations during weight downloads, and reduces the overall clock cycles required for complete MAC operations.
Smart Images

Figure 0007894418000002 
Figure 0007894418000003 
Figure 0007894418000004
Abstract
Description
Technical Field
[0001] Field of the Invention The present invention relates to a circuit that can be used to perform in-memory or near-memory calculations such as multiply-and-accumulate (MAC) operations or other operations for summing products.
Background Art
[0002] Description of Related Art In circuits used for certain calculations based on neuromorphic (neuro-morphological) computing systems, machine learning systems, and linear algebra, the multiply-and-accumulate function or the sum-of-products function can be an important component. Such functions can be expressed as follows:
Number
[0003] In this equation, each product term is the product of the variable input X , ,
[0005] and the weight W i The weight W i can vary between terms and, for example, corresponds to the coefficient of the variable input X i .
[0004] The sum-of-products function can be realized as the operation of a circuit using a cross-point array architecture, in which the electrical characteristics of the cells of the array give rise to this function.
[0005] These architectures can be implemented in digital computing-in-memory (dCIM) systems and digital near-memory-computing (dNMC) systems to perform the multiply-accumulate (MAC) operations described in the above formulas. Traditionally, in these systems, a subgroup of a product is implemented by a corresponding adder tree (e.g., an accumulator). As a result, there are many adder trees because each subgroup (of a larger group) has its own corresponding adder tree. Adder trees occupy a relatively large layout area and are therefore space-intensive. Traditionally, in these systems, adder trees occupy an undesirable amount of space. For simplicity, dCIM systems are also included in dNMC systems below.
[0006] In addition, the dCIM system requires an undesirably long time to download memory content (e.g., weights) as a result of the need to toggle some or all word lines (WLs), which leads to slower operation. Specifically, the dCIM system has to pause multiplication operations for several subgroups while downloading content (e.g., weights), which degrades performance. [Overview of the Initiative] [Problems that the invention aims to solve]
[0007] Therefore, it is desirable to provide a dCIM system that has an adder tree with a reduced number of elements and can perform MAC operations, etc., while downloading new content (e.g., weights). [Means for solving the problem]
[0008] In one preferred example, an in-memory computing circuit is provided. This in-memory computing circuit may include one or more input lines, an array of memory cells, a multiplier circuit connected to the array of memory cells and one or more input lines, and an accumulator circuit, where one or more input lines receive M input data elements, M is an integer greater than 0, the array of memory cells includes one or more subgroups, each subgroup in one or more subgroups stores M input data elements, the multiplier circuit is configured to multiply the M input data elements by M stored data elements in a selected subgroup from one or more subgroups, and to provide a multiplier output having M data elements, the accumulator circuit includes an accumulator input of M data elements connected to the multiplier output and is configured to produce the sum of the M data elements of the multiplier output, and the multiplier circuit provides the multiplication results from one or more subgroups to the multiplier output (sequentially).
[0009] In a further preferred example, the multiplication circuit may include M tristate multipliers connected to the multiplier output for each subgroup in one or more subgroups.
[0010] In another preferred example, M tristate multipliers can be replaced with M tristate NOR gates.
[0011] In one preferred example, one or more subgroups may include a first subgroup that stores M storage data elements and a second subgroup that stores M storage data elements, wherein M tristate multipliers for the first subgroup are enabled by a first timing signal to multiply M input data elements by the M storage data elements of the first subgroup, and M tristate multipliers for the second subgroup are enabled by a second timing signal to multiply M input data elements by the M storage data elements of the second subgroup, and the second timing signal is supplied at a different time than the first timing signal so that the M tristate multipliers for the second subgroup are enabled at a different time than the M tristate multipliers for the first subgroup.
[0012] In a further preferred example, one or more subgroups may include a first subgroup that stores M memory data elements in M memory circuits, and a second subgroup that stores M memory data elements in M memory circuits, wherein the first subgroup is connected to a first word line and the second subgroup is connected to a second word line.
[0013] In another preferred example, a specific memory circuit among the M memory circuits of a first subgroup and a specific memory circuit among the M memory circuits of a second subgroup can share a common line to control the storage of their respective data elements, the specific memory circuit of the first subgroup stores a specific data element in response to a first word line activating the first subgroup, and the specific memory circuit of the second subgroup stores a specific data element in response to a second word line activating the second subgroup.
[0014] In one preferred example, a common line shared by a specific memory circuit in the first subgroup and a specific memory circuit in the second subgroup may include a bit line (BL).
[0015] In a further preferred example, a specific data element may be written to a specific memory circuit among the M memory circuits of the second subgroup at least one of the following: (i) while M tristate multipliers for the first subgroup are enabled by a first timing signal to multiply M input data elements by M storage data elements of the first subgroup to provide a multiplier output having M data elements, and (ii) while an accumulator circuit receives and accumulates the multiplier output having M data elements.
[0016] In other preferred examples, the multiplier outputs may include a first output line and a second output line, the first output line being shared by the output of one tristate multiplier for a first subgroup and the output of one tristate multiplier for a second subgroup, the second output line being shared by the output of another tristate multiplier for a first subgroup and the output of another tristate multiplier for a second subgroup, and depending on whether the M tristate multipliers for the first subgroup are enabled by a timing control signal and the M tristate multipliers for the second subgroup are not enabled by a timing control signal, outputs related to the first subgroup are provided to the accumulating circuit via the first and second output lines, and depending on whether the M tristate multipliers for the second subgroup are enabled by a timing control signal and the M tristate multipliers for the first subgroup are not enabled by a timing control signal, outputs related to the second subgroup are provided to the accumulating circuit via the first and second output lines.
[0017] In one preferred example, one or more subgroups may include a first subgroup storing M storage data elements, a second subgroup storing M storage data elements, a third subgroup storing M storage data elements, and a fourth subgroup storing M storage data elements, wherein the first and second subgroups are connected to a first word line, and the third and fourth subgroups are connected to a second word line, and M tristate multipliers for the first subgroup are enabled by a first timing signal to multiply M input data elements by the M storage data elements of the first subgroup, and M tristate multipliers for the second subgroup are enabled by a first timing signal to multiply M input data elements by the M storage data elements of the first subgroup, and M tristate multipliers for the second subgroup are enabled by a first timing signal Two tristate multipliers are enabled by a timing signal to multiply M input data elements by M storage data elements of the second subgroup, M tristate multipliers for the third subgroup are enabled by a third timing signal to multiply M input data elements by M storage data elements of the third subgroup, and M tristate multipliers for the fourth subgroup are enabled by a fourth timing signal to multiply M input data elements by M storage data elements of the fourth subgroup, where L is an integer that can represent the total number of storage data elements of the first subgroup and the second subgroup, where M = L / 2.
[0018] In another preferred example, the multiplier output may include a first output line, which is shared by the output of one tristate multiplier for the first subgroup, the output of one tristate multiplier for the second subgroup, the output of one tristate multiplier for the third subgroup, and the output of one tristate multiplier for the fourth subgroup.
[0019] In one preferred example, the multiplier output can include a second output line, and the second output line is shared by the output of another one tri-state multiplier for the first subgroup, the output of another one tri-state multiplier for the second subgroup, the output of another one tri-state multiplier for the third subgroup, and the output of another one tri-state multiplier for the fourth subgroup.
[0020] In a further preferred example, the M memory data elements of each subgroup in one or more subgroups can be written to each subgroup using bit lines.
[0021] In another preferred example, the M memory data elements of each subgroup in one or more subgroups can be written to each subgroup using a sense amplifier connected to the bit lines.
[0022] In one preferred example, one or more subgroups can include a first subgroup storing M memory data elements and a second subgroup storing M memory data elements. During the first clock cycle, the first subgroup multiplies M input data elements by M memory data elements, and M memory data elements are written to the second subgroup. During the second clock cycle, the second subgroup multiplies M input data elements by M memory data elements, and M memory data elements are written to the first subgroup.
[0023] In a further preferred example, one or more subgroups can include a first subgroup storing M memory data elements and a second subgroup storing M memory data elements. During a specific clock cycle, an accumulation circuit accumulates the output related to the first subgroup, and during a subsequent clock cycle, the accumulation circuit accumulates the output related to the second subgroup.
[0024] In another preferred example, the accumulation circuit can be pipelined.
[0025] In one preferred example, the multiplication circuit can include M pass gates connected to a shared M-bit multiplier connected to the multiplier output for each of one or more subgroups.
[0026] In another preferred example, the multiplication circuit can be enabled by a timing control signal to provide a multiplication result. In one preferred example, the timing control signal can include a first timing signal and a second timing signal, and the first timing signal is supplied at a time different from the second timing signal.
[0027] In a further preferred example, a method of performing an operation is provided. This method can be executed using an in-memory computing circuit, which includes (i) an array of memory cells, (ii) a multiplication circuit connected to the array of memory cells and one or more input lines, and (iii) an accumulation circuit. The array of memory cells includes one or more subgroups, each subgroup stores M stored data elements, M is an integer greater than 0, and the accumulation circuit includes accumulator inputs for M data elements connected to the multiplier output. Further, this method can include the steps of obtaining M input data elements from one or more input lines, multiplying, by the multiplication circuit, the M input data elements by the M stored data elements of a selected subgroup of one or more subgroups to provide a multiplier output having M data elements, wherein the multiplication circuit is enabled by a timing control signal to sequentially provide multiplication results from subgroups within one or more subgroups to the multiplier output, and generating, by the accumulation circuit, a sum of the M data elements of the multiplier output.
[0028] In other preferred examples, an in-memory calculation circuit is provided. This in-memory calculation circuit may include a first subgroup of the circuit, a second subgroup of the circuit, a multiplier circuit, and an accumulator circuit, wherein the first subgroup of the circuit is connected to a first word line and configured to store a first set of weights, the second subgroup of the circuit is connected to a second word line and configured to store a second set of weights, and the multiplier circuit (i) multiplies the input to the first set of weights in response to a first timing signal, (ii) provides a first output, and (iii) multiplies the input to the second set of weights in response to a second timing signal, (i v) A second output is provided, and the multiplication of the second set of weights is enabled at a different time than the multiplication of the first set of weights is enabled, and the accumulating circuit is shared by the first and second subgroups and is configured to (i) receive and accumulate the first output in response to the multiplication of the first set of weights being enabled by the first timing signal, and (ii) receive and accumulate the second output in response to the multiplication of the second set of weights being enabled by the second timing signal.
[0029] In one preferred example, the common line described above may include a bit line bar line (BLB: a negative logic bit line).
[0030] In a further preferred example, the common line may further include a reference voltage line (VREF).
[0031] In another preferred example, the array of memory cells may include latches.
[0032] Other aspects and advantages of the present invention can be understood by considering the drawings, detailed description, and claims that follow. [Brief explanation of the drawing]
[0033] [Figure 1]This diagram illustrates a conventional static random access memory (SRAM)-based dCIM containing multiple subgroups of multiplier units, each subgroup having its own corresponding adder tree (accumulator). [Figure 2] Figure 1 shows an example of a subgroup of a conventional SRAM-based dCIM system. [Figure 3] This figure shows a dCIM system that includes a tristate NOR gate, enabling the sharing of an adder tree among various subgroups performing MAC operations. [Figure 4] Figure 3 shows enlarged views of parts of each of the three different subgroups of the dCIM system. [Figure 5] Figure 3 shows the dCIM system, where content (data) is written to subgroup 0, and one or more other subgroups process the input and calculate the output. [Figure 6] This figure shows a pipeline adder tree used in one embodiment of the dCIM system. [Figure 7] This diagram illustrates a dCIM system, in which two tristate NOR gates from two different subgroups share one output line and one word line. [Figure 8] Figure 7 shows the adder tree of the dCIM system. [Figure 9] This diagram illustrates a dCIM system, in which four NOR gates from four different subgroups share one output line and one word line. [Figure 10] This figure shows the calculation time series for subgroups 0 (903) to 7 (910) in Figure 9. [Figure 11] This diagram illustrates the dCIM system, which uses bit lines and a reference voltage (Vref) to write content. [Figure 12] This diagram illustrates the dCIM system, which uses a sense amplifier to write content. [Figure 13] This figure shows the timing chart for executing MAC calculations using the dCIM system. [Figure 14] This figure shows a dCIM system that uses a pass gate instead of the tristate NOR gate shown in Figure 3. [Figure 15] This is a simplified block diagram of an integrated circuit device that includes a memory array configured for in-memory operations using signed (or unsigned) inputs and weights. [Modes for carrying out the invention]
[0034] Detailed explanation A detailed description of embodiments of the present invention is provided with reference to Figures 1 to 15.
[0035] Figure 1 shows a conventional SRAM (static random access memory) based dCIM system, which includes multiple subgroups of multiplication units, each subgroup having its own corresponding adder tree (accumulator). The dCIM system in Figure 1 is considered SRAM-based because the weights are stored in and read from the SRAM.
[0036] Specifically, Figure 1 shows a conventional SRAM-based dCIM system 100, which includes multiple subgroups within a single group, such as subgroup 102, with each subgroup utilizing a corresponding adder tree. For example, subgroup 102 receives input and word line signals from an input activation driver and an SRAM WL driver 106, performs mathematical operations, and uses the adder tree 104 to perform cumulative operations. Each subgroup includes a memory circuit for storing weights and a circuit for performing mathematical operations such as multiplying the weights by the input.
[0037] In this example, weights are stored using a memory device such as a 6-transistor (6T) SRAM cell, and the weights are multiplied by an input (IN<0:255>) using a multiplier such as a 4-transistor (4T) NOR gate. Furthermore, as shown in the figure, the input activation driver and the SRAM WL driver receive inputs IN<0:255> on the line and word line signals on the word line WL<0:255>. Each of the inputs IN<0:255> is received by subgroup 102 (and other subgroups), thereby allowing the stored weights to be multiplied by these inputs. The word line WL<0:255> can be used to access the memory device for reading and writing (e.g., to write weights to the memory device and / or to read weights from the memory device). The output of the multiplication operation of subgroup 102 is provided to the adder tree 104 as input 4b. The adder tree 104 can combine various inputs in operation 5b to provide a single output. In this example, a single subgroup with 4 columns and 256 rows of cells can implement Ini×Wi<0:3>[i=0~255] for a combination of a 1024-bit (4 bits per row) output and a 1024-bit input adder tree to complete the MAC operation. This multiplication operation can be performed by a NOR gate or any other type of circuit capable of performing multiplication (or other types of mathematical) operations.
[0038] As illustrated, this conventional SRAM-based dCIM system 100 requires a separate adder tree for each subgroup. In this example, there are 64 subgroups, and therefore 64 adder trees, and these adder trees occupy a larger amount of physical space within the conventional SRAM-based dCIM system 100. Specifically, one of the problems with the conventional SRAM-based dCIM system 100 is that adder trees cannot be shared with other subgroups. Consequently, as the number of adder trees increases based on the number of subgroups, the overall layout size of the SRAM-based dCIM 100 also increases.
[0039] Figure 2 shows an example of a subgroup of the conventional SRAM-based dCIM system shown in Figure 1.
[0040] Specifically, Figure 2 shows a subgroup 200 that receives word line input 204 (WL<0:255>), input 205 (IN_B<0:255>), bit line input 206 (BL<0:3>), and bit line bar input 208 (BLB<0:3>), where bit line input 206 (BL<0:3>) and bit line bar input 208 (BLB<0:3>) are used to program the memory device 202 where the weights are stored. Furthermore, subgroup 200 includes a NOR gate for multiplying the stored weights by input 205. The result of the multiplication is provided to the adder tree 210. As described above with reference to Figure 1, each subgroup 200 communicates with the corresponding adder tree 210. Although Figure 2 simply shows a single subgroup and one adder tree, as shown in Figure 1, an SRAM-based dCIM requires at least 64 adder trees.
[0041] To update the subgroup's memory device 202 and store a new set of weights (e.g., weight values), all word lines WL<0:255> are sequentially enabled to update the entire contents of the SRAM. Updating memory device 202 with a new set of weights is expensive due to the time required to store the new values. Furthermore, since updating memory device 202 affects the data received by the adder tree 210, the entire MAC operation for the entire subgroup must be stopped while memory device 202 is being updated. This further slows down the performance of the SRAM-based dCIM.
[0042] The disclosed technology addresses these shortcomings by providing an SRAM-based dCIM with improved performance and reduced layout area.
[0043] Specifically, compared to the systems in Figures 1 and 2, the disclosed technology uses multipliers with tristate outputs, for example, by replacing general NOR gates with tristate NOR gates, and each tristate NOR gate in each subgroup can be controlled by a single separate timing signal (e.g., a clock signal, etc.). Alternatively, the tristate NOR gates can be replaced by any kind of logic gate that can perform a multiplication operation between the input and the stored data (latched data bits) and / or can be controlled / enabled by a timing signal. The tristate output (of the tristate NOR gate) has a high impedance state when not in the active state. This makes it possible to connect multiple tristate NOR gate multipliers to each input of an accumulator, where only the tristate NOR gate multipliers selected and activated by a separate timing signal operate on the input. This structure allows multiple subgroups of an adder tree to be shared, resulting in a reduced layout area.
[0044] Furthermore, the disclosed technology allows each subgroup to be aligned along a WL direction orthogonal to the BL and BLB directions, enabling faster content downloads into each subgroup by enabling a single WL, as a single WL can activate the entire subgroup for downloading. The physical orientation of the WL direction and the BL and BLB directions can be changed, allowing the WL direction and the BL and BLB directions to have non-orthogonal orientations. The result is improved performance compared to the system in Figure 1, as it eliminates the need to stop MAC operations when downloading new weights. Moreover, the structure of the disclosed technology allows the output of one subgroup to be processed by the operation of a single adder tree. The disclosed technology provides an architecture in which other subgroups can continue processing (e.g., performing multiplication operations) while one subgroup is downloading content (e.g., weights), which further improves performance. The disclosed technology can further be implemented in a structure that uses a single WL to activate / enable multiple subgroups of a single group, thereby allowing a single group to contain two, four, eight, or even more subgroups using multiple timing signals (dividing one group into such subgroups). These timing signals can be based on any clock cycle or various clock signals, and do not need to be based on only one clock cycle. This structure enables multiplication operations of 1 / 2, 1 / 4, 1 / 8, or fewer per clock cycle, which reduces the number of adder tree inputs, and further reduces the layout size of the adder tree.
[0045] Furthermore, an adder tree may require, for example, seven accumulating (adding) layers to go from, for example, 128 inputs to a single output (e.g., one layer receives 128 inputs, the next receives 64, the next receives 32, the next receives 16, the next receives 8, the next receives 4, and the next receives 2 to provide the final single output). The time required to fully complete this number of accumulating layers may be longer than the time required for a NOR gate to complete a multiplication operation. For this reason, an adder tree can become a bottleneck for MAC operations. Therefore, the disclosed technique can implement a pipelined adder tree separated into several stages, with buffers or latches between the stages. This allows the adder tree to store primary output data from the previous stage of the adder tree and act as a pipeline for subsequently received inputs. As a result, each stage can operate in one clock cycle, which coincides with the clock cycle required to complete a multiplication operation for one subgroup. This prevents delays that could occur if the adder tree waits for the cumulative operation to complete for all levels before receiving new inputs. As a result, the total clock cycles for the entire MAC operation can be reduced. In other words, the adder tree is divided into several stages, with buffers (latches) inserted between each stage, so that the operation of one stage does not affect the others, thereby allowing the adder tree to operate in a pipeline flow where new inputs are received every clock cycle. As described above, the technology disclosed herein can also be implemented in the form of a near-memory computing system.
[0046] The structure and operation that enable these features described above will be explained below with reference to Figures 3 to 15.
[0047] Figure 3 shows a dCIM system that includes tristate NOR gates, which enable the sharing of the adder tree among various subgroups performing MAC operations.
[0048] Specifically, Figure 3 shows a dCIM system containing multiple (N) subgroups of circuits, including subgroup 0 302, subgroup 1 304, and subgroup N 306 (where N is an integer greater than 0). An adder tree 308 (e.g., an accumulator circuit) is connected to the outputs of all subgroups (e.g., subgroup 0 302, subgroup 1 304, ..., subgroup N 306), so that all subgroups share the adder tree 308. Furthermore, subgroup 0 302 is connected to the word line WL. <0> 310 is connected, and subgroup 1 304 is word line WL <1> Connected to 312, subgroup N 306 is word line WL <n>It is connected to 314. Each of subgroup 0 302, subgroup 1 304, and subgroup N 306 (hereinafter referred to as subgroups 302, 304, and 306) is individually activated or enabled to store content by their respective word lines 310, 312, and 314. A group can be referred to as a collection of subgroups such as subgroups 302, 304, and 306 that share the same adder tree 308.
[0049] As illustrated, each of subgroups 302, 304, and 306 may include memory circuits and / or multiplication (multiplier) circuits. Furthermore, subgroups 302, 304, and 306 may be referred to as arrays of memory cells, so that each subgroup includes memory cells that store data elements, and the multiplication circuits are connected to the arrays of memory cells and one or more input lines, which provide M (or any other number) input data elements on one or more input lines, on one or more contents, and / or on one or more multipliers. For example, subgroup 302 includes memory circuits 316, 317, and 318 (e.g., arrays of memory cells). Each subgroup may include M (or any other number) stored data elements. The memory circuits 316, 317, and 318 can be any type of memory, including latches, sense amplifier (SA) latches, SRAM, DRAM (dynamic random access memory), other types of volatile memory, and even NVM (non-volatile memory). (SA latches can be used to detect and store data in a memory array, for example, the memory array 902 in Figure 9.) This applies to any of the memory circuits (arrays of memory cells) described herein with respect to Figures 3-15. For example, the memory circuits 316, 317, and 318 can be composed of 6 transistors (6T). For example, a 6T SRAM connected to a true and complementary bit line by activating the word line has an additional output connected to a multiplier, which allows the input bits to be multiplied by the data stored in the cell without activating the word line. Other types of memory circuits can also be implemented. Subgroups 304 and 306 include similar memory circuits.Bit line BL0 320 and bit line bar line BL0B 322 (for example, a common line) are connected to word line WL <0> Used in conjunction with the activation of 310, content can be written to the memory circuit 316 (for example, content can be downloaded to the memory circuit 316 for programming the dCIM system 300), and the word line WL <0> 310 can be connected to each of the memory circuits of subgroup 302. As shown in the figure, BL0 320 and BL0B 322 are word lines WL <0> 310 can be orthogonal to this. Other orientations of word lines, bit lines, and bit bar lines can be realized. In addition, BL0 320 and BL0B 322 are connected to the corresponding memory circuits of subgroups 304 and 306, thereby sharing BL0 320 and BL0B 322 with the memory circuits of subgroups 302, 304, and 306. As shown in the figure, there is essentially a sequence of memory circuits of subgroups 302, 304, and 306 connected by BL0 320 and BL0B 322.
[0050] BL1 342 and BL1B 326 are connected to memory circuit 317 to write content to memory circuit 317 and to the corresponding memory circuits of subgroups 304 and 306, thereby allowing BL1 324 and BL1B 326 (e.g., a common line) to be shared by the memory circuits of subgroups 302, 304, and 306. BLm328 and BLmB330 (e.g., a common line) are connected to memory circuit 318 to write content to memory circuit 318 and to the corresponding memory circuits of subgroups 304 and 306, thereby allowing BLm328 and BLmB330 to be shared by the memory circuits of subgroups 302, 304, and 306. Activation of various word lines 310, 312, and 314 controls which subgroup the content is written to.
[0051] Semantically, a multiplication circuit can be referred to as part of a subgroup, or as being for a subgroup, but in practice it is not part of that subgroup. For example, subgroup 302 may include multiplication circuits such as tristate NOR gates 332, 334, and 336 (also referred to as tristate multipliers). As shown in Figure 3 and subsequent figures, memory circuits 316, 317, and 318 may have read and write ports connected to bit lines and bit bar lines, and may have separate read ports connected to tristate NOR gates. Each subgroup contains m tristate NOR gates. Tristate NOR gate 322 is connected to memory circuit 316, so that tristate NOR gate 322 can retrieve weights (content, input / storage data elements) stored in memory circuit 316 and multiply the retrieved weights by inputs such as input 0 received on input line 0 338. Tristate NOR gate 332 outputs the multiplied value only when enabled by a timing signal (e.g., first timing signal time 0). A timing control signal can include several different timing signals, such as a first timing signal time 0, a second timing signal time 1, and an Nth timing signal time N, or other timing signals described herein, and / or can control the transmission of different timing signals. Furthermore, a timing control signal can be connected to a multiplier circuit, and these multiplier circuits can be enabled to sequentially provide multiplication results from one or more subgroups to the multiplier output. A timing control signal can be said to select a particular subgroup (for example, the first timing signal time 0 can be said to select subgroup 302, the second timing signal time 1 can be said to select subgroup 304, and the Nth timing signal time N can be said to select subgroup 306).
[0052] Input 0 can be received at the same time as or approximately at the same time as a timing signal (e.g., a first timing signal time 0) (on the input line from the input driver). As shown in the figure, input 0 can be received at input B of the tristate NOR gate 332, and the weights can be received at weight B of the tristate NOR gate 332. The output of the multiplication performed by the NOR gate 332 is provided on the output line "output 0 340" (e.g., the first output line) and received by the adder tree 308. The output line "output 0 340" is shared by the respective tristate NOR gates of subgroups 302, 304, and 306. As shown in the figure, the sequence of tristate NOR gates, including the tristate NOR gate 332, extending through subgroups 302, 304, and 306, shares the same output line "output 0 340". However, at time 0, i.e., the time when timing signal time 0 enables the tristate NOR gate 332, the only output provided to output line "output 0 340" is from the tristate NOR gate 332 because the other tristate NOR gates in the other subgroups 304 and 306 are not enabled.
[0053] Similarly, the tristate NOR gate 334 is connected to the memory circuit 317, so that the tristate NOR gate 334 can retrieve the weights (contents) stored in the memory circuit 317 and multiply the retrieved weights by an input such as input 1 received on the input line. The tristate NOR gate 334 can only output (or perform) the multiplication value after being enabled by a timing signal (e.g., first timing signal time 0). Input 1 can be received at the same time as or approximately at the same time as the timing signal (e.g., first timing signal time 0). As shown in the figure, input 1 can be received at input B of the tristate NOR gate 334, and the weights can be received at weight B of the tristate NOR gate 334. The output of the multiplication performed by the tristate NOR gate 334 is provided on the output line "output 1 342" (e.g., second output line) and received by the adder tree 308. Output line "Output 1 342" is shared by the respective tristate NOR gates of subgroups 302, 304, and 306, just as described above with respect to output line "Output 0 340".
[0054] Similarly, the tristate NOR gate 336 is connected to the memory circuit 318, so that the tristate NOR gate 336 can retrieve the weights (contents) stored in the memory circuit 318 and multiply the retrieved weights by an input such as input m received on the input line. The tristate NOR gate 336 can only output (or perform the multiplication) the multiplied value after being enabled by a timing signal (e.g., first timing signal time 0). Input m can be received at the same time as or approximately at the same time as the timing signal (e.g., first timing signal time 0). As shown in the figure, input m can be received at input B of the tristate NOR gate 336, and the weight can be received at weight B of the tristate NOR gate 336. The output of the multiplication performed by the tristate NOR gate 336 is provided on the output line "output m 344" and received by the adder tree 308. Just as described above regarding output line "output 0 340", output line "output m 344" is shared by the respective tristate NOR gates of subgroups 302, 304, and 306. There can be M (or any other number) output lines per subgroup.
[0055] The tristate NOR gate of subgroup 304 can be enabled by a timing signal (e.g., timing signal time 1) and can receive inputs 0, 1 through m simultaneously or approximately simultaneously. In a similar manner to subgroup 302, the circuit of subgroup 304 can multiply the weights by the inputs and provide the output to the adder tree 308. Furthermore, the tristate NOR gate of subgroup 306 can be enabled by a timing signal (e.g., the nth timing signal time N) and can receive inputs 0, 1 through m simultaneously or approximately simultaneously. In a similar manner to subgroups 302 and 304, the circuit of subgroup 306 can multiply the weights by the inputs and provide the output to the adder tree 308. As shown in the figure, subgroup 302 takes one clock cycle to provide the outputs "output 0 340" through "output m 344", and therefore, in the Nth subsequent clock cycle, subgroup 306 provides the outputs "output 0 340" through "output m 344". The multiplications performed by each of the subgroups 302, 304, and 306 are controlled by different timing signals. In this specification, a single timing signal may comprise multiple distinct timing signals. Other timing signal schemes can be used to enable and control the tristate NOR gates described herein.
[0056] The outputs of each subgroup 302, 304, and 306 are cumulatively provided as MAC output 346, so that the output of subgroup 302 is provided as MAC output 346, then the output of subgroup 304 is provided as MAC output 346, and finally the output of subgroup 306 is provided as MAC output 346.
[0057] Subgroup 302 can be referred to as the first subgroup of the circuit, and the first word line (e.g., word line WL) <0> It is connected to 310) and configured to (i) store the first set of weights, (ii) multiply the first set of weights by the input in accordance with the first timing signal (e.g., time 0), and (iii) provide the first output (e.g., output 0 340 to output m 344). Subgroup 304 can be referred to as the second group of the circuit and is the second word line (e.g., WL). <1> Connected to 312), it is configured to (i) store the second set of weights, (ii) multiply the input to the second set of weights in accordance with a second timing signal (e.g., time 1), and (iii) provide a second output (e.g., output 0 340 to output m 344), wherein the multiplication of the second set of weights is enabled at a different time than the multiplication of the first set of weights is enabled. The adder tree 308 can be called an accumulation circuit, shared by a first subgroup (e.g., subgroup 302) and a second subgroup (e.g., subgroup 304), and is configured to receive and accumulate (i) a first output in accordance with the activation of the multiplication of the first set of weights by a first timing signal (e.g., time 0), and (ii) a second output in accordance with the activation of the multiplication of the second set of weights by a second timing signal (e.g., time 1).
[0058] Furthermore, the memory circuits 316, 317, and 318 of subgroup 302 may be referred to as first memory circuits, and the tristate NOR gates 332, 334, and 336 may be referred to as first multiplier circuits, the first memory circuits being connected to a first word line and configured to store (and write to) a first set of weights, the first multiplier circuits being enabled by a first timing signal and multiplying the input by the first set of weights to provide a first output. In addition, the memory circuit of subgroup 304 may be referred to as the second memory circuit, and the tristate NOR gate of subgroup 304 may be referred to as the second multiplier circuit, the second memory circuit being connected to a second word line and configured to store (and write to) a second set of weights, the second multiplier circuit being enabled by a second timing signal to multiply the input by the second set of weights to provide a second output, the second multiplier circuit being enabled at a different time than the first multiplier circuit being enabled, so that an accumulator circuit (e.g., adder tree 308) is configured to receive and accumulate (i) a first output corresponding to the first multiplier circuit being enabled by the first timing signal, and (ii) a second output corresponding to the second multiplier circuit being enabled by the second timing signal. The accumulator circuit 308 may include an accumulator input of N data elements connected to the multiplier output (e.g., the output of a multiplier circuit such as a tristate NOR gate). The cumulative circuit 308 can also generate the sum of the N data elements of the multiplier output.
[0059] The dCIM system 300 in Figure 3 can receive data elements for storage in the array of memory cells from a non-volatile memory (NVM) array such as NOR flash memory, NAND flash memory, etc. This type of dCIM system 300 can be referred to as an NVM-based dCIM system. Furthermore, the dCIM system 300 in Figure 3 can receive data elements for storage in the array of memory cells from a volatile memory array such as dynamic random access memory (DRAM), SRAM, etc. These types of dCIM systems can be referred to as an SRAM-based dCIM system, a DRAM-based dCIM system, and / or a volatile memory-based dCIM system, etc. Any of the dCIM systems described herein with respect to Figures 3-15 can be either an NVM-based dCIM system or a volatile memory-based dCIM system as described above.
[0060] Figure 4 shows enlarged views of parts of each of the three different subgroups of the dCIM system shown in Figure 3.
[0061] Specifically, Figure 4 shows enlarged views of portions of subgroups 302, 304, and 306. As shown in the figure, the word line WL <0> 310 is connected to the memory circuit 316 of subgroup 302, and the word line WL <1> 312 is connected to the memory circuit 402 of subgroup 304, and the word line WL <n>314 is connected to the memory circuit 404 of subgroup 306. The tristate NOR gate 332 of subgroup 302 receives a first timing signal (time 0) and is enabled by the first timing signal. It receives the stored weights from the memory circuit 316, receives input 0 on input line "input 0 338" at or around the time the first timing signal (time 0) is received, multiplies the received weights by input 0 on input line "input 0 338", and provides the output on output line "output 0 340". As mentioned above, the memory circuit 316 and the tristate NOR gate 332 are merely examples, and different types of memory circuits for storage, different types of gates for performing mathematical operations, etc., can be implemented using different technologies.
[0062] Similarly, the tristate NOR gate 406 of subgroup 304 receives the second timing signal (time 1) and is enabled by the second timing signal, receives the stored weight from the memory circuit 402, receives input 0 on input line "input 0 338" at the time or approximately at the time the second timing signal (time 1) is received, multiplies the received weight by input 0, and provides the output on output line "output 0 340". Furthermore, the tristate NOR gate 408 of subgroup 306 receives the Nth timing signal (time N) and is enabled by the Nth timing signal, receives the stored weight from the memory circuit 404, receives input 0 on input line "input 0 338" at the time or approximately at the time the Nth timing signal (time N) is received, multiplies the received weight by input 0, and provides the output on output line "output 0 340".
[0063] Figure 5 shows the dCIM system from Figure 3, where content (data such as N memory data elements) is written to subgroup 0, and one or more other subgroups process the input and calculate the output. Throughout this document, the content written to the memory circuit may be referred to as weights, memory data elements, etc. The content can be written to various subgroups in the operation of downloading sets of weights.
[0064] Specifically, Figure 5 shows the same dCIM system 500 as the dCIM system 300 described with reference to Figure 3, and therefore redundant explanations are omitted. In addition to what was described with reference to Figure 3, Figure 5 shows that in operation 502, content is written to the memory circuit 316 of subgroup 302, in operation 504, content is written to the memory circuit 317 of subgroup 302, and in operation 506, content is written to the memory circuit 318 of subgroup 302. The content (data) to be written may include weights, such as the first set of weights. In this example, the memory circuits of subgroups 304 and 306 may already have weights stored (for example, the memory circuit of subgroup 304 may have the second set of weights stored, and the memory circuit of subgroup 306 may have the Nth set of weights stored). While content is being written to memory circuits 316, 317, and 318 of subgroup 302, multiplication circuits (e.g., tristate NOR gates) of subgroup 304 or subgroup 306 can perform the multiplication element of the MAC operation and provide the output to an adder tree (e.g., an accumulator).
[0065] Figure 6 shows a pipelined adder tree used in the embodiment of the dCIM system shown in Figure 3.
[0066] Specifically, Figure 6 shows an adder tree 600 (e.g., an accumulator circuit) that receives 128 outputs on the output line connected to the tristate NOR gate of the dCIM system in Figure 3. Outputs "Output 0 602", "Output 1 604" through "Output 126 606", and "Output 127 608" are received as input 601 of the adder tree 600. The adder tree 600 has seven layers, which include a first adder layer 614 that adds 128 inputs to provide 64 results, a second adder layer 616 that adds 64 results to provide 32 results, a third adder layer 618 that adds 32 results to provide 16 results, a fourth adder layer 620 that adds 16 results to provide 8 results, a fifth adder layer 622 that adds 8 results to provide 4 results, a sixth adder layer 624 that adds 4 results to provide 2 results, and a seventh adder layer 626 that adds 2 results to provide a single output to path 628. The single output provided on path 628 is mathematically equivalent to equation 630 and is the result of the entire MAC operation (for a particular subgroup such as subgroup 302 in Figure 2).
[0067] As illustrated, the adder tree 600 includes buffer latches 610 and 612 for pipeline stages, which temporarily store intermediate and final results of the cumulative operation. For example, buffer latch 610 stores 16 results provided by the third adder layer 618, and buffer latch 612 stores 2 results provided by the third adder layer 624. Furthermore, as illustrated, it takes one clock cycle to receive 128 inputs and store 16 results in buffer latch 610, one clock cycle to retrieve the 16 results stored in buffer latch 610, perform the cumulative operation and store 2 results in buffer latch 612, and one clock cycle to retrieve the 2 results stored in buffer latch 612, perform the cumulative operation and provide a single output to path 628. The adder tree 600 is structured such that, for example, while subgroup 304 is performing a multiplication operation, it can receive 128 outputs from subgroup 302 in Figure 3 at a given clock cycle (time) and store them in buffer latch 610. In a subsequent clock cycle (subsequent time), while the previous contents of buffer latch 610 (based on the operation of subgroup 302) are added and stored in buffer latch 612, 128 outputs can be received from subgroup 304 and stored in buffer latch 610. In yet another subsequent clock cycle (even later time), while the previous contents of buffer latch 610 (based on the operation of subgroup 304) are added and stored in buffer latch 612, and while the previous contents of buffer latch 612 (based on the operation of subgroup 302) are added and provided as a single output on path 628, 128 inputs can be received from subgroup 306 and stored in buffer latch 610.The adder tree 600 can be said to have three stages: the first stage (stage 0) includes receiving an input and storing it in buffer latch 610; the second stage (stage 1) includes processing the data from buffer latch 610 and storing it in buffer latch 612; and the third stage (stage 2) includes processing the data from buffer latch 612 and providing a single output on path 628.
[0068] Pipeline operation continues as long as the subgroup of the dCIM system continues to hold the written content and continues to perform multiplication operations. This pipeline operation allows MAC operations to continue without interruption because it eliminates the need for the dCIM system to wait for the adder tree 600 to complete the accumulation and provide the result. Therefore, the adder tree 600 can have an inherently faster clock and higher processing power. For example, as illustrated and described above, the adder tree 600 actually requires 3 clock cycles to take 128 inputs and provide a single output. Therefore, without buffer latches 610 and 612, a multiplication operation that would require only 1 clock cycle would have to wait for the adder tree to complete the accumulation operation. As a result, the dCIM system with this adder tree structure operates significantly faster.
[0069] Instead, the adder tree 600 (or any other adder tree described herein) can operate without a buffer latch. Furthermore, the buffer latch can be any kind of circuit or component that can store data. Moreover, the adder tree 600 (or any other adder tree described herein) can be a counter that performs counting (counting the number of "1"s).
[0070] Figure 7 shows a dCIM system in which two tristate NOR gates from two different subgroups share one output line and one word line.
[0071] Specifically, Figure 7 shows a dCIM system 700 similar to the one described with reference to Figure 3, and therefore a redundant explanation is omitted. In the dCIM system 700 in Figure 7, subgroup 0 702 and subgroup 1 703 share the same word line WL. <0> 310 is connected, and subgroups 2 704 and 3 705 share the same word line WL. <1> Connected to 312, subgroups N-1 706 and N 707 share the same word line WL. <n>It differs from the dCIM system in Figure 3 in that it is connected to 314. The dCIM system 700 in Figure 7 includes the same word lines 310, 312, and 314, the same memory circuits 316, 317, and 318, the same bit lines and bit bar lines 320, 322, 324, 326, 328, and 330 as described above with reference to Figure 3, and the same tristate NOR gates 322, 324, and 326.
[0072] The dCIM system 700 in Figure 7 also differs in that two tristate NOR gates from two different subgroups 702 and 703 share one output line "output 0 340" and one word line 310. For example, tristate NOR gate 332 from subgroup 702 and tristate NOR gate 334 from subgroup 703 both provide an output to output line "output 0 340". Similarly, two of the tristate NOR gates from two different subgroups 704 and 705 provide an output to the same output line "output 0 340", and two of the tristate NOR gates from two different subgroups 706 and 707 provide an output to the same output line "output 0 340".
[0073] In this example, when 128 tristate NOR gates are located in subgroups 702 and 703, 64 tristate NOR gates are members of subgroup 702, and 64 tristate NOR gates are members of subgroup 703. Furthermore, in this example, tristate NOR gate 332 is a member of subgroup 702, and tristate NOR gates 334 and 336 are members of subgroup 703. In addition, 64 tristate NOR gates are members of subgroup 704, and 64 tristate NOR gates are members of subgroup 705. Similarly, 64 tristate NOR gates are members of subgroup 706, and 64 tristate NOR gates are members of subgroup 707.
[0074] As shown in the diagram, at time 00 (e.g., the first clock cycle), the tristate NOR gates, members of subgroup 702, provide 64 outputs on the output lines "output 0 340" to "output m / 2 708". In this example, m (or L) = 128, meaning that there are 128 tristate NOR gates for the combination of subgroups 702 and 703, resulting in 64 (128 / 2) outputs. This subgroup architecture of the dCIM system 700 results in 64 outputs per clock cycle, in contrast to the 128 outputs per clock cycle described above with reference to the dCIM system 300 in Figure 3. At time 01 (e.g., the second clock cycle), the tristate NOR gates of subgroup 703 are enabled and multiply the 64 outputs. As a result, subgroups 702 and 703 are utilized for two clock cycles, which provides an additional clock cycle for writing content to other groups while subgroups 702 and 703 are performing multiplication, and also enables the use of smaller adder trees (accumulators), as described below with reference to Figure 8. Content can be written simultaneously to multiple subgroups connected to the same word line.
[0075] Returning to Figure 7, at time 10 (e.g., the third clock cycle), the tristate NOR gate, a member of subgroup 704, provides 64 outputs on the output lines "output 0 340" to "output m / 2 708". At time 11 (e.g., the fourth clock cycle), the tristate NOR gate of subgroup 705 is enabled and provides the 64 outputs multiplied together. At time N0 (e.g., the (N-1)th clock cycle), the tristate NOR gate, a member of subgroup 706, provides 64 outputs on the output lines "output 0 340" to "output m / 2 708". At time N1 (e.g., the Nth clock cycle), the tristate NOR gate of subgroup 707 is enabled and provides the 64 outputs multiplied together. This architecture shown in Figure 7 can be modified to include eight or more subgroups per group. For example, the number of "adjacent" NOR gates in subgroups 702 and 703 sharing the same output line can be 2, 4, 8, or any 2. N It can be divided into n parts, where N is an integer.
[0076] Figure 8 shows the adder tree of the dCIM system in Figure 7.
[0077] As described above, the dCIM system 700 in Figure 7 provides, for example, 64 outputs per clock cycle, in contrast to the 128 outputs per clock cycle of the dCIM system 300 in Figure 3. The adder tree 800 in Figure 8 has 64 inputs and receives the 64 outputs of the dCIM system 700 in Figure 7.
[0078] Specifically, the adder tree 800 receives outputs "Output 0 802", "Output 1 804" through "Output 62 806", and "Output 63 808". These outputs are received as 64 inputs 801 of the adder tree 800. This adder tree 800 has one less layer than the adder tree 600 in Figure 6. For example, adder tree 800 includes (i) a first adder tree layer 814 that receives 64 inputs and provides 32 results, (ii) a second adder tree layer 816 that receives 32 results and provides 16 results, (iii) a third adder tree layer 818 that receives 16 results and provides 8 results, (iv) a buffer latch 810 that receives and temporarily stores 8 results, (v) a fourth adder tree layer 820 that receives the 8 temporarily stored results and provides 4 results, (vi) a fifth adder tree layer 822 that receives 4 results and provides 2 results, and (vii) a sixth adder tree layer 824 that receives 2 results and provides a single result. Similar to adder tree 600 in Figure 6, adder tree 800 provides a single result on path 628, which is mathematically equivalent to equation 630 and is the result of MAC operations on a particular subgroup.
[0079] As illustrated, the adder tree 800 completes the cumulative operation in 2 clock cycles, requires 1 clock cycle to receive the input and store the result in the buffer latch 810 (stage 0), and requires 1 clock cycle to retrieve the stored data from the buffer latch 810 and provide a single output (stage 1). This is the same pipeline structure as described above for the adder tree 600, except that there is only one buffer latch in contrast to two, and that it requires only 2 clock cycles in contrast to 3 clock cycles to complete the cumulative operation. The reduction in the number of buffer latches and the number of clock cycles required is because the adder tree 800 receives only 64 inputs in contrast to the 128 inputs of the adder tree 600. As a result, the number of gate counts (i.e., gates that perform the cumulative operation) of the adder tree 800 is about half the number of gate counts of the adder tree 600, and thus the size of the adder tree 800 is about half the size of the adder tree 600. This adder tree 800, with fewer inputs, is a result of the subgroup structure and shared output lines described above, as shown in Figure 7.
[0080] Figure 9 shows a dCIM system in which four NOR gates from four different subgroups share one output line and one word line.
[0081] The dCIM system 900 in Figure 9 is the same as the dCIM system 700 in Figure 7, and redundant explanations of the components described with reference to Figure 7 are omitted.
[0082] The dCIM system 900 in Figure 9 differs from the dCIM system in Figure 7 in that there are four different subgroups of tristate transistors that share one output line and one word line. Specifically, in Figure 9, the four subgroups (i.e., subgroup 0 903, subgroup 1 904, subgroup 2 905, and subgroup 3 906, hereafter referred to as subgroups 903, 904, 905, and 906) share the same word line WL. <0> It indicates that it is connected to 310, and that four subgroups (i.e., subgroup 4 907, subgroup 5 908, subgroup 6 909, and subgroup 7 910; hereafter subgroups 907, 908, 909, and 910) are on the same word line WL <1> Further, it is shown that it is connected to 312. Additional sets of four subgroups up to N-3, N-2, N-1, and N can be connected to their respective word lines. The dCIM system 900 in Figure 9 includes the same word lines 310 and 312, the same memory circuit, and the same tristate NOR gate as described with respect to Figures 3 and 7.
[0083] Subgroup 903 includes a memory circuit and a tristate NOR gate associated with time 0, subgroup 904 includes a memory circuit and a tristate NOR gate associated with time 1, subgroup 905 includes a memory circuit and a tristate NOR gate associated with time 2, subgroup 906 includes a memory circuit and a tristate NOR gate associated with time 3, subgroup 907 includes a memory circuit and a tristate NOR gate associated with time 4, subgroup 908 includes a memory circuit and a tristate NOR gate associated with time 5, subgroup 909 includes a memory circuit and a tristate NOR gate associated with time 6, and subgroup 910 includes a memory circuit and a tristate NOR gate associated with time 7.
[0084] The bit line configuration in Figure 9 differs from that in Figure 7. Specifically, bit lines BL0 912 and BL1 914 and Vref916 (e.g., reference voltage line) are implemented to write content to specific memory circuits in subgroups 903, 904, 907, and 908 (e.g., memory circuits in the first column on the left). As shown in the illustration, Vref916 is the word line WL <0> Memory circuits above and below 310, and word line WL <1> While 312 is connected to the memory circuits above and below, bit line BL1 914 is connected to word line WL <0> Above 310 and the word line WL <1> 312 is connected to the memory circuit above, and bit line BL0 912 is connected to word line WL <0> Below 310 and the word line WL <1> It is connected to the memory circuit below 312. Furthermore, bit lines BL2 920 and BL3 922 and Vref918 are implemented to write content to specific memory circuits of subgroups 905, 906, 909, and 910 (for example, the memory circuits in the second column to the right of the first column). As shown in the figure, Vref918 is connected to word line WL <0> Memory circuits above and below 310, and word line WL <1> While 312 is connected to the memory circuits above and below, bit line BL3 922 is connected to word line WL <0> Above 310 and the word line WL <1> Connected to the memory circuit above 312, bit line BL2 920 is connected to word line WL <0> Below 310 and the word line WL <1> It is connected to the memory circuit below 312. Bit lines BL510 926 and BL511 928 and Vref924 are implemented to write content to specific memory circuits of subgroups 905, 906, 909, and 910 (e.g., memory circuits in the rightmost column). As shown in the diagram, Vref924 is the word line WL <0> Memory circuits above and below 310, and word line WL <1> While 312 is connected to the upper and lower memory circuits, bit line BL511 928 is connected to word line WL. <0> Above 310 and the word line WL <1> Connected to the memory circuit above 312, bit line BL510 926 is word line WL <0> Below 310 and the word line WL <1> It is connected to the memory circuit below 312.
[0085] The dCIM system 900 in Figure 9 differs from the dCIM system 700 in that the four tristate NOR gates of subgroups 903, 904, 905, and 906 and the four tristate NOR gates of subgroups 907, 908, 909, and 910 share a single output line "Output 0 936". For example, the tristate NOR gates associated with times 0, 1, 2, and 3 of subgroups 903, 904, 905, and 906 can all provide outputs to output line "Output 0 936". Similarly, the tristate NOR gates associated with times 4, 5, 6, and 7 of subgroups 907, 908, 909, and 910 provide outputs to the same output line "Output 0 936". Input lines 930, 932, and 934 provide inputs to the various tristate NOR gates.
[0086] In this example, with 512 tristate NOR gates in subgroups 903, 904, 905, and 906, 128 tristate NOR gates are members of subgroup 903, 128 tristate NOR gates are members of subgroup 904, 128 tristate NOR gates are members of subgroup 905, and 128 tristate NOR gates are members of subgroup 906. In addition, 128 tristate NOR gates are members of subgroup 907, 128 tristate NOR gates are members of subgroup 908, 128 tristate NOR gates are members of subgroup 909, and 128 tristate NOR gates are members of subgroup 910. Other configurations are possible depending on the desired number of outputs per clock cycle. In this example in Figure 9, there are 128 outputs. If there were 64 outputs, the number of tristate NOR gates per subgroup would be reduced to 64.
[0087] As shown in the diagram, at time 0 (e.g., the first clock cycle), the tristate NOR gate, a member of subgroup 903, outputs 128 outputs on the output lines "output 0 936" to "output 127 938". At time 1 (e.g., the second clock cycle), the tristate NOR gate of subgroup 904 is enabled and outputs the 128 outputs multiplied by each other. At time 2 (e.g., the third clock cycle), the tristate NOR gate of subgroup 905 is enabled and outputs the 128 outputs multiplied by each other. At time 3 (e.g., the fourth clock cycle), the tristate NOR gate of subgroup 906 is enabled and outputs the 128 outputs multiplied by each other. As a result, subgroups 903, 904, 905, and 906 are used for four clock cycles, which allows for additional clock cycles to write content to other subgroups connected to other word lines.
[0088] At time 4 (for example, the 5th clock cycle), a tristate NOR gate, a member of subgroup 907, provides 128 outputs on output lines "output 0 936" to "output 127 938". At time 5 (for example, the 6th clock cycle), a tristate NOR gate, a member of subgroup 908, provides 128 outputs on output lines "output 0 936" to "output 127 938". At time 6 (for example, the 7th clock cycle), a tristate NOR gate, a member of subgroup 909, provides 128 outputs on output lines "output 0 936" to "output 127 938". At time 7 (for example, the 8th clock cycle), a tristate NOR gate, a member of subgroup 910, provides 128 outputs on output lines "output 0 936" to "output 127 938".
[0089] Figure 9 further illustrates that the contents of memory array 902 are written to the storage circuits of subgroups 903, 904, 905, 906, 907, 908, 909, and 910. As previously mentioned, the contents (data elements) can be any kind of NVM and volatile memory. The contents (data elements) of memory array 902 can be written to the storage circuits of subgroups 903, 904, 905, and 906 while subgroups 903, 904, 905, and 906 are performing multiplication operations, and the contents can be written to the storage circuits of subgroups 907, 908, 909, and 910 while subgroups 903, 904, 905, and 906 are performing multiplication operations. Writing can be done to any subgroup that does not share a word line with an active subgroup. As shown in the figure, a single word line (e.g., WL) can be used. <0> 310) The two memory circuits above can share one output line (e.g., output 0 936), connect to separate bit lines (e.g., BL0 912 and BL1 914), and connect to one reference line (e.g., Vref 916), which makes it possible to store different data in the two memory circuits.
[0090] Figure 10 shows the time series of operations for subgroups 0 903 to 7 910 in Figure 9. In Figure 10, subgroup 0 903 is identified as "subgroup 0", subgroup 1 904 is identified as "subgroup 1", and so on.
[0091] Specifically, Figure 10 shows a chart 1000 containing a time sequence of a cycle operation from time 0 to time 7, and then starting again from time 0. The top of the time sequence identifies the clock signal 1002. Below the corresponding time sequence of the clock signal 1002, chart 1000 shows a column 1004 of subgroups, indicating which subgroups are performing and / or outputting mathematical operations. Below the column of subgroups, chart 1000 shows a pipeline 1006 of the adder tree, which indicates which stage of the adder tree is active. Since Figure 9 shows 128 outputs, the adder tree 600 in Figure 6 can be implemented to have three stages: the first stage (stage 0) includes receiving an input and storing the result in a buffer latch 610; the second stage (stage 1) includes processing the data from the buffer latch 610 and storing the result in a buffer latch 612; and the third stage (stage 2) includes processing the data from the buffer latch 612 and providing a single output to path 628.
[0092] As shown in the diagram, the adder tree pipeline 1006 indicates which stage of the adder tree is processing the input / data. Furthermore, this chart shows a download process 1008 that downloads content, which includes writing content to the memory circuits of various subgroups. For example, from time 0 to time 3, while subgroups 0, 1, 2, and 3 are performing multiplication and / or output operations, content is downloaded and written to the memory circuits of subgroups 4, 5, 6, and 7. Similarly, from time 4 to time 7, while subgroups 4, 5, 6, and 7 are performing multiplication and / or output operations, content is downloaded and written to the memory circuits of subgroups 0, 1, 2, and 3. Then, again, while subgroups 0, 1, 2, and 3 are performing multiplication and / or output operations using the content written between times 4 and 7, content is downloaded and written to the memory circuits of subgroups 4, 5, 6, and 7. The download and write time may take two or more cycles. Therefore, the advantage of the dCIM system 900 in Figure 9 is that there are four clock cycles available to write content to a specific set of subgroups (in this example, one set of subgroups contains four subgroups). This structure reduces or eliminates the time spent waiting for content writing to finish before initiating multiplication and / or output operations.
[0093] Figure 11 shows a dCIM system in which content is written using bit lines and a reference voltage (Vref).
[0094] Specifically, Figure 11 shows a dCIM system 1100 similar to the one described with reference to Figure 5, and therefore a redundant explanation is omitted. The dCIM system 1100 differs from the dCIM system 500 in Figure 5 in that the bit bar lines are replaced with Vref lines. Specifically, Figure 11 shows that Vref lines 1102, 1104, and 1106 are connected to the memory circuit for writing to the memory circuit. Replacing bit bar lines with Vref lines has the advantage of doubling the number of SA latches (sense amplifiers and latches) and thus doubling the sensing speed. For example, if there are 1024 memory BLs in the memory array, the 1024 memory BLs are divided into 512 SRAM BLs and 512 SRAM BLBs, which means that only 512 bits can be sensed at a time. If there are 1024 memory blocks connected to 1024 SRAM blocks, and 1024 BLBs connected to Vrefs, then 1024 bits of data can be sensed at once.
[0095] Figure 12 shows a dCIM system in which content is written using a sense amplifier.
[0096] Specifically, Figure 12 shows a dCIM system 1200 similar to the one described with reference to Figure 5, and therefore a redundant explanation is omitted. Figure 12 shows a memory array 1202 that stores content, and here a read operation 1210 is performed to read content from the memory array 1202 into sense amplifiers SA0 1204, SA1 1206, and SAm 1208 via the respective bit lines and bit bar lines BL0, BL0B, BL1, BL1B, BLm, and BLmB. As mentioned above, the memory array 1202 that stores content (data elements) can be any type of NVM and volatile memory. The dCIM system 1200 in Figure 12 differs from the dCIM system 500 in Figure 5 in that it includes sense amplifiers for providing content to various subgroups. As shown, a write operation 1212 is performed to write content from the sense amplifiers to the storage circuits of the various subgroups. Specifically, sense amplifier SA0 1206 provides content on SA DL0 line 1214 and SA DL0B line 1216 and writes this content to a specific memory circuit, sense amplifier SA1 1206 provides content on SA DL1 line 1218 and SA DL1B line 1220 and writes this content to a specific memory circuit, and sense amplifier SAm1208 provides content on SA DLm line 1222 and SA DLmB line 1224 and writes this content to a specific memory circuit. Alternatively, instead of writing content using bar lines (i.e., SA DL0B line 1216, SA DL1B line 1220, SA DLmB line 1224), content can be written using Vref. In contrast to programming memory circuits directly from bit lines and bit bar lines, using sense amplifiers enables faster programming because sense amplifiers can supply higher voltage signals than those that can be supplied on bit lines and bit bar lines.
[0097] Figure 13 shows a timing chart for executing MAC calculations using the dCIM system 500 shown in Figure 5.
[0098] Specifically, the timing chart 1300 shows that the input can remain stable until all weights have been multiplied by the input. For example, Figure 13 shows the active state subgroup 1302, the time series 1304 from time 0 to time N, and the input sequence (input column) 1306 of input 0 338 for the dCIM system 500 of Figure 5. This timing chart can be modified to match other dCIM systems described herein.
[0099] As illustrated, in Figure 5, N=64, so there are 64 subgroups. The active subgroup 1302 is shown to proceed through the entire subgroup in this chronological order. The input sequence 1306 on input 0 338 remains the same (of Xi0) and is multiplied by weights Wi0, Wi1, Wi2, Wi3, Wi4, Wi5, Wi6, and Wi7 from time 0 to time 7. Next, at time 8, the input changes to Xi1 and is multiplied by weights Wi0, Wi1, Wi2, Wi3, Wi4, Wi5, Wi6, and Wi7 from time 8 to time 15. This process continues until Xi7 is multiplied by weight Wi7 at time 63. As described herein, while a subgroup connected to a particular word line is performing a multiplication operation and / or output operation, the weights of other subgroups connected to other word lines are updated. This series of operations shown in Figure 13 can be referred to as "input stabilization and sequential serial input in product operation".
[0100] Figure 14 shows a dCIM system that uses a path gate instead of the tristate NOR gate shown in Figure 3.
[0101] Specifically, Figure 14 shows structure 1400, which is similar to the structure in Figure 4, and its redundant explanation is omitted. However, structure 1400 replaces the tristate NOR gates 332, 406, and 408 in Figure 4 with path gates 1402, 1404, and 1406. Just as in Figure 4, there can be N path gates for each subgroup. At time 0 (when timing signals time 0 and time 0B are received), path gate 1402 allows the content stored in memory circuit 316 to be passed to NOR gate 1408, which multiplies the content stored in memory circuit 316 by the input 1410 provided at time 0. N NOR gates can exist within structure 1400, and each of the N NOR gates corresponds to a specific sequence of path gates. NOR gate 1408 provides the output "output 0 1412" to the adder tree (accumulator). At time 1 (upon receiving timing signals time 1 and time 1B), pass gate 1404 allows the content stored in memory circuit 402 to be passed to NOR gate 1408, and NOR gate 1408 multiplies the content stored in memory circuit 402 by the input 1410 provided at time 1. At time N (upon receiving timing signals time N and time NB), pass gate 1406 allows the content stored in memory circuit 404 to be passed to NOR gate 1408, and NOR gate 1408 multiplies the content stored in memory circuit 404 by the input 1410 provided at time N. This implementation eliminates the need for tristate NOR gates and the space they occupy.
[0102] Figure 15 shows a simplified block diagram of an integrated circuit device that includes a memory array configured for in-memory computation with signed (or unsigned) inputs and weights.
[0103] Specifically, Figure 15 shows an integrated circuit device 1500 including a memory array 1560, which is configured for signed in-memory computations for CIM operations such as signed (or unsigned) multiply-accumulate (MAC) operations, performed by the dCIM system described herein. The integrated circuit device 1500 can be implemented on a single chip or on a multi-chip module.
[0104] Device 1500 includes input / output circuits 1506 for communicating control signals, data, addresses, and commands with other data processing resources such as a CPU (central processing unit) or memory controller.
[0105] Input / output data is provided to the controller 1510 and the cache 1590 on bus 1591. Addresses are provided to the decoder 1542 and the controller 1510 on bus 1593. Buses 1591 and 1593 can also be operationally connected to data sources within an integrated circuit device 1500, such as a general-purpose processor, a dedicated processor, or a combination of modules providing system-on-a-chip functionality.
[0106] The memory array 1560 may include an array of memory cells in the form of a NOR or NAND architecture, such that the memory cells are arranged in the form of columns along bit lines and rows along word lines, and the memory cells in a given column are connected in parallel between the bit lines and a reference source. The reference source may comprise a ground terminal or a power line connected to a bias power supply on the reference source side. These memory cells may comprise charge trap transistor cells arranged in the form of a 3D (three-dimensional) structure. The memory array 1560 having in-memory (or near-memory) computation can be configured as described above with respect to Figures 3-14 and can perform the computations described above with respect to Figures 3-14.
[0107] The bit lines can be connected to the global bit line 1565 by a block selection circuit, which is configured for selectable connections to the page buffer 1580 and the CIM sense circuit 1570.
[0108] In the illustrated embodiment, the page buffer 1580 is connected to the cache 1590 by a bus 1585. The page buffer 1580 includes memory elements (which can be various types of memory arrays) and sensing circuits for memory operations, including read and write operations. For flash memory, including dielectric charge trap memory and floating-gate charge trap memory, the write operation includes program and erase operations.
[0109] The driver circuit 1540 is connected to the word line 1545 in the array 1560 and applies a word line voltage to the selected word line in response to the decoder 1542 which decodes an address on the bus 1593, or, in computational operation, in response to input data stored in the input buffer 1541.
[0110] The controller 1510 is coupled to the cache 1590 and the memory array 1560, and is also coupled to other peripheral circuits used in memory access operations and memory computation operations.
[0111] The controller 1510, for example using a state machine, controls the application of power supply voltage and current generated or supplied through the voltage or current sources within block 1520 for memory operation and CIM operation.
[0112] The controller 1510 includes a control logic circuit that includes a control register and a state register, and can be implemented using a dedicated logic circuit that includes a state machine and combinational logic circuit known in the present art. In an alternative embodiment, the control logic circuit comprises a general-purpose processor, which can be implemented on the same integrated circuit and executes a computer program to control the operation of the device. In yet another embodiment, a combination of a dedicated logic circuit and a general-purpose processor can be used to implement the control logic circuit.
[0113] Array 1560 includes memory cells arranged in columns and rows, where memory cells in columns are connected to corresponding bit lines and memory cells in rows are connected to corresponding word lines. Array 1560 is programmable to store signed coefficients (weights Wi) in multiple sets of memory cells.
[0114] In CIM mode, the word line driver circuit 1540 or the driver circuit 1540 may include a driver (referred to as an input driver or input activation driver) configured to operate a signed or unsigned input Xi from the input buffer 1541. The driver circuit 1540 may be separate from the word line driver circuit. The CIM sense circuit 1570 is configured to sense the difference between the first and second currents on each bit line in a selected pair of bit lines and to generate an output for the selected pair of bit lines as a function of the difference. This output can be provided to memory elements in the page buffer 1580 and to the cache 1590.
[0115] In one embodiment, the first subgroup may include a first memory circuit and a first multiplier circuit, the first memory circuit being connected to a first word line and programmable to store a first set of weights, the first multiplier circuit being enabled by a first timing signal to multiply the input by the first set of weights to provide a first output, the second subgroup may include a second memory circuit and a second multiplier circuit, the second memory circuit being connected to a second word line and programmable to store a second set of weights, the second multiplier circuit being enabled by a second timing signal to multiply the input by the second set of weights to provide a second output, the second multiplier circuit being enabled at a different time than the time the first multiplier circuit is enabled, and the accumulator circuit receives and accumulates (i) a first output corresponding to the first multiplier circuit being enabled by the first timing signal, and (ii) a second output corresponding to the second multiplier circuit being enabled by the second timing signal.
[0116] In one embodiment, the in-memory calculation circuit may include a first output line and a second output line, the first output line being shared by the output of one of the first multiplier circuits of a first subgroup and the output of one of the second multiplier circuits of a second subgroup, the second output line being shared by the output of another of the first multiplier circuits of a first subgroup and the output of another of the second multiplier circuits of a second subgroup, the first output of the first subgroup being provided to the accumulator circuit via the first and second output lines depending on whether the first multiplier circuit is enabled and the second multiplier circuit is not enabled, and the second output of the second subgroup being provided to the accumulator circuit via the first and second output lines depending on whether the second multiplier circuit is enabled and the first multiplier circuit is not enabled.
[0117] In a further embodiment, a specific memory circuit among the first memory circuits of a first subgroup and a specific memory circuit among the second memory circuits of a second subgroup can share a common programming line for controlling the storage of their respective weights, wherein the specific memory circuit of the first subgroup is programmed to store a specific weight from a first set of weights in response to a first word line activating the first subgroup, and the specific memory circuit of the second subgroup is programmed to store a specific weight from a second set of weights in response to a second word line activating the second subgroup.
[0118] In one embodiment, the in-memory calculation circuit may include a multiplication circuit, a first subgroup of the circuit, a second subgroup of the circuit, and an accumulation circuit, wherein the multiplication circuit is configured to receive an input, multiply it, and provide an output; the first subgroup of the circuit is connected to a first word line and is configured to (i) store a first set of weights and (ii) provide the first set of weights to the multiplication circuit in response to a first timing signal enabling a first pass gate; the second subgroup of the circuit is connected to a second word line and is configured to (i) store a second set of weights and (ii) enable a second pass gate in response to a second timing signal Depending on the state, the circuit is configured to provide a second set of weights to the multiplier circuit, the provision of the second set of weights is enabled at a different time than the provision of the first set of weights, and the accumulation circuit is shared by the first subgroup and the second subgroup, and is configured to (i) receive and accumulate a first output received from the multiplier circuit in response to the first pass gate being enabled by the first timing signal, and (ii) receive and accumulate a second output received from the multiplier circuit in response to the second pass gate being enabled by the second timing signal.
[0119] The realization of the memory array can be based on charge and wrapping memory cells, such as floating gate memory cells that can include polysilicon charge and wrapping layers, or dielectric charge and wrapping memory cells that can include silicon nitride charge and wrapping layers. In various embodiments of the technology described herein, other types of memory technology can be applied.
[0120] Other implementations of the methods described in this section may include a non-temporary computer-readable storage medium that stores executable instructions for a processor to perform any of the methods described above, and yet another implementation of the methods described in this section may include a system including memory and one or more processors, the one or more processors being operable to execute instructions stored in memory to perform any of the methods described above.
[0121] Any data structures and codes described above or referenced above can be stored in numerous implementations on computer-readable storage media, which can be any device or medium capable of storing code and / or data for use by a computer system. These devices or mediums include, but are not limited to, volatile memory, non-volatile memory, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), magnetic disks, magnetic tapes, CDs (compact discs), DVDs (digital versatile discs or digital video discs), or other media capable of storing computer-readable media, whether currently known or to be developed.
[0122] An example of a processor is a hardware unit (equipped with hardware circuitry such as one or more active devices) that is enabled to execute program code. A processor optionally comprises one or more controllers and / or state machines. Processors can be realized by application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and / or custom design techniques. Processors can be manufactured using integrated circuits, and optical and quantum technologies. A processor employs one or more architectural techniques, such as sequential (e.g., von Neumann) processing or very long instruction word processing. A processor employs one or more microarchitectures that execute multiple instructions one at a time or in parallel. A processor can be directed to multipurpose users (and / or) purpose-specific users (such as signal, audio, video, and / or graphics users). A processor can have fixed functions or, for example, programmable functions. A processor comprises one or more of the following: registers, memory, logic units, arithmetic units, and graphics units. The term "processor" includes both singular and plural forms, such as multiprocessors and / or collections of processors.
[0123] The logic circuits described herein can be implemented using a processor programmed with a computer program, which is stored in memory accessible to the computer system and can be executed by the processor, by dedicated logic hardware including field-programmable integrated circuits, and by a combination of the dedicated logic hardware and the computer program. From all the flowcharts herein, it will be understood that many of the steps can be combined, executed in parallel, or executed in different orders without affecting the function being implemented. In some cases, as the reader will understand, rearranging the steps will achieve the same result only if other specific modifications are also made. In other cases, as the reader will understand, rearranging the steps will achieve the same result only if certain conditions are met. Furthermore, it will be understood that the flowcharts herein show only the steps relevant to understanding the invention, and that many additional steps to achieve other functions can be performed before, after, and between the illustrated steps.
[0124] The present invention is disclosed by reference to the preferred embodiments and examples detailed above, but it should be understood that these examples are intended to be illustrative, not restrictive. Modifications and combinations will be readily conceivable to those skilled in the art, and these modifications and combinations are intended to be in the spirit of the invention and within the scope of the following claims.< / n> < / n> < / n>
Claims
1. One or more input lines that receive M input data elements, where M is an integer greater than 0. One or more input lines, An array of memory cells comprising one or more subgroups, wherein the one or more subgroups Each subgroup in the loop is an array of memory cells that store M memory data elements, The array of memory cells and the one or more input lines are connected to the M input data The element contains the M entries within the selected subgroup from the one or more subgroups. It is configured to multiply millions of data elements to provide a multiplier output having M data elements. A multiplication circuit and The multiplier output is connected to the accumulator input of the M data elements, Includes an accumulation circuit configured to generate the sum of the M data elements of the multiplier output. A memory-based calculation circuit, The multiplication circuit provides the multiplication result from one or more subgroups to the multiplier output. The multiplication circuit is enabled by a timing control signal and provides the multiplication result in a memory-based calculation circuit.
2. The multiplication circuit applies to each subgroup in the one or more subgroups, the multiplier The in-memory computation according to claim 1, including M tristate multipliers connected to the output. Road.
3. The claim states that the M tristate multipliers are M tristate NOR gates. The memory-based calculation circuit described in 2.
4. The one or more subgroups mentioned above are a first subgroup that stores M memory data elements and , including a second subgroup that stores M memory data elements, The M tristate multipliers for the first subgroup are configured to perform the first timing signal. Therefore, it is enabled, and the first subgroup is set to the M input data elements. Multiply the M memory data elements, The M tristate multipliers for the second subgroup are configured to handle the second timing signal. Therefore, it is enabled, and the M input data elements are set to the second subgroup Multiply the M memory data elements and the M tristable for the second subgroup The multiplier is at a different time than the M tristate multipliers for the first subgroup. The second timing signal and the first timing signal are set to enable the second timing signal. The memory-based calculation circuit according to claim 2, wherein the components are supplied at different times.
5. the one or more sub-groups store M memory data elements in M memory circuits, the first sub-group and the second sub-group that store M memory data elements in M memory circuits and the first sub-group is connected to a first word line, the second sub-group is connected to a second word line, the in-memory calculation circuit according to claim 2 . **Claim 6** a specific memory circuit among the M memory circuits of the first sub-group and a specific memory circuit among the M memory circuits of the second sub-group share a common line for controlling the storage of respective data elements, the specific memory circuit of the first sub-group stores a specific data element in response to the first word line activating the first sub-group, the specific memory circuit of the second sub-group stores a specific data element in response to the second word line activating the second sub-group, the in-memory calculation circuit according to claim 5 . **Claim 7** the common line shared by the specific memory circuit of the first sub-group and the specific memory circuit of the second sub-group includes a bit line (BL), the in-memory calculation circuit according to claim 6 . **Claim 8** (i) While the M tristate multipliers for the first sub-group are enabled by a first timing signal to multiply the M input data elements by the M memory data elements of the first sub-group and provide a multiplier output having the M data elements, and (ii) while the accumulation circuit receives and accumulates the multiplier output having the M data elements, a specific data element is written into a specific memory circuit among the M memory circuits of the second sub-group, the in-memory calculation circuit according to claim 5 . **Claim 9** This is enabled, and the M tristate powers for the second subgroup In accordance with the fact that the calculator is not enabled by the timing control signal, The output associated with one subgroup is routed through the first output line and the second output line to the cumulative Provided to the circuit, The M tristate multipliers for the second subgroup control the timing control signal The M tristaes for the first subgroup are enabled by and Depending on whether the multiplier is enabled by the timing control signal, The output related to the second subgroup is transmitted via the first output line and the second output line. A memory-based calculation circuit according to claim 5, provided for a cumulative circuit.
10. The one or more subgroups mentioned above are a first subgroup that stores M memory data elements and , a second subgroup that stores M memory data elements, and a second subgroup that stores M memory data elements It includes a third subgroup and a fourth subgroup that stores M memory data elements, The first subgroup and the second subgroup are connected to the first word line, The third subgroup and the fourth subgroup are connected to the second word line, The M tristate multipliers for the first subgroup are configured to perform the first timing signal. Therefore, it is enabled, and the M input data elements are set to the first subgroup. Multiply the M stored data elements, The M tristate multipliers for the second subgroup are configured to handle the second timing signal. Therefore, it is enabled, and the M input data elements are set to the preceding second subgroup. Multiply the M stored data elements, The M tristate multipliers for the third subgroup are configured to handle the third timing signal. Therefore, it is enabled, and the M input data elements are set to the third subgroup before Multiply the M stored data elements, The M tristate multipliers for the fourth subgroup are configured to handle the fourth timing signal. Therefore, it is enabled, and the M input data elements are set to the preceding fourth subgroup. Multiply the M stored data elements, L is the M storage data elements of the first subgroup and the front of the second subgroup. An integer representing the total number of M stored data elements, where M = L / 2, as described in claim 2. The built-in memory-based calculation circuit.
11. The multiplier output includes a first output line, The first output line is the output of one of the tristate multipliers for the first subgroup. and the output of one of the tristate multipliers for the second subgroup, and the third sub The output of one of the tristate multipliers for the group and one for the fourth subgroup The memory according to claim 10, which is shared with the output of the tristate multiplier. calculation circuit.
12. The multiplier output includes a second output line, The second output line is for another tristate multiplier for the first subgroup. The output, the output of another tristate multiplier for the second subgroup, and the The output of another tristate multiplier for the third subgroup, and the fourth subgroup The output of the other tristate multiplier for the loop is shared by the output of the other tristate multiplier, claim 11. The memory-based calculation circuit described above.
13. In each of the one or more subgroups, the M storage data elements of each subgroup are The memo according to claim 1 is written to each of the subgroups using bit lines. Internal calculation circuit.
14. In each of the one or more subgroups, the M storage data elements of each subgroup are The data is written to each of the aforementioned subgroups using a sense amplifier connected to the bit line. The memory-based calculation circuit according to claim 1.
15. The one or more subgroups mentioned above are a first subgroup that stores M memory data elements and , including a second subgroup that stores M memory data elements, During the first clock cycle, the first subgroup has the M storage data elements The M input data elements are multiplied, and the second subgroup contains the M stored data elements It has been written, During the second clock cycle, the second subgroup has the M storage data elements The M input data elements are multiplied, and the first subgroup contains the M stored data elements A memory-based calculation circuit according to claim 1, wherein the following is written to it.
16. The one or more subgroups mentioned above are a first subgroup that stores M memory data elements and , including a second subgroup that stores M memory data elements, During a specific clock cycle, the accumulating circuit outputs the output associated with the first subgroup Cumulatively, During the subsequent clock cycle, the cumulative circuit outputs the output associated with the second subgroup. A memory-based calculation circuit according to claim 1, which accumulates the values.
17. The memory-based calculation circuit according to claim 16, wherein the cumulative circuit is pipelined.
18. The multiplication circuit applies to each subgroup in the one or more subgroups, the multiplier The claim includes M path gates connected to a shared M bit multiplier connected to the output. The memory-based calculation circuit described in item 1.
19. The timing control signal includes a first timing signal and a second timing signal, Claim 1: The timing signal is supplied at a different time than the second timing signal. The memory-based calculation circuit described above.
20. The in-memory computation according to claim 7, wherein the common line further includes a bit line bar line (BLB). circuit.
21. The memory calculation cycle according to claim 7, wherein the common line further includes a reference voltage line (VREF) Road.
22. The memory-based computing circuit according to claim 1, wherein the array of memory cells includes a latch.
23. A method for performing calculations using an in-memory calculation circuit, wherein the in-memory calculation circuit is (i ) An array of memory cells comprising one or more subgroups, wherein the one or more subgroups Each subgroup in the loop stores M memory data elements, where M is an integer greater than 0. (ii) an array of memory cells and one or more input lines connected to the array of memory cells (iii) A multiplier circuit and an accumulator input of M data elements connected to the multiplier output. In a method including a cumulative circuit that includes force, A step of obtaining M input data elements from one or more input lines, The multiplication circuit multiplies the M input data elements by one or more subgroups Multiply the M memory data elements within our selected subgroup to obtain M data elements A step of providing a multiplier output having elements, wherein the multiplier circuit is a timing control signal. It is enabled by and the multiplication result from one or more subgroups is multiplied by the multiplication Steps include sequentially providing the output to the device, The cumulative circuit generates the sum of the M data elements of the multiplier output. Top and A method that includes this.
24. The first subgroup of the circuit is connected to the first word line and configured to store the first set of weights. Loops and The second subgroup of the circuit is connected to the second word line and configured to store the second set of weights. Loops and Multiplication circuits and, The system comprises a cumulative circuit shared by the first subgroup and the second subgroup. A memory-based calculation circuit, The multiplication circuit (i) multiplies the input by the first set of weights in accordance with the first timing signal. (ii) provides a first output, and (iii) inputs the second set of weights according to the second timing signal. (iv) is configured to multiply by the second set of weights and provide a second output, wherein the multiplication of the second set of weights is the The multiplication of one set of weights is enabled at a different time than the time when it is enabled. 、 The cumulative circuit (i) multiplies the first set of weights by the first timing signal In response to being put into a cable state, the first output is received and accumulated, and (ii) the second set of weights In response to the multiplication being enabled by the second timing signal, A memory-based calculation circuit configured to receive and accumulate the second output.