In-memory computing circuit and method of performing operation thereof
By using three-state NAG multiplier and timing control signals in the digital memory computing system, multiple subgroups are allowed to share the adder tree, and by optimizing the word line direction, the problems of large number of adder trees and low download efficiency in the prior art are solved, and more efficient multiplication accumulation operations and content downloads are achieved.
Patent Information
- Application Number
- CN202411358863.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-01
- Filing Date
- 2024-09-27
- Publication Date
- 2025-07-01
AI Technical Summary
When the existing digital memory computing system performs multiplication and accumulation operations, the existence of multiple adder trees leads to high space occupation and is inefficient when downloading new content (such as weights), and the multiplication operations of multiple subgroups must be stopped.
Using a multiplier with a three-state output, by replacing the general NATO gate with a three-state NATO gate, the three-state NATO gate of each subgroup is controlled by a separate timing signal, allowing multiple subgroups to share an adder tree, and arranging the subgroups by word line direction perpendicular to the bit line and complementary bit line to improve download efficiency.
Reduces the number of adder trees, reduces the layout area, and improves the efficiency of downloading new content, which improves the overall performance of MAC operations.
Smart Images

Figure CN120233980A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to performing in-memory or near-memory computations, such as multiply-accumulate (MAC) or other product-sum operations. Background Art
[0002] In neural network computing systems, machine learning systems, and circuits for certain types of linear algebra computations, the multiply-accumulate or product-sum function can be an important building block. These functions can be expressed as follows:
[0003]
[0004] In this formula, each product term is the product of a variable input X i and a weight W i The weight W i can vary between terms, for example, as coefficients corresponding to the variable input X i
[0005] The product-sum function can be implemented using a cross-point array architecture as a circuit operation, where the electrical characteristics of the storage cells of the array implement the function.
[0006] These architectures can be implemented in digital in-memory computing (dCIM) systems and digital near-memory computing (dNMC) systems to perform the multiply-accumulate (MAC) operations described in the above formula. In known systems, a product subgroup is accompanied by a corresponding adder tree (e.g., an accumulator). Thus, there are many adder trees because each subgroup (in a larger group) has its own corresponding adder tree. The adder trees occupy a relatively large layout area and are thus costly in terms of space. In known systems, the adder trees occupy a non-required amount of space. For brevity, the term dCIM system hereinafter also includes dNMC systems.
[0007] In addition, in these digital in-memory computing systems, the time required to download memory content (such as weights) is undesirably long due to the need to switch multiple or all word lines (WL), resulting in slower performance. Specifically, in digital in-memory computing systems, during the download of content (such as weights), the multiplication operations of multiple subgroups must be stopped, which reduces performance.
[0008] Therefore, there is a need to provide a digital in-memory computing system with a reduced number of adder trees that can perform MAC operations or similar operations while also downloading new content (such as weights). Summary of the Invention
[0009] In one embodiment, a in-memory computing circuit is provided. The in-memory computing circuit may include: one or more input lines that receive M input data elements, where M is an integer greater than zero; a memory cell array including one or more subgroups, each subgroup of the one or more subgroups storing M stored data elements; a plurality of multiplication circuits connected to the memory cell array and the one or more input lines and configured to multiply the M input data elements by the M stored data elements in a selected subgroup of the one or more subgroups and provide a multiplier output having M data elements; and an accumulation circuit including accumulator inputs for the M data elements, the M data elements being connected to the multiplier output, and the accumulation circuit being configured to produce a sum of the M data elements of the multiplier output. Wherein the plurality of multiplication circuits provide multiplication results from the one or more subgroups to the multiplier output.
[0010] In another embodiment, for each subgroup of the one or more subgroups, the plurality of multiplication circuits include M tri-state multipliers connected to the multiplier output.
[0011] In another embodiment, the M tri-state multipliers are M tri-state NOR gates.
[0012] In one embodiment, the one or more subgroups include a first subgroup storing M stored data elements and a second subgroup storing M stored data elements. The M tri-state multipliers of the first subgroup are enabled by a first timing signal to multiply the M input data elements by the M stored data elements of the first subgroup. The M tri-state multipliers of the second subgroup are enabled by a second timing signal to multiply the M input data elements by the M stored data elements of the second subgroup, such that the M tri-state multipliers of the second subgroup are enabled at a different time than the M tri-state multipliers of the first subgroup, and the second timing signal is provided at a different time than the first timing signal.
[0013] In another embodiment, the one or more subgroups include a first subgroup storing M stored data elements in M storage circuits and a second subgroup storing M stored data elements in M storage circuits. The first subgroup is connected to a first word line. The second subgroup is connected to a second word line.
[0014] In another embodiment, a specific storage circuit of the M storage circuits in the first subgroup and a specific storage circuit of the M storage circuits in the second subgroup share a common line for controlling the storage of respective data elements. The specific storage circuit in the first subgroup activates the first subgroup to store a specific data element depending on the first word line. The specific storage circuit in the second subgroup activates the second subgroup to store a specific data element depending on the second word line.
[0015] In one embodiment, the common line shared by the specific storage circuit of the first subgroup and the specific storage circuit of the second subgroup includes a bit line.
[0016] In another embodiment, the specific storage circuit of the M storage circuits in the second subgroup writes a specific data element in at least one of the following cases: (i) the M three-state multipliers in the first subgroup are enabled by a first timing signal to multiply the M input data elements by the M stored data elements in the first subgroup to provide a multiplier output having M data elements, and (ii) the accumulation circuit receives and accumulates the multiplier output having M data elements.
[0017] In another embodiment, the multiplier output includes a first output line and a second output line. The first output line is shared by the output of a three-state multiplier in the first subgroup and the output of a three-state multiplier in the second subgroup. The second output line is shared by the output of another three-state multiplier in the first subgroup and the output of another three-state multiplier in the second subgroup. The output related to the first subgroup is provided to the accumulation circuit through the first and second output lines depending on the M three-state multipliers in the first subgroup being enabled by a timing control signal while the M three-state multipliers in the second subgroup are not enabled by the timing control signal. The output related to the second subgroup is provided to the accumulation circuit through the first and second output lines depending on the M three-state multipliers in the second subgroup being enabled by the timing control signal while the M three-state multipliers in the first subgroup are not enabled by the timing control signal.
[0018] In one embodiment, the one or more subgroups include a first subgroup storing M stored data elements, a second subgroup storing M stored data elements, a third subgroup storing M stored data elements, and a fourth subgroup storing M stored data elements. The first subgroup and the second subgroup are connected to a first output line. The third subgroup and the fourth subgroup are connected to a second output line. The M three-state multipliers of the first subgroup are enabled by a first timing signal to multiply the M input data elements by the M stored data elements of the first subgroup. The M three-state multipliers of the second subgroup are enabled by a second timing signal to multiply the M input data elements by the M stored data elements of the second subgroup. The M three-state multipliers of the third subgroup are enabled by a third timing signal to multiply the M input data elements by the M stored data elements of the third subgroup. The M three-state multipliers of the fourth subgroup are enabled by a fourth timing signal to multiply the M input data elements by the M stored data elements of the fourth subgroup. L is an integer representing the total number of the M stored data elements of the first subgroup and the M stored data elements of the second subgroup, and where M = L / 2.
[0019] In another embodiment, the multiplier output includes a first output line. The first output line is shared by the output of one three-state multiplier of the first subgroup, the output of one three-state multiplier of the second subgroup, the output of one three-state multiplier of the third subgroup, and the output of one three-state multiplier of the fourth subgroup.
[0020] In one embodiment, the multiplier output includes a second output line. The second output line is shared by the output of another three-state multiplier of the first subgroup, the output of another three-state multiplier of the second subgroup, the output of another three-state multiplier of the third subgroup, and the output of another three-state multiplier of the fourth subgroup.
[0021] In another embodiment, the M stored data elements of each respective subgroup of the one or more subgroups are written to each respective subgroup using bit lines.
[0022] In another embodiment, the M stored data elements of each respective subgroup of the one or more subgroups are written to each respective subgroup using sense amplifiers connected to the bit lines.
[0023] In one embodiment, the one or more subgroups include a first subgroup storing M stored data elements and a second subgroup storing M stored data elements. During a first clock cycle, the first subgroup multiplies the M stored data elements by the M input data elements, and the second subgroup has the M stored elements written therein. During a second clock cycle, the second subgroup multiplies the M stored data elements by the M input data elements, and the first subgroup has the M stored elements written therein.
[0024] In another embodiment, the one or more subgroups include a first subgroup storing M stored data elements and a second subgroup storing M stored data elements. During a particular clock cycle, the accumulation circuit accumulates the output associated with the first subgroup. During a subsequent clock cycle, the accumulation circuit accumulates the output associated with the second subgroup.
[0025] In another embodiment, the accumulation circuit is pipelined.
[0026] In one embodiment, for each of the one or more subgroups, the multiplication circuit includes M pass gates, the M pass gates are connected to a shared M-bit multiplier, and the shared M-bit multiplier is connected to the multiplier output.
[0027] In another embodiment, the multiplication circuit is enabled by a timing control signal to provide the multiplication result. In one embodiment, the timing control signal includes a first timing signal and a second timing signal such that the first timing signal and the second timing signal are provided at different times.
[0028] In another embodiment, a method of performing an operation using an in-memory computing circuit is provided. The in-memory computing circuit includes (i) an array of memory cells including one or more subgroups, each of the one or more subgroups storing M stored data elements, where M is an integer greater than zero, (ii) a multiplication circuit connected to the array of memory cells and one or more input lines, and (iii) an accumulation circuit including an accumulator input for M data elements, the accumulator input being connected to the multiplier output. Further, the method may include obtaining M input data elements from the one or more input lines; multiplying, by the multiplication circuit, the M input data elements by the M stored data elements in a selected subgroup of the one or more subgroups to provide a multiplier output having M data elements, wherein the multiplication circuit is enabled by a timing control signal to sequentially provide multiplication results from the one or more subgroups to the multiplier output; and generating, by the accumulation circuit, a sum of the M data elements of the multiplier output.
[0029] In another embodiment, an in-memory computing circuit is provided. The in-memory computing circuit may include: a first sub-group circuit connected to a first word line and configured to store a first set of weights; a second sub-group circuit connected to a second word line and configured to store a second set of weights; a multiplication circuit configured to (i) multiply the first set of weights by an input according to a first timing signal, (ii) provide a first output, (iii) multiply the second set of weights by the input according to a second timing signal, and (iv) provide a second output, wherein the multiplication of the second set of weights is enabled at a different time than the multiplication of the first set of weights; and an accumulation circuit shared by the first sub-group and the second sub-group and configured to receive and accumulate (i) the first output obtained by enabling the multiplication of the first set of weights by the first timing signal, and (ii) the second output obtained by enabling the multiplication of the second set of weights by the second timing signal.
[0030] In one embodiment, the common line further includes complementary bit lines.
[0031] In another embodiment, the common line further includes a reference voltage line.
[0032] In another embodiment, the memory cell array includes a plurality of latches.
[0033] Other aspects and advantages of the present invention can be seen in reviewing the drawings, the detailed description, and the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Illustrates a known SRAM-based digital in-memory computing system that includes a plurality of multiplication operation sub-groups, where each sub-group has its own corresponding adder tree (accumulator).
[0035] Figure 2 Description Figure 1 A exemplary sub-group of the known SRAM-based digital in-memory computing system in.
[0036] Figure 3 Illustrates a digital in-memory computing system including tri-state NOR gates that allows sharing of an adder tree between various sub-groups performing MAC operations.
[0037] Figure 4 Description Figure 3 An enlarged view of three different sub-group portions of the digital in-memory computing system in.
[0038] Figure 5 Description Figure 3 of the digital in-memory computing system, where content (data) is written to sub-group_0 while one or more other sub-groups are processing the input and calculating the output.
[0039] Figure 6 Describe the adder tree pipeline used in an example of a computing system in a digital memory Figure 3
[0040] Figure 7 Describe a computing system in a digital memory where two three-state NOR gates from two different subgroups share an output line and an adder tree.
[0041] Figure 8 Describe Figure 7 the adder tree of the computing system in the digital memory in
[0042] Figure 9 Describe a computing system in a digital memory where four NOR gates from four different subgroups share an output line and an adder tree.
[0043] Figure 10 Describe Figure 9 the operation time series from subgroup_0 903 to subgroup_7 910 in
[0044] Figure 11 Describe a computing system in a digital memory where content is written using bit lines and a reference voltage (Vref).
[0045] Figure 12 Describe a computing system in a digital memory where content is written using a sense amplifier.
[0046] Figure 13 Describe the timing diagram for performing MAC operations using a computing system in a digital memory.
[0047] Figure 14 Describe a computing system in a digital memory that uses transmission gates instead of Figure 3 the three-state NOR gates shown in
[0048] Figure 15 Describe a simplified block diagram of an integrated circuit device that includes a memory array configured for in-memory computing with signed (or unsigned) inputs and weights.
[0049] Description of reference numerals:
[0050] 4b: Input
[0051] 5b: Operation
[0052] 100: Digital memory-based computing system of SRAM type
[0053] 102, 200: Subgroups
[0054] 104, 210: Adder tree
[0055] 106: Input start driver and SRAM word line driver
[0056] 202: Storage
[0057] 205, IN_B<0>, IN_B<1>, IN_B<2>, IN_B<3>: Inputs (or input lines)
[0058] 204, WL<0:255>: Word lines (or inputs, input lines)
[0059] IN<0:255>: Input lines (or inputs)
[0060] 206, BL<0>, BL<1>, BL_<2>, BL<3>: Bit lines (inputs)
[0061] 207: NOR gate
[0062] 208, BLB<0>, BLB<1>, BLB_<2>, BL<3>: Complementary bit lines (complementary inputs)
[0063] 300, 500, 700, 900, 1100, 1200: Digital in - body computing systems
[0064] 302, 304, 306: Sub - groups_0 to sub - group N
[0065] 308: Adder tree
[0066] 310, 312, 314, WL<0>, WL<1>, WL <n>: Word line
[0067] 316, 317, 318, 402, 404: Storage circuit
[0068] 332, 334, 336, 406, 408: NOR gate
[0069] 320, 324, 328, BL0, BL1, BLm: Bit line
[0070] 322, 326, 330, BL0B, BL1B, BLmB: Complementary bit line
[0071] 338, 340, 342, 344, out0, out1, outm: Output line (or output)
[0072] 346: MAC output
[0073] 502, 504, 506: Operation
[0074] 600: Adder tree
[0075] 601: Input of adder tree
[0076] 610, 612: Buffer latch
[0077] 602, 606, 608, out0~out127: Output line (or output)
[0078] 614, 616, 618, 620, 622, 624, 626: First to seventh adder layers
[0079] 628: Path
[0080] 630: Equation (for calculating MAC)
[0081] 702, 703, 704, 705, 706, 707: Subgroup_0, …, Subgroup_N
[0082] 708, outm / 2: Output line (or output)
[0083] 800: Adder tree
[0084] 801: Input of adder tree
[0085] 810: Buffer latch
[0086] 802, 804, 806, 806, out0~out63: Output line(or output)
[0087] 814, 816, 818, 820, 822, 824: First to sixth adder layers
[0088] 902: Memory array
[0089] 903~910: Subgroup_0, …, Subgroup_7
[0090] 912, 914, 920, 922, 926, 928, BL0~B:511: Bit lines
[0091] 916, 918, 924, Vref: Reference voltage lines
[0092] 930, 932, 934: Input lines (or inputs)
[0093] 936, 938, out0, out127: Output lines (or outputs)
[0094] 1000: Chart
[0095] 1002: Clock signal
[0096] 1004: Subgroup sequence
[0097] 1006: Adder tree pipeline
[0098] 1008: Download program
[0099] 1102, 1104, 1106: Reference voltage lines
[0100] 1202: Memory array
[0101] 1210: Read operation
[0102] 1212: Write operation
[0103] 1204, 1206, 1208, SA0, SA1, SAm: Sense amplifiers
[0104] 1214, 1218, 1222, SADL0, SADL1, SADLm: Sense amplifier download lines
[0105] 1216, 1220, 1224, SADL0B, SADL1B, SADLmB: Complementary sense amplifier download lines
[0106] 1300: Timing diagram
[0107] 1302: Start subgroup
[0108] 1304: Time series
[0109] 1306: Input sequence (input_0 etc.)
[0110] 1400: Structure
[0111] 1402, 1404, and 1406: Transmission gates
[0112] 1410: Input
[0113] 1408: NOR gate
[0114] 1412: Output out0
[0115] 1500: Integrated circuit device
[0116] 1505: Input / output circuit
[0117] 1510: Controller
[0118] 1520: Bias setting
[0119] 1540: Driver
[0120] 1541: Input buffer
[0121] 1542: Decoder
[0122] 1545: Word line
[0123] 1560: Memory array
[0124] 1565: Bit line
[0125] 1570: CIM sensing circuit
[0126] 1580: Page buffer
[0127] 1590: Cache
[0128] 1585, 1591, 1593: Bus
[0129] Time_0 to Time_N, Time_00, Time_01, Time_11, Time_N0, Time_N1: Timing signals (or clock signals)
[0130] time_0 to time_N, time_00, time_01, time_11, time_N0, time_N1: Time Detailed implementation manners
[0131] Refer to Figures 1-15 Provide a detailed description of the embodiments of the present invention.
[0132] Figure 1 Describe a known digital memory-based computing system with SRAM type, which includes multiple subgroups of multipliers, and each subgroup has its corresponding adder tree (accumulator). Figure 1 The in-digital-memory computing system is considered SRAM-based because the weights are stored and retrieved by SRAM.
[0133] Specifically, Figure 1 illustrate a known SRAM-based in-digital-memory computing system 100, which includes a plurality of subgroups of a group, such as subgroup 102, where each subgroup utilizes a corresponding adder tree. For example, subgroup 102 receives inputs and word line signals from an input start driver and an SRAM word line driver 106 to perform mathematical operations and utilizes an adder tree 104 to perform accumulation operations. Each subgroup includes a storage circuit for storing weights and a circuit for performing mathematical operations, such as multiplying the weights by the inputs.
[0134] In this example, a storage device such as a 6-transistor (6T) SRAM memory cell is used to store the weights, and a multiplier such as a 4-transistor (4T) NOR gate is used to multiply the weights by the inputs (IN<0:255>). Additionally, as Figure 1 shown, the input start driver and the SRAM word line driver receive the inputs on the IN<0:255> lines and receive the sense line signals on the word lines WL<0:255>. Each input IN<0:255> is received by subgroup 102 (and other subgroups) so that they can be multiplied by the stored weights. The word lines WL<0:255> can be used to access the storage device for reading and writing (e.g., writing the weights to the storage device and / or reading the weights from the storage device). The output of the multiplication operation of subgroup 102 is provided as an input 4b to the adder tree 104. The adder tree 104 can combine various inputs in operation 5b and then provide a single output. In this example, a subgroup with four columns and 256 rows of memory cells implements Ini × Wi<0:3> [i = 0~255] to obtain a 1024-bit output (one 4-bit output per row), and combined with a 1024-input-bit adder tree, can complete the MAC operation. The multiplication operation can be performed by a NOR gate or other circuit types capable of performing multiplication (or other types of mathematics) operations.
[0135] As Figure 1 shown, this known SRAM-based in-digital-memory computing system 100 requires a separate adder tree to be configured for each subgroup. In this example, there are 64 subgroups, and thus 64 adder trees, which occupy a large amount of physical space in the known SRAM-based in-digital-memory computing system 100. Specifically, one of the problems of this known SRAM-based in-digital-memory computing system 100 is that the adder trees cannot be shared among other subgroups. Therefore, as the number of adder trees increases based on the number of subgroups, the overall layout size of the SRAM-based in-digital-memory computing system 100 also increases.
[0136] Figure 2 Description Figure 1 A sub - group example of a computing system within a digital memory of the SRAM type is known in the art.
[0137] Specifically, Figure 2 A sub - group 200 is described that receives an input 204 (WL<0:255>), an input 205 (IN_B<0:255>), an input 206 (BL<0:3>), and a complementary line input 208 (BLB<0:3>), where the input 206 (BL<0:3>) and the complementary line input 208 (BLB<0:3>) are used to program a memory 202 that stores weights. In addition, the sub - group 200 includes a NOR gate 207 for multiplying the stored weights by the input 205. The multiplication result is provided to an adder tree 210. As discussed above with reference to Figure 1 Each sub - group 200 communicates with a corresponding adder tree 210. Figure 2 Only a single sub - group and one adder tree are described, but as Figure 1 shown, at least 64 adder trees are required for computing within a digital memory of the SRAM type.
[0138] To update the memory 202 of the sub - group to store a new set of weights (e.g., the values of the weights), all word lines WL<0:255> are sequentially enabled to update the content of all SRAMs. Due to the time required to store the new values, it is costly to update the memory 202 with a new set of weights. In addition, since updating the memory 202 affects the data received by the adder tree 210, the entire MAC operation of all sub - groups must be stopped when updating the memory 202. This further reduces the efficiency of computing within a digital memory of the SRAM type.
[0139] The technology of the present invention solves these drawbacks by demonstrating computing within a digital memory of the SRAM type with reduced layout area and improved performance.
[0140] Specifically, compared with Figure 1 and Figure 2 Compared with the system of [[[ID=]], the technology of the present invention uses a multiplier with a tri-state output. For example, by replacing a general NOR gate with a tri-state NOR gate, where each subgroup of tri-state NOR gates can be controlled by a separate timing signal (e.g., a clock signal, etc.). Alternatively, the tri-state NOR gate can be replaced by any type of logic gate that can perform the product operation of the input and stored data (latched data bits) and / or can be controlled / enabled according to a timing signal. The tri-state output of the tri-state NOR gate has a high-impedance state when not enabled. This allows many tri-state NOR gate multipliers to be connected to each input of the accumulator, and only the enabled tri-state NOR gate multiplier selected by the separate timing signal will drive the input. This structure allows multiple subgroups to share an adder tree, thereby reducing the layout area.
[0141] In addition, the technology of the present invention can arrange each subgroup along the word line direction perpendicular to the bit line and complementary bit line directions. This can download content to each subgroup faster by enabling a word line because the entire subgroup is in an enabled state when using a single word line for downloading. The physical direction of the word line direction and the bit line and complementary bit line directions can vary, such that the word line direction and the bit line and complementary bit line directions can also have non-perpendicular directions. Compared with Figure 1 the system of [[[ID=]], this brings improved performance because the entire MAC operation does not have to stop when downloading new weights. In addition, the structure of the technology of the present invention allows the output of one subgroup to be processed in one adder tree operation. The present invention implements an architecture such that when one subgroup is downloading content (e.g., weights), other subgroups can continue to process (e.g., perform multiplication operations), which further improves performance. The technology of the present invention can also implement a structure that allows a single word line to be used to start / enable multiple subgroups of a group, such that a group can include (be divided into) two, four, eight, or even more subgroups using multiple timing signals. The timing signals can be based on any clock period or various clock signals, which do not need to be based on only one clock period. This structure will enable 1 / 2, 1 / 4, 1 / 8, or fewer multiplication operations to be performed per clock cycle, which allows reducing the number of adder tree inputs, further resulting in a reduction in the layout size of the adder tree.
[0142] In addition, an adder tree may require, for example, seven accumulation (addition) layers to go from 128 inputs to a single output, which could be 7 (e.g., one layer receives 128 inputs, the next layer receives 64 inputs, the next layer receives 32 inputs, the next layer receives 16 inputs, the next layer receives 8 inputs, the next layer receives 4 inputs, the next layer receives 2 inputs, and then provides the final single output). The time taken for these accumulation layers to complete fully may be longer than the time taken for a NOR gate to perform a multiplication operation. Thus, the adder tree may become a bottleneck for MAC operations. Therefore, the techniques of the present invention can implement a pipelined adder tree that is divided into several stages with buffers or latches between the stages. This allows the adder tree to store the temporary output data from the previous stage of the adder tree and act as a pipeline for continuously receiving inputs. As a result, each stage can operate within one clock cycle, which is consistent with the clock cycle required to complete a sub-group multiplication operation. This prevents the delay that may occur due to waiting for the adder tree to complete the accumulation operation for all layers before receiving new inputs. Thus, the overall clock cycle for the entire MAC operation can be reduced. In other words, since the adder tree is divided into several stages with buffers (latches) inserted between each stage, the operation of one stage does not affect the operation of another stage, enabling the adder tree to operate in a pipeline process of receiving new inputs in each clock cycle. As described above, the techniques of the present invention can also be implemented in a near-memory computing system.
[0143] The structure and operation for implementing the above features will be described below with reference to Figures 3-15 description.
[0144] Figure 3 A digital in-memory computing system incorporating tri-state NOR gates is described that allows sharing of an adder tree among various sub-groups performing MAC operations.
[0145] Specifically, Figure 3 a digital in-body computing system 300 is described that includes a plurality (N) of sub-group circuits, including sub-group_0 302, sub-group_1 304, and sub-group_N 306 (N is an integer greater than zero). An adder tree 308 (e.g., an accumulation circuit) is connected to the outputs of all sub-groups (e.g., sub-group_0 302, sub-group_1 304,..., sub-group_N 306) such that all sub-groups share the adder tree 308. In addition, sub-group_0 302 is connected to word line WL<0> 310, sub-group_1 304 is connected to word line WL<1> 312, and sub-group_N 306 is connected to word line WL <n>314. Each of subgroup_0 302, subgroup_1 304, and subgroup_N 306 (hereinafter referred to as subgroups 302, 304, and 306) is individually activated or enabled via its respective word lines 310, 312, and 314 to store content. A group can be referred to as a set of subgroups sharing the same adder tree 308, such as subgroups 302, 304, and 306.
[0146] As Figure 3 shown, each of subgroups 302, 304, and 306 may include storage circuits and / or multiplication (multiplier) circuits. Additionally, subgroups 302, 304, and 306 can be referred to as memory cell arrays such that each subgroup includes memory cells storing data elements, where the multiplier circuit is connected to the memory cell array and one or more input lines that provide M (or other quantity) input data elements, content, and / or multipliers on the one or more input lines. For example, subgroup 302 includes storage circuits 316, 317, and 318 (e.g., memory cell arrays). Each subgroup may include M (or other quantity) storage circuits storing M storage data elements. The storage circuits 316, 317, and 318 can be any type of memory, including but not limited to latches, sense amplifier (SA) latches, SRAM, DRAM, other types of volatile memory, and even NVM (SA latches can be used to sense data in a memory array, such as Figure 9 the memory array 902, and store the sensed data). This also applies to any storage circuit (memory cell array) described herein with respect to Figures 3-15 For example, the storage circuits 316, 317, and 318 can be configured as six transistors (6T). For example, by enabling the word line, the 6T SRAM memory cell is connected to true and complementary bit lines and has an additional output connected to the multiplier circuit to multiply the data stored in the memory cell by the input bit without enabling the word line. Other types of storage circuits can also be used to implement. Subgroups 304 and 306 include similar storage circuits. The bit line BL0 320 and complementary bit line BL0B 322 (e.g., common line) can be used in combination with the enabled word line WL<0> 310 to write content to the storage circuit 316 (e.g., content can be downloaded into it to program the computing system 300 within the digital memory), and the word line WL<0> 310 is connected to each storage circuit of subgroup 302. As Figure 3 As shown, bit line BL0 320 and complementary bit line BL0B 322 can be perpendicular to word line WL<0> 310. Other orientations of the word line, bit line, and complementary bit line can be implemented. Additionally, bit line BL0 320 and bit line BL0B 322 are connected to corresponding storage circuits of sub-groups 304 and 306 such that bit line BL0 320 and complementary bit line BL0B 322 are shared by the storage circuits of sub-groups 302, 304, and 306. As Figure 3 shown, there is essentially a column of storage circuits of sub-groups 302, 304, and 306 connected by bit line BL0 320 and complementary bit line BL0B 322.
[0147] Bit line BL1 324 and complementary bit line BL1B 326 are connected to storage circuit 317 to write content into storage circuit 317, and to write content into corresponding storage circuits of sub-groups 304 and 306 such that bit line BL1 324 and complementary bit line BL1B 326 (e.g., common line) are shared by the storage circuits of sub-groups 302, 304, and 306. Bit line BLm 328 and complementary bit line BLmB 330 (e.g., common line) are connected to storage circuit 318 to write content into storage circuit 318, and to write content into corresponding storage circuits of sub-groups 304 and 306 such that bit line BLm 328 and complementary bit line BLmB 330 are shared by the storage circuits of sub-groups 302, 304, and 306. Enabling various word lines 310, 312, and 314 controls which sub-groups have content written into them.
[0148] Semantically, the multiplication circuit can be referred to as being part of a sub-group, or can be referred to as being for a sub-group, but is not actually part of the sub-group. For example, sub-group 302 can include multiplication circuits such as tri-state NAND gates 332, 334, and 336 (also referred to as tri-state multipliers). As Figure 3 As shown in the subsequent figures, storage circuits 316, 317, and 318 can have read / write interfaces connected to bit lines and complementary bit lines, and can have separate read interfaces connected to tri-state NOR gates. Each subgroup includes m tri-state NOR gates. Tri-state NOR gate 332 is connected to storage circuit 316 such that tri-state NOR gate 332 can obtain the weight (content, input / storage data element) stored in storage circuit 316 and multiply the obtained weight by an input, such as input_0 received on input_0 line 338. Tri-state NOR gate 332 outputs the multiplication result only when enabled by a timing signal (e.g., the first timing signal Time_0). The timing control signal can include several different timing signals and / or control the transmission of different timing signals, such as the first timing signal Time_0, the second timing signal Time_1, and the Nth timing signal Time_N, or any other timing signal described herein. Additionally, the timing control signal can be connected to the multiplication circuit to enable them to sequentially provide the multiplication results from one or more subgroups of subgroups to the multiplier output. It can be said that the timing control signal selects a specific subgroup (e.g., it can be said that the first timing signal Time_0 selects subgroup 302, the second timing signal Time_1 selects subgroup 304, and the Nth timing signal Time_N selects subgroup 306).
[0149] Input_0 can be received from an input driver on an input line at approximately the same time as the timing signal (e.g., the first timing signal Time_0). As Figure 3 shown, input_0 can be received at IN_B of tri-state NOR gate 332, and the weight can be received at W_B of tri-state NOR gate 332. The multiplication output performed by tri-state NOR gate 332 is provided on output line out0 340 (e.g., the first output line), which is received by adder tree 308. Output line out0 340 is shared by the tri-state NOR gates of each subgroup 302, 304, and 306. As Figure 3 shown, the column of tri-state NOR gates including tri-state NOR gate 332 extends through subgroups 302, 304, and 306 and shares the same output line out0 340. However, at time time_0, i.e., the time when the timing signal enables tri-state NOR gate 332, the only output provided to output line out0 340 is from tri-state NOR gate 332 because the other tri-state NOR gates of other subgroups 304 and 306 are not enabled.
[0150] Similarly, the tri-state NOR gate 334 is connected to the storage circuit 317 such that the tri-state NOR gate 334 can obtain the weight (content) stored in the storage circuit 317 and multiply the obtained weight by the input, such as the input_1 received on the input line. The tri-state NOR gate 334 outputs (or performs) the multiplication only when enabled by a timing signal (e.g., the first timing signal Time_0). The input_1 can be received at approximately the same time as the timing signal (e.g., the first timing signal Time_0). As Figure 3 shown, the input_1 can be received at IN_B of the tri-state NOR gate 334, and the weight can be received at W_B of the tri-state NOR gate 334. The multiplication output performed by the tri-state NOR gate 334 is provided on the output line out1 342 (e.g., the second output line), which is received by the adder tree 308. The output line out1 342 is shared by the tri-state NOR gates of each subgroup 302, 304, and 306, as described above for the output line out0 340.
[0151] Similarly, the tri-state NOR gate 336 is connected to the storage circuit 318 such that the tri-state NOR gate 336 can obtain the weight (content) stored in the storage circuit 318 and multiply the obtained weight by the input, such as the input_m received on the input line. The tri-state NOR gate 336 outputs (or performs) the multiplication only when enabled by a timing signal (e.g., the first timing signal Time_0). The input_m can be received at approximately the same time as the timing signal (e.g., the first timing signal Time_0). As Figure 3 shown, the input_m can be received at IN_B of the tri-state NOR gate 336, and the weight can be received at W_B of the tri-state NOR gate 336. The multiplication output performed by the tri-state NOR gate 336 is provided on the output line outm 344, which is received by the adder tree 308. The output line outm 344 is shared by the tri-state NOR gates of each subgroup 302, 304, and 306, as described above for the output line out0 340. Each subgroup can have M (or other number) of output lines.
[0152] The tri-state NOR gates of subgroup 304 can be enabled by a timing signal (e.g., the second timing signal Time_1) and can receive the input_0, input_1 to input_m at approximately the same time. In the same way as subgroup 302, the circuit of subgroup 304 multiplies the weight by the input to provide an output to the adder tree 308. Additionally, the tri-state NOR gates of subgroup 306 can be enabled by a timing signal (e.g., the Nth timing signal Time_N) and can receive the input_0, input_1 to input_m at approximately the same time. In the same way as subgroup 302 and 304, the circuit of subgroup 306 multiplies the weight by the input to provide an output to the adder tree 308. As Figure 3 As shown, subgroup 302 requires 1 clock cycle to provide output out0 340 to output outm 344, then in the next clock cycle subgroup 304 provides output out0 340 to output outm 344, and then in the subsequent Nth clock cycle subgroup 306 provides output out0 340 to output outm 344. The multiplications performed by each of subgroups 302, 304, and 306 are controlled by different timing signals. When referring to timing signals herein, the timing signals may include multiple individual timing signals. Other timing signal schemes may be used to enable and control the tri-state or NOT gates described herein.
[0153] The outputs of each of subgroups 302, 304, and 306 are accumulated and provided as the MAC output 346, such that the output of subgroup 302 is provided as the MAC output 346, then the output of subgroup 304 is provided as the MAC output 346, and finally the output of subgroup 306 is provided as the MAC output 346.
[0154] Subgroup 302 may be referred to as a first subgroup circuit connected to a first word line (e.g., WL<0> 310), which is configured to (i) store a first set of weights, (ii) multiply the first set of weights by an input according to a first timing signal (e.g., at Time_0), and (iii) provide a first output (out0 340 to outm 344). Subgroup 304 may be referred to as a second subgroup circuit connected to a second word line (e.g., WL<1> 312), which is configured to (i) store a second set of weights, (ii) multiply the second set of weights by the input according to a second timing signal (e.g., at Time_1), and (iii) provide a second output (out0 340 to outm 344), wherein the multiplication of the second set of weights is enabled at a different time from the multiplication of the first set of weights. The adder tree 308 may be referred to as an accumulation circuit shared by the first subgroup (e.g., subgroup 302) and the second subgroup (e.g., subgroup 304), which is configured to receive and accumulate (i) the first output enabling the multiplication of the first set of weights according to the first timing signal (e.g., Time_0), and (ii) the second output enabling the multiplication of the second set of weights according to the second timing signal (e.g., Time_1).
[0155] In addition, the storage circuits 316, 317, and 318 of subgroup 302 may be referred to as the first storage circuits, and the three-state NAND gates 332, 334, and 336 may be referred to as the first multiplication circuits, where the first storage circuits are connected to the first word line and configured (writeable in sequence) to store the first set of weights, and the first multiplication circuits are enabled by the first timing signal to multiply the first set of weights by the input to provide the first output. In addition, the storage circuits of subgroup 304 may be referred to as the second storage circuits, and the three-state NAND gates of subgroup 304 may be referred to as the second multiplication circuits, where the second storage circuits are connected to the second word line and configured (writeable in sequence) to store the second set of weights, and the second multiplication circuits are enabled by the second timing signal to multiply the second set of weights by the input to provide the second output, and the second multiplication circuits are enabled at a different time from the time when the first multiplication circuits are enabled, such that the accumulation circuit (e.g., adder tree 308) receives and accumulates (i) the first output of the first multiplication circuits enabled according to the first timing signal, and (ii) the second output of the second multiplication circuits enabled according to the second timing signal. The accumulation circuit 308 may include accumulator inputs for N data elements connected to the multiplier outputs (e.g., the outputs of the multiplication circuits, such as three-state NAND gates). The accumulation circuit 308 may also generate the sum of the N data elements of the multiplier outputs.
[0156] Figure 3 The in-memory computing system 300 in a digital memory can receive data elements for storage in a memory cell array (e.g., storage circuits) from a non-volatile memory (NVM) array (such as NOR flash memory, NAND flash memory, etc.). This type of in-memory computing system 300 can be referred to as an NVM-type in-memory computing system. In addition, Figure 3 The in-memory computing system 300 in a digital memory can receive data elements for storage in a memory cell array from a volatile memory array (such as dynamic random access memory (DRAM), SRAM, etc.). These types of in-memory computing systems can be referred to as SRAM-type in-memory computing systems, DRAM-type in-memory computing systems, and / or volatile memory-type in-memory computing systems, etc. Regarding Figures 3-15 Any in-memory computing system described herein can be any of the above NVM-type in-memory computing systems and volatile memory-type in-memory computing systems.
[0157] Figure 4 Illustration Figure 3 An enlarged view of three different subgroup portions in the in-memory computing system.
[0158] Specifically, Figure 4 An enlarged view 400 illustrating the portions of subgroups 302, 304, and 306. As Figure 4 As shown, the word line WL<0> 310 is connected to the storage circuit 316 of the subgroup 302, the word line WL<1> 312 is connected to the storage circuit 402 of the subgroup 304, and the word line WL <n>314 is connected to the storage circuit 404 of subgroup 306. The tri-state NOR gate 332 of subgroup 302 receives and is enabled by the first timing signal (Time_0), receives the stored weight from the storage circuit 316, receives input_0 on the input line (input_0) 338 at or near the time of receiving the first timing signal (Time_0), multiplies the received weight by the input_0 on the input line (input_0) 338 to provide an output on the output line out0 340. As described above, the storage circuit 316 and the tri-state NOR gate 332 are just examples, and different types of storage circuits can be used for storage and gates, etc. to perform mathematical operations, and different technologies can be used for implementation.
[0159] Similarly, the tri-state NOR gate 406 of subgroup 304 receives and is enabled by the second timing signal (Time_1), receives the stored weight from the storage circuit 402, receives input_0 on the input line (input_0) 338 at or near the time of receiving the second timing signal (Time_1), multiplies the received weight by input_0 to provide an output on the output line out0 340. In addition, the tri-state NOR gate 408 of subgroup 306 receives and is enabled by the Nth timing signal (Time_N), receives the stored weight from the storage circuit 404, receives input_0 on the input line (input_0) 338 at or near the time of receiving the Nth timing signal (Time_N), multiplies the received weight by input_0 to provide an output on the output line out0 340.
[0160] Figure 5 Description Figure 3 A computing system within a digital memory as described, where content (data, such as N stored data elements) is written to subgroup_0 while one or more other subgroups are processing inputs and calculating outputs. Throughout this document, the content written to the storage circuit can be referred to as weights, stored data elements, etc. The content can be written to various subgroups during the operation of downloading a set of weights.
[0161] Specifically, Figure 5 The described computing system 500 within a digital memory is the same as the computing system 300 within a digital memory described with reference to Figure 3 and the repeated description is omitted here. Except for the content described with reference to Figure 3 the content described, Figure 5 It is also described that in operation 502, content is written into the storage circuit 316 of subgroup 302, in operation 504, content is written into the storage circuit 317 of subgroup 302, and in operation 506, content is written into the storage circuit 318 of subgroup 302. The written content (data) may include weights, such as a first set of weights. In this example, the storage circuits of subgroups 304 and 306 may have already stored weights (e.g., the storage circuit of subgroup 304 may have stored a second set of weights, and the storage circuit of subgroup 306 may have stored an Nth set of weights). When the storage circuits 316, 317, and 318 of subgroup 302 are writing content, the multiplication circuits (e.g., tri-state NAND gates) of subgroup 304 or subgroup 306 may perform the multiplication part of the MAC operation to provide an output to an adder tree (e.g., accumulator).
[0162] Figure 6 It is described in Figure 3 the adder tree pipeline used in a digital memory-based computing system embodiment.
[0163] Specifically, Figure 6 it describes an adder tree 600 (e.g., an accumulation circuit) that receives 128 outputs on the output lines connected to tri-state NAND gates in a digital memory-based computing system from Figure 3 . The outputs out0 602, out1 604 to out126 606, and out127 608 are received as inputs 601 to the adder tree 600. The adder tree 600 has seven layers, including a first adder layer 614 that adds 128 inputs to provide 64 results, a second adder layer 616 that adds 64 results to provide 32 results, a third adder layer 618 that adds 32 results to provide 16 results, a fourth adder layer 620 that adds 16 results to provide 8 results, a fifth adder layer 622 that adds 8 results to provide 4 results, a sixth adder layer 624 that adds 4 results to provide 2 results, and a seventh adder layer 626 that adds 2 results to provide a single output to path 628. The single output provided on path 628 is mathematically equivalent to equation 630, which is the result of the entire MAC operation (for a specific subgroup, such as Figure 3 subgroup 302).
[0164] As Figure 6 shown, the adder tree 600 includes buffer latches 610 and 612 for pipeline stages that temporarily store the intermediate and final results of the accumulation operation. For example, buffer latch 610 stores the 16 results provided by the third adder layer 618, and buffer latch 612 stores the 2 results provided by the sixth adder layer 624. Additionally, as Figure 6 As shown, it takes one clock cycle to receive 128 inputs and store 16 results in buffer latch 610, one clock cycle to retrieve the 16 stored results from buffer latch 610, perform an accumulation operation, and store 2 results in buffer latch 612, and 1 clock cycle to retrieve the 2 stored results from buffer latch 612, perform an accumulation operation, and provide a single output to path 628. The structure of adder tree 600 enables, for example, that at a certain clock cycle (time), 128 outputs can be received from subgroup 302 of Figure 3 and stored in buffer latch 610 while subgroup 304 is performing a multiplication operation. In the subsequent clock cycle (subsequent time), 128 outputs can be received from subgroup 304 and stored in buffer latch 610 while the previous content of buffer latch 610 (based on the operation of subgroup 302) is added and stored in buffer latch 612. In an even more subsequent clock cycle (even more subsequent time), 128 outputs can be received from subgroup 306 and stored in buffer latch 610 while the previous content of buffer latch 610 (based on the operation of subgroup 304) is added and stored in buffer latch 612, and the previous content of buffer latch 612 (based on the operation of subgroup 302) is added and provided as a single output to path 628. It can be said that adder tree 600 has three stages. The first stage (stage 0) includes receiving inputs and storing them in buffer latch 610, the second stage (stage 1) includes processing data from buffer latch 610 and storing it in buffer latch 612, and the third stage (stage 2) includes processing data from buffer latch 612 and providing a single output to path 628. Figure 3 This pipelined operation continues as the subgroups of the computing system within the digital memory continue to write content and continue to perform multiplication operations. This pipelined operation allows the MAC operation to continue without interruption because it eliminates the need for the computing system within the digital memory to wait for adder tree 600 to complete the accumulation and provide the result. Therefore, adder tree 600 is essentially capable of having a faster clock and higher throughput. For example, as described above, adder tree 600 actually requires three clock cycles to receive 128 inputs and provide a single output. Thus, without buffer latches 610 and 612, a multiplication operation that only requires one clock cycle would have to wait three clock cycles for adder tree to complete the accumulation operation. As a result, the computing system within the digital memory with this adder tree structure operates significantly faster.
[0165] When the subgroups of the computing system within the digital memory continue to write content and continue to perform multiplication operations, this pipelined operation continues. This pipelined operation allows the MAC operation to continue without interruption because it eliminates the need for the computing system within the digital memory to wait for adder tree 600 to complete the accumulation and provide the result. Therefore, adder tree 600 is essentially capable of having a faster clock and higher throughput. For example, as described above, adder tree 600 actually requires three clock cycles to receive 128 inputs and provide a single output. Thus, without buffer latches 610 and 612, a multiplication operation that only requires one clock cycle would have to wait three clock cycles for adder tree to complete the accumulation operation. As a result, the computing system within the digital memory with this adder tree structure operates significantly faster.
[0166] Alternatively, adder tree 600 (or any other adder tree described herein) may operate without buffer latches. Additionally, a buffer latch may be any type of circuit or element that can store data. Additionally, adder tree 600 (or any other adder tree described herein) may be a counter that performs a population count (i.e., counts the number of "1"s).
[0167] Figure 7 Describes a computing system within a digital memory where two tri-state NOR gates of two different subgroups share an output line and a word line.
[0168] Specifically, Figure 7 Describes a computing system 700 within a digital memory similar to that described with reference to Figure 3 which is omitted here for brevity. Figure 7 The computing system 700 within the digital memory of Figure 3 differs from the computing system 300 within the digital memory of <n>314。 Figure 7 The in-digital-memory computing system 700 includes the same word lines 310, 312, and 314, the same storage circuits 316, 317, and 318, the same bit lines and complementary bit lines 320, 322, 324, 326, 328, and 330, and the same tri-state NOR gates 332, 334, and 336 as discussed with reference to Figure 3 The in-digital-memory computing system 700 is also different because two tri-state NOR gates of two different subgroups 702 and 703 share an output line out0 340 and a word line 310. For example, the tri-state NOR gate 332 of subgroup 702 and the tri-state NOR gate 334 of subgroup 703 both provide outputs to the output line out0 340. Similarly, two tri-state NOR gates of two different subgroups 704 and 705 provide outputs to the same output line out0 340, and two tri-state NOR gates of two different subgroups 706 and 707 provide outputs to the same output line out 340.
[0169] Figure 7 In this example, if there are 128 tri-state NOR gates within subgroups 702 and 703, 64 tri-state NOR gates are members of subgroup 702 and 64 tri-state NOR gates are members of subgroup 703. Additionally, in this example, the tri-state NOR gate 332 is a member of subgroup 702, and the tri-state NOR gates 334 and 336 are members of subgroup 703. Also, 64 tri-state NOR gates are members of subgroup 704, and 64 tri-state NOR gates are members of subgroup 705. Similarly, 64 tri-state NOR gates are members of subgroup 706, and 64 tri-state NOR gates are members of subgroup 707.
[0170] As
[0171] shown, at time time_00 (e.g., the first clock cycle), the member tri-state NOR gates of subgroup 702 provide 64 outputs on output line out0 340 to output line outm / 2 708. In this example, m (or L) = 128, meaning that subgroups 702 and 703 combined have 128 tri-state NOR gates and there are 64 (128 / 2) outputs. Due to this subgroup architecture of the in-digital-memory computing system 700, there are 64 outputs per clock cycle, rather than as described above with reference to Figure 7 Figure 3 Each clock cycle of the digital memory-integrated computing system 300 discussed has 128 outputs. At time time_01 (e.g., the second clock cycle), the tri-state NOR gates of subgroup 703 are enabled for multiplication and provide 64 outputs. Thus, subgroups 702 and 703 are utilized for 2 clock cycles, which can provide additional clock cycles for writing content to other groups while subgroups 702 and 703 are performing multiplication, and also allows the use of a smaller adder tree (accumulator), as discussed below with reference to Figure 8 Content can be written to multiple subgroups connected to the same word line simultaneously.
[0172] Returning to Figure 7 , at time time_10 (e.g., the third clock cycle), the member tri-state NOR gates of subgroup 704 provide 64 outputs on output lines out0 340 to outm / 2 708. At time time_11 (e.g., the fourth clock cycle), the tri-state NOR gates of subgroup 705 are enabled for multiplication and provide 64 outputs. At time time_N0 (e.g., the (N - 1)th clock cycle), the member tri-state NOR gates of subgroup 706 provide 64 outputs on output lines out0 340 to outm / 2 708. At time time_N1 (e.g., the Nth clock cycle), the tri-state NOR gates of subgroup 707 are enabled for multiplication and provide 64 outputs. Figure 7 The architecture illustrated in N can be modified such that each group includes 8 or more subgroups. For example, the number of "adjacent" NOR gates of subgroups 702 and 703 sharing the same output line can be 2, 4, 8, or any number that is a power of 2, where N is an integer.
[0173] Figure 8 Illustrated Figure 7 is the adder tree of the digital memory-integrated computing system.
[0174] As described above, Figure 7 the digital memory-integrated computing system 700 provides, for example, 64 outputs per clock cycle, rather than Figure 3 the 128 outputs per clock cycle of the digital computing storage system 300. Figure 8 The adder tree 800 has 64 inputs to receive Figure 7 the 64 outputs of the digital memory-integrated computing system 700.
[0175] Specifically, the adder tree 800 receives output out0 802, outputs out1 804 to out62 806, and output out63 808. These outputs are received as the 64 inputs 801 of the adder tree 800. This adder tree 800 is smaller than Figure 6 adder tree 600 has one less layer. For example, adder tree 800 includes (i) a first adder layer 814 that receives 64 inputs and provides 32 results, (ii) a second adder layer 816 that receives 32 results and provides 16 results, (iii) a third adder layer 818 that receives 16 results and provides 8 results, (iv) a buffer latch 810 that receives and temporarily stores 8 results, (v) a fourth adder layer 820 that receives the 8 temporarily stored results and provides 4 results, (vi) a fifth adder layer 822 that receives 4 results and provides 2 results, and (vii) a sixth adder layer 824 that receives 2 results and provides a single result. Similar to Figure 6 adder tree 600, adder tree 800 provides a single result on path 628, where the single result is mathematically equivalent to equation 630, which is the result of the entire MAC operation for a particular subgroup.
[0176] As Figure 8 shown, adder tree 800 completes the accumulation operation in two clock cycles, where one clock cycle (phase 0) is used to receive the inputs and store the results in buffer latch 810, and the other clock cycle (phase 1) is used to retrieve the stored data from buffer latch 810 and provide a single output. This is the same pipeline structure as discussed above for adder tree 600, except that there is only one buffer latch instead of two, and only two clock cycles are required to complete the accumulation operation instead of three clock cycles. The reduction in the number of buffer latches and the required clock cycles is because adder tree 800 receives only 64 inputs instead of 128 inputs for adder tree 600. As a result, the number of gates (i.e., the gates that perform the accumulation) in adder tree 800 is approximately half that of adder tree 600, making adder tree 800 approximately half the size of adder tree 600. This adder tree 800 with fewer inputs can be implemented as a result of the subgroup structure and shared output lines, as discussed above with reference to Figure 7 discussed.
[0177] Figure 9 illustrates a digital computing storage system in which four NOR gates of four different subgroups share an output line and one.
[0178] Figure 9 digital in-memory computing system 900 is similar to Figure 7 digital in-memory computing system 700, and the repeated description of the components described with reference to Figure 7 is omitted here.
[0179] Figure 9 digital in-memory system 900 is similar to Figure 7 The in-digital-memory system 700 is different in that there are four different groups of three-state transistor tubes sharing an output line and a word line. Specifically, Figure 9 It is described that four subgroups (i.e., subgroup_0 903, subgroup_1 904, subgroup_2 905, and subgroup_3 906, hereinafter referred to as subgroups 903, 904, 905, and 906) are connected to the same word line WL<0> 310, and it is further described that four subgroups (i.e., subgroup_4 907, subgroup_5 908, subgroup_6 909, and subgroup_7 910, hereinafter referred to as subgroups 907, 908, 909, and 910). An additional set of four subgroups, up to subgroups N-3, N-2, N-1, and N, can be connected to their respective word lines. Figure 9 The in-digital-memory computing system 900 includes the Figure 3 and Figure 7 same word lines 310 and 312, the same storage circuits, and the same three-state NOR gates as discussed.
[0180] Subgroup 903 includes a storage circuit and a three-state NOR gate related to time time_0, subgroup 904 includes a storage circuit and a three-state NOR gate related to time time_1, subgroup 905 includes a storage circuit and a three-state NOR gate related to time time_2, subgroup 906 includes a storage circuit and a three-state NOR gate related to time time_3, subgroup 907 includes a storage circuit and a three-state NOR gate related to time time_4, subgroup 908 includes a storage circuit and a three-state NOR gate related to time time_5, subgroup 909 includes a storage circuit and a three-state NOR gate related to time time_6, and subgroup 910 includes a storage circuit and a three-state NOR gate related to time time_7.
[0181] Figure 9 The bit-line configuration of Figure 7 is different. Specifically, bit lines BL0 912 and BL1 914 and Vref 916 (e.g., a reference voltage line) are implemented to write content to certain storage circuits (e.g., the storage circuits in the leftmost column) of subgroups 903, 904, 907, and 908. As Figure 9 As shown, the reference voltage line Vref 916 is connected to the storage circuits above and below the word line WL<0> 310 and the storage circuits above and below the word line WL<1> 312, while the bit line BL1 914 is connected to the storage circuits above the word line WL<0> 310 and above the word line WL<1> 312, and the bit line BL0 912 is connected to the storage circuits below the word line WL<0> 312 and below the word line WL<1> 312. In addition, the bit lines BL2 920 and BL3 922 and the reference voltage line Vref918 are implemented to write content to certain storage circuits of the subgroups 905, 906, 909, and 910 (e.g., the storage circuits in the second column to the right of the first column). As Figure 9 As shown, the reference voltage line Vref 918 is connected to the storage circuits above and below the word line WL<0> 310 and the storage circuits above and below the word line WL<1> 312, while the bit line BL3 922 is connected to the storage circuits above the word line WL<0> 310 and above the word line WL<1> 312, and the bit line BL2 920 is connected to the storage circuits below the word line WL<0> 312 and below the word line WL<1> 312. The bit lines BL510 926 and BL511 928 and the reference voltage line Vref 924 are implemented to write content to certain storage circuits of the subgroups 905, 906, 909, and 910 (e.g., the storage circuits in the rightmost column). As Figure 9 As shown, the reference voltage line Vref 924 is connected to the storage circuits above and below the word line WL<0> 310 and the storage circuits above and below the word line WL<1> 312, while the bit line BL511 928 is connected to the storage circuits above the word line WL<0> 310 and above the word line WL<1> 312, and the bit line BL510 926 is connected to the storage circuits below the word line WL<0> 312 and below the word line WL<1> 312.
[0182] Figure 9 The in-memory computing system 900 in the digital memory is also different from the in-memory computing system 700 because the four three-state NOR gates in the subgroups 903, 904, 905, and 906 and the four three-state NOR gates in the subgroups 907, 908, 909, and 910 share an output line out0 936. For example, the three-state NOR gates related to the time time_0~time3 in the subgroups 903, 904, 905, and 905 can all provide outputs to the output line out0 904. Similarly, the three-state NOR gates related to the time ~time7 in the subgroups 907, 908, 909, and 910 provide outputs to the same output line out0 936. The input lines 930, 932, and 934 provide inputs to various three-state NOR gates.
[0183] In this example, if there are 512 three-state NOR gates in subgroups 903, 904, 905, and 906, 128 three-state NOR gates are members of subgroup 903, 128 three-state NOR gates are members of subgroup 904, 128 three-state NOR gates are members of subgroup 905, and 128 three-state NOR gates are members of subgroup 906. Additionally, 128 three-state NOR gates are members of subgroup 907, 128 three-state NOR gates are members of subgroup 908, 128 three-state NOR gates are members of subgroup 909, and 128 three-state NOR gates are members of subgroup 910. Other configurations are possible, depending on the number of outputs required per clock cycle. In Figure 9 's example, there are 128 outputs. If there are 64 outputs, the number of three-state NOR gates in each subgroup will be reduced to 64.
[0184] As Figure 9 shown, at time_0 (e.g., the first clock cycle), the three-state NOR gates that are members of subgroup 903 provide 128 outputs on output lines out0 936 to out127 938. At time_1 (e.g., the second clock cycle), the three-state NOR gates of subgroup 904 are enabled to perform multiplication and provide 128 outputs. At time_2 (e.g., the third clock cycle), the three-state NOR gates of subgroup 905 are enabled to perform multiplication and provide 128 outputs. At time_3 (e.g., the fourth clock cycle), the three-state NOR gates of subgroup 906 are enabled to perform multiplication and provide 128 outputs. Thus, subgroups 903, 904, 905, and 906 are utilized for four clock cycles, which allows additional clock cycles for writing content to other subgroups connected to others.
[0185] At time_4 (e.g., the fifth clock cycle), the three-state NOR gates that are members of subgroup 907 provide 128 outputs on output lines out0 936 to out127 938. At time_5 (e.g., the sixth clock cycle), the three-state NOR gates that are members of subgroup 908 provide 128 outputs on output lines out0 936 to out127 938. At time_6 (e.g., the seventh clock cycle), the three-state NOR gates that are members of subgroup 909 provide 128 outputs on output lines out0 936 to out127 938. At time_7 (e.g., the eighth clock cycle), the three-state NOR gates that are members of subgroup 910 provide 128 outputs on output lines out0 936 to out127 938.
[0186] Figure 9 Further describe that the content of the memory array 902 is written into the storage circuits of subgroups 903, 904, 905, 906, 907, 908, 909, and 910. As previously described, the memory array 902 that stores the content (data element) can be any type of non-volatile memory and volatile memory. The content (data element) of the memory array 902 can be written into the storage circuits of subgroups 903, 904, 905, and 906 when subgroups 907, 908, 909, and 910 perform multiplication operations, and the content can be written into the storage circuits of subgroups 907, 908, 909, and 910 when subgroups 903, 904, 905, and 906 perform multiplication operations. Any subgroup that does not share a word line with the active subgroup can be written. As Figure 9 shown, two storage circuits on a single word line (e.g., WL<0> 310) can share an output line (e.g., out0 936), can be connected to different bit lines (e.g., BL0 912 and BL1 914), and can share a reference line (e.g., Vref 916), which allows different data to be stored in the two storage circuits.
[0187] Figure 10 Describe Figure 9 the operation time sequence of subgroups_0 903 to subgroups_7 910 in Figure 10 Identify subgroup_0 903 as "subgroup 0", subgroup_1 904 as "subgroup 1", and so on.
[0188] Specifically, Figure 10 describe the graph 1000 that includes a time sequence that loops from time time_0 to time time_7 and then returns to time time_0 to start. The clock signal 1002 is identified at the top of the time sequence. Below the corresponding time sequence of the clock signal 1002, the graph 1000 describes the subgroup sequence 1004, which indicates which subgroup is performing and / or outputting a mathematical operation. Below the subgroup sequence 1004, the graph 1000 describes the adder tree pipeline 1006, which indicates which stage of the adder tree is in the startup state. Since Figure 9 describe 128 outputs, an adder tree 600 can be implemented Figure 6 to have 3 stages. The first stage (stage 0) includes receiving an input and storing the result in the buffer latch 610. The second stage (stage 1) includes processing the data from the buffer latch 610 and storing the result in the buffer latch 612. The third stage (stage 2) includes processing the data from the buffer latch 612 and providing a single output to the path 628.
[0189] As Figure 10 As shown, the adder tree pipeline 1006 indicates which stages of the adder tree 600 are processing the input / data. Additionally, the diagram illustrates the download program 1008 for the downloaded content, which includes writing the content to the storage circuits of various subgroups. For example, from time time_0 to time time_3, the content is downloaded and written to the storage circuits of subgroups 4, 5, 6, and 7, while subgroups 0, 1, 2, and 3 are performing multiplication and / or output operations. Similarly, from time time_4 to time time_7, the content is downloaded and written to the storage circuits of subgroups 0, 1, 2, and 3, while subgroups 4, 5, 6, and 7 are performing multiplication and / or output operations. Then again from time time_0 to time time_3, new content is downloaded and written to the storage circuits of subgroups 4, 5, 6, and 7, while subgroups 0, 1, 2, and 3 are performing multiplication and / or output operations using the content written during time time_4 to time time_7. The download and write times may take more than one clock cycle. Therefore, Figure 9 One advantage of the in-digital-memory computing system 900 is that there are four clock cycles available for writing content to a specific set of subgroups (in this example, a set of subgroups includes four subgroups). This structure reduces or eliminates the time spent waiting for the content to be written before starting the multiplication and / or output operations.
[0190] Figure 11 Describes a digital computing storage system in which bit lines and reference voltage lines (Vref) are used to write content.
[0191] Specifically, Figure 11 Describes an in-digital-memory computing system 1100 similar to that described with reference to Figure 5 here, the repeated description is omitted. The in-digital-memory computing system 1100 is different from the in-digital-memory computing system 500 of Figure 5 in that the complementary bit lines are replaced by reference voltage lines (Vref). Specifically, Figure 11 Describes that the reference voltage lines (Vref) 1102, 1104, and 1106 are connected to the storage circuits for writing. The advantage of replacing the complementary bit lines with reference voltage lines (Vref) is that the number of SA-latches (sense amplifiers and latches) can be doubled, thus doubling the sensing speed. For example, if there are 1024 memory bit lines in the memory array, the 1024 memory bit lines are divided into 512 SRAM bit lines and 512 SRAM complementary bit lines, so that only 512 bits of data can be sensed at a time. If 1024 memory bit lines are connected to 1024 SRAM bit lines and 1024 complementary bit lines are connected to the reference voltage line (Vref), then 1024 bits of data can be sensed at a time.
[0192] Figure 12 A digital memory computing system for writing content using sense amplifiers is described.
[0193] Specifically, Figure 12 a digital memory in - system computing system 1200 similar to that described with reference to Figure 5 is described, and the repeated description is omitted here. Figure 12 A memory array 1202 for storing content is described, in which a read operation 1210 is performed to read the content from the memory array 1202 into sense amplifiers SA0 1204, SA1 1206, and SAm 1208 through respective bit lines and complementary bit lines BL0, BL0B, BL1, BL1B, BLm, and BLmB. As previously mentioned, the memory array 1202 for storing content (data elements) can be any type of non - volatile memory and volatile memory. Figure 12 The digital memory in - system computing system 1200 Figure 5 differs from the digital memory in - system computing system 500 in that it includes sense amplifiers that provide content for various subgroups. As Figure 12 shown, a write operation 1212 is performed to write the content from the sense amplifiers into the storage circuits of various subgroups. Specifically, sense amplifier SA0 1206 provides content on the SA DL0 line (sense amplifier download line) 1214 and the SA DL0B line (complementary sense amplifier download line) 1216 to write the content into certain storage circuits, sense amplifier SA1 1206 provides content on the SA DL1 line 1218 and the SA DL1B line 1220 to write the content into certain storage circuits, and sense amplifier SAm 1208 provides content on the SA DLm line 1222 and the SA DLmB line 1224 to write the content into certain storage circuits. Alternatively, a reference voltage line Vref can be used instead of using complementary lines (i.e., SA DL0B line 1216, SA DL1B line 1220, and SA DLmB line 1224) to write the content. Using sense amplifiers instead of directly programming the storage circuits from the bit lines and complementary bit lines allows for faster programming because the sense amplifiers can provide higher voltage signals than the bit lines and complementary bit lines can provide.
[0194] Figure 13 A timing diagram for performing MAC operations using the Figure 5 digital memory in - system computing system 500 is described.
[0195] Specifically, the timing diagram 1300 shows that the inputs can remain static until they are multiplied by all the weights. For example, Figure 13 a startup subgroup 1302, a time series 1304 from Time_0 to Time_N are shown, and Figure 5 Input_0 338 in the input sequence 1306 of the in - digital - memory computing system 500. This timing diagram can be modified to accommodate other in - digital - memory computing systems described herein.
[0196] As Figure 13 shown, with respect to Figure 5 , N = 64, so there are 64 sub - groups. The start sub - group 1302 illustrates the sub - groups in sequence according to the time series. The input_0 338 on the input sequence 1306 remains the same (Xi0) and is multiplied by each of the weights Wi0, Wi1, Wi2, Wi3, Wi4, Wi5, Wi6, and Wi7 from time time_0 to time time_7. Then at time time_8, the input changes to Xi1 and is multiplied by each of the weights Wi0, Wi1, Wi2, Wi3, Wi4, Wi5, Wi6, and Wi7 from time time_8 to time time_15. This process continues until the input Xi7 is multiplied by the weight Wi7 at time time_63. As described herein, when a sub - group connected to a particular input line performs multiplication and / or output operations, the weights of other sub - groups connected to other input lines can be updated. Figure 13 This sequence of operations illustrated in
[0197] Figure 14 can be referred to as "input - stationary and sequential - serial - input in product operations". Figure 3 illustrates an in - digital - memory system that uses transmission gates instead of the
[0198] tri - state NOR gates shown in Figure 14 Specifically, Figure 4 illustrates structure 1400, which is similar to Figure 4 and thus the repeated description is omitted. However, structure 1400 replaces the tri - state NOR gates 332, 406, and 408 in Figure 4 Similarly, each subgroup can have N transmission gates. At time time_0 (when receiving timing signals Time_0 and Time_0B), transmission gate 1402 allows the content stored in storage circuit 316 to be passed to NOR gate 1408, and NOR gate 1408 multiplies the content stored in storage circuit 316 with input 1410 provided to it at time time_0. There can be N NOR gates in structure 1400, where each of the N NOR gates corresponds to a specific column of the transmission gates. NOR gate 1408 provides output out0 1412 to an adder tree (accumulator). At time time_1 (when receiving timing signals Time_1 and Time_1B), transmission gate 1404 allows the content stored in storage circuit 402 to be passed to NOR gate 1408, and NOR gate 1408 multiplies the content stored in storage circuit 402 with input 1410 provided to it at time time_1. At time time_N (when receiving timing signals Time_N and Time_NB), transmission gate 1406 allows the content stored in storage circuit 404 to be passed to NOR gate 1408, and NOR gate 1408 multiplies the content stored in storage circuit 404 with input 1410 provided to it at time time_N. This implementation eliminates the need for tri-state NOR gates and the space they occupy.
[0199] Figure 15 Simplified block diagram illustrating an integrated circuit device that includes a memory array configured for in-memory computing with signed (or unsigned) inputs and weights.
[0200] Specifically, Figure 15 Illustrating an integrated circuit device 1500 that includes a memory array 1560 arranged for signed in-memory computing for CIM operations such as signed (or unsigned) multiply-accumulate (MAC) operations performed by the digital in-memory computing system described herein. The integrated circuit device 1500 can be implemented on a single chip or on a multi-chip module.
[0201] Device 1500 includes input / output circuit 1505 for communicating control signals, data, addresses, and commands with other data processing resources such as a CPU or a memory controller.
[0202] Input / output data is applied to bus 1591 and then to controller 1510 and cache 1590. Additionally, addresses are applied to bus 1593 and then to decoder 1542 and controller 1510. Meanwhile, bus 1591 and bus 1593 are operably connected to data sources inside the integrated circuit device 1500, such as a general-purpose processor or a special-purpose application circuit, or a combination of modules providing system-on-chip functionality.
[0203] The memory array 1560 may include a memory array in a NOR architecture or an AND architecture, such that memory cells are arranged in rows along bit lines and in columns along word lines, and the memory cells in a given row are connected in parallel between the bit line and a source reference. The source reference may include a ground terminal or a source line connected to a source-side bias resource. The memory cells may include charge-trapping transistor memory cells arranged in a 3D structure. The memory array 1560 with in-memory (or near-memory) computing capabilities may be configured and operate in the manner described in Figures 3-14 as described.
[0204] The word lines may be connected through a block selection circuit to global bit lines 1565 and configured to selectively connect to a page buffer 1580 and a CIM sensing circuit 1570.
[0205] In the illustrated embodiment, the page buffer 1580 is connected to a cache 1590 through a bus 1585. The page buffer 1580 includes storage elements (which may be various types of memory arrays) for storing operations and sensing circuits, and such storage operations include read and write operations. For flash memories including dielectric charge-trapping memories and floating-gate charge-trapping memories, the write operations include programming and erasing operations.
[0206] The drive circuit 1540 is coupled to word lines 1545 in the memory array 1560 and applies a word line voltage to selected word lines according to the decoding of an address on a bus 1593 by a decoder 1542, or in a computing operation, according to input data stored in an input buffer 1541.
[0207] The controller 1510 is coupled to the cache 1590 and the memory array 1560, as well as other peripheral circuits for memory access and in-memory computing operations.
[0208] The controller 1510 controls, for example using a state machine, the application of supply voltages and currents generated or provided through blocks 1520 via a voltage supply or a current source, which are for storage operations and CIM operations.
[0209] The controller 1510 includes control and status registers, and control logic that may be implemented using special-purpose logic circuits known to those skilled in the art (including state machines and combinational logic). In alternative embodiments, the control logic includes a general-purpose processor that may be implemented on the same integrated circuit, and the processor executes a computer program to control the operation of the device. In other embodiments, the control logic may be implemented using a combination of special-purpose logic circuits and general-purpose processors.
[0210] The memory array 1560 includes memory cells arranged in rows and columns, where the memory cells in a row are connected to corresponding bit lines and the memory cells in a column are connected to corresponding word lines. The memory array 1560 is programmable to store signed coefficients (weights Wi) in groups of memory cells.
[0211] In the CIM mode, the word line driver circuit 1540 or the driver circuit 1540 may include a driver (referred to as an input driver or input activation driver) configured to drive a signed or unsigned input Xi from the input buffer 1541. The driver circuit 1540 may be separate from the word line driver circuit. The CIM sensing circuit 1570 is configured to sense the difference between the first and second currents on each of the selected bit line pairs and generate an output of the selected bit line pair that is a function of the difference. These outputs may be applied to the storage elements in the page buffer 1580 and the cache 1590.
[0212] In one embodiment, the first subgroup may include a first storage circuit and a first multiplication circuit, where the first storage circuit is connected to a first word line and is programmable to store a first set of weights, and the first multiplication circuit is enabled by a first timing signal to multiply the first set of weights by an input to provide a first output, where the second subgroup includes a second storage circuit and a second multiplication circuit, where the second storage circuit is connected to a second word line and is programmable to store a second set of weights, and the second multiplication circuit is enabled by a second timing signal to multiply the second set of weights by the input to provide a second output, the second multiplication circuit is enabled at a different time than the first multiplication circuit, and where the accumulation circuit receives and accumulates (i) the first output depending on the first multiplication circuit being enabled by the first timing signal and (ii) the second output depending on the second multiplication circuit being enabled by the second timing signal.
[0213] In one embodiment, the in-memory computing circuit may include a first output line and a second output line, where the first output line is shared by an output of one of the multiplication circuits in the first multiplication circuit of the first subgroup and an output of one of the multiplication circuits in the second multiplication circuit of the second subgroup, where the second output line is shared by an output of the other multiplication circuit in the first multiplication circuit of the first subgroup and an output of the other multiplication circuit in the second multiplication circuit of the second subgroup, where the first output of the first subgroup depends on the first multiplication circuit being enabled while the second multiplication circuit is not enabled and is provided to the accumulation circuit through the first and second output lines, and where the second output of the second subgroup depends on the second multiplication circuit being enabled while the first multiplication circuit is not enabled and is provided to the accumulation circuit through the first and second output lines.
[0214] In another embodiment, a particular storage circuit in the first storage circuit of the first subgroup and a particular storage circuit in the second storage circuit of the second subgroup may share a common programming line for controlling the storage of their respective weights, wherein the particular storage circuit of the first subgroup is programmed to store a particular weight in the first set of weights depending on the first subgroup being activated by the first word line, and wherein the particular storage circuit of the second subgroup is programmed to store a particular weight in the second set of weights depending on the second subgroup being activated by the second word line.
[0215] In one embodiment, an in-memory computing circuit may include a multiplication circuit configured to receive and multiply inputs and provide an output, a first subgroup circuit connected to a first word line and configured to (i) store a first set of weights and (ii) enable a first transmission gate depending on a first timing signal to provide the first set of weights to the multiplication circuit, a second subgroup circuit connected to a second word line and configured to (i) store a second set of weights and (ii) enable a second transmission gate depending on a second timing signal to provide the second set of weights to the multiplication circuit, the enabling time of the second set of weights being different from the enabling time of the first set of weights, and an accumulation circuit shared by the first subgroup and the second subgroup and configured to receive and accumulate (i) a first output received from the multiplication circuit depending on the first transmission gate being enabled by the first timing signal and (ii) a second output received from the multiplication circuit depending on the second transmission gate being enabled by the second timing signal.
[0216] The implementation of the memory array may be based on charge-trapping memory cells, such as floating memory cells that may include a polysilicon charge-trapping layer, or dielectric charge-trapping memory cells that may include a silicon nitride charge-trapping layer. Other types of memory technologies may be applied to various embodiments of the techniques described herein.
[0217] Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Another implementation of the methods described in this section may include a system that includes a memory and one or more processors operable to execute instructions stored in the memory to perform any of the methods described above.
[0218] According to many embodiments, any of the data structures and code described or referenced above are stored on a computer-readable storage medium, which can be any device or medium that can store code and / or data for use by a computer system. This includes, but is not limited to, volatile memory, non-volatile memory, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), magnetic and optical storage devices such as disk drives, tapes, CDs (compact discs), DVDs (digital versatile discs or digital video discs), or other media now known or later developed that are capable of storing computer-readable media.
[0219] An example of a processor is a hardware unit (e.g., a hardware circuit that includes one or more active devices, etc.) capable of executing program code. The processor may optionally include one or more controllers and / or state machines. The processor can be implemented according to application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), and / or customized design techniques. The processor can be manufactured according to integrated circuit, optical, and quantum technologies. The processor uses one or more architectural techniques, such as sequential (e.g., von Neumann) processing, very long instruction word (VLIW) processing. The processor uses one or more microarchitectural techniques, such as executing one instruction at a time or executing instructions in parallel, e.g., through one or more pipelines. The processor is for general-purpose (and / or) special-purpose (such as signal, audio, video, and / or graphics purposes). The processor is fixed-function or variable-function, e.g., according to programming. The processor includes any one or more registers, memories, logic units, arithmetic units, and graphics units. The term processor is intended to include a single processor as well as multiple processors, such as multiprocessors and / or processor clusters.
[0220] The logic described herein can be implemented in the following ways: using a processor to execute a computer program stored in a computer system accessible memory, using dedicated logic hardware (including field programmable integrated circuits), and combinations of dedicated logic hardware and computer programs. For all flowcharts herein, it should be understood that many steps can be combined, executed in parallel, or executed in a different order without affecting the functions implemented. In some cases, the reader will understand that a rearrangement of steps can achieve the same result only when certain other changes are made. In other cases, the reader will understand that a rearrangement of steps can achieve the same result only when certain conditions are met. In addition, it should be understood that the flowcharts herein only show the steps relevant to understanding the present invention, and it should be understood that many additional steps can be executed before, after, and between the shown steps to complete other functions.
[0221] Although the present invention is disclosed by reference to the above preferred embodiments and examples, it should be understood that these examples are illustrative rather than restrictive. Modifications and combinations that can be easily conceived by those skilled in the art will be within the spirit of the present invention and the scope of the appended claims.< / n> < / n> < / n> < / n>
Claims
1. A memory computing circuit, comprising: One or more input lines receiving M input data elements, where M is an integer greater than zero; A memory cell array comprising one or more subgroups, each of the one or more subgroups storing M storage data elements; a plurality of multiplication circuits connected to the memory cell array and the one or more input lines and configured to multiply the M input data elements with the M stored data elements in a selected subgroup of the plurality of subgroups and provide a multiplier output having M data elements; as well as an accumulation circuit including an accumulator input of the M data elements, the M data elements being connected to the multiplier output and the accumulation circuit being configured to produce a sum of the M data elements of the multiplier output, Wherein the plurality of multiplication circuits provide multiplication results from the one or more subgroups to the multiplier output.
2. The in-memory computing circuit of claim 1 , wherein for each subgroup of the one or more subgroups, the plurality of multiplication circuits comprises M tri-state multipliers connected to the multiplier output. 3 . The in-memory computing circuit according to claim 2 , wherein the M three-state multipliers are M three-state NOR gates.
4. The in-memory computing circuit according to claim 2, wherein the one or more subgroups include a first subgroup storing M stored data elements and a second subgroup storing M stored data elements, wherein the M tri-state multipliers of the first subgroup are enabled by a first timing signal to multiply the M input data elements with the M stored data elements of the first subgroup, and The M three-state multipliers of the second subgroup are enabled by a second timing signal to multiply the M input data elements with the M stored data elements of the second subgroup, so that the M three-state multipliers of the second subgroup are enabled at a different time from the M three-state multipliers of the first subgroup, and the second timing signal and the first timing signal are provided at different times.
5. The in-memory computing circuit according to claim 2, wherein the one or more subgroups include a first subgroup storing M stored data elements in M storage circuits and a second subgroup storing M stored data elements in M storage circuits, wherein the first subgroup is connected to a first word line, and The second subgroup is connected to a second word line.
6. The in-memory computing circuit according to claim 5, wherein a specific storage circuit of the M storage circuits of the first subgroup and a specific storage circuit of the M storage circuits of the second subgroup share a common line for controlling storage of respective data elements, wherein the specific storage circuits of the first subgroup rely on the first word line to enable the first subgroup to store specific data elements, and The specific storage circuits of the second subgroup rely on the second word line to activate the second subgroup to store specific data elements.
7. The in-memory computing circuit of claim 6, wherein the common line shared by the particular storage circuits of the first subgroup and the particular storage circuits of the second subgroup comprises a bit line.
8. The in-memory computing circuit of claim 5 , wherein a specific storage circuit of the M storage circuits of the second subgroup writes a specific data element in at least one of the following circumstances: (i) the M three-state multipliers of the first subgroup are enabled by a first timing signal to multiply the M input data elements with the M stored data elements of the first subgroup to provide a multiplier output having M data elements, and (ii) the accumulation circuit receives and accumulates the multiplier output having M data elements.
9. The in-memory computing circuit according to claim 5, wherein the multiplier output comprises a first output line and a second output line, wherein the first output line is shared by an output of one tri-state multiplier of the first subgroup and an output of one tri-state multiplier of the second subgroup, wherein the second output line is shared by an output of another three-state multiplier of the first subgroup and an output of another three-state multiplier of the second subgroup, wherein the output associated with the first subgroup is provided to the accumulation circuit via the first and second output lines depending on the M tri-state multipliers of the first subgroup being enabled by a timing control signal while the M tri-state multipliers of the second subgroup are not enabled by the timing control signal, and The output associated with the second subgroup depends on the M tri-state multipliers of the second subgroup being enabled by the timing control signal while the M tri-state multipliers of the first subgroup are not enabled by the timing control signal, and is provided to the accumulation circuit through the first and second output lines.
10. The in-memory computing circuit according to claim 2, wherein the one or more subgroups include a first subgroup storing M stored data elements, a second subgroup storing M stored data elements, a third subgroup storing M stored data elements, and a fourth subgroup storing M stored data elements, wherein the first subgroup and the second subgroup are connected to a first output line, wherein the third subgroup and the fourth subgroup are connected to a second output line, wherein the M tri-state multipliers of the first subgroup are enabled by a first timing signal to multiply the M input data elements with the M stored data elements of the first subgroup, wherein the M tri-state multipliers of the second subgroup are enabled by a second timing signal to multiply the M input data elements by the M stored data elements of the second subgroup, wherein the M tri-state multipliers of the third subgroup are enabled by a third timing signal to multiply the M input data elements with the M stored data elements of the third subgroup, wherein the M tri-state multipliers of the fourth subgroup are enabled by a fourth timing signal to multiply the M input data elements with the M stored data elements of the fourth subgroup, and Wherein L is an integer representing the total number of the M stored data elements of the first subgroup and the M stored data elements of the second subgroup, and wherein M = L / 2.
11. The in-memory computing circuit according to claim 10, wherein the multiplier output comprises a first output line, and The first output line is shared by an output of a tri-state multiplier of the first subgroup, an output of a tri-state multiplier of the second subgroup, an output of a tri-state multiplier of the third subgroup, and an output of a tri-state multiplier of the fourth subgroup.
12. The in-memory computing circuit according to claim 11, wherein the multiplier output comprises a second output line, and The second output line is shared by the output of another three-state multiplier of the first subgroup, the output of another three-state multiplier of the second subgroup, the output of another three-state multiplier of the third subgroup and the output of another three-state multiplier of the fourth subgroup.
13. The in-memory computing circuit of claim 1, wherein the M stored data elements of each respective subgroup of the one or more subgroups are written to each respective subgroup using a word line.
14. The in-memory computing circuit of claim 1, wherein the M storage elements of each respective subgroup of the one or more subgroups are written to each respective subgroup using a sense amplifier connected to a word line.
15. The in-memory computing circuit according to claim 1, wherein the one or more subgroups include a first subgroup storing M stored data elements and a second subgroup storing M stored data elements, in, during a first clock cycle, the first subgroup multiplies the M stored data elements by the M input data elements and the second subgroup has the M stored elements written thereto, and Wherein, during a second clock cycle, the second subgroup multiplies the M stored data elements with the M input data elements and the first subgroup has the M stored elements written thereto.
16. The in-memory computing circuit according to claim 1, wherein the one or more subgroups include a first subgroup storing M stored data elements and a second subgroup storing M stored data elements, in, During a specific clock cycle, the accumulation circuit accumulates outputs associated with the first subgroup, and During successive clock cycles, the accumulation circuit accumulates outputs associated with the second subgroup.
17. The in-memory computation circuit of claim 16, wherein the accumulation circuit is pipelined.
18. The in-memory computing circuit of claim 1, wherein for each subgroup of the one or more subgroups, the multiplication circuit comprises M pass gates connected to a shared M-bit multiplier connected to the multiplier output.
19. The in-memory computing circuit according to claim 1, wherein the multiplication circuit is enabled by a timing control signal to provide the multiplication result. 20 . The in-memory computing circuit according to claim 19 , wherein the timing control signal comprises a first timing signal and a second timing signal, such that the first timing signal and the second timing signal are provided at different times.
21. The in-memory computing circuit of claim 7, wherein the common line further comprises a complementary bit line (BLB).
22. The in-memory computing circuit of claim 7, wherein the common line further comprises a reference voltage line (Vref).
23. The in-memory computing circuit of claim 1, wherein the memory cell array comprises a plurality of latches.
24. A method of performing an operation using an in-memory computation circuit, the in-memory computation circuit comprising (i) a memory cell array including one or more subgroups, each of the one or more subgroups storing M stored data elements, M being an integer greater than zero, (ii) a multiplication circuit coupled to the memory cell array and one or more input lines, and (iii) an accumulation circuit including an accumulator input for the M data elements, the accumulator input coupled to a multiplier output, the method comprising: Obtain M input data elements from the one or more input lines; multiplying, by the multiplication circuit, the M input data elements with M stored data elements in a selected subgroup of the one or more subgroups to provide a multiplier output having M data elements, wherein the multiplication circuit is enabled by a timing control signal to provide multiplication results from the one or more subgroups to the multiplier output in sequence; as well as The sum of the M data elements output by the multiplier is produced by the accumulation circuit.
25. An in-memory computing circuit comprising: A first subgroup circuit connected to the first word line and configured to store a first set of weights; A second subgroup circuit connected to the second word line and configured to store a second set of weights; A multiplication circuit, configured to (i) multiply the first set of weights by an input according to a first timing signal, (ii) provide a first output, (iii) multiply the second set of weights by the input according to a second timing signal, and (iv) provide a second output, wherein the multiplication of the second set of weights is enabled at a different time than the multiplication of the first set of weights is enabled; as well as An accumulation circuit is shared by the first subgroup and the second subgroup and is configured to receive and accumulate (i) the first output enabled by the first timing signal according to the multiplication of the first set of weights, and (ii) the second output enabled by the second timing signal according to the multiplication of the second set of weights.