Design for efficient near-memory-computing and digital computing-in-memory
The in-memory computing circuit with shared adder trees and pipelined architecture addresses the layout and performance issues of dCIM systems by reducing adder tree count and enabling simultaneous multiplication and download operations, thereby improving efficiency.
Patent Information
- Application Number
- JP2024187853
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-01
- Filing Date
- 2024-10-25
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2044-10-25
AI Technical Summary
Conventional digital computing-in-memory (dCIM) systems face challenges with large layout area due to numerous adder trees and slow performance due to long download times for memory contents, which necessitate stopping sub-groups of multiplication operations.
The proposed solution involves an in-memory computing circuit with shared adder trees among sub-groups, utilizing tri-state NOR gates and pipelined adder trees, enabling simultaneous multiplication and download operations across sub-groups, reducing layout area and improving performance.
This approach reduces the number of adder trees, allows faster download of memory contents, and enhances overall system performance by enabling continuous processing without interruption during weight updates.
Smart Images

Figure 2025105454000001_ABST
Abstract
Description
Technical Field
[0001] Field of the Invention The present invention relates to a circuit that can be used to perform in-memory or near-memory calculations such as multiply-and-accumulate (MAC) operations or other summation-of-products operations.
Background Art
[0002] Description of Related Art In circuits used for certain calculations based on neuromorphic (neuro-morphological) computing systems, machine learning systems, and linear algebra, a multiply-and-sum function or a sum-of-products function can be an important component. Such functions can be expressed as follows:
Number
[0003] In this equation, each product term is the product of the variable input X i and the weight W i The weight W i can vary between terms and, for example, corresponds to the coefficient of the variable input X i .
[0004] The sum-of-products function can be realized as the operation of a circuit using a cross-point array architecture, in which the electrical characteristics of the cells of the array give rise to this function.
[0005] These architectures can be implemented within a digital computing-in-memory (dCIM) system and as a digital near-memory-computing (dNMC) system to perform the multiply-accumulate (MAC) operations described in the above equations. Conventionally, in these systems, one sub-group of products is realized by a corresponding adder tree (e.g., an accumulator). As a result, there are a large number of adder trees because each sub-group (of a larger group) has its own corresponding adder tree. Adder trees have a relatively large layout area and are thus expensive in terms of space. Conventionally, in these systems, adder trees occupy an undesired amount of space. For simplicity, hereinafter, the dCIM system includes the dNMC system.
[0006] In addition, in a dCIM system, as a result of the need to toggle some or all of the word lines (WL), the time required to download memory contents (e.g., weights) is undesirably long, which results in slower operation. Specifically, in a dCIM system, during the download of contents (e.g., weights), some sub-groups of multiplication operations must be stopped, which degrades performance. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0007] Therefore, it is desirable to provide a dCIM system that has a reduced number of adder trees and can perform MAC operations and the like while also downloading new contents (e.g., weights). MEANS FOR SOLVING THE PROBLEM
[0008] In one preferred example, an in-memory computing circuit is provided. This in-memory computing circuit can include one or more input lines, an array of memory cells, a multiplication circuit connected to the array of memory cells and the one or more input lines, and an accumulation circuit. The one or more input lines receive M input data elements, where M is an integer greater than 0. The array of memory cells includes one or more sub-groups, and each sub-group in the one or more sub-groups stores M input data elements. The multiplication circuit is configured to multiply the M input data elements by M stored data elements within a selected sub-group of the one or more sub-groups and to provide a multiplier output having M data elements. The accumulation circuit includes an accumulator input of M data elements connected to the multiplier output and is configured to generate a sum of the M data elements of the multiplier output. The multiplication circuit provides (sequentially) the multiplication result from the one or more sub-groups to the multiplier output.
[0009] In a further preferred example, the multiplication circuit can include M tri-state multipliers connected to the multiplier output for each sub-group in the one or more sub-groups.
[0010] In another preferred example, the M tri-state multipliers can be M tri-state NOR (negative OR) gates.
[0011] In one preferred example, one or more subgroups can include a first subgroup that stores M memory data elements and a second subgroup that stores M memory data elements. The M tri-state multipliers for the first subgroup are enabled by a first timing signal to multiply the M input data elements by the M memory data elements of the first subgroup. The M tri-state multipliers for the second subgroup are enabled by a second timing signal to multiply the M input data elements by the M memory data elements of the second subgroup. The second timing signal is supplied at a different time from the first timing signal so that the M tri-state multipliers for the second subgroup are enabled at a time different from the M tri-state multipliers for the first subgroup.
[0012] In a further preferred example, one or more subgroups can include a first subgroup that stores M memory data elements in M memory circuits and a second subgroup that stores M memory data elements in M memory circuits. The first subgroup is connected to a first word line, and the second subgroup is connected to a second word line.
[0013] In another preferred example, a specific memory circuit among the M memory circuits of the first subgroup and a specific memory circuit among the M memory circuits of the second subgroup share a common line to control the storage of respective data elements. The specific memory circuit of the first subgroup stores a specific data element in response to the first word line activating the first subgroup, and the specific memory circuit of the second subgroup stores a specific data element in response to the second word line activating the second subgroup.
[0014] In one preferred example, the common line shared by the specific memory circuit of the first subgroup and the specific memory circuit of the second subgroup can include a bit line (BL).
[0015] In a further preferred example, during at least one of (i) when M tri-state multipliers for a first subgroup are enabled by a first timing signal to multiply M input data elements by M stored data elements of the first subgroup to provide a multiplier output having M data elements, and (ii) when an accumulation circuit receives and accumulates the multiplier output having M data elements, a specific data element can be written to a specific memory circuit among the M memory circuits of the second subgroup.
[0016] In another preferred example, the multiplier output can include a first output line and a second output line. The first output line is shared by the output of one tri-state multiplier for the first subgroup and the output of one tri-state multiplier for the second subgroup. The second output line is shared by the output of another tri-state multiplier for the first subgroup and the output of another tri-state multiplier for the second subgroup. In response to the M tri-state multipliers for the first subgroup being enabled by a timing control signal and the M tri-state multipliers for the second subgroup not being enabled by the timing control signal, the output related to the first subgroup is provided to the accumulation circuit via the first and second output lines. In response to the M tri-state multipliers for the second subgroup being enabled by the timing control signal and the M tri-state multipliers for the first subgroup not being enabled by the timing control signal, the output related to the second subgroup is provided to the accumulation circuit via the first and second output lines.
[0017] In one preferred example, one or more subgroups can include a first subgroup that stores M memory data elements, a second subgroup that stores M memory data elements, a third subgroup that stores M memory data elements, and a fourth subgroup that stores M memory data elements. The first subgroup and the second subgroup are connected to a first word line, the third subgroup and the fourth subgroup are connected to a second word line. M tri-state multipliers for the first subgroup are enabled by a first timing signal to multiply M input data elements by the M memory data elements of the first subgroup. M tri-state multipliers for the second subgroup are enabled by a second timing signal to multiply M input data elements by the M memory data elements of the second subgroup. M tri-state multipliers for the third subgroup are enabled by a third timing signal to multiply M input data elements by the M memory data elements of the third subgroup. M tri-state multipliers for the fourth subgroup are enabled by a fourth timing signal to multiply M input data elements by the M memory data elements of the fourth subgroup. L is an integer that can represent the total number of the M memory data elements of the first subgroup and the M memory data elements of the second subgroup, where M = L / 2.
[0018] In another preferred example, the multiplier output can include a first output line, and the first output line is shared by the output of one tri-state multiplier for the first subgroup, the output of one tri-state multiplier for the second subgroup, the output of one tri-state multiplier for the third subgroup, and the output of one tri-state multiplier for the fourth subgroup.
[0019] In one preferred example, the multiplier output can include a second output line, and the second output line is shared by the output of another one tri-state multiplier for the first subgroup, the output of another one tri-state multiplier for the second subgroup, the output of another one tri-state multiplier for the third subgroup, and the output of another one tri-state multiplier for the fourth subgroup.
[0020] In a further preferred example, the M memory data elements of each subgroup in one or more subgroups can be written to each subgroup using bit lines.
[0021] In another preferred example, the M memory data elements of each subgroup in one or more subgroups can be written to each subgroup using a sense amplifier connected to the bit lines.
[0022] In one preferred example, one or more subgroups can include a first subgroup storing M memory data elements and a second subgroup storing M memory data elements. During the first clock cycle, the first subgroup multiplies M input data elements by M memory data elements, and M memory data elements are written to the second subgroup. During the second clock cycle, the second subgroup multiplies M input data elements by M memory data elements, and M memory data elements are written to the first subgroup.
[0023] In a further preferred example, one or more subgroups can include a first subgroup storing M memory data elements and a second subgroup storing M memory data elements. During a specific clock cycle, an accumulation circuit accumulates the output related to the first subgroup, and during a subsequent clock cycle, the accumulation circuit accumulates the output related to the second subgroup.
[0024] In another preferred example, the accumulation circuit can be pipelined.
[0025] In one preferred example, the multiplication circuit can include M pass gates connected to a shared M-bit multiplier connected to the multiplier output for each subgroup in one or more subgroups.
[0026] In another preferred example, the multiplication circuit can be enabled by a timing control signal to provide a multiplication result. In one preferred example, the timing control signal can include a first timing signal and a second timing signal, and the first timing signal is supplied at a time different from the second timing signal.
[0027] In a further preferred example, a method of performing an operation is provided. This method can be executed using an in-memory computing circuit, which includes (i) an array of memory cells, (ii) a multiplication circuit connected to the array of memory cells and one or more input lines, and (iii) an accumulation circuit. The array of memory cells includes one or more subgroups, each subgroup stores M stored data elements, M is an integer greater than 0, and the accumulation circuit includes accumulator inputs for M data elements connected to the multiplier output. Further, this method can include the steps of obtaining M input data elements from one or more input lines, multiplying, by the multiplication circuit, the M input data elements by the M stored data elements of a selected subgroup among one or more subgroups to provide a multiplier output having M data elements, wherein the multiplication circuit is enabled by a timing control signal to sequentially provide multiplication results from subgroups within one or more subgroups to the multiplier output, and generating, by the accumulation circuit, a sum of the M data elements of the multiplier output.
[0028] In another preferred example, an in-memory computing circuit is provided. This in-memory computing circuit can include a first subgroup of circuits, a second subgroup of circuits, a multiplication circuit, and an accumulation circuit. The first subgroup of circuits is connected to a first word line and is configured to store a first set of weights. The second subgroup of circuits is connected to a second word line and is configured to store a second set of weights. The multiplication circuit is configured to (i) multiply an input by the first set of weights in response to a first timing signal, (ii) provide a first output, (iii) multiply the input by the second set of weights in response to a second timing signal, and (iv) provide a second output. The multiplication of the second set of weights is enabled at a time different from the time when the multiplication of the first set of weights is enabled. The accumulation circuit is shared by the first subgroup and the second subgroup and is configured to (i) receive and accumulate the first output in response to the multiplication of the first set of weights being enabled by the first timing signal, and (ii) receive and accumulate the second output in response to the multiplication of the second set of weights being enabled by the second timing signal.
[0029] In one preferred example, the common line can include a bit line bar line (BLB: negative logic bit line).
[0030] In a further preferred example, the common line can further include a reference voltage line (VREF).
[0031] In another preferred example, an array of memory cells can include latches.
[0032] Other aspects and advantages of the present invention can be understood by considering the following drawings, detailed description, and claims.
Brief Description of the Drawings
[0033]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Mode for Carrying Out the Invention
[0034] Detailed Description A detailed description of embodiments of the present invention will be provided with reference to FIGS. 1 to 15.
[0035] FIG. 1 shows a conventional SRAM (static random access memory)-based dCIM system, which includes a plurality of subgroups of multiplier units, each subgroup having its corresponding adder tree (accumulator). The dCIM system of FIG. 1 is considered to be SRAM-based because the weights are stored in and read from the SRAM.
[0036] Specifically, FIG. 1 shows a conventional SRAM-based dCIM system 100, which includes a plurality of subgroups of one group, such as subgroup 102, each subgroup utilizing a corresponding adder tree. For example, subgroup 102 receives inputs and word line signals from an input activation (activation, enabling) driver and an SRAM WL driver 106, executes mathematical operations, and executes an accumulation operation using adder tree 104. Each subgroup includes a storage circuit for storing weights and a circuit for executing mathematical operations such as multiplying an input by a weight.
[0037] In this example, weights are stored using a memory device such as a six-transistor (6T) SRAM cell, and an input (IN<0:255>) is multiplied by the weights using a multiplier such as a four-transistor (4T) NOR gate. Further, as shown, an input activation driver and an SRAM WL driver receive the input IN<0:255> on the line and receive word line signals on word lines WL<0:255>. Each of the inputs IN<0:255> is received by subgroup 102 (as well as other subgroups), whereby the stored weights can be multiplied by these inputs. The word lines WL<0:255> are used to access the memory device for reading and writing (e.g., to write weights to the memory device and / or to read weights from the memory device). The output of the multiplication operation of subgroup 102 is provided as input 4b to adder tree 104. Adder tree 104 can combine various inputs to provide a single output in operation 5b. In this example, one subgroup having 4 columns and 256 rows of cells can implement Ini×Wi<0:3>[i = 0~255] for combination with an adder tree of 1024-bit (4 bits per row) output and 1024 input bits to complete the MAC operation. This multiplication operation can be performed by a NOR gate or other types of circuits capable of performing multiplication (or other types of mathematics) operations.
[0038] As shown, this conventional SRAM-based dCIM system 100 requires a separate adder tree for each subgroup. Here, in this example, there are 64 subgroups, and thus 64 adder trees, and these adder trees occupy a larger amount of physical space within the conventional SRAM-based dCIM system 100. Specifically, one of the problems 1 of the conventional SRAM-based dCIM system 100 is that the adder trees cannot be shared among other subgroups. Therefore, as the count of adder trees increases based on the number of subgroups, the overall layout size of the SRAM-based dCIM 100 also increases.
[0039] FIG. 2 shows an example of a subgroup of the conventional SRAM-based dCIM system of FIG. 1.
[0040] Specifically, FIG. 2 shows a subgroup 200 that receives word line inputs 204 (WL<0:255>), inputs 205 (IN_B<0:255>), bit line inputs 206 (BL<0:3>), and bit line bar inputs 208 (BLB<0:3>). The bit line inputs 206 (BL<0:3>) and bit line bar inputs 208 (BLB<0:3>) are for programming a storage device 202 in which weights are stored. Further, subgroup 200 includes NOR gates for multiplying the stored weights by input 205. The result of the multiplication is provided to an adder tree 210. As described above with reference to FIG. 1, each subgroup 200 communicates with a corresponding adder tree 210. FIG. 2 shows only a single subgroup and one adder tree, but as shown in FIG. 1, SRAM-based dCIM requires at least 64 adder trees.
[0041] To update the storage device 202 of the subgroup to store a new set of weights (e.g., weight values), all word lines WL<0:255> are sequentially enabled to update all the contents of the SRAM. Updating the storage device 202 with a new set of weights is expensive due to the time required to store the new values. Further, updating the storage device 202 affects the data received by the adder tree 210, so the entire MAC operation of all subgroups must be stopped while the storage device 202 is being updated. This further slows down the performance of the SRAM-based dCIM.
[0042] The disclosed technology addresses these drawbacks by providing an SRAM-based dCIM with a reduced layout area and improved performance.
[0043] Specifically, compared to the systems of FIGS. 1 and 2, the disclosed technique uses a multiplier with a tri-state output by, for example, replacing a common NOR gate with a tri-state NOR gate, and each subgroup of tri-state NOR gates can be controlled by one individual timing signal (e.g., a clock signal, etc.). Alternatively, the tri-state NOR gate can be replaced with any kind of logic gate that can perform a product operation between an input and stored data (latched data bits) and / or can be controlled / enabled by a timing signal. The tri-state output (of the tri-state NOR gate) has a high impedance state when not made active. This allows a number of tri-state NOR gate multipliers to be connected to each input of the accumulator, where only the tri-state NOR gate multiplier selected by an individual timing signal and made active operates on the input. This structure allows sharing by multiple subgroups of one adder tree, resulting in a reduced layout area.
[0044] Furthermore, the disclosed technology can arrange each of the subgroups along the WL direction orthogonal to the BL and BLB directions, which enables faster download of content into each subgroup by enabling one WL, because the entire subgroup can be made active for download using a single WL. The physical orientation of the WL direction and the BL and BLB directions can be changed, so that the WL direction and the BL and BLB directions can also have a non-orthogonal orientation. The result is improved performance compared to the system of FIG. 1, because it is not necessary to stop the MAC operation when downloading new weights. Furthermore, the structure of the disclosed technology enables processing the output of one subgroup in the operation of one adder tree. The disclosed technology realizes an architecture in which when one subgroup is downloading content (e.g., weights), other subgroups can continue processing (e.g., performing multiplication operations), which further improves performance. The disclosed technology can further be realized in a structure that makes a plurality of subgroups of one group active / enabled using a single WL, so that one group can include two, four, eight, or even more subgroups using a plurality of timing signals (dividing one group into such subgroups). These timing signals can be based on any clock cycle or various clock signals and do not have to be based on only one clock cycle. This structure enables 1 / 2, 1 / 4, 1 / 8, or fewer multiplication operations per clock cycle, which enables reducing the number of adder tree inputs, which further results in a reduction in the layout size of the adder tree.
[0045] Furthermore, an adder tree may require, for example, seven accumulation (addition) levels to reach a single output from, for example, 128 inputs (e.g., one level receives 128 inputs, the next level receives 64 inputs, the next level receives 32 inputs, the next level receives 16 inputs, the next level receives 8 inputs, the next level receives 4 inputs, and the next level receives 2 inputs to provide a final single output). The time required to complete this number of accumulation levels can be longer than the time required for a NOR gate to complete a multiplication operation. Therefore, the adder tree can be a bottleneck for MAC operations. Thus, the disclosed technique can implement a pipelined adder tree separated into several stages, having buffers or latches between the stages. This allows the adder tree to store the primary output data from the previous-stage adder tree and function as a pipeline for continuously received inputs. As a result, each stage can operate in one clock cycle, which matches the clock cycle required to complete a multiplication operation for one subgroup. This prevents the delay that could occur by waiting for the accumulation operation to complete for all levels before the adder tree receives a new input. As a result, the overall clock cycle of the entire MAC operation can be reduced. In other words, since the adder tree is divided into several stages and buffers (latches) are inserted between each stage, the operation of one stage does not affect other stages, thereby enabling the adder tree to operate in a pipeline flow that receives new inputs every clock cycle. As described above, the technique disclosed herein can also be implemented in the form of a near-memory computing system.
[0046] The structure and operation enabling these features described above will be described below with reference to FIGS. 3 to 15.
[0047] FIG. 3 shows a dCIM system including tri-state NOR gates, which enable sharing of an adder tree among various subgroups performing MAC operations.
[0048] Specifically, FIG. 3 shows a dCIM system including a plurality (N) of sub-groups of circuits, the sub-groups including sub-group 0 302, sub-group 1 304, and sub-group N 306 (N is an integer greater than 0). An adder tree 308 (e.g., an accumulation circuit) is connected to the outputs of all the sub-groups (e.g., sub-group 0 302, sub-group 1 304, ..., sub-group N 306), whereby all the sub-groups share the adder tree 308. Further, sub-group 0 302 is connected to word line WL<0> 310, sub-group 1 304 is connected to word line WL<1> 312, and sub-group N 306 is connected to word line WL <n>It is connected to 314. Each of subgroup 0 302, subgroup 1 304, and subgroup N 306 (hereinafter referred to as subgroups 302, 304, and 306) is individually activated or enabled by its respective word lines 310, 312, and 314 to store content. One group can be referred to as a set of subgroups such as subgroups 302, 304, and 306 that share the same adder tree 308.
[0049] As shown, each of subgroups 302, 304, and 306 can include a memory circuit and / or a multiplication (multiplier) circuit. Further, subgroups 302, 304, and 306 can be referred to as an array of memory cells, whereby each subgroup includes memory cells that store data elements, and the multiplication circuit is connected to the array of memory cells and one or more input lines, and these input lines provide M (or some other number) of input data elements on one or more input lines, on one or more contents, and / or on one or more multipliers. For example, subgroup 302 includes memory circuits 316, 317, and 318 (e.g., an array of memory cells). Each subgroup can include M (or some other number) of stored data elements. Memory circuits 316, 317, and 318 can be any type of memory, and these memories can be latches, sense amplifier (SA) latches, SRAMs, DRAMs (dynamic random access memories), other types of volatile memories, and even NVMs (non-volatile memories), but are not limited thereto (SA latches can be used to detect data in a memory array, such as memory array 902 in FIG. 9, and store the detected data). This applies to any of the memory circuits (arrays of memory cells) described in this specification with respect to FIGS. 3 to 15. For example, memory circuits 316, 317, and 318 can be composed of 6 transistors (6T). For example, a 6T SRAM connected to complementary bit lines that are made active by activating a word line has an additional output connected to a multiplier, whereby an input bit can be multiplied by the data stored in the cell without activating the word line. Other types of memory circuits can also be realized. Subgroups 304 and 306 include similar memory circuits.Bit line BL0 320 and bit line bar line BL0B 322 (e.g., a common line) can be used in combination with the activation of word line WL<0> 310 to write content into memory circuit 316 (e.g., content can be downloaded into memory circuit 316 for programming of dCIM system 300), and word line WL<0> 310 can be connected to each of the memory circuits in subgroup 302. As shown, BL0 320 and BL0B 322 can be orthogonal to word line WL<0> 310. Other orientations of word lines, bit lines, and bit line bar lines can be realized. In addition, BL0 320 and BL0B 322 are connected to the corresponding memory circuits in subgroups 304 and 306, whereby BL0 320 and BL0B 322 are shared by the memory circuits in subgroups 302, 304, and 306. As shown, there is basically a column of memory circuits in subgroups 302, 304, and 306 connected by BL0 320 and BL0B 322.
[0050] To write content into memory circuit 317 and to write content into the corresponding memory circuits in subgroups 304 and 306, BL1 342 and BL1B 326 are connected to memory circuit 317, whereby BL1 324 and BL1B 326 (e.g., a common line) are shared by the memory circuits in subgroups 302, 304, and 306. To write content into memory circuit 318 and to write content into the corresponding memory circuits in subgroups 304 and 306, BLm 328 and BLmB 330 (e.g., a common line) are connected to memory circuit 318, whereby BLm 328 and BLmB 330 are shared by the memory circuits in subgroups 302, 304, and 306. Activations of various word lines 310, 312, and 314 control which subgroup the content is written into.
[0051] Meaningfully, the multiplication circuit can be referred to as part of a subgroup, or can be referred to as for a subgroup, but in fact it is not part of that subgroup. For example, subgroup 302 can include multiplication circuits such as tristate NOR gates 332, 334, and 336 (also referred to as tristate multipliers). As shown in FIG. 3 and subsequent figures, memory circuits 316, 317, and 318 can have read ports and write ports connected to bit lines and bit line bar lines, and can have separate read ports connected to tristate NOR gates. Each subgroup includes m tristate NOR gates. Tristate NOR gate 322 is connected to memory circuit 316, whereby tristate NOR gate 322 can acquire the weight (content, input / memory data element) stored in memory circuit 316 and multiply the acquired weight by an input such as input 0 received on input 0 line 338. Tristate NOR gate 332 outputs a multiplication value only when enabled by a timing signal (e.g., first timing signal time 0). The timing control signal can include several different timing signals such as the first timing signal time 0, the second timing signal time 1, and the Nth timing signal time N, or other timing signals described in this specification, and / or can control the transmission of different timing signals. Further, the timing control signal can be connected to the multiplication circuits to enable these multiplication circuits and sequentially provide the multiplication results from the subgroups among one or more subgroups to the multiplier output. The timing control signal can be said to select a particular subgroup (e.g., the first timing signal time 0 can be said to select subgroup 302, the second timing signal time 1 can be said to select subgroup 304, and the Nth timing signal time N can be said to select subgroup 306).
[0052] Input 0 can be received at the same time or approximately the same time as a timing signal (e.g., the first timing signal time 0) on the input line from the input driver. As shown, input 0 can be received at input B of tri-state NOR gate 332, and the weight can be received at weight B of tri-state NOR gate 332. The output of the multiplication performed by NOR gate 332 is provided on the output line "output 0 340" (e.g., the first output line) and is received by adder tree 308. Output line "output 0 340" is shared by the tri-state NOR gates of each of subgroups 302, 304, and 306. As shown, the column of tri-state NOR gates including tri-state NOR gate 332 that extends through subgroups 302, 304, and 306 shares the same output line "output 0 340". However, at time 0, i.e., the time when the timing signal time 0 enables tri-state NOR gate 332, the only output provided on output line "output 0 340" is from tri-state NOR gate 332, because the other tri-state NOR gates of the other subgroups 304 and 306 are not enabled.
[0053] Similarly, the tri-state NOR gate 334 is connected to the memory circuit 317, whereby the tri-state NOR gate 334 can obtain the weight (content) stored in the memory circuit 317 and multiply the obtained weight by an input such as input 1 received on the input line. The tri-state NOR gate 334 can output (or perform) the multiplication value only when enabled by a timing signal (e.g., the first timing signal time 0). Input 1 can be received at the same time or approximately at the same time as the timing signal (e.g., the first timing signal time 0). As shown in the figure, input 1 can be received at input B of the tri-state NOR gate 334, and the weight can be received at weight B of the tri-state NOR gate 334. The output of the multiplication performed by the tri-state NOR gate 334 is provided on the output line "output 1 342" (e.g., the second output line) and received by the adder tree 308. The output line "output 1 342" is shared by each of the tri-state NOR gates of the subgroups 302, 304, and 306, just as described above for the output line "output 0 340".
[0054] Similarly, the tri-state NOR gate 336 is connected to the memory circuit 318, whereby the tri-state NOR gate 336 can obtain the weights (contents) stored in the memory circuit 318 and multiply the obtained weights by an input such as the input m received on the input line. The tri-state NOR gate 336 can output (or perform) the multiplication value only when enabled by a timing signal (e.g., the first timing signal time 0). The input m can be received at the same time or approximately at the same time as the timing signal (e.g., the first timing signal time 0). As shown in the figure, the input m can be received at input B of the tri-state NOR gate 336, and the weights can be received at weight B of the tri-state NOR gate 336. The output of the multiplication performed by the tri-state NOR gate 336 is provided on the output line "output m 344" and received by the adder tree 308. Just as described above for the output line "output 0 340", the output line "output m 344" is shared by the tri-state NOR gates of each of the subgroups 302, 304, and 306. There can be M (or some other number) of output lines for each subgroup.
[0055] The tri-state NOR gates of subgroup 304 can be enabled by a timing signal (e.g., timing signal time 1), and can receive input 0, input 1 to input m at the same time or approximately the same time. In the same way as subgroup 302, the circuit of subgroup 304 can multiply the inputs by weights and provide the outputs to adder tree 308. Further, the tri-state NOR gates of subgroup 306 can be enabled by a timing signal (e.g., the Nth timing signal time N), and can receive input 0, input 1 to input m at the same time or approximately the same time. In the same way as subgroups 302 and 304, the circuit of subgroup 306 can multiply the inputs by weights and provide the outputs to adder tree 308. As shown, subgroup 302 requires 1 clock cycle to provide outputs "output 0 340" to "output m 344". Thus, in the Nth subsequent clock cycle, subgroup 306 provides outputs "output 0 340" to "output m 344". The multiplications performed by each of subgroups 302, 304, and 306 are controlled by different timing signals. When referring to a timing signal in this specification, one timing signal may include a plurality of separate timing signals. Other timing signal schemes can be used to enable and control the tri-state NOR gates described in this specification.
[0056] The outputs of each of subgroups 302, 304, and 306 are accumulated and provided as the MAC output 346, whereby the output of subgroup 302 is provided as the MAC output 346, then the output of subgroup 304 is provided as the MAC output 346, and finally, the output of subgroup 306 is provided as the MAC output 346.
[0057] Sub-group 302 can be referred to as the first sub-group of the circuit, is connected to the first word line (e.g., word line WL<0> 310), and is configured to (i) store the first set of weights, (ii) multiply an input by the first set of weights in response to the first timing signal (e.g., time 0), and (iii) provide a first output (e.g., outputs 0 340 to outputs m 344). Sub-group 304 can be referred to as the second group of the circuit, is connected to the second word line (e.g., WL<1> 312), and is configured to (i) store the second set of weights, (ii) multiply an input by the second set of weights in response to the second timing signal (e.g., time 1), and (iii) provide a second output (e.g., outputs 0 340 to outputs m 344), and the multiplication of the second set of weights is enabled at a time different from the time when the multiplication of the first set of weights is enabled. The adder tree 308 can be referred to as an accumulation circuit, is shared by the first sub-group (e.g., sub-group 302) and the second sub-group (e.g., sub-group 304), and is configured to receive and accumulate (i) a first output corresponding to the multiplication of the first set of weights being enabled by the first timing signal (e.g., time 0), and (ii) a second output corresponding to the multiplication of the second set of weights being enabled by the second timing signal (e.g., time 1).
[0058] Furthermore, the memory circuits 316, 317, and 318 of subgroup 302 can be referred to as a first memory circuit, the tri-state NOR gates 332, 334, and 336 can be referred to as a first multiplication circuit, the first memory circuit is connected to a first word line and is configured to store a first set of weights (and can be written to store the first set of weights), the first multiplication circuit is enabled by a first timing signal to multiply an input by the first set of weights and provide a first output. In addition, the memory circuits of subgroup 304 can be referred to as a second memory circuit, the tri-state NOR gates of subgroup 304 can be referred to as a second multiplication circuit, the second memory circuit is connected to a second word line and is configured to store a second set of weights (and can be written to store the second set of weights), the second multiplication circuit is enabled by a second timing signal to multiply an input by the second set of weights and provide a second output, and the second multiplication circuit is enabled at a time different from the time when the first multiplication circuit is enabled, whereby an accumulation circuit (e.g., adder tree 308) is configured to receive and accumulate (i) a first output in response to the first multiplication circuit being enabled by the first timing signal, and (ii) a second output in response to the second multiplication circuit being enabled by the second timing signal. The accumulation circuit 308 can include accumulator inputs for N data elements connected to the multiplier output (e.g., the output of a multiplication circuit such as a tri-state NOR gate). The accumulation circuit 308 can also generate a sum of the N data elements of the multiplier output.
[0059] The dCIM system 300 of FIG. 3 can receive data elements for storage in an array of memory cells from a non-volatile memory (NVM) array such as a NOR flash memory, a NAND (negative AND) flash memory, and the like. This type of dCIM system 300 can be referred to as an NVM-based dCIM system. Further, the dCIM system 300 of FIG. 3 can receive data elements for storage in an array of memory cells from a volatile memory array such as a dynamic random access memory (DRAM), a SRAM, and the like. These types of dCIM systems can be referred to as SRAM-based dCIM systems, DRAM-based dCIM systems, and / or volatile memory-based dCIM systems, and the like. Any of the dCIM systems described herein with respect to FIGS. 3-15 can be either of the above-described NVM-based dCIM systems and volatile memory-based dCIM systems.
[0060] FIG. 4 shows an enlarged view of a portion of each of three different subgroups of the dCIM system of FIG. 3.
[0061] Specifically, FIG. 4 shows an enlarged view of a portion of each of subgroups 302, 304, and 306. As shown, word line WL<0> 310 is connected to the memory circuit 316 of subgroup 302, word line WL<1> 312 is connected to the memory circuit 402 of subgroup 304, and word line WL <n>314 is connected to the memory circuit 404 of subgroup 306. The tri-state NOR gate 332 of subgroup 302 receives the first timing signal (time 0) and is enabled by the first timing signal, receives the stored weight from the memory circuit 316, receives the input 0 on the input line "input 0 338" at or around the time when the first timing signal (time 0) is received, multiplies the received weight by the input 0 on the input line "input 0 338", and provides the output on the output line "output 0 340". As described above, the memory circuit 316 and the tri-state NOR gate 332 are merely examples, and different types of memory circuits for storage, different types of gates for performing mathematical operations, etc. can be realized using different technologies.
[0062] Similarly, the tri-state NOR gate 406 of subgroup 304 receives the second timing signal (time 1) and is enabled by the second timing signal, receives the stored weight from the memory circuit 402, receives the input 0 on the input line "input 0 338" at or around the time when the second timing signal (time 1) is received, multiplies the received weight by the input 0, and provides the output on the output line "output 0 340". Further, the tri-state NOR gate 408 of subgroup 306 receives the Nth timing signal (time N) and is enabled by the Nth timing signal, receives the stored weight from the memory circuit 404, receives the input 0 on the input line "input 0 338" at or around the time when the Nth timing signal (time N) is received, multiplies the received weight by the input 0, and provides the output on the output line "output 0 340".
[0063] FIG. 5 shows the dCIM system of FIG. 3, where content (data such as N memory data elements) is written to subgroup 0 and one or more other subgroups are processing inputs and calculating outputs. Throughout this document, the content written to the memory circuit may be referred to as weights, memory data elements, etc. The content can be written to various subgroups in the operation of downloading a set of weights.
[0064] Specifically, FIG. 5 shows the same dCIM system 500 as the dCIM system 300 described with reference to FIG. 3, and thus redundant explanations are omitted. In addition to what has been described with reference to FIG. 3, FIG. 5 shows that content is written to the memory circuit 316 of the subgroup 302 in operation 502, content is written to the memory circuit 317 of the subgroup 302 in operation 504, and content is written to the memory circuit 318 of the subgroup 302 in operation 506. The content (data) to be written can include weights such as the first set of weights. In this example, the memory circuits of the subgroups 304 and 306 can already store weights (for example, the memory circuit of the subgroup 304 can store the second set of weights, and the memory circuit of the subgroup 306 can store the Nth set of weights). While content is being written to the memory circuits 316, 317, and 318 of the subgroup 302, the multiplication circuits (e.g., tri-state NOR gates) of the subgroup 304 or the subgroup 306 can perform the multiplication element of the MAC operation and provide the output to an adder tree (e.g., an accumulator).
[0065] FIG. 6 shows a pipelined adder tree used in an embodiment of the dCIM system of FIG. 3.
[0066] Specifically, FIG. 6 shows an adder tree 600 (e.g., an accumulation circuit) that receives 128 outputs on an output line connected to the tri-state NOR gate of the dCIM system of FIG. 3. Output “Output 0 602”, output “Output 1 604” to output “Output 126 606” and output “Output 127 608” are received as inputs 601 to the adder tree 600. The adder tree 600 has seven layers, and these layers include a first adder layer 614 that adds 128 inputs to provide 64 results, a second adder layer 616 that adds the 64 results to provide 32 results, a third adder layer 618 that adds the 32 results to provide 16 results, a fourth adder layer 620 that adds the 16 results to provide 8 results, a fifth adder layer 622 that adds the 8 results to provide 4 results, a sixth adder layer 624 that adds the 4 results to provide 2 results, and a seventh adder layer 626 that adds the 2 results to provide a single output to path 628. The single output provided on path 628 is mathematically equivalent to equation 630 and is the result of the entire MAC operation (for a specific subgroup such as subgroup 302 of FIG. 2).
[0067] As shown, adder tree 600 includes latch buffers 610 and 612 for pipeline stages, which temporarily store intermediate and final results of accumulation operations. For example, latch buffer 610 stores 16 results provided by the third adder layer 618, and latch buffer 612 stores two results provided by the third adder layer 624. Further, as shown, it takes one clock cycle to receive 128 inputs and store 16 results in latch buffer 610, one clock cycle to obtain the 16 results stored in latch buffer 610, perform an accumulation operation, and store two results in latch buffer 612, and one clock cycle to obtain the two results stored in latch buffer 612, perform an accumulation operation, and provide a single output on path 628. Adder tree 600 is structured such that, for example, during the time when subgroup 304 is performing a multiplication operation, it can receive 128 outputs from subgroup 302 in FIG. 3 at a specific clock cycle and store them in latch buffer 610. In a subsequent clock cycle, while adding the previous contents of latch buffer 610 (based on the operation of subgroup 302) and storing them in latch buffer 612, it can receive 128 outputs from subgroup 304 and store them in latch buffer 610. In a further subsequent clock cycle, while adding the previous contents of latch buffer 610 (based on the operation of subgroup 304) and storing them in latch buffer 612, and while adding the previous contents of latch buffer 612 (based on the operation of subgroup 302) and providing them as a single output on path 628, it can receive 128 inputs from subgroup 306 and store them in latch buffer 610.The adder tree 600 can be said to have three stages. The first stage (stage 0) involves receiving inputs and storing them in buffer latch 610. The second stage (stage 1) involves processing the data from buffer latch 610 and storing it in buffer latch 612. The third stage (stage 2) involves processing the data from buffer latch 612 and providing a single output onto path 628.
[0068] While the subgroup of the dCIM system continues to have the written content and continues to perform multiplication operations, the pipeline operation continues. This pipeline operation makes it possible to continue the MAC operation without interruption, because it eliminates the need for the dCIM system to wait for the adder tree 600 to complete the accumulation and provide the result. Therefore, the adder tree 600 can essentially have a higher clock speed and higher processing power. For example, as illustrated and described above, the adder tree 600 actually requires 3 clock cycles to obtain 128 inputs and provide a single output. Therefore, without buffer latches 610 and 612, the multiplication operation that only requires 1 clock cycle has to wait for the adder tree to complete the accumulation operation. As a result, the dCIM system with this adder tree structure operates significantly faster.
[0069] Instead, the adder tree 600 (or any other adder tree described herein) can operate without buffer latches. Furthermore, the buffer latches can be any type of circuit or component that can store data. Moreover, the adder tree 600 (or any other adder tree described herein) can be a counter that performs a count (counting the number of "1"s).
[0070] FIG. 7 shows a dCIM system in which two tri-state NOR gates of two different subgroups share one output line and one word line.
[0071] Specifically, FIG. 7 shows a dCIM system 700 similar to that described with reference to FIG. 3, and thus redundant explanations are omitted. In the dCIM system 700 of FIG. 7, subgroup 0 702 and subgroup 1 703 are connected to the same word line WL<0> 310, subgroup 2 704 and subgroup 3 705 are connected to the same word line WL<1> 312, and subgroup N-1 706 and subgroup N 707 are connected to the same word line WL <n>It differs from the dCIM system of FIG. 3 in terms of the points connected to 314. The dCIM system 700 of FIG. 7 includes the same word lines 310, 312, and 314, the same memory circuits 316, 317, and 318, the same bit lines and bit line bar lines 320, 322, 324, 326, 328, and 330, and the same tri-state NOR gates 322, 324, and 326 as those described above with reference to FIG. 3.
[0072] The dCIM system 700 of FIG. 7 also differs in that two tri-state NOR gates of two different subgroups 702 and 703 share one output line "Output 0 340" and one word line 310. For example, the tri-state NOR gate 332 of subgroup 702 and the tri-state NOR gate 334 of subgroup 703 both provide outputs to the output line "Output 0 340". Similarly, two of the tri-state NOR gates of two different subgroups 704 and 705 provide outputs to the same output line "Output 0 340", and two of the tri-state NOR gates of two different subgroups 706 and 707 provide outputs to the same output line "Output 0 340".
[0073] In this example, when 128 tri-state NOR gates exist within subgroups 702 and 703, 64 tri-state NOR gates are members of subgroup 702 and 64 tri-state NOR gates are members of subgroup 703. Furthermore, in this example, the tri-state NOR gate 332 is a member of subgroup 702, and the tri-state NOR gates 334 and 336 are members of subgroup 703. In addition, 64 tri-state NOR gates are members of subgroup 704, and 64 tri-state NOR gates are members of subgroup 705. Similarly, 64 tri-state NOR gates are members of subgroup 706, and 64 tri-state NOR gates are members of subgroup 707.
[0074] As shown, at time 00 (e.g., the first clock cycle), the tri-state NOR gates that are members of subgroup 702 provide 64 outputs on output lines "Output 0 340" to "Output m / 2 708". In this example, m (or L) = 128, meaning that there are 128 tri-state NOR gates for the combination of subgroups 702 and 703 and 64 (128 / 2) outputs. This subgroup architecture of the dCIM system 700 results in 64 outputs per clock cycle, as opposed to 128 outputs per clock cycle described above with reference to the dCIM system 300 of FIG. 3. At time 01 (e.g., the second clock cycle), the tri-state NOR gates of subgroup 703 are enabled to multiply and provide 64 outputs. As a result, subgroups 702 and 703 are utilized over two clock cycles, which allows for additional clock cycles to write content to other groups while subgroups 702 and 703 are performing multiplication and also enables the use of a smaller adder tree (accumulator), as will be described below with reference to FIG. 8. Content can be written simultaneously to multiple subgroups connected to the same word line.
[0075] Returning to FIG. 7, at time 10 (e.g., the third clock cycle), the tri-state NOR gates that are members of subgroup 704 provide 64 outputs on output lines "Output 0 340" to "Output m / 2 708". At time 11 (e.g., the fourth clock cycle), the tri-state NOR gates of subgroup 705 are enabled to multiply and provide 64 outputs. At time N0 (e.g., the (N - 1)th clock cycle), the tri-state NOR gates that are members of subgroup 706 provide 64 outputs on output lines "Output 0 340" to "Output m / 2 708". At time N1 (e.g., the Nth clock cycle), the tri-state NOR gates of subgroup 707 are enabled to multiply and provide 64 outputs. This architecture shown in FIG. 7 can be modified to include eight or more subgroups per group. For example, the number of "adjacent" NOR gates in subgroups 702 and 703 that share the same output line can be two, four, eight, or any N number of two, where N is an integer.
[0076] FIG. 8 shows the adder tree of the dCIM system of FIG. 7.
[0077] As described above, the dCIM system 700 of FIG. 7 provides, for example, 64 outputs per clock cycle, in contrast to 128 outputs per clock cycle of the dCIM system 300 of FIG. 3. The adder tree 800 of FIG. 8 has 64 inputs and receives the 64 outputs of the dCIM system 700 of FIG. 7.
[0078] Specifically, adder tree 800 receives outputs "Output 0 802", "Output 1 804" to "Output 62 806", and "Output 63 808". These outputs are received as 64 inputs 801 of adder tree 800. This adder tree 800 has one less hierarchy than adder tree 600 of FIG. 6. For example, adder tree 800 includes (i) a first adder tree layer 814 that receives 64 inputs and provides 32 results, (ii) a second adder tree layer 816 that receives 32 results and provides 16 results, (iii) a third adder tree layer 818 that receives 16 results and provides 8 results, (iv) a buffer latch 810 that receives 8 results and stores them temporarily, (v) a fourth adder tree layer 820 that receives the 8 temporarily stored results and provides 4 results, (vi) a fifth adder tree layer 822 that receives 4 results and provides 2 results, and (vii) a sixth adder tree layer 824 that receives 2 results and provides a single result. Similar to adder tree 600 of FIG. 6, adder tree 800 provides a single result onto path 628, and this single result is mathematically equivalent to equation 630 and is the result of a specific subgroup of MAC operations.
[0079] As shown, adder tree 800 completes the accumulation operation in two clock cycles, requires one clock cycle to receive an input and store the result in buffer latch 810 (stage 0), and requires one clock cycle to retrieve the stored data from buffer latch 810 and provide a single output (stage 1). This has the same pipeline structure as described above for adder tree 600, except that there is only one buffer latch as opposed to two buffer latches, and except that it requires only two clock cycles as opposed to three clock cycles to complete the accumulation operation. The reduction in the number of buffer latches and the number of clock cycles required is because adder tree 800 receives only 64 inputs as opposed to 128 inputs for adder tree 600. As a result, the gate count of adder tree 800 (i.e., the gates that perform the accumulation) is approximately half the gate count of adder tree 600, and thus the size of adder tree 800 is approximately half the size of adder tree 600. This adder tree 800 with fewer inputs is the result of the subgroup structure and the sharing of output lines described above with reference to FIG. 7.
[0080] FIG. 9 shows a dCIM system in which four NOR gates of four different subgroups share one output line and one word line.
[0081] The dCIM system 900 of FIG. 9 is similar to the dCIM system 700 of FIG. 7, and redundant descriptions of the components described with reference to FIG. 7 are omitted.
[0082] The dCIM system 900 of FIG. 9 differs from the dCIM system of FIG. 7 in that there are four different subgroups of tri-state transistors that share one output line and one word line. Specifically, FIG. 9 shows that four subgroups (i.e., subgroup 0 903, subgroup 1 904, subgroup 2 905, and subgroup 3 906, hereinafter subgroups 903, 904, 905, and 906) are connected to the same word line WL<0> 310, and further shows that four subgroups (i.e., subgroup 4 907, subgroup 5 908, subgroup 6 909, and subgroup 7 910, hereinafter subgroups 907, 908, 909, and 910) are connected to the same word line WL<1> 312. Additional sets of four subgroups up to subgroups N-3, N-2, N-1, and N can be connected to each word line. The dCIM system 900 of FIG. 9 includes the same word lines 310 and 312, the same memory circuits, and the same tri-state NOR gates as described with respect to FIGS. 3 and 7.
[0083] Subgroup 903 includes a memory circuit and a tri-state NOR gate associated with time 0, subgroup 904 includes a memory circuit and a tri-state NOR gate associated with time 1, subgroup 905 includes a memory circuit and a tri-state NOR gate associated with time 2, subgroup 906 includes a memory circuit and a tri-state NOR gate associated with time 3, subgroup 907 includes a memory circuit and a tri-state NOR gate associated with time 4, subgroup 908 includes a memory circuit and a tri-state NOR gate associated with time 5, subgroup 909 includes a memory circuit and a tri-state NOR gate associated with time 6, and subgroup 910 includes a memory circuit and a tri-state NOR gate associated with time 7.
[0084] The bit line configuration of FIG. 9 is different from that of FIG. 7. Specifically, bit lines BL0 912, BL1 914, and Vref916 (e.g., a reference voltage line) are implemented to write content to specific memory circuits in subgroups 903, 904, 907, and 908 (e.g., memory circuits in the first left column). As shown, Vref916 is connected to the memory circuits above and below word line WL<0> 310 and the memory circuits above and below word line WL<1> 312, while bit line BL1 914 is connected to the memory circuits above word line WL<0> 310 and above word line WL<1> 312, and bit line BL0 912 is connected to the memory circuits below word line WL<0> 310 and below word line WL<1> 312. Further, bit lines BL2 920, BL3 922, and Vref918 are implemented to write content to specific memory circuits in subgroups 905, 906, 909, and 910 (e.g., memory circuits in the second column on the right side of the first column). As shown, Vref918 is connected to the memory circuits above and below word line WL<0> 310 and the memory circuits above and below word line WL<1> 312, while bit line BL3 922 is connected to the memory circuits above word line WL<0> 310 and above word line WL<1> 312, and bit line BL2 920 is connected to the memory circuits below word line WL<0> 310 and below word line WL<1> 312. Bit lines BL510 926, BL511 928, and Vref924 are implemented to write content to specific memory circuits in subgroups 905, 906, 909, and 910 (e.g., memory circuits in the rightmost column). As shown, Vref924 is connected to the memory circuits above and below word line WL<0> 310 and the memory circuits above and below word line WL<1> 312, while bit line BL511 928 is connected to the memory circuits above word line WL<0> 310 and above word line WL<1> 312, and bit line BL510 926 is connected to the memory circuits below word line WL<0> 310 and below word line WL<1> 312.
[0085] The dCIM system 900 of FIG. 9 is different from the dCIM system 700 in that four tri-state NOR gates of subgroups 903, 904, 905, and 906 and four tri-state NOR gates of subgroups 907, 908, 909, and 910 share a single output line "Output 0 936". For example, the tri-state NOR gates related to times 0, 1, 2, and 3 of subgroups 903, 904, 905, and 906 can all provide outputs to the output line "Output 0 936". Similarly, the tri-state NOR gates related to times 4, 5, 6, and 7 of subgroups 907, 908, 909, and 910 provide outputs to the same output line "Output 0 936". Input lines 930, 932, and 934 provide inputs to various tri-state NOR gates.
[0086] In this example, if there are 512 tri-state NOR gates in subgroups 903, 904, 905, and 906, then 128 tri-state NOR gates are members of subgroup 903, 128 tri-state NOR gates are members of subgroup 904, 128 tri-state NOR gates are members of subgroup 905, and 128 tri-state NOR gates are members of subgroup 906. In addition, 128 tri-state NOR gates are members of subgroup 907, 128 tri-state NOR gates are members of subgroup 908, 128 tri-state NOR gates are members of subgroup 909, and 128 tri-state NOR gates are members of subgroup 910. Other configurations are possible depending on the desired number of outputs per clock cycle. In this example of FIG. 9, there are 128 outputs. If there were 64 outputs, the number of tri-state NOR gates per subgroup would be reduced to 64.
[0087] As shown, at time 0 (e.g., the first clock cycle), the tri-state NOR gates that are members of subgroup 903 output 128 outputs onto output lines "Output 0 936" to "Output 127 938". At time 1 (e.g., the second clock cycle), the tri-state NOR gates of subgroup 904 are enabled to multiply and output 128 outputs. At time 2 (e.g., the third clock cycle), the tri-state NOR gates of subgroup 905 are enabled to multiply and output 128 outputs. At time 3 (e.g., the fourth clock cycle), the tri-state NOR gates of subgroup 906 are enabled to multiply and output 128 outputs. As a result, subgroups 903, 904, 905, and 906 are utilized for four clock cycles, which enables additional clock cycles for writing content to other subgroups connected to other word lines.
[0088] At time 4 (e.g., the fifth clock cycle), the tri-state NOR gates that are members of subgroup 907 provide 128 outputs onto output lines "Output 0 936" to "Output 127 938". At time 5 (e.g., the sixth clock cycle), the tri-state NOR gates that are members of subgroup 908 provide 128 outputs onto output lines "Output 0 936" to "Output 127 938". At time 6 (e.g., the seventh clock cycle), the tri-state NOR gates that are members of subgroup 909 provide 128 outputs onto output lines "Output 0 936" to "Output 127 938". At time 7 (e.g., the eighth clock cycle), the tri-state NOR gates that are members of subgroup 910 provide 128 outputs onto output lines "Output 0 936" to "Output 127 938".
[0089] FIG. 9 further shows that the content of the memory array 902 is written into the memory circuits of the subgroups 903, 904, 905, 906, 907, 908, 909, and 910. As described above, the content (data element) can be any type of NVM and volatile memory. While the subgroups 907, 908, 909, and 910 are performing multiplication operations, the content (data element) of the memory array 902 can be written into the memory circuits of the subgroups 903, 904, 905, and 906, and while the subgroups 903, 904, 905, and 906 are performing multiplication operations, the content can be written into the memory circuits of the subgroups 907, 908, 909, and 910. It can be written into any subgroup that does not share a word line with the active subgroup. As shown, two memory circuits on a single word line (e.g., WL<0> 310) can share one output line (e.g., output 0 936), can be connected to separate bit lines (e.g., BL0 912 and BL1 914), and can be connected to one reference line (e.g., Vref 916), which enables storing different data in the two memory circuits.
[0090] FIG. 10 shows the operation time series of the subgroups 0 903 to 7 910 of FIG. 9. FIG. 10 identifies the subgroup 0 903 as "subgroup 0", the subgroup 1 904 as "subgroup 1", and so on.
[0091] Specifically, FIG. 10 shows a chart (diagram) 1000 including a time series (time sequence) of a cycle operation starting from time 0 to time 7 and then starting from time 0 again. The top of the time series identifies a clock signal 1002. Below the corresponding time series of the clock signal 1002, the chart 1000 shows a column 1004 of subgroups, indicating which subgroups are performing and / or outputting mathematical operations. Below the column of subgroups, the chart 1000 shows a pipeline 1006 of an adder tree, and the pipeline 1006 indicates which stage of the adder tree is in an active state. Since FIG. 9 shows 128 outputs, the adder tree 600 in FIG. 6 can be implemented to have three stages. The first stage (stage 0) includes receiving an input and storing the result in a buffer latch 610. The second stage (stage 1) includes processing the data from the buffer latch 610 and storing the result in a buffer latch 612. The third stage (stage 2) includes processing the data from the buffer latch 612 and providing a single output to a path 628.
[0092] As shown, pipeline 1006 of the adder tree indicates which stage of the adder tree is processing the input / data. Further, this chart shows download process 1008 for downloading content, and download process 1008 includes writing the content to various subgroups of memory circuits. For example, from time 0 to time 3, while subgroups 0, 1, 2, and 3 are performing multiplication operations and / or output operations, the content is downloaded and written to the memory circuits of subgroups 4, 5, 6, and 7. Similarly, from time 4 to time 7, while subgroups 4, 5, 6, and 7 are performing multiplication operations and / or output operations, the content is downloaded and written to the memory circuits of subgroups 0, 1, 2, and 3. Next, again, while subgroups 0, 1, 2, and 3 are performing multiplication operations and / or output operations using the content written during time 4 - 7, the content is downloaded and written to the memory circuits of subgroups 4, 5, 6, and 7. The download and write time may require two or more cycles. Thus, the advantage of the dCIM system 900 of FIG. 9 is that there are four clock cycles available for writing content to a particular set of subgroups (in this example, one set of subgroups includes four subgroups). This structure reduces or eliminates the time spent waiting for the content writing to finish before starting the multiplication operation and / or output operation.
[0093] FIG. 11 shows a dCIM system in which content is written using bit lines and a reference voltage (Vref).
[0094] Specifically, FIG. 11 shows a dCIM system 1100 similar to that described with reference to FIG. 5, and thus redundant explanations are omitted. The dCIM system 1100 is different from the dCIM system 500 of FIG. 5 in that the bit line bar lines are replaced by Vref lines. Specifically, FIG. 11 shows that the Vref lines 1102, 1104, and 1106 are connected to the memory circuit for writing to the memory circuit. Replacing the bit line bar lines with Vref lines has the advantage of doubling the number of SA-latches (sense amplifiers and latches) and increasing the sensing speed by a factor of two. For example, when 1024 memory BLs are present in the memory array, the 1024 memory BLs are divided into 512 SRAM BLs and 512 SRAM BLBs, and only 512 bits can be sensed at a time. When there are 1024 memory BLs connected to 1024 SRAM BLs and 1024 BLBs connected to Vref, 1024 bits of data can be sensed at a time.
[0095] FIG. 12 shows a dCIM system in which content is written using a sense amplifier.
[0096] Specifically, FIG. 12 shows a dCIM system 1200 similar to that described with reference to FIG. 5, and thus redundant explanations are omitted. FIG. 12 shows a memory array 1202 for storing content, and here, a read operation 1210 is executed to read the content from the memory array 1202 into sense amplifiers SA0 1204, SA1 1206, and SAm 1208 via respective bit lines and bit line bar lines BL0, BL0B, BL1, BL1B, BLm, and BLmB. As described above, the memory array 1202 for storing content (data elements) can be any type of NVM and volatile memory. The dCIM system 1200 in FIG. 12 is different from the dCIM system 500 in FIG. 5 in that it includes sense amplifiers for providing content to various subgroups. As shown, a write operation 1212 is executed to write content from the sense amplifiers to various subgroups of memory circuits. Specifically, the sense amplifier SA0 1206 provides content on the SA DL0 line 1214 and the SA DL0B line 1216 and writes this content to a specific memory circuit, the sense amplifier SA1 1206 provides content on the SA DL1 line 1218 and the SA DL1B line 1220 and writes this content to a specific memory circuit, and the sense amplifier SAm 1208 provides content on the SA DLm line 1222 and the SA DLmB line 1224 and writes this content to a specific memory circuit. Instead of using bar lines (i.e., SA DL0B line 1216, SA SL1B line 1220, SA DLmB line 1224) to write content, content can be written using Vref. In contrast to programming the memory circuit directly from the bit lines and bit line bar lines, using sense amplifiers enables faster programming because the sense amplifiers can supply a higher voltage signal than can be supplied on the bit lines and bit line bar lines.
[0097] FIG. 13 shows a timing chart for performing a MAC operation using the dCIM system 500 of FIG. 5.
[0098] Specifically, the timing chart 1300 indicates that the input can remain stable until all weights are multiplied by the input. For example, FIG. 13 shows the active subgroup 1302 of the dCIM system 500 of FIG. 5, the time series 1304 from time 0 to time N, and the input sequence (input column) 1306 of input 0 338. This timing chart can be modified to fit other dCIM systems described herein.
[0099] As shown, with respect to FIG. 5, N = 64 such that there are 64 subgroups. The active subgroup 1302 indicates progression through the entire subgroup in the order of this time series. The input sequence 1306 on input 0 338 remains the same (Xi0) and is multiplied by weights Wi0, Wi1, Wi2, Wi3, Wi4, Wi5, Wi6, and Wi7 from time 0 to time 7. Next, at time 8, the input changes to Xi1 and is multiplied by weights Wi0, Wi1, Wi2, Wi3, Wi4, Wi5, Wi6, and Wi7 from time 8 to time 15. This process continues until at time 63, weight Wi7 is multiplied by Xi7. As described herein, while a subgroup connected to a particular word line is performing a multiplication operation and / or an output operation, the weights of other subgroups connected to other word lines are updated. This series of operations shown in FIG. 13 can be referred to as "input stability and sequential serial input in the product operation".
[0100] FIG. 14 shows a dCIM system that uses pass gates instead of the tri-state NOR gates shown in FIG. 3.
[0101] Specifically, FIG. 14 shows a structure 1400 similar to the structure of FIG. 4, and its redundant description is omitted. However, in structure 1400, the tri-state NOR gates 332, 406, and 408 of FIG. 4 are replaced with pass gates 1402, 1404, and 1406. Just like in FIG. 4, N pass gates can exist for each subgroup. At time 0 (when receiving the timing signals time 0 and time 0B), pass gate 1402 enables the content stored in memory circuit 316 to be passed to NOR gate 1408, and NOR gate 1408 multiplies the content stored in memory circuit 316 by the input 1410 provided at time 0. N NOR gates can exist within structure 1400, and each of the N NOR gates corresponds to a specific column of pass gates. NOR gate 1408 provides the output "output 0 1412" to the adder tree (accumulator). At time 1 (when receiving the timing signals time 1 and time 1B), pass gate 1404 enables the content stored in memory circuit 402 to be passed to NOR gate 1408, and NOR gate 1408 multiplies the content stored in memory circuit 402 by the input 1410 provided at time 1. At time N (when receiving the timing signals time N and time NB), pass gate 1406 enables the content stored in memory circuit 404 to be passed to NOR gate 1408, and NOR gate 1408 multiplies the content stored in memory circuit 404 by the input 1410 provided at time N. This implementation eliminates the need for tri-state NOR gates and the space occupied by tri-state NOR gates.
[0102] FIG. 15 shows a simplified block diagram of an integrated circuit device including a memory array configured for in-memory computing with signed (or unsigned) inputs and weights.
[0103] Specifically, FIG. 15 shows an integrated circuit device 1500 including a memory array 1560, which is configured for signed in-memory computing for CIM operations such as signed (or unsigned) multiply-accumulate (MAC) operations executed by the dCIM system described herein. The integrated circuit device 1500 can be implemented on a single chip or on a multi-chip module.
[0104] Device 1500 includes an input / output circuit 1506 for communication of control signals, data, addresses, and commands with other data processing resources such as a CPU (central processing unit) or a memory controller.
[0105] Input / output data is provided on bus 1591 to controller 1510 and to cache 1590. Also, addresses are provided on bus 1593 to decoder 1542 and to controller 1510. Also, buses 1591 and 1593 can be operatively connected to internal data sources of integrated circuit device 1500 such as a general-purpose processor or a dedicated processor, or a combination of modules providing, for example, system-on-chip functionality.
[0106] Memory array 1560 can include an array of memory cells in the form of columns along bit lines and rows along word lines in the form of a NOR architecture or a NAND architecture, and the memory cells within a given column are connected in parallel between the bit line and a reference source. The reference source can comprise a ground terminal or a power line connected to a bias power supply on the reference source side. These memory cells can comprise charge trap transistor cells arranged in the form of a 3D (three-dimensional) structure. Memory array 1560 having in-memory (or near-memory) computing can be configured as described above with respect to FIGS. 3-14 and can execute the operations described above with respect to FIGS. 3-14.
[0107] The bit lines can be connected to the global bit line 1565 by a block selection circuit, and the global bit line 1565 is configured for selectable connection to the page buffer 1580 and to the CIM sense circuit 1570.
[0108] The page buffer 1580 in the illustrated embodiment is connected to the cache 1590 by a bus 1585. The page buffer 1580 includes memory elements (which can be various types of memory arrays) and a sensing circuit for memory operations including read and write operations. For flash memories including dielectric charge trap memories and floating gate charge trap memories, the write operation includes a program operation and an erase operation.
[0109] The driver circuit 1540 is connected to the word lines 1545 in the array 1560 and applies a word line voltage to the selected word line in response to a decoder 1542 that decodes an address on the bus 1593 or, in a computing operation, in response to input data stored in the input buffer 1541.
[0110] The controller 1510 is coupled to the cache 1590 and the memory array 1560 and is also coupled to other peripheral circuits used in memory access operations and memory computing operations.
[0111] The controller 1510 controls, for example using a state machine, the application of power supply voltages and currents generated by or supplied through voltage or current sources within the block 1520 for memory operations and CIM operations.
[0112] The controller 1510 includes a control register and a status register, and includes a control logic circuit that can be implemented using a dedicated logic circuit including a state machine and combinational logic circuits known in the current art. In an alternative embodiment, the control logic circuit comprises a general-purpose processor, which can be implemented on the same integrated circuit and executes a computer program to control the operation of the device. In yet another embodiment, a combination of a dedicated logic circuit and a general-purpose processor can be utilized to implement the control logic circuit.
[0113] The array 1560 includes memory cells arranged in columns and rows, where the memory cells within a column are connected to corresponding bit lines, and the memory cells within a row are connected to corresponding word lines. The array 1560 is programmable to store signed coefficients (weights Wi) in multiple sets of memory cells.
[0114] In the CIM mode, the word line driver circuit 1540 or the driver circuit 1540 can include a driver (referred to as an input driver or an input activation driver) configured to operate on the signed or unsigned input Xi from the input buffer 1541. The driver circuit 1540 can be separate from the word line driver circuit. The CIM sense circuit 1570 is configured to sense the difference between a first current and a second current on each bit line in a selected pair of bit lines and generate an output for the selected pair of bit lines as a function of the difference. This output can be provided to storage elements within the page buffer 1580 and to the cache 1590.
[0115] In one embodiment, the first subgroup can include a first memory circuit and a first multiplication circuit. The first memory circuit is connected to a first word line and is programmable to store a first set of weights. The first multiplication circuit is enabled by a first timing signal to multiply an input by the first set of weights and provide a first output. The second subgroup includes a second memory circuit and a second multiplication circuit. The second memory circuit is connected to a second word line and is programmable to store a second set of weights. The second multiplication circuit is enabled by a second timing signal to multiply an input by the second set of weights and provide a second output. The second multiplication circuit is enabled at a time different from the time when the first multiplication circuit is enabled. The accumulation circuit receives and accumulates (i) the first output in response to the first multiplication circuit being enabled by the first timing signal and (ii) the second output in response to the second multiplication circuit being enabled by the second timing signal.
[0116] In one embodiment, the in-memory computing circuit can include a first output line and a second output line. The first output line is shared by the output of one of the multiplication circuits of the first multiplication circuit of the first subgroup and the output of one of the multiplication circuits of the second multiplication circuit of the second subgroup. The second output line is shared by the output of the other multiplication circuit of the first multiplication circuit of the first subgroup and the output of the other multiplication circuit of the second multiplication circuit of the second subgroup. The first output of the first subgroup is provided to the accumulation circuit via the first and second output lines in response to the first multiplication circuit being enabled and the second multiplication circuit not being enabled. The second output of the second subgroup is provided to the accumulation circuit via the first and second output lines in response to the second multiplication circuit being enabled and the first multiplication circuit not being enabled.
[0117] In a further embodiment, a specific memory circuit among the first memory circuits of the first subgroup and a specific memory circuit among the second memory circuits of the second subgroup can share a common programming line for controlling the storage of their respective weights. The specific memory circuit of the first subgroup is programmed to store a specific weight among the first set of weights in response to the first word line activating the first subgroup. The specific memory circuit of the second subgroup is programmed to store a specific weight among the second set of weights in response to the second word line activating the second subgroup.
[0118] In one embodiment, an in-memory computing circuit can include a multiplication circuit, a first subgroup of circuits, a second subgroup of circuits, and an accumulation circuit. The multiplication circuit is configured to receive and multiply inputs to provide an output. The first subgroup of circuits is connected to a first word line and is configured to (i) store a first set of weights and (ii) provide the first set of weights to the multiplication circuit in response to a first timing signal enabling a first pass gate. The second subgroup of circuits is connected to a second word line and is configured to (i) store a second set of weights and (ii) provide the second set of weights to the multiplication circuit in response to a second timing signal enabling a second pass gate. The enabling of the provision of the second set of weights is at a time different from the time when the enabling of the provision of the first set of weights is enabled. The accumulation circuit is shared by the first subgroup and the second subgroup and is configured to (i) receive and accumulate a first output received from the multiplication circuit in response to the first pass gate being enabled by the first timing signal and (ii) receive and accumulate a second output received from the multiplication circuit in response to the second pass gate being enabled by the second timing signal.
[0119] The realization of the memory array can be based on charge and wrapping memory cells, such as floating gate memory cells that can include polysilicon charge and a wrapping layer, or dielectric charge and wrapping memory cells that can include silicon nitride charge and a wrapping layer. In various embodiments of the techniques described herein, other types of memory technologies can be applied.
[0120] Other realizations of the methods described in this section can include a non-transitory computer-readable storage medium that stores instructions executable for a processor to perform any of the methods described above, and still other realizations of the methods described in this section can include a system that includes a memory and one or more processors, where the one or more processors are operable to execute instructions stored in the memory to perform any of the methods described above.
[0121] Any data structures and code described above or referenced above can be stored on a computer-readable storage medium in a number of realizations, and these storage media can be any device or medium that can store code and / or data for use by a computer system. This device or medium can include, but is not limited to, volatile memory, non-volatile memory, application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), magnetic disks, magnetic tapes, CDs (compact discs), DVDs (digital versatile discs or digital video discs), or other media that can store computer-readable media that are currently known or later developed.
[0122] An example of a processor is a hardware unit that is enabled to execute program code (e.g., having a hardware circuit such as one or more active devices). The processor optionally includes one or more controllers and / or state machines. The processor can be realized by application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), and / or custom design techniques. The processor can be manufactured by integrated circuits, and optical and quantum technologies. The processor uses one or more architecture techniques such as sequential (e.g., von Neumann) processing, very long instruction word processing. The processor uses one or more microarchitectures to execute multiple instructions one by one or in parallel. The processor can be directed to general-purpose users and / or specific-purpose users (such as signal, voice, video, and / or graphics users). The processor has fixed functionality or variable functionality, e.g., by programming. The processor includes any one or more of registers, memory, logic units, arithmetic units, and graphics units. The term "processor" means a processor in the singular as well as processors in the plural, including multi-processors and / or aggregates of processors.
[0123] The logic circuits described in this specification can be implemented using a processor programmed with a computer program, which is stored in a memory accessible to a computer system and is executable by the processor, by dedicated logic hardware including a field programmable integrated circuit, and by a combination of dedicated logic hardware and a computer program. It is understood that, for all the flowcharts in this specification, many of the steps can be combined, executed in parallel, or executed in a different order without affecting the functions achieved. In some cases, as will be understood by the reader, rearranging the steps will achieve the same result only if other specific changes are also made. In other cases, as will be understood by the reader, rearranging the steps will achieve the same result only if certain conditions are satisfied. Further, it is understood that the flowcharts in this specification show only the steps relevant to the understanding of the present invention, and that numerous additional steps for achieving other functions can be executed before, after, and between the steps shown.
[0124] The present invention has been disclosed by reference to the preferred embodiments and examples detailed above, but it should be understood that these examples are intended in an illustrative rather than a limiting sense. Those skilled in the art can readily conceive of changes and combinations, and it is contemplated that these changes and combinations will fall within the spirit of the present invention and the scope of the following claims.< / n> < / n> < / n>
Claims
1. One or more input lines for receiving M input data elements, where M is an integer greater than 0, and An array of memory cells including one or more sub - groups, wherein each sub - group in the one or more sub - groups stores M memory data elements, and A multiplication circuit connected to the array of memory cells and the one or more input lines, configured to multiply the M input data elements by the M memory data elements in a selected sub - group of the one or more sub - groups to provide a multiplier output having M data elements, and An accumulator circuit including an accumulator input for the M data elements connected to the multiplier output, configured to generate a sum of the M data elements of the multiplier output, a memory - in computing circuit comprising The multiplication circuit provides a multiplication result from the one or more sub - groups to the multiplier output, a memory - in computing circuit.
2. The memory - in computing circuit according to claim 1, wherein the multiplication circuit includes M tri - state multipliers connected to the multiplier output for each sub - group in the one or more sub - groups.
3. The memory - in computing circuit according to claim 2, wherein the M tri - state multipliers are M tri - state NOR gates.
4. The one or more sub - groups include a first sub - group storing M memory data elements and a second sub - group storing M memory data elements, The M tri - state multipliers for the first sub - group are enabled by a first timing signal to multiply the M input data elements by the M memory data elements of the first sub - group, The M tri - state multipliers for the second sub - group are enabled by a second timing signal to multiply the M input data elements by the M memory data elements of the second sub - group, and the second timing signal is supplied at a different time from the first timing signal such that the M tri - state multipliers for the second sub - group are enabled at a different time from the M tri - state multipliers for the first sub - group. The memory - in computing circuit according to claim 2.
5. The one or more subgroups include a first subgroup that stores M memory data elements in M memory circuits, and a second subgroup that stores M memory data elements in M memory circuits, the first subgroup is connected to a first word line, the second subgroup is connected to a second word line, the in-memory computing circuit according to claim 2. **Claim 6** a specific memory circuit among the M memory circuits of the first subgroup and a specific memory circuit among the M memory circuits of the second subgroup share a common line for controlling the storage of respective data elements, the specific memory circuit of the first subgroup stores a specific data element in response to the first word line activating the first subgroup, the specific memory circuit of the second subgroup stores a specific data element in response to the second word line activating the second subgroup, the in-memory computing circuit according to claim 5. **Claim 7** the common line shared by the specific memory circuit of the first subgroup and the specific memory circuit of the second subgroup includes a bit line (BL), the in-memory computing circuit according to claim 6. **Claim 8** (i) While the M tri-state multipliers for the first subgroup are enabled by a first timing signal to multiply the M input data elements by the M memory data elements of the first subgroup to provide a multiplier output having the M data elements, and (ii) while the accumulation circuit receives and accumulates the multiplier output having the M data elements, a specific data element is written to a specific memory circuit among the M memory circuits of the second subgroup, the in-memory computing circuit according to claim 5. **Claim 9** the multiplier output includes a first output line and a second output line, the first output line is shared by the output of one of the tri-state multipliers for the first subgroup and the output of one of the tri-state multipliers for the second subgroup, the second output line is shared by the output of another one of the tri-state multipliers for the first subgroup and the output of another one of the tri-state multipliers for the second subgroup, In response to the M tri-state multipliers for the first subgroup being enabled by a timing control signal and the M tri-state multipliers for the second subgroup not being enabled by the timing control signal, the output related to the first subgroup is provided to the accumulation circuit via the first output line and the second output line. The memory-integrated circuit according to claim 5, wherein in response to the M tri-state multipliers for the second subgroup being enabled by the timing control signal and the M tri-state multipliers for the first subgroup not being enabled by the timing control signal, the output related to the second subgroup is provided to the accumulation circuit via the first output line and the second output line.
10. The one or more subgroups include a first subgroup storing M stored data elements, a second subgroup storing M stored data elements, a third subgroup storing M stored data elements, and a fourth subgroup storing M stored data elements. The first subgroup and the second subgroup are connected to a first word line. The third subgroup and the fourth subgroup are connected to a second word line. The M tri-state multipliers for the first subgroup are enabled by a first timing signal to multiply the M input data elements by the M stored data elements of the first subgroup. The M tri-state multipliers for the second subgroup are enabled by a second timing signal to multiply the M input data elements by the M stored data elements of the second subgroup. The M tri-state multipliers for the third subgroup are enabled by a third timing signal to multiply the M input data elements by the M stored data elements of the third subgroup. The M tri-state multipliers for the fourth subgroup are enabled by a fourth timing signal to multiply the M input data elements by the M stored data elements of the fourth subgroup. L is an integer representing the total number of the M memory data elements of the first subgroup and the M memory data elements of the second subgroup, and M = L / 2. The in-memory computing circuit according to claim 2.
11. The multiplier output includes a first output line, The first output line is shared by the output of one of the tri-state multipliers for the first subgroup, the output of one of the tri-state multipliers for the second subgroup, the output of one of the tri-state multipliers for the third subgroup, and the output of one of the tri-state multipliers for the fourth subgroup. The in-memory computing circuit according to claim 10.
12. The multiplier output includes a second output line, The second output line is shared by the output of another one of the tri-state multipliers for the first subgroup, the output of another one of the tri-state multipliers for the second subgroup, the output of another one of the tri-state multipliers for the third subgroup, and the output of another one of the tri-state multipliers for the fourth subgroup. The in-memory computing circuit according to claim 11.
13. The M memory data elements of each subgroup in the one or more subgroups are written into their respective subgroups using bit lines. The in-memory computing circuit according to claim 1.
14. The M memory data elements of each subgroup in the one or more subgroups are written into their respective subgroups using sense amplifiers connected to bit lines. The in-memory computing circuit according to claim 1.
15. The one or more subgroups include a first subgroup storing M memory data elements and a second subgroup storing M memory data elements, During a first clock cycle, the first subgroup multiplies the M input data elements by the M memory data elements, and the M memory data elements are written into the second subgroup. During a second clock cycle, the second subgroup multiplies the M input data elements by the M memory data elements, and the M memory data elements are written into the first subgroup. The in-memory computing circuit according to claim 1.
16. The one or more subgroups include a first subgroup that stores M memory data elements and a second subgroup that stores M memory data elements, During a specific clock cycle, the accumulation circuit accumulates the output related to the first subgroup, During a subsequent clock cycle, the accumulation circuit accumulates the output related to the second subgroup, The in-memory computing circuit according to claim 1. **Claim 17** The in-memory computing circuit according to claim 16, wherein the accumulation circuit is pipelined. **Claim 18** The in-memory computing circuit according to claim 1, wherein the multiplication circuit includes M pass gates connected to a shared M-bit multiplier connected to the multiplier output for each subgroup in the one or more subgroups. **Claim 19** The in-memory computing circuit according to claim 1, wherein the multiplication circuit is enabled by a timing control signal to provide a multiplication result. **Claim 20** The in-memory computing circuit according to claim 19, wherein the timing control signal includes a first timing signal and a second timing signal, and the first timing signal is supplied at a different time from the second timing signal. **Claim 21** The in-memory computing circuit according to claim 7, wherein the common line further includes a bit line bar line (BLB). **Claim 22** The in-memory computing circuit according to claim 7, wherein the common line further includes a reference voltage line (VREF). **Claim 23** The in-memory computing circuit according to claim 1, wherein the array of memory cells includes latches. **Claim 24** A method of performing an operation using an in-memory computing circuit, wherein the in-memory computing circuit includes: (i) an array of memory cells including one or more subgroups, each subgroup in the one or more subgroups storing M memory data elements, where M is an integer greater than 0; (ii) a multiplication circuit connected to the array of memory cells and one or more input lines; and (iii) an accumulation circuit including accumulator inputs for M data elements connected to the multiplier output. In the method, Obtaining M input data elements from the one or more input lines, A step of multiplying, by the multiplication circuit, the M input data elements by the M storage data elements within a selected sub-group of the one or more sub-groups to provide a multiplier output having M data elements, wherein the multiplication circuit is enabled by a timing control signal to sequentially provide multiplication results from the one or more sub-groups to the multiplier output; A step of generating, by the accumulation circuit, a sum of the M data elements of the multiplier output; A method including the above.
25. A first sub-group of circuits connected to a first word line and configured to store a first set of weights; A second sub-group of circuits connected to a second word line and configured to store a second set of weights; A multiplication circuit; An in-memory computing circuit comprising the accumulation circuit shared by the first sub-group and the second sub-group, wherein the multiplication circuit is configured to: (i) multiply an input by the first set of weights in response to a first timing signal, (ii) provide a first output, (iii) multiply an input by the second set of weights in response to a second timing signal, and (iv) provide a second output, and the multiplication of the second set of weights is enabled at a time different from the time when the multiplication of the first set of weights is enabled; The in-memory computing circuit, wherein the accumulation circuit is configured to: (i) receive and accumulate the first output in response to the multiplication of the first set of weights being enabled by the first timing signal, and (ii) receive and accumulate the second output in response to the multiplication of the second set of weights being enabled by the second timing signal.
Citation Information
Patent Citations
Inner product arithmetic circuit
JP1993324695A
In-memory computation circuit and method
JP2022018112A
Computing device and method using multiplier-accumulator
JP2023039419A
Compute in memory
US20220244916A1
Compute in memory (CIM) memory array
US20220375508A1