Memory circuit and method of forming the same

CN122822003APending Publication Date: 2026-09-25TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610794229.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-11-10
Filing Date
2026-06-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

虽然可以使用传统的计算机硬件在软件中处理大量数据,但现有的计算机硬件对于一些数据处理应用来说效率低下

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822003A_ABST
    Figure CN122822003A_ABST
Patent Text Reader

Abstract

A memory circuit includes a plurality of first multiply-accumulate (MAC) circuit groups physically arranged along a first lateral direction and operable to form a first output lane. Each first MAC circuit group includes a first adder, a first MAC circuit and a second MAC circuit physically sandwiching the first adder along the first lateral direction, and a first memory array and a second memory array physically sandwiching the first MAC circuit and the second MAC circuit along the first lateral direction. Embodiments of the present application also disclose a method of forming a memory circuit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this application relate to memory circuits and methods of forming the same. Background Technology

[0002] With advancements in modern semiconductor manufacturing processes and the ever-increasing volume of data generated daily, the demand for storing and processing massive amounts of data is growing, thus creating an incentive to find improved methods for storing and processing such data. While traditional computer hardware can be used to process large amounts of data in software, existing computer hardware is inefficient for some data processing applications. Summary of the Invention

[0003] According to one aspect of the embodiments of this application, a memory circuit is provided, comprising: a first number A of memory arrays, each of the memory arrays including a plurality of memory cells, each of the plurality of memory cells being configured to store a corresponding weighted data element; a second number B of multiply-accumulate (MAC) circuits; a third number C of adders; a fourth number D of word line (WL) driver circuits; a fifth number E of input drivers; a sixth number F of control circuits; and a seventh number G of input / output (I / O) circuits; wherein the number A, B, C, D, E and G are each configurable based on at least one of the number X of input data elements or the number Z of output channels, wherein the number X and the number Z are each arbitrary positive integers.

[0004] According to another aspect of the embodiments of this application, a memory circuit is provided, including: a plurality of first multiplication-accumulation (MAC) circuit groups physically arranged along a first lateral direction and operably forming a first output channel; wherein each of the first MAC circuit groups includes: a first adder, a first MAC circuit and a second MAC circuit physically sandwiched between the first adder along the first lateral direction, and a first memory array and a second memory array physically sandwiched between the first MAC circuit and the second MAC circuit along the first lateral direction.

[0005] According to another aspect of the embodiments of this application, a method for forming a memory circuit is provided, comprising: arranging a plurality of first multiply-accumulate (MAC) circuit groups along a first lateral direction, wherein each of the plurality of first MAC circuit groups includes: a first adder; a first MAC circuit and a second MAC circuit, physically sandwiching the first adder between them along the first lateral direction; and a first memory array and a second memory array, physically sandwiching the first MAC circuit and the second MAC circuit between them along the first lateral direction; arranging a plurality of input drivers relative to the plurality of first MAC circuit groups along a second lateral direction perpendicular to the first lateral direction, wherein the first input driver and the second input driver among the plurality of input drivers are physically arranged along the second lateral direction relative to the first MAC circuit and the second MAC circuit of a corresponding one in the first MAC circuit groups; and arranging a plurality of word line (WL) driver circuits relative to the plurality of first MAC circuit groups along the second lateral direction, wherein the first WL driver circuit and the second WL driver circuit among the plurality of WL driver circuits are physically arranged along the second lateral direction relative to the first memory array and the second memory array of a corresponding one in the first MAC circuit groups. Attached Figure Description

[0006] The various aspects of this disclosure are best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be emphasized that, in accordance with standard industry practice, the various parts are not drawn to scale and are for illustrative purposes only. In fact, the dimensions of the various parts may be arbitrarily increased or decreased for clarity of discussion.

[0007] Figure 1 An example neural network according to some embodiments is shown.

[0008] Figure 2 An example planar layout diagram of a memory circuit according to some embodiments is shown.

[0009] Figure 3 , Figure 4 , Figure 5 , Figure 6 , Figure 7 , Figure 8 , Figure 9 , Figure 10 , Figure 11 and Figure 12 Example plan views of memory circuits according to some embodiments are shown.

[0010] Figure 13 An example flowchart of a method for forming a memory circuit according to some embodiments is shown.

[0011] Figure 14 Example computer systems for implementing various embodiments of the present disclosure are shown.

[0012] Figure 15 The process of forming a standard cell structure based on a Graphical Database System (GDS) file according to some embodiments is illustrated.

[0013] Figure 16 These are diagrams of example computer systems according to some embodiments that can implement various embodiments of the present disclosure.

[0014] Figure 17 This is a diagram of an exemplary method for circuit fabrication according to some embodiments. Detailed Implementation

[0015] The following disclosure provides numerous different embodiments or examples for implementing various features of this disclosure. Specific embodiments or examples of components and arrangements are described below to simplify this disclosure. Of course, these are merely examples and not intended to be limiting. For example, in the following description, forming a first component above or on a second component can include embodiments where the first and second components are in direct contact, and can also include embodiments where an additional component can be formed between the first and second components, such that the first and second components are not in direct contact. Furthermore, reference numerals and / or letters may be repeated in various examples. This repetition is for simplicity and clarity and does not in itself indicate a relationship between the various embodiments and / or configurations discussed.

[0016] Furthermore, for ease of description, this document may use spacing terms such as “below,” “under,” “lower,” “above,” “upper,” etc., to describe the relationship between one element or component and another, as shown in the figures. In addition to the orientations shown in the figures, spacing terms are intended to include different orientations of the device during use or operation. The device may be positioned in other ways (rotated 90 degrees or in other orientations), and the spacing descriptors used herein may be interpreted accordingly.

[0017] In this regard, machine learning has become an effective method for analyzing and extracting value from such massive amounts of data. Generally speaking, machine learning is a field of computer science that involves algorithms that allow computers to "learn" (e.g., improve the performance of a task) without explicit programming. Machine learning can involve different techniques for analyzing data to improve a task. One such technique, such as deep learning, is based on neural networks. However, machine learning performed on traditional computer systems can involve excessive data transfer between memory and processor, resulting in high power consumption and slow computation time.

[0018] Computation in memory (CIM) (also known as processing in memory) involves performing computational operations within a memory array. In other words, computational operations are performed directly on data read from memory cells, rather than transferring data to a digital processor for processing. By avoiding transferring some data to the digital processor, the bandwidth limitations associated with transferring data back and forth between the processor and memory in traditional computer systems are reduced.

[0019] One application of this type of CIM is artificial intelligence (AI), particularly machine learning. For example, a computing system (such as a CIM system) can use multi-layered computing nodes, where lower layers perform calculations based on the results of calculations performed by higher layers. These calculations may sometimes rely on the calculation of dot products and absolute differences of vectors, typically performed by performing MAC (operations) on parameters, input data, and weights. The term "MAC" can refer to multiplication-accumulation, multiplication / accumulation, or multiplication accumulator, and generally refers to an operation involving the multiplication of two values ​​and the accumulation of a series of multiplications.

[0020] In machine learning applications, CIM systems are often configured to process dot product multiplications of large numbers of data elements (e.g., input data elements and weight data elements), followed by addition (or accumulation) of these dot products. Traditionally, CIM systems configured to produce MAC results of data elements are often limited to a fixed configuration. For example, existing CIM systems are designed to retrieve (or input) a fixed number of input data elements, retrieve (or input) a fixed number of weight data elements, or generate (or output) a fixed number of MAC results. When any of these numbers change, the corresponding configuration of the CIM system needs to be changed accordingly. This can adversely affect the flexibility of CIM system design. Therefore, existing CIM systems are not entirely satisfactory in some aspects.

[0021] This disclosure provides various embodiments of a computation in memory (CIM) system or circuit capable of processing multiple input data elements and multiple weighted data elements. As disclosed herein, the CIM circuit can perform in-memory computations (e.g., multiplication-accumulation (MAC) operations) on a configurable number of input data elements and a configurable number of weighted data elements to generate a configurable number of MAC results. In some embodiments, the disclosed CIM circuit can be flexibly configured or otherwise constructed from seven circuitry or functional components based on the number of input data elements, the number of weighted data elements, and the number of MAC results. For example, these seven functional components may include multiple memory arrays, multiple MAC circuits, multiple adders, multiple word line (WL) driver circuits, multiple input drivers, multiple input / output (I / O) circuits, and multiple control circuits.

[0022] As a non-limiting example, the disclosed CIM circuit may include multiple first MAC circuit groups and multiple second MAC circuit groups physically arranged on opposite sides of corresponding I / O circuits. Each of the first / second MAC circuit groups may include (e.g., internally) an adder sandwiched between a pair of MAC circuits, which in turn are sandwiched between a pair of memory arrays. This physically arranged first / second MAC circuit groups and corresponding I / O circuits can be operatively used as at least a portion of one of multiple output channels of the CIM circuit. By arranging these multiple output channels relative to multiple input drivers, multiple WL driver circuits, and control circuitry, the disclosed CIM circuit may have a number of its MAC results, a number of its processable input data elements, and a number of its processable weighted data elements, all of which can be configured based on at least one of the number of first MAC circuit groups, the number of second MAC circuit groups, and the number of I / O circuits. This flexible “configurability” advantageously simplifies the design of CIM circuits, which can have a variable number of input data elements, a variable number of weighted data elements, etc.

[0023] Figure 1 An example neural network 100 according to various embodiments is shown. As shown, the neural network 100 includes four layers 110, 220, 130, and 140, where layers 110 and 140 are referred to as the input layer and output layer, respectively, and layers 220 to 130 are all referred to as hidden layers. Each layer may contain multiple neurons. Generally, the hidden layers of the neural network 100 can be roughly considered as neuron layers, each neuron layer receiving (e.g., weighted) outputs from neurons in the preceding layer in a mesh-like interconnection structure between layers. Connections from the output of a particular preceding neuron to the input of another subsequent neuron are set according to the influence or effect of the preceding neuron on the subsequent neuron (for simplicity, only one neuron 101 and connection are labeled). Figure 1 In the illustrative example, the output value of the preceding neuron is multiplied by its connection weight with the following neuron to determine the specific stimulus presented by the preceding neuron to the following neuron.

[0024] The total input stimulus to a neuron corresponds to the total stimulus across all its weighted input connections. According to various implementations, if the total input stimulus to a neuron exceeds a certain threshold, the neuron is triggered to perform some kind of linear or nonlinear mathematical function on its input stimulus. The output of the mathematical function corresponds to the neuron's output, which is then multiplied by the corresponding weights of the neuron's output connections to its subsequent neurons.

[0025] Generally, the more connections between neurons, the more neurons per layer, and / or the more layers of neurons, the higher the intelligence the network can achieve. Therefore, neural networks used in practical, real-world artificial intelligence applications are typically characterized by a large number of neurons and numerous connections between them. Consequently, processing information through neural networks involves a significant amount of computation (not only of neuron output functions but also of weighted connections).

[0026] As mentioned above, although neural networks can be implemented entirely in software as program code instructions that execute on one or more traditional general-purpose central processing unit (CPU) or graphics processing unit (GPU) cores, the read / write activity between the CPU / GPU cores and system memory required to perform all the computations is very intensive. In the millions or billions of computations required to implement neural networks, the overhead and energy associated with repeatedly reading large amounts of data from system memory, processing that data by the CPU / GPU cores, and then writing the results back to system memory are, in some respects, not entirely satisfactory.

[0027] Figure 2 An example block diagram is shown of a memory system or memory circuit 200 employing CIM technology suitable for implementing various embodiments. In some embodiments, Figure 2 The block diagram can correspond to the physical layout (or floor plan) of a memory circuit 200 that can be constructed from seven functional components. Each of these functional components is associated with... Figure 2 The reference figures are related and used for illustrative purposes only. It should be understood that... Figure 2 The block diagram has been simplified, so the memory circuit 200 may include any of a variety of other components (e.g., a global adder, which will be discussed below) while still remaining within the scope of this disclosure.

[0028] like Figure 2 As shown, the memory circuit 200 can be composed of functional components 211, 212, 213, 214, 215, 216, and 217. In some embodiments, components 211 to 217 can respectively represent a memory array, a MAC circuit, a local adder, an I / O circuit, a WL driver circuit, an input driver, and a control circuit. Hereinafter, components 211 to 217 can be referred to as memory array 211, MAC circuit 212, local adder 213, I / O circuit 214, WL driver circuit 215, input driver 216, and control circuit 217. Figure 3As shown, according to some embodiments, each local adder 213 may be sandwiched between a pair of MAC circuits 212, and the MAC circuits 212 may be further sandwiched between a pair of memory arrays 211. This physical arrangement of a local adder 213, a pair of MAC circuits 212, and a pair of memory arrays 211 can operatively form a MAC circuit set (MCS).

[0029] Refer again Figure 2 Multiple MCSs can be arranged along a first lateral direction (e.g., the X direction). Furthermore, a first number of MCSs and a second number of MCSs can be arranged along the X direction on opposite sides of the I / O circuit 214. Such MCSs, together with their respective I / O circuits 214, can operatively form output channels. In some embodiments, multiple output channels can be arranged along a second lateral direction (e.g., the Y direction) relative to multiple WL driver circuits 215, multiple input drivers 216, and control circuitry 217. Furthermore, each WL driver circuit 215 can be arranged along the Y direction relative to a corresponding subset of the memory array 211 (across output channels); each input driver 216 can be arranged along the Y direction relative to a corresponding subset of the MAC circuitry (across output channels); and the control circuitry 217 can be arranged along the Y direction relative to the I / O circuitry 214 (across output channels).

[0030] Control circuit 217 can accurately and timely generate signals for the operation of memory circuit 200. For example, control circuit 217 can generate or receive a clock (CLK) signal that oscillates periodically between rising and falling edges. The CLK signal can be provided to I / O circuits 214 simultaneously or individually. In another example, control circuit 217 can generate a control (CTRL) signal to activate one or more I / O circuits 214 while deactivating the remaining I / O circuits 214. Typically, based at least on the CLK and CTRL signals, each I / O circuit 214 can receive a MAC result, which can be the sum of multiple intermediate MAC results from the corresponding MCS, and output the MAC result through the corresponding output channel.

[0031] Although not shown, control circuitry 217 can also control the operation of WL driver circuitry 215 and input driver 216. For example, WL driver circuitry 215 can assert at least one WL for each coupled memory array 211 based on a corresponding control signal (e.g., a decoded address signal) received from control circuitry 217. Data elements (e.g., weighted data elements) stored in memory cells of memory array 211 coupled to the asserted WL can be read out or otherwise received by a corresponding one of MAC circuits 212. Before, simultaneously with, or after receiving weighted data elements from memory array 211, for example via bit lines (BL), MAC circuitry 212 can receive multiple input data elements (XIN) from the corresponding input driver 216. In some embodiments, each memory array 211 may have WL and BL extending along the Y and X directions, respectively. In other words, WL may be arranged parallel to each other along the X direction, and BL may be arranged parallel to each other along the Y direction.

[0032] use Figure 2 The MCS shown is a representative example. At least one memory level (WL) of memory array 211A can be asserted by WL driver circuit 215A, and at least one memory level (WL) of memory array 212B can be asserted by WL driver circuit 215B. Therefore, MAC circuit 212A can receive or retrieve weighted data elements stored in corresponding memory cells of memory array 211A coupled to the WL asserted by WL driver circuit 215A; and MAC circuit 212B can receive or retrieve weighted data elements stored in corresponding memory cells of memory array 211B coupled to the WL asserted by WL driver circuit 215B. Furthermore, MAC circuit 212A can receive multiple input data elements (e.g., four input data elements) from input driver 216A; and MAC circuit 212B can receive multiple input data elements (e.g., four input data elements) from input driver 216B.

[0033] After receiving weighted data elements and input data elements from memory array 211A and input driver 216A respectively, MAC circuit 212A can perform MAC operations on the received weighted data elements and input data elements. For example, MAC circuit 212A can multiply each input data element by the corresponding weighted data element to generate a corresponding partial product, and then sum the partial products as an intermediate MAC result. Similarly, after receiving weighted data elements and input data elements from memory array 211B and input driver 216B respectively, MAC circuit 212B can perform MAC operations on the received weighted data elements and input data elements. For example, MAC circuit 212B can multiply each input data element by the corresponding weighted data element to generate a corresponding partial product, and then sum all partial products as another intermediate MAC result. Then, local adder 213A can sum the intermediate MAC results to generate an intra-level MAC result. In this embodiment, where each of the MAC circuits 212A and 212B can process or otherwise receive up to four input data elements, each MCS can process up to eight input data elements simultaneously. In other words, each MCS can simultaneously provide up to eight intra-level MAC results.

[0034] The term "intra-level MAC result" can refer to the MAC result generated by the local adder 213 of the MCS (MAC circuitry group). As discussed below, multiple global adders can be tiled or otherwise inserted into appropriate locations (e.g., output channels) between different MCSs. According to some embodiments of this disclosure, these global adders can each sum or forward the results of different intra-level MAC results to generate inter-level MAC results. The term "intra-level MAC result" can refer to the MAC result output by the local adder of the MCS, which can be a single intra-level MAC result or a combination of multiple different intra-level MAC results, and then processed by the global adder (e.g., forwarded or combined with one or more other intra-level MAC results). The I / O circuitry 214 can then sum multiple such intra-level MAC results to provide MAC results through the corresponding output channels. Furthermore, each global adder can receive MAC results through multiple global adders. Based on the number (n) of global adders, the global adder can be referred to as the (n+1)th level adder.

[0035] based on Figure 2The layout diagram shown illustrates that any memory circuit employing CIM technology can be constructed using a corresponding configurable number of memory arrays 211, MAC circuits 212, local adders 213, I / O circuits 214, WL driver circuits 215, input drivers 216, and control circuits 217. The memory circuit may include “A” memory arrays 211, “B” MAC circuits 212, “C” local adders 213, “D” WL driver circuits, “E” input drivers, “F” control circuits 217, and “G” I / O circuits 214. In some embodiments, the numbers A, B, C, D, E, and G can all be configured based on at least one of the number of input data elements (X) or the number of output channels (Z). In other words, both the number X and the number Z can be any positive integer (e.g., specified or changed by the customer). Even so, the designer can... Figure 2 (and Figures 4-14 The floor plan adapts effortlessly to these specified or updated quantities.

[0036] Refer again Figure 2 In a planar layout, for example, the number of subsets of A memory arrays 211 coupled to (or physically positioned relative to the corresponding WL driver circuit 215 along the Y direction) of D WL driver circuits 215 is equal to the number Z. In another example, the number of subsets of B MAC circuits 212 coupled to (or physically positioned relative to the corresponding input driver 216 along the Y direction) of E input drivers 216 is equal to the number Z. In yet another example, the number of subsets of A memory arrays 211 coupled to or belonging to (or physically positioned relative to the WL driver circuit 215 along the Y direction) of Z output channels is equal to the number D. In yet another example, the number of subsets of B MAC circuits 212 coupled to or belonging to (or physically positioned relative to the input driver 216 along the Y direction) of Z output channels is equal to the number E. In another example, each of the input drivers 216 can provide a subset of X input data elements (e.g., 4 input data elements) to a subset of B MAC circuits 212 coupled to Z output channels. Furthermore, each of the D WL driver circuits 215 can include a plurality of (Y) WL drivers configured to activate or assert Y WLs for each of A subsets of memory arrays, wherein the subsets of memory arrays are coupled to Z output channels respectively. Additionally, each of the C local adders 213 can sum the intra-stage MAC results provided by the first and second subsets of the B MAC circuits 212, wherein the first and second MAC circuits are coupled to a corresponding one of the Z output channels.

[0037] Figure 4 , Figure 5 , Figure 6 , Figure 7 , Figure 8 , Figure 9 , Figure 10 , Figure 11 , Figure 12 , Figure 13 and Figure 14 Various example plan views (or layouts) of memory circuits 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, and 1400 according to some embodiments are shown respectively. Figures 4 to 14 In the examples, each of the memory circuits 400 to 1400 includes an output channel (Z=1) with the ability to process a corresponding number (X) of input data elements (XIN). For example, memory circuits 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, and 1400 can process 8, 16, 24, 32, 40, 48, 56, 64, 72, 80, and 144 input data elements, respectively.

[0038] First refer to Figure 4 The memory circuit 400 may include an MCS 410 coupled to the corresponding I / O circuit 420 (which includes a local adder, a pair of MAC circuits, and a pair of memory arrays, such as...). Figure 3 (As shown). MCS 410 can process up to 8 input data elements. In some embodiments, since MCS 410 is located on one side of I / O circuit 420 and does not need to process more than 8 input data elements, MCS may not be located on the other side of I / O circuit 420. MCS 410 (or its local adder) can provide one or more intra-stage MAC results to global adder 430, as shown by symbolic arrow 401. Global adder 430 can provide these intra-stage MAC results to I / O circuit 420, as shown by symbolic arrow 403. Furthermore, when it is not necessary to process more than 8 input data elements, I / O circuit 420 can receive a signal of all zeros from the other side, as shown by symbolic arrow 405.

[0039] Next reference Figure 5 The memory circuit 500 may include two MCSs 510 and 520 coupled to the corresponding I / O circuits 530 (each including a local adder, a pair of MAC circuits, and a pair of memory arrays, such as...). Figure 3(As shown). MCS 510 and 520 can each process up to 8 input data elements. In some embodiments, since MCS 510 and 520 are located on one side of I / O circuit 530 and do not need to process more than 16 input data elements, an MCS may not be located on the other side of I / O circuit 520. MCS 510 (or its local adder) can provide one or more first-level intra-MAC results to global adder 540, as shown by symbolic arrow 501; and MCS 520 (or its local adder) can provide one or more second-level intra-MAC results to global adder 542, as shown by symbolic arrow 503.

[0040] Global adder 542 can provide these second-stage intra-MAC results (received from MCS 520) to global adder 544, as shown by symbolic arrow 505. These results are further provided to global adder 546, as shown by symbolic arrow 507, and further to global adder 540, as shown by symbolic arrow 509. Global adder 540 can then sum the first-stage intra-MAC results and the second-stage intra-MAC results (received via arrows 501 and 509, respectively) and provide the sum to I / O circuitry 530, as shown by symbolic arrow 511. In some embodiments, global adders 542, 544, 546, and 540 may sometimes be referred to as first-stage adder, second-stage adder, third-stage adder, and fourth-stage adder, respectively. Furthermore, since it is not necessary to process more than 16 input data elements, I / O circuitry 530 can receive a signal of all zeros from the other side, as shown by symbolic arrow 513.

[0041] Next reference Figure 6 The memory circuit 600 may include three MCSs 610, 620, and 630 coupled to the corresponding I / O circuits 640 (each including a local adder, a pair of MAC circuits, and a pair of memory arrays, such as...). Figure 3 (As shown). MCS 610, 620, and 630 can each process up to 8 input data elements. In some embodiments, if MCS 610 to 630 are located on one side of I / O circuit 640 and it is not necessary to process more than 24 input data elements, an MCS may not be located on the other side of I / O circuit 640. MCS 610 (or its local adder) can provide one or more first-level intra-MAC results to global adder 650, as shown by symbolic arrow 601; MCS 620 (or its local adder) can provide one or more second-level intra-MAC results to global adder 652, as shown by symbolic arrow 603; and MCS 630 (or its local adder) can provide one or more third-level intra-MAC results to global adder 652, as shown by symbolic arrow 605.

[0042] Global adder 652 sums the MAC results within the second and third stages (received via arrows 603 and 605, respectively) and provides the sum to global adder 654, as shown by arrow 607, further to global adder 656, as shown by arrow 609, and further to global adder 650, as shown by arrow 611. Global adder 650 then adds the MAC result within the first stage (received via arrow 601) to the sum of the MAC results within the second and third stages (received via arrow 611) and provides the sum to I / O circuitry 640, as shown by arrow 613. In some embodiments, global adders 652, 654, 656, and 650 may sometimes be referred to as first-stage adder, second-stage adder, third-stage adder, and fourth-stage adder, respectively. Furthermore, when it is not necessary to process more than 24 input data elements, I / O circuitry 640 can receive a signal of all zeros from the other side, as shown by arrow 615.

[0043] First refer to Figure 7 The memory circuit 700 may include four MCSs 710, 720, 730, and 740 coupled to the corresponding I / O circuits 750 (each including a local adder, a pair of MAC circuits, and a pair of memory arrays, such as...). Figure 3 (As shown). MCS 710, 720, 730, and 740 can each process up to 8 input data elements. In some embodiments, if MCS 710 to 740 are located on one side of I / O circuit 750 and it is not necessary to process more than 32 input data elements, an MCS may not be located on the other side of I / O circuit 750. MCS 710 (or its local adder) can provide one or more first-level intra-MAC results to global adder 760, as shown by symbolic arrow 701; MCS 720 (or its local adder) can provide one or more second-level intra-MAC results to global adder 762, as shown by symbolic arrow 703; MCS 730 (or its local adder) can provide one or more third-level intra-MAC results to global adder 762, as shown by symbolic arrow 705; and MCS 740 (or its local adder) can provide one or more fourth-level intra-MAC results to global adder 766, as shown by symbolic arrow 707.

[0044] Global adder 762 can sum the MAC results of the second and third stages (received via arrows 703 and 705 respectively) and provide the sum to global adder 764, as shown by arrow 709. Global adder 764 can also receive the MAC result of the fourth stage from global adder 766, as shown by arrow 711. Global adder 764 can add the sums of the MAC results of the second and third stages (received via arrow 709) and the MAC result of the fourth stage (received via arrow 711), and provide the sum of the MAC results of the second to fourth stages to global adder 768, as shown by arrow 713, and further provide it to global adder 760, as shown by arrow 715. Then, global adder 760 can add the sums of the MAC results of the first stage (received via arrow 701) and the MAC results of the second to fourth stages (received via arrow 715), and provide the sum of the MAC results of the first to fourth stages to I / O circuit 750, as shown by arrow 717. In some embodiments, the global adders 762 / 766, 764, 768, and 760 may sometimes be referred to as first-stage adders, second-stage adders, third-stage adders, and fourth-stage adders, respectively. Furthermore, since it is not necessary to process more than 32 input data elements, the I / O circuit 750 can receive a signal of all zeros from the other side, as indicated by symbol arrow 719.

[0045] Next reference Figure 8 The memory circuitry 800 may include five MCSs 810, 820, 830, 840, and 850 coupled to the corresponding I / O circuitry 860 (each including a local adder, a pair of MAC circuits, and a pair of memory arrays, such as...). Figure 3 (As shown). MCS 810, 820, 830, 840 and 850 can each process up to 8 input data elements. In some embodiments, since MCS 810 to 850 are located on one side of I / O circuit 860 and do not need to process more than 40 input data elements, an MCS may not be located on the other side of I / O circuit 860. MCS 810 (or its local adder) can provide one or more first-level intra-MAC results to global adder 870, as shown by symbolic arrow 801; MCS 820 (or its local adder) can provide one or more second-level intra-MAC results to global adder 872, as shown by symbolic arrow 803; MCS 830 (or its local adder) can provide one or more third-level intra-MAC results to global adder 872, as shown by symbolic arrow 805; MCS 840 (or its local adder) can provide one or more fourth-level intra-MAC results to global adder 876, as shown by symbolic arrow 807; and MCS 850 (or its local adder) can provide one or more fifth-level intra-MAC results to global adder 876, as shown by symbolic arrow 809.

[0046] Global adder 872 can add the MAC results within the second and third levels (received via arrows 803 and 805 respectively), and provide the sum of the MAC results within the second and third levels to global adder 874, as shown by symbolic arrow 811. Global adder 876 can add the MAC results within the fourth and fifth levels (received via arrows 807 and 809 respectively), and provide the sum of the MAC results within the fourth and fifth levels to global adder 874, as shown by symbolic arrow 813. Global adder 874 can add the sum of the MAC results within the second and third levels (received via arrow 811) to the sum of the MAC results within the fourth and fifth levels (received according to arrow 813), and provide the sum of the MAC results within the second to fifth levels to global adder 878, as shown by symbolic arrow 815, and further provide it to global adder 870, as shown by symbolic arrow 817. The global adder 870 can then add the MAC results within the first stage (received via arrow 801) and the MAC results within the second to fifth stages (received via arrow 817), and provide the sum to the I / O circuit 860, as indicated by symbol arrow 819. In some embodiments, the global adders 872 / 876, 874, 878, and 870 may sometimes be referred to as the first-stage adder, second-stage adder, third-stage adder, and fourth-stage adder, respectively. Furthermore, when it is not necessary to process more than 40 input data elements, the I / O circuit 860 can receive a signal of all zeros from the other side, as indicated by symbol arrow 821.

[0047] Next reference Figure 9 The memory circuitry 900 may include six MCSs 910, 920, 930, 940, 950, and 960 (each including a local adder, a pair of MAC circuits, and a pair of memory arrays) coupled to the corresponding I / O circuitry 970. Figure 3(As shown). MCSs 910, 920, 930, 940, 950, and 960 can each handle up to 8 input data elements. In some embodiments, if MCSs 910 to 960 are located on one side of I / O circuit 970 and it is not necessary to handle more than 48 input data elements, an MCS may not be located on the other side of I / O circuit 970. MCS 910 (or its local adder) can provide one or more first-level intra-MAC results to global adder 980, as shown by symbolic arrow 901; MCS 920 (or its local adder) can provide one or more second-level intra-MAC results to global adder 982, as shown by symbolic arrow 903; MCS 930 (or its local adder) can provide one or more third-level intra-MAC results to global adder 982, as shown by symbolic arrow 905; MCS 940 (or its local adder) can provide one or more fourth-level intra-MAC results to global adder 986, as shown by symbolic arrow 907; MCS 950 (or its local adder) can provide one or more fifth-level intra-MAC results to global adder 986, as shown by symbolic arrow 909; and MCS 960 (or its local adder) can provide one or more sixth-level intra-MAC results to global adder 990, as shown by symbolic arrow 911.

[0048] Global adder 982 can add the MAC results of the second and third levels (received via arrows 903 and 905 respectively), and provide the sum of the MAC results of the second and third levels to global adder 984, as shown by symbolic arrow 913. Global adder 986 can add the MAC results of the fourth and fifth levels (received via arrows 907 and 909 respectively), and provide the sum of the MAC results of the fourth and fifth levels to global adder 984, as shown by symbolic arrow 915. Global adder 984 can add the sum of the MAC results of the second and third levels (received via arrow 913) to the sum of the MAC results of the fourth and fifth levels (received via arrow 915), and provide the sum of the MAC results of the second to fifth levels to global adder 988, as shown by symbolic arrow 917. Global adder 988 can also receive the MAC result within the sixth level from global adder 992, as shown by symbolic arrow 919, where global adder 992 can receive the MAC result within the sixth level from global adder 990 beforehand (as shown by symbolic arrow 921). Then, global adder 988 can add the MAC result within the sixth level (received via arrow 919) to the sum of the MAC results within the second to fifth levels (received via arrow 917) and provide the sum to global adder 980, as shown by symbolic arrow 923. Global adder 980 can then add the MAC result within the first level (received via arrow 901) to the sum of the MAC results within the second to sixth levels (received via arrow 923) and provide the sum to I / O circuit 970, as shown by symbolic arrow 925. In some embodiments, global adders 982 / 986 / 990, 984, 988, and 980 may sometimes be referred to as first-level adders, second-level adders, third-level adders, and fourth-level adders, respectively. Furthermore, without needing to process more than 48 input data elements, the I / O circuit 970 can receive a signal of all zeros from the other side, as indicated by symbol arrow 927.

[0049] First refer to Figure 10 The memory circuit 1000 may include seven MCSs 1010, 1020, 1030, 1040, 1050, 1060, and 1070 coupled to the corresponding I / O circuits 1080 (each including a local adder, a pair of MAC circuits, and a pair of memory arrays, such as...). Figure 3(As shown). MCSs 1010, 1020, 1030, 1040, 1050, 1060, and 1070 can each process up to eight input data elements. In some embodiments, since MCSs 1010 to 1070 are located on one side of the I / O circuit 1080 and do not need to process more than 56 input data elements, an MCS may not be located on the other side of the I / O circuit 1080. MCS 1010 (or its local adder) can provide one or more first-level intra-MAC results to global adder 1082, as shown by symbolic arrow 1001; MCS 1020 (or its local adder) can provide one or more second-level intra-MAC results to global adder 1084, as shown by symbolic arrow 1003; MCS 1030 (or its local adder) can provide one or more third-level intra-MAC results to global adder 1084, as shown by symbolic arrow 1005; MCS 1040 (or its local adder) can provide one or more fourth-level intra-MAC results to global adder 1088, as shown by symbolic arrow 1007; MCS 1050 (or its local adder) can provide one or more fifth-level intra-MAC results to global adder 1088, as shown by symbolic arrow 1009; MCS 1060 (or its local adder) can provide one or more MAC results at level 6 to global adder 1092, as indicated by symbolic arrow 1011; and MCS 1070 (or its local adder) can provide one or more MAC results at level 7 to global adder 1092, as indicated by symbolic arrow 1013.

[0050] Global adder 1084 can add the MAC results within the second and third levels (received via arrows 1003 and 1005 respectively), and provide the sum of the MAC results within the second and third levels to global adder 1086, as shown by symbolic arrow 1015. Global adder 1088 can add the MAC results within the fourth and fifth levels (received via arrows 1007 and 1009 respectively), and provide the sum of the MAC results within the fourth and fifth levels to global adder 1086, as shown by symbolic arrow 1017. Global adder 1086 can add the sum of the MAC results within the second and third levels (received via arrow 1015) to the sum of the MAC results within the fourth and fifth levels (received according to arrow 1017), and provide the sum of the MAC results from the second to the fifth levels to global adder 1090, as shown by symbolic arrow 1019. Global adder 1092 can sum the MAC results within the sixth and seventh stages (received via arrows 1011 and 1013 respectively) and provide the sum of the MAC results within the sixth and seventh stages to global adder 1094, as indicated by symbolic arrow 1021. Global adder 1090 can also receive the sum of the MAC results within the sixth and seventh stages from global adder 1094, as indicated by symbolic arrow 1023. Therefore, global adder 1090 can then add the sum of the MAC results within the sixth and seventh stages (received via arrow 1023) to the sum of the MAC effects within the second to fifth stages (received via arrow 1019) and provide the sum to global adder 1082, as indicated by symbolic arrow 1025. Global adder 1082 can then add the MAC results within the first stage (received via arrow 1001) to the sum of the MAC results within the second to seventh stages (received via arrow 1025) and provide the sum to I / O circuit 1080, as indicated by symbolic arrow 1027. In some embodiments, the global adders 1084 / 1088 / 1092, 1086 / 1094, 1090, and 1082 may sometimes be referred to as first-stage adders, second-stage adders, third-stage adders, and fourth-stage adders, respectively. Furthermore, since it is not necessary to process more than 56 input data elements, the I / O circuit 1080 can receive a signal of all zeros from the other side, as indicated by symbolic arrow 1029.

[0051] Next reference Figure 11 The memory circuit 1100 may include eight MCSs 1110, 1112, 1114, 1116, 1118, 1120, 1122, and 1124 (each including a local adder, a pair of MAC circuits, and a pair of memory arrays) coupled to the corresponding I / O circuits 1130. Figure 3(As shown). MCSs 1110, 1112, 1114, 1116, 1118, 1120, 1122, and 1124 can each process up to eight input data elements. In some embodiments, if MCSs 1110 to 1124 are located on one side of the I / O circuit 1130 and it is not necessary to process more than 64 input data elements, an MCS may not be located on the other side of the I / O circuit 1130. MCS 1110 (or its local adder) can provide one or more first-level intra-MAC results to global adder 1132, as shown by symbolic arrow 1101; MCS 1112 (or its local adder) can provide one or more second-level intra-MAC results to global adder 1134, as shown by symbolic arrow 1103; MCS 1114 (or its local adder) can provide one or more third-level intra-MAC results to global adder 1134, as shown by symbolic arrow 1105; MCS 1116 (or its local adder) can provide one or more fourth-level intra-MAC results to global adder 1138, as shown by symbolic arrow 1107; MCS 1118 (or its local adder) can provide one or more fifth-level intra-MAC results to global adder 1138, as shown by symbolic arrow 1109; MCS 1120 (or its local adder) can provide one or more MAC results at level 6 to global adder 1142, as shown by symbolic arrow 1111; MCS 1122 (or its local adder) can provide one or more MAC results at level 7 to global adder 1142, as shown by symbolic arrow 1113; and MCS 1124 (or its local adder) can provide one or more MAC results at level 8 to global adder 1146, as shown by symbolic arrow 1115.

[0052] Global adder 1134 can add the MAC results within the second and third levels (received via arrows 1103 and 1105 respectively), and provide the sum of the MAC results within the second and third levels to global adder 1136, as shown by symbolic arrow 1117. Global adder 1138 can add the MAC results within the fourth and fifth levels (received via arrows 1107 and 1109 respectively), and provide the sum of the MAC results within the fourth and fifth levels to global adder 1136, as shown by symbolic arrow 1119. Global adder 1136 can add the sum of the MAC results within the second and third levels (received via arrow 1117) to the sum of the MAC results within the fourth and fifth levels (received via arrow 1119), and provide the sum of the MAC results from the second to the fifth levels to global adder 1140, as shown by symbolic arrow 1121. Global adder 1142 can add the MAC results within the sixth and seventh levels (received via arrows 1111 and 1113 respectively), and provide the sum of the MAC results within the sixth and seventh levels to global adder 1144, as indicated by symbolic arrow 1123. Global adder 1144 can also receive the MAC result within the eighth level from global adder 1146, as indicated by symbolic arrow 1125. Therefore, global adder 1144 can then add the sum of the MAC results within the sixth and seventh levels (received via arrow 1123) and the MAC result within the eighth level (received via arrow 1125), and provide the sum to global adder 1140, as indicated by symbolic arrow 1127. Global adder 1140 can then add the sum of the MAC results within the second to fifth levels (received via arrow 1121) and the sum of the MAC results within the sixth to eighth levels (received according to arrow 1127), and provide the sum to global adder 1132, as indicated by symbolic arrow 1129. Global adder 1132 can then add the MAC result within the first stage (received via arrow 1101) and the sum of the MAC results within the second to eighth stages (received via arrow 1129), and provide the sum to I / O circuitry 1130, as indicated by symbolic arrow 1131. In some embodiments, global adders 1134 / 1138 / 1142 / 1146, 1136 / 1144, 1140, and 1132 may sometimes be referred to as first-stage adder, second-stage adder, third-stage adder, and fourth-stage adder, respectively. Furthermore, when it is not necessary to process more than 64 input data elements, I / O circuitry 1130 can receive a signal of all zeros from the other side, as indicated by symbolic arrow 1133.

[0053] Next reference Figure 12 The memory circuit 1200 may include nine MCSs 1210, 1212, 1214, 1216, 1218, 1220, 1222, 1224 and 1226 coupled to the corresponding I / O circuits 1130 (each including a local adder, a pair of MAC circuits and a pair of memory arrays, such as...). Figure 3 (As shown). MCSs 1210, 1212, 1214, 1216, 1218, 1220, 1222, 1224, and 1226 can each process up to 8 input data elements. In some embodiments, since MCSs 1210 to 1226 are located on one side of I / O circuit 1228 and do not need to process more than 72 input data elements, an MCS may not be located on the other side of I / O circuit 1228. MCS 1210 (or its local adder) can provide one or more first-level intra-MAC results to global adder 1230, as shown by symbolic arrow 1201; MCS 1212 (or its local adder) can provide one or more second-level intra-MAC results to global adder 1232, as shown by symbolic arrow 1203; MCS 1214 (or its local adder) can provide one or more third-level intra-MAC results to global adder 1232, as shown by symbolic arrow 1205; MCS 1216 (or its local adder) can provide one or more fourth-level intra-MAC results to global adder 1236, as shown by symbolic arrow 1207; MCS 1218 (or its local adder) can provide one or more fifth-level intra-MAC results to global adder 1236, as shown by symbolic arrow 1209; MCS MCS 1220 (or its local adder) can provide one or more MAC results at level 6 to global adder 1240, as shown by symbolic arrow 1211; MCS 1222 (or its local adder) can provide one or more MAC results at level 7 to global adder 1240, as shown by symbolic arrow 1213; MCS 1224 (or its local adder) can provide one or more MAC results at level 8 to global adder 1244, as shown by symbolic arrow 1215; and MCS 1226 (or its local adder) can provide one or more MAC results at level 9 to global adder 1244, as shown by symbolic arrow 1217.

[0054] Global adder 1232 can add the MAC results within the second and third levels (received via arrows 1203 and 1205 respectively), and provide the sum of the MAC results within the second and third levels to global adder 1234, as shown by symbolic arrow 1219. Global adder 1236 can add the MAC results within the fourth and fifth levels (received via arrows 1207 and 1209 respectively), and provide the sum of the MAC results within the fourth and fifth levels to global adder 1234, as shown by symbolic arrow 1221. Global adder 1234 can add the sum of the MAC results within the second and third levels (received via arrow 1219) to the sum of the MAC results within the fourth and fifth levels (received via arrow 1221), and provide the sum of the MAC results from the second to the fifth levels to global adder 1238, as shown by symbolic arrow 1223. Global adder 1240 can add the MAC results within the sixth and seventh levels (received via arrows 1211 and 1213 respectively), and provide the sum of the MAC results within the sixth and seventh levels to global adder 1242, as indicated by symbolic arrow 1225. Global adder 1244 can add the MAC results within the eighth and ninth levels (received via arrows 1215 and 1217 respectively), and provide the sum of the MAC results within the eighth and ninth levels to global adder 1242, as indicated by symbolic arrow 1227. Therefore, global adder 1242 can then add the sum of the MAC results within the sixth and seventh levels (received via arrow 1225) to the sum of the MAC results within the eighth and ninth levels (received from arrow 1227), and provide the sum to global adder 1238, as indicated by symbolic arrow 1229. Then, global adder 1238 can add the sum of the MAC results in the second to fifth stages (received via arrow 1223) to the sum of the MAC results in the sixth to ninth stages (received via arrow 1229), and provide the sum to global adder 1230, as indicated by symbolic arrow 1231. Global adder 1230 can then add the MAC results in the first stage (received via arrow 1201) to the sum of the MAC results in the second to ninth stages (received via arrow 1231), and provide the sum to I / O circuit 1228, as indicated by symbolic arrow 1233. In some embodiments, global adders 1232 / 1236 / 1240 / 1244, 1234 / 1242, 1238, and 1230 may sometimes be referred to as first-stage adders, second-stage adders, third-stage adders, and fourth-stage adders, respectively. Furthermore, when it is not necessary to process more than 72 input data elements, the I / O circuit 1228 can receive a signal of all zeros from the other side, as shown by symbol arrow 1235.

[0055] Next reference Figure 13 The memory circuit 1300 may include the memory circuit 1200 ( Figure 12), and also includes the MCS 1310 (including a local adder, a pair of MAC circuits and a pair of memory arrays, such as Figure 3 (As shown) and a global adder 1320 arranged between I / O circuit 1228 and MCS 1310. MCS 1310 can similarly handle up to 8 input data elements. In some embodiments, the MCS of memory circuit 1200 is located on one side of I / O circuit 1228, and MCS 1310 is located on the other side of I / O circuit 1228, allowing memory circuit 1300 to handle up to 80 input data elements. (Continued discussion) Figure 12 The MCS1310 (or its local adder) can provide one or more 10th-level intra-MAC results to the global adder 1320, as indicated by symbolic arrow 1301. The global adder 1320 can provide the 10th-level intra-MAC results to the I / O circuit 1228.

[0056] Next reference Figure 14 The memory circuit 1400 may include memory circuit 1200 other than the I / O circuit 1228. Figure 12 The first of the series, for example, 1200A, and the memory circuit 1200 other than the I / O circuit 1228. Figure 12 The second one in the series, for example, 1200B. As described above, each of the memory circuits 1200A and 1200B can process up to 72 input data elements. The memory circuit 1400 also includes I / O circuit 1410, with memory circuits 1200A and 1200B arranged on opposite sides of I / O circuit 1410. Therefore, the memory circuit 1400 can process up to 144 input data elements.

[0057] Figure 15 A flowchart of a method 1500 for forming a memory circuit using CIM technology according to one or more embodiments of the present disclosure is shown. For example, at least some operations (or steps) of method 1500 can be used to form one of the aforementioned memory circuits based on a corresponding layout (or arrangement). However, it should be noted that method 1500 is merely an example and is not intended to limit the present disclosure. Therefore, it is understood that... Figure 1 Additional operations can be provided before, during, and after method 1500. And only a few of these additional operations will be briefly described here.

[0058] Method 1500 begins with operation 1510, arranging a plurality of first MAC circuit groups (first MCS) along a first lateral direction, wherein each of the first MAC circuit groups includes: a first adder; a first MAC circuit and a second MAC circuit, physically clamping the first adder along the first lateral direction; and a first memory array and a second memory array, physically clamping the first MAC circuit and the second MAC circuit along the first lateral direction.

[0059] Method 1500 proceeds to operation 1520, arranging a plurality of second MAC circuit groups (second MCS) along a first lateral direction, wherein each of the second MAC circuit groups includes: a second local adder, a third MAC circuit and a fourth MAC circuit that physically clamp the second local adder along the first lateral direction, and a third memory array and a fourth memory array that physically clamp the third MAC circuit and the fourth MAC circuit along the first lateral direction.

[0060] Using a planar layout of memory circuit 200 as a representative example, a first MAC circuit group may be disposed on one side of the first I / O circuit 214 along the X direction, and a second MAC circuit group may be disposed on one side of the second I / O circuit 214 along the X direction. The first and second I / O circuits 214 may be arranged opposite to each other along the Y direction. In some embodiments, the first MAC circuit group is operatively configured to form at least a portion of a first output channel, and the second MAC circuit group is operatively configured to form at least a portion of a second output channel.

[0061] Method 1500 continues to operation 1530, where a plurality of input drivers are arranged relative to the first and second MAC circuit groups along a second lateral direction perpendicular to the first lateral direction. In some embodiments, the first and second of the plurality of input drivers are physically arranged along the second lateral direction relative to a corresponding first MAC circuit and a corresponding second MAC circuit in the first MAC circuit group, and are physically arranged along the second lateral direction relative to a corresponding third MAC circuit and a corresponding fourth MAC circuit in the second MAC circuit group, respectively.

[0062] Method 1500 continues to operation 1540, whereby a plurality of word line (WL) driver circuits are physically arranged relative to the first and second MAC circuits along a second lateral direction. In some embodiments, the first and second WL driver circuits of the plurality of WL driver circuits are respectively physically arranged in a first memory array and a second memory array of a corresponding one of the physically arranged relative to the first MAC circuit along the second lateral direction, and respectively in a third memory array and a fourth memory array of a corresponding one of the physically arranged relative to the second MAC circuit along the second lateral direction.

[0063] Continuing the example above, the first MAC circuit and the second MAC circuit of each first MAC circuit group can be aligned with the first and second input drivers 216 along the Y direction, respectively; and the first memory array and the second memory array of each first MAC circuit group can be aligned with the first and second WL drivers 215 along the Y direction, respectively. Each first MAC circuit group can be aligned with a corresponding one of the second MAC circuit groups along the Y direction. Furthermore, the third MAC circuit and the fourth MAC circuit of the corresponding second MAC circuit group can be aligned with the first input driver 216 and the second input driver 216 along the Y direction, respectively; and the third memory array and the fourth memory array of the corresponding second MAC circuit group can be aligned with the first WL driver 215 and the second WL driver 215 along the Y direction, respectively.

[0064] Figure 16 This is an illustration of an example computer system 1600 according to some embodiments, in which various embodiments of the present disclosure may be implemented. Computer system 1600 may be any known computer capable of performing the functions and operations described herein. For example, but not limited to, computer system 1600 may be capable of selecting various standard units to be optimized and placing these standard units in desired locations, such as EDA tools. For example, computer system 1600 may be used to perform one or more operations in method 1500.

[0065] Computer system 1600 includes one or more processors (also referred to as central processing units or CPUs), such as processor 1604. Processor 1604 is connected to communication infrastructure or bus 1606. Computer system 1600 also includes input / output devices 1603, such as monitors, keyboards, pointing devices, etc., which communicate with the communication infrastructure or bus 1606 via input / output interfaces 1602. EDA tools can receive instructions through input / output devices 1603 to perform the functions and operations described herein, such as... Figure 15 Method 1500. The computer system 1600 also includes main memory or primary memory 1608, such as random access memory (RAM). Main memory 1608 may include one or more levels of cache. Control logic (e.g., computer software) and / or data are stored in main memory 1608. In some embodiments, control logic (e.g., computer software) and / or data may include the above-described... Figure 15 Method 1500 describes one or more operations.

[0066] The computer system 1600 may also include one or more secondary storage devices or secondary memories 1610. Secondary memories 1610 may include, for example, a hard disk drive 1612 and / or a removable storage device or drive 1614. The removable storage drive 1614 may be a floppy disk drive, a magnetic tape drive, an optical disc drive, an optical storage device, a magnetic tape backup device, and / or any other storage device / drive.

[0067] The removable storage drive 1614 can interact with the removable storage unit 1618. The removable storage unit 1618 includes a computer-usable or readable storage device on which computer software (control logic) and / or data are stored. The removable storage unit 1618 can be a floppy disk, magnetic tape, optical disc, DVD, optical storage disc, and / or any other computer data storage device. The removable storage drive 1614 reads from and / or writes to the removable storage unit 1618.

[0068] According to some embodiments, secondary memory 1610 may include other means, tools, or other methods for allowing computer system 1600 to access computer programs and / or other instructions and / or data. Such means, tools, or other methods may include, for example, removable storage unit 1622 and interface 1620. Examples of removable storage unit 1622 and interface 1620 may include program cartridges and cartridge interfaces (e.g., found in video game devices), removable memory chips (e.g., EPROM or PROM) and associated sockets, memory sticks and USB ports, memory cards and associated memory card slots, and / or any other removable storage devices and associated interfaces. In some embodiments, secondary memory 1610, removable storage unit 1618, and / or removable storage unit 1622 may include the above-mentioned... Figure 15 Method 1500 describes one or more operations.

[0069] Computer system 1600 may also include a communication or network interface 1624. Communication interface 1624 enables computer system 1600 to communicate and interact with any combination of remote devices, remote networks, remote entities, etc. (individually and collectively referred to as reference numeral 1628). For example, communication interface 1624 may allow computer system 1600 to communicate with remote device 1628 via communication path 1626, which may be wired and / or wireless, and may include any combination of LAN, WAN, Internet, etc. Control logic and / or data may be transmitted to and from computer system 1600 via communication path 1626.

[0070] The operations described in the foregoing embodiments can be implemented in various configurations and architectures. Therefore, some or all of the operations described in the foregoing embodiments, such as... Figure 15 Method 1500 and Figure 17 Method 1700 (described below) can be performed in hardware, software, or both. In some embodiments, a tangible device or article of manufacture including a tangible computer-usable or readable medium on which control logic (software) is stored is also referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 1600, main memory 1608, secondary memory 1610, and removable storage units 1618 and 1622, and tangible articles of manufacture embodying any combination thereof. When executed by one or more data processing devices (such as computer system 1600), this control logic causes these data processing devices to operate as described herein. In some embodiments, computer system 1600 is equipped with software to perform operations in photomask and circuit fabrication, such as... Figure 17 The method 1700 is shown below. In some embodiments, the computer system 1600 includes hardware / devices for manufacturing photomasks and circuit fabrication. For example, the hardware / devices may be connected to or be part of element 1628 (remote device, network, entity) of the computer system 1600.

[0071] Figure 17 This is an illustration of an exemplary method 1700 for circuit fabrication according to some embodiments. The operation of method 1700 may also be performed in different orders and / or variations. Variations of method 1700 should also be within the scope of this disclosure.

[0072] In operation 1701, a GDS file is provided. The GDS file can be generated by an EDA tool and contains a standard cell structure that has been optimized using the disclosed method. The operations described in 1701 can be performed by an EDA tool, for example, running on a computer system (such as computer system 1600 described above).

[0073] In operation 1702, a photomask is formed based on the GDS file. In some embodiments, the GDS file provided in operation 1701 is used in a tape-out operation to generate a photomask for fabricating one or more integrated circuits. In some embodiments, a circuit layout included in the GDS file can be read and transferred onto a quartz or glass substrate to form an opaque pattern corresponding to the circuit layout. The opaque pattern can be made of, for example, chromium or other suitable metal. Operation 1702 can be performed by a photomask manufacturer, where the circuit layout is read using suitable software (e.g., EDA tools) and transferred onto a substrate using suitable printing / deposition tools. The photomask reflects the circuit layout / features contained in the GDS file.

[0074] In operation 1703, one or more circuits are formed based on the photomask generated in operation 1702. In some embodiments, the photomask is used to form a pattern / structure of the circuits contained in the GDS file. In some embodiments, various manufacturing tools (e.g., photolithography equipment, deposition equipment, and etching equipment) are used to form features of the one or more circuits.

[0075] In one aspect of this disclosure, a memory circuit is disclosed. The circuit includes: a first quantity A of memory arrays, each of which includes a plurality of memory cells, each of which is configured to store a corresponding weighted data element; a second quantity B of multiply-accumulate (MAC) circuitry; a third quantity C of adders; a fourth quantity D of word line (WL) driver circuitry; a fifth quantity E of input drivers; a sixth quantity F of control circuitry; and a seventh quantity G of input / output (I / O) circuitry; wherein each of quantities A, B, C, D, E, and G is configurable based on at least one of the quantity X of input data elements or the quantity Z of output channels, wherein both quantity X and quantity Z are arbitrary positive integers.

[0076] In some embodiments, each of the I / O circuits, along with a subset of the memory array, a subset of the MAC circuitry, and a subset of the adders, is physically arranged along a first lateral direction.

[0077] In some embodiments, a subset of I / O circuitry, a subset of memory array, a subset of MAC circuitry, and a subset of adders may operatively form a corresponding output channel in the output channels.

[0078] In some embodiments, along a first lateral direction, a subset of adders is inserted between a first MAC circuit and a second MAC circuit in a subset of MAC circuits, and the first MAC circuit and the second MAC circuit are also inserted between a first memory array and a second memory array in a subset of memory arrays.

[0079] In some embodiments, each of the input drivers is physically positioned relative to another subset of the MAC circuitry along a second lateral direction perpendicular to the first lateral direction, each of the WL driver circuitry is physically positioned relative to another subset of the memory array, and each of the control circuitry is physically positioned relative to a subset of the input / output circuitry.

[0080] In some embodiments, the number of subsets of the A memory arrays coupled to one of the D WL driver circuits is equal to the number Z.

[0081] In some embodiments, the number of subsets of B MAC circuits coupled to one of the E input drivers is equal to the number Z.

[0082] In some embodiments, the number of subsets of the A memory arrays coupled to one of the Z output channels is equal to the number D.

[0083] In some embodiments, the number of subsets of B MAC circuits coupled to one of the Z output channels is equal to the number E.

[0084] In some embodiments, each of the input drivers is configured to provide a subset of X input data elements to a subset of B MAC circuits, each coupled to a Z output channel.

[0085] In some embodiments, each of the D WL driver circuits includes a number of Y WL drivers, the Y WL drivers being configured to activate Y WLs in each of a subset of A memory arrays respectively coupled to Z output channels.

[0086] In some embodiments, each of the C adders is configured to sum the MAC results provided by a subset of the first and second MAC circuits of the B MAC circuits coupled to a corresponding one of the Z output channels.

[0087] In another aspect of this disclosure, a memory circuit is disclosed. The circuit includes: a plurality of first multiply-accumulate (MAC) circuit groups physically arranged along a first lateral direction and operably forming a first output channel. Each of the first MAC circuit groups includes: a first adder; a first MAC circuit and a second MAC circuit, physically sandwiching the first adder along the first lateral direction; and a first memory array and a second memory array, physically sandwiching the first MAC circuit and the second MAC circuit along the first lateral direction.

[0088] In some embodiments, the circuit further includes: a first input / output (I / O) circuit, wherein a plurality of first MAC circuit groups are physically arranged on a first side of the first I / O circuit along a first lateral direction.

[0089] In some embodiments, the circuit further includes: a plurality of second MAC circuit groups physically arranged along a first lateral direction and operably forming a first output channel; wherein each of the second MAC circuit groups includes: a second adder, a third MAC circuit and a fourth MAC circuit physically sandwiched between the second adder along the first lateral direction, and a third memory array and a fourth memory array physically sandwiched between the third MAC circuit and the fourth MAC circuit along the first lateral direction; wherein the plurality of second MAC circuit groups are physically arranged on a second side of the first I / O circuit along the first lateral direction.

[0090] In some embodiments, the circuit further includes: a plurality of second MAC circuit groups physically arranged along a first lateral direction, spaced apart from the plurality of first MAC circuit groups along a second lateral direction perpendicular to the first lateral direction, and operably forming a second output channel; wherein each of the second MAC circuit groups includes: a second adder, a third MAC circuit and a fourth MAC circuit physically sandwiched between the second adder along the first lateral direction, and a third memory array and a fourth memory array physically sandwiched between the third MAC circuit and the fourth MAC circuit along the first lateral direction.

[0091] In some embodiments, the circuit further includes: a plurality of input drivers, wherein a first input driver and a second input driver of the plurality of input drivers are physically arranged along a second lateral direction relative to a first MAC circuit and a second MAC circuit of a corresponding one in the first MAC circuit group; and a plurality of word line (WL) driver circuits, wherein a first WL driver circuit and a second WL driver circuit of the plurality of WL driver circuits are physically arranged along a second lateral direction relative to a first memory array and a second memory array of a corresponding one in the first MAC circuit group.

[0092] In another aspect of this disclosure, a method for forming a memory circuit is disclosed. The method includes arranging a plurality of first multiply-accumulate (MAC) circuit groups along a first lateral direction. Each of the plurality of first MAC circuit groups includes: a first adder; a first MAC circuit and a second MAC circuit, physically sandwiching the first adder along the first lateral direction; and a first memory array and a second memory array, physically sandwiching the first MAC circuit and the second MAC circuit along the first lateral direction. The method includes arranging a plurality of input drivers relative to the plurality of first MAC circuit groups along a second lateral direction perpendicular to the first lateral direction, wherein the first input driver and the second input driver of the plurality of input drivers are respectively physically arranged along the second lateral direction relative to the first MAC circuit and the second MAC circuit of a corresponding one of the first MAC circuit groups. The method includes arranging a plurality of word line (WL) driver circuits relative to the plurality of first MAC circuit groups along the second lateral direction, wherein the first WL driver circuit and the second WL driver circuit of the plurality of WL driver circuits are respectively physically arranged along the second lateral direction relative to the first memory array and the second memory array of a corresponding one of the first MAC circuit groups.

[0093] In some embodiments, the method further includes: arranging a plurality of second MAC circuit groups along a first lateral direction, the plurality of second MAC circuit groups being spaced apart from a plurality of first MAC circuit groups along a second lateral direction, wherein each of the plurality of second MAC circuit groups includes: a second adder, a third MAC circuit and a fourth MAC circuit, the second adder being physically sandwiched in the middle along the first lateral direction, and a third memory array and a fourth memory array, the third MAC circuit and the fourth MAC circuit being physically sandwiched in the middle along the first lateral direction.

[0094] In some embodiments, a plurality of first MAC circuit groups are operably configured to form a first output channel, and a plurality of second MAC circuit groups are operably configured to form a second output channel.

[0095] As used herein, the terms “about” and “approximately” generally refer to the value of a given quantity that can vary depending on the specific technology node associated with the subject semiconductor device. Based on a specific technology node, the term “about” can refer to a given quantity of value, for example, varying within a range of 10-30% of the value (e.g., +10%, ±20%, or ±30% of the value).

[0096] The foregoing outlines features of several embodiments to enable those skilled in the art to better understand various aspects of this disclosure. Those skilled in the art will understand that they can readily use this disclosure as the basis for designing or modifying other processes and structures to achieve the same purposes and / or advantages of the embodiments described herein. Those skilled in the art will also recognize that such equivalent structures do not depart from the spirit and scope of this disclosure, and that various changes, substitutions, and modifications can be made to them within this disclosure without departing from its spirit and scope.

Claims

1. A memory circuit, comprising: A first quantity A of memory arrays, each of which includes a plurality of memory cells, each of which is configured to store a corresponding weight data element; Multiplication-accumulation circuit for the second quantity B; Adder of the third quantity C; The fourth quantity D word line driver circuit; The fifth quantity E is the input driver; The control circuit for the sixth quantity F; and The input / output circuit for the seventh quantity G; The quantities A, B, C, D, E, and G are each configurable based on at least one of the quantity X of input data elements or the quantity Z of output channels, wherein the quantity X and the quantity Z are each arbitrary positive integers.

2. The memory circuit according to claim 1, wherein, Each of the input / output circuits is physically arranged along a first lateral direction with a subset of the memory array, a subset of the multiplication-accumulation circuit, and a subset of the adder.

3. The memory circuit according to claim 2, wherein, The input / output circuitry, a subset of the memory array, a subset of the multiplication-accumulation circuitry, and a subset of the adder operably form a corresponding output channel among the output channels.

4. The memory circuit according to claim 2, wherein, Along the first lateral direction, one of the subsets of adders is inserted between the first and second multiplication-accumulation circuits in the subset of the multiplication-accumulation circuits, and the first and second multiplication-accumulation circuits are also inserted between the first and second memory arrays in the subset of the memory arrays.

5. The memory circuit according to claim 2, wherein, Along a second lateral direction perpendicular to the first lateral direction, each of the input drivers is physically positioned relative to another subset of the multiplication-accumulation circuitry, each of the word line driver circuitry is physically positioned relative to another subset of the memory array, and each of the control circuitry is physically positioned relative to a subset of the input / output circuitry.

6. The memory circuit according to claim 1, wherein, Each of the input drivers is configured to provide a subset of X input data elements to a subset of B multiplication-accumulation circuits, each coupled to one of Z output channels.

7. A memory circuit, comprising: Multiple first multiplication-accumulation circuit groups are physically arranged along a first lateral direction and operably form a first output channel; Each of the first multiplication-accumulation circuit groups includes: First adder, The first multiplication-accumulation circuit and the second multiplication-accumulation circuit physically sandwich the first adder between them along the first lateral direction. A first memory array and a second memory array physically sandwich the first multiplication-accumulation circuit and the second multiplication-accumulation circuit in the middle along the first lateral direction.

8. The memory circuit according to claim 7, further comprising: Multiple second multiplication-accumulation circuit groups are physically arranged along the first lateral direction, separated from the multiple first multiplication-accumulation circuit groups along a second lateral direction perpendicular to the first lateral direction, and operably form a second output channel; Each of the second multiplication-accumulation circuit groups includes: Second adder, The third and fourth multiplication-accumulation circuits physically sandwich the second adder between them along the first lateral direction. The third memory array and the fourth memory array physically sandwich the third multiplication-accumulation circuit and the fourth multiplication-accumulation circuit in the middle along the first lateral direction.

9. The memory circuit according to claim 8, further comprising: A plurality of input drivers, wherein a first input driver and a second input driver of the plurality of input drivers are physically arranged along the second lateral direction relative to a corresponding first multiply-accumulate circuit and a second multiply-accumulate circuit in the first multiply-accumulate circuit group; as well as Multiple word line driver circuits, wherein a first word line driver circuit and a second word line driver circuit are physically arranged along the second lateral direction relative to the first memory array and the second memory array of a corresponding one of the first multiplication-accumulation circuit groups.

10. A method of forming a memory circuit, comprising: A plurality of first multiplication-accumulation circuit groups are arranged along a first lateral direction, wherein each of the plurality of first multiplication-accumulation circuit groups includes: First adder, The first multiplication-accumulation circuit and the second multiplication-accumulation circuit physically sandwich the first adder between them along the first lateral direction. The first memory array and the second memory array physically sandwich the first multiplication-accumulation circuit and the second multiplication-accumulation circuit in the middle along the first lateral direction; A plurality of input drivers are arranged relative to the plurality of first multiply-accumulate circuit groups along a second lateral direction perpendicular to the first lateral direction, wherein the first input driver and the second input driver among the plurality of input drivers are physically arranged relative to a corresponding first multiply-accumulate circuit and a second multiply-accumulate circuit in the first multiply-accumulate circuit group, respectively, along the second lateral direction; and Multiple word line driver circuits are arranged along the second lateral direction relative to the plurality of first multiply-accumulate circuit groups, wherein the first word line driver circuit and the second word line driver circuit of the plurality of word line driver circuits are physically arranged along the second lateral direction relative to the first memory array and the second memory array of a corresponding one of the first multiply-accumulate circuit groups.