SEQUENTIAL HYBRID ACCUMULATOR LAYOUT PLAN FOR COMPUTING-IN-MEMORY

By employing a sequential and hybrid layout for CIM circuits, the challenges of over-densification are mitigated, resulting in reduced metal layer usage, lower manufacturing costs, and improved performance and scalability.

DE102024135603A1Pending Publication Date: 2025-07-03TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102024135603
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-22
Filing Date
2024-12-02
Publication Date
2025-07-03

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A memory device may include a memory array, a first computing unit, and a second computing unit. The memory array may include a plurality of memory cells for storing weights for a neural network. The first computing unit may be configured to receive the stored weights from the plurality of memory cells and generate a first partial sum according to the stored weights. The second computing unit may be configured to receive the stored weights from the plurality of memory cells and the first partial sum and generate a second partial sum according to the stored weights and the first partial sum. The second computing unit may be sequentially connected to the first computing unit.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 616,925, filed January 2, 2024, entitled "SYSTEM and METHOD FOR FLOOR PLAN FOR COMPUTE-IN-MEMORY," which is incorporated by reference into the present application for all purposes. BACKGROUND

[0002] Memory devices are integral components of electronic systems that store data in a manner that allows for rapid access and modification. Traditionally, memory devices have been designed to store binary information in the form of "0s" and "1s" across a large array of memory cells. Due to manufacturing variations and design constraints, these cells often have unbalanced physical structures that lead to mismatches in their electrical properties. Compute-in-memory (CIM) technology integrates processing capabilities directly into memory arrays, enabling faster data processing by reducing the distance data must travel between storage and processing units. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] Aspects of the present disclosure can best be understood by reference to the following detailed description when taken in conjunction with the accompanying drawings. It should be noted that, in accordance with industry practice, various features are not drawn to scale. Indeed, the dimensions of various features may be arbitrarily exaggerated or reduced for the sake of clarity of illustration. Fig. 1 shows a block diagram of a memory device 100 according to some embodiments of the present disclosure. Fig. 2 shows a detailed schematic representation of the memory device 100 of Fig. 1 according to some embodiments of the present disclosure. Fig. 3 shows a detailed schematic representation of the memory device 100 of Fig. 1 according to some embodiments of the present disclosure. Fig. 4 shows a detailed schematic representation of the memory device 100 of Fig. 2 according to some embodiments of the present disclosure. Fig. Figure 5 shows a detailed schematic representation of the memory device 100 of Fig. 3 according to some embodiments of the present disclosure. Fig. 6 is a flowchart of an exemplary method for manufacturing the memory device 100 of Fig. 1 according to some embodiments of the present disclosure. Fig. 7 is a flowchart of an exemplary method of manufacturing the memory device 100 according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0004] The following description provides many different embodiments or examples for implementing various features of the provided subject matter. Specific examples of components and arrangements are described below to facilitate the present invention. These are, of course, merely examples and are not intended to be limiting. For example, the fabrication of a first element over or on top of a second element in the following description may include embodiments in which the first and second elements are fabricated in direct contact, and may also include embodiments in which additional elements may be fabricated between the first and second elements such that the first and second elements are not in direct contact. Furthermore, in the present invention, reference numerals and / or letters may be repeated in the various examples.This repetition is for simplicity and clarity and does not, in itself, prescribe any relationship between the various embodiments and / or configurations discussed.

[0005] In addition, spatially relative terms such as "beneath," "under," "lower," "above," "upper," and the like may be used herein to conveniently describe the relationship of one element or structure to one or more other elements or structures illustrated in the figures. The spatially relative terms are intended to encompass other orientations of the device in use or operation, in addition to the orientation illustrated in the figures. The device may be oriented differently (rotated 90 degrees or in a different orientation), and the spatially relative descriptors used herein may be interpreted accordingly.

[0006] In the traditional floor plan of a Compute-In-Memory (CIM) macro, the design layout often presents significant challenges for accumulator routing (e.g., both global and local accumulators). This layout typically results in severe over-densification due to the extremely complex and intertwined paths for electrical connections between different components. Over-densification not only compromises signal integrity and can potentially increase crosstalk between circuits, but also leads to inefficient use of metal layers. This inefficiency dramatically increases the need for a higher number of metal layers or more complex metal stack configurations to meet dense routing requirements.This not only leads to a more complicated manufacturing process but also increases production costs and can negatively impact the overall performance and scalability of the CIM system. Therefore, optimizing the layout to mitigate excessive routing densification and reduce metal layer utilization is a critical design consideration for improving CIM architectures.

[0007] To improve the design of compute-in-memory (CIM) circuits, it is critical to address the challenge of over-densification, which can be mitigated by implementing a more rational layout that promotes efficient traceability. By reorganizing the layout, signal transmission paths can be simplified, reducing the complexity and overlap of traces that contribute to over-densification. Furthermore, these optimization efforts aim to reduce the reliance on multiple metal layers for local and global routing, which not only simplifies the manufacturing process but also reduces the physical thickness and manufacturing cost of the chip.A careful balance between circuit density and routing simplicity can lead to more scalable and cost-effective CIM designs while significantly reducing the number of required metal layers. The present disclosure not only mitigates excessive routing densification but also improves the overall performance and reliability of the CIM architecture.

[0008] Traceability, in the context of integrated circuit (IC) design, refers to the ease with which electrical connections between different components on a chip can be successfully and efficiently established during the layout phase. It can be a measure of how easily and efficiently the metal wires (also known as "routes") can be placed within the layers of an IC without causing signal integrity or manufacturing problems or violating design rules. A design with high traceability, for example, means that there is sufficient space for all necessary connections and that the risk of causing short circuits, crosstalk, or other problems that can exist if the routes are too dense or poorly organized is minimal.Factors that affect traceability include the number of traces in a metal layer, the complexity of the circuit, the number of layers available for traceability, the accuracy of the manufacturing process, and the effectiveness of the design tools used to create the layout. Improving traceability is essential to ensure that a semiconductor chip can be reliably manufactured at scale and to optimize its performance and energy efficiency. This is especially important as memory circuits become increasingly complex and dense with the continuous advancement of semiconductor technology.

[0009] In a standard layout for a compute-in-memory (CIM) macro, local accumulators are strategically placed between CIM banks to facilitate multiply-accumulate (MAC) operations, which are fundamental to the macro's computational tasks. Despite the logic behind their placement, these local accumulators operate by performing parallel accumulations, which can result in a significant challenge: over-densification. This over-densification occurs because multiple parallel signals must converge on the local accumulator, resulting in a dense and complex network of interconnects. As a result, a higher number of metal layers is often required to accommodate all the required routing paths.For example, the design may require up to 84 traces for some routes and 70 traces for others, highlighting the extensive metal layer utilization required to maintain signal integrity and functionality in these dense routing environments. This not only increases the complexity of the chip design but also impacts the manufacturing process, potentially leading to higher costs and scalability issues.

[0010] The present disclosure provides various embodiments of a memory device that address such issues (e.g., traceability). For example, the memory device disclosed herein comprises a memory array, a first compute unit, and a second compute unit. The second compute unit is sequentially connected to the first compute unit. In some embodiments, by using a fully sequential layout, the total routing traces can be optimized to 36 traces, and fewer metal layers can be used. In some embodiments, by using a hybrid layout (e.g., partially sequential and partially parallel layout), the total routing traces can be kept low to 52 traces, and fewer metal layers can be used.

[0011] Fig. Figure 1 shows a block diagram of a memory device 100 according to some embodiments of the present disclosure. The memory device 100 includes a memory array 120, a first compute unit 140, and a second compute unit 160. In some embodiments, the memory device 100 may include a memory array 120, a first compute unit 140, a second compute unit 160, and a global compute unit 180.

[0012] The memory array 120 may include a plurality of memory cells. The plurality of memory cells may store weights for a neural network. One or more peripheral circuits (not shown) may be disposed in one or more regions peripheral to the memory array 120 or within the memory array 120. The memory cells and the peripheral circuits may be connected by word lines and / or complementary bit lines BL and BLB, and data may be read from and written to the memory bit cells via the complementary bit lines BL and BLB. Various voltage combinations applied to the word lines and bit lines may define a read, erase, or write (programming) operation on the memory bit cells.In some embodiments, various types of non-volatile or volatile memory technologies may be integrated into the architecture of the memory array 120, including, but not limited to, static random-access memory (SRAM), resistive random-access memory (ReRAM), magnetoresistive random-access memory (MRAM), and phase-change random access memory (PCRAM).

[0013] Deep learning uses neural networks to implement artificial intelligence. These networks comprise numerous interconnected processing nodes that facilitate machine learning by analyzing sample data. For example, there is a system designed for object recognition: it can process thousands of images of objects, such as trucks, to perceive and learn the visual structures that correspond to the object in new images. The structure of neural networks is typically in layers, and data flows through these layers in a single forward direction. Each node in the network can have connections to multiple nodes in the subsequent layer to which it sends data, and it can also have connections to numerous nodes in the previous layer from which it receives data.

[0014] In the neural network, a node assigns a numerical value, called a "weight," to its connections. When the node is activated, it can multiply incoming data by this weight and add the products of all its connections, resulting in a single numerical output. If the output falls below a certain threshold, the node can prevent it from moving forward to the next layer. Conversely, if the output exceeds the threshold, the node can transmit this sum to the nodes it is connected to in the subsequent layer. In a deep learning system, a neural network model is stored in memory, and computational logic in a processor performs multiply-accumulate (MAC) calculations on the parameters (e.g., weights) stored in the memory.In some embodiments, the weights may be stored in the plurality of memory cells in the memory array 120.

[0015] In some embodiments, the first computation unit 140 may be configured to receive a plurality of inputs 122a, 122b from the plurality of memory cells 120. The plurality of inputs 122a, 122b may include the stored weights from the plurality of memory cells 120 and / or an input activation vector element. In some embodiments, the first computation unit 140 may be a local accumulator, a full adder, a half adder, a summation register, a partial sum register, and / or an accumulation circuit. In some embodiments, weights (W) or input activation vector elements may be stored in a sub-matrix of the memory matrix 120. Each output of the sub-matrix may be an input to a computation unit 140. For example, the first computation unit 140 may be configured to receive the stored weights from the plurality of memory cells 120.In some embodiments, the first computing unit 140 is inserted between sub-arrays of memory cells 120. The first computing unit 140 and the sub-arrays may be connected on a plurality of local bitlines. In certain embodiments, at least one further computing unit may be inserted between the memory array 120 and the first computing unit 140. In some embodiments, the first computing unit 140 may be configured to generate a first partial sum 142 according to the stored weights 122a, 122b. In certain embodiments, the first computing unit 140 may be configured to generate a first partial sum 142 according to the stored weights and a partial sum from the at least one further computing unit.

[0016] In some embodiments, the second computation unit 160 may be configured to receive a plurality of inputs 122c, 122d from the plurality of memory cells 120. The second computation unit 160 may be configured to receive the first partial sum 142 from the first computation unit 140. The plurality of inputs 122c, 122d may include the stored weights from the plurality of memory cells 120 and / or an input activation vector element. In some embodiments, the second computation unit 160 may be a local accumulator, a full adder, a half adder, a summation register, a partial sum register, and / or an accumulation circuit. In some embodiments, weights (W) or input activation vector elements may be stored in a sub-matrix of the memory matrix 120. Each output of the sub-matrix may be an input to a computation unit 160.For example, the second computing unit 160 may be configured to receive the stored weights from the plurality of memory cells 120. In some embodiments, the second computing unit 160 is inserted between sub-arrays of memory cells 120. The second computing unit 160 and the sub-arrays may be connected on a plurality of local bitlines. In certain embodiments, at least one further computing unit may be inserted between the memory array 120 and the second computing unit 160. In some embodiments, the second computing unit 160 may be configured to generate a second partial sum 162 according to the stored weights 122c, 122d and the first partial sum 142.In certain embodiments, the second computing unit 160 may be configured to generate a second partial sum 162 according to the stored weights, the first partial sum 142, and a partial sum from the at least one further computing unit.

[0017] In some embodiments, the first computing unit 140 may generate the first partial sum 142 by multiplying an input activation vector element by the weights stored in a sub-matrix of memory cells 120. The second computing unit 160 may generate the second partial sum 162 by multiplying an input activation vector element by the weights stored in a sub-matrix of memory cells 120. In some embodiments, each of the plurality of memory cells 120 may include a plurality of wordlines. A multiplication of an input activation vector element on one of the plurality of wordlines by a weight stored in the plurality of memory cells may be calculated by accessing a sub-matrix of memory cells across the plurality of wordlines.

[0018] In some embodiments, the second computing unit 160 may be sequentially connected to the first computing unit 140. In some embodiments, the first computing unit 140 and the second computing unit 160 may be sequentially connected in a same metallization layer / metal layer. In some embodiments, the first computing unit 140 may be directly connected to the second computing unit 160.

[0019] In some embodiments, a memory device 100 may implement a fully sequential layout. Implementing the fully sequential layout in the design of memory devices may result in more optimal use of routing traces. By arranging components in a sequential manner, the total number of routing traces required may be significantly reduced. This more streamlined approach facilitates a more organized and less overly dense routing layout, allowing for fewer metal layers to be used in the circuit. Furthermore, with fewer layers, electrical path lengths are shortened, which may improve signal integrity and potentially the operating speed of the circuit.

[0020] In some embodiments, a memory device 100 may implement a hybrid layout (e.g., a partially sequential layout and a partially parallel layout). By implementing a hybrid layout in IC design that combines both sequential and parallel elements, a significant improvement in traceability and layer utilization can be achieved. By selectively applying sequential routing where feasible and parallel routing where necessary, the overall number of routing traces can be effectively reduced. This balanced approach mitigates the excessive routing densification typically associated with parallel designs while still maintaining the design compactness that purely sequential layouts can impede.The result is a layout that requires fewer metal layers and can significantly reduce complexity and manufacturing costs. This hybrid layout therefore provides an efficient way to optimize the chip's routing infrastructure and contributes to a more streamlined manufacturing process and improved overall circuit performance.

[0021] In some embodiments, the global computation unit 180 may be configured to accumulate at least one partial sum 182 (e.g., first partial sum 142 and / or second partial sum 162) of multiplications from the first computation unit 140 and the second computation unit 160. In some embodiments, the global computation unit 180 may be a global accumulator, a full adder, a half adder, a summation register, a partial sum register, and / or an accumulation circuit. In some embodiments, at least one further computation unit may be inserted between the global computation unit 180 and the second computation unit 160. In certain embodiments, the global computation unit 180 may be sequentially connected to the first computation unit 140, the second computation unit 160, and the at least one further computation unit.

[0022] Fig. 2 shows a detailed schematic representation of the memory device 100 of Fig. 1 according to some embodiments of the present disclosure. Fig. 4 shows a detailed schematic representation of the memory device 100 of Fig. 2 according to some embodiments of the present disclosure. The memory device 100 may include a memory array 120, a first computing unit 140, a second computing unit 160, and a third computing unit 220. In some embodiments, the memory device 100 may include a memory array 120, a first computing unit 140, a second computing unit 160, a third computing unit 220, and a global computing unit 180. The memory devices 100 of the Fig. 2 and Fig. 4 are substantially similar to the memory device 100 of Fig. 1. The specific operations of similar elements, which have already been discussed in detail in the preceding paragraphs, are omitted here for the sake of brevity, unless the cooperative relationship with the elements mentioned in the Fig. 2 and Fig. 4 elements shown must be presented. In the Fig. 2 and Fig. 4 assumes that there are 16 partial sums to be accumulated from 16 CIM banks, where each partial sum is 16 bits.

[0023] In some embodiments, the third computation unit 220 may be configured to receive a plurality of inputs 122e, 122f from the plurality of memory cells 120. The third computation unit 220 may be configured to receive the second partial sum 162 from the second computation unit 160. The plurality of inputs 122e, 122f may include the stored weights from the plurality of memory cells 120 and / or an input activation vector element. In some embodiments, the third computation unit 220 may be a local accumulator, a full adder, a half adder, a summation register, a partial sum register, and / or an accumulation circuit. In some embodiments, weights (W) or input activation vector elements may be stored in a sub-matrix of the memory matrix 120. Each output of the sub-matrix may be an input to a computation unit 220.For example, the third computing unit 220 may be configured to receive the stored weights from the plurality of memory cells 120. In some embodiments, the third computing unit 220 is inserted between sub-arrays of memory cells 120. The third computing unit 220 and the sub-arrays may be connected on a plurality of local bitlines. In certain embodiments, at least one further computing unit may be inserted between the memory array 120 and the third computing unit 220. In some embodiments, the third computing unit 220 may be configured to generate a third partial sum 222 according to the stored weights 122e, 122f and the second partial sum 162.In certain embodiments, the third computing unit 220 may be configured to generate a third partial sum 222 according to the stored weights, the second partial sum 162, and a partial sum from the at least one further computing unit.

[0024] In some embodiments, the third computing unit 220 may be sequentially connected to the first computing unit 140 and the second computing unit 160. In some embodiments, the third computing unit 220, the first computing unit 140, and the second computing unit 160 may be sequentially connected in the same metallization layer / metal layer. In some embodiments, the third computing unit 220 may be directly connected to the second computing unit 160. In some embodiments, a memory device 100 may implement a fully sequential layout (e.g., 140, 160, and 220). Implementing the fully sequential layout in the design of memory devices may result in more optimal use of routing traces. By arranging the components in a sequential manner, the required total number of routing traces may be significantly reduced (e.g., 36 routing traces).This more streamlined approach facilitates a more organized and less dense routing layout, allowing for fewer metal layers to be used in the circuit. Furthermore, with fewer layers, electrical path lengths are shortened, which can improve signal integrity and potentially the circuit's operating speed.

[0025] As in the Fig. 2 and Fig. 4, a first memory cell 120a (e.g., CIM 0) may utilize 16 routing traces to transmit / deliver neural network data to the first local accumulator 140. A second memory cell 120b (e.g., CIM 1) may utilize 16 routing traces to transmit / deliver neural network data to the first local accumulator 140. The first local accumulator 140 may receive / compile the neural network data from the first memory cell 120a and the second memory cell 120b. The first local accumulator 140 may generate a first partial sum 142 and transmit the first partial sum 142 to the second local accumulator 160 using 17 routing traces. A third memory cell 120c (e.g., CIM 2) may utilize 16 routing traces to transmit / deliver neural network data to the second local accumulator 160. A fourth memory cell 120d (e.g.,CIM 3) may utilize 16 routing traces to transmit / deliver neural network data to the second local accumulator 160. The second local accumulator 160 may receive the neural network data from the third memory cell 120c and the fourth memory cell 120d. The second local accumulator 160 may generate a second partial sum 162. The second local accumulator 160 may transmit the second partial sum 162 to a subsequent component (e.g., a third local accumulator 200) using 18 routing traces. The third local accumulator 220 may receive the neural network data from the memory cells. The third local accumulator 220 may generate a third partial sum 222. The third local accumulator 220 may transmit the third partial sum 222 to a subsequent element using 19 routing traces.In this case, the highest number of routing traces in this system is 36 routing traces for CIM 7 210. Compared to the conventional routing layout design that uses up to 84 routing traces, the present disclosure provides a memory device that reduces excessive routing densification and minimizes the use of metal layers.

[0026] In some embodiments, the first computing unit 140 may be sequentially connected to the second computing unit 160 by using routing traces in the same metallization layer (e.g., metal layer N). In some embodiments, multiple metal layers (e.g., metal layers N+1, N+2) may be used. The first memory cell 120a and the second memory cell 120b may be in the same metal layer. The first memory cell 120a and the first computing unit 140 may be in different metallization layers. For example, the first memory cell 120a and the second memory cell 120b may be fabricated in metal layer N. The first computing unit 140, the second computing unit 160, and the third computing unit 222 may be fabricated in metal layer N+1.

[0027] Fig. 3 shows a detailed schematic representation of the memory device 100 of Fig. 1 according to some embodiments of the present disclosure. Fig. Figure 5 shows a detailed schematic representation of the memory device 100 of Fig. 3 according to some embodiments of the present disclosure. The memory device 100 may include a memory array 120, a first computing unit 140, a second computing unit 160, a fourth computing unit 320, and a fifth computing unit 340. In some embodiments, the memory device 100 may include a memory array 120, a first computing unit 140, a second computing unit 160, a fourth computing unit 320, a fifth computing unit 340, and a global computing unit 180. The memory devices 100 of the Fig. 3 and Fig. 5 are substantially similar to the memory device 100 of Fig. 1. The specific operations of similar elements, which have already been discussed in detail in the preceding paragraphs, are omitted here for the sake of brevity, unless the cooperative relationship with the elements mentioned in the Fig. 3 and Fig. 5 elements shown must be presented. In the Fig. 3 and Fig. 5 assumes that there are 16 partial sums to be accumulated from 16 CIM banks, where each partial sum is 16 bits.

[0028] In some embodiments, the fourth computation unit 320 may be configured to receive a plurality of inputs 122a, 122b from the plurality of memory cells 120. The plurality of inputs 122a, 122b may include the stored weights from the plurality of memory cells 120 and / or an input activation vector element. In some embodiments, the fourth computation unit 320 may be a local accumulator, a full adder, a half adder, a summation register, a partial sum register, and / or an accumulation circuit. In some embodiments, weights (W) or input activation vector elements may be stored in a sub-matrix of the memory matrix 120. Each output of the sub-matrix may be an input to a computation unit 320. For example, the fourth computation unit 320 may be configured to receive the stored weights from the plurality of memory cells 120.In some embodiments, the fourth computing unit 320 is inserted between sub-arrays of memory cells 120. The fourth computing unit 320 and the sub-arrays may be connected on a plurality of local bitlines. In certain embodiments, at least one further computing unit may be inserted between the memory array 120 and the fourth computing unit 320. In some embodiments, the fourth computing unit 320 may be configured to generate a fourth partial sum 322 according to the stored weights 122a, 122b. In certain embodiments, the fourth computing unit 320 may be configured to generate a fourth partial sum 322 according to the stored weights and a partial sum from the at least one further computing unit.

[0029] In some embodiments, the fifth computation unit 340 may be configured to receive a plurality of inputs 122c, 122d from the plurality of memory cells 120. The plurality of inputs 122c, 122d may include the stored weights from the plurality of memory cells 120 and / or an input activation vector element. In some embodiments, the fifth computation unit 340 may be a local accumulator, a full adder, a half adder, a summation register, a partial sum register, and / or an accumulation circuit. In some embodiments, weights (W) or input activation vector elements may be stored in a sub-matrix of the memory matrix 120. Each output of the sub-matrix may be an input to a computation unit 340. For example, the fifth computation unit 340 may be configured to receive the stored weights from the plurality of memory cells 120.In some embodiments, the fifth computing unit 340 is inserted between sub-arrays of memory cells 120. The fifth computing unit 340 and the sub-arrays may be connected on a plurality of local bitlines. In certain embodiments, at least one further computing unit may be inserted between the memory array 120 and the fifth computing unit 340. In some embodiments, the fifth computing unit 340 may be configured to generate a fifth partial sum 342 according to the stored weights 122c, 122d. In certain embodiments, the fifth computing unit 340 may be configured to generate a fifth partial sum 342 according to the stored weights and a partial sum from the at least one further computing unit.

[0030] In some embodiments, the first computing unit 140 may be connected in parallel to the fourth computing unit 320 and the fifth computing unit 340 to receive the fourth partial sum 322 and the fifth partial sum 342. The first computing unit 140 may receive the fourth partial sum 322 from the fourth computing unit 320 and the fifth partial sum 342 from the fifth computing unit 340. The first computing unit 140 may generate the first partial sum 142 according to the fourth partial sum 322 and the fifth partial sum 342. In some embodiments, the fourth computing unit 320 and the first computing unit 140 may be in different metallization layers / metal layers. The fourth computing unit 322 and the fifth computing unit 340 may be in the same metallization layer / metal layer.In some embodiments, the fourth computing unit 320 may be connected to the first computing unit 140 using a via structure 360. The fifth computing unit 340 may be connected to the first computing unit 140 via the via structure 360.

[0031] In some embodiments, the second computation unit 160 may be configured to receive a plurality of inputs 122e, 122f, 122g, 122h from the plurality of memory cells 120. The second computation unit 160 may be configured to receive the first partial sum 142 from the first computation unit 140. The plurality of inputs 122e, 122f, 122g, 122h may include stored weights from the plurality of memory cells 120 and / or an input activation vector element. In some embodiments, the second computation unit 160 may be a local accumulator, a full adder, a half adder, a summation register, a partial sum register, and / or an accumulation circuit. In some embodiments, weights (W) or input activation vector elements may be stored in a sub-matrix of the memory matrix 120. Each output of the submatrix can be an input to a computing unit 160.For example, the second computing unit 160 may be configured to receive the stored weights 122e, 122f, 122g, 122h from the plurality of memory cells 120. In some embodiments, the second computing unit 160 is inserted between sub-arrays of memory cells 120. The second computing unit 160 and the sub-arrays may be connected on a plurality of local bitlines. In certain embodiments, at least one further computing unit may be inserted between the memory array 120 and the second computing unit 160. In some embodiments, the second computing unit 160 may be configured to generate a second partial sum 162 according to the stored weights 122e, 122f, 122g, 122h and the first partial sum 142.In certain embodiments, the second computing unit 160 may be configured to generate a second partial sum 162 according to the stored weights, the first partial sum 142, and a partial sum from the at least one further computing unit.

[0032] In some embodiments, the second computing unit 160 may be sequentially connected to the first computing unit 140. In some embodiments, the first computing unit 140 and the second computing unit 160 may be sequentially connected in the same metallization layer / metal layer (e.g., metal layer N, N+1, or N+2). In some embodiments, the first computing unit 140 may be directly connected to the second computing unit 160.

[0033] In some embodiments, a memory device 100 may implement a hybrid layout (e.g., a partially sequential layout and a partially parallel layout). By implementing a hybrid layout in the IC design that combines both sequential elements (e.g., the fourth compute unit 320, the fifth compute unit 340, and the first compute unit 140) and parallel elements (e.g., the first compute unit 140 and the second compute unit 160), a significant improvement in traceability and layer utilization can be achieved. By selectively applying sequential routing where feasible and parallel routing where required, the total number of routing traces can be effectively reduced (e.g., 52 routing traces).This balanced approach mitigates the excessive routing densification typically associated with parallel designs while still maintaining the design compactness that can be hampered by purely sequential layouts. The result is a layout that requires fewer metal layers and can significantly reduce complexity and manufacturing costs. This hybrid layout therefore provides an efficient way to optimize the chip's routing infrastructure and contributes to a more streamlined manufacturing process and better overall circuit performance.

[0034] As in the Fig. 3 and Fig. 5, a first memory cell 120a (e.g., CIM 0) may utilize 16 routing traces to transmit / deliver neural network data to the fourth local accumulator 320. A second memory cell 120b (e.g., CIM 1) may utilize 16 routing traces to transmit / deliver neural network data to the fourth local accumulator 320. The fourth local accumulator 320 may receive / compile the neural network data from the first memory cell 120a and the second memory cell 120b. The fourth local accumulator 320 may generate a fourth partial sum 322 and transmit the fourth partial sum 3222 to the first local accumulator 140 using 17 routing traces in another metal layer. A third memory cell 120c (e.g., CIM 2) may utilize 16 routing traces to transmit / deliver neural network data to the fifth local accumulator 340. A fifth memory cell 120d (e.g.,CIM 3) may use 16 routing traces to transmit / deliver neural network data to the fifth local accumulator 340. The fifth local accumulator 340 may receive the neural network data from the third memory cell 120c and the fourth memory cell 120d. The fifth local accumulator 340 may generate a fifth partial sum 342. The fifth local accumulator 342 may transmit the fifth partial sum 342 to the first computing unit 140 using 17 routing traces. The first local accumulator 140 may receive the neural network data from the memory cells 120. The first local accumulator 140 may generate a first partial sum 142 according to the fourth partial sum 322 and the fifth partial sum 342. The first local accumulator 140 may transmit the first partial sum 142 to a subsequent element (e.g., the second computing unit 160) using 18 routing traces.The second local accumulator 160 may receive the first partial sum 142 and the neural network data from the memory cells. The second local accumulator 160 may generate a second partial sum 162. The second local accumulator 160 may transfer the second partial sum 162 to a subsequent component (e.g., a global accumulator 180) using 19 routing traces. In this case, the highest number of routing traces in this system is 52 routing traces for CIM 7 210. Compared to the conventional routing layout design that uses up to 84 routing traces, the present disclosure provides a memory device that reduces excessive routing densification and minimizes the use of metal layers.

[0035] In some embodiments, the first computing unit 140 may be connected in parallel to the fourth computing unit 320 and the fifth computing unit 340 using the routing traces and a via structure 360. In some embodiments, the fourth computing unit 320 and the first computing unit 140 may be in different metallization layers (e.g., metal layer N, N+1, and N+2). The fourth computing unit 320 and the fifth computing unit 340 may be in the same metallization layer (e.g., metal layer N). In certain embodiments, the first computing unit 140, the fourth computing unit 320, and the fifth computing unit 340 may all be in the same metallization layer (e.g., metal layer N). In some embodiments, the first computing unit 140 may be connected in the same metallization layer (e.g.,Metal layer N) may be sequentially connected to the second computing unit 160. In some embodiments, multiple metal layers (e.g., metal layer N+1, N+2) may be used. The first memory cell 120a and the second memory cell 120b may be in the same metal layer. The first memory cell 120a and the fourth computing unit 320 (or the fifth computing unit 340) may be in different metallization layers. For example, the first memory cell 120a and the second memory cell 120b may be fabricated in metal layer N. The fourth computing unit 320 and the fifth computing unit 340 may be fabricated in metal layer N+1. The first computing unit 140 and the second computing unit 160 may be fabricated in metal layer N+2.

[0036] Fig. 6 is a flow diagram of an exemplary method 600 for manufacturing the memory device 100 of Fig. 1 according to some embodiments of the present disclosure. It is understood that Fig. 6 has been simplified for a better understanding of the principles of the present disclosure. Accordingly, it should be noted that further processes before, during, and after the method of Fig. 6 may be provided and that some other processes may only be briefly described in this disclosure.

[0037] With reference to Fig. 6, operation 602 may provide a substrate. The substrate may have a front side and a back side that are opposite each other. The substrate may be a semiconductor substrate, such as a bulk semiconductor substrate, a semiconductor-on-insulator (SOI) substrate, or the like, which may be doped (e.g., with a p-type or an n-type dopant) or undoped.

[0038] The method 600 then proceeds to operation 604, in which a memory cell array 120 having a plurality of subarrays of memory cells is fabricated on the front side of the substrate. Each subarray of memory cells may be arranged adjacent to any other subarray of memory cells. In some embodiments, a subarray of the memory array 120 may store weights (W) or input activation vector elements for a neural network. Each output of the subarray may be an input to a computation unit.

[0039] The method 600 then proceeds to operation 606, in which a first local accumulator 140 is fabricated on a first metal layer on the front side of the substrate. The first local accumulator 140 may generate a first partial sum 142 by multiplying an input activation vector element by the weights stored in a subarray of memory cells 120.

[0040] The method 600 then proceeds to operation 608, in which a second local accumulator 160 is fabricated on the first metal layer. The first computation unit 140 may be a local accumulator, a full adder, a half adder, a summing register, a partial sum register, and / or an accumulation circuit. The first local accumulator 140 may be sequentially connected to the second local accumulator 160 on the first metal layer. The second computation unit 160 may be a local accumulator, a full adder, a half adder, a summing register, a partial sum register, and / or an accumulation circuit. In some embodiments, the first computation unit 140 may be directly connected to the second computation unit 160. The second computation unit 160 may be configured to receive multiple inputs from the plurality of memory cells 120.The second computing unit 160 may be configured to receive the first partial sum 142 from the first computing unit 140. In some embodiments, the second computing unit 160 may be configured to generate a second partial sum according to the stored weights and the first partial sum.

[0041] Then, the method 600 may continue with an operation in which a third local accumulator is fabricated on the first metal layer. The third local accumulator may be sequentially connected to the first local accumulator and the second local accumulator on the first metal layer. The third computation unit may be a local accumulator, a full adder, a half adder, a summation register, a partial sum register, and / or an accumulation circuit. In some embodiments, the third computation unit may be sequentially connected to the first computation unit 140 and the second computation unit 160. In some embodiments, the third computation unit may be connected in parallel to the first computation unit 140.

[0042] In some embodiments, the method 600 may continue with an operation in which an interconnect structure is fabricated on the backside of the substrate. The interconnect structure may be connected to the array of memory cells 120, the first local accumulator 140, and the second local accumulator 160 on the frontside of the substrate. There may be a plurality of first via structures extending vertically through the substrate from its frontside to its backside. In certain embodiments, the array of memory cells 120, the first local accumulator 140, and the second local accumulator 160 may also be fabricated on the backside of the substrate. The first local accumulator 140 may be sequentially connected to the second local accumulator 160 in the interconnect structure.

[0043] Fig. 7 is a flowchart of an exemplary method 700 for manufacturing the memory device 100 according to some embodiments of the present disclosure. It should be understood that Fig. 7 has been simplified for a better understanding of the principles of the present disclosure. Accordingly, it should be noted that further processes before, during, and after the method of Fig. 7 may be provided and that some other processes may only be briefly described in this disclosure.

[0044] With reference to Fig. 7, operation 705 may provide a substrate. The substrate may have a first side and a second side that are opposite each other. The substrate may be a semiconductor substrate, such as a bulk semiconductor substrate, a semiconductor-on-insulator (SOI) substrate, or the like, which may be doped (e.g., with a p-type or an n-type dopant) or undoped.

[0045] Then, the method 700 continues with operation 710, in which a plurality of first transistors and a second transistor are fabricated on the first side of the substrate. The plurality of first transistors may be a hardware component that stores data. Each of the plurality of first transistors may have a p-type conductivity or an n-type conductivity. In one aspect, the plurality of first transistors may be embodied as a semiconductor memory device or an array of memory cells 120. In some embodiments, the second transistor may be for a head device on the first side of the substrate. The second transistor may have a p-type conductivity or an n-type conductivity.

[0046] Then, the method 700 continues with operation 715, in which a metal structure (e.g., word lines and / or bit lines, interconnect) is formed on the first side of the substrate. The plurality of first transistors and the metal structure (e.g., word lines and / or bit lines) may be for a plurality of memory cells 120 on the first side of the substrate. In some embodiments, a conductor structure may be formed on the first side of the substrate. The conductor structure may be configured to supply the supply voltage to the first transistors. A plurality of via structures may be formed on the first side of the substrate. The via structures may be configured to electrically connect a source / drain terminal of the transistor to the conductor structure.In some embodiments, the first local battery 140 and the second local battery 160 may also be fabricated in the metal structure. The first local battery 140 may be sequentially connected to the second local battery 160 in the interconnect structure.

[0047] In the present disclosure, accumulators in a compute-in-memory (CIM) circuit are strategically placed in a sequential map / layout, enabling them to effectively aggregate partial sums (PSUMs). This methodical arrangement serves to mitigate excessive routing densification by streamlining the paths signals must traverse, thus reducing network complexity. Furthermore, it reduces the reliance on multiple metal layers, contributes to a simpler, lower-cost manufacturing process, and potentially improves signal integrity due to shorter interconnect lengths.

[0048] The present disclosure also presents a hybrid layout approach that combines both sequential and parallel layouts for accumulators accumulating PSUMs. This design also aims to mitigate excessive routing densification by combining the advantages of a sequential system (lower complexity and metal utilization) with the higher connectivity of parallel configurations. This hybrid model provides a balanced solution that can adapt to changing design constraints and performance requirements, providing a versatile framework for maintaining signal integrity and reducing crosstalk, even in densely packed circuit architectures.

[0049] As used herein, the terms "about" and "approximately" generally indicate the value of a given quantity, which may vary depending on the particular technology node associated with the semiconductor device in question. Depending on the particular technology node, the term "about" may indicate a value of a given quantity that varies, for example, within 10-30% of the value (e.g., +10%, ±20%, or ±30% of the value).

[0050] Features of various embodiments have been described above so that those skilled in the art can better understand aspects of the present disclosure. Those skilled in the art will appreciate that they can readily use the present disclosure as a basis for designing or modifying other methods and structures for achieving the same objectives and / or obtaining the same benefits as the embodiments presented herein. Those skilled in the art will also appreciate that such equivalent interpretations do not depart from the spirit and scope of the present disclosure and that they may make various changes, substitutions, and alterations herein without departing from the spirit and scope of the present disclosure. QUOTES CONTAINED IN THE DESCRIPTION

[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature

[0000] US 63 / 616,925

[0001]

Claims

[1] Storage device with: a memory matrix having a plurality of memory cells for storing weights for a neural network; a first computing unit configured to receive the stored weights from the plurality of memory cells and to generate a first partial sum according to the stored weights; and a second computing unit configured to receive the stored weights from the plurality of memory cells and the first partial sum and to generate a second partial sum according to the stored weights and the first partial sum, wherein the second computing unit is sequentially connected to the first computing unit. [2] The memory device of claim 1, wherein the first computing unit and the second computing unit are in a same metallization layer. [3] The storage device according to claim 1 or 2, wherein the first computing unit is directly connected to the second computing unit. [4] Storage device according to one of the preceding claims, comprising: a third computing unit configured to receive the stored weights from the plurality of memory cells and the second partial sum and to generate a third partial sum according to the stored weights and the second partial sum, wherein the third computing unit is sequentially connected to the first computing unit and the second computing unit. [5] The memory device of claim 4, wherein the third computing unit and the first computing unit are in a same metallization layer. [6] Storage device according to one of the preceding claims, comprising: a fourth computing unit configured to receive the stored weights from the plurality of memory cells and to generate a fourth partial sum according to the stored weights; a fifth computing unit configured to receive the stored weights from the plurality of memory cells and generate a fifth partial sum according to the stored weights, wherein the first computing unit is connected in parallel to the fourth computing unit and the fifth computing unit to receive the fourth partial sum and the fifth partial sum and to generate the first partial sum. [7] The memory device of claim 6, wherein the fourth computing unit and the first computing unit are in different metallization layers and the fourth computing unit and the fifth computing unit are in a same metallization layer. [8] The memory device of claim 6, wherein the fourth computing unit is connected to the first computing unit via a via structure, wherein the fifth computing unit is connected to the first computing unit via the via structure. [9] Storage device according to one of the preceding claims, comprising: a global arithmetic unit configured to accumulate partial sums of multiplications from the first arithmetic unit and the second arithmetic unit. [10] A memory device according to any one of the preceding claims, wherein the first arithmetic unit generates the first partial sum by multiplying an input activation vector element by the weights stored in a sub-matrix of memory cells. [11] A memory device according to any one of the preceding claims, wherein the second arithmetic unit generates the second partial sum by multiplying an input activation vector element by the weights stored in a sub-matrix of memory cells. [12] A memory device according to any preceding claim, wherein each of the plurality of memory cells comprises a plurality of word lines, and wherein a multiplication of an input activation vector element by a weight stored in the plurality of memory cells is calculated by accessing a sub-array of memory cells via the plurality of word lines. [13] Storage device comprising: a substrate; a matrix of memory cells having a plurality of sub-matrices of memory cells, fabricated on the substrate and configured to store weights for a neural network; a first local accumulator fabricated on a first metal layer and configured to receive the stored weights from the sub-arrays of memory cells and generate a first partial sum according to the stored weights; a second local accumulator fabricated on the first metal layer and configured to receive the stored weights from the sub-arrays of memory cells and the first partial sum and to generate a second partial sum according to the stored weights and the first partial sum, wherein the second local accumulator is sequentially connected to the first local accumulator on the first metal layer. [14] A storage device according to claim 13, comprising: a third local accumulator fabricated on the first metal layer and configured to receive the stored weights from the sub-arrays of the memory cells and the second partial sum, and to generate a third partial sum according to the stored weights and the second partial sum, wherein the third local accumulator is sequentially connected to the first local accumulator and the second local accumulator on the first metal layer. [15] A storage device according to claim 13 or 14, comprising: a fourth local accumulator fabricated on a second metal layer and configured to receive the stored weights from the sub-arrays of memory cells and to generate a fourth partial sum according to the stored weights; a fifth arithmetic unit fabricated on a second metal layer and configured to receive the stored weights from the sub-arrays of memory cells and generate a fifth partial sum according to the stored weights, wherein the first local accumulator is connected in parallel to the fourth local accumulator and the fifth local accumulator to receive the fourth partial sum and the fifth partial sum and to generate the first partial sum. [16] The memory device of claim 15, wherein the fourth local accumulator and the first local accumulator are in different metal layers, and the fourth local accumulator and the fifth local accumulator are in a same metallization layer. [17] The memory device of claim 15 or 16, wherein the fourth local accumulator is connected to the first local accumulator via a via structure, and wherein the fifth local accumulator is connected to the first local accumulator via the via structure. [18] A method of manufacturing a memory device comprising the following steps: Providing a substrate; Producing a matrix of memory cells having a plurality of sub-matrices of memory cells on a front side of the substrate; Producing a first local accumulator on a first metal layer on the front side of the substrate; and Producing a second local accumulator on the first metal layer, wherein the first local accumulator is sequentially connected to the second local accumulator on the first metal layer. [19] A method according to claim 18, comprising: Producing a third local accumulator formed on the first metal layer, wherein the third local accumulator is sequentially connected to the first local accumulator and the second local accumulator on the first metal layer.

Citation Information

Patent Citations

  • US-PATENTANMELDUNGNR.63/616,925