3D Integrated Ultra-High Bandwidth Memory
By placing the memory die under or on one side of the computing die in an AI processing system, and using tight micro bump spacing and high-density through-silicon (TSV) interconnects, the problem of limited I/O bandwidth between the computing die and the memory die is solved, improving system performance and reducing power consumption.
Patent Information
- Application Number
- CN202080035818.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-31
- Filing Date
- 2020-05-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-05-14
AI Technical Summary
In existing AI processing systems, the I/O bandwidth between the computing die and the memory die is limited, resulting in increased latency and high power consumption.
High bandwidth interconnects are achieved by placing the memory die under or on one side of the computing die, tight micro bump spacing and high-density through-silicon (TSV).
It improves the performance of AI processing systems, reduces the heat dissipation problem of the computing die, and reduces power consumption by reducing the direct correlation between TSV density and I/O density.
Smart Images

Figure CN113826202B_ABST
Abstract
Description
[0001] Priority Claim
[0002] This application claims the benefit of priority of U.S. Patent Application No. 16,428,885, entitled "3D INTEGRATED ULTRAHIGH-BANDWIDTH MEMORY", filed on May 31, 2019, which is incorporated herein by reference in its entirety. Background Art
[0003] Artificial intelligence (AI) is a broad field of hardware and software computing in which data is analyzed, classified, and then decisions are made about the data. For example, over time, a large amount of data is used to train a model that describes the classification of data for a certain attribute or multiple attributes. The process of training a model requires a large amount of data and processing power to analyze the data. When training a model, the weights or weight factors are modified based on the output of the model. Once the weights of the model are calculated to a high confidence level (e.g., 95% or higher) by repeatedly analyzing the data and modifying the weights to obtain the desired results, the model is considered "trained". Then the trained model with fixed weights is used to make decisions about new data. Training a model and then applying the trained model to new data is a hardware-intensive activity. There is a desire to reduce the latency of training a model and using a trained model, and to reduce the power consumption of such AI processor systems.
[0004] The background art description provided herein is for the purpose of generally presenting the context of the present disclosure. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art by virtue of being included in this section. Brief Description of the Drawings
[0005] Embodiments of the present disclosure will be more fully understood from the following detailed description and the accompanying drawings of various embodiments of the present disclosure. However, the present disclosure should not be considered limited to the specific embodiments, but is only for explanation and understanding.
[0006] Figure 1 Shows a high-level architecture of an artificial intelligence (AI) machine including a computing die on top of a memory die according to some embodiments.
[0007] Figure 2 Shows an architecture of a computing block of a computing die on top of a memory die according to some embodiments.
[0008] Figure 3A Shows a cross-section of a package where the computing die is below the memory die, resulting in limited I / O (input-output) bandwidth and thermal issues of the computing die.
[0009] Figure 3BShows a cross-section of a package where a compute die is below a memory die, and where the compute die is through-via perforated with high-density through-silicon vias (TSVs) to couple to bumps between the compute die and the memory die.
[0010] Figure 3C Shows a cross-section of a package where high-bandwidth memory (HBM) is on either side of a compute die, resulting in limited I / O bandwidth due to peripheral constraints on the number of I / Os.
[0011] Figure 4A Shows a cross-section of a package including a compute block according to some embodiments, the compute block including a compute die (e.g., an inference logic die) above a dynamic random access memory (DRAM) die.
[0012] Figure 4B Shows a cross-section of a package including a compute block according to some embodiments, the compute block including a compute die (e.g., an inference logic die) above a stack of memory dies and a controller logic die.
[0013] Figure 4C Shows a cross-section of a package including a compute block according to some embodiments, the compute block including a compute die on a memory (e.g., DRAM) that also serves as an interposer.
[0014] Figure 5A Shows a cross-section of a package including an AI machine according to some embodiments, the AI machine including a system-on-chip (SOC) having a compute block that includes a compute die on DRAM.
[0015] Figure 5B Shows a cross-section of a package including an AI machine according to some embodiments, the AI machine including an SOC having a compute block that includes a compute die on DRAM, a processor, and solid-state memory.
[0016] Figure 5C Shows a cross-section of multiple packages on a circuit board according to some embodiments, where one of the packages includes a compute die on a memory die and another of the packages includes a graphics processing unit.
[0017] Figure 6A Shows a unit cell (or processing element (PE)) of a compute die according to some embodiments, the compute die being configured to couple to a memory die below it.
[0018] Figure 6B Shows a unit cell of a memory die according to some embodiments, the memory die being configured to couple to a compute die above it.
[0019] Figure 7A Shows a computing die including Figure 6A a plurality of unit cells according to some embodiments.
[0020] Figure 7B Shows a memory die including Figure 6B a plurality of unit cells according to some embodiments.
[0021] Figure 8 Shows a cross-section of a top view of a computing die according to some embodiments, the computing die having microbumps on the side for connection to a memory along a horizontal plane.
[0022] Figure 9 Shows a cross-section of a top view of a computing die according to some embodiments, having microbumps on the top and bottom of the computing die for connection to a memory die along a vertical plane of a package.
[0023] Figure 10A Shows a cross-section of a memory die under a computing die according to some embodiments.
[0024] Figure 10B Shows a cross-section of a computing die above a memory die according to some embodiments.
[0025] Figure 11A Shows a cross-section of a memory die having 2x2 tiles under a computing die according to some embodiments.
[0026] Figure 11B Shows a cross-section of a computing die having 2x2 tiles above a memory die according to some embodiments.
[0027] Figure 12 Shows a method of forming a package having a computing die on a memory die according to some embodiments.
[0028] Figure 13 Shows a memory architecture of a portion of a memory die according to some embodiments.
[0029] Figure 14 Shows a bank group in a memory die according to some embodiments.
[0030] Figure 15 Shows a memory channel or block in a memory die according to some embodiments.
[0031] Figure 16 Shows an apparatus for partitioning a memory die into a plurality of channels according to some embodiments.
[0032] Figure 17Illustrates an apparatus showing wafer-to-wafer bonding or Cu-Cu hybrid bonding with microbumps according to some embodiments.
[0033] Figure 18 Illustrates an apparatus showing wafer-to-wafer bonding with a stack of memory cells according to some embodiments, where a first memory wafer of the stack is directly connected to a compute wafer.
[0034] Figure 19 Illustrates an apparatus showing wafer-to-wafer bonding with a stack of memory cells according to some embodiments, where a first memory wafer of the stack is indirectly connected to a compute wafer. Detailed Description
[0035] Due to peripheral constraints, existing packaging technologies that stack dynamic random access memory (DRAM) on top of compute dies result in limited I / O bandwidth. These peripheral constraints come from the vertical interconnections or pillars between the package substrate and the DRAM die. Additionally, placing the compute die below the DRAM causes heat dissipation problems for the compute die because any heat sink is closer to the DRAM and farther from the compute die. Even when wafer-to-wafer bonding of DRAM and compute dies is used in the package, it results in excessive perforations in the compute die because the compute die is stacked below the DRAM. These perforations are caused by through-silicon vias (TSVs) that couple C4 bumps adjacent to the compute die to microbumps, copper-to-copper (Cu-to-Cu) pillars, or hybrid copper-to-copper pillars between the DRAM die and the compute die. When the DRAM die is located above the compute die in a wafer-to-wafer configuration, the TSV density is directly aligned with the die-to-die I / O count, which is substantially similar to the number of microbumps (or copper-to-copper pillars) between the DRAM die and the compute die. Additionally, placing the compute die below the DRAM die in a wafer-to-wafer coupled stack causes heat problems for the compute die because the heat sink is closer to the DRAM die and farther from the compute die. Placing the memory as high-bandwidth memory (HBM) on either side of the compute die does not solve the bandwidth problem for the stacked compute and DRAM dies because the bandwidth is limited by peripheral constraints on the number of I / Os from each side of the HBM and the compute die.
[0036] Some embodiments describe a packaging technology to improve the performance of AI processing systems, resulting in ultra-high bandwidth AI processing systems. In some embodiments, an integrated circuit package is provided that includes: a substrate; a first die on the substrate and a second die stacked on the first die, where the first die includes memory and the second die includes compute logic. In some embodiments, the first die includes a dynamic access memory (DRAM) having bit cells, where each bit cell includes an access transistor and a capacitor.
[0037] In other embodiments, the DRAM below the computing die may replace or supplement other fast access memories, such as ferroelectric RAM (FeRAM), static random access memory (SRAM), and other non-volatile memories (e.g., flash memory, NAND, magnetic RAM (MRAM), Fe-SRAM, Fe-DRAM, and other resistive RAM (Re-RAM), etc.). The memory of the first die may store input data and weight factors. The computing logic of the second die is coupled to the memory of the first die. The second die may be an inference die that applies the fixed weights of the trained model to the input data to generate an output. In some embodiments, the second die includes a processing core (or processing entity (PE)) having a matrix multiplier, an adder, a buffer, etc. In some embodiments, the first die includes a high bandwidth memory (HBM). The HBM may include a controller and a memory array.
[0038] In some embodiments, the second die includes an application specific integrated circuit (ASIC) that can train the model by modifying the weights and also use the model on new data with fixed weights. In some embodiments, the memory includes DRAM. In some embodiments, the memory includes SRAM (static random access memory). In some embodiments, the memory of the first die includes MRAM (magnetic random access memory). In some embodiments, the memory of the first die includes Re-RAM (resistive random access memory). In some embodiments, the substrate is an active interposer and the first die is embedded in the active interposer. In some embodiments, the first die itself is an active interposer.
[0039] In some embodiments, the integrated circuit package is a package for a system on a chip (SOC). The SOC may include a computing die on top of a memory die; an HBM, and a processor die that is coupled to the memory die adjacent to it (e.g., on top of or on the side of the processor die). In some embodiments, the SOC includes a solid state memory die.
[0040] The technical effects of the packaging technologies of various embodiments are numerous. For example, by placing the memory die below the computing die, or by placing one or more memory dies on one (or multiple) sides of the computing die, the performance of the AI system is improved. By placing the memory below the computing die, the heat dissipation problem associated with the computing die being far from the heat sink is solved. The ultra-high bandwidth between the memory and the computing die is achieved through the tight microbump pitch between the two dies. In existing systems, the bottom die is highly perforated with TSVs to transmit signals to and from the active devices of the computing die via the microbumps to the active devices of the memory die. By placing the memory die below the computing die such that their active devices are placed close to each other (e.g., face-to-face), the perforation requirements of the bottom die are greatly reduced. This is because the relationship between the number of microbumps and the TSVs is decoupled. For example, the die-to-die I / O density is independent of the TSV density. The TSVs of the memory die are used to provide power and ground, as well as signals from devices external to the package. From the various embodiments and the drawings, other technical effects will be apparent.
[0041] In the following description, many details are discussed to provide a more thorough explanation of the embodiments of the present disclosure. However, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present disclosure.
[0042] Note that in the corresponding drawings of the embodiments, signals are represented by lines. Some lines may be thicker to indicate more constituent signal paths, and / or have arrows at one or more ends to indicate the primary information flow direction. Such indications are not intended to be restrictive. Instead, these lines are used in conjunction with one or more exemplary embodiments to facilitate easier understanding of the circuit or logic unit. Any represented signal (as dictated by design needs or preferences) can actually include one or more signals that can propagate in either direction and can be implemented with any suitable type of signal scheme.
[0043] The term "device" can generally refer to an apparatus depending on the context in which the term is used. For example, a device can refer to a stack of layers or structures, a single structure or layer, a connection of various structures having active and / or passive elements, etc. Generally, a device is a three-dimensional structure having a plane in the x-y direction and a height in the z direction of the x-y-z Cartesian coordinate system. The plane of the device can also be the plane of the apparatus that includes the device.
[0044] Throughout the specification and in the claims, the term "connected" means a direct connection between the things being connected, e.g., an electrical connection, a mechanical connection, or a magnetic connection, without any intermediate device.
[0045] The term "coupled" means directly or indirectly connected, e.g., a direct electrical connection, a mechanical connection, or a magnetic connection between the things being connected, or an indirect connection through one or more passive or active intermediate devices.
[0046] The term "adjacent" herein generally refers to a position of one thing being near another thing (e.g., immediately adjacent to or near one or more things between them) or adjoining another thing (e.g., abutting another thing).
[0047] The term "circuit" or "module" may refer to one or more passive and / or active components arranged to cooperate with each other to provide a desired function.
[0048] The term "signal" may refer to at least one of a current signal, a voltage signal, a magnetic signal, or a data / clock signal. The meanings of "a", "an", and "the" include plural references. The meaning of "in" includes "in" and "on".
[0049] The term "scaling" generally refers to converting a design (schematic and layout) from one process technology to another and then reducing in layout area. The term "scaling" generally also refers to shrinking the layout and devices within the same technology node. The term "scaling" may also refer to adjusting (e.g., slowing down or speeding up - i.e., downscaling or upscaling respectively) a signal frequency relative to another parameter (e.g., power supply level).
[0050] The terms "substantially", "near", "substantially", "close to", and "about" generally refer to within + / - 10% of a target value. For example, unless otherwise specified in the context of its use, the terms "substantially equal", "about equal", and "substantially equal" mean that there are only accidental variations between the things so described. In the art, such variations generally do not exceed + / - 10% of a predetermined target value.
[0051] Unless otherwise specified, the use of ordinal adjectives "first", "second", and "third", etc. to describe a common object only indicates different instances of the same object being referred to, and is not intended to imply that the objects so described must be in a given sequence in terms of time, space, ranking, or any other way.
[0052] For the purposes of this disclosure, the phrases "A and / or B" and "A or B" mean (A), (B), or (A and B). For the purposes of this disclosure, the phrase "A, B, and / or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).
[0053] The terms "left", "right", "front", "back", "top", "bottom", "upper", "lower", etc. (if any) in the specification and claims are used for descriptive purposes and not necessarily to describe a permanent relative position. For example, the terms "upper", "lower", "front", "back", "top", "bottom", "above", "below", and "on" as used herein refer to the relative position of a component, structure, or material with respect to other reference components, structures, or materials within the device, where such physical relationships are of interest. These terms are used herein for descriptive purposes only and are primarily used in the context of the z-axis of the device and can thus be relative to the orientation of the device. Thus, if the device is oriented upside down relative to the context of the provided figures, the first material "above" the second material in the context of the figures provided herein can also be "below" the second material. In the context of materials, a material disposed on or under another material can be in direct contact or can have one or more intervening materials. Additionally, a material disposed between two materials can be in direct contact with both layers or can have one or more intervening layers. In contrast, the first material "on" the second material is in direct contact with the second material. Similar distinctions are made in the context of component assemblies.
[0054] The term "between" can be used in the context of the z-axis, x-axis, or y-axis of the device. A material between two other materials can be in contact with one or both of these materials, or it can be separated from the other two materials by one or more intervening materials. Thus, a material "between" two other materials can be in contact with either of the other two materials, or it can be coupled to the other two materials through intervening materials. A device between two other devices can be directly connected to one or both of these devices, or it can be separated from the other two devices by one or more intervening devices.
[0055] Here, the term "backend" or BE generally refers to that part of the die which is opposite to the "front end" or FE and where the IC (integrated circuit) package is coupled to the IC die bumps. For example, the advanced metal layers (e.g., metal layer 6 and above in a ten-metal stacked die) and the corresponding vias closer to the die package are considered part of the backend of the die. In contrast, the term "front end" generally refers to that part of the die which includes the active region (e.g., where transistors are fabricated) and the lower metal layers and corresponding vias closer to the active region (e.g., metal layer 5 and below in the ten-metal stacked die example).
[0056] It should be noted that those elements of a figure having the same reference numeral (or name) as elements of any other figure can operate or function in any manner similar to the manner described, but not limited thereto.
[0057] Figure 1 FIG. 1 shows a high-level architecture of an artificial intelligence (AI) machine 100 including a computing die located on top of a memory die, according to some embodiments. The AI machine 100 includes: a computing block 101 or processor having a random access memory (RAM) 102 and computing logic 103; a static random access memory (SRAM) 104, a main processor 105, a dynamic random access memory (DRAM) 106, and a solid state memory or drive (SSD) 107. In some embodiments, some or all components of the AI machine 100 are encapsulated in a single package forming a system on a chip (SOC). In some embodiments, the computing block 101 is encapsulated in a single package and then coupled to the processor 105 and memories 104, 106, and 107 on a printed circuit board (PCB). In various embodiments, the computing block 101 includes a dedicated computing die 103 or microprocessor. In some embodiments, the RAM 102 is a DRAM that forms a dedicated memory / cache for the dedicated computing die 103. The DRAM can be an embedded DRAM (eDRAM), e.g., a 1T-1C (one transistor and one capacitor)-based memory. In some embodiments, the RAM 102 is a ferroelectric RAM (Fe-RAM).
[0058] In some embodiments, the computing die 103 is dedicated to applications such as artificial intelligence, graphics processing, and data processing algorithms. In some embodiments, the computing die 103 also has logic computing blocks, e.g., for multipliers and buffers, special data memory blocks including DRAM (e.g., buffers). In some embodiments, the DRAM 102 has weights and inputs stored in an ordered manner to improve computing efficiency. The interconnections between the processor 105 or dedicated processor 105, SRAM 104, and the computing die 103 are optimized for high bandwidth and low latency. In some embodiments, the SRAM 104 is replaced by Fe-RAM. Figure 1 The architecture allows for efficient packaging to reduce energy / power / cost and provides ultra-high bandwidth between the DRAM 102 and the computing block 101.
[0059] In some embodiments, the RAM 102 includes a DRAM that is partitioned to store input data (or data to be processed) 102a and weight factors 102b. In some embodiments, the RAM 102 includes Fe-RAM. For example, the RAM 102 includes FE-DRAM or FE-SRAM. In some embodiments, the input data 103a is stored in a separate memory (e.g., a separate memory die), and the weight factor 102b is stored in a separate memory (e.g., a separate memory die).
[0060] In some embodiments, the computing logic 103 includes a matrix multiplier, an adder, cascading logic, buffers, and combinational logic. In various embodiments, the computing logic 103 performs a multiplication operation on the input 102a and the weights 102b. In some embodiments, the weights 102b are fixed weights. For example, a processor 105 (e.g., a graphics processing unit (GPU), a field programmable gate array (FPGA) processor, an application specific integrated circuit (ASIC) processor, a digital signal processor (DSP), an AI processor, a central processing unit (CPU), or any other high-performance processor) computes the weights of a trained model. Once the weights are computed, they are stored in the memory 102b. In various embodiments, the input data to be analyzed using the trained model is processed by the computing block 101 with the computed weights 102b to generate an output (e.g., a classification result).
[0061] In some embodiments, the SRAM 104 is ferroelectric-based SRAM. For example, a six-transistor (6T) SRAM bit cell with ferroelectric transistors is used to implement non-volatile Fe-SRAM. In some embodiments, the SSD 107 includes NAND flash cells. In some embodiments, the SSD 107 includes NOR flash cells. In some embodiments, the SSD 107 includes multi-threshold NAND flash cells.
[0062] In various embodiments, the non-volatility of the Fe-RAM is used to introduce new features, e.g., the security, functional safety, and faster restart time of the architecture 100. The non-volatile Fe-RAM is a low-power RAM that provides fast access to data and weights. The Fe-RAM 104 can also be used as a fast storage device for the inference die 101 (or accelerator), which typically has low capacity and fast access requirements.
[0063] In various embodiments, a Fe-RAM (Fe-DRAM or Fe-SRAM) includes a ferroelectric material. The ferroelectric (FE) material can be in a transistor gate stack or in a capacitor of the memory. The ferroelectric material can be any suitable low-voltage FE material that allows the FE material to switch its state with a low voltage (e.g., 100 mV). In some embodiments, the FE material includes a perovskite of the ABO3 type, where "A" and "B" are two cations of different sizes, and "O" is oxygen, which is an anion bonded to the two cations. Generally, the size of the atoms of A is larger than that of B atoms. In some embodiments, the perovskite can be doped (e.g., doped with La or lanthanum). In various embodiments, when the FE material is a perovskite, the conductive oxide is of the AA’BB’O3 type. A’ is a dopant at the atomic site A, which can be an element from the lanthanide series. B’ is a dopant at the atomic site B, which can be one of the transition metal elements, particularly Sc, Ti, V, Cr, Mn, Fe, Co, Ni, Cu, Zn. A’ may have the same valence as site A but a different ferroelectric susceptibility.
[0064] In some embodiments, the FE material includes a hexagonal ferroelectric of the h-RMnO3 type, where R is a rare earth element, i.e., cerium (Ce), dysprosium (Dy), erbium (Er), europium (Eu), gadolinium (Gd), holmium (Ho), lanthanum (La), lutetium (Lu), neodymium (Nd), praseodymium (Pr), promethium (Pm), samarium (Sm), scandium (Sc), terbium (Tb), thulium (Tm), ytterbium (Yb), and yttrium (Y). The ferroelectric phase is characterized by the buckling of layered MnO5 polyhedra, accompanied by the displacement of Y ions, resulting in a net polarization. In some embodiments, the hexagonal FE includes one of the following: YMnO3 or LuFeO3. In various embodiments, when the FE material includes a hexagonal ferroelectric, the conductive oxide is of the A2O3 type (e.g., In2O3, Fe2O3) and of the ABO3 type, where "A" is a rare earth element and B is Mn.
[0065] In some embodiments, the FE material includes improper FE materials. An improper ferroelectric is a ferroelectric in which the primary order parameter is an order mechanism such as strain or buckling of the atomic order. Examples of improper FE materials are materials of the LuFeO3 class or superlattices of ferroelectric and paraelectric materials (PbTiO3 (PTO) and SnTiO3 (STO), and LaAlO3 (LAO) and STO, respectively). For example, a superlattice of [PTO / STO]n or [LAO / STO]n, where "n" ranges from 1 to 100. Although the various embodiments are described herein with reference to ferroelectric materials for storing charge states, the embodiments are also applicable to paraelectric materials. In some embodiments, the memory 104 includes a DRAM instead of a Fe-RAM.
[0066] Figure 2 Shows an architecture of a computing block 200 (e.g., 101) including a computing die located on top of a memory die according to some embodiments. Figure 2 The architecture shows an architecture for a dedicated computing die, where the RAM memory buffers for the input and weights are split on die 1, and the logic and optional memory buffers are split on die 2.
[0067] In some embodiments, the memory die (e.g., die 1) is located below the computing die (e.g., die 2) such that a heat sink or thermal solution is adjacent to the computing die. In some embodiments, the memory die is embedded in an interposer. In some embodiments, in addition to its basic memory function, the memory die also acts as an interposer. In some embodiments, the memory die is a high bandwidth memory (HBM) that includes multiple memory dies in a stack and a controller that controls the read and write functions to the stack of memory dies. In some embodiments, the memory die includes a first die 201 for storing input data and a second die 202 for storing weight factors. In some embodiments, the memory die is a single die that is partitioned such that a first partition 201 of the memory die is for storing input data and a second partition 202 of the memory die is for storing weights. In some embodiments, the memory die includes DRAM. In some embodiments, the memory die includes FE-SRAM or FE-SRAM. In some embodiments, the memory die includes MRAM. In some embodiments, the memory die includes SRAM. For example, the memory partitions 201 and 202, or the memory dies 201 and 202 include one or more of the following: DRAM, FE-SRAM, FE-DRAM, SRAM, and / or MRAM. In some embodiments, the input data stored in the memory partition or die 201 is data that will be analyzed by a trained model with fixed weights stored in the memory partition or die 202.
[0068] In some embodiments, the compute die includes a matrix multiplier 203, logic 204, and a temporary buffer 205. The matrix multiplier 203 performs a multiplication operation on input data “X” and weights “W” to generate an output “Y”. This output can be further processed by the logic 204. In some embodiments, the logic 204 performs: threshold operations, pooling and dropout operations, and / or concatenation operations to complete AI logic primitive functions. In some embodiments, the output of the logic 204 (e.g., the processed output “Y”) is temporarily stored in the buffer 205. In some embodiments, the buffer 205 is a memory, such as one or more of the following: DRAM, Fe-SRAM, Fe-DRAM, MRAM, resistive RAM (Re-RAM), and / or SRAM. In some embodiments, the buffer 205 is part of a memory die (e.g., die 1). In some embodiments, the buffer 205 performs the function of a retimer. In some embodiments, the output of the buffer 205 (e.g., the processed output ‘Y’) is used to modify weights in the memory partition or die 202. In one such embodiment, the compute block 200 operates not only as an inference circuit but also as a training circuit for training a model. In some embodiments, the matrix multiplier 203 includes an array of multiplier units, where DRAMs 201 and 202 each include an array of memory bit cells, and each multiplier unit is coupled to a corresponding memory bit cell of DRAM 201 and / or DRAM 202. In some embodiments, the compute block 200 includes an interconnect structure coupled to the array of multiplier units such that each multiplier unit is coupled to the interconnect structure.
[0069] The architecture 200 reduces memory access to the compute die (e.g., die 2) by providing data locality for weights, inputs, and outputs. In one example, data from and to the AI compute block (e.g., the matrix multiplier 203) is locally processed within the same package unit. The architecture 200 also isolates memory and logic operations onto a memory die (e.g., die 1) and a logic die (e.g., die 2) respectively, thus allowing optimized AI processing. The isolated dies allow for increased die yield. The high-capacity memory process for die 1 allows for reduced power in the external interconnects of the memory, reduced integration cost, and also results in a smaller footprint.
[0070] Figure 3AA cross-section of package 300 (also referred to as package configuration 300) is shown, where the compute die is below the memory die, resulting in limited I / O bandwidth and thermal issues for the compute die. In some embodiments, an integrated circuit (IC) package component is coupled to a circuit board 301. In some embodiments, the circuit board 301 can be a printed circuit board (PCB) made of an electrically insulating material such as an epoxy laminate. For example, the circuit board 301 can include an electrically insulating layer made of materials such as: phenolic cotton paper material (e.g., FR-1), cotton paper and epoxy material (e.g., FR-3), woven glass material laminated together with epoxy (FR-4), epoxy glass / paper (e.g., CEM-1), epoxy glass composite, polytetrafluoroethylene woven glass cloth (e.g., PTFE CCL), or other polytetrafluoroethylene-based prepreg materials. In some embodiments, layer 301 is a package substrate and is part of the IC package component.
[0071] The IC package component can include a substrate 302, a compute die 303, and a memory die 304. In this case, the memory die 304 is located above the compute die 304. Here, the compute die 303 is coupled to the memory die 304 through pillar interconnects (e.g., copper pillars). The memory die 303 communicates with the compute die 304 through these pillar interconnects. The pillar interconnects are embedded in a dielectric 318 (or a sealant 318).
[0072] The package substrate 302 can be a coreless substrate. For example, the package substrate 302 can be a "bump-less build-up layer" (BBUL) component that includes multiple "bump-less" build-up layers. Here, the term "bump-less build-up layer" generally refers to a layer of substrate and components embedded without using solder or other attachment means that may be regarded as "bumps". However, various embodiments are not limited to the BBUL-type connection between the die and the substrate, but can be used for any suitable flip-chip substrate. One or more build-up layers can have material properties that can be changed and / or optimized for reliability, warpage reduction, etc. The package substrate 302 can be composed of polymer, ceramic, glass, or semiconductor materials. The package substrate 302 can be a traditional core substrate and / or an interposer. The package substrate 302 includes active and / or passive devices embedded therein.
[0073] The upper side of the package substrate 302 is coupled to the compute die 303 via C4 bumps. The opposite lower side of the package substrate 302 is coupled to the circuit board 301 through package interconnect 317. The package interconnect 316 can couple the electrical wiring features 317 provided on the second side of the package substrate 302 to the corresponding electrical wiring features 315 on the circuit board 301.
[0074] Here, the term "C4" bump (also known as Controlled Collapse Chip Connection) provides a mechanism for interconnecting semiconductor devices. These bumps are commonly used in flip-chip packaging technology, but are not limited to this technology.
[0075] The package substrate 302 may have electrical wiring features formed therein to route electrical signals between the computing die 303 (and / or memory die 304) and the circuit board 301 and / or other electrical components external to the IC package assembly. The package interconnects 316 and die interconnects 310 include any of a variety of suitable structures and / or materials, including, for example, bumps, pillars, or balls formed using metals, alloys, solderable materials, or combinations thereof. The electrical wiring features 315 may be arranged in a ball grid array ("BGA") or other configuration. The computing die 303 and / or memory die 304 include two or more dies embedded in a sealant 318. Here, a heat sink 315 and associated fins are coupled to the memory die 304.
[0076] In this example, the computing die 303 is coupled to the memory die 304 in a front-to-back configuration (e.g., the "front" or "active" side of the memory die 304 is coupled to the "back" or "inactive" side of the computing die 303). The back-end (BE) interconnect layer 303a of the computing die 303 and the active devices 303b are closer to the C4 bumps than the DRAM die 304. The BE interconnect layer 304a and active devices 304b (e.g., transistors) of the DRAM die 304 are closer to the computing die 303 than the heat sink 315.
[0077] In this example, the stack of the DRAM die 304 on top of the computing die 303 is not wafer-to-wafer bonded. This is evident from the different surface areas of the two dies. Pillars such as TSVs are used for communication between the circuit board 301, the computing die 303, and the DRAM die 304. This particular package configuration has limited I / O bandwidth because the DRAM die 304 and the computing die 303 communicate via the pillars in the periphery. Signals from the computing die 303 are routed via the C4 bumps and through the substrate 302 and the pillars, and then reach the active devices 304b via the BE 304a of the DRAM die 304. This long path, along with the limited number of pillars and C4 bumps, limits the overall bandwidth of the AI system. Additionally, since the computing die 303 is not directly coupled to the heat sink 315, this configuration also suffers from thermal issues. Although the heat sink 315 is shown as a thermal solution, other thermal solutions may also be used. For example, in addition to or instead of the heat sink 315, a fan, liquid cooling, etc. may be used.
[0078] Figure 3BShows a cross-section of package 320, where compute die 303 is below memory die 304, and where compute die 303 is perforated with high-density through-silicon vias (TSVs) to couple to bumps between compute die 303 and memory die 304. In this example, compute die 303 and DRAM die 304 are wafer-to-wafer bonded via solder balls or micro-bumps 310 or any suitable technology. The configuration of package 320 results in a higher bandwidth than the configuration of package 320. This is because the peripheral routing via columns is replaced by direct routing between bump 310 and TSV 303c. In this package configuration, bottom die 303 is highly perforated with TSVs 303b to transmit signals to and from the active devices of compute die 303 via micro-bumps 310 to the active devices of memory die 304. This perforation is due to a direct correlation between the number of bumps 310 and TSVs 303b. In this case, the number of TSVs is the same as the number of bumps 310. To increase the bandwidth, a larger number of bumps and TSVs need to be added. However, increasing the TSVs limits the wiring in compute die 303. Similar to Figure 3A the configuration of, here package configuration 320 also suffers from thermal issues because compute die 303 is not directly coupled to heat sink 315.
[0079] Figure 3C Shows a cross-section of package 330, where high-bandwidth memory (HBM) is located on either side of compute die 303, resulting in limited I / O bandwidth due to peripheral constraints on the number of I / Os. In this case, the memory dies are not stacked on compute die 303 but are placed adjacent or laterally close to compute die 303 as HBMs 334 and 335. The bandwidth of this configuration is limited by the peripheral constraints in region 226 between the bumps 310 of HBM 334 / 335 and compute die 303. Therefore, the memory access energy is higher than that of package configuration 320 because the memory access is non-uniformly distributed. In this configuration, the number of channels is limited by the number of peripheral I / O counts in region 336.
[0080] Figure 4A Shows a cross-section of a package 400 (herein referred to as package configuration 400) including a compute block according to some embodiments, the compute block including a compute die (e.g., an inference logic die) above a dynamic random access memory (DRAM) die. Compared to Figure 3A the package configuration of -C, this particular topology enhances the overall performance of the AI system by providing ultra-high bandwidth. Compared to Figure 3BIn contrast, here the DRAM die 401 is located beneath the compute die 402, and the two dies are wafer-to-wafer bonded via microbumps 403, copper-to-copper (Cu-to-Cu) pillars, hybrid copper-to-copper pillars 403. In some embodiments, the copper-to-copper pillars are made of copper pillars formed on each wafer substrate to be bonded together. In various embodiments, a conductive material (e.g., nickel) is coated between the copper pillars of the two wafer dies.
[0081] Dies 401 and 402 are bonded such that their respective BE layers and active devices 401a / b and 403a / b face each other. Thus, the transistors between the two dies are closest to the location where die-to-die bonding occurs. This configuration reduces latency because the active devices 401a and 402a are closer to each other compared to Figure 3B the active devices 301a and 302a.
[0082] Compared with Figure 3B the configuration of. The TSV 401c is decoupled from the microbumps (or copper-to-copper pillars). For example, the number of TSVs 401c is not directly related to the number of microbumps 403. Thus, the memory die TSV via requirements are minimized because the die-to-die I / O density is independent of the TSV density. The ultra-high bandwidth also comes from the tight microbump pitch. In some embodiments, the microbump pitch 403 is tighter than Figure 3B the microbump pitch 310 because the DRAM 401 is not via-perforated at the same pitch as in Figure 3B the compute die 302. For example, in Figure 3B , the microbump density depends on the TSV pitch of the compute die 302 and the overall signal routing design. The package configuration 400 does not have such a limitation.
[0083] Here, the DRAM die 401 is via-perforated to form a few TSVs 401c that transfer DC signals such as power and ground from the substrate 302 to the compute die 402. External signals (e.g., external to the package 400) can also be routed to the compute die 402 via the TSVs 401c. Most of all communication between the compute die 402 and the DRAM die 401 occurs via the microbumps 403 or the face-to-face interconnects 403. In various embodiments, the compute die 402 is not via-perforated because TSVs may not be needed. Even if any additional dies (not shown) are routed to the top of the compute die 402 using TSVs, the number of these TSVs is independent of the number of microbumps 403 because they may not have to be the same number. In various embodiments, the TSVs 401c pass through the active region or layer (e.g., the transistor region) of the DRAM die 401.
[0084] In various embodiments, the compute die 402 includes the logic portion of the inference die. The inference die or chip is used to apply inputs and fixed weights associated with a trained model to generate an output. By separating the memory 401 associated with the inference die 402, AI performance is improved. Additionally, this topology allows for better utilization of thermal solutions such as the heat sink 315, which dissipates heat from the power consumption source, the inference die 402. Although the memory for die 401 is shown as DRAM 401, different types of memory may also be used. For example, in some embodiments, the memory 402 may be one or more of the following: FE-SRAM, FE-DRAM, SRAM, MRAM, resistive RAM (Re-RAM), embedded DRAM (e.g., 1T-1C based memory), or combinations thereof. Using FE-SRAM, MRAM, or Re-RAM allows for low-power and high-speed memory operation. This allows the memory die 401 to be placed below the compute die 402 to more efficiently utilize the thermal solution for the compute die 402. In some embodiments, the memory die 401 is a high bandwidth memory (HBM).
[0085] In some embodiments, the compute die 402 is an application specific integrated circuit (ASIC), a processor, or some combination of these functions. In some embodiments, one or both of the memory die 401 and the compute die 402 may be embedded in a sealant (not shown). In some embodiments, the sealant may be any suitable material, e.g., epoxy-based laminate substrates, other dielectric / organic materials, resins, epoxy resins, polymer adhesives, silicones, acrylics, polyimides, cyanates, thermoplastics, and / or thermosets.
[0086] The memory circuits of some embodiments may also have active and passive devices on the front side of the die. The memory die 401 may have a first side S1 and a second side S2 opposite the first side S1. The first side S1 may be the side of the die that is generally referred to as the "inactive" or "back" side of the die. The back side of the memory die 401 may include active or passive devices, signal and power routing, etc. The second side S2 may include one or more transistors (e.g., access transistors) and may be the side of the die that is generally referred to as the "active" or "front" side of the die. The second side S2 of the memory die 401 may include one or more electrical routing features 310. The compute die 402 may include an "active" or "front" side having one or more electrical routing features connected to microbumps 403. In some embodiments, the electrical routing features may be pads, solder balls, or any other suitable coupling technique.
[0087] Compared with the packaging configuration 320, the thermal problem is alleviated here because the heat sink 315 is directly attached to the computing die 402, which generates most of the heat in this packaging configuration. Although Figure 4A The embodiment of is shown as a wafer-to-wafer bond between dies 401 and 402. However, in some embodiments, these dies can also be bonded using wafer-to-die bonding techniques. Compared with the packaging configuration 320, a higher bandwidth is achieved between the DRAM die 401 and the computing die 402 because more channels are available between the memory die 401 and the computing die 402. In addition, compared with the memory access energy of the packaging configuration 320, the memory access energy is reduced because the memory access is direct and uniform, rather than indirect and distributed. Due to the local access of the processing elements (PEs) of the computing die 402 to the memory in the die, the latency is reduced compared to the latency in the packaging configuration 320. The close and direct connection between the computing die 402 and the memory die 401 allows the memory of the memory die 401 to act as a cache memory that can be accessed quickly.
[0088] In some embodiments, the IC packaging component can include, for example, a combination of flip chip and wire bonding techniques, an interposer, a multi-chip packaging configuration, including a system-on-chip (SoC) for routing electrical signals and / or a package-on-package (PoP) configuration.
[0089] Figure 4B A cross-section of a package 420 (also referred to herein as packaging configuration 420) including a computing block is shown, the computing block including a computing die (e.g., an inference logic die) above a stack of memory dies and a controller logic die. Compared with the packaging configuration 400, the stack of memory dies is located below the computing die 402 here. The stack of memory dies includes die 401, which can include memory (e.g., cache) and controller circuits (e.g., row / column controllers and decoders, read and write drivers, sense amplifiers, etc.). Below die 401, the stacked memory die 403 1-N , where die 4031 is adjacent to die 401 and die 403 N is adjacent to the substrate 302, and where "N" is an integer greater than 1. In some embodiments, each die in the stack is wafer-to-wafer bonded via microbumps or copper-to-copper hybrid pillars. In various embodiments, each memory die 403 1-N has its active devices away from the C4 bumps and more towards the active devices of 402a.
[0090] However, in some embodiments, the memory die 403 1-NIt can be flipped such that the active device faces the substrate 302. In some embodiments, the connection between the compute die 402 and the first memory die 401 (or a controller die with memory) is face-to-face and can provide a higher bandwidth for this interface compared to the interfaces with other memory dies in the stack. Signals and power can be transferred from the compute die 402 to the C4 bumps through the TSVs of the memory die. The TSVs between the various memory dies can transfer signals between the dies in the stack or transfer power (and ground) to the C4 bumps. In some embodiments, the communication channels between the compute die 402 or across the memory dies in the stack are connected through TSVs and micro-bumps or wafer-to-wafer Cu hybrid bonding. Although Figure 4B the embodiments shown have the memory as DRAM, the memory can be embedded DRAM, SRAM, flash memory, Fe-RAM, MRAM, Fe-SRAM, Re-RAM, etc. or a combination thereof.
[0091] In some embodiments, the variable pitch TSVs (e.g., TSV 401c) between the memory dies (e.g., 401 and / or 403 1-N ) enable a large number of I / Os between the dies, resulting in distributed bandwidth. In some embodiments, the stacked memory dies connected through a combination of TSVs and bonding between the dies (e.g., using micro-bumps or wafer-to-wafer bonding) can transfer power and signals. In some embodiments, the variable pitch TSVs achieve high density on the bottom die (e.g., die 401), where the I / Os are achieved with a closer pitch, while the power lines and / or ground lines are achieved with relaxed pitch TSVs.
[0092] Figure 4C FIG. shows a cross-section of a package 430 (also referred to as package configuration 430) including a compute block according to some embodiments, the compute block including a compute die on a memory (e.g., DRAM) that also serves as an interposer. In some embodiments, the compute die 402 is embedded in the sealant 318. In some embodiments, the sealant 318 can be any suitable material, e.g., epoxy-based laminate substrates, other dielectric / organic materials, resins, epoxy resins, polymer adhesives, silicones, acrylics, polyimides, cyanates, thermoplastics, and / or thermosets.
[0093] Compared with the package configuration 400, here the memory die 401 is removed and integrated in the interposer 432, such that the memory provides both storage functionality and the functionality of the interposer. This configuration allows for a reduction in packaging cost. The interconnect 403 (e.g., C4 bumps or micro-bumps) now electrically couples the compute die 402 to the memory 432. The memory 432 can include DRAM, embedded DRAM, flash memory, FE-SRAM, FE-DRAM, SRAM, MRAM, Re-RAM, or a combination thereof. The same advantages of Figure 4A are also achieved in this embodiment. In some embodiments, the memory die 401 is embedded in a substrate or an interposer 302.
[0094] In some embodiments, the compute die and two or more memories are positioned along a plane of the package, and the memory also serves as an interposer. In some embodiments, the memory interposer 432 is replaced by a three-dimensional (3D) RAM stack that also serves as an interposer. In some embodiments, the 3D memory stack is a stack of DRAM, embedded DRAM, MRAM, Re-RAM, or SRAM.
[0095] Figure 5AA cross-section of a package 500 including an AI machine is shown, the AI machine including a system-on-chip (SOC) having a computing block, the computing block including a compute die on memory. The package 500 includes a processor die 506 coupled to a substrate or interposer 302. Two or more memory dies 507 (e.g., memory 104) and 508 (e.g., memory 106) are stacked on the processor die 506. The processor die 506 (e.g., 105) can be any one of the following: a central processing unit (CPU), a graphics processing unit (GPU), a DSP, a field programmable gate array (FPGA) processor, or an application specific integrated circuit (ASIC) processor. The memory (RAM) dies 507 and 508 can include DRAM, embedded DRAM, FE-SRAM, FE-DRAM, SRAM, MRAM, Re-RAM, or a combination thereof. In some embodiments, the RAM dies 507 and 508 can include HBM. In some embodiments, one of the memories 104 and 106 is implemented as HBM in die 405. The memory in the HBM die 505 includes any one or more of the following: DRAM, embedded DRAM, FE-SRAM, FE-DRAM, SRAM, MRAM, Re-RAM, or a combination thereof. A heat sink 315 provides a thermal management solution for the various dies in the sealant 318. In some embodiments, a solid state drive (SSD) 509 is located outside the first package component including the heat sink 315. In some embodiments, the SSD 509 includes one of the following: NAND flash memory, NOR flash memory, or any other type of non-volatile memory, e.g., DRAM, embedded DRAM, MRAM, FE-DRAM, FE-SRAM, Re-RAM, etc.
[0096] Figure 5B A cross-section of a package 520 including an AI machine is shown, the AI machine including an SOC having a computing block, the computing block including a compute die on memory, a processor, and solid state memory. The package 520 is similar to the package 500, but incorporates the SSD 509 into a single package under a common heat sink 315. In this case, the single package SOC provides an AI machine that includes the ability to generate a training model and then use the trained model to generate an output for different data.
[0097] Figure 5CShows a cross-section 530 of multiple packages on a circuit board according to some embodiments, where one of the packages includes a computing die on a memory die, and another of the packages includes a graphics processing unit. In this example, an AI processor such as CPU 525 (GPU, DSP, FPGA, ASIC, etc.) is coupled to a substrate 301 (e.g., a printed circuit board (PCB)). Here, two packages are shown - one with a heat sink 526 and the other with a heat sink 527. The heat sink 526 is a dedicated thermal solution for the GPU chip 525, while the heat sink 527 provides a thermal solution for the computing block (dies 402 and 304) with HBM505.
[0098] Figure 6A Shows a unit cell (or processing element (PE)) 600 of a computing die 402 according to some embodiments, where the computing die 402 is configured to be coupled to a memory die 401 below it. In some embodiments, the PE 600 includes a matrix multiplication unit (MMU) 601, registers 602, a system bus controller 603, an east / west (E / W) bus 604, a north / south (N / S) bus 605, a local memory controller 606, and a die-to-die I / O interface 607. The MMU 601 serves the same function as the multiplier 103, and the registers 602 are used to hold the input 102a and the weights 102b. The system bus controller 603 controls data and control communication via the E / W bus 604 and the N / W bus 605. The local memory controller 606 controls the selection of the input and weights and the associated read and write drivers. The die-to-die I / O interface communicates with the memory unit cell below.
[0099] Figure 6B Shows a unit cell 620 of a memory die 401 according to some embodiments, where the memory die 401 is configured to be coupled to a computing die 402 above it. The memory unit cell 600 includes an array of bit cells, where each array can be a unit array cell. In this example, a 4x4 unit array is shown, where each unit array (e.g., array 0,0; array 0,4; array 4,0; array 4,4) includes a plurality of bit cells arranged in rows and columns. However, any NxM array can be used for the unit array, where "N" and "M" are integers that can be the same or different numbers. The bit cells of each array can be accessed by a row address decoder. The bit cells of each array can be read and written using adjacent read / write control and drivers. The unit cell 600 includes control and refresh logic 626 to control the reading and writing of the bit cells of the array. The unit cell 600 includes a die-to-die I / O interface 627 for communicating with the die-to-die I / O interface 607 of the PE 600.
[0100] Figure 7A shows a plurality of unit cells 600 including Figure 6A in accordance with some embodiments N,M (wherein in this example "N" and "M" are 4) of a compute die 700 (e.g., 402). Note that "N" and "M" can be any numbers, depending on the desired architecture. The compute die 700 includes I / O interfaces and memory channels along its periphery. The PEs 600 N,M can be accessed by a network-on-chip (NoC) including routers, drivers, and interconnects 701a and 701b. In some embodiments, both sides (or more) have memory channels (MC) 702 including MC1 to MC4. In some embodiments, the compute die 700 includes channels 703 compliant with double data rate (DDR) (e.g., DDR CH1, DDR CH2, DDR CH3, DDR CH4). However, the embodiments are not limited to DDR-compliant I / O interfaces. Other low-power and fast interfaces can also be used. In some embodiments, the compute die 700 includes PCIe (Peripheral Component Interconnect Express) and / or SATA (Serial ATA) interfaces 704. Other serial or parallel I / O interfaces can also be used. In some embodiments, additional general-purpose I / O (GPIO) interfaces 705 are added along the periphery of the compute die 700. Each PE is above the corresponding memory unit cell. The architecture of the compute die 700 allows the memory of the memory die 401 to be decomposed into the required number of channels, and helps to increase bandwidth, reduce latency, and reduce access energy.
[0101] Figure 7B shows a plurality of unit cells 620 including Figure 6B in accordance with some embodiments N,M (wherein in this example "N" and "M" are 4) of a memory die 720. In some embodiments, the memory die 720 communicates with the compute die 700 above it via GPIO 725. In other embodiments, other types of I / O can be used to communicate with the compute die 700.
[0102] Figure 8 shows a cross-section of a top view 800 of a compute die 402 in accordance with some embodiments, which has microbumps on the sides for connection to memory along a horizontal plane. The shaded regions 801 and 802 on either side of the compute die 402 include microbumps 803 for connection to memory on either side of the compute die 402. The microbumps 804 can be used to connect to the substrate 302 or the interposer 302.
[0103] Figure 9Shows a cross-section of a top view 900 of a computing die 402 according to some embodiments, having microbumps at the top and bottom of the computing die for connection to a memory die along a vertical plane of the package. The shaded regions 901 and 902 on the upper and lower portions of the computing die 402 include microbumps 903 for connection to an upper memory and a lower memory, respectively. Microbumps 904 can be used for connection to the substrate 302 or the interposer 302.
[0104] Figure 10A Shows a cross-section 1000 of a memory die (e.g., 401) below a computing die 402 according to some embodiments. The pitch of the memory die 401 is "L" x "W". The cross-section 1000 shows a strip of TSVs for connection to the computing die 402. The shaded strip 1001 carries signals, while the strips 1002 and 1003 carry power and ground lines. The strip 1004 supplies a power signal 1005 and a ground signal 1006 to memory cells within a row. The TSV 1008 connects a signal (e.g., a word line) to a memory bit cell.
[0105] Figure 10B Shows a cross-section 1020 of a computing die (e.g., 402) above a memory die (e.g., 401) according to some embodiments. The TSV 1028 can be coupled to the TSV 1008, while the strip 1024 is above the strip 1004. The TSVs 1025 and 1026 are coupled to the TSVs 1005 and 1006, respectively.
[0106] Figure 11A Shows a cross-section 1100 of a memory die 401 with 2x2 sharding below a computing die according to some embodiments. While Figure 10A the memory die 401 shown has a single shard, here 2x2 sharding is used to organize the memory. This allows for a clear division of the memory for storing data and weights. Here, the shards are indicated by shards 1101. The embodiments are not limited to the 2x2 sharding and MxN sharding organizations (where M and N are integers that can be equal or different).
[0107] Figure 11B Shows a cross-section 1120 of a computing die with 2x2 sharding above a memory die according to some embodiments. Similar to the memory die 401, the computing die 402 can also be divided into shards. According to some embodiments, each shard 1121 is similar to Figure 10B the computing die 402. This organization of the computing die 402 allows different training models with different input data and weights to be run simultaneously or in parallel.
[0108] Figure 12FIG. 1200 is a flow chart of a method of forming a package of a computing block according to some embodiments, the computing block including a computing die (e.g., an inference logic die) located above a memory die. The blocks in FIG. 1200 are shown in a particular order. However, the order of the various processing steps may be modified without changing the essence of the embodiments. For example, some processing blocks may be processed simultaneously, while others may be executed out of order.
[0109] At block 1201, a substrate (e.g., 302) is formed. In some embodiments, the substrate 302 is a package substrate. In some embodiments, the substrate 302 is an interposer (e.g., an active or passive interposer). At block 1202, a first die (e.g., 401) is formed on the substrate. In some embodiments, forming the first die includes a dynamic random access memory (DRAM) having bit cells, where each bit cell includes an access transistor and a capacitor. At block 1203, a second die (e.g., computing die 402) is formed and stacked on the first die, where forming the second die includes forming computing logic coupled to the memory of the first die. In some embodiments, forming the computing logic includes forming an array of multiplier units, and where the DRAM includes an array of memory bit cells.
[0110] At block 1204, an interconnect structure is formed. At block 1205, the interconnect structure is coupled to the array of multiplier units such that each multiplier unit is coupled to the interconnect structure. In some embodiments, the DRAM is partitioned into a first partition that can operate as a buffer; and a second partition that stores weight factors.
[0111] In some embodiments, the method of FIG. 1200 includes: receiving data from the first partition and the second partition by the computing logic; and providing an output of the computing logic to a logic circuit. In some embodiments, forming the computing logic includes forming ferroelectric logic. In some embodiments, the computing logic is operable to multiply at least two matrices. In some embodiments, the method of forming the substrate includes forming active or passive devices. In some embodiments, the method includes: forming a third die (e.g., a logic die or a memory) on the substrate. In some embodiments, the method includes coupling the third die on the substrate. In some embodiments, the method includes a fourth die that includes a dynamic random access memory (DRAM); and stacking the fourth die on the third die. In some embodiments, the method includes coupling a heat sink to the second die.
[0112] In some embodiments, the method includes: coupling an AI processor to the DRAM of a first die includes performing a wafer-to-wafer bond of the first die and a second die; or, coupling an AI processor to the DRAM of a first die includes coupling the first die and the second die via microbumps. In some embodiments, the method includes: forming a first die includes forming through-silicon vias (TSVs) in the first die, wherein the number of TSVs is substantially less than the number of microbumps. In some embodiments, the method includes: coupling the first die and the second die via microbumps includes coupling the first die and the second die such that the active devices of the first die and the active devices of the second die are closer to the microbumps than a heat sink. In some embodiments, the method includes: providing a power supply and a ground supply for the TSVs. In some embodiments, the method includes: coupling a device external to the apparatus via the TSVs, wherein the second die is independent of the TSVs. In some embodiments, the method includes: forming a first die on a substrate includes coupling the first die to the substrate via C4 bumps. In some embodiments, the method includes forming a network-on-chip (NoC) on the first die or the second die. In some embodiments, the method includes coupling a heat sink to the second die.
[0113] In some embodiments, forming an AI includes forming an array of multiplier units, and wherein the DRAM includes an array of memory bit cells, and wherein the AI processor is operable to multiply at least two matrices. In some embodiments, the method includes: forming an interconnect structure; and coupling the interconnect structure to the array of multiplier units such that each multiplier unit is coupled to the interconnect structure. In some embodiments, the DRAM is partitioned into a first partition operable as a buffer; and a second partition for storing weight factors, wherein the method includes: receiving data from the first partition and the second partition by compute logic; and providing an output of the AI processor to the logic circuitry.
[0114] Figure 13 A memory architecture 1300 of a portion of a memory die 401 according to some embodiments is shown. In some embodiments, the memory organization uses fine-grained banks. These fine-grained banks use smaller arrays and subarrays. In this example, smaller array sizes (e.g., 128x129 or 256x257) are used to improve the speed of certain applications. In some embodiments, wide bus access is used to reduce unwanted activation energy costs. In some embodiments, memory banks can be constructed with a larger number of subarrays. Similarly, subarrays with a larger number of arrays can also be used.
[0115] Figure 14Shows a bank group 1400 in a memory die 401 according to some embodiments. In some embodiments, a bank group (BGn) may include a plurality of fine-grained banks. For example, a bank may include a cache bank to allow a 1T-SRAM type interface for DRAM or embedded DRAM (eDRAM) refresh timing management from a timing perspective. Refresh timing management is combined with DRAM to provide a high-bandwidth, low-latency interface that can hide periodic refresh requirements in the background while not interfering with normal read / write access to memory blocks. In some embodiments, the memory die 401 may include redundant banks for remapping. In some embodiments, different numbers of active banks can be implemented within a bank group by using or organizing more or fewer fine-grained banks. In some embodiments, memory bank refresh (e.g., for eDRAM or DRAM) can occur separately. In some embodiments, logic is provided for intelligent refresh using cache banks.
[0116] Figure 15 Shows a memory channel 1500 or block in a memory die according to some embodiments. The memory channel may include one or more bank groups. In some embodiments, intermediate blocks are used to facilitate data width sizing and / or to order prefetching for each memory access to allow matching the I / O speed to any inherent speed limitations within the memory bank.
[0117] Figure 16 Shows a memory die 1600 divided into multiple channels according to some embodiments. In various embodiments, the bottom memory die 401 includes multiple memory sub-blocks per die. Each sub-block provides independent wide-channel access to the top compute die 402. In some embodiments, the bottom die itself may also include a network-on-chip (NoC) to facilitate communication between different memory sub-blocks.
[0118] Figure 17 Shows a device 1700 according to some embodiments, which shows a wafer-to-wafer bond with micro-bumps or Cu-Cu hybrid bonding. As discussed herein, the memory wafer has through-silicon vias (TSVs) for interfacing with C4 bumps (or the package side). In some embodiments, the memory wafer is thinned after bonding to reduce the length of the TSVs from the memory die 401 to the compute die 402. As a result, a closer TSV pitch is achieved, thereby reducing IR drop and reducing latency (resulting in a higher operating speed).
[0119] Figure 18FIG. 0 shows apparatus 1800 in accordance with some embodiments, which shows a wafer-to-wafer bond of a stack of memory cells, where a first memory wafer of the stack is directly connected to a compute wafer. In this example, the first memory wafer (having a memory or controller die 401) is directly connected to the compute wafer (having a compute die 402). This face-to-face bond allows for a greater number of I / O channels. In some embodiments, the memory wafer is thinned after the bond to reduce the length of the TSVs from the memory die 401 to the compute die 402. As a result, a closer TSV pitch is achieved, thereby reducing IR drop and decreasing latency (resulting in a higher operating speed).
[0120] Figure 19 FIG. 4 shows a wafer-to-wafer bond of a stack of memory cells for apparatus 1900 in accordance with some embodiments, where a first memory wafer of the stack is indirectly connected to a compute wafer. In this example, the stack of wafers (pressed into dies) is not face-to-face connected. For example, in this example, the active devices of the dies do not face each other.
[0121] References in the specification to "an embodiment", "one embodiment", "some embodiments", or "other embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments is included in at least some embodiments, but not necessarily all embodiments. References to "an embodiment", "one embodiment", or "some embodiments" do not necessarily refer to the same embodiments everywhere. If the specification states that a component, feature, structure, or characteristic "may", "might", or "could" be included, that particular component, feature, structure, or characteristic is not necessarily required to be included. If the specification or claim refers to "a" or "an" element, this does not mean there is only one element. If the specification or claim refers to "additional" elements, more than one additional element is not excluded.
[0122] Furthermore, particular features, structures, functions, or characteristics may be combined in any suitable manner in one or more embodiments. For example, wherever a particular feature, structure, function, or characteristic associated with two embodiments is not mutually exclusive, the first embodiment may be combined with the second embodiment.
[0123] Although the present disclosure has been described in connection with its particular embodiments, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. Embodiments of the present disclosure are intended to embrace all such alternatives, modifications, and variations that fall within the broad scope of the appended claims.
[0124] In addition, for simplicity of illustration and discussion, and so as not to obscure the present disclosure, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the figures presented. Further, to avoid obscuring the present disclosure and also in view of the fact that the details of the implementation of such block diagram arrangements are highly dependent on the platform in which the present disclosure is to be implemented, the arrangements may be shown in block diagram form (i.e., these details should be within the capabilities of those of ordinary skill in the art). In cases where specific details (e.g., circuits) are set forth to describe example embodiments of the present disclosure, it should be apparent to those of ordinary skill in the art that the present disclosure may be practiced without these specific details or with variations of these specific details. Accordingly, the description is to be regarded as illustrative rather than restrictive.
[0125] The following examples are provided to illustrate various embodiments. These examples may be combined with other examples. Accordingly, various embodiments may be combined with other embodiments without changing the scope of the invention.
[0126] Example 1: A device, comprising: a substrate; a first die on the substrate, wherein the first die includes a dynamic random access memory (DRAM) having bit cells, wherein each bit cell includes an access transistor and a capacitor; and a second die stacked on the first die, wherein the second die includes a computing block coupled to the DRAM of the first die.
[0127] Example 2: The device of Example 1, wherein the first die and the second die are wafer-to-wafer bonded or die-to-wafer bonded.
[0128] Example 3: The device of Example 1, wherein the first die and the second die are coupled via at least one of the following: microbumps, copper-to-copper hybrid bonding, or wire bonding.
[0129] Example 4: The device of Example 3, wherein the first die includes through-silicon vias (TSVs), wherein the number of TSVs is substantially less than the number of microbumps.
[0130] Example 5: The device of Example 4, wherein the TSVs include power lines and ground lines, and lines coupling to devices external to the device.
[0131] Example 6: The device of Example 4, wherein the second die is independent of the TSVs.
[0132] Example 7: The device of Example 3, wherein the first die and the second die are coupled such that the active devices of the first die and the active devices of the second die are closer to the microbumps than to a heat sink.
[0133] Example 8: The device of Example 1, wherein the first die is coupled to the substrate via C4 bumps.
[0134] Example 9: The apparatus of Example 1, wherein the first die or the second die includes a Network-on-Chip (NoC).
[0135] Example 10: The apparatus of Example 1, wherein the compute die includes an array of multiplier units, and wherein the DRAM includes an array of memory bit cells.
[0136] Example 11: The apparatus of Example 10 includes an interconnect structure coupled to the array of multiplier units such that each multiplier unit is coupled to the interconnect structure.
[0137] Example 12: The apparatus of Example 1, wherein the DRAM is partitioned into a first partition operable as a buffer; and a second partition for storing weight factors.
[0138] Example 13: The apparatus of Example 12, wherein the compute die receives data from the first partition and the second partition, and wherein the output of the compute logic is received by a logic circuit.
[0139] Example 14: The apparatus of Example 12, wherein the AI processor is operable to multiply at least two matrices.
[0140] Example 15: The apparatus of Example 1, wherein the substrate includes active or passive devices.
[0141] Example 16: The apparatus of Example 1, wherein a third die is on the substrate, and wherein a fourth die includes a DRAM stacked on the third die.
[0142] Example 17: The apparatus of Example 1, wherein a heat sink is coupled to the second die.
[0143] Example 18: The apparatus of Example 1, wherein the DRAM includes an embedded DRAM (eDRAM).
[0144] Example 19: The apparatus of Example 1, wherein the compute die includes one of the following: FPGA, ASIC, CPU, AI processor, DSP, or GPU.
[0145] Example 20: A method includes: forming a substrate; forming a first die on the substrate, wherein forming the first die includes forming a dynamic random access memory (DRAM) having bit cells; and forming a second die, wherein forming the second die includes forming an artificial intelligence (AI) processor; and stacking the second die on the first die, wherein stacking the second die on the first die includes coupling the AI processor to the DRAM of the first die.
[0146] Example 21: The method of Example 20, wherein: coupling the AI processor to the DRAM of the first die includes performing a wafer-to-wafer bond between the first die and the second die; or, coupling the AI processor to the DRAM of the first die includes only microbumps coupling the first die and the second die; forming the first die includes forming through-silicon vias (TSVs) in the first die, wherein the number of TSVs is substantially less than the number of microbumps; and coupling the first die and the second die via the microbumps includes coupling the first die and the second die such that the active devices of the first die and the active devices of the second die are closer to the microbumps than to a heat sink.
[0147] Example 22: The method of Example 20 includes: providing a power supply and a ground supply for the TSVs; coupling devices external to the apparatus via the TSVs, wherein the second die is independent of the TSVs; forming the first die on a substrate includes coupling the first die to the substrate via C4 bumps; forming a network-on-chip (NoC) on the first die or the second die; and coupling a heat sink to the second die.
[0148] Example 23: The method of Example 20, wherein forming the AI includes forming an array of multiplier units, and wherein the DRAM includes an array of memory bit units, and wherein the AI processor is operable to multiply at least two matrices.
[0149] Example 24: The method of Example 20 includes: forming an interconnect structure; and coupling the interconnect structure to the array of multiplier units such that each multiplier unit is coupled to the interconnect structure.
[0150] Example 25: The method of Example 20, wherein the DRAM is partitioned into a first partition operable as a buffer; and a second partition for storing weight factors, wherein the method includes: receiving data from the first partition and the second partition by computing logic; and providing an output of the AI processor to a logic circuit.
[0151] Example 26: A system, comprising: a first memory including non-volatile memory (NVM) cells; a second memory, wherein the first memory is coupled to the second memory; a third memory, coupled to the first memory; a first processor, coupled to the second memory; and a second processor coupled to the third memory and the first processor, wherein the second processor includes: a substrate; a first die on the substrate, wherein the first die includes a memory having bit cells; and a second die stacked on the first die, wherein the second die includes a computing block coupled to the memory of the first die.
[0152] Example 27: The system of Example 26, wherein: the first die and the second die are wafer-to-wafer bonded or die-to-wafer bonded; the first die and the second die are coupled via microbumps; the first die includes through-silicon vias (TSVs), wherein the number of TSVs is substantially less than the number of microbumps; the TSVs include power lines and ground lines, and lines that couple to devices external to the device; the second die is independent of the TSVs; and the first die and the second die are coupled such that the active devices of the first die and the active devices of the second die are closer to the microbumps than to a heat sink.
[0153] Example 28: The system of Example 26, wherein the memory of the second processor includes one of the following: DRAM, flash memory, eDRAM, MRAM, ReRAM, SRAM, or FeRAM.
[0154] Example 29: A device, comprising: a substrate; a first die on the substrate, wherein the first die includes a memory having bit cells; and a second die stacked on the first die, wherein the second die includes a computing block that is coupled to the memory of the first die.
[0155] Example 30: The device of Example 29, wherein the second die includes one of the following: FPGA, ASIC, CPU, AI processor, DSP, or GPU.
[0156] Example 31: The device of Example 29, wherein the memory includes one of the following: DRAM, flash memory, eDRAM, MRAM, ReRAM, SRAM, or FeRAM.
[0157] Example 32: A device, comprising: a substrate; a stack of memory dies, including a first die and a second die, the first die including a memory having bit cells, the second die including controller logic, a cache, or memory, wherein one of the dies in the stack is on the substrate; and a computing die stacked on the second die in the stack of memory dies.
[0158] Example 33: The device of Example 32, wherein the memory includes one of the following: DRAM, flash memory, eDRAM, MRAM, ReRAM, SRAM, or FeRAM.
[0159] Example 34: The device of Example 32, wherein the first die and the computing die are wafer-to-wafer bonded or die-to-wafer bonded.
[0160] Example 35: The device of Example 32, wherein the first die and the second die are coupled via at least one of the following: microbumps, copper-to-copper hybrid bonding, or wire bonding.
[0161] Example 36: The apparatus of Example 32, wherein the first die and the compute die are coupled via at least one of the following: microbumps, copper-to-copper hybrid bonding, or wire bonding.
[0162] Example 37: The apparatus of Example 36, wherein the die on the substrate in the stack includes through-silicon vias (TSVs), where the number of TSVs is substantially less than the number of microbumps, copper-to-copper hybrid bonding, or wire bonding.
[0163] Example 38: The apparatus of Example 32, wherein the compute die is independent of the TSVs.
[0164] Example 39: The apparatus of Example 32, wherein at least one of the die in the stack or the compute die includes a network-on-chip (NoC).
[0165] Example 40: The apparatus of Example 32, wherein the compute die includes one of the following: FPGA, ASIC, CPU, AI processor, DSP, or GPU.
[0166] Example 41: An apparatus, comprising: a substrate; a stack of memory dies including a first die and a second die, the first die including a memory having bit cells, the second die including controller logic, cache, or memory, wherein one of the dies in the stack is on the substrate; and an artificial intelligence processor die stacked on the second die in the stack of memory dies.
[0167] Example 42: The apparatus of Example 41, wherein the memory includes one of the following: DRAM, flash memory, eDRAM, MRAM, ReRAM, SRAM, or FeRAM.
[0168] Example 43: The apparatus of Example 41, wherein the first die and the compute die are wafer-to-wafer bonded or die-to-wafer bonded.
[0169] Example 44: The apparatus of Example 41, wherein the first die and the second die are coupled via at least one of the following: microbumps, copper-to-copper hybrid bonding, or wire bonding.
[0170] Example 45: The apparatus of Example 41, wherein the first die and the artificial intelligence processor die are coupled via at least one of the following: microbumps, copper-to-copper hybrid bonding, or wire bonding.
[0171] Example 46: The apparatus of Example 45, wherein the die on the substrate in the stack includes through-silicon vias (TSVs), where the number of TSVs is substantially less than the number of microbumps, copper-to-copper hybrid bonding, or wire bonding.
[0172] Example 47: The apparatus of Example 41, wherein the artificial intelligence processor die is independent of the TSVs.
[0173] Example 48: A system comprising: a first memory including non-volatile memory (NVM) cells; a second memory, wherein the first memory is coupled to the second memory; a third memory coupled to the first memory; a first processor coupled to the second memory; and a second processor coupled to the third memory and the first processor, wherein the second processor includes: a substrate; a stack of memory dies including a first die and a second die, the first die including a memory having bit cells, the second die including controller logic, a cache, or a memory, wherein one of the dies in the stack is on the substrate; and a compute die stacked on the second die in the stack of memory dies.
[0174] Example 49: The system of Example 48, wherein the memory of the first die includes one of the following: DRAM, flash memory, eDRAM, MRAM, ReRAM, SRAM, or FeRAM.
[0175] Example 50: The system of Example 17, wherein: the first die and the compute die are wafer-to-wafer bonded or die-to-wafer bonded; the first die and the second die are coupled via at least one of microbumps, copper-to-copper hybrid bonding, or wire bonding; the first die and the compute die are coupled via at least one of microbumps, copper-to-copper hybrid bonding, or wire bonding; and wherein the die in the stack on the substrate includes through-silicon vias (TSVs), wherein the number of TSVs is substantially less than the number of microbumps, copper-to-copper hybrid bonding, or wire bonding.
[0176] Example 51: The system of Example 48, wherein the compute die is independent of the TSVs.
[0177] Example 52: The system of Example 48, wherein at least one of the dies in the stack or the compute die includes a network-on-chip (NoC).
[0178] Example 53: The system of Example 48, wherein the compute die includes one of the following: FPGA, ASIC, CPU, AI processor, DSP, or GPU.
[0179] A summary is provided that allows the reader to determine the nature and gist of the technical disclosure. When submitting the summary, it should be understood that it will not be used to limit the scope or meaning of the claims. The appended claims are hereby incorporated into the detailed description, where each claim stands on its own as a separate embodiment.
Claims
1. An apparatus having a high bandwidth memory, the apparatus comprising: A substrate; A first die on the substrate, wherein the first die includes a dynamic random access memory (DRAM) having bit cells, and wherein each bit cell includes an access transistor and a capacitor; and A second die stacked on the first die, wherein the second die includes a computing block coupled to the DRAM of the first die, Wherein the first die and the second die are coupled via at least one of: micro - bumps, copper - to - copper hybrid bonding, or wire bonding, wherein the first die includes through - silicon vias (TSVs), and wherein the number of TSVs is substantially less than the number of micro - bumps.
2. The device according to claim 1, wherein The first die and the second die are wafer - to - wafer bonded or die - to - wafer bonded.
3. The device according to claim 1, wherein The TSVs include power lines and ground lines, and lines coupling to devices external to the apparatus.
4. The device according to claim 1, wherein, The second die is independent of the TSVs.
5. The apparatus according to claim 1, wherein The first die and the second die are coupled such that, compared to a heat sink, the active devices of the first die and the active devices of the second die are closer to the micro - bumps.
6. The apparatus according to claim 1, wherein The second die includes an array of multiplier units, and wherein the DRAM includes an array of memory bit cells.
7. The apparatus according to claim 6, comprising an interconnect structure coupled to the array of multiplier units such that each multiplier unit is coupled to the interconnect structure.
8. The apparatus according to claim 1, wherein The DRAM is partitioned into a first partition capable of operating as a buffer; and a second partition for storing weight factors.
9. The apparatus according to claim 8, wherein, The second die is configured to receive data from the first partition and the second partition, and wherein the output of the computing logic is received by a logic circuit.
10. The apparatus according to claim 8, wherein, An AI processor is operable to multiply at least two matrices.
11. The device according to any one of claims 1 to 10, wherein, The first die is coupled to the substrate via C4 bumps.
12. The device according to any one of claims 1 to 10, wherein, The first die or the second die includes a network - on - chip (NoC).
13. The device according to any one of claims 1 to 10, wherein The substrate includes active or passive devices.
14. The apparatus according to any one of claims 1 to 10, wherein, A third die is on the substrate, and wherein a fourth die includes a DRAM stacked on the third die.
15. The device according to any one of claims 1 to 10, wherein A heat sink is coupled to the second die.
16. The device according to any one of claims 1 to 10, wherein, The DRAM includes an embedded DRAM (eDRAM).
17. The device according to any one of claims 1 to 10, wherein The second die includes one of: an FPGA, an ASIC, a CPU, an AI processor, a DSP, or a GPU.
18. A method for forming a high bandwidth memory, the method comprising: Forming a substrate; Forming a first die on the substrate, wherein forming the first die includes forming a dynamic random access memory (DRAM) having bit cells; Forming a second die, wherein forming the second die includes forming an artificial intelligence (AI) processor; and Stacking the second die on the first die, wherein stacking the second die on the first die includes coupling the AI processor to the DRAM of the first die, Among them, coupling the AI processor to the DRAM of the first die includes wafer-to-wafer bonding of the first die and the second die or coupling the first die and the second die via microbumps. Among them, forming the first die includes forming through-silicon vias (TSVs) in the first die, and among them, the number of TSVs is substantially less than the number of microbumps.
19. The method according to claim 18, wherein: Coupling the first die and the second die via microbumps includes coupling the first die and the second die such that they are coupled so that the active devices of the first die and the active devices of the second die are closer to the microbumps than to the heat sink.
20. The method according to claim 18, includes: Providing a power supply and a ground supply for the TSVs; Coupling devices outside the device via the TSVs, wherein the second die is independent of the TSVs; Forming the first die on the substrate includes coupling the first die to the substrate via C4 bumps; Forming an on-chip network (NoC) on the first die or the second die; and Coupling a heat sink to the second die.
21. The method according to claim 18, wherein Forming the AI includes forming an array of multiplier units, and among them, the DRAM includes an array of memory bit units, and among them, the AI processor is operable to multiply at least two matrices.
22. The method according to claim 21, includes: Forming an interconnect structure; And Coupling the interconnect structure to the array of multiplier units such that each multiplier unit is coupled to the interconnect structure.
23. The method according to any one of claims 18 to 22, wherein, The DRAM is divided into a first partition that can operate as a buffer; And a second partition for storing weight factors, wherein the method includes: Receiving data from the first partition and the second partition by computing logic; and Providing the output of the AI processor to the logic circuit.
24. A system, includes: A first memory, which includes non-volatile memory (NVM) units; A second memory, wherein the first memory is coupled to the second memory; A third memory, which is coupled to the first memory; A first processor, which is coupled to the second memory; and A second processor, which is coupled to the third memory and the first processor, wherein the second processor includes: A substrate; A first die on the substrate, wherein the first die includes a memory having bit units; and A second die stacked on the first die, wherein the second die includes a computing block, and the computing block is coupled to the memory of the first die, Wherein the first die and the second die are coupled via at least one of the following: microbumps, copper-to-copper hybrid bonding, or wire bonding, wherein the first die includes through-silicon vias (TSVs), and wherein the number of TSVs is substantially less than the number of microbumps.
25. The system according to claim 24, wherein: The first die and the second die are wafer-to-wafer bonded or die-to-wafer bonded; The TSV includes a power line and a ground line, and lines that couple to devices external to the device; The second die is independent of the TSV; and the first die and the second die are coupled such that the active devices of the first die and the active devices of the second die are closer to the microbumps than a heat sink.
26. The system according to any one of claims 24 to 25, wherein, The memory of the second processor includes one of the following: DRAM, flash memory, eDRAM, MRAM, ReRAM, SRAM, or FeRAM.
27. The system according to any one of claims 24 to 25, wherein, A heat sink is coupled to the second die.
Citation Information
Patent Citations
Memory expansion structure in multi-path accessible semiconductor memory device
US20070208902A1
Distributed on-chip decoupling apparatus and method using package interconnect
US20130320560A1