Artificial intelligence processor including three-dimensional stack memory

The integration of a three-dimensional stack memory with FeRAM in AI processors addresses latency and power consumption issues by optimizing matrix multiplication and reducing power consumption, thereby enhancing AI system performance.

JP2025118837APending Publication Date: 2025-08-13KEPLER COMPUTING INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2025080359
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-03-18
Filing Date
2025-05-13
Publication Date
2025-08-13

AI Technical Summary

Technical Problem

Existing artificial intelligence (AI) processors face high latency and power consumption due to the hardware-intensive processes of training and applying trained models, necessitating a reduction in computing latency and power consumption.

Method used

The integration of a three-dimensional stack memory with ferroelectric RAM (FeRAM) and computational logic, where FeRAM is used for storing input data and weighting coefficients, and the computational logic applies fixed weights to generate outputs, optimizing matrix multiplication and reducing power consumption.

Benefits of technology

This configuration enhances AI system performance by making matrix multiplication 15-20 times faster and reducing power consumption by an order of magnitude compared to SRAM-based memory, while also reducing interconnect energy and external memory bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025118837000001_ABST
    Figure 2025118837000001_ABST
Patent Text Reader

Abstract

To provide an artificial intelligence processor including a three-dimensional stack memory.SOLUTION: An IC package 300 includes a substrate 301, a first die 303 on the substrate, and a second die 304 stacked on the first die. The first die includes a memory, and the second die includes a calculation logic. The first die includes a RAM such as a ferroelectric RAM (FeRAM) having a bit cell, an SRAM, or a DRAM. Each bit cell includes an access transistor, and a capacitor including a ferroelectric material. The access transistor is bonded to the ferroelectric material. The memory of the first die stores input data and a weighting coefficient. The calculation logic of the second die is bonded to the memory of the first die. The second die is an inference die to apply the fixed weight of the trained model to the input data and generate output. In one example, the second die is a training die to enable learning of weight.SELECTED DRAWING: Figure 3A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Priority claims This application claims priority to U.S. patent application Ser. No. 16 / 357,265, filed March 18, 2019, entitled "Artificial Intelligence Processor with Three-Dimensional Stacked Memory," which is incorporated by reference in its entirety for all purposes.

[0002] This application relates to an artificial intelligence processor that includes a three-dimensional stack memory. [Background technology]

[0003] Artificial intelligence (AI) is a broad area of hardware and software computing that analyzes and classifies data, and then makes decisions about the data. For example, a model describing the classification of data for a specific property or properties is trained over time using large amounts of data. The process of training a model requires large amounts of data and processing power to analyze the data. Once the model is trained, weights or weight coefficients are changed based on the model's output. By repeatedly analyzing data and changing the weights to achieve expected results, the model is considered "trained" when the model weights are calculated to a high confidence level (e.g., 95% or higher). This trained model with fixed weights is then used to make decisions about new data. Training a model and then applying the trained model to new data are hardware-intensive activities. There is a desire to reduce the latency of computing and using the trained model and to reduce the power consumption of such AI processor systems.

[0004] The background discussion provided herein is intended to generally present the context for the present disclosure. Unless otherwise stated herein, the material described in this paragraph is not prior art to the claims of this application, and no admission of prior art is made by inclusion in this paragraph. [Brief explanation of the drawings]

[0005] Embodiments of the present disclosure will be more fully understood from the detailed description given below and the accompanying drawings of various embodiments of the present disclosure, which should not be construed as limiting the disclosure to the particular embodiments, but are for purposes of illustration and understanding only. [Figure 1] 1 illustrates a high-level architecture of an artificial intelligence (AI) machine that includes a compute die positioned on top of a memory die, according to some embodiments. [Figure 2] 1 illustrates the architecture of a compute block including a compute die positioned on top of a memory die, according to some embodiments. [Figure 3A] 1 illustrates a cross section of a package including a compute block according to some embodiments, the compute block including a compute die (e.g., an inference logic die) on top of a memory die. [Figure 3B] 1 illustrates a cross section of a package including a compute block according to some embodiments, the compute block including a compute die (e.g., an inference logic die) on top of a stack of memory dies and controller logic dies. [Figure 3C] 1 illustrates a cross section of a package including a compute block according to some embodiments, the compute block including a compute die on top of a memory that also functions as an interposer. [Figure 3D] 1 illustrates a cross section of a package including a compute block according to some embodiments, the compute block including a compute die between memory dies in a horizontal stack along the plane of the package. [Figure 3E]1 illustrates a cross section of a package including a compute block according to some embodiments, the compute block including a compute die and two or more memories along the plane of the package. [Figure 3F] 1 illustrates a cross section of a package including a compute block including a compute die on an interposer, the interposer including a memory die embedded therein, according to some embodiments. [Figure 3G] 1 illustrates a cross section of a package including a compute block including a compute die and two or more memories along the plane of the package, with the memory also functioning as an interposer, according to some embodiments. [Figure 3H] 1 illustrates a cross section of a package including a compute block according to some embodiments, the compute block including a compute die on top of a 3D ferroelectric memory that also functions as an interposer. [Figure 4A] 1 illustrates a cross section of a package containing an AI machine, including a system-on-chip (SOC) with a computational block, which includes a computational die above memory, according to some embodiments. [Figure 4B] 1 illustrates a cross section of a package containing an AI machine, including a SOC with a compute block, according to some embodiments, the compute block including a compute die over memory, a processor, and solid-state memory. [Figure 5] 1 illustrates a cross section of multiple packages on a circuit board, one of the packages including a compute die above a memory die, and another of the packages including a graphics processing unit, according to some embodiments. [Figure 6] 1 illustrates a cross section of a top view of a compute die with micro-humps on the sides for connecting with memory along horizontal planes, according to some embodiments. [Figure 7] 1 illustrates a cross section of a top view of a computational die with microbumps on the top and bottom of the computational die for connecting with a memory die along the vertical side of the package, according to some embodiments. [Figure 8A] 1 illustrates a cross section of a memory die underlying a compute die, according to some embodiments. [Figure 8B] 1 illustrates a cross section of a compute die overlying a memory die, according to some embodiments. [Figure 9A] 1 illustrates a cross section of a memory die containing 2x2 tiles underneath a compute die, according to some embodiments. [Figure 9B] 1 illustrates a cross section of a compute die containing 2x2 tiles above a memory die, according to some embodiments. [Figure 10] 1 illustrates a method for forming a package. DETAILED DESCRIPTION OF THE INVENTION

[0006] Some embodiments describe packaging techniques for improving the performance of AI processing systems. In some embodiments, an integrated circuit package is provided, the integrated circuit package including a substrate, a first die on the substrate, and a second die stacked on the first die, the first die including memory and the second die including computational logic. In some embodiments, the first die includes a ferroelectric random access memory (FeRAM) having bit cells, each including an access transistor and a capacitor including a ferroelectric material, the access transistor coupled to the ferroelectric material. The FeRAM may be a ferroelectric dynamic random access memory (FeDRAM) or a ferroelectric static random access memory (FeSRAM). The memory of the first die may store input data and weighting coefficients. The computational logic of the second die is coupled to the memory of the first die. The second die may be an inference die that applies fixed weights of a trained model to input data to generate output. In some embodiments, the second die includes a processing core (or processing entity (PE)) having a matrix multiplier, an adder, a buffer, etc. In some embodiments, the first die includes a high-bandwidth memory (HBM), which may include a controller and a memory array.

[0007] In some embodiments, the second die includes an application specific integrated circuit (ASIC) that can train the model by changing the weights and also use the model on new data with fixed weights. In some embodiments, the memory includes SRAM (static random access memory). In some embodiments, the memory of the first die includes MRAM (magnetic random access memory). In some embodiments, the memory of the first die includes Re-RAM (resistive random access memory). In some embodiments, the substrate is an active interposer, and the first die is embedded in the active interposer. In some embodiments, the first die itself is an active interposer.

[0008] In some embodiments, the integrated circuit package is a package for a system-on-chip (SOC). The SOC may include a computational die on top of a memory die (HBM) and a processor die coupled to a memory die adjacent to the computational die (e.g., on top of or to the side of the processor die). In some embodiments, the SOC includes a solid-state memory die.

[0009] The packaging techniques of various embodiments have many technical advantages. For example, placing a memory die under a compute die or placing one or more memory dies to the side of a compute die improves the performance of an AI system. In some embodiments, using Fe-RAM for memory makes the matrix multiplication process by the compute die 15-20 times faster than traditional matrix multiplication. Furthermore, using Fe-RAM reduces the power consumption of an AI system by an order of magnitude compared to SRAM-based memory. Using Fe-RAM reduces interconnect energy, reduces external memory bandwidth requirements, reduces circuit complexity, and reduces the cost of the computing system. Other technical advantages will be apparent from the various embodiments and figures.

[0010] In the following description, numerous details are discussed to provide a more thorough explanation of embodiments of the present disclosure. However, it will be apparent to those skilled in the art that embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring embodiments of the present disclosure.

[0011] Note that in the corresponding drawings of the embodiments, signals are represented by lines. Some lines may be thicker to indicate more constituent signal paths and / or include arrows at one or more ends to indicate the direction of primary information flow. Such indications are not intended to be limiting. Rather, the lines are used in connection with one or more exemplary embodiments to facilitate easier understanding of a circuit or logic unit. As dictated by design needs or preferences, the represented signals may actually include one or more signals that may travel in either direction and may be implemented with any suitable type of signaling scheme.

[0012] The term "device" may generally refer to an apparatus depending on the context of the use of the term. For example, a device may refer to a stack of layers or structures, a single structure or layer, a connection of various structures with active and / or passive elements, etc. Generally, a device is a three-dimensional structure that includes a plane along the xy direction of an xyz Cartesian coordinate system and a height along the z direction. The plane of a device may also be the plane of an instrument that includes the device.

[0013] Throughout this specification and in the claims, the term "connected" means a direct connection, such as an electrical, mechanical, or magnetic connection, between the things connected, without any intermediate devices.

[0014] The term "coupled" means a direct or indirect connection, such as a direct electrical, mechanical, or magnetic connection between the things connected, or an indirect connection through one or more passive or active intermediary devices.

[0015] As used herein, the term "adjacent" generally refers to the location of one thing next to (e.g., immediately adjacent to or nearby with one or more things between them) or adjacent to (e.g., abutting) another thing.

[0016] The term "circuit" or "module" may refer to one or more passive and / or active components arranged to cooperate with each other to provide a desired functionality.

[0017] The term "signal" may refer to at least one current signal, voltage signal, magnetic signal, or data / clock signal. The meanings of "a," "an," and "the" include plural references. The meaning of "in" includes "in" and "on."

[0018] The term "scaling" generally refers to converting a design (schematic and layout) from one process technology to another, followed by reducing the layout area. The term "scaling" also generally refers to downsizing of layouts and devices within the same technology node. The term "scaling" can also refer to adjusting (e.g., slowing down or speeding up—i.e., scaling down or scaling up, respectively) another parameter, for example, signal frequency relative to power supply levels.

[0019] The terms "substantially," "close," "approximately," "near," and "about" generally refer to within + / - 10% of a target value. For example, unless otherwise specified by the express context of their use, the terms "substantially equal," "approximately equal," and "approximately equal" refer to only incidental variations between those so described. In the art, such variations are typically no more than + / - 10% of a given target value.

[0020] The use of ordinal adjectives such as "first," "second," and "third" to describe a common object merely indicates that different instances of the same object are being referred to, unless otherwise specified, and is not intended to imply that the objects so described need be in a particular order, temporally, spatially, ranked, or otherwise.

[0021] For purposes of this disclosure, the phrases "A and / or B" and "A or B" mean (A), (B), or (A and B). For purposes of this disclosure, the phrase "A, B, and / or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).

[0022] In the detailed description and claims, terms such as "left," "right," "front," "rear," "top," "bottom," "above," "below," and the like (if any) are used for descriptive purposes and not necessarily to describe permanent relative positions. For example, as used herein, the terms "above," "below," "front," "rear," "top," "bottom," "above," "below," and "on" refer to the relative position of a component, structure, or material within a device relative to another referenced component, structure, or material. Such physical relationships are noteworthy. These terms are used herein for descriptive purposes only and primarily in the context of the device's z-axis, and therefore may relate to the orientation of the device. Thus, a first material "above" a second material in the context of the figures provided herein could also be "below" the second material if the device were oriented upside down relative to the context of the figures provided. In the context of materials, one material disposed above or below another material may be in direct contact or may have one or more intervening materials. Furthermore, a material disposed between two materials may be in direct contact with the two layers or may have one or more intervening layers. In contrast, a first material "on" a second material is in direct contact with that second material. Similar distinctions can be made in the context of component assembly.

[0023] The term "between" may be used in the context of the z-axis, x-axis, or y-axis of a device. A material that is between two other materials may be in contact with one or both of those materials, or may be separated from both of the other two materials by one or more intervening materials. Thus, a material that is "between" two other materials may be in contact with either of the other two materials, or may be coupled to the other two materials via an intervening material. A device that is between two other devices may be directly connected to one or both of those devices, or may be separated from both of the other two devices by one or more intervening devices.

[0024] Here, the term "back end" generally refers to the section of a die opposite the "front end" where an IC (integrated circuit) package is bonded to the IC die bumps. For example, higher-level metal layers (e.g., metal layers 6 or higher in a 10-metal stack die) and corresponding vias near the die package are considered part of the back end of the die. Conversely, the term "front end" generally refers to the section of the die that includes the active area (e.g., where transistors are fabricated) and the lower-level metal layers and corresponding vias near the active area (e.g., metal layers 5 or lower in a 10-metal stack die example).

[0025] It is pointed out that elements of a figure having the same reference number (or name) as elements of another figure may operate or function in a similar manner as described, but is not limited to such.

[0026] FIG. 1 illustrates a high-level architecture of an artificial intelligence (AI) machine 100, including a compute die positioned on top of a memory die, according to some embodiments. The AI machine 100 includes a compute block 101 or processor with random access memory (RAM) 102 and computational logic 103, static random access memory (SRAM) 104, a main processor 105, dynamic random access memory (DRAM) 106, and a solid-state memory or drive (SSD) 107. In some embodiments, some or all components of the AI machine are packaged in a single package to form a system-on-chip (SOC). In some embodiments, the compute block 101 is packaged in a single package and then coupled to the processor 105 and memories 104, 106, and 107 on a printed circuit board (PCB). In various embodiments, the compute block 101 includes a dedicated compute die 103 or microprocessor. In some embodiments, the RAM 102 is ferroelectric RAM (Fe-RAM), which forms a special memory / cache for the dedicated compute die 103. In some embodiments, the compute die 103 is specialized for applications such as artificial intelligence, graph processing, and data processing algorithms. In some embodiments, the compute die 103 further includes logic computation blocks, e.g., for multipliers and buffers, and specialized data memory blocks (e.g., buffers), including Fe-RAM. In some embodiments, the FE-RAM 102 stores weights and inputs in order to improve computational efficiency. The interconnections between the processor 105 or dedicated processor 105, the FE-SRAM 104, and the compute die 103 are optimized for high bandwidth and low latency. The architecture of FIG. 1 enables efficient packaging, resulting in energy / power / cost savings.

[0027] In some embodiments, RAM 102 includes SRAM partitioned to store input data (or data to be processed) 102a and weighting factors 102b. In some embodiments, RAM 102 includes Fe-RAM. For example, RAM 102 includes FE-DRAM or FE-SRAM. In some embodiments, input data 103a is stored in a separate memory (e.g., a separate memory die), and weighting factors 102b are stored in a separate memory (e.g., a separate memory die).

[0028] In some embodiments, the computation logic 103 includes a matrix multiplier, an adder, concatenation logic, a buffer, and combination logic. In various embodiments, the computation logic 103 performs a multiplication operation on the inputs 102a and the weights 102b. In some embodiments, the weights 102b are fixed weights. For example, the processor 105 (e.g., a graphics processing unit (GPU), an AI processor, a central processing unit (CPU), or any other high-performance processor) calculates the weights of the training model. Once the weights are calculated, they are stored in the memory 102b. In various embodiments, input data to be analyzed using the trained model is processed using the weights 102b calculated by the computation block 101 to generate an output (e.g., a classification result).

[0029] In some embodiments, SRAM 104 is a ferroelectric-based SRAM. For example, a six-transistor (6T) SRAM bit cell with a ferroelectric transistor is used to implement a non-volatile Fe-SRAM. In some embodiments, SSD 107 includes NAND flash cells. In some embodiments, SSD 107 includes NOR flash cells. In some embodiments, SSD 107 includes multi-threshold NAND flash cells.

[0030] In various embodiments, non-volatile Fe-RAM is used to introduce new features such as security, functional safety, and faster restart times for architecture 100. Non-volatile Fe-RAM is a low-power RAM that provides fast access to data and weights. Fe-RAM 104 also serves as fast storage for inference die 101 (or accelerators), which typically have low capacity and fast access requirements.

[0031] In various embodiments, the Fe-RAM (Fe-DRAM or Fe-SRAM) includes a ferroelectric material. The ferroelectric (FE) material can be in the gate stack of a transistor or the capacitor of a memory. The ferroelectric material can be any suitable low-voltage FE material that allows the FE material to switch its state with a low voltage (e.g., 100 mV). In some embodiments, the FE material includes a perovskite of type ABO3, where "A" and "B" are two cations of different sizes, and "O" is oxygen, an anion that bonds to both cations. Generally, the size of the A atom is larger than the size of the B atom. In some embodiments, the perovskite can be doped (e.g., with La or a lanthanide). In various embodiments, when the FE material is a perovskite, the conductive oxide is of type AA'BB'O3. A' is a dopant at atomic site A and can be an element of the lanthanide series. B' is a dopant for atomic site B and can be a transition metal element, especially Sc, Ti, V, Cr, Mn, Fe, Co, Ni, Cu, or Zn. A' has the same valence as site A and may have a different ferroelectric polarizability.

[0032] In some embodiments, the FE material comprises a hexagonal ferroelectric of type h-RMnO3, where R is a rare earth element, i.e., cerium (Ce), dysprosium (Dy), erbium (Er), europium (Eu), gadolinium (Gd), holmium (Ho), lanthanum (La), lutetium (Lu), neodymium (Nd), praseodymium (Pr), promethium (Pm), samarium (Sm), scandium (Sc), terbium (Tb), thulium (Tm), ytterbium (Yb), and yttrium (Y). The ferroelectric phase is characterized by buckling of layered MnO5 polyhedra with displacement of Y ions, which causes a net electric polarization. In some embodiments, the hexagonal FE comprises one of YMnO3 or LuFeO3. In various embodiments, when the FE material comprises a hexagonal ferroelectric, the conductive oxide is of the A2O3 (e.g., In2O3, Fe2O3) and ABO3 type, where "A" is a rare earth element and B is Mn.

[0033] In some embodiments, the FE material includes an unsuitable FE material. An unsuitable ferroelectric is one whose primary order parameter is an ordering mechanism such as distortion or buckling of atomic order. Examples of unsuitable FE materials include the LuFeO3 class of materials, or superlattices of ferroelectric and paraelectric materials (PbTiO3 (PTO) and SnTiO3 (STO), respectively, and LaAlO3 (LAO) and STO, respectively). For example, a [PTO / STO]n or [LAO / STO]n superlattice, where "n" is between 1 and 100. While various embodiments herein are described with respect to ferroelectric materials for storing charge states, the embodiments are also applicable to paraelectric materials.

[0034] 2 illustrates the architecture of a compute block 200 (e.g., 101) that includes a compute die positioned on top of a memory die, according to some embodiments. The architecture in FIG. 2 illustrates a dedicated compute die architecture where RAM memory buffers for inputs and weights are split on die 1 and logic and optional memory buffers are split on die 2.

[0035] In some embodiments, the memory die (e.g., die 1) is positioned below the computational die (e.g., die 2) such that a heat sink or thermal solution is adjacent to the computational die. In some embodiments, the memory die is embedded in an interposer. In some embodiments, the memory die operates as an interposer in addition to its basic memory functionality. In some embodiments, the memory die is a high-bandwidth memory (HBM) that includes multiple dies of memory in a stack and a controller for controlling read and write functions to the stack of memory dies. In some embodiments, the memory die includes a first die 201 for storing input data and a second die 202 for storing weighting factors. In some embodiments, the memory die is a single die that is partitioned such that a first partition 201 of the memory die is used to store input data and a second partition 202 of the memory die is used to store weights. In some embodiments, the memory die includes FE-DRAM. In some embodiments, the memory die includes FE-SRAM. In some embodiments, the memory die includes MRAM. In some embodiments, the memory die includes SRAM. For example, memory partitions 201 and 202 or memory dies 201 and 202 include one or more of FE-SRAM, FE-DRAM, SRAM, and / or MRAM. In some embodiments, the input data stored in memory partition or die 201 is data that is analyzed by a trained model using fixed weights stored in memory partition or die 202.

[0036] In some embodiments, the computation die includes a matrix multiplier 203, logic 204, and a temporary buffer 205. The matrix multiplier 203 performs a multiplication operation on input data “X” and weights “W” to generate an output “Y.” This output may be further processed by the logic 204. In some embodiments, the logic 204 performs thresholding operations, pooling and dropout operations, and / or concatenation operations to complete AI logic primitive functions. In some embodiments, the output of the logic 204 (e.g., the processed output “Y”) is temporarily stored in a buffer 205. In some embodiments, the buffer 205 is a memory such as one or more of Fe-SRAM, Fe-DRAM, MRAM, Resistive RAM (Re-RAM), and / or SRAM. In some embodiments, the buffer 205 is part of a memory die (e.g., die 1). In some embodiments, the buffer 205 performs the function of a re-timer. In some embodiments, the output of buffer 205 (e.g., processed output "Y") is used to modify weights within memory partition or die 202. In one such embodiment, computation block 200 operates not only as an inference circuit but also as a training circuit for training the model. In some embodiments, matrix multiplier 203 includes an array of multiplier cells, and FeRAMs 201 and 202 each include an array of memory bit cells, with each multiplier cell coupled to a corresponding memory bit cell in FE-RAM 201 and FE-RAM 202. In some embodiments, computation block 200 includes interconnect fibers coupled to the array of multiplier cells, such that each multiplier cell is coupled to an interconnect fiber.

[0037] Architecture 200 provides reduced memory accesses for the compute die (e.g., die 2) by providing data locality for weights, inputs, and outputs. In one example, data to and from the AI compute block (e.g., matrix multiplier 203) is processed locally within the same packaging unit. Architecture 200 also separates memory and logic operations onto a memory die (e.g., die 1) and a logic die (e.g., die 2), respectively, enabling optimized AI processing. Desegregated die increases die yield. Die 1's large memory process also reduces the power of the external interconnect to memory, reducing integration costs and enabling a smaller footprint.

[0038] FIG. 3A shows a cross section of a package 300 including a compute block according to some embodiments, where the compute block includes a compute die (eg, an inference logic die) on top of a memory die.

[0039] In some embodiments, an integrated circuit (IC) package assembly is coupled to circuit board 301. In some embodiments, circuit board 301 can be a printed circuit board (PCB) constructed from an electrically insulating material such as an epoxy laminate. For example, circuit board 301 can include an electrically insulating layer constructed from a material such as a phenolic cotton paper material (e.g., FR-1), a cotton paper and epoxy material (e.g., FR-3), a woven glass material laminated together using epoxy resin (FR-4), a glass / paper with epoxy resin (e.g., CEM-1), a glass composite with epoxy resin, a glass fabric with polytetrafluoroethylene (e.g., PTFE CCL), or other polytetrafluoroethylene-based prepreg material. In some embodiments, layer 301 is a package substrate and is part of the IC package assembly.

[0040] The IC package assembly may include a substrate 302, a memory die 303 (e.g., die 1 in FIG. 2 ), and a computation die 304 (e.g., die 2 in FIG. 2 ). In various embodiments, the memory die 303 is located below the computation die 304. This particular topology improves the overall performance of the AI system. In various embodiments, the computation die 304 includes the logic portion of the inference die. The inference die or chip is used to apply fixed weights and inputs associated with trained models to generate outputs. Separating the memory 3003 associated with the inference die 304 improves AI performance. Furthermore, such a topology allows for better use of thermal solutions, such as a heat sink 315, which dissipates heat from power dissipation sources, such as the inference die 304. In various embodiments, the memory 303 may be one or more of FE-SRAM, FE-DRAM, SRAM, MRAM, resistive RAM (Re-RAM), or a combination thereof. Using FE-SRAM, MRAM, or Re-RAM enables low-power, high-speed memory operations. This allows the memory die 303 to be placed below the compute die 304 to more efficiently use the thermal solution for the compute die 304. In some embodiments, the memory die 303 is a high bandwidth memory (HBM).

[0041] In some embodiments, the computational die 304 is an application specific circuit (ASIC), a processor, or some combination of such functionality. In some embodiments, one or both of the memory die 303 and the computational die 304 may be embedded in an encapsulation material 318. In some embodiments, the encapsulation material 318 may be any suitable material, such as an epoxy-based build-up substrate, other dielectric / organic materials, resin, epoxy, polymer adhesive, silicone, acrylic, polyimide, cyanate ester, thermoplastic resin, and / or thermoset resin.

[0042] In some embodiments, the memory die 303 may have a first side S1 and a second side S2 opposite the first side S1. In some embodiments, the first side S1 may be the side of the die commonly referred to as the “inactive” or “back” side of the die. In some embodiments, the back side of the memory die 303 may include active or passive devices, signal and power routing, etc. In some embodiments, the second side S2 may include one or more transistors (e.g., access transistors) and is the side of the die commonly referred to as the “active” or “front” side of the die. The memory circuitry of some embodiments may also have active and passive devices on the front side of the die. In some embodiments, the second side S2 of the memory die 303 may include one or more electrical routing features 310. In some embodiments, the computational die 304 may include an “active” or “front” side with one or more electrical routing features 312. In some embodiments, the electrical routing features 312 may be bond pads, microbumps, solder balls, or any other suitable bonding technique.

[0043] In some embodiments, the memory die 302 may include one or more through-silicon vias (TSVs) that couple the substrate 302 to the computational die 304 via the electrical routing features 312. For example, the computational die 304 is coupled to the memory die 303 by a die interconnect. In some embodiments, the inter-die interconnect may be a solder bump, a copper pillar, or other conductive feature. In some embodiments, an interface layer (not shown) may be provided between the memory die 303 and the computational die 304. The memory die 303 may be coupled to the computational die 304 using TSVs. In some embodiments, interconnect pillars with corresponding solder balls are used to connect the memory die 303 to the computational die 304. In some embodiments, the interface layer (not shown) may be or include an underfill, adhesive, dielectric, or other layer of material. In some embodiments, the interface layer may perform various functions, such as providing mechanical strength, conductivity, heat dissipation, or adhesion.

[0044] In some embodiments, package substrate 303 may be a coreless substrate. For example, package substrate 302 may be a "bumpless" build-up layer (BBUL) assembly including multiple "bumpless" build-up layers. Here, the term "bumpless build-up layer" generally refers to a substrate and layers of components embedded therein, without the use of solder or other attachment means that may be considered "bumps." However, various embodiments are not limited to BBUL-type connections between the die and the substrate and may be used with any suitable flip-chip substrate. In some embodiments, one or more build-up layers may have material properties that may be modified and / or optimized for reliability, warpage reduction, etc. In some embodiments, package substrate 504 may be composed of a polymer, ceramic, glass, or semiconductor material. In some embodiments, package substrate 302 may be a conventional core substrate and / or interposer. In some embodiments, package substrate 302 includes active and / or passive devices embedded therein.

[0045] In some embodiments, a top side of the package substrate 302 is coupled to the second surface S2 of the memory die 303 and / or the electrical routing feature 310. In some embodiments, an opposite bottom side of the package substrate 302 is coupled to the circuit board 301 by a package interconnect 317. In some embodiments, the package interconnect 316 can couple the electrical routing feature 317 located on the second side of the package substrate 304 to a corresponding electrical routing feature 315 on the circuit board 301.

[0046] In some embodiments, package substrate 504 may have electrical routing features formed therein for routing electrical signals between memory die 303 (and / or computational die 304) and circuit board 301 and / or other electrical components external to the IC package assembly. In some embodiments, package interconnect 316 and die interconnect 310 comprise any of a wide variety of suitable structures and / or materials, including, for example, bumps, pillars, or balls formed using metals, alloys, solderable materials, or combinations thereof. In some embodiments, electrical routing features 315 may be arranged in a ball grid array (BGA) or other configuration.

[0047] In some embodiments, the computational die 304 is coupled to the memory die 303 in a front-to-back configuration (e.g., the “front” or “active” side of the computational die 303 is coupled to the “back” or “inactive” side S1 of the memory die 303). In some embodiments, the dies may be coupled to each other in a front-to-front, back-to-back, or side-to-side arrangement. In some embodiments, one or more additional dies may be coupled to the memory die 303, the computational die 304, and / or the package substrate 302. In some embodiments, the IC package assembly may include a multi-chip package configuration including, for example, a combination of flip-chip and wire bonding technologies, interposers, system-on-chip (SOC) and / or package-on-package (PoP) configurations for routing electrical signals.

[0048] In some embodiments, memory die 303 and computational die 304 may be a single die. In some embodiments, memory die 303 is an HBM including two or more dies, where the two or more dies include a controller die and a memory die. In some embodiments, the computational die may further include two dies. For example, buffer 205 may be a separate memory die coupled near surface S1 of memory die 303, and matrix multiplication and other computational units may be in separate dies. In one example, memory die 303 and / or computational die 304 may be a wafer (or portion of a wafer) having two or more dies formed thereon. In some embodiments, memory die 303 and / or computational die 304 include two or more dies embedded in encapsulation material 318. In some embodiments, the two or more dies are positioned side-by-side, vertically stacked, or in any other suitable arrangement.

[0049] In various embodiments, a heat sink 315 and associated fins are coupled to the compute die 304. While a heat sink 315 is shown as the thermal solution, other thermal solutions may be used. For example, a fan, liquid cooling, etc. may be used in addition to or instead of the heat sink 315.

[0050] FIG. 3B illustrates a cross section of a package 320 including a computational block according to some embodiments, where the computational block includes a computational die (e.g., an inference logic die) on top of a stack of memory dies and a controller logic die. To avoid obscuring the embodiments of package 320, differences between packages 300 and 320 will be discussed. Here, memory die 303 is replaced with a controller die 323 and a stack of memory dies (RAMs) 324a and 324b. In some embodiments, controller die 323 is a memory controller that includes read logic, write logic, column and row multiplexers, error correction logic, an interface with RAMs 324a / b, an interface with computational die 304, and an interface with substrate 302. In various embodiments, memory dies 324a / b are disposed or stacked on top of controller die 323. In some embodiments, RAM 324a / b is one or more of FE-SRAM, FE-DRAM, SRAM, MRAM, Re-RAM, or a combination thereof. In some embodiments, RAM die 324a is used to store inputs, while RAM die 324b is used to store weights. In some embodiments, either RAM die 324a / b can include memory for buffer 205. Although the embodiment of FIG. 3B shows two RAM dies, any number of RAM dies can be stacked on top of controller die 323.

[0051] 3C shows a cross section of a package 330 including a compute block according to some embodiments, where the compute block includes a compute die 304 on top of a memory that also functions as an interposer. Compared to package 300, here the memory die 303 is removed and integrated into interposer 332, so that the memory provides not only the storage function but also the interposer function. This configuration can reduce the cost of the package. Here, interconnects 310 electrically couple the compute die 304 to the memory 332. The memory 332 can include FE-SRAM, FE-DRAM, SRAM, MRAM, Re-RAM, or a combination thereof.

[0052] FIG. 3D shows a cross section of a package 340 including a computational block according to some embodiments, where the computational die is located between memory dies in a horizontal stack along the plane of the package. Compared to package 300, here computational die 304 is positioned between memories 343 and 345, and RAM die 343 is coupled to substrate 302 via interconnect 310. In various embodiments, computational die 304 communicates with RAM dies 343 and 345 through both its front and back sides via interconnects 311a and 311b, respectively. This embodiment allows computational die 304 to efficiently use its real estate by applying active devices to its front and back ends. RAM dies 343 / 345 may include FE-SRAM, FE-DRAM, SRAM, MRAM, Re-RAM, or a combination thereof. In some embodiments, RAM die 343 is used to store inputs, while RAM die 345 is used to store weights. In some embodiments, either RAM die 343 or 345 can include memory for buffer 205. Although the embodiment of Figure 3D shows two RAM dies, any number of RAM dies can be stacked above and below compute die 304.

[0053] FIG. 3E shows a cross section of a package 350 including a compute block according to some embodiments, where the compute block includes a compute die and two or more memories along the plane of the package. Compared to package 300, here, the compute die 304 is in the center, and memory dies 354 and 355 are on either side of the compute die 304. In some embodiments, the memory die surrounds the compute die 304. AI processing is memory-intensive. Such an embodiment allows the compute die 304 to access memory from its four sides. In this case, a heat sink 315 is coupled to the memory dies 354 and 355 and the compute die 304. The RAM dies 354 and 355 may include FE-SRAM, FE-DRAM, SRAM, MRAM, Re-RAM, or a combination thereof. The RAM dies 354 and 355 may include HBM. Each HBM includes two or more memory dies and a controller. In some embodiments, the RAM die 354 is used to store inputs, while the RAM die 355 is used to store weights. In some embodiments, either RAM die 354 or 355 can include memory for buffer 205. Although the embodiment of FIG. 3E shows two RAM dies, any number of RAM dies can be positioned along the sides of compute die 304.

[0054] FIG. 3F illustrates a cross section of a package 360 including a compute block with a compute die on an interposer containing a memory die embedded therein, according to some embodiments. Compared to package 300, here the memory die 363 is embedded in the substrate or interposer 302. This embodiment allows for a reduced z-height of the package and also reduces latency between the compute die 304 and other devices coupled to the substrate 301. The RAM die 363 may include FE-SRAM, FE-DMAM, SRAM, MRAM, Re-RAM, or a combination thereof. The RAM die 363 may include an HBM. Each HBM includes two or more memory dies and a controller. While the embodiment of FIG. 3F illustrates one RAM die 363, any number of RAM dies can be embedded in the interposer 302.

[0055] FIG. 3G shows a cross section of package 370 including a compute block including a compute die and two or more memories along the plane of the package, with the memory also functioning as an interposer, according to some embodiments. Compared to package 350, here, memories 374 and 375 on the sides of the compute die are RAM (e.g., SRAM, Fe-RAM, MRAM, or Re-RAM). In various embodiments, interposer 302 is replaced with memory acting as an interposer. The memory can be either Fe-RAM, MRAM, Re-RAM, or SRAM. In some embodiments, the memory in the interposer is a three-dimensional (3D) Fe-RAM stack that also functions as an interposer. In some embodiments, the 3D memory stack is a stack of MRAM, Re-RAM, or SRAM.

[0056] 3H shows a cross section of a package 380 including a compute block according to some embodiments, where the compute block includes a compute die on top of a 3D ferroelectric memory that also functions as an interposer. Compared to package 330, in various embodiments, memory interposer 332 is replaced with a three-dimensional (3D) Fe-RAM stack that also functions as an interposer. In some embodiments, the 3D memory stack is a stack of MRAM, Re-RAM, or SRAM.

[0057] FIG. 4A illustrates a cross section of a package 400 including an AI machine, including a system-on-chip (SOC) with a computation block, which includes a computation die on top of a memory, according to some embodiments. The package 400 includes a processor die 406 coupled to a substrate or interposer 302. Two or more memory dies 407 (e.g., memory 104) and 408 (e.g., memory 106) are stacked on the processor die 406. The processor die 406 (e.g., 105) can be a central processing unit (CPU), a graphics processing unit (GPU), or an application-specific integrated circuit (ASIC). The memory (RAM) dies 407 and 408 can include FE-SRAM, FE-DRAM, SRAM, MRAM, Re-RAM, or a combination thereof. In some embodiments, the RAM dies 407 and 408 can include HBM. In some embodiments, one of the memories 104 and 106 is implemented as HBM within the die 405. The memory in the HBM die 405 includes any one or more of FE-SRAM, FE-DRAM, SRAM, MRAM, Re-RAM, or a combination thereof. A heat sink 315 provides a thermal management solution for the various dies within an encapsulation material 318. In some embodiments, a solid-state drive (SSD) 409 is positioned outside the first package assembly, which includes the heat sink 315. In some embodiments, the SSD 409 includes one of NAND flash memory, NOR flash memory, or any other type of non-volatile memory, such as MRAM, FE-DRAM, FE-SRAM, Re-RAM, etc.

[0058] 4B shows a cross section of a package 420 including an AI machine, including a SOC with a compute block, which includes a compute die on top of memory, a processor, and solid-state memory, according to some embodiments. Package 420 is similar to package 400, but for incorporating SSD 409 in a single package under a common heat sink 315. In this case, the single packaged SOC provides an AI machine that includes the ability to generate a training model and use the training model on different data to generate outputs.

[0059] 5 illustrates a cross section 500 of multiple packages on a circuit board according to some embodiments, where one of the packages includes a compute die above a memory die and another of the packages includes a graphics processing unit. In this example, an AI processor, such as a CPU 505, is coupled to a substrate 201 (e.g., a PCB). Two packages are shown here, one with a heat sink 506 and the other with a heat sink 507. Heat sink 506 is a dedicated thermal solution for the GPU chip 505, while heat sink 507 provides a thermal solution for the compute blocks (dies 303 and 304) that include HBM 305.

[0060] 6 shows a cross section of a top view 600 of a computational die 304 with micro-humps on the sides for connecting to memory along horizontal planes, according to some embodiments. Shaded areas 601 and 602 on either side of the computational die 304 include micro-bumps 603 (e.g., 310) used to connect to memory on either side of the computational die 304. For example, as shown in FIG. 3E, HBMs 354 and 355 are coupled to the computational die 304 via micro-bumps 603. Micro-bumps 604 can be used to connect to the substrate 302 or the interposer 302.

[0061] 7 illustrates a cross section of a top view 700 of a computational die 304 having microbumps on the top and bottom of the computational die for connecting with memory dies along the vertical plane of the package, according to some embodiments. Shaded regions 701 and 702 on the upper and lower sections of the computational die 304 include microbumps 703 (e.g., 311a and 311b) used to connect to the upper and lower memory dies 345 and 343, respectively. For example, as shown in FIG. 3E, FE-RAMs 343 and 345 are coupled to the computational die 304 via microbumps 311a and 311b, respectively. Microbumps 704 can be used to connect to the substrate 302 or the interposer 302.

[0062] 8A shows a cross section 800 of a memory die (e.g., 303 or 333) underlying a compute die 304, according to some embodiments. The pitch of memory die 303 is L×W. Cross section 800 shows the strips of TSVs used to connect to compute die 304. Shaded strip 801 carries signals, while strips 802 and 803 carry power and ground lines. Strip 804 provides power and ground signals 805 and 806 to memory cells in a row. TSV 808 connects signals (e.g., word lines) to the memory bit cells.

[0063] 8B shows a cross section 820 of a compute die (e.g., 304) overlying a memory die (e.g., 303), according to some embodiments. TSV 828 can be coupled to TSV 808, and strip 824 is overlying strip 804. TSVs 825 and 826 are coupled to TSVs 805 and 806, respectively.

[0064] Figure 9A shows a cross section 900 of a memory die 303 below the compute die, including 2x2 tiles, according to some embodiments. While memory die 202 in Figure 8A shows a single tile, here 2x2 tiles are used to organize the memory. This allows for a clean partitioning of the memory for storing data and weights. Here, the tiles are shown by tile 901. Embodiments are not limited to 2x2 tile and MxN tile configurations (where M and N are integers that may be equal or different).

[0065] 9B illustrates a cross section 920 of a compute die including 2x2 tiles on top of a memory die, according to some embodiments. Similar to memory 303, compute die 304 can also be partitioned into tiles. Each tile 921 is like compute die 304 of FIG. 8B, according to some embodiments. Such an organization of compute die 304 allows different training models to be run simultaneously or in parallel using different input data and weights.

[0066] 10 illustrates a flowchart 1000 of a method for forming a package of computational blocks including a computational die (e.g., an inference logic die) on a memory die, according to some embodiments. The blocks of flowchart 1000 are shown in a particular order. However, the order of various processing steps can be changed without changing the essence of the embodiment. For example, some processing blocks can be processed concurrently, while other blocks can be performed out of order.

[0067] In block 1001, a substrate (e.g., 302) is formed. In some embodiments, the substrate 302 is a packaging substrate. In some embodiments, the substrate 302 is an interposer (e.g., an active or passive interposer). In block 1002, a first die (e.g., 303) is formed on the substrate. In some embodiments, forming the first die includes a ferroelectric random access memory (FeRAM) having bit cells, each bit cell including an access transistor and a capacitor including a ferroelectric material, the access transistor coupled to the ferroelectric material. In block 1003, a second die (e.g., computation die 304) is formed and stacked on the first die, and forming the second die includes forming computation logic coupled to the memory of the first die. In some embodiments, forming the computation logic includes forming an array of multiplier cells, where the FeRAM includes an array of memory bit cells.

[0068] At block 1004, an interconnect fiber is formed. At block 1005, the interconnect fiber is coupled to the array of multiplier cells such that each multiplier cell is coupled to an interconnect fiber. In some embodiments, the FeRAM is partitioned into a first partition operable as a buffer and a second partition for storing weight coefficients.

[0069] In some embodiments, the method of flowchart 1000 includes computational logic receiving data from the first partition and the second partition, and providing an output of the computational logic to a logic circuit. In some embodiments, forming the computational logic includes forming ferroelectric logic. In some embodiments, the computational logic is operable to multiply at least two matrices. In some embodiments, the method of forming a substrate includes forming an active or passive device. In some embodiments, the method includes forming a third die (e.g., a logic die or a memory) on the substrate. In some embodiments, the method includes bonding the third die to the substrate. In some embodiments, the method includes a fourth die including dynamic random access memory (DRAM), and the method includes stacking the fourth die on the third die. In some embodiments, the method includes bonding a heat sink to the second die.

[0070] References herein to "an embodiment," "one embodiment," "some embodiments," or "other embodiments" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least some embodiments, but not necessarily in all embodiments. Various occurrences of "an embodiment," "one embodiment," or "some embodiments" do not necessarily all refer to the same embodiments. When the specification describes a component, feature, structure, or characteristic as "may include," "may include," or "can include," that particular component, feature, structure, or characteristic need not be included. When the specification or claims refer to "a" or "an" element, this does not mean that there is only one element. When the specification or claims refer to "additional" elements, this does not exclude the presence of multiple additional elements.

[0071] Furthermore, particular features, structures, functions, or characteristics may be combined in any suitable manner in one or more embodiments. For example, a first embodiment may be combined with a second embodiment whenever particular features, structures, functions, or characteristics associated with the two embodiments are not mutually exclusive.

[0072] While the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of such embodiments will be apparent to those skilled in the art in light of the foregoing description. The embodiments of the present disclosure are intended to embrace all such alternatives, modifications, and variations as fall within the scope of the appended claims.

[0073] Additionally, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the presented figures for ease of illustration and description, and so as not to obscure the present disclosure. Furthermore, arrangements may be shown in block diagram form to avoid obscuring the present disclosure and in recognition of the fact that details regarding the implementation of such block diagram arrangements will vary greatly depending on the platform on which the present disclosure is implemented (i.e., such details should be well within the purview of one skilled in the art). Where specific details (e.g., circuits) are shown to illustrate exemplary embodiments of the present disclosure, it should be apparent to one skilled in the art that the present disclosure can be practiced without or with these specific details altered. Thus, the description should be considered illustrative, rather than limiting.

[0074] The following examples are provided to illustrate various embodiments. These examples can be combined with other examples. Thus, various embodiments can be combined with other embodiments without changing the scope of the invention.

[0075] Example 1: An apparatus including: a substrate; a first die on the substrate, the first die including ferroelectric random access memory (FeRAM) having bit cells, each bit cell including an access transistor and a capacitor including a ferroelectric material, the access transistor coupled to the ferroelectric material; and a second die stacked on the first die, the second die including computational logic coupled to the memory of the first die.

[0076] Example 2: The apparatus of Example 1, wherein the computational logic includes an array of multiplier cells and the FeRAM includes an array of memory bit cells.

[0077] Example 3: The apparatus of Example 2, including an interconnecting fiber coupled to the array of multiplier cells, each multiplier cell coupled to an interconnecting fiber.

[0078] Example 4: The apparatus of example 1, wherein the memory is partitioned into a first partition operable as a buffer and a second partition for storing weighting factors.

[0079] Example 5: The apparatus of Example 4, wherein the computation logic can receive data from the first partition and the second partition, and an output of the computation logic is received by the logic circuit.

[0080] Example 6: The device of example 4, wherein the computational logic includes ferroelectric logic.

[0081] Example 7: The apparatus of example 4, wherein the computational logic is operable to multiply at least two matrices.

[0082] Example 8: The apparatus of example 1, wherein the substrate comprises an active or passive device.

[0083] Example 9: The apparatus of Example 1, wherein a third die is bonded onto the substrate, and a fourth die including a dynamic random access memory (DRAM) is stacked on top of the third die.

[0084] Example 10: The apparatus of Example 1, wherein a heat sink is coupled to the second die.

[0085] Example 11: A method, the method including: forming a substrate; forming a first die on the substrate, the forming the first die including a ferroelectric random access memory (FeRAM) having bit cells, each bit cell including an access transistor and a capacitor including a ferroelectric material, the access transistor coupled to the ferroelectric material; and forming a second die stacked on the first die, the forming the second die including forming computational logic coupled to the memory of the first die.

[0086] Example 12: The method of Example 11, wherein forming the computational logic includes forming an array of multiplier cells, and the FeRAM includes an array of memory bit cells.

[0087] Example 13: The method of Example 12, comprising: forming an interconnect fiber; and coupling the interconnect fiber to the array of multiplier cells such that each multiplier cell is coupled to the interconnect fiber.

[0088] Example 14: The method of Example 11, wherein the FeRAM is partitioned into a first partition operable as a buffer and a second partition for storing weighting factors.

[0089] Example 15: The method of Example 14, comprising: the computation logic receiving data from the first partition and the second partition; and providing an output of the computation logic to the logic circuit.

[0090] Example 16: The method of Example 14, wherein forming the computational logic includes forming ferroelectric logic, the computational logic operable to multiply at least two matrices.

[0091] Example 17: The method of Example 11, wherein forming the substrate includes forming an active or passive device.

[0092] Example 18: The method of Example 11, comprising: forming a third die; bonding the third die onto a substrate; forming a fourth die including dynamic random access memory (DRAM); and stacking the fourth die on the third die.

[0093] Example 19: The method of Example 11, comprising coupling a heat sink to the second die.

[0094] Example 20: A system including: a first memory including non-volatile memory cells; a second memory including a dynamic random access memory (DRAM), wherein the first memory is coupled to the second memory; a third memory including a ferroelectric random access memory (FeRAM), wherein the third memory is coupled to the first memory; a first processor coupled to the second memory; and a second processor coupled to the third memory and the first processor; the second processor including: a substrate; a first die on the substrate, wherein the first die includes ferroelectric random access memory (FeRAM) having bit cells, each bit cell including an access transistor and a capacitor including a ferroelectric material, the access transistor coupled to the ferroelectric material; and a second die stacked on the first die, wherein the second die includes a multiplier coupled to the memory of the first die.

[0095] Example 21: The system of Example 20, wherein the multiplier includes an array of multiplier cells, and the FeRAM includes an array of memory bit cells, each multiplier cell coupled to a corresponding memory bit cell.

[0096] Example 22: The system of Example 21, wherein the second processor includes an interconnecting fiber coupled to the array of multiplier cells, each multiplier cell being coupled to an interconnecting fiber.

[0097] Example 23: An apparatus including: an interposer; a first die on the interposer, the first die including random access memory (RAM) having bit cells; and a second die stacked on the first die, the second die including a matrix multiplier coupled to the memory of the first die.

[0098] Example 24: The apparatus of Example 23, wherein the matrix multiplier includes an array of multiplier cells, the RAM includes an array of memory bit cells, and each multiplier cell is coupled to a corresponding memory bit cell.

[0099] Example 25: The apparatus of Example 23, wherein the second die includes a logic circuit coupled to the matrix multiplier.

[0100] Example 26: The apparatus of Example 25, wherein the second die includes a buffer coupled to the logic circuit, the buffer coupled to the memory.

[0101] Example 27: The device of Example 23, wherein the memory includes a ferroelectric random access memory (FeRAM) having bit cells, each bit cell including an access transistor and a capacitor including a ferroelectric material, and the access transistor is coupled to the ferroelectric material.

[0102] Example 28: The device of Example 23, wherein the memory includes a static random access memory (SRAM) having bit cells.

[0103] Example 29: The apparatus of Example 23, wherein a heat sink is coupled to the second die.

[0104] Example 30: The apparatus of Example 23, wherein the interposer includes a memory coupled to the second die.

[0105] Example 31: An apparatus including: an interposer; a first die on the interposer, the first die including random access memory (RAM) having bit cells including a ferroelectric material; a second die adjacent to the first die and on the interposer, the second die including computational logic electrically coupled to the memory of the first die; and a third die on the interposer, the third die including RAM having bit cells, the third die adjacent to the second die.

[0106] Example 32: The apparatus of Example 31, wherein the interposer includes a RAM electrically coupled to the second die.

[0107] Example 33: The device of example 31, wherein the RAM of the third die comprises a ferroelectric material.

[0108] Example 34: The apparatus of example 31, wherein the computational logic includes a matrix multiplier including an array of multiplier cells.

[0109] Example 35: The apparatus of Example 31, wherein the second die includes a logic circuit coupled to the matrix multiplier.

[0110] Example 36: The apparatus of Example 35, wherein the second die includes a buffer coupled to the logic circuit, and the buffer is coupled to the first or second die.

[0111] Example 37: The device of Example 31, wherein at least one bit cell of the first die includes an access transistor and a capacitor including a ferroelectric material, the access transistor being coupled to the ferroelectric material.

[0112] Example 38: The device of Example 31, wherein the RAM of the third die includes a static random access memory (SRAM) having bit cells.

[0113] Example 39: The apparatus of Example 31, comprising a heat sink coupled to the first, second, and third dies.

[0114] Example 40: A method, the method including: forming an interposer; forming a first die on the interposer, where forming the first die includes forming random access memory (RAM) having bit cells comprising a ferroelectric material; forming a second die on the interposer adjacent to the first die, where forming the second die includes forming computational logic and electrically coupling the memory of the first die to the computational logic; forming a third die on the interposer, where forming the third die includes forming RAM having bit cells; and positioning the third die adjacent to the second die.

[0115] Example 41: The method of Example 40, wherein forming the interposer includes forming a RAM in the interposer, and the method includes electrically coupling the RAM to the second die.

[0116] Example 42: The method of example 40, wherein the RAM of the third die comprises a ferroelectric material.

[0117] Example 43: The method of example 40, wherein forming the computational logic includes forming a matrix multiplier including an array of multiplier cells.

[0118] Example 44: The method of Example 43, wherein forming the second die includes forming a logic circuit, and the method includes coupling the logic circuit to a matrix multiplier.

[0119] Example 45: The method described in Example 44, wherein the step of forming the second die includes the steps of forming a buffer; and coupling the buffer to the logic circuit; and the method includes the step of coupling the buffer to the first or second die.

[0120] Example 46: The method of Example 40, wherein at least one bitcell of the first die includes an access transistor and a capacitor including a ferroelectric material, the access transistor coupled to the ferroelectric material.

[0121] Example 47: The method of Example 40, wherein forming the RAM of the third die includes forming a static random access memory (SRAM) having bit cells.

[0122] Example 48: The method of Example 40, comprising coupling a heat sink to the first, second, and third dies.

[0123] Example 49: A system including: a first memory including non-volatile memory cells; a second memory including dynamic random access memory (DRAM), wherein the first memory is coupled to the second memory; a third memory including ferroelectric random access memory (FeRAM), wherein the third memory is coupled to the first memory; a first processor coupled to the second memory; and a second processor coupled to the third memory and the first processor; the second processor including: an interposer; a first die on the interposer, wherein the first die includes random access memory (RAM) having bit cells including a ferroelectric material; a second die adjacent to the first die and on the interposer, wherein the second die includes computational logic electrically coupled to the memory of the first die; and a third die on the interposer, wherein the third die includes RAM having bit cells, and the third die is adjacent to the second die.

[0124] Example 50: The system of Example 49, wherein the interposer includes a RAM electrically coupled to the second die, and the RAM of the third die includes a ferroelectric material.

[0125] An Abstract is provided to allow the reader to ascertain the nature and gist of the technical disclosure. The Abstract is submitted with the understanding that it will not be used to limit the scope or meaning of the claims. The following claims are incorporated into the Detailed Description, with each claim standing on its own as a separate embodiment.

[0126] Below, the contents of the claims as originally filed are described as examples. [Example 1] A device, the device comprising: A substrate; a first die on the substrate, the first die including a ferroelectric random access memory (FeRAM) having bit cells, each bit cell including an access transistor and a capacitor including a ferroelectric material, the access transistor coupled to the ferroelectric material; a second die stacked on top of the first die, the second die including computational logic coupled to the memory of the first die; device. [Example 2] 2. The apparatus of claim 1, wherein the computational logic includes an array of multiplier cells and the FeRAM includes an array of memory bit cells. [Example 3] 3. The apparatus of example 2, including an interconnecting fiber coupled to the array of multiplier cells, each multiplier cell being coupled to the interconnecting fiber. [Example 4] 2. The apparatus of claim 1, wherein the memory is partitioned into a first partition operable as a buffer and a second partition for storing weighting factors. [Example 5] 5. The apparatus of example 4, wherein the computation logic is capable of receiving data from the first partition and the second partition, and an output of the computation logic is received by a logic circuit. [Example 6] 5. The apparatus of example 4, wherein the computational logic comprises ferroelectric logic. [Example 7] 5. The apparatus of example 4, wherein the calculation logic is operable to multiply at least two matrices. [Example 8] 2. The apparatus of claim 1, wherein the substrate comprises an active or passive device. [Example 9] 9. The device of any one of claims 1 to 8, wherein a third die is bonded to the substrate and a fourth die including a dynamic random access memory (DRAM) is stacked on top of the third die. [Example 10] 10. The apparatus of example 1, wherein a heat sink is coupled to the second die. [Example 11] 1. A method, comprising: forming a substrate; forming a first die on the substrate, the first die including a ferroelectric random access memory (FeRAM) having bit cells, each bit cell including an access transistor and a capacitor including a ferroelectric material, the access transistor coupled to the ferroelectric material; forming a second die stacked on top of the first die, wherein forming the second die includes forming computational logic coupled to the memory of the first die; method. [Example 12] 12. The method of claim 11, wherein forming the computational logic includes forming an array of multiplier cells, and the FeRAM includes an array of memory bit cells. [Example 13] forming an interconnect fiber; and coupling the interconnect fiber to the array of multiplier cells such that each multiplier cell is coupled to the interconnect fiber. [Example 14] 12. The method of example 11, wherein the FeRAM is partitioned into a first partition operable as a buffer and a second partition for storing weighting factors. [Example 15] the computation logic receiving data from the first partition and the second partition; and providing an output of the computational logic to a logic circuit. [Example 16] 15. The method of example 14, wherein forming the computational logic includes forming ferroelectric logic, the computational logic operable to multiply at least two matrices. [Example 17] 12. The method of example 11, wherein forming the substrate comprises forming an active or passive device. [Example 18] forming a third die; bonding the third die onto the substrate; forming a fourth die including a dynamic random access memory (DRAM); stacking the fourth die on top of the third die; 12. The method of claim 11, comprising: coupling a heat sink to the second die. [Example 19] 1. A system comprising: a first memory including nonvolatile memory cells; a second memory including a dynamic random access memory (DRAM), wherein the first memory is coupled to the second memory; and a third memory coupled to the first memory, the third memory comprising a ferroelectric random access memory (FeRAM); a first processor coupled to the second memory; a second processor coupled to the third memory and the first processor; The second processor A substrate; a first die on the substrate, the first die including a ferroelectric random access memory (FeRAM) having bit cells, each bit cell including an access transistor and a capacitor including a ferroelectric material, the access transistor coupled to the ferroelectric material; a second die stacked on top of the first die, the second die including a multiplier coupled to the memory of the first die; system. [Example 20] the multiplier includes an array of multiplier cells, and the FeRAM includes an array of memory bit cells; Each multiplier cell is coupled to a corresponding memory bit cell; 20. The system of claim 19, wherein the second processor includes an interconnecting fiber coupled to the array of multiplier cells, each multiplier cell being coupled to the interconnecting fiber. [Example 21] A device, the device comprising: an interposer; a first die on the interposer, the first die including random access memory (RAM) having bit cells; a second die stacked on top of the first die, the second die including a matrix multiplier coupled to the memory of the first die; device. [Example 22] 22. The apparatus of claim 21, wherein the matrix multiplier includes an array of multiplier cells, the RAM includes an array of memory bit cells, and each multiplier cell is coupled to a corresponding memory bit cell. [Example 23] 22. The apparatus of claim 21, wherein the second die includes a logic circuit coupled to the matrix multiplier. [Example 24] 24. The apparatus of claim 23, wherein the second die includes a buffer coupled to the logic circuit, the buffer coupled to the memory. [Example 25] the memory comprises a ferroelectric random access memory (FeRAM) having bit cells, each bit cell comprising an access transistor and a capacitor comprising a ferroelectric material, the access transistor coupled to the ferroelectric material; or 22. The device of claim 21, wherein the memory comprises a static random access memory (SRAM) having bit cells. [Example 26] 22. The apparatus of claim 21, wherein a heat sink is coupled to the second die, and the interposer includes a memory coupled to the second die. [Example 27] A device, the device comprising: an interposer; a first die on the interposer, the first die including random access memory (RAM) having bit cells including a ferroelectric material; a second die adjacent to the first die and on the interposer, the second die including computational logic electrically coupled to the memory of the first die; a third die on the interposer, the third die including RAM having bit cells, the third die being adjacent to the second die; device. [Example 28] 28. The apparatus of claim 27, wherein the interposer includes a RAM electrically coupled to the second die. [Example 29] The device of Example 27, wherein the RAM of the third die includes a ferroelectric material or the RAM of the third die includes a static random access memory (SRAM) having bit cells. [Example 30] the computation logic includes a matrix multiplier including an array of multiplier cells; the second die includes logic circuitry coupled to the matrix multiplier; 28. The apparatus of example 27, wherein a second die includes a buffer coupled to the logic circuit, the buffer being coupled to the first or second die. [Example 31] 28. The apparatus of claim 27, wherein at least one of the bit cells of the first die includes an access transistor and a capacitor including the ferroelectric material, the access transistor coupled to the ferroelectric material. [Example 32] 33. The apparatus of any one of Examples 27 to 32, comprising a heat sink coupled to the first, second, and third dies. [Example 33] 1. A system comprising: a first memory including nonvolatile memory cells; a second memory including a dynamic random access memory (DRAM), wherein the first memory is coupled to the second memory; and a third memory coupled to the first memory, the third memory comprising a ferroelectric random access memory (FeRAM); a first processor coupled to the second memory; a second processor coupled to the third memory and the first processor; The second processor an interposer; a first die on the interposer, the first die including random access memory (RAM) having bit cells including a ferroelectric material; a second die adjacent to the first die and on the interposer, the second die including computational logic electrically coupled to the memory of the first die; a third die on the interposer, the third die including RAM having bit cells, the third die being adjacent to the second die; system. [Example 34] 34. The system of embodiment 33, wherein the interposer includes RAM electrically coupled to the second die, and the RAM of the third die includes a ferroelectric material.

Claims

1. 1. An IC package assembly, the IC package assembly comprising: a substrate having a top surface; and a stack of dies on the substrate, the die stack comprising: a first die comprising a processing core and a first memory; and a second die vertically stacked with the first die, the first memory comprising a first memory type; and the second die comprising a second memory, the second memory comprising a second memory type, the first die and the second die being connected via inter-die copper pillars that are not solder bumps, the first die comprising through-silicon vias coupled to or part of the inter-die copper pillars, the IC package assembly further comprising: a first silicon structure adjacent to the second die but not vertically stacked with the second die; and a second silicon structure adjacent to the second die but not vertically stacked with the second die. an IC package assembly, wherein the first silicon structure and the second silicon structure are on opposite sides of the second die, the first silicon structure including a high bandwidth memory (HBM) and the second silicon structure including a ferroelectric random access memory (FeRAM), the FeRAM for storing input data and weighting coefficients for the processing core.

2. 2. The IC package assembly of claim 1, wherein the first die has a first surface and the second die has a second surface, the first surface facing the second surface such that the first surface completely overlaps the second surface or the second surface completely overlaps the first surface, the first surface having a first width and the second surface having a second width, and the first width is substantially equal to the second width.

3. 2. The IC package assembly of claim 1, wherein the first die is on top of the second die, or the second die is on top of the first die.

4. 10. The IC package assembly of claim 1, wherein said first memory type and said second memory type comprise ferroelectric materials.

5. 10. The IC package assembly of claim 1, wherein the first die includes ferroelectric logic.

6. 10. The IC package assembly of claim 1, wherein the first die includes ferroelectric logic and ferroelectric memory.

7. 2. The IC package assembly of claim 1, wherein said second memory comprises a ferroelectric material and said first memory comprises an SRAM.

8. 2. The IC package assembly of claim 1, wherein said first memory comprises an SRAM and said second memory comprises an SRAM.

9. 10. The IC package assembly of claim 1, further comprising a heat spreader over the stack of dies.

10. 1. A method of forming an IC package assembly, the method comprising: forming a substrate having a top surface; and forming a stack of dies on the substrate, wherein forming the stack of dies comprises: forming a first die including a processing core and a first memory, the first memory including a first memory type; forming a second die vertically stacked with the first die, the second die including a second memory, the second memory including a second memory type; and connecting the first die and the second die via inter-die copper pillars that are not solder bumps, the first memory of the first die including through-silicon vias coupled to or part of the inter-die copper pillars. The method of forming the IC package assembly comprises: disposing a first silicon structure adjacent to the second die but not vertically stacked with the second die; and disposing a second silicon structure adjacent to the second die but not vertically stacked with the second die, wherein the first silicon structure and the second silicon structure are on opposite sides of the second die, the first silicon structure including a high bandwidth memory (HBM), and the second silicon structure including a ferroelectric random access memory (FeRAM), the FeRAM for storing input data and weighting factors for the processing core.

11. 11. The method of claim 10, wherein the first die has a first surface and the second die has a second surface, the first surface facing the second surface such that the first surface completely overlaps the second surface or the second surface completely overlaps the first surface, the first surface having a first width and the second surface having a second width, and the first width is substantially equal to the second width.

12. The method of claim 10 , wherein the first die is on top of the second die, or the second die is on top of the first die.

13. 11. The method of claim 10, wherein the first memory type and the second memory type include ferroelectric materials.

14. The method of claim 10 , wherein the first die includes ferroelectric logic.

15. The method of claim 10 , wherein the first die includes ferroelectric logic and ferroelectric memory.

16. 11. The method of claim 10, wherein the second memory comprises a ferroelectric material and the first memory comprises an SRAM.

17. The method of claim 10 , wherein the first memory comprises an SRAM and the second memory comprises an SRAM.

Citation Information

Patent Citations

  • Semiconductor integrated circuit device

    JP2006324430A

  • Information processing device, semiconductor chip, information processing method, and program

    JP2015176158A

  • Semiconductor device

    JP2017022352A

  • Artificial intelligence processor with three-dimensional stacked memory

    JP2023164436A

  • Top-side connector interface for processor packaging

    US20170354031A1