Multi-stack computing chip and memory architecture
By stacking DRAM on a computing chip and using interconnects to implement a multi-stacked computing chip and memory architecture, the limitations of computing chips and memory architectures in the existing technology in terms of efficient data access and power consumption are resolved, and high-bandwidth, low-power memory access is achieved.
Patent Information
- Application Number
- CN202380093679.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2023-12-21
- Publication Date
- 2025-09-16
AI Technical Summary
The architecture of existing computing chips and memories has limitations in terms of efficient data access and power consumption, making it difficult to achieve fast, low-power, high-bandwidth memory access.
It adopts a multi-stack computing chip and memory architecture, directly stacking dynamic random access memory (DRAM) on the computing chip, and using interconnects to couple the computing stacks to achieve coherent shared memory across multiple computing stacks, and using silicon interposers, silicon bridges, glass interposers or organic packaging to achieve efficient data transmission.
It provides high-bandwidth, low-power memory access, enables fast data access to memory by computing chips, reduces the physical and topological distance of data communication paths, and improves the efficiency of memory access.
Smart Images

Figure CN120660195A_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims priority to U.S. non-provisional application No. 18 / 390,893, filed on December 20, 2023, and entitled “Multi-Stack Compute Chip and Memory Architecture,” which in turn claims priority under 35 U.S.C. §119(e) to U.S. provisional patent application No. 63 / 484,183, filed on February 9, 2023, and entitled “Multi-Stack Compute Chip and Memory Architecture,” the entire disclosures of which are hereby incorporated by reference. Background Art
[0003] For example, computing technology is constantly advancing, as demonstrated by the expansion of machine learning into many different fields. As computing applications advance, computer hardware is also advancing, such as providing improved processing and memory that is faster, higher performance, consumes less power, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Figure 1 is a block diagram of a non-limiting exemplary system having multiple stacks of computing chips and memory interconnected to form a package, according to some implementations.
[0005] Figure 2 is a block diagram of a non-limiting example of a multi-stack package having stacks arranged in a grid.
[0006] Figure 3 is a block diagram of a non-limiting example of a multi-stack package having stacks arranged in an array.
[0007] Figure 4 is a block diagram of a non-limiting example in which a multi-stack package is expanded through connections to additional multi-stack packages.
[0008] Figure 5 Non-limiting examples of different topologies for connecting stacks of multi-stack packages with interconnects are depicted.
[0009] Figure 6 Depicted are non-limiting examples of arrangements with stacks of computing chips and memory according to some implementations.
[0010] Figure 7 Depicted are procedures in an example implementation of a multi-stacked compute chip and memory architecture.
[0011] Figure 8 Depicted is a process in an example implementation for fabricating a multi-stack package. DETAILED DESCRIPTION
[0012] Overview
[0013] A multi-stacked compute chip and memory architecture is described. According to the described technology, a package includes multiple compute stacks, and in some variations, each compute stack includes at least one compute chip and memory (e.g., at least one memory die). As an example, the memory is a stacked memory, such as a stack of dynamic random access memory (DRAM), which is three-dimensionally (3D) stacked directly above (or below) the compute chip to form a stack. This provides the compute chip with high-bandwidth access to data in the corresponding stacked memory in a power-efficient manner. The package also includes one or more interconnects that couple the compute stack to at least one other compute stack for sharing memory in a coherent manner across multiple compute stacks. Through the interconnects and the multiple stacked memories, the package provides such shared, coherent memory with non-uniform memory access (NUMA) characteristics.
[0014] In some aspects, the technology described herein relates to an apparatus comprising: a plurality of compute stacks, wherein a first compute stack and a second compute stack of the plurality of compute stacks each comprise at least one compute chip and a memory; and one or more interconnects coupling the first compute stack to at least the second compute stack for shared memory.
[0015] In some aspects, the technology described herein relates to an apparatus in which the memory is one or more memory dies.
[0016] In some aspects, the technology described herein relates to an apparatus in which the memory of a first computing stack includes dynamic random access memory (DRAM) and the memory of a second computing stack includes non-volatile memory.
[0017] In some aspects, the technology described herein relates to an apparatus in which the memory of the first compute stack includes dynamic random access memory (DRAM) and non-volatile memory.
[0018] In some aspects, the technology described herein relates to an apparatus wherein a memory of a first computing stack includes a first portion and a second portion, wherein the first portion is embedded in at least one computing chip, and wherein the second portion is separate from and communicatively coupled to the at least one computing chip.
[0019] In some aspects, the technology described herein relates to an apparatus in which at least one computing chip includes at least one central processing unit, graphics processing unit, field programmable gate array, accelerator, or digital signal processor.
[0020] In some aspects, the technology described herein relates to an apparatus wherein a first computing stack further includes one or more memory request units adapted to perform at least one of: sending and receiving data over one or more interconnects between memories of the first computing stack and a second computing stack; providing coherent shared memory across memories of the plurality of computing stacks; or providing coherent shared memory across memories of a subset of the plurality of computing stacks.
[0021] In some aspects, the technology described herein relates to a device in which at least one compute chip of a first compute stack is configured to access data from a memory of the first compute stack faster than data from a memory of a second compute stack, wherein the at least one compute chip of the first compute stack accesses data from the memory of the second compute stack through one or more interconnects.
[0022] In some aspects, the technology described herein relates to a device in which one or more interconnects are disposed on or within at least one of a silicon interposer, a silicon bridge, a glass interposer, an organic package, or a silicon photonic interconnect.
[0023] In some aspects, the technology described herein relates to an apparatus in which a plurality of compute stacks are interconnected with one or more interconnects in an array topology.
[0024] In some aspects, the technology described herein relates to an apparatus in which a plurality of compute stacks are interconnected with one or more interconnects in a mesh topology.
[0025] In some aspects, the technology described herein relates to a device, wherein the device is a multi-stack package that is communicatively coupled to at least one additional multi-stack package.
[0026] In some aspects, the technology described herein relates to a device in which memory of a first computing stack is disposed in a stacked arrangement above or below at least one computing chip of the first computing stack, and the stacked arrangement is disposed on a substrate of the device.
[0027] In some aspects, the technology described herein relates to an apparatus in which a memory of a first computing stack is disposed in a side-by-side arrangement with at least one computing chip of the first computing stack.
[0028] In some aspects, the technology described herein relates to a device in which the side-by-side arrangement is provided on a substrate of the device.
[0029] In some aspects, the technology described herein relates to a device in which a side-by-side arrangement is disposed in a stacked arrangement above or below a circuit die of a first computing stack, and the stacked arrangement is disposed on a substrate of the device.
[0030] In some aspects, the technology described herein relates to an apparatus wherein a circuit die includes at least one of: a memory controller, a cache, a data fabric, a network on a chip (NoC), or a memory interface circuit.
[0031] In some aspects, the technology described herein relates to a system-level package (SoP) comprising: a plurality of computing stacks, wherein: each of the plurality of computing stacks comprises at least one computing chip and a memory; and at least a first computing stack of the plurality of computing stacks is coupled to at least a second computing stack and a third computing stack of the plurality of computing stacks via an interconnect for sharing memory; and one or more interfaces for coupling the system-level package to at least one external device.
[0032] In some aspects, the technology described herein relates to a system-in-package (SoP) in which at least one external device includes at least one of: an integrated circuit, external memory, a motherboard, or an additional system-in-package.
[0033] In some aspects, the technology described herein relates to a method for manufacturing a multi-stack package, the method comprising: forming a plurality of computing stacks, wherein a first computing stack and a second computing stack in the plurality of computing stacks each include at least one computing chip and a memory; and disposing the plurality of computing stacks on a substrate, the first computing stack and the second computing stack being electrically connected on the substrate via one or more interconnects for sharing memory.
[0034] Figure 1 is a block diagram of a non-limiting exemplary system 100 having multiple stacks of computing chips and memory interconnected to form a package, according to some implementations. Specifically, the illustrated example depicts two views of the system 100, including a top-down "bird's eye" view and a side cutaway view. It should be understood that the components are not drawn to scale in either view, and that the sizes and positioning of the components relative to each other may vary in implementations. Additionally, the number of components of the system 100 may vary in variations without departing from the spirit or scope of the described technology.
[0035] According to the described techniques, the system 100 is or includes a multi-stack package 102 having a plurality of stacks 104 (e.g., at least a first stack 104 and a second stack 104). The illustration includes ellipses to indicate that in one or more implementations, the multi-stack package 102 includes more than two stacks 104, some examples of which are shown in FIG. Figures 2 to 5 Depicted in.
[0036] Multiple stacks 104 include computing chips 106 and memory 108. For example, each stack 104 includes at least a computing chip 106 and memory 108, such as a memory die. In variations, the stack 104 includes more than one computing chip 106 (e.g., two or more computing chips) and / or more memory than a single memory die (e.g., multiple memory dies and / or at least one memory die and additional memory embedded in the computing chip 106). As an example, the stack 104 includes one or more computing chips 106 and stacked memory 108. In one or more implementations, the memory 108 is stacked directly on top of the one or more computing chips 106. In one or more implementations, the memory 108 is stacked directly below the one or more computing chips 106. In one or more implementations, the computing chip 106 is disposed between a first portion of the memory 108 and a second portion of the memory 108, such that one of those portions is stacked directly on top of the one or more computing chips 106 and the other portion is stacked directly below the one or more computing chips 106. Alternatively or in addition, the memory 108 and the computing chips 106 are interleaved. It should be understood that in variations, one or more computing chips 106 and memory 108 (e.g., one or more memory dies) are arranged in a different manner to form the stack 104 without departing from the described technology.
[0037] The computing chips 106 and memories 108 of a particular stack 104 are coupled to each other using any one or more of a variety of wired or wireless connection types. Exemplary wired connections include, but are not limited to, one or more memory channels, buses (e.g., data buses), interconnects, through-silicon vias, data links (e.g., 1024 data links), traces, photonic interconnects, and planes, to name a few. The stacked arrangement of computing chips 106 and memories 108 provides computing chips 106 with high-bandwidth access to data in memories 108 within a corresponding stack 104, and provides that access with reduced power consumption. This is due to shorter data communication paths relative to configurations where the memories 108 are physically and / or topologically farther away from the computing chips 106, as well as shorter data communication paths relative to memories 108 in another stack 104.
[0038] The illustrated example also depicts an interconnect 110. The interconnect 110 couples (e.g., communicatively couples) the stacks 104 of the multi-stack package 102. According to the described techniques, a stack 104 is connected to at least one other stack 104 via one or more interconnects 110. Via those interconnects 110, the system 100 transfers data between the stacks 104. The interconnects 110 also enable the other stacks 104 to access data maintained in the memory 108 of the stacks 104. In at least one scenario, for example, one or more interconnects 110 enable the computing chips 106 of a first stack 104 to access data loaded into the memory 108 of at least a second stack 104. Based on this, the multi-stack package 102 is configured to use the memory 108 of the multiple stacks 104 to provide coherent shared memory. For example, the multi-stack package 102 uses the memory 108 of a subset of the stacks (e.g., less than all of the multiple stacks) to provide such coherent shared memory. Alternatively, the multi-stack package 102 uses at least a portion of the memory 108 of all stacks 104 to provide such coherent shared memory.
[0039] In one or more specific implementations, the compute chip 106 is an electronic circuit that performs various operations on and / or uses data in a memory 108 (such as data in the memory 108 of the stack 104 of compute chips 106 or data in the memory 108 of at least one other stack 104). By way of example and not limitation, such operations are associated with programs, applications, and / or threads (not shown). In accordance with the described techniques, the compute chip 106 of a given stack is any one or more of a variety of processing units, such as a central processing unit (CPU), a graphics processing unit (GPU), an accelerator, an accelerated processing unit (APU), a parallel accelerated processor, a digital signal processor, an artificial intelligence (AI) or machine learning accelerator, a field programmable gate array (FPGA), etc. In variations, without departing from the spirit or scope of the described techniques, the compute chip 106 corresponds to one or more different types of components, such as a cache.
[0040] Although a single computing chip 106 is illustrated in each stack 104 of the multi-stack package 102, the stacks 104 may optionally include any number of computing chips 106 of the same or different types. In one or more implementations, each stack 104 of the multi-stack package 102 includes one or more computing chips 106 of the same type, e.g., each stack includes the same processing unit (subject to manufacturing variations) and / or a combination of the same processing units (subject to manufacturing variations). However, in other implementations, at least one stack 104 has one or more computing chips 106 that are different from at least one other stack 104, e.g., the computing chips 106 of a first stack 104 are CPUs, and the computing chips 106 of a second stack 104 are different types of CPUs or GPUs.
[0041] The memory 108 is a device or system for storing information, such as for immediate use in a device, for example, by the computing chip 106 of the corresponding stack 104 or by the computing chip 106 of at least one other stack 104. In one or more specific implementations, the memory 108 corresponds to a semiconductor memory, in which data is stored in memory cells on one or more integrated circuits. In at least one example, the memory 108 corresponds to or includes a volatile memory, examples of which include random access memory (RAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), phase change memory (PCM), memristor, static random access memory (SRAM), etc. In variations, the memory 108 is packaged or configured in any of a variety of different ways.
[0042] Another example of a memory configuration includes low-power double data rate (LPDDR), also known as LPDDR SDRAM, which is a type of synchronous dynamic random access memory. In variations, LPDDR consumes less power than other types of memory and / or has a form factor suitable for mobile computers and devices (such as mobile phones). Examples of LPDDR include, but are not limited to, low-power double data rate 2 (LPDDR2), low-power double data rate 3 (LPDDR3), low-power double data rate 4 (LPDDR4), and low-power double data rate 5 (LPDDR5).
[0043] In at least one variation, the memory 108 is a stacked memory, an example of which is stacked DRAM. Alternatively or additionally, the memory 108 corresponds to or includes a non-volatile memory, examples of which include ferroelectric RAM, magnetoresistive RAM, flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM). It should be understood that the memory 108 can be configured in a variety of ways without departing from the spirit or scope of the described technology.
[0044] In one or more variations, at least one stack 104 of the multi-stack package 102 includes a cache (or more than one cache) in addition to or in lieu of the memory 108. Furthermore, while discussed throughout as having one or more compute chips 106 and memory 108 (e.g., one or more memory dies), in variations, the stacks 104 of the multi-stack package 102 include additional and / or different components stacked relative to each other according to the described techniques.
[0045] Although not shown, in one or more implementations, the memory controller manages access of the compute chip 106 to the memory 108 , such as by sending read and write requests to the memory 108 and receiving responses from the memory 108 .
[0046] As described above, the multi-stack package 102 includes one or more interconnects 110 connecting the stacks 104. For example, the interconnect 110 connects at least two stacks 104 and is configured to route data between the at least two stacks 104. In other words, the interconnect 110 enables data transmission and / or exchange between the at least two stacks. Broadly speaking, the interconnect 110 is a component, system and / or device through which data can be transmitted between endpoints (e.g., between at least two stacks 104). In variations, the multi-stack package 102 includes multiple interconnects 110. In one or more variations, the interconnect 110 is implemented on or at least partially within the following components: a silicon interposer (e.g., a passive silicon interposer or an active silicon interposer), a glass interposer, a silicon bridge, an organic package, or a photonic interconnect (e.g., a silicon photonic interconnect), to name a few. Alternatively or additionally, the interconnect 110 is implemented as a bus (e.g., a data bus), a data link (e.g., a 1024 data link), a trace, and / or a plane. In variations, the interconnect 110 is configured in different ways without departing from the spirit or scope of the described technology.
[0047] According to the described techniques, the interconnect 110 connects the stacks 104 of the multi-stack package 102 in any of a variety of topologies, such as a two-dimensional (i.e., 2D) grid, to facilitate data communication between the stacks 104. Using the interconnect 110, the system 100 is able to implement coherent shared memory across the individual memories 108 of the multiple stacks 104 of the multi-stack package 102. Figures 2 to 5 Exemplary topologies with different dimensions are discussed in more detail.
[0048] The illustrated example also depicts a memory request unit 112. In one or more implementations, the memory request unit 112 is a logic block configured to manage memory 108 (such as managing memory 108 of a single stack 104, memory 108 of more than one stack 104), and / or coordinate with at least one additional memory request unit 112 (e.g., of another stack 104). In one or more implementations, one or more of the stacks 104 include multiple memory request units, where the individual memory request units perform different operations to implement the techniques discussed above and below. In one variation, for example, at least one of the stacks 104 includes a first memory request unit that sends and / or receives data over the interconnect 110, and also includes a second memory request unit adapted to provide a coherent shared memory.
[0049] In the illustrated example, each of the stacks 104 is depicted as having a corresponding memory request unit 112. However, in one or more variations, each stack 104 includes more than one memory request unit as discussed immediately above, for example, to perform a specialized subset of memory-based operations. In at least one variation, the memory request unit 112 manages memory 108 for more than one stack 104, such that at least one of the stacks 104 does not include a memory request unit 112. For example, one memory request unit 112 manages memory 108 for a subset of the plurality of stacks 104 (but at least two stacks 104). In another example, one memory request unit 112 manages memory 108 for all stacks 104 of the multi-stack package 102. It should be understood that the number of memory request units 112 implemented for the multi-stack package 102 varies among the variations.
[0050] In addition to implementing different numbers of memory request units 112 for the multi-stack package 102, in variations, the memory request units 112 are implemented in different components of the multi-stack package 102. In at least one variation, for example, as in the illustrated example, the memory request units 112 are implemented in the computing chip 106. However, in variations, the memory request units 112 may be implemented in other components. For example, the memory request units 112 are configured as dedicated circuitry integral to (e.g., soldered to) each stack (e.g., but separate from the computing chip 106), which performs the various operations discussed above and below. Alternatively or in addition, the memory request units 112 are configured as microcontrollers, such as those provided on a die integral to each stack or provided on a die of the multi-stack package 102, that run firmware to perform the various operations discussed above and below. In one or more specific implementations, the memory request units 112 are shared among the multiple stacks 104. Memory request unit 112 may be implemented in one or more of various components in accordance with the described techniques.
[0051] In one or more implementations, one or more memory request units 112 operate the memory 108 of multiple stacks 104 (or the memory 108 of at least a subset of the stacks 104) of the multi-stack package 102 as shared coherent memory. Those memory request units 112 do so, for example, based on one or more memory management techniques. For example, one or more of the memory request units 112 expose the memory 108 across the multiple stacks 104 as coherent shared memory with non-uniform memory access (NUMA) characteristics. Such memory has NUMA characteristics because, from the location of a given compute chip 106, data in the memory 108 of the same stack 104 as a given compute chip 106 can be accessed more quickly than data in the memory 108 of a different stack (e.g., across at least one interconnect 110). This is due, at least in part, to differences in physical and topological distances and the number of interfaces and components across which data in another stack 104 (e.g., a remote stack) may be accessed.
[0052] Additionally or alternatively, one or more of the memory request units in the memory request unit 112 are configured to cause high-speed data movement between one or more of the stacks 104 (e.g., a subset of the stacks). In at least one variation, one or more of the memory request units in the memory request unit 112 implement separate data movers. In at least one variation, these data movers may be configured to implement a message passing interface (MPI) having separate levels for one or more of the stacks 104 (e.g., a subset of the stacks). In one or more implementations, the data movers utilize one or more formats and / or standards other than MPI.
[0053] The multi-stack package 102 optionally includes one or more additional controllers to link to additional devices, such as a Peripheral Component Interconnect Express (PCIe) controller, a Serial Advanced Technology Attachment (SATA) controller, a Universal Serial Bus (USB) controller, a Serial Peripheral Interface (SPI) controller, a Low Pin Count (LPC) controller, HyperTransport (HT), a Compute Express Link (CXL), etc. In variations, the multi-stack package 102 includes one or more additional components not depicted in this example, such as an interface (e.g., for connecting to components and / or systems external to the multi-stack package 102), a controller, a system manager, and an optical component, to name a few. As an example, the multi-stack package 102 is configured to connect to and communicate with at least one other multi-stack package 102 or a different component or system using one or more such additional components (e.g., one or more silicon photonic interconnects and / or a network interface card (NIC)).
[0054] The system 100 is configured for incorporation into a device or apparatus. By way of example and not limitation, examples of different types of devices or apparatuses that may be incorporated into the system 100 include servers, personal computers (e.g., desktop or tower computers), smartphones or other wireless phones, tablets or tablet computers, notebook computers, laptop computers, wearable devices (e.g., smart watches, augmented reality headsets or devices, virtual reality headsets or devices), entertainment devices (e.g., game consoles, portable gaming devices, streaming media players, digital video recorders, music or other audio playback devices, televisions, set-top boxes), Internet of Things (IoT) devices, automotive computers, computers for other types of vehicles (e.g., scooters, electric bicycles, motorcycles), systems on chips (SoCs), system-in-package (SoPs) (sometimes referred to as system-in-package (SiPs)), and the like.
[0055] Figure 2 is a block diagram of a non-limiting example 200 of a multi-stack package having stacks arranged in a grid.
[0056] In the illustrated example 200, the multi-stack package 102 includes a plurality of stacks 104, which can be configured in various ways as discussed above. The stacks 104 are connected to the other stacks 104 of the multi-stack package 102 in a grid (e.g., a two-dimensional (2D) grid). As discussed above, the interconnect 110 connects the stacks 104 of the multi-stack package 102. As also described above, the interconnect 110 enables the memory 108 of a single stack 104 to be shared with one or more additional stacks 104. In this manner, the computing chip 106 of a single stack 104 can access data in the memory 108 of one or more additional stacks 104 of the multi-stack package 102 as well as data in the memory 108 of the same stack 104 via the interconnect 110.
[0057] In this example, the multi-stack package 102 is depicted as including interface components 202. These interface components 202 support interaction with different devices and / or systems, such as interaction with additional multi-stack packages 102, other integrated circuits, external memory, a motherboard, and the like. Examples of interface components 202 include, but are not limited to, a network interface card (NIC), a photonic interconnect component, one or more sockets, a peripheral component interconnect express (PCIe) component, a serial advanced technology attachment (SATA) component, a universal serial bus (USB) component, a serial peripheral interface (SPI) component, a low pin count (LPC) component, a compute express link (CXL) component, and the like. In one or more implementations, the interconnect 110 also connects the stack 104 to the interface components 202. However, in at least one variation, the stack 104 is connected to the interface components 202 (or other components) using a coupling different from the coupling used to connect the stack 104 (connecting one stack to another).
[0058] In one or more implementations, the system or device includes additional and / or different types of memory external to the multi-stack package 102. Alternatively or in addition, the system or device includes additional and / or different types of memory within one or more of the multi-stack packages 102, wherein the additional and / or different types of memory are separate from the memory 108 in the stack 104. For example, the one or more additional memories are integrated within the package and are not part of the stack 104. Instead, those one or more additional memories are positioned “outside” of the interface assembly 202 relative to the multi-stack package 102, e.g., on a side of at least one interface assembly 202 opposite the multi-stack package 102. With respect to example 200, for example, from the depicted top-down perspective, one or more of those additional memories may be positioned to the left of the leftmost interface assembly 202, above the topmost interface assembly 202, to the right of the rightmost interface assembly 202, and / or below the bottommost interface assembly 202. In the context of different arrangements of the stacks 104 of the multi-stack packages 102, consider Figure 3 .
[0059] Figure 3 is a block diagram of a non-limiting example 300 of a multi-stack package having stacks arranged in an array.
[0060] In the illustrated example 300, the multi-stack package 102 includes a plurality of stacks 104, which can be configured in various ways as discussed above. The stacks 104 are connected to the other stacks 104 of the multi-stack package 102 in an array (e.g., a one-dimensional (1D) array). As discussed throughout, an interconnect 110 connects the stacks 104 of the multi-stack package 102. As also described above, the interconnect 110 enables the memory 108 of a single stack 104 to be shared with one or more additional stacks 104. In this way, the computing chip 106 of a single stack 104 can access data in the memory 108 of one or more additional stacks 104 of the multi-stack package 102 as well as data in the memory 108 of the same stack 104 via the interconnect 110.
[0061] Figures 2 to 3 The examples depicted in FIG. 1 are merely illustrative, and the layout of the stacks 104 of the multi-stack packages 102 may vary in variations without departing from the spirit or scope of the described technology.
[0062] Figure 4 is a block diagram of a non-limiting example 400 in which a multi-stack package is expanded through connections to additional multi-stack packages.
[0063] The illustrated example 400 includes a plurality of multi-stack packages 102. The multi-stack packages 102 are connected via a communicative coupling 402. This illustrates a scenario in which a multi-stack package 102 is expanded by connecting it to one or more additional multi-stack packages 102. In one or more implementations, this enables shared coherent memory across multiple multi-stack packages 102 in addition to shared coherent memory across multiple stacks of a single package.
[0064] Figure 5 Depicted are non-limiting examples 500 of different topologies for connecting stacks of multi-stack packages with interconnects.
[0065] Specifically, the illustrated example 500 includes a first topology 502, a second topology 504, a third topology 506, and a fourth topology 508. Each of the topologies depicts a plurality of stacks 104 and interconnects 110 connecting the stacks. The first topology 502 is an array of stacks 104 and is an example of a one-dimensional (i.e., 1D) topology. The second topology 504 is a grid of stacks 104 (e.g., four stacks 104) and is an example of a two-dimensional (i.e., 2D) topology. The third topology 506 is another grid of stacks 104 (e.g., eight stacks 104) and is an example of a three-dimensional (i.e., 3D) topology. The fourth topology 508 is another grid of stacks 104 (e.g., sixteen stacks 104) and is an example of a four-dimensional (i.e., 4D) topology. In one or more specific implementations, the dimensionality of the topology is based on the number of other computing stacks to which a single computing stack is connected via the interconnect 110. In at least one variation, for example, the dimensionality is based on the number of compute stacks to which the compute stack that is connected to the least number of other compute stacks in the topology is connected via interconnects 110. As an example, each compute stack 104 of the third topology 506 is connected to three other compute stacks, and thus, the third topology 506 is a 3D topology. In contrast, each compute stack 104 of the fourth topology 508 is connected to four other compute stacks, and thus, the fourth topology 508 is a 4D topology. In other words, the dimensionality of the topology is related to the number of interconnects between a given stack and the other stacks (e.g., how many other stacks a given stack is connected to). For example, a stack 104 having interconnects 110 to N different stacks corresponds to an ND topology. As another specific example (not depicted), if a given stack 104 is connected to ten (10) other stacks 104 via interconnects, then the topology corresponds to a 10D topology.
[0066] It should be understood that, in variations, the interconnect 110 and stack 104 may be arranged in a variety of different topologies of higher or lower dimensions without departing from the spirit or scope of the described technology.
[0067] Figure 6Depicted is a non-limiting example 600 of an arrangement having a stack of computing chips and memory according to some implementations.
[0068] The illustrated example 600 depicts an example of a stack 602 having an arrangement of hardware components for use in at least one variation according to the described techniques. As an example, the components of one or more of the first stack 104 or the second stack 104 are arranged similar to the stack 602 in example 600. It should be understood that the components of the stacks of the multi-stack package can be arranged in a different manner without departing from the spirit or scope of the described techniques.
[0069] In the illustrated example 600, stack 602 includes computing chip 106, memory 108, and circuit die 604. In one or more implementations, stack 602 or one or more components of stack 602 (e.g., circuit die 604) are coupled to or otherwise integrated with a substrate 606, an example of which is a system-on-chip (SOC) substrate. By way of example and not limitation, substrate 606 corresponds to multi-stack package 102.
[0070] It is worth noting that the arrangement in the illustrated example 600 is different from Figure 1 106. In system 100, computing chip 106 and memory 108 are arranged in a vertically stacked arrangement, e.g., memory 108 is disposed “on top” of computing chip 106. In contrast, in example 600, computing chip 106 and memory 108 have a side-by-side arrangement, such that memory 108 is disposed next to computing chip 106, rather than vertically on top of or below computing chip 106. It should be understood that in one or more implementations, different stacks of a multi-stack package include stacks having different component arrangements, such as stacks having components disposed “on top” of computing chip 106. Figure 1 and a second stack having an arrangement as depicted in example 600. However, in at least one variation, all stacks have the same component arrangement, e.g., all stacks of a multi-stack package have an arrangement similar to that depicted in the illustrated example 600.
[0071] In the illustrated example 600, the circuit die includes a memory request unit 112. In one or more implementations, the memory request unit 112 is implemented using different and / or additional components than the circuit die 604. The circuit die 604 can be configured as and / or configured with any of a variety of semiconductor components, examples of which include, but are not limited to, a memory controller, a cache, a data fabric, a network on chip (NoC), and a memory interface circuit, to name a few. In one or more implementations, the stack 602 includes multiple circuit dies.
[0072] Although a single computing chip 106 and memory 108 are depicted in this example 600, in one or more variations, multiple layers of computing chips 106 and memory 108 are stacked on the circuit die 604 such that at least a second layer having computing chips 106 and memory 108 is stacked on the first layer to form a "higher" stack.
[0073] Figure 7 Depicted are processes in an example 700 implementation of a multi-stacked compute chip and memory architecture.
[0074] Data in a first memory is accessed for use by a computing chip (block 702). According to the principles discussed herein, a first memory and a computing chip are coupled together to form a first stack of a multi-stack package. As an example, data in a memory 108 of a first stack 104 of a multi-stack package 102 is accessed by a memory controller (not shown) of the first stack 104 for use by a computing chip 106 of the first stack 104. For example, the data is accessed for use by the computing chip 106 to execute instructions of an application (e.g., an operating system, a machine learning-based task, etc.).
[0075] Data in the second memory is accessed for use by the computing chip (block 704). According to the principles discussed herein, the second memory is disposed in a second stack of the multi-stack package, and the second stack is communicatively coupled to the first stack via one or more interconnects of the multi-stack package. As an example, data in the memory 108 of the second stack 104 of the multi-stack package 102 is accessed by a memory controller (not shown) of the second stack 104 for use by the computing chip 106 of the first stack 104. In at least one embodiment, the data is communicated from the second stack 104 to the first stack 104 across the interconnect 110. The computing chip 106 of the first stack 104 then uses (e.g., processes) the data obtained from the memory 108 of the second stack 104. For example, the computing chip 106 of the first stack 104 executes instructions (e.g., of an application) associated with the data obtained from the memory 108 of the first stack 104.
[0076] Figure 8 Depicted is a process in an example 800 implementation of manufacturing a multi-stack package.
[0077] A plurality of computing stacks are formed (block 802). In accordance with the principles herein, a first computing stack and a second computing stack in the plurality of computing stacks include at least one computing chip and a memory. As an example, a plurality of stacks 104 are formed, and the plurality of stacks include at least a first stack 104 and a second stack 104. In this example, the first stack 104 and the second stack 104 are each formed to include a computing chip 106 and a memory 108. In one or more implementations, one or more of the stacks are formed by disposing the memory 108 on top of or below the computing chip 106 in a stacked arrangement, an example of which is shown in FIG. Figure 1 Alternatively or additionally, one or more stacks may be formed by arranging the memory 108 side by side with the computing chip 106, an example of which is shown in FIG. Figure 6 The particular stack of compute chips 106 and memory 108 may be electrically connectable to enable communication between those components in various ways without departing from the spirit or scope of the described technology, such as utilizing through-silicon vias or any of the various connections discussed above.
[0078] A plurality of computing stacks are disposed on a substrate (block 804). According to the principles herein, the computing stacks are disposed on a substrate such that a first computing stack and a second computing stack are electrically connected on the substrate via one or more interconnects for sharing memory. As an example, a plurality of computing stacks including a first computing stack 104 and a second computing stack 104 are disposed on a substrate such as substrate 606. In variations, the plurality of computing stacks are disposed in different topologies, examples of which include but are not limited to Figure 5 On substrate 606, first computing stack 104 and second computing stack 104 are electrically connected, such as via one or more interconnects in interconnects 110. Across interconnects 110, first computing stack 104 and second computing stack 104 communicate to share respective memories 108, such as for use by another stack of computing dies 106 having NUMA characteristics, for example.
[0079] It should be understood that many variations are possible based on the disclosure herein. Although the above features and elements are described in particular combinations, each feature or element can be used alone without the other features and elements, or in various combinations with or without the other features or elements.
[0080] The various functional units illustrated in the figures and / or described herein (including, where appropriate, the multi-stack package 102, the plurality of stacks 104, the computing chip 106, the memory 108, the interconnect 110, and the memory request unit 112) are implemented in any of a variety of different ways, such as hardware circuits, software or firmware executed on a programmable processor, or any combination of two or more of hardware, software, and firmware. The provided methods are implemented in any of a variety of devices such as a general-purpose computer, a processor, or a processor core. For example, suitable processors include general-purpose processors, special-purpose processors, conventional processors, digital signal processors (DSPs), graphics processing units (GPUs), parallel acceleration processors, multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and / or a state machine.
[0081] In one or more specific implementations, the methods or processes provided herein are implemented in a computer program, software, or firmware incorporated into a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM disks), and digital versatile disks (DVDs).
Claims
1. A device, comprising: a plurality of computing stacks, wherein a first computing stack and a second computing stack of the plurality of computing stacks each include at least one computing chip and a memory; and One or more interconnects coupling the first compute stack to at least the second compute stack for sharing the memory.
2. The apparatus of claim 1, wherein the memory is one or more memory dies.
3. The apparatus of claim 1 , wherein the memory of the first computing stack comprises dynamic random access memory (DRAM) and the memory of the second computing stack comprises non-volatile memory.
4. The apparatus of claim 1 , wherein the memory of the first computing stack comprises dynamic random access memory (DRAM) and non-volatile memory.
5. The apparatus of claim 1 , wherein the memory of the first computing stack comprises a first portion and a second portion, wherein the first portion is embedded in the at least one computing chip, and wherein the second portion is separate from and communicatively coupled to the at least one computing chip.
6. The apparatus of claim 1, wherein the at least one computing chip comprises at least one central processing unit, a graphics processing unit, a field programmable gate array, an accelerator, or a digital signal processor.
7. The apparatus of claim 1 , wherein the first computational stack further comprises one or more memory request units adapted to perform at least one of: sending and receiving data via the one or more interconnects between the memories of the first and second compute stacks; providing coherent shared memory across the memories of the plurality of compute stacks; or Coherent shared memory is provided across the memories of the subset of the plurality of compute stacks.
8. An apparatus according to claim 1, wherein the at least one computing chip of the first computing stack is configured to access data from the memory of the first computing stack faster than data from the memory of the second computing stack, wherein the at least one computing chip of the first computing stack accesses the data from the memory of the second computing stack through the one or more interconnects.
9. The apparatus of claim 1, wherein the one or more interconnects are disposed on or within at least one of a silicon interposer, a silicon bridge, a glass interposer, an organic package, or a silicon photonic interconnect.
10. The apparatus of claim 1, wherein the plurality of compute stacks and the one or more interconnects are interconnected in an array topology.
11. The apparatus of claim 1 , wherein the plurality of compute stacks and the one or more interconnects are interconnected in a mesh topology. 12 . The device of claim 1 , wherein the device is a multi-stack package communicatively coupleable to at least one additional multi-stack package.
13. The device of claim 1, wherein the memory of the first computing stack is disposed in a stacked arrangement above or below the at least one computing chip of the first computing stack, and the stacked arrangement is disposed on a substrate of the device.
14. The apparatus of claim 1, wherein the memory of the first computing stack and the at least one computing chip of the first computing stack are disposed in a side-by-side arrangement.
15. The device of claim 14, wherein the side-by-side arrangement is provided on a substrate of the device.
16. The apparatus of claim 14, wherein the side-by-side arrangement is disposed in a stacked arrangement above or below a circuit die of the first computing stack, and the stacked arrangement is disposed on a substrate of the apparatus.
17. The apparatus of claim 16, wherein the circuit die comprises at least one of: Memory controller; Cache; Data organization; Network on chip (NoC); or Memory interface circuit.
18. A system-in-package (SoP), comprising: Multiple computing stacks, where: Each computing stack of the plurality of computing stacks includes at least one computing chip and a memory; and At least a first computing stack of the plurality of computing stacks is coupled to at least a second computing stack and a third computing stack of the plurality of computing stacks via an interconnect for sharing the memory; and One or more interfaces for coupling the system-in-package to at least one external device.
19. The system-in-package (SoP) of claim 18, wherein the at least one external device comprises at least one of: integrated circuit; External storage; motherboard; or Additional system-in-package.
20. A method for manufacturing a multi-stack package, the method comprising: forming a plurality of computing stacks, wherein a first computing stack and a second computing stack of the plurality of computing stacks each include at least one computing chip and a memory; as well as The plurality of computing stacks are disposed on a substrate, and the first computing stack and the second computing stack are electrically connected on the substrate via one or more interconnects for sharing the memory.
Citation Information
Cited By
Multi-stack compute chip and memory architecture
US12688131B2