Application of 3D-DRAM Chiplets to Large Language Models

The computing package with low-power dies and 3D memory stacks addresses the inefficiency of current packages by providing high-bandwidth processing for large-language models, optimizing power usage and thermal management.

JP2025523323APending Publication Date: 2025-07-23GOOGLE LLC
View PDF 22 Cites 0 Cited by

Patent Information

Application Number
JP2024537997
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-06-16
Filing Date
2023-10-24
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Current high-performance computing packages are not optimal for serving as machine-learning accelerators for large-language models due to their high computational intensity and high-power design, which limits their efficiency and suitability for low-computational-intensity applications.

Method used

A computing package design featuring multiple low-power computing dies and 3D memory stacks, organized into compute-memory chiplets, which are connected via input/output dies for efficient data transmission and operation within thermal and spatial constraints.

Benefits of technology

The design enables high-bandwidth processing suitable for large-language models by reducing power consumption and accommodating more chiplets per package, enhancing bandwidth without increasing operational costs, while maintaining reliability through spare dies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523323000001_ABST
    Figure 2025523323000001_ABST
Patent Text Reader

Abstract

The systems and methods disclosed herein provide high-bandwidth processing using multiple compute-memory chiplets. A computing package may be configured to have multiple compute-memory chiplets to perform processing operations related to large language models. The compute-memory chiplets may be configured to operate using small, low-power computing dies that are operable efficiently for low-intensity workloads.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application is a continuation of U.S. Patent Application No. 18 / 210,846, filed on Jun. 16, 2023, which claims the benefit of the filing date of U.S. Provisional Patent Application No. 63 / 505,725, filed on Jun. 2, 2023, the disclosures of which are incorporated herein by reference.

Background Art

[0002] Background High - performance computing can be performed using computing packages having multiple high - bandwidth memory (“HBM”) dies. However, these packages are configured to operate for applications with high computational intensity through the use of high - power computing dies. However, current packages are not optimal for serving as machine - learning accelerators for large - language models due to their high computational intensity and high - power design.

Summary of the Invention

Problems to be Solved by the Invention

[0003] Brief Summary What is needed is a high - bandwidth memory package design that is configured to operate efficiently as a machine - learning accelerator for applications with low computational intensity, such as for large - language models (LLMs) or other models limited by memory bandwidth. Aspects of the present disclosure are directed to a machine - learning accelerator having a 3D memory die implemented using multiple chiplets to improve bandwidth. The machine - learning accelerator is further designed for low computational intensity to reduce the power required for the operation of the package. Additionally, aspects of the present disclosure enable the machine - learning accelerator to operate on a package that is limited with respect to thermal constraints and the amount of space available for computing and memory components.

Means for Solving the Problem

[0004] According to an aspect of the present disclosure, a computing package can include a package substrate and one or more computing clusters disposed on the package substrate. Each of the one or more computing clusters includes a plurality of arithmetic-memory stacks that communicate with an input / output die. Each arithmetic-memory stack can include a plurality of memory dies stacked with a low-power arithmetic die. The input / output die can be configured to transmit data regarding the plurality of arithmetic-memory stacks via one or more peripheral component interconnects.

[0005] In another aspect of the present disclosure, each low-power arithmetic die of the plurality of arithmetic-memory stacks can be configured to operate with a power supply of about 40 W or less.

[0006] In yet another aspect, for a particular arithmetic-memory stack among the plurality of arithmetic-memory stacks, the low-power arithmetic die has a footprint on the package substrate that is less than 30% larger than the footprint of the plurality of memory dies.

[0007] In still another aspect of the present disclosure, the computing package can further include a plurality of computing clusters disposed on the package substrate. The input / output die for each computing cluster is connected to one or more input / output dies of other computing clusters disposed on the package substrate.

[0008] In yet another aspect of the present disclosure, the package substrate can include four computing clusters. Each computing cluster includes four or more arithmetic-memory stacks. Further, each computing cluster can include at least one inactive spare arithmetic-memory stack.

[0009] In other aspects of the present disclosure, the package substrate can include two computing clusters, and each computing cluster can include eight or more compute-memory stacks. Additionally, each computing cluster can include at least one inactive spare compute-memory stack.

[0010] In still further aspects of the present disclosure, the plurality of compute-memory stacks are stacked 3D-DRAM dielets, and the computing package is configured to operate as a large model processing unit.

[0011] In yet other aspects of the present disclosure, the input / output die can be further configured to communicate with at least one of an external DRAM or an external Remote Direct Memory Access (RDMA) for interconnecting with other computing packages.

[0012] In other aspects of the present disclosure, a computing method can include receiving a processing command in one or more computing clusters disposed on a package substrate, and performing a computing operation based on the processing command using a plurality of compute-memory stacks that communicate with an input / output die, where each compute-memory stack includes a low-power compute die and a plurality of memory dies stacked thereon, and the method can further include transmitting data from the input / output die via one or more peripheral component interconnects.

[0013] In yet another aspect of the present disclosure, each low-power compute die of the plurality of compute-memory stacks can be configured to operate with a power supply of about 40W or less.

[0014] In still further aspects of the present disclosure, for a particular compute-memory stack among the plurality of compute-memory stacks, the low-power compute die can have a footprint on the package substrate that is less than 30% larger than the footprint of the plurality of memory dies.

[0015] In yet other aspects of the present disclosure, performing a computing operation can further include performing the computing operation using a plurality of computing clusters disposed on a package substrate, wherein input / output dies for each computing cluster are connected to one or more input / output dies of other computing clusters disposed on the package substrate.

[0016] In other aspects of the present disclosure, the input / output die can be further configured to communicate with at least one of an external DRAM or an external remote direct memory access (RDMA) for interconnecting with other computing packages.

[0017] In yet other aspects of the present disclosure, a large-scale model processing unit can include one or more computing packages connected via a peripheral component interconnect, each computing package can include a package substrate and one or more computing clusters disposed on the package substrate, each of the one or more computing clusters includes a plurality of compute-memory stacks that communicate with an input / output die, each compute-memory stack includes a low-power compute die and a plurality of memory dies stacked thereon, and the input / output die is configured to transmit data regarding the plurality of compute-memory stacks via one or more peripheral component interconnects.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

[0019] Detailed Description This technology generally relates to high - bandwidth processing in low - compute - intensity environments. For example, systems and methods for use in large - model processing are disclosed, which may include the use of machine - learning accelerators contributing to large - language models (LLMs). Large - model processing may refer to the ability to achieve general - purpose understanding and / or generation by training one or more machine - learning models with large amounts of data. An example of a large - model processing unit can be an LLM. A machine - learning accelerator may include multiple chiplets each having a 3D memory die. A machine - learning accelerator can be designed to leverage the low compute intensity associated with LLM processing by including multiple low - power, small compute dies on each of the multiple chiplets.

[0020] According to an implementation form, the disclosed systems and methods provide high - bandwidth processing using multiple compute - memory chiplets. According to one aspect, a computing package is configured to have multiple compute - memory chiplets for performing processing operations related to large - language models. According to another aspect, a compute - memory chiplet is configured to operate using small, low - power computing dies that are efficiently operable for low - compute - intensity workloads.

[0021] FIG. 1 is a block diagram 100 of a multi-chip package 101 operable as a machine learning accelerator according to an aspect of the present disclosure. The multi-chip package 101 may include a substrate 110 on which a plurality of computing clusters 102a-d are disposed. Each of the clusters 102a-d may include input / output (IO) dies 108a-d and a plurality of compute-memory chiplets 104a-e. Each chiplet 104 may include a compute die 120 and a 3D stack of a plurality of memory dies 122. For example, as shown in block diagram 100, chiplet 104a of cluster 102a has a compute die 120 stacked with a plurality of memory dies 122. The other chiplets 104b-e of cluster 102a and the chiplets 104a-e of the other clusters 102b-d may include similar sets of memory dies 122 stacked on the compute die 120.

[0022] Each of the chiplets 104a-e within the cluster 102 is connected to a corresponding IO die 108 within the cluster 102. For example, as shown in block diagram 100, chiplet 104a is connected to IO die 108a via a connection 131. Similar connections exist between IO die 108a and each of the other chiplets 104b-e of cluster 102a. Further, the clusters 102a-d of the multi-chip package 101 may be connected to each other via the IO dies 108a-d. For example, each of the IO dies 108a-d may be connected to all of the other IO dies 108a-d via a connection 130. For the sake of brevity, only a portion of the connections 130 in block diagram 100 are identified by reference numerals. Each of the chiplets 104a-e within the cluster 102 may also be configured to communicate directly with each other via additional connections (not shown).

[0023] The IO dies 108a - d of each cluster 102a - d may also be configured to transmit to components or devices external to the multi - chip package 101. For example, the connection parts 132a - d may be configured such that the IO dies 108a - d can communicate with an external dynamic random access memory (DRAM) or an external core - to - core interconnect (ICI). Further, each of the IO dies 108a - d may be configured to transmit data via a peripheral component interconnect, for example, peripheral component interconnect express (PCIe) connection parts 134a - d. According to an implementation form, the input / output die is configured to transmit data regarding one or more compute - memory stacks via one or more peripheral component interconnects, enabling efficient data communication between multi - chip packages such that a computing package including one or more multi - chip packages can operate efficiently as a machine learning accelerator. As an example of a peripheral component interconnect, as shown in FIG. 1, each of the PCIe connection parts 134a - d may be configured to transmit data regarding all the chiplets 104a - e within a given cluster. Thus, the multi - chip package 101 can be configured to communicate with other multi - chip packages and to operate as a machine learning accelerator that is part of a large - language model server.

[0024] FIG. 2 is a block diagram 200 of a partial side view of a multi-chip package 101 showing compute dies 120 and memory dies 122 for chiplets 104a-c. Each of the chiplets 104a-c includes a plurality of memory dies 122 stacked on the compute die 120. Each compute die 120 can be directly connected to a substrate 110, such as a circuit board, via an electrical connection portion 112. The plurality of memory dies 122 may be a 3D stack of HBM dies stacked on the compute die 120, and each memory die 122 may be electrically connected to an adjacent memory die 122 in the stack such that HBM electrical signals can pass through the chiplet 104. As shown in block diagram 200, the compute die 120 and the memory dies 122 can be configured to include one or more through-silicon vias (TSV210) extending through the stack of dies within the chiplet 104. The electrical connection between the compute die 120 and the plurality of memory dies may exist via the TSV210. Each of the chiplets 104a-c is shown as including four memory dies 122, but the number of memory dies 122 on a given chiplet can vary based on the capacity required for a given application and based on the total number of chiplets 104 included in the package 101. In particular, if more chiplets are used per package 101, the number of memory dies 122 required for one chiplet can be fewer. Thus, increasing the number of chiplets 104 can lower the height of the package 101 and reduce the thermal constraints and cooling requirements that would exist if a larger stack of memory dies 122 were used.

[0025] By each chiplet 104 including its own arithmetic die 120, the chiplet 104 can be designed using an arithmetic die that is smaller than would be required if only a single arithmetic die were used for all memory dies 122 within the cluster or within the package 101. According to aspects of the present disclosure, the arithmetic die 120 of each chiplet 104 can be designed to have the same or a similar footprint as the memory die 122. For example, in FIG. 200, the arithmetic die 120 is shown as having the same footprint as the memory die 122 with which they are stacked. Alternatively, the arithmetic die 120 may have a footprint on the package substrate that is less than 30% larger than the footprint of the memory die 122 with which they are stacked. Thus, the package 101 may be capable of accommodating a large number of small and low-power chiplets 104. Further, the stacked configuration of the arithmetic die 120 and the memory die 122 enables the chiplets 104 of the package 101 to be directly connected to the substrate 110 without the need to use an interposer.

[0026] For example, FIG. 3 is a block diagram 300 of a package 301 that does not use the arithmetic-memory chiplet 104 described in connection with FIGS. 1 and 2. In particular, the arithmetic die 320 and the memory stack 310 are each connected to an interposer 340, and the interposer 340 itself is connected to a substrate 330 via an electrical connection portion 352. The arithmetic die 320 communicates with a plurality of memory dies 322 within the memory stack 310 via the interposer 340. Thus, the number, size, and arrangement of the arithmetic die 320 and the memory die 322 are limited by the interposer 340 used. Even further, the package 301 is configured such that a single arithmetic die 320 performs all of the processing for the entire package 301, and the single arithmetic die 320 is required to access all of the 10 memory dies 322 within the memory stack 310. In this configuration, the arithmetic die 320 has a footprint that is more than twice the size of the memory die 322.

[0027] Returning to FIG. 1, package 101 may have dimensions such that it can fit within other packages on an existing substrate. The specific numbers described herein are merely exemplary and not intended to be limiting, but by way of example only, package 101 may have a length and width of approximately 60 mm × 60 mm with respect to the substrate, and the arithmetic dies 120 may each have a footprint of approximately 100 mm 2 and the I / O dies 108 may each have a footprint of approximately 150 mm 2 . The memory die 122 may be a 3D stacked dynamic random access memory (DRAM) die with a bandwidth of approximately 800 GB / second, a capacity of approximately 4 GB, and a thermal design power (TDP) of approximately 10 W per die. Each of the arithmetic dies 120 may have a thermal design power of approximately 40 W or less and a processing capacity of approximately 32 TFLOPs, and may also have approximately 32 MB of SRAM. The thermal design power of the I / O dies 108a - d may be approximately 35 W TDP.

[0028] Package 101 may enable high-bandwidth communication among the IO dies 108. As an example, the total IO bandwidth per IO die 108 can be up to several TB / second, for example, up to 4TB / second. For example, a bandwidth of 400GB / second may be allocated to each compute die 120 within cluster 102, or a bandwidth of 400GB / second may be allocated to each IO die 108 within package 101. As another example, when high-bandwidth IO die communication is not required, the total IO bandwidth per IO die 108 can be lower, for example, about 400GB / second. For example, for each IO die 108, a bandwidth of 160GB / second may be allocated to the compute die 120 within cluster 102, a bandwidth of 150GB / second may be allocated to other IO dies 108 within package 101, a bandwidth of 16GB / second may be allocated to the host, for example, a host device supporting PCIe Gen5, a bandwidth of 64GB / second may be allocated to external DRAM, and a bandwidth of 50GB / second may be allocated to external RDMA. The IO die 108 can include light computations for smart routing, such as aggregation of partial sums computed by a memory-compute stack connected to the IO die 108, and can be connected to external components for additional processing.

[0029] According to aspects of the present disclosure, each cluster 102 of the package 101 can be configured such that at least one die 104 within the cluster 102 is initially designated as a cold spare. For example, four out of the five dies 104a - e of cluster 102a may be designated as active, while the fifth die 104e is designated as an inactive spare. Thus, the IO die 108 communicates only with dies 104a - d for the processing operations of the machine learning accelerator. However, the IO die may also be configured to receive and transmit diagnostic information regarding the operation of each of the active dies 104a - d and to replace any problematic die 104a - d with the spare die 104e. For example, if the active die 104a is determined to have experienced a failure or is not operating correctly for other reasons, die 104a can be re - designated from an active die to an inactive die, and the spare die 104e can be re - designated from an inactive spare die to an active die. The package 101 can be designed to increase reliability in that the operation of the machine learning accelerator as a whole is not impaired even if a failure occurs in any one of the dies 104. In this configuration, the package 101 will have four spare dies and sixteen active dies.

[0030] In an example where each of the 16 active computing dies 120 has a TDP of 40W, each of the 16 DRAM dies has a TDP of 10W, and each of the 4 IP dies has a TDP of 35W, the total TDP of the package 101 as a whole is 620W. However, even when using low-power computing dies 120 with a TDP of approximately 10 - 40W, by using multiple compute-memory chiplets, the package 101 can increase the overall bandwidth compared to a conventional HBM configuration and execute large language model processing. This increased bandwidth is a function of the number of chiplets 104 used within the package. Since the computing dies 120 are small-sized and low-power consuming, it becomes possible to use a larger number of chiplets 104 per package. On the other hand, because the computational intensity of large language model processing is low, the efficient operation of the small and low-power computing dies 120 is not hindered.

[0031] In the case of large language models, the computational intensity of the processing is orders of magnitude lower compared to other machine learning processes. With the current architectures for machine learning accelerators, overprovisioning by more than 10 times is possible for this level of low computational intensity. However, in relation to the package 101, due to the size (100mm 2 ) and thermal design power (10 - 40W) of the computing dies 120, it becomes possible to use a larger number of compute-memory chiplets 104. By using this larger number of compute-memory chiplets 104, the available bandwidth for applications with low computational intensity can be increased without increasing the operating cost compared to conventional machine learning accelerators.

[0032] According to aspects of the present disclosure, the package 101 can be configured to operate in a manner such that multiple levels of sharding are performed when distributing processing to various dielets 104. For example, when a large language model is sharded to 16GB in relation to a conventional machine learning accelerator, the system disclosed herein can be configured to execute the same sharding algorithm where each compute-memory stack is a shard, but the difference is that the memory capacity per shard (4GB) is smaller than the shard size (16GB) used by the conventional machine learning accelerator described above.

[0033] According to aspects of the present disclosure, a machine learning accelerator can utilize one or more packages having different dimensions and compute-memory dielets and other components of different specifications than those described above. For example, FIG. 4 is a block diagram 400 of another exemplary package 401 that can be implemented as part of a machine learning accelerator for a large language model. The package 401 can be configured to operate in a manner similar to the package 101 described above. However, as shown in FIG. 4, the package 401 includes two computing clusters 402a - b. Each cluster 402 includes an IO die 408 and nine dielets 404a - i.

[0034] Each dielet 404a - e within clusters 402a - b can be respectively connected to the corresponding IO die 408a - b within clusters 402a - b. For example, as shown in the block diagram 400, dielet 404a is connected to IO die 408a via connection 431. Similar connections exist between IO die 408a and each of the other dielets 404b - i of cluster 402a. Further, clusters 402a - b of the multi-chip package 401 can be connected to each other via IO dies 408a and 408b. For example, IO dies 408a and 408b can be connected via connection 430. Each dielet 404a - i within cluster 402 can also be configured to communicate directly with each other via additional connections (not shown).

[0035] The IO dies 408a and 408b of each of the clusters 402a and 402b may also be configured to transmit to components or devices external to the multi-chip package 401. For example, the connections 432a-b may be configured such that the IO dies 408a-b can communicate with an external dynamic random access memory (DRAM) or an external core interconnect (ICI). Further, the IO dies 408a-b may each be configured to transmit data via the peripheral component interconnect express (PCIe) connections 434a-b. Each of the PCIe connections 434a-b shown in FIG. 4 may be configured to transmit data regarding all of the dielets 104a-i within a given cluster. Thus, the multi-chip package 401 may be configured to communicate with other multi-chip packages and may be configured to operate as a machine learning accelerator that is part of a large language model server.

[0036] The architecture of the package 401 may be such that it includes a different number of spares than the example described above for the package 101 that provided four spare dielets 104 (one per cluster) and sixteen active dielets 104 (four per cluster). For example, the package 401 may include one spare per cluster 402a and 402b. Thus, the package 401 may operate with sixteen active dielets 404 and four non-active spare dielets 404 at any given time.

[0037] The specific numbers described herein are merely exemplary and are not intended to be limiting, but by way of example, the package 401 may include a 3D stacked dynamic random access memory (DRAM) die with a bandwidth of about 800 GB / second, a capacity of 4 GB, and a TDP of 10 W per DRAM stack. The package 401 may have a footprint of 100 mm for one die 2, may further include an arithmetic die with a TDP of 20W, a processing power of 32 TFLOPs, and 32 MB of SRAM. The IO dies 408a - b may have a footprint of about 150 mm for one die 2 , and may have a TDP of 50W. The total IO bandwidth per IO die 408 may be about 450 GB / second, for example, 144 GB / second to the arithmetic die, 128 GB / second to other IO dies, 32 GB / second to a host such as PCIe Gen5, 64 GB / second to external DRAM, and / or 50 GB / second to external RDMA. The package 401 may have an overall TDP of about 580W, with the arithmetic dies using a total of about 320W, the memory dies using about 160W, and the IO dies using about 100W.

[0038] FIG. 5 is a block diagram 500 of another package 501 operable as a machine learning accelerator. In the case of package 501, the clusters 502a - d include memory stacks 504a - e that include a plurality of memory dies 522 but no computing dies. Instead, the computing operations and input / output operations are performed by the arithmetic and IO dies 508a - d. In addition to the substrate 510, the package 501 includes interposers 550a - d for each of the clusters 502a - d. In this case, each cluster 502a - d may include an interposer 550a - d connected to the substrate 510. The arithmetic and IO dies 508 and the plurality of memory stacks 504a - e are then connected to one of the interposers 550a - d.

[0039] Each memory stack 504a - e within cluster 502 is connected to the corresponding compute and IO die 508 within cluster 502. For example, as shown in block diagram 500, memory stack 504a is connected to IO die 508a via connection part 531. Similar connection parts exist between compute and IO die 508a and each of the other memory stacks 504b - e of cluster 502a. Further, clusters 502a - d of package 501 can be connected to each other via compute and IO dies 508a - d. For example, each compute and IO die 508a - d may be connected to all other compute and IO dies 508a - d via connection part 530. For the sake of brevity, only a part of the connection parts 530 in block diagram 500 are specified by reference numbers.

[0040] The compute IO dies 508a - d of each cluster 502a - d may also be configured to perform transmissions to components or devices external to package 501. For example, connection parts 532a - d can be configured such that compute and IO dies 508a - d can communicate with an external dynamic random access memory (DRAM) or an external core - to - core interconnect (ICI). Further, each of the compute and IO dies 508a - d can be configured to transmit data via a peripheral component interconnect express (PCIe) connection part 534a - d. Each PCIe connection part 134a - d shown in FIG. 5 can be configured to transmit data regarding all memory stacks 504a - e within a given cluster 502. Thus, the multi - chip package 501 can be configured to communicate with other multi - chip packages and can be configured to operate as a machine learning accelerator that is part of a large - language - model server.

[0041] The specific numbers described herein are merely exemplary and not intended to be limiting. As a mere example, package 501 may have a length and width of approximately 80 mm × 80 mm with respect to the substrate, and compute and IO dies 508a - d may be approximately 200 - 300 mm 2may each have a footprint, and the interposers 550a - d may each have a footprint of about 1000 mm 2 may each have a footprint. The memory die 122 can be a 3D stacked dynamic random access memory (DRAM) die with a bandwidth of about 800 GB / second, a capacity of about 16 GB, and a TDP of about 35 W per stack. The compute and IO dies 508a - d may each have a TDP of about 150 W, a processing power of about 128 TFLOPs, and may have about 128 MB of SRAM. The thermal design power of the IO dies 108a - d can be about 35 W TDP. Further, each cluster can have four active memory stacks and one non - active spare stack. The total TDP of the 16 active memory stacks 104 can be about 560 W, and the total TDP of the four compute and IO dies can be about 600 W, but the interposer and the package may further require 100 W. Thus, in this example, the total TDP of the package 501 can be 1300 W.

[0042] Figure 6 is a flowchart of an exemplary process 600 for performing computing operations according to an aspect of the present disclosure. The exemplary process 600 can be executed on a system of one or more processors in one or more locations, such as the multi - chip package 101 depicted in FIG. 1.

[0043] As shown in block 610, the multi - chip package 101 receives a processing command in one or more computing clusters disposed on a package substrate. The processing command can include any instructions for processing data to perform any of a variety of computing operations. For example, these computing operations can include loading data into a computing circuit, moving data to one or more processing elements of a computing circuit, processing data by one or more processing elements, and pushing data out of the computing circuit.

[0044] As shown in block 620, the multi-chip package 101 executes computing operations based on processing commands using a plurality of compute-memory stacks that communicate with input / output dies. Each compute-memory stack includes a low-power compute die and a plurality of memory dies stacked thereon. Each low-power compute die can be configured to operate with a power supply of about 40W or less. For a particular compute-memory stack, the low-power compute die can have a footprint on the package substrate that is less than 30% larger than the footprint of the plurality of memory dies.

[0045] Execution of the computing operations can further include executing the computing operations using a plurality of computing clusters disposed on the package substrate. The input / output die for each computing cluster can be connected to one or more input / output dies of other computing clusters disposed on the package substrate. For example, the package substrate can include four computing clusters, and each computing cluster can include four or more active compute-memory stacks. As another example, the package substrate can include two computing clusters, and each computing cluster can include eight or more active compute-memory stacks. Each computing cluster can include at least one inactive spare compute-memory stack. The plurality of compute-memory stacks can be stacked 3D-DRAM chiplets.

[0046] As shown in block 630, the multi-chip package 101 transmits data from the input / output die via one or more peripheral component interconnects. The input / output die can be further configured to communicate with at least one of an external DRAM or an external RDMA for interconnecting with other computing packages.

[0047] Unless otherwise specified, the foregoing alternative examples are not mutually exclusive and may be implemented in various combinations to achieve their respective advantages. Since these and other variations and combinations of the features described above may be utilized without departing from the subject matter defined by the claims, the foregoing description of the embodiments should be understood as illustrative rather than as a limitation of the subject matter defined by the claims. Further, the provision of the examples described herein and phrases expressed as "such as" and "including" should not be construed as limiting the subject matter of the claims to specific examples, but rather these examples are intended to illustrate only one of many possible embodiments. Further, the same reference numerals in different drawings may identify the same or similar elements.

Claims

1. A package substrate, One or more computing clusters disposed on the package substrate, and Comprising, Each of the one or more computing clusters includes a plurality of arithmetic-memory stacks that communicate with an input / output die, Each arithmetic-memory stack includes a low-power arithmetic die and a plurality of memory dies stacked thereon, The input / output die is configured to transmit data regarding the plurality of arithmetic-memory stacks via one or more peripheral component interconnects, A computing package.

2. The computing package according to claim 1, wherein each low-power arithmetic die of the plurality of arithmetic-memory stacks is configured to operate with a power supply of about 40 W or less.

3. For a particular arithmetic-memory stack from the plurality of arithmetic-memory stacks, the low-power arithmetic die has a footprint on the package substrate that is less than 30% larger than the footprint of the plurality of memory dies. The computing package according to claim 1.

4. The computing package according to claim 1, further comprising a plurality of computing clusters disposed on the package substrate, and the input / output die for each computing cluster is connected to one or more input / output dies of other computing clusters disposed on the package substrate.

5. The computing package according to claim 4, wherein the package substrate includes four computing clusters, and each computing cluster includes four or more arithmetic-memory stacks.

6. The computing package according to claim 4, wherein the package substrate includes two computing clusters, and each computing cluster includes eight or more arithmetic-memory stacks.

7. The computing package according to claim 1, wherein each computing cluster includes at least one inactive spare arithmetic-memory stack.

8. The computing package according to claim 1, wherein the plurality of arithmetic-memory stacks are stacked 3D-DRAM chiplets.

9. The computing package according to claim 1, configured to operate as a large-scale model processing unit.

10. The computing package according to claim 1, wherein the input / output die is further configured to communicate with at least one of an external DRAM or an external remote direct memory access (RDMA) for interconnecting with other computing packages. **Claim 11** A computing method, comprising: receiving a processing command in one or more computing clusters disposed on a package substrate; and executing a computing operation based on the processing command using a plurality of arithmetic-memory stacks that communicate with an input / output die, each arithmetic-memory stack including a low-power arithmetic die and a plurality of memory dies stacked thereon; The computing method further includes: transmitting data from the input / output die via one or more peripheral component interconnects. **Claim 12** The method according to claim 11, wherein each low-power arithmetic die of the plurality of arithmetic-memory stacks is configured to operate with a power supply of about 40 W or less. **Claim 13** For a particular arithmetic-memory stack of the plurality of arithmetic-memory stacks, the low-power arithmetic die has a footprint on the package substrate that is less than 30% larger than the footprint of the plurality of memory dies. **Claim 14** Executing the computing operation further includes executing the computing operation using a plurality of computing clusters disposed on the package substrate, and the input / output die for each computing cluster is connected to one or more input / output dies of other computing clusters disposed on the package substrate. **Claim 15** The method according to claim 14, wherein the package substrate includes four computing clusters, and each computing cluster includes four or more arithmetic-memory stacks. **Claim 16** The method according to claim 14, wherein the package substrate includes two computing clusters, and each computing cluster includes eight or more active arithmetic-memory stacks. **Claim 17** The method according to claim 11, wherein each computing cluster includes at least one inactive spare arithmetic-memory stack.

18. The method of claim 11, wherein the plurality of compute-memory stacks are stacked 3D-DRAM dielets.

19. The method of claim 11, wherein the input / output die is further configured to communicate with at least one of an external DRAM or an external remote direct memory access (RDMA) for interconnecting with other computing packages.

20. A large model processing unit comprising one or more computing packages connected via a peripheral component interconnect, each computing package comprising a package substrate, and one or more computing clusters disposed on the package substrate and comprising each of the one or more computing clusters includes a plurality of compute-memory stacks that communicate with an input / output die, each compute-memory stack includes a plurality of memory dies stacked with a low-power compute die, the input / output die is configured to transmit data regarding the plurality of compute-memory stacks via one or more of the peripheral component interconnects.

Citation Information

Patent Citations

  • Integrated circuit device, electronic equipment, board card and computing method

    CN112686379A

  • Data transmission method between chiplets

    CN112732631A

  • Packaging structure, device, board card and method for arranging integrated circuit

    CN114330201A

  • A device enabling simultaneous communication between an interface die and multiple die stacks, an interleaved conductive path in a stack-type device, and a method for forming and operating the same.

    JP2013524519A

  • Intelligent high bandwidth memory system and logic dies therefor

    JP2019036298A