Three-dimensional stacking processing system
By adopting a combination of multiple processor core particles, input/output modules and active intermediary layers in the processing system, and using three-dimensional stacking technology, the limitations of memory capacity, bandwidth, power efficiency and form factor in the prior art are solved, and a high-performance memory and processing solution is realized.
Patent Information
- Application Number
- CN202080102801.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-17
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-09-17
AI Technical Summary
Existing near-memory processing (PNM) solutions based on high bandwidth memory (HBM) have limited memory capacity, bandwidth, power efficiency and large form factor problems.
By using a combination of multiple processor cores, input/output modules and active intermediary layers in the processing system, the memory die is coupled with the logical die using three-dimensional stacking technology to achieve high memory capacity, high memory bandwidth, low power consumption and small form factor.
It achieves high memory capacity, improves memory bandwidth, reduces power consumption, and reduces form factor, meeting the needs of high-performance memory and processing solutions.
Smart Images

Figure CN115868023B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of integrated circuit technology, and more particularly to a three-dimensional stacked processing system. Background Art
[0002] Computing systems have made significant contributions to the progress of modern society and are used in many applications to obtain favorable results. Many devices, such as desktop personal computers (PCs), laptops, tablets, netbooks, smartphones, servers, etc., have promoted the improvement of productivity and cost reduction in communication and data analysis in most fields of entertainment, education, business, and science. Many technologies, such as high-end graphics, 40G / 100G Ethernet, exascale high-performance computing, etc., require high-density memory, high memory bandwidth, and low-power solutions.
[0003] Reference Figure 1A And Figure 1B show a high-bandwidth memory (HBM)-based near-memory processing (PNM) solution according to conventional techniques. The HBM PNM solution may include a three-dimensional (3D) stack of memory dies 110-125 and a logic die 130. The logic die 130 and the memory dies 110 to 125 are of the same size, so the logic die size is limited by the memory die size. In addition, the memory die manufacturing yield limits the size of the memory dies 110-125, and thus also limits the die size of the logic die 130.
[0004] The 3D stack of the memory dies 110-125 and the logic die 130 can be integrated with a processing unit or other system-on-chip 140 using an interposer 145. For example, in conventional 2.5D integration, the 3D stack of the memory dies 110 to 125 and the logic chip 130 can be integrated with a GPU using an interposer 145.
[0005] As Figure 1B shown, the interposer 145 can provide physical layer (PHY) logic 150, such as address command logic, data (DQ) line transceiver logic, and signal connectivity test logic. The interposer 145 can also provide through-silicon via regions 155 for coupling power to the memory dies 110-125. The interposer 145 can also provide a test design region 160 including direct access connections.
[0006] Now refer to Figure 2, multiple HBM PNMs 210-220 can be further integrated on a printed circuit board (PCB) 230 using a conventional interconnection such as a Peripheral Component Interconnect Express (PCIe) fan-out switch 240. However, conventional HBM PNN solutions are characterized by limited memory capacity, limited memory bandwidth, low power efficiency, and a large form factor. Accordingly, there is a continuing need for improved high-performance memory and processing solutions. SUMMARY OF THE INVENTION
[0007] The present disclosure may best be understood by reference to the following description and drawings, which are used to illustrate embodiments of the present disclosure directed to a processing system including a plurality of processor dies and input / output circuitry directly coupled to each of the plurality of processor dies. The processing system may be characterized by high memory capacity, high memory bandwidth, low power consumption, and a small form factor.
[0008] In one embodiment, a processing system may include a plurality of processor dies, an input / output module, and an interposer. Each of the plurality of processor dies may include a plurality of memory dies coupled to a corresponding logic die. The interposer may be configured to couple each processor die to one or more other processor dies of the plurality of processor dies. The interposer may also be configured to couple each processor die to the input / output module. The interposer may also be configured to couple the input / output module to a plurality of external contacts of the processing system.
[0009] In another embodiment, a processing system may include a plurality of processor dies and an active interposer. Each of the plurality of processor dies may include a plurality of memory dies coupled to a corresponding logic die. The active interposer may include input / output circuitry. The active interposer may be configured to couple each processor die to one or more other processor dies of the plurality of processor dies. The active interposer may also be configured to couple the plurality of processor dies to the input / output circuitry. The active interposer may also be configured to couple the input / output circuitry to a plurality of external contacts of the processing system.
[0010] This Summary of the Invention is provided to introduce in a simplified form a few concepts that are further described in detail in the Detailed Description below. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Embodiments of the present disclosure are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals indicate like elements and in which:
[0012] Figure 1A and 1BShows a high - bandwidth memory (HBM) - based near - memory processing (PNM) solution according to conventional techniques.
[0013] Figure 2 Shows an HBM - based PNM interconnection solution according to conventional techniques.
[0014] Figure 3 Shows a processing system according to some aspects of the present disclosure.
[0015] Figure 4 Shows a processing system according to some aspects of the present disclosure. Detailed Description
[0016] Reference will now be made in detail to embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. While the present disclosure will be described in conjunction with these embodiments, it should be understood that they are not intended to limit the present disclosure to these embodiments. On the contrary, the present invention is intended to cover alternatives, modifications, and equivalents that may be included within the scope of the present invention as defined by the appended claims. Furthermore, in the following detailed description of the present disclosure, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it should be understood that the present disclosure may be practiced without these specific details. In other instances, well - known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure some aspects of the present disclosure.
[0017] Some embodiments of the present disclosure below are presented in the form of routines, modules, logic blocks, and other symbolic representations of operations on data within one or more electronic devices. The description and representation are means by which those skilled in the art most effectively convey the substance of their work to others skilled in the art. Routines, modules, logic blocks, etc. are generally considered herein to be self - consistent sequences of processes or instructions that result in a desired outcome. These processes include physical operations on physical quantities. Typically, although not necessarily, these physical operations take the form of electrical or magnetic signals capable of being stored, transmitted, compared, and otherwise manipulated in an electronic device. For convenience and in accordance with common usage, in the context of the embodiments of the present disclosure, these signals are referred to as data, bits, values, elements, symbols, characters, terms, numbers, strings, etc.
[0018] However, it should be borne in mind that these terms are to be interpreted as referring to physical operations and quantities and are merely convenient labels and will be further interpreted in accordance with terms commonly used in the art. Unless explicitly stated otherwise from the discussion below, it should be understood that by the discussion of the present disclosure, the discussion using terms such as "receiving" refers to the actions and processes of an electronic device (such as an electronic computing device that manipulates and transforms data). Data is represented as physical (e.g., electrical) quantities in the logic circuits, registers, memory, etc. of an electronic device and is transformed into other data similarly represented as physical quantities in the electronic device.
[0019] In this application, the use of conjunctive connectives is intended to include conjunctions. The use of definite or indefinite articles is not intended to denote cardinality. In particular, reference to "the" or "a" object is intended to denote one of potentially multiple such objects. The use of terms such as "comprising" and "including" specifies the presence of the stated elements, but does not preclude the presence or addition of one or more other elements and / or groups thereof. It should also be understood that although terms such as first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used herein to distinguish one element from another. For example, without departing from the scope of the embodiments, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. It should also be understood that when an element is referred to as being "coupled" to another element, it may be directly or indirectly connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly connected" to another element, no intervening elements are present. It should also be understood that the term "and / or" includes any and all combinations of one or more related elements. It should also be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting.
[0020] Reference Figure 3 , a processing system is shown in accordance with some aspects of the present disclosure. The processing system 300 may include a plurality of processor chiplets 305 - 340, an input / output module 345, and an interposer 350. The interposer 350 may couple each processor chiplet 305 - 340 to one or more other processor chiplets among the plurality of processor chiplets 305 - 340. The interposer 300 may also couple the plurality of processor chiplets 305 - 340 to the input / output module 345. The interposer 350 may also couple the input / output module 345 to a plurality of external contacts 355 of the processing system 300. In one implementation, the plurality of processor chiplets may be a plurality of near-memory processing (PNM) chiplets, a plurality of processors, and near-memory architectures, etc. A chiplet may be an integrated circuit die and / or a group of integrated circuit dies that are designed to work together with other similar chiplets to form a larger and more complex chip. In the case of near-memory processing or near-memory, memory and logic are incorporated into an integrated circuit package. In one implementation, the input / output module may be an input / output chiplet. In one implementation, the plurality of processor chiplets, the input / output module, and the interposer may be embodied in a system-in-package (SiP), a multi-chip module (MCM), a chip stack, etc.
[0021] Each processor die 305-340 may include a plurality of memory dies 360-375 coupled to a corresponding logic die 380. The plurality of memory dies 360-375 and the corresponding logic die 380 may be implemented as a three-dimensional (3D) die stack. In one embodiment, the memory dies 360-375 may be random access memory (RAM) dies, such as but not limited to DDR3, DDR4, GDDR5, etc. In one embodiment, each of the memory dies 360-375 may include a plurality of memory banks, each located in one of a plurality of memory channels. For example, memory die 360 may include eight banks in each of two memory channels. Additionally, the banks may be further organized into sub-banks. In one embodiment, each of the memory dies 360-375 may be organized as a corresponding memory slice. In one embodiment, the plurality of memory dies 360-375 and the corresponding logic die 380 may implement high bandwidth memory (HBM). In another embodiment, the plurality of memory dies 360-375 and the corresponding logic die 380 may implement hybrid memory cube (HMC). In a non-limiting example, an HBM 3D stacked device may include 4 to 8 memory dies of 2, 4, or 8 GB. An exemplary HBM 3D stacked device may achieve 1 to 2 Gbps per pin and a bandwidth of 128 to 256 Gbps. Each of the logic dies 380 may include computing logic. For example, each of the logic dies 380 may implement one or more processing units, one or more graphics processing units, one or more encoder / decoder engines, one or more artificial intelligence (AI) engines, one or more digital signal processors (DSPs), etc. Each logic die 380 may also include through-silicon via regions, physical layers, design for test (DFT) regions, etc.
[0022] In one embodiment, the plurality of memory dies 360-375 may be coupled together by through-silicon vias (TSVs) or a combination of through-silicon vias and an array of micro-bumps (uBUMPs). In one embodiment, the plurality of memory dies 360-375 may be further coupled to the corresponding logic die 380 by through-silicon vias or a combination of through-silicon vias and an array of micro-bumps (uBUMPs). The through-silicon vias or the combination of through-silicon vias and an array of micro-bumps (uBUMPs) may also provide a power potential to the plurality of memory dies 360-375. For example, the through-silicon vias may couple an external power potential of 1.2 volts (V) to the plurality of memory dies 360-375.
[0023] The input / output module 345 can be configured to provide one or more communication interfaces. For example, the input / output module 345 can provide one or more Peripheral Component Interconnect Express (PCIe) interfaces, one or more Double Data Rate (DDR) interfaces, and one or more open source interfaces, etc. The input / output module 345 can be configured to provide 1024-bit data access to multiple processor dies 305-340, and the access granularity is 256 bytes. In one implementation, the input / output module 345 and the processor dies 305-340 can be coupled together in a mesh network topology. In other implementations, the input / output module 345 and the processor dies 305-340 can be coupled together in a star, ring, bus, daisy chain, or other similar topologies. In one implementation, each processor die 305-340 can be directly connected to the input / output module 345 through an interposer layer 350. The interposer layer 350 can also directly connect multiple processor dies 305-340 to the output / input module 345. The input / output module 345 can also include computing logic, such as but not limited to one or more host controllers, one or more arbitration engines, one or more Design for Test (DFT) engines, one or more data compression engines, one or more error correction code engines, etc.
[0024] In one implementation, multiple processor dies 305-340 can be of the same size, multiple memory dies 360-375 can have the same memory density, and the logic die 380 can provide the same function. In such an implementation, the processing system 300 can be homogeneous. In another implementation, multiple processor dies 305-340 can be of different sizes, or a subset of the processor dies can be of different sizes. For example, the processor dies can include memory dies 360-375 with different memory densities, and / or the logic die 380 can perform different functions or subsets of functions can be different among the logic dies 380 of different processor dies. In this implementation, the processing system 300 is heterogeneous.
[0025] The processing system 300 can also include a substrate 385 configured to couple the processing system 300 to one or more other circuits. The substrate 385 can include a first set of ball array contacts 390 for coupling to the interposer layer 350, and a second set of ball array contacts 355 for coupling to one or more other circuits, chips, modules, devices, etc.
[0026] Now refer to Figure 4, showing a processing system in accordance with some aspects of the present disclosure. The processing system 400 may include a plurality of processor dies 405-445 and an active interposer 450. The active interposer 450 may include input / output circuitry. The active interposer 450 may couple each processor die 405-445 to one or more other processor dies among the plurality of processor dies 405-445. The input / output circuitry of the active interposer 450 may also be coupled to the plurality of processor dies 405-445. The output / input circuitry of the active interposer 450 may also be coupled to a plurality of external contacts 455 of the processing system 400. In one implementation, the plurality of processor dies may be a plurality of near-memory processing (PNM) dies, a plurality of multi-processors and near-memory architectures, etc. A die may be an integrated circuit die and / or a set of integrated circuit dies that are designed to work together with other similar dies to form a larger and more complex chip. In the case of near-memory processing or near-memory, memory and logic are incorporated into an integrated circuit package. In one implementation, the input / output module may be an input / output die. In one implementation, the plurality of processor dies and the active interposer including the input / output circuitry may be embodied in a system-in-package (SiP), a multi-chip module (MCM), a chip stack, etc.
[0027] Each processor die 405-445 may include a plurality of memory dies 460-475 coupled to a corresponding logic die 480. The plurality of memory dies 460-475 and the corresponding logic die 480 may be implemented as a three-dimensional (3D) die stack. In one embodiment, the memory dies 460-475 may be random access memory (RAM) dies, such as but not limited to DDR3, DDR4, GDDR5, etc. In one implementation, each memory die 460-475 may include a plurality of memory banks, each located in one of a plurality of memory channels. For example, memory die 460 may include eight banks in each of two memory channels. Additionally, the banks may be further organized into sub-banks. In one embodiment, each memory die 460-475 may be organized as a corresponding memory slice. In one embodiment, the plurality of memory dies 460-475 and the corresponding logic die 480 may implement high bandwidth memory (HBM). In another embodiment, the plurality of memory dies 460-475 and the corresponding logic die 480 may implement hybrid memory cube (HMC). In a non-limiting example, an HBM 3D stacked device may include 4 to 8 memory dies of 2, 4, or 8 GB. An exemplary HBM 3D stacked device may achieve 1 to 2 Gbps per pin and a bandwidth of 128 to 256 Gbps. Each logic die 480 may include computing logic. For example, each logic die 480 may implement one or more processing units, one or more graphics processing units, one or more encoder / decoder engines, one or more artificial intelligence (AI) engines, one or more digital signal processors (DSPs), etc. Each logic die 480 may also include through-silicon via regions, physical layers, design for test (DFT) regions, etc.
[0028] In one embodiment, the plurality of memory dies 460-475 may be coupled together by through-silicon vias (TSVs) or a combination of through-silicon vias and an array of micro-bumps (uBUMPs). In one embodiment, the plurality of memory dies 460-475 may be further coupled to the corresponding logic die 480 by through-silicon vias or a combination of through-silicon vias and an array of micro-bumps (uBUMPs). The through-silicon vias or the combination of through-silicon vias and an array of micro-bumps (uBUMPs) may also provide a power potential to the plurality of memory dies 460-475. For example, the through-silicon vias may couple an external power potential of 1.2 volts (V) to the plurality of memory dies 460-475.
[0029] The input / output circuitry of the active interposer 450 may be configured to provide one or more communication interfaces. For example, the input / output circuitry may provide one or more Peripheral Component Interconnect Express (PCIe) interfaces, one or more Double Data Rate (DDR) interfaces, and one or more open source interfaces, etc. The input / output circuitry of the active interposer 450 may be configured to provide 1024-bit data access to multiple processor dies 405-445, and the access granularity is 256 bytes. In one embodiment, the input / output circuitry of the active interposer 450 and the processor dies 405-445 may be coupled together in a mesh network topology. In a mesh topology, the input / output circuitry and each processor die 405-445 may be directly coupled together, dynamically and non-hierarchically coupled to as many other processor dies as possible, and cooperate with each other to effectively route data, instructions, control signals, etc. In other embodiments, the input / output circuitry and the processor dies 405-445 may be coupled together in a star, ring, bus, daisy chain, or other similar topology. The input / output circuitry of the active interposer 450 may further include computing logic, such as but not limited to one or more host controllers, one or more arbitration engines, one or more Design for Test (DFT) engines, one or more data compression engines, one or more error correction code engines, etc.
[0030] In one embodiment, the multiple processor dies 405-445 may be of the same size, the multiple memory dies 460-475 may have the same memory density, and the logic die 480 may provide the same function. In such an embodiment, the processing system 400 may be homogeneous. In another embodiment, the multiple processor dies 405-445 may be of different sizes, or a subset of the processor dies may be of different sizes. For example, the processor dies may include memory dies 460-475 with different memory densities, and / or the logic die 480 may perform different functions or subsets of functions may be different between the logic dies of different processor dies. In this embodiment, the processing system 400 is heterogeneous.
[0031] The processing system 400 may further include a substrate 485 configured to couple the processing system 400 to one or more other circuits. The substrate 485 may include a first set of ball array contacts 490 for coupling to the interposer 450, and a second set of ball array contacts 455 for coupling to one or more other circuits, chips, modules, devices, etc. In one embodiment, fine pitch ball grid array (FBGA) contacts 495 may be utilized to couple the active interposer 450 to the substrate 485. Ball grid array (BGA) contacts 455 may be utilized to couple the substrate 485 to one or more other circuits, chips, modules, devices, etc.
[0032] According to some aspects of the present disclosure, through-silicon via coupling of multiple memory dies and one or more logic dies can advantageously provide high-bandwidth communication channels. For example, in a processing system according to some aspects of the present disclosure, a bandwidth of 1024 Gbps or 256 GBps can be achieved. In contrast, the bandwidth of a DDR3 memory card is limited to 8 - 64 Gbps. Through-silicon via coupling of multiple memory dies and one or more logic dies can also advantageously provide high memory capacity by overcoming the scaling limitations in memory through stacking memory dies. For example, a processing system utilizing three-dimensional die stacking can achieve a memory capacity of 128 GB. In contrast, a DDR3 memory card is limited to 16 GB. Similarly, a processing system according to some aspects of the present disclosure can also advantageously provide a small form factor. For example, for a processing system utilizing through-silicon via coupling in a 3D memory die stack, a package size of 42 mm 2 can be achieved. In contrast, a 3D memory stack utilizing wire bonding has a package size of 117 mm 2 . A processing system according to some aspects of the present disclosure can also advantageously provide improved power efficiency. For example, a processing system utilizing through-silicon vias can consume 3.3 watts (W) when operating at a speed of 128 GB / s. In contrast, a 3D memory stack utilizing wire bonding can consume 6.4 W. Some aspects of the present disclosure also advantageously support higher memory capacity and inter-die communication.
[0033] For purposes of illustration and description, the foregoing description of specific embodiments of the present disclosure has been presented. They are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed, and many modifications and variations are apparent in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of the present disclosure and its practical application, thereby enabling others skilled in the art to best utilize the present disclosure and various embodiments with various modifications suitable for the particular purposes contemplated. The scope of the invention is intended to be defined by the appended claims and their equivalents.
Claims
1. A processing system, wherein, The processing system includes: a plurality of processor dies, each processor die including a plurality of memory dies coupled to a corresponding logic die; an input / output module; and an interposer that couples each processor die to one or more other processor dies among the plurality of processor dies, couples the plurality of processor chips to the input / output module, and couples the input / output module to a plurality of external contacts of the processing system; wherein the plurality of memory dies and the corresponding logic die of each processor die are arranged in a three-dimensional stack.
2. The processing system according to claim 1, wherein The plurality of processor dies and the input / output module are coupled together through the interposer in a mesh topology.
3. The processing system according to claim 1, wherein Each processor die is directly coupled to the input / output module through the interposer.
4. The processing system according to claim 1, wherein, It further includes a substrate that couples the plurality of processor dies, the input / output module, and the interposer to the external contacts of the processing system.
5. The processing system according to claim 4, wherein The interposer is coupled to the substrate through a fine-pitch ball grid array.
6. The processing system according to claim 4, wherein, The external contacts of the processing system include a fine-pitch ball grid array arranged on the substrate.
7. The processing system according to claim 1, wherein, The plurality of memory dies and the corresponding logic die in each processor die are coupled together through through-silicon vias.
8. The processing system according to claim 7, wherein, An external power potential is coupled to the plurality of memory dies in each processor die through through-silicon vias.
9. The processing system according to claim 1, wherein, Through a combination of through-silicon vias and a microbump array, the plurality of memory dies and the corresponding logic die in each processor die are coupled together.
10. The processing system according to claim 9, wherein, An external power potential is coupled to the plurality of memory dies in each processor die through through-silicon vias.
11. The processing system according to claim 1, wherein, The corresponding logic die of each processor die is coupled to the interposer through a microbump array.
12. The processing system according to claim 1, wherein, The plurality of processor dies include a plurality of near-memory processor dies.
13. The processing system according to claim 1, wherein, The plurality of processor dies include a plurality of processors and near-memory structures.
14. The processing system according to claim 1, wherein, The plurality of processor dies, the input / output module, and the interposer are a system-in-package.
15. The processing system according to claim 1, wherein, The input / output module includes an input / output die.
16. A processing system, wherein, The processing system includes: a plurality of processor dies, each processor die including a plurality of memory dies coupled to a corresponding logic die; and an active interposer that includes input / output circuitry, wherein the active interposer couples each processor die to one or more other processor dies among the plurality of processor dies, couples the plurality of processor dies to the input / output circuitry, and couples the input / output circuitry to a plurality of external contacts of the system; wherein the plurality of memory dies and the corresponding logic die of each processor die are arranged in a three-dimensional stack.
17. The processing system according to claim 16, wherein, The plurality of processor dies and the input / output circuitry of the active interposer are coupled together through the active interposer in a mesh topology.
18. The processing system according to claim 16, wherein, Each processor die is directly coupled to the input / output circuitry of the active interposer.
19. The processing system according to claim 16, further comprising a substrate that couples the plurality of processor dies and the active interposer to the external contacts of the processing system.
20. The processing system according to claim 19, wherein, The active interposer is coupled to the substrate through a fine-pitch ball grid array.
21. The processing system according to claim 19, wherein, The external contacts of the processing system include a fine-pitch ball grid array arranged on the substrate.
22. The processing system according to claim 16, wherein, The multiple memory dies and the corresponding logic dies in each processor die are coupled together through through-silicon vias.
23. The processing system according to claim 22, wherein, An external power supply potential is coupled to the multiple memory dies in each processor die through through-silicon vias.
24. The processing system according to claim 16, wherein Through a combination of through-silicon vias and micro-bump arrays, the multiple memory dies and the corresponding logic dies in each processor die are coupled together.
25. The processing system according to claim 24, wherein, An external power supply potential is coupled to the multiple memory dies in each processor die through through-silicon vias.
26. The processing system according to claim 16, wherein, The corresponding logic die of each processor die is coupled to the active interposer through a micro-bump array.
27. The processing system according to claim 16, wherein, The multiple processor dies include multiple near-memory processor dies.
28. The processing system according to claim 16, wherein, The multiple processor dies include multiple processors and near-memory structures.
29. The processing system according to claim 16, wherein, The multiple processor dies and the active interposer including the input / output module are a system-in-package.
Citation Information
Patent Citations
3D stacked integrated circuits having failure management
US10666264B1
Configuration of multi-die modules with through-silicon vias
US20190332561A1