Wafer-level heterogeneous integration for artificial intelligence accelerators and systems
By integrating diverse computing components on individual wafers and deploying them in immersion-cooled systems, the wafer-level heterogeneous integration method enhances computational density and addresses the limitations of traditional computing systems, particularly for AI applications.
Patent Information
- Application Number
- PCT/US2024/059364
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-07
- Filing Date
- 2024-12-10
- Publication Date
- 2025-06-26
AI Technical Summary
Traditional computing systems, such as GPU servers, face limitations in thermal efficiency, power efficiency, and bandwidth due to their PCB-integrated format, leading to sub-optimal computational densities that cannot keep pace with modern computational needs for AI training and inference.
The implementation of wafer-level heterogeneous integration, which involves integrating diverse computing components like GPUs, CPUs, switches, and memory on individual wafers, referred to as systems-on-a-wafer (SoWs), to create high-density computing systems. These SoWs are then disposed in distributed computing systems, such as immersion-cooled systems, to enhance computational density, thermal efficiency, and bandwidth.
This approach enables significantly enhanced computational density, improved space and thermal management, and increased bandwidth, effectively addressing the limitations of traditional computing systems and meeting the demands of modern AI applications.
Smart Images

Figure US2024059364_26062025_PF_FP_ABST
Abstract
Description
WAFER-LEVEL HETEROGENEOUS INTEGRATION FOR ARTIFICIALINTELLIGENCE ACCELERATORS AND SYSTEMSCROSS-REFERENCE TO RELATED APPLICATION S)
[0001] This application claims the priority benefit, under 35 U.S.C. 119(e), of U.S. Provisional Patent Application No. 63 / 611,410, filed December 18, 2023 and entitled “Wafer-Level Heterogeneous Integration for Al Accelerators and Systems,” U.S. Provisional Patent Application No. 63 / 643,728, filed May 7, 2024 and entitled “Wafer-Level Heterogeneous Integration for Al Accelerators and Systems,” U.S. Provisional Patent Application No. 63 / 549,123, filed February 2, 2024 and entitled “Universal Grid for Wafer Level Heterogeneous Integration,” and U.S. Provisional Patent Application No. 63 / 556,770, filed February 22, 2024 and entitled “Wafer Level Integration for System on Wafer with Optical Engine,” each of which are incorporated herein by reference in their entirety for all purposes.BACKGROUND
[0002] As feature sizes and transistor sizes have decreased for computing hardware such as integrated circuits (ICs) including chips and semiconductor dies, the amount of heat generated by a single chip, such as a microprocessor, has increased. Computing hardware that has traditionally been air cooled has evolved to levels of power consumption requiring more heat dissipation than can be provided by air alone. In some cases, immersion cooling of ICs in a tank containing a coolant liquid is employed to maintain ICs at appropriate operating temperatures.
[0003] One type of immersion cooling is two-phase immersion cooling, in which heat from a semiconductor die is high enough to boil the coolant liquid. The boiling creates a coolant-liquid vapor in the tank, which is condensed by cooling coils back to liquid form. Heat from the semiconductor dies can then be sunk into the liquid-to-gas and gas-to-liquid phase transitions of the coolant liquid with the result that the semiconductor dies are kept at an acceptable temperature.
[0004] This improved heat removal capability has led to new computing hardware architectures to increase computational density. However, this increased computational density has shifted computational and data transfer bottlenecks from the realm of IC processing power to the realmof memory latency and bandwidth. Traditional memory architectures have failed to keep pace with increased computational density and power efficiency requirements, and a new solution is needed.SUMMARY
[0005] Traditional computing systems such as graphics processing unit (GPU) servers are often embodied as varied computing components disposed on and interconnected through a printed circuit board (PCB). This PCB -integrated format imposes limitations on thermal efficiency, power efficiency, and bandwidth availability, and may result in a relatively large physical footprint (e.g., a server may have twice the overall area as the constituent components combined with the PCB included). This large footprint and sub-optimal efficiency and bandwidth result in computational densities that are being rapidly outpaced by modem computational needs for applications such as artificial intelligence (Al) training and inference.
[0006] The present technology provides methods and systems for wafer-level integration of heterogeneous computing components (e.g., GPUs, central processing units (CPUs), switches, network interface cards (NICs), memory, etc.) on individual wafers, also referred to as systems- on-a-wafer (SoWs). These SoWs may be disposed in distributed computing systems such as immersion cooled computing systems, and such computing systems may include SoWs with plurality of different arrangements and functionalities. For example, an exemplary computing system utilizing the present technology may include 80 GPU SoWs (each having 128 or more GPUs), 16 SoWs functioning as CPU head nodes for task and computational load distribution to the GPU SoWs, and 4 SoWs configured for networking and data transport communicatively coupled to the 80 GPU SoWs and 16 CPU head node SoWs.
[0007] The present technology may enable greatly enhanced computational density through improved packaging (e.g., components may be placed more closely together than on a PCB), space efficiency, waste heat disposal efficiency, and bandwidth.
[0008] In some aspects, the techniques described herein relate to a system-on-a-wafer (SoW), the SoW including: a substrate; a first computing component disposed on the substrate and a second computing component disposed on the substrate; and a plurality of interconnects disposed in the substrate in a regular pattern and communicatively coupling the first computing component and the second computing component.
[0009] In some aspects, the techniques described herein relate to a system, further including: a plurality of third computing components, analogous to the first computing component, disposed on the substrate and communicatively coupled to at least one another through at least some of the plurality of interconnects; a plurality of fourth computing components, analogous to the second computing component, disposed on the substrate and communicatively coupled to at least one another through at least some of the plurality of interconnects; wherein: the first computing component and the plurality of third computing components are disposed on the substrate according to the regular pattern; and the second computing component and the plurality of fourth computing components are disposed on the substrate according to the regular pattern.
[0010] In some aspects, the techniques described herein relate to a system, wherein the first computing component and the plurality of third computing components include graphics processing units (GPUs).
[0011] In some aspects, the techniques described herein relate to a system, further including: a chip connector communicatively coupled to the first computing component and the second computing component and configured to transfer data between at least the first computing component, the second computing component, and a device external to the SoW.
[0012] In some aspects, the techniques described herein relate to a system, wherein the device external to the SoW includes a second SoW.
[0013] In some aspects, the techniques described herein relate to a system, wherein the first computing component and the second computing component are analogous.
[0014] In some aspects, the techniques described herein relate to a system, further including: a power module electrically coupled to the first computing component and the second computing component; wherein the power module is configured to provide power to the first computing component and the second computing component.
[0015] In some aspects, the techniques described herein relate to a system, wherein the second computing component includes at least one of a network interface card (NIC), a central processing unit (CPU), a switch, an optical engine, or a memory module.
[0016] In some aspects, the techniques described herein relate to a system, wherein the substrate includes a semiconductor wafer.
[0017] In some aspects, the techniques described herein relate to a system, wherein the semiconductor wafer has a substantially circular shape.
[0018] In some aspects, the techniques described herein relate to a system, wherein the semiconductor wafer has a substantially rectilinear shape.
[0019] In some aspects, the techniques described herein relate to a server including: a plurality of systems-on-a-wafer (SoWs), each SoW including: a semiconductor wafer; a plurality of chips mounted on the semiconductor wafer; a power module; one or more chip connectors; and a wiring layer that electrically couples each chip to the power module and to at least one of the one or more chip connectors; and a plurality of inter-wafer connectors, each inter-wafer connector electrically coupling at least one of the one or more chip connectors of one SoW to at least one of the one or more chip connectors of another SoW.
[0020] In some aspects, the techniques described herein relate to a server, further including an inter-SoW connector that electrically connects at least a first chip connector of a first SoW of the plurality of SoWs to at least a second chip connector of a second SoW of the plurality of SoWs.
[0021] In some aspects, the techniques described herein relate to a server, further including a plurality of inter-SoW connectors, each inter-SoW connector electrically connecting a respective chip connector of a respective SoW to another respective chip connector of another respective SoW.
[0022] In some aspects, the techniques described herein relate to a server, wherein each SoW has the same configuration as each other SoW.
[0023] In some aspects, the techniques described herein relate to a server, wherein each SoW of the plurality of SoWs has a different configuration than each other SoW of the plurality of SoWs.
[0024] In some aspects, the techniques described herein relate to a server, wherein the plurality of chips in each SoW includes at least one of: logic chips, switches, memory modules, network interface controllers (NICs), or input / output (I / O) controllers.
[0025] In some aspects, the techniques described herein relate to a server, wherein the plurality of chips in each SoW includes the logic chips, the switches, the memory modules, the NICs, and the I / O controllers.
[0026] In some aspects, the techniques described herein relate to a server, wherein the logic chips include at least one of: graphics processing units, central processing units, data processing units, or tensor processing units.
[0027] In some aspects, the techniques described herein relate to a server, wherein the plurality of chips in each SoW are selected and arranged to form an artificial intelligence training server.
[0028] In some aspects, the techniques described herein relate to a server, wherein the plurality of chips in each SoW are selected and arranged to form an artificial intelligence inference server.
[0029] In some aspects, the techniques described herein relate to a server, wherein the wiring layer includes a redistribution layer that includes electrical wiring electrically coupled to each of: the chips, to the power module, and to the chip connectors.
[0030] In some aspects, the techniques described herein relate to a server, wherein the wiring layer forms an integrated fan-out packaging connection.
[0031] In some aspects, the techniques described herein relate to a server, wherein the wiring layer includes: a first redistribution layer that includes first electrical wiring that is electrically coupled to the power module and to the chip connectors; a second redistribution layer that includes first electrical wiring electrically coupled to the chips; and a plurality of local silicon interconnects that is electrically coupled to the first and second electrical wiring.
[0032] In some aspects, the techniques described herein relate to a server, further including at least one thermal module in direct physical contact with and / or in thermal communication with the SoWs.
[0033] In some aspects, the techniques described herein relate to a server, wherein the at least one thermal module includes an immersion cooler.
[0034] In some aspects, the techniques described herein relate to a high performance computing system on a wafer (SoW) including: a wafer-scale chip integration unit having two opposing faces and including a plurality of semiconductor computing chips and respective interconnects; a thermal module coupled to a first face of said wafer-scale chip integration unit; and an optical engine coupled to a second face of said wafer-scale chip integration unit.
[0035] In some aspects, the techniques described herein relate to a system, wherein said waferscale chip integration unit includes said plurality of semiconductor computing chips in a chip layer of said unit, and the respective interconnects include local silicon interconnects (LSI) ina redistribution layer (RDL) of said unit providing conduction pathways to said semiconductor computing chips.
[0036] In some aspects, the techniques described herein relate to a system, wherein said semiconductor computing chips include any of: central processing units (CPU), general processing units (GPU), and input / output (IO) units.
[0037] In some aspects, the techniques described herein relate to a system, wherein said thermal module includes a two-phase immersion cooling system in thermal communication with said wafer-scale chip integration unit.
[0038] In some aspects, the techniques described herein relate to a system, wherein said optical engine includes any of: a linear drive pluggable optics (LPO) engine, a co-packaged optics (CPO) engine, and a near packaged optics (NPO) engine.
[0039] In some aspects, the techniques described herein relate to a system, wherein said optical engine is coupled to said wafer-scale chip integration unit through any of: a hybrid bond layer, a bump bond layer, and a plug-in or push socket.
[0040] In some aspects, the techniques described herein relate to a system, further including an integrated connector interface disposed on and in data communication with said wafer-scale chip integration unit.
[0041] In some aspects, the techniques described herein relate to a system, further including a power module in electrical communication with said wafer-scale chip integration unit.
[0042] In some aspects, the techniques described herein relate to a system, further including a data switch unit in data communication with said wafer-scale chip integration unit.
[0043] In some aspects, the techniques described herein relate to a system, wherein said waferscale chip integration unit includes a heterogeneous unit having a plurality of different types of semiconductor computing chips.
[0044] In some aspects, the techniques described herein relate to a high performance computing system on a wafer (SoW) including: a wafer substrate; a plurality of computing hardware chips (chips) disposed on a surface of said wafer substrate and defining a spatial arrangement between and among said chips; and a plurality of electrical interconnects placing said chips in electrical communication according to a layout of said chips on the surface of said wafer substrate; wherein : said electrical interconnects include respective positions with respect to one another and with respect to the surface of said wafer substrate; said electricalinterconnects are disposed in a regular geometric arrangement and have a spatial density on a portion of the surface of said wafer substrate; and said spatial density is greater than a minimum required density to support said electrical communication according to said layout of said chips.
[0045] In some aspects, the techniques described herein relate to a system, wherein said regular geometric arrangement includes a regularly spaced grid of rows and columns on the surface of said wafer.
[0046] In some aspects, the techniques described herein relate to a system, wherein said spatial density is a first spatial density and said portion of the surface is a first portion of the surface, and the system further includes a second portion of said surface on which a second spatial density of interconnects is disposed.
[0047] In some aspects, the techniques described herein relate to a system, wherein said plurality of computing hardware chips includes a plurality of processing chips including any of: central processing units (CPU) and general processing units (GPU).
[0048] In some aspects, the techniques described herein relate to a system, wherein said plurality of computing hardware chips includes a plurality of data storage units including any of: DRAM, SRAM or other memory storage hardware chips.
[0049] In some aspects, the techniques described herein relate to a system, wherein said electrical interconnects include redistribution layer (RDL) and local silicon interconnect (LSI) interconnects.
[0050] In some aspects, the techniques described herein relate to a system, wherein said wafer substrate includes a carrier silicon layer.
[0051] In some aspects, the techniques described herein relate to a system, wherein said chips are configured and arranged in a layer substantially parallel to said wafer substrate and wherein said chips lie in spatial separation separated by a molding material therebetween.
[0052] In some aspects, the techniques described herein relate to a system, wherein said electrical interconnects include a conducting metal material disposed in a dielectric layer of said SoW substantially parallel to said wafer substrate.
[0053] In some aspects, the techniques described herein relate to a system, further including a plurality of stacked integrated circuits (3DIC) on said computing hardware chips.
[0054] In some aspects, the techniques described herein relate to a system, further including an immersion cooling assembly in thermal communication with said computing hardware chips and operational to remove thermal energy generated by said chips.
[0055] In some aspects, the techniques described herein relate to a system, wherein said first spatial density in said first portion of said surface is greater than said second spatial density in said second portion of said surface.
[0056] In some aspects, the techniques described herein relate to a system, wherein said arrangement and said spatial density are such that said wafer substrate can support a plurality of computing hardware chips thereon and wherein only a subset of said plurality of electrical interconnects are required for interconnection of said chips while other electrical interconnects are not required.
[0057] All combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are part of the inventive subject matter disclosed herein. The terminology used herein that also may appear in any disclosure incorporated by reference should be accorded a meaning most consistent with the particular concepts disclosed herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The skilled artisan will understand that the drawings primarily are for illustrative purposes and are not intended to limit the scope of the inventive subject matter described herein. The drawings are not necessarily to scale; in some instances, various aspects of the inventive subject matter disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like reference characters generally refer to like features (e.g., functionally similar and / or structurally similar elements).
[0059] FIG. l is a top view of a system-on-a-wafer (SoW) according to an embodiment.
[0060] FIG. 2 is a cross section of the SoW illustrated in FIG. 1 according to an embodiment
[0061] FIG. 3 is a cross section of the SoW illustrated in FIG. 1 according to another embodiment.
[0062] FIG. 4 is a cross section of a structure, according to an embodiment, formed during a first step of a chip-last manufacturing process for manufacturing the embodiment illustrated in FIG. 3.
[0063] FIG. 5 is a cross section of a structure, according to an embodiment, formed during a second step of a chip-last manufacturing process for manufacturing the embodiment illustrated in FIG. 3.
[0064] FIG. 6 is a cross section of a structure, according to an embodiment, formed during a third step of a chip-last manufacturing process for manufacturing the embodiment illustrated in FIG. 3.
[0065] FIG. 7 is a cross section of a structure, according to an embodiment, formed during a fourth step of a chip-last manufacturing process for manufacturing the embodiment illustrated in FIG. 3.
[0066] FIG. 8 is a top view of an SoW according to another embodiment.
[0067] FIG. 9 is a top view of an SoW according to another embodiment.
[0068] FIG. 10 is an isometric view of a server according to an embodiment.
[0069] FIG. 11 is an isometric view of a server according to another embodiment.
[0070] FIG. 12 is an isometric view of a server according to another embodiment.
[0071] FIG. 13 is a conceptual diagram that illustrates how the functionality of an Al training server can be divided among several SoWs.
[0072] FIG. 14 is a cross section of a server according to an embodiment.
[0073] FIG. 15 is a conceptual diagram of an immersion cooler according to an embodiment.
[0074] FIG. 16 illustrates a comparison of areas between a circular wafer and a panel wafer for use as SoWs in accordance with the present technology.
[0075] FIG. 17 illustrates a graph comparing number of chips per wafer vs. chip area in mm2.
[0076] FIG. 18 illustrates a comparison of chip density between a circular wafer SoW and a panel wafer SoW.
[0077] FIG. 19 illustrates an exemplary computing system on a wafer.
[0078] FIG. 20 illustrates an exemplary computing system on a wafer.
[0079] FIG. 21 illustrates a cross section of an exemplary SoW and method of making the same.
[0080] FIG. 22 illustrates alternative RDL / LSI configurations for supporting a SoW.
[0081] FIG. 23 illustrates several uniform electrical interconnect density embodiments on a wafer.
[0082] FIG. 24 illustrates several different arrangements of electrical interconnect densities on a wafer.
[0083] FIG. 25 illustrates a cross section of an exemplary SoW including a wafer-scale chip integration with RDL / LSI layers.
[0084] FIG. 26 illustrates a multi-layered SoW with an optical engine (OE) directly coupled to a wafer-scale chip integration using a bump bond layer.
[0085] FIG. 27 illustrates a multi-layered SoW with an optical engine (OE) coupled to a waferscale chip integration using a socket.
[0086] FIG. 28 depicts aspects of an immersion cooling system for dissipating heat from one or more heat-generating components such as semiconductor die packages via immersion cooling.DETAILED DESCRIPTION
[0087] Following below are more detailed descriptions of various concepts related to, and implementations of, a SoW and servers including SoWs. It should be appreciated that various concepts introduced above and discussed in greater detail below may be implemented in multiple ways. Examples of specific implementations and applications are provided primarily for illustrative purposes so as to enable those skilled in the art to practice the implementations and alternatives apparent to those skilled in the art.
[0088] The figures and example implementations described below are not meant to limit the scope of the present implementations to a single embodiment. Other implementations are possible by way of interchange of some or all of the described or illustrated elements. Moreover, where certain elements of the disclosed example implementations may be partially or fully implemented using known components, in some instances only those portions of such known components that are necessary for an understanding of the present implementations aredescribed, and detailed descriptions of other portions of such known components are omitted so as not to obscure the present implementations.
[0089] FIG. 1 is a top view of an SoW 10 according to an embodiment. The SoW 10 includes a semiconductor wafer 100, a plurality of chips 110, one or more power modules 120, and one or more chip connectors 130.
[0090] The semiconductor wafer 100 comprises a semiconducting material such as silicon, gallium arsenide, sapphire, silicon carbine, indium phosphide, gallium nitride, germanium, and / or another semiconducting material.
[0091] The chips 110 are mounted on or above the semiconductor wafer 100. The chips 110 can be multi-functional and can include logic chips, switches, memory modules (e.g., high- bandwidth memory (HBM)), network interface controllers (NICs), input / output (I / O) controllers, and / or other chips. Examples of logic chips include graphics processing units (GPUs), central processing units (CPUs), data processing units (DPUs), tensor processing units (TPUs), systems on a chip (SoCs), and the like. An example of a switch chip is disclosed in Provisional Application No. 63 / 123,476, titled “Compute Express Link Switch With Integrated Optical Engine,” filed on October 31, 2023, which is hereby incorporated by reference in its entirety for all purposes.
[0092] A memory module may be an IC configured to store data and may include a dynamic random access memory (DRAM) module, a static random access memory (SRAM) module, a flash memory module, a solid-state drive (SSD), a non-volatile random access memory (NVRAM) module, a read-only memory (ROM) module (such as a floating-gate ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), one-time programmable ROM (OTPROM), or the like), or any suitable type of memory module.
[0093] The chips 110 can be selected in any combination to provide an overall function or architecture for the SoW 10. For example, the chips 110 can be selected such that the SoW 10 can function as server or a portion of a server. Examples of servers that the SoW 10 can function as can include an artificial intelligence (Al) training server, an Al inference server, a web server, a database server, an email server, a web proxy server, a domain name server (DNS), an application server, and / or another server.
[0094] The chips 110 are electrically coupled to one or more modules or structures that are mounted on or above the semiconductor wafer 100 and that are electrically coupled to the chips110. The modules / structures can include a power module 120 and / or a chip connector 130. The power module 120 can include a voltage converter that can modulate an external supply voltage to a chip supply voltage that is at an appropriate level to power some or all of the chips 110. The chip connector 130 provides an electrical connection point to an external electrical component and / or to another SoW. One or more of the modules can be located directly above one or more of the chips 110 but are not illustrated in FIG. 1 for illustration purposes only so as to not obscure the chips 110.
[0095] FIG. 2 illustrates a cross section of SoW 10 through plane 20 in FIG 1 according to an embodiment. In this embodiment, a redistribution layer 200 may be located between the chips 110 and the modules mounted on or above the semiconductor wafer 100, such as the power module 120 and the chip connector 130. The redistribution layer 200 includes electrical wiring 210 that electrically connects each chip 110 to the power module 120, to the chip connector 130, and / or to another module mounted on or above the semiconductor wafer 100. The electrical wiring 210 can include one or more levels including a multilevel wiring structure. The electrical wiring 210 can include or define an Integrated Fan-Out (InFO) packaging connection. The redistribution layer 200 includes an electrical insulator 220, such as silicon dioxide, that electrically isolates the wires and levels of the electrical wiring 210. The redistribution layer 200 can alternately be referred to as a wiring layer.
[0096] The structure illustrated in FIG. 2 can be manufactured by mounting the chips 110 on (e.g., directly or indirectly) the semiconductor wafer 100. Next, an epoxy and / or reflow material 230 may be formed on and / or between the chips 110. The epoxy / reflow material 230 can be planarized. The redistribution layer 200 may be formed on the epoxy / reflow material 230 and / or on the chips 110. The electrical wiring 210 can be formed by photolithography and etching of the electrical insulator 220 to form patterns and metal deposition into the patterns. The power module 120, the chip connector 130, and / or other modules are mounted on redistribution layer 200 and electrically coupled to the appropriate leads or terminals of the electrical wiring 210. This manufacturing process can be referred to as chip-first manufacturing process.
[0097] FIG. 3 is a cross section of SoW 10 through plane 20 in FIG 1 according to another embodiment. In this embodiment, first and second redistribution layers 301, 302 are located between the chips 110 and the modules mounted on or above the semiconductor wafer 100, such as the power module 120 and the chip connector 130. A plurality of local silicon interconnects (LSIs) 310 are located between the first and second redistribution layers 301,302. The LSIs 310 are electrically coupled to the chips 110 through first electrical wiring 311 in the first redistribution layer 301. The LSIs 310 are electrically coupled to the power module 120, to the chip connector 130, and / or to another module mounted on or above the semiconductor wafer 100 through second electrical wiring 312 in the second redistribution layer 302. Thus, the chips 110 are electrically coupled to the to the power module 120, to the chip connector 130, and / or to another module mounted on or above the semiconductor wafer 100 through the first electrical wiring 311, the second electrical wiring 312, and the LSIs 310. The first and second redistribution layers 301, 302 also include a respective electrical insulator that can be the same as electrical insulator 220. The first and second redistribution layers 301, 302 can be formed in the same or similar manner as the redistribution layer 200. As in the embodiment of FIG. 2, an epoxy / reflow material 230 is formed on and / or between the chips 110.
[0098] The embodiment illustrated in FIG. 3 can be formed in a chip-last manufacturing process. The chip-last manufacturing process can also be referred as a chip-on-wafer (CoW) manufacturing process.
[0099] FIG. 4 is a cross section of a structure 40, according to an embodiment, formed during a first step of a chip-last manufacturing process for manufacturing the embodiment illustrated in FIG. 3. The structure 40 includes a sacrificial semiconductor wafer 400, the first and second redistribution layers 301, 302, and the LSI 310. The sacrificial semiconductor wafer 400 can be formed of the same material or a different material than the semiconductor wafer 100. The second redistribution layer 302 may be deposited and formed on the sacrificial semiconductor wafer 400. Next, the LSIs 310 are deposited and / or formed on the second redistribution layer 302. The first redistribution layer 301 may be deposited and formed on the LSIs 310. The first and second redistribution layers 301, 302 can be formed in the same manner as discussed above with respect to the redistribution layer 200. Traces 303 may provide communicative connection through redistribution layers 301 and 302.
[0100] FIG. 5 is a cross section of a structure 50, according to an embodiment, formed during a second step of a chip-last manufacturing process for manufacturing the embodiment illustrated in FIG. 3. Structure 50 may be analogous to or the same as structure 40 except that in structure 50 the chips 110 are mounted and electrically coupled to first electrical wiring 311 of the first redistribution layer 301.
[0101] FIG. 6 is a cross section of a structure 60, according to an embodiment, formed during a third step of a chip-last manufacturing process for manufacturing the embodiment illustrated in FIG. 3. Structure 60 may be analogous to or the same as structure 50 except that in structure 60 the epoxy / reflow material 230 may be formed on and / or between the chips 110 and the semiconductor wafer may be mounted on the chips 110. The epoxy / reflow material 230 can be formed before or after the semiconductor wafer is mounted on the chips 110.
[0102] FIG. 7 is a cross section of a structure 70, according to an embodiment, formed during a fourth step of a chip-last manufacturing process for manufacturing the embodiment illustrated in FIG. 3. Structure 70 may be analogous to or the same as structure 60 except that structure 70 is flipped over compared to structure 60 and the sacrificial semiconductor wafer 400 is removed.
[0103] A fifth step is to mount the power module 120, the chip connector 130, and / or another module onto the second redistribution layer 302. The power module 120, the chip connector 130, and / or another module are electrically coupled to the second electrical wiring 312 in the second redistribution layer 302.
[0104] FIG. 8 is a top view of a SoW 80 according to another embodiment. SoW 80 may be analogous to or the same as SoW 10 except that the chips 110 in SoW 80 include CPUs, GPUs, and switches. This configuration / architecture can function as an Al inference server 800 or a portion of an Al inference server 800. In a specific example, the SoW 80 can include 128 GPU chips, 8 CPU chips, and 4 switch chips. The configuration in this specific example can replace (e.g., function as) a conventional Peripheral Component Interconnect Express (PCIE) card for a conventional Al inference server, where the PCIE card includes the same number of GPU, CPU, and switch chips (i.e., 128 GPU chips, 8 CPU chips, and 4 switch chips). In some embodiments, the SoW 80 can include 256 GPU chips, 16 CPU chips, and 8 switch chips in which case the SoW 80 can replace (e.g., function as) two or more PCIE cards for a conventional Al inference server. Additionally or alternatively, the SoW 80 can include memory (e.g., HBM) chips, NIC chips, and / or other chips. Additional or fewer numbers of each type of chip 110 can be included in other embodiments.
[0105] A technical advantage of the configuration / architecture of SoW 80 is that the physical size of the SoW 80 may for example replace six, eight, or more PCIE cards and / or that additional components (e.g., chips) can be included in the SoW 80 compared to a conventional PCIE card.
[0106] FIG. 9 is a top view of a SoW 90 according to another embodiment. SoW 90 may be analogous to or the same as SoW 10 except that the chips 110 in SoW 90 include CPUs, GPUs, switches, NICs, I / Os, and memory (e.g., HBM). The I / O chips and the memory modules are represented as 921 and 922, respectively. This configuration / architecture can function as an Al training server 900 or a portion of an Al training server 900. In a specific example, the SoW 90 can include 8 GPU chips, 2 CPU chips, 9 NIC chips, 2 switch chips, 18 memory modules, and 16 VO chips. The configuration in this specific example can replace (e.g., function as) a conventional Peripheral Component Interconnect Express (PCIE) card for a conventional Al training server, where the PCIE card includes the same number of GPU, CPU, and switch chips (i.e., 8 GPU chips, 2 CPU chips, 9 NIC chips, 2 switch chips, 18 memory modules, and 16 VO chips). In some embodiments, the SoW 90 can include 16 GPU chips, 4 CPU chips, 18 NIC chips, 4 switch chips, 16 memory modules, and 32 VO chips in which case the SoW 90 can replace (e.g., function as) two or more PCIE cards (for example, six, eight, or more PCIe cards) for a conventional Al training server. Additionally or alternatively, the SoW 80 can include memory (e.g., HBM) chips, NIC chips, and / or other chips. Additional or fewer numbers of each type of chip 110 can be included in other embodiments.
[0107] A technical advantage of the configuration / architecture of SoW 90 is that the memory modules 922 are located in close physical proximity (e.g., at the millimeter scale and / or at the micron scale) to the GPU chips, which improves processing speed and / or reduces latency.
[0108] FIG. 10 is an isometric view of a server 1000 according to an embodiment. The server 1000 may include a plurality of SoWs 1010 that are contained within a housing 1020. Example dimensions of the SoWs 1010 and of the housing 1020 are included, for example to compare with corresponding dimensions of a conventional server 1030. As illustrated, multiple SoWs 1010 can be placed laterally with respect to each other and / or can be stacked vertically within the housing 1020. In some embodiments, each SoW 1010 is configured in the same manner. In other embodiments, one, some, or all of the SoWs 1010 is / are different than other SoWs 1010. In a specific example, each SoW 1010 can be configured the same as and / or can include the same number and type of chips as in the conventional server 1030. Thus, the server 1000 can have multiple times the computing power as the conventional server 1000 while maintaining the same (or substantially the same) physical dimensions and footprint.
[0109] Each SoW 1010 can be independently electrically connected to a common circuit board (or another electrical lead / connection). Alternatively, some or all of the SoWs 1010 canbe electrically connected to one or more other SoWs 1010. Each SoW 1010 can be the same as or different than SoW 10 and / or SoW 90.
[0110] The server 1000 can be an Al server such as an Al training server or an Al inference server.
[0111] FIG. 11 is an isometric view of a server 1100 according to another embodiment. The server 1100 may be analogous to or the same as server 1000 except that in server 1100 some of the SoWs 1010 are electrically connected to each other through an inter-SoW (or interwafer) connector 1110. The inter-SoW connector 1110 is electrically connected to a respective chip connector 130 of the connected SoWs 1010. The inter-SoW connector 1110 can include wires, cables, and / or other electrical connections. The inter-SoW connector 1110 can be used to connect SoWs 1010 that are laterally and / or vertically displaced with respect to each other. The SoWs 1010 in server 1100 have the same configuration / architecture (e.g., the same numbers of chips, the same types of chips, and / or the same physical arrangement of chips).
[0112] The server 1100 can be an Al server such as an Al training server or an Al inference server.
[0113] FIG. 12 is an isometric view of a server 1200 according to another embodiment. The server 1200 may be analogous to or the same as server 1100 except that in server 1200 at least some of the SoWs 1010 have different configurations / architectures (e.g., different numbers of chips, different types of chips, and / or a different physical arrangement of chips).
[0114] The server 1200 can be an Al server such as an Al training server or an Al inference server.
[0115] FIG. 13 is a conceptual diagram that illustrates how the functionality of an Al training server 1300 can be divided among several SoWs 1310 (including SoW 1310a, SoW 1310b, SoW 1310c, and SoW 13 lOd). Some or all of the SoWs 1310 can be electrically connected to each other through one or more inter-SoW connectors, each of which can be the same as inter- SoW connector 1110. Each SoW 1310 can be the same as or different than SoW 10, SoW 90, and / or SoW 1010. In particular, FIG. 13 illustrates the flexibility afforded by the present technology, whereby a given group of components may be disposed in a variety of different combinations on different SoWs. For example, FIG. 13 shows an embodiment where all GPUs of Al training server 1300 are disposed on a single SoW 1310c, while in an alternative embodiment, the GPUs of Al training server 1300 are split among two SoWs, SoW 13 lOd (which includes a plurality of GPUs, a Cedar Module, and a CPU), and SoW 1310a (whichincludes a plurality of GPUs and a Cedar Module). This flexibility may enable improved hardware efficiency, provide cost savings by eliminating unnecessary components that might otherwise be included in off-the-shelf systems such as GPU servers or CPU head nodes, and give system designers more leeway to include or exclude components that may optimize a particular computing system based on given design criteria.
[0116] FIG. 14 is a cross section of a server 1400 according to an embodiment. Server 1400 includes one or more SoWs 1410 and one or more thermal modules 1420. The thermal module(s) 1420 is / are configured to reduce the heat produced by the SoW(s) 1410 during operation. The thermal module(s) 1420 can include fluid coolers (e.g., air and / or liquid coolers), cooling fins, heat exchangers, and / or other cooling units. Liquid coolers can include immersion coolers. The thermal module(s) 1420 can be in direct physical contact with the SoW(s) 1410. Additionally or alternatively, the thermal module(s) 1420 can be in thermal communication with the SoW(s) 1410. In an aspect, server 1400 may be disposed within an immersion cooling liquid as part of an immersion cooling system, e.g., immersion cooling system 2800 as depicted in FIG. 28.
[0117] Examples of thermal modules are disclosed in PCT Application No. PCT / US23 / 67058, titled “Electronic Package Construction For Immersion Cooling of Integrated Circuits,” filed on May 16, 2023 and / or U.S. Provisional Application No. US 63 / 597,586, titled “Systems for Thermal Management in Three-Dimensional Integrated Circuits,” filed on November 9, 2023, which are hereby incorporated by reference in their entirety for all purposes.
[0118] Server 1400 may be analogous to or the same as any exemplary server disclosed herein, such as server 1000, server 1100, and / or server 1200. The server 1400 can be an Al server such as an Al training server or an Al inference server.
[0119] FIG. 15 is a conceptual diagram of an immersion cooling system 1500 in which a plurality of SoWs 1510 (which may be analogous to or the same as SoWs 1310 or other SoWs described herein) are disposed in an immersion cooling container 1520. In particular, FIG. 15 shows a computational density enabled by the present technology. Because computing components may be placed closer together than on traditional multi-chip module (MCM) printed circuit boards (PCBs), an overall computing density of
[0120] FIG. 16 illustrates a comparison between a circular wafer 1610 that may be used as an SoW vs. a panel wafer 1620 that may be used as an SoW in accordance with the presenttechnology. An SoW in accordance with the present technology may be any suitable shape, including circular, ovular, square, rectangular, irregularly shaped, triangular, curved, rounded, or the like. In particular, a panel wafer 1620 may have the benefit of a larger functional area 1622 (e.g., surface area on which computing hardware including integrated circuits, chips, memory modules, chip connectors, power modules, or other components may be disposed for a given maximum dimension) compared to a functional area 1612 of circular wafer 1610.
[0121] For example, each of circular wafer 1610 and panel wafer 1620 may have maximum principal dimensions (e.g., horizontal and vertical extents) of 300 mm. However, circular wafer 1610 may have a functional area 1612 of 210 mm by 210 mm (i.e., 44,100 mm2) while panel wafer 1620 may have a functional area 1622 of 270 mm by 270 mm (i.e., 72,900 mm2) - an increase of 65% over the functional area 1612 of circular wafer 1610. This larger functional area 1622 may allow for a higher chip density, computing performance, bandwidth, and functionality on a given wafer as well as increasing chip (and associated computational) density in a computing system using panel wafer SoWs instead of circular wafer SoWs.
[0122] Equivalently, for a fixed functional area (e.g., a functional area of about 210 mm by about 210 mm, or about 44,100 mm2), a panel wafer 1620 may have a smaller overall footprint than circular wafer 1610 (e.g., a maximum principal dimension of panel wafer 1620 may be smaller than a maximum principal dimension of circular wafer 1610). For example, as illustrated in FIG. 16, circular wafer 1610 may have a functional area 1612 of about 210 mm by about 210 mm, or about 44,100 mm2and a maximum dimension of 300 mm. Assuming the same margins between an edge of panel wafer 1620 and the functional area 1622 of 15 mm per side, a panel wafer 1620 having the same functional area may require a maximum dimension of only 240 mm. This additional space may be utilized for additional computing components in a server, to reduce an overall footprint of a server, reduce a latency between SoWs in a computing system, or the like.
[0123] Further, semiconductor dies, chips, and processors are typically square in profile, which means that a greater percentage of the total area of a square panel wafer may be utilized to include chips than a corresponding circular wafer SoW of the same maximum principal dimensions. For example, FIG. 16 shows circular wafer 1610 having a diameter of 300 mm and panel wafer 1620 having a length and width of 300 mm. The total area of circular wafer 1610 is itx1502= 70,686 mm2, while the total area of panel wafer 1620 is 3002= 90,000 mm2. Accordingly, the functional area 1612 of circular wafer 1610 is 44,100 mm2and the usable percentage of circular wafer 1610 is 44,100 70,686 = 62.4%. Likewise, the functional area1622 of panel wafer 1620 is 2702= 72,900, and the usable percentage of panel wafer 1620 is 72,900 mm290,000 mm2= 81%. This increase in functional area for an exemplary panel wafer 1620 means better overall system performance, improved computational density, higher computational throughput, greater potential revenues from selling computing power, and various additional improvements.
[0124] FIG. 17 illustrates a graph 1700 comparing the number of chips per wafer vs. individual chip area in mm2for a panel wafer SoW and a circular wafer SoW of the same maximum dimension. Graph 1700 shows for every chip area between 100 mm2and 800 mm2(and assuming equal spacing and analogous arrangement of chips on the respective wafers), a panel wafer SoW can include more chips than a circular wafer SoW of the same maximum dimension. For example, graph 1700 shows that for a chip area of 200 mm2, an exemplary circular wafer SoW may include 220 chips, while an exemplary panel wafer SoW may include 364 chips, a 65% increase for the panel wafer SoW over the circular wafer SoW.
[0125] FIG. 18 illustrates a comparison of chip density between a circular wafer SoW 1800a vs. a panel wafer SoW 1800b. Each SoW may have the same maximum dimension (e.g., about 300 mm) and same spacing between analogous components (e.g., a same distance between adjacent chips). Additionally, each SoW may include analogous or substantially identical components (e.g., chips, memory modules, chip connectors, and chip interconnects). However, in an example, panel wafer SoW 1800b may include 96 chips 1810b, whereas circular wafer SoW 1800a may only include 60 chips 1810a, an increase of 60% for panel wafer SoW 1800b. Each additional chip 1810b may increase a computational performance (e.g., computational speed), reliability (by enabling a portion of chips 1810b to be kept in reserve if one or more other chips 1810b fail or malfunction), computational flexibility (e.g., allowing additional tasks to be run at a given time by panel wafer SoW 1800b as compared to circular wafer SoW 1800a), or the like.
[0126] Circular wafer SoW 1800a may include a plurality of chips 1810a (labelled C0-C59). Each chip 1810a may be communicatively coupled to one or more memory modules 1812a. Each chip of chips 1810a may additionally be communicatively coupled to one or more other chips of chips 1810a via one or more chip interconnects 1822a. A plurality of chips 1810a may be arranged into one or more groups (e.g., pairs of rows such as group C0-C19). Each group of chips 1810a may be communicatively coupled to one or more other groups via group interconnects 1820a, which may be particularly configured to transfer data generated by a plurality of chips with high bandwidth and low latency. For example, group interconnects1820a may provide the same data transfer speeds and latencies as one or more chip interconnects 1822a, but with a high enough bandwidth to support data transfers from a plurality of chips at the same time. For instance, group interconnects 1820a may provide enough bandwidth to support data transfers at or about the maximum for two or more, four or more, eight or more, ten or more, 15 or more, 20 or more, 25 or more, 50 or more, 100 or more, or any suitable number of chips at a time.
[0127] Circular wafer SoW 1800a may additionally include one or more memory modules 1812a, each of which is communicatively coupled to a respective chip of chips 1810a. Each of one or more memory modules 1812a may be analogous to other memory modules described herein.
[0128] Circular wafer SoW 1800a may additionally include one or more chip connectors 1830a, which may provide an electrical connection point to an external electrical component and / or to another SoW in a manner analogous to chip connector 130 or other suitable chip connector in accordance with the present technology. Chip connector 1830a may provide a low-latency connection having a bandwidth that exceeds that of group interconnects 1820a and / or one or more chip interconnects 1822a by a suitable margin, such as by a factor of about two, a factor of about four, a factor of about eight, a factor of about ten, a factor of about 15, a factor of about 20, a factor of about 25, a factor of about 40, a factor of about 50, a factor of about 75, a factor of about 100, a factor of about 500, or any suitable factor.
[0129] Panel wafer SoW 1800b may include a plurality of components analogous to those described with respect to circular wafer SoW 1800a. For example, panel wafer SoW 1800b may include a plurality of chips 1810b, a plurality of memory modules 1812b, a plurality of group interconnects 1820b, a plurality of chip interconnects 1822b, and one or more chip connectors 1830b. Panel wafer SoW 1800b may include an equal or greater number of analogous components to circular wafer SoW 1800a. For example, each group of chips of panel wafer SoW 1800b may include 24 chips as compared to 20 chips per group of circular wafer SoW 1800a, and panel wafer SoW 1800b may include four groups of 24 chips as compared to three groups of 20 chips on circular wafer SoW 1800a. The extra group in panel wafer SoW 1800b may necessitate one or more additional group interconnects 1820a as well as additional memory modules 1812b. One or more chip connectors 1830b may further have a higher bandwidth than one or more chip connectors 1830a to account for the larger amount of data being processed by additional chips C60-C95 of panel wafer SoW 1800b.
[0130] FIG. 19 illustrates an exemplary computing system on a wafer (SoW) 1900 or part thereof comprising a semiconductor (e.g., silicon) based wafer or substrate 1902 on or in which are disposed a plurality of computing system hardware components 1910 such as processing circuits, central processing units (CPU), graphics processing units (GPU), data storage units (DRAM or otherwise as suits a given implementation). In addition, other electrical and electronic connections, interface points, buses and data communication channels can be fabricated or installed in / and or on the SoW 1900 and may have respective conducting interface(s) for the SoW components 1910 with respect to substrate 1902.
[0131] In some examples, which are not limiting of the present disclosure, the components 1910 on substrate 1902 in SoW 1900 are arranged in rows 1912 and columns 1914, or a substantial number of components 1910 may be so arranged on the surface of substrate 1902. In an aspect, the present disclosure describes wafer integrations that improve the efficiency and flexibility of construction of such SoW systems, especially in the case of heterogeneous integrations and integrations involving various types of components 1910 on substrate 1902. The present disclosure and SoW 1900 may be specifically directed to certain high performance computing architectures such as those performing or supporting artificial intelligence (Al) integrated circuits (IC), stacked three-dimensional integrated circuits (3DIC) as described by the present applicant in this and other disclosures, which are hereby incorporated by reference.
[0132] FIG. 20 illustrates an exemplary computing system on a wafer SoW 2000. As before, the system may include a substrate 2002 or platform for SoW 2000. A plurality of varying (heterogeneous) components may be integrated onto substrate 2002 of SoW 2000. These components may for example include one or more of: CPUs 2010, GPUs 2020, network interface cards (NICs 2030), switch(es) 2040, electrical or data connections 2050, input-output (VO) interface(s) 2060, memory (e.g., DRAM) units and other components on bump and / or hybrid bonded connections such as HBM 2070.
[0133] It is noted that some such components on the SoW may have a same or similar dimension or size or footprint. But other components may be of different dimensions, sizes or footprints as shown illustratively in the figure.
[0134] FIG. 21 illustrates a cross section of an exemplary SoW 2100 and method of making the same 3100. The whole system is disposed on or fabricated on a carrier silicon substrate 2120. A redistribution layer (RDL) interposer 2122 is fabricated in a dielectric material (e.g.,silicon oxide SiO2) at block 3110. A plurality of local silicon interconnects (LSI) 2128 between the respective computing, memory and circuit components of the SoW 2100.
[0135] In an aspect, one or more integrated circuits or “chips” 2130 are placed at block 3112 on the RDL interposer. As mentioned earlier, chips 2130 may include one or more types of circuits including but not limited to computing circuitry, data storage circuitry, communications circuitry and other chips as called for by a particular SoW design. These chips 2130 may be bump and / or hybrid bonded to the portions of the SoW 2100 on which they are placed and on RDL interposer parts 2122, 2124, 2126. The LSI 2128 may provide signal, data or electrical interconnection functions among one or more components and chips in the SoW 2100.
[0136] A molding material 2142 or layer is disposed among chips 2130 and a second carrier silicon layer 2140 is provided on the chips and molding material as shown at block 3114. In this scenario, the system is sandwiched by carrier silicon substrates 2120 and 2140.
[0137] At block 3116, the carrier silicon substrate 2120 may be removed (e.g., by chemical or etching means) as shown. The resulting product after block 3116 is a SoW 2100 on a carrier Si in one example and upon which are a plurality of chips (e.g., CPUs, GPUs, DRAM memory and so on) interconnected by RDL and LSI as needed. It can be appreciated that in a heterogeneous system design said interconnections and RDL are arranged and configured to suit the chips 2130 which they support. Thus, the size, number and position and geometry of these chips 2130 would dictate the position and arrangement and configuration of said RDL and LSI.
[0138] FIG. 22 illustrates two possible configurations and / or arrangements “A” and “B” of RDL and LSI connection points in an SoW with heterogeneous integrations. An SoW is formed according to configuration “A” whereby RDL and corresponding LSI 2244 are disposed on substrate 2242 according to and only where these connections are required to support the chips and components 2212 of SoW 2202. As can be understood, the wafer and RDL / LSI configuration 2240 is generally limited to chip placements that are reached and supported by the specific placement of the RDL / LSI points 2244 or positions on substrate 2242. If the nature, geometry, number, or configuration of chips 2212 are changed, a corresponding change in the configuration of RDL / LSI points 2244 on substrate 2242 may be advantageous.
[0139] In an aspect addressed by this disclosure, we consider a grid of RDL / LSI connections arranged in any useful configuration on substrate 2222. In an example, the configuration of theRDL / LSI connections is done according to a rectilinear or Cartesian layout having rows 2225 and columns 2255. The number and density of the grid of connections may be determined by the nature and design of the SoW system it is to support. It is to be appreciated that any other layout of RDL / LSI connections can be prepared as well. For example, the layout may comprise regular geometric arrangements, irregular arrangements and / or combinations thereof. The layout may also be configured and arranged in groupings or clusters as described below. The layout may include row and column embodiments, hexagonal (honeycomb) arrangements, circular or polar coordinate system-based arrangements, or other configurations as suitable for a given SoW and chip placement scenario.
[0140] Configuration “B” may support the same chip and SoW components 2212 on the regular RDL / LSI grid 2223 of substrate 2222, where suitable connections may be provided for the operation of SoW 2202. Here, some or many RDL / LSI connections of RDL / LSI grid 2223 may be unused. But RDL / LSI configuration 2220 may be capable of supporting other SoW chip or component configurations as well, e.g., by connecting as needed at other available RDL / LSI connections on RDL / LSI grid 2223.
[0141] In other words, if we consider a base grid of RDL / LSI connections designed to accommodate more than one SoW chip design, such wafers may be considered “universal” in that they are capable of interchangeably supporting a variety of chip sizes and placements thereon, and are not only capable of supporting one arrangement.
[0142] In view of the present discussion, we note that the SoW systems described herein are suited for very high-performance computing applications as mentioned, including Al, language systems, image processing systems and other computing systems having a large amount and density of computing hardware and resources. In an aspect, the present applicant anticipates and accommodates designs that include 3DIC stacks of processors, memory devices and other chips on SoW. These systems may require significant cooling and heat removal capabilities.
[0143] FIG. 23 illustrates three possible RDL / LSI grid density embodiments on respective wafers 2300A, 2300B and 2300C. Wafer 2300A includes a first configuration of connections arranged in rows and columns as discussed above and supporting a certain range of chip placements thereon. Wafer 2300B may include a grid of RDL / LSI connections being sparser, less dense, or more spaced apart from the example of 2300 A and thus providing an alternative set of chip placements compared to the example of 2300A. Wafer 2300C may include a greater grid density or reduced spacing of RDL / LSI points compared to that of 2300A, andaccordingly, supporting potentially the same or a super set of chip placements and designs in a SoW as compared to that of wafer 2300 A.
[0144] Each wafer 2300A, 2300B, and 2300C may provide one or more advantages as compared to the other wafers. For example, the relative sparseness of wafer 2300B may have a lower manufacturing cost as compared to wafers 2300A and / or 2300C. Conversely, an increased density of RDL / LSI connections on wafers 2300A and 2300C may provide an optionality for a wider range of hardware components as compared to wafer 2300B.
[0145] FIG. 24 illustrates embodiments of a multi-grid universal wafer configured and arranged to support a corresponding plurality of SoW designs. FIG. 24 illustrates Wafer 2400 A, which may include two different RDL / LSI grid densities whereby the central region comprises a first grid density or spacing, and a second outer region comprises a second grid density or spacing. Wafer 2400B includes an RDL / LSI grid configuration in a first half of wafer 2400B and a second RDL / LSI grid configuration in a second half of the wafer. In wafer 2400C we see four quadrants of the wafer, each having a different RDL / LSI grid configuration to support corresponding different chip and component placements thereon.
[0146] The embodiments shown in FIG. 24 illustrate how a variety of hardware components may be accommodated on various wafers based on different RDL / LSI grids disposed in each wafer, as well as how a substrate can be manufactured to a particular specification and enable a plurality of layouts for computing hardware. This flexibility enables application-specific designs, and may additionally enable designers to optimize SoW computing, thermal, and space performance, including for multiple different SoWs within a single server or immersion cooled computing system.
[0147] Additionally, FIG. 24 illustrates how different components may be disposed in different portions (e.g., quadrants or sections) of wafers based on corresponding RDL / LSI arrangements. For example, wafer 2400B may be disposed in an immersion-cooled server in an orientation as depicted in FIG. 24, with a first density or arrangement of RDL / LSI connections on an upper half of the wafer, while the bottom half of the wafer has a second density or arrangement of RDL / LSI connections. In an aspect, the upper half may particularly support one or more GPUs, which may be the primary generators of waste heat for wafer 2400B and therefore generate relatively large amounts of immersion cooling vapor. If components were disposed on wafer 2400B above the GPUs, those components may receive a reduced amount of cooling while the GPUs are creating a relatively large amount of immersion coolingvapor, because the immersion cooling vapor may tend to displace immersion cooling liquid and may not have equally optimal heat transfer properties as immersion cooling liquid. Thus, the design flexibility enabled by the present technology provides advantages over traditional computing hardware designs.
[0148] Recognizing that the illustrated examples are for the sake of explanation and not limitation, wafer designs comprising RDL / LSI arrangements can include portions or subregions of an entire wafer surface, where each sub-region or portion has a different density or geometric configuration of universal or grid like positioning of the RDL / LSI elements.
[0149] As mentioned before, it is noted that the depicted examples are not limiting, and that in practice any number of respective universal RDL / LSI distributions on a wafer are conceivable that suit a given SoW design.
[0150] FIG. 25 illustrates an exemplary computing system on wafer (SoW 2500) based on a wafer-scale chip integration in the form of a vertical stack, or 3DIC as it is sometimes referred, as shown in the cross-sectional view of the drawing. Aspects of the SoW 2500 and its fabrication are described by the present applicant and the present disclosure comprehends structures and methods for implementing the present invention accordingly (See, e.g., U.S. Provisional Patent Application No. 63 / 605,344, filed on December 1, 2023 and entitled “Stacked Semiconductor Structure with Direct Electrical Connections;” each of which are incorporated herein by reference). The SoW 2500 comprises heterogeneous electrical and electronic components and integrated circuitry in said wafer-scale chip integration portion coupled to a power module 2510 and a connector 2520 unit disposed on and in electrical or signal communication with said wafer-scale chip integration.
[0151] The power module 2510 may provide electrical power to drive and power one or more portions of SoW 2500. The connector 2520 unit may provide signal or data communication functions to the circuits and chips of the wafer-scale chip integration portion of SoW 2500. Power module 2510 may convert, regulate and control the flow of electrical energy to provide stable power output to one or more components of SoW 2500. Power module 2510 may include different types of power semiconductor devices, such as metal oxide semiconductor field-effect transistors (MOSFETs), insulated gate bipolar transistors (IGBTs), silicon carbide (SiC), etc., to meet applications with different power and efficiency requirements.
[0152] Power module 2510 may enable the design and assembly process of the power supply system to be simplified, and may improve the reliability and efficiency of the system.Depending on the power delivery and functionality requirements, the power module can be further integrated with RLC (resistor (R), an inductor (L), and a capacitor (C)) passive devices, clocks, memory, and voltage regulator module (VRM) to fulfill the overall system requirements.
[0153] The wafer-scale chip integration may comprise a heterogeneous chip integration of a plurality of different types of chips. The wafer-scale chip integration may be a wafer-scale chip integration unit and which may be coupled to and sandwiched between other components in the SoW above and below (or coupled to a first and a second face of said chip integration unit). In the chip integration unit there may be more than one layer coupled to one another is such a sandwich configuration and including a chip layer where the chips or chiplets are disposed and interconnected with one another and / or external layers and devices through local silicon interconnects (LSI) in a redistribution layer (RDL).
[0154] A thermal module 2530 is coupled to and in thermal communication with said waferscale chip integration and the portions of SoW 2500 generally and integrated therewith as appropriate to provide temperature regulation and heat removal capability to SoW 2500. As described herein, some or all computing circuits and processing units of the present SoW 2500 may generate significant thermal energy in use in the context of high-performance computing systems. Such thermal modules 2530 can be coupled to the present system for thermal energy management (See, e.g., U.S. Provisional Patent Application No. 63 / 597,586, filed on Nov. 9, 2023 and entitled “Systems for Thermal Management in Three-Dimensional Integrated Circuits,” which is incorporated herein by reference in its entirety for all purposes). The thermal module 2530 may comprise or consist of an immersion cooling system as described herein.
[0155] FIG. 26 illustrates an exemplary computing system-on-a-wafer (SoW 2600) in crosssection and indicating some typical or exemplary parts thereof. As mentioned before, the system may comprise a heterogeneous apparatus having a wafer-scale chip integration, e.g., implemented as a 3D stack of chips, chiplets 2606, redistribution layers 2602, 2604 and connecting layers such as hybrid or bump connections 2605. Through silicon vias (TSV) and local silicon interconnects (LSI) are aspects that may be used in one or more embodiments. We shall incorporate herein the descriptions and references above regarding useful structures, devices, circuits and methods of manufacturing the same such as wafer-scale chip integration 2601 and thermal module 2630. Connector 2620 may provide communicative coupling between one or more components of SoW 2600 and another device external to SoW 2600, such as a second SoW or plurality of additional SoWs.
[0156] As an aspect, the wafer-scale system integration described may be combined with said power and thermal modules to eliminate the need for a PCB type system substrate within the package and overall computing system. The result can provide improved power efficiency, high bandwidth, and greater chip integration density compared to conventional designs. This may in an aspect yield improved computing and data processing capabilities needed in very high-performance computing applications. In a specific aspect, the present SoW designs and methods provide new system technology co-optimizations (STCO) platforms usable for chip / chiplet heterogeneous integrations. These can in turn be applied to computing tiles, I / O dies, switching modules, etc.
[0157] SoW 2600 may include optical engine (OE) 2625, which may be coupled to the other parts of the SoW and integration by way of any suitable method, including but not limited to direct bump connection, hybrid connection and / or pop-in socket connectors. These connections can thus provide signal or data communication for input / output (IO) purposes.
[0158] Power module 2610 is provided to supply electrical power to one or more components of SoW 2600, and may be electrically coupled thereto as appropriate for a given implementation, including using hybrid or bump bonds 2611 as conducting power connections, buses or lines.
[0159] FIG. 27 illustrates an exemplary high performance SoW system 2700 in cross-section, which shows a wafer-scale chip integration 2701 and thermal module 2730 and which may be similar to those described herein in some embodiments. Power, connectivity, and optical engine (OE 2725) components are disposed on the wafer-scale chip integration 2701 and / or coupled thereto as described, e.g., using hybrid or bump bonds.
[0160] In this example, OE 2725 is coupled to the other portions of the system 2700 by way of a socket 2720 allowing OE 2725 to be pressure placed into said socket 2720 to achieve the necessary electrical connection points between OE 2725 and conduction points to chip integration 2701.
[0161] OE 2725 in this and other embodiments may comprise any appropriate optical interface units for a given implementation. Without limitation, the present system and method may comprise linear drive pluggable optics (LPO) engines, co-packaged optics (CPO), near packaged optics (NPO) or equivalent or evolutions thereof, as well as combinations of such engines as useful for a given system and use.
[0162] Immersion cooling systems may provide particular advantage to SoWs and 3DICstacks due to the lower surface area to volume ratio of a 3DIC stack compared to the individual components of the 3DIC stack (e.g., a bonded logic IC and memory module will have a lower surface area to volume ratio than the combined surface area to volume ratio of the physically separated logic IC and memory module, thus resulting in a lower total surface area for the 3DIC stack) as well as the additional heat generated by state of the art logic ICs. This lower surface area to volume ratio means waste heat generated by the 3DIC stack may not be as efficiently dissipated and may require better cooling performance than air cooling can provide. Two-phase immersion cooling in particular can provide this additional heat removal required by 3DIC stacks.
[0163] One or more semiconductor die packages 2805 may be 3DIC stacks in accordance with the present technology. For example, one or more semiconductor die packages 2805 may include a logic IC and at least one memory module bonded to the logic IC using a hybrid bond or micro-bump bond.
[0164] FIG. 28 depicts aspects of an immersion cooling system 2800 for dissipating heat from one or more heat-generating components such as semiconductor die packages 2805 via immersion cooling. Each package 2805 can include one or more semiconductor dies that produce heat when the system is in operation. The immersion cooling system 2800 in the illustrated example of FIG. 28 is a two-phase immersion cooling system, though the invention may also be implemented in a single-phase immersion cooling system.
[0165] Immersion cooling system 2800 includes a container such as tank 2820 filled, at least in part, with immersion cooling liquid 2864. The immersion cooling system 2800 can further include at least one chiller 2880 that flows a heat-transfer fluid through at least one condenser tube 2870 that is disposed in the tank 2820 and headspace 2808. Condenser tubes 2870 and chiller 2880 may be part of a heat exchanger. The packages 2805 can be mounted on one or more printed circuit boards (PCBs) 2857 that are immersed, at least in part, in the immersion cooling liquid 2864. Immersion-cooling system 2800 may further include a filter 2875 disposed adjacent to the tank 2820.
[0166] Filter 2875 may include a filtration media, a housing, and a pump configured to force immersion cooling liquid 2864 through filter 2875 to remove contaminants, particulates, or other impurities that may be added to immersion cooling liquid 2864 during use. Filter 2875 may be housed outside of tank 2820 while being in fluidic communication with immersion cooling liquid 2864 in tank 2820. Alternatively, filter 2875 may be submerged within immersion cooling liquid 2864 inside of tank 2820.
[0167] Immersion cooling liquid 2864 may be a hydrocarbon, a fluoroketone, an oil, or a similar dielectric liquid that will act as an insulator while simultaneously transferring heat from package 2805 more efficiently than air. Examples of immersion cooling liquid 2864 are Novec™ 649, Novec™ 7000, and Novec™ 7100 produced by 3M™. An exemplary immersion cooling liquid 2864 used in accordance with embodiments of the present invention may have a dielectric constant baseline value of about 1.8-2 at a frequency of about 1 kHz.
[0168] In an embodiment of the invention, immersion cooling liquid 2864 may be considered unacceptably contaminated if the dielectric constant and / or dielectric loss tangent of immersion cooling fluid being used in an immersion cooling system 2800 differs by a threshold amount as compared to unused or pure immersion cooling liquid 2864. For example, immersion cooling liquid 2864 may be considered unacceptably contaminated or degraded if the dielectric constant and / or dielectric loss tangent differs by a threshold of 10% or more as compared to unused or pure immersion cooling liquid 2864. In an embodiment, a dielectric constant and / or dielectric loss tangent variation threshold may be 20%, 15%, 5%, 3%, 1%, or any suitable threshold.
[0169] Contamination of the immersion cooling liquid 2864 and resulting changes to dielectric constant and / or dielectric loss tangent may alter or negatively impact operation of components within immersion cooling liquid 2864 including semiconductor die(s) 2850. An altered dielectric constant and / or dielectric loss tangent may result in undesirable cross-talk between components on a PCB, additional noise or reduction in signal strength transmitted along exposed wires of a PCB or semiconductor die(s) 2850 submerged in immersion fluid, and / or signal dissipation through the immersion cooling liquid 2864. Signal loss may be severe enough that two elements may be effectively represented as being separated by an open circuit despite being physically connected. In an embodiment, a dielectric constant and / or dielectric loss tangent variation threshold may be selected based on an observed or inferred effect on one or more submerged semiconductor die(s) 2850. For example, an increase in PCIe bit error rate above an error rate baseline may be correlated with an increase in dielectric constant and / or dielectric loss tangent above a dielectric constant and / or dielectric loss tangent baseline. Accordingly, operation of semiconductor die(s) 2850 may be throttled or suspended when a dielectric constant and / or dielectric loss tangent of immersion cooling liquid 2864 exceeds a predetermined threshold.
[0170] Changes to dielectric constant and / or dielectric loss tangent may be caused by contaminants within immersion cooling liquid 2864. In some cases, changes to dielectricconstant and / or dielectric loss tangent may be reversed by filtering the contaminants from immersion cooling liquid 2864. In some embodiments, upon detecting an increase in dielectric constant and / or dielectric loss tangent of immersion cooling liquid 2864, controller 2802 may instruct filter 2875 to increase filtration throughput or notify a user that an immersion cooling liquid 2864 filtration media may need to be replaced. If a dielectric constant and / or dielectric loss tangent exceeds a predetermined threshold, controller 2802 may throttle or shut down one or more semiconductor die(s) 2850, generate a notification that immersion cooling liquid 2864 should be replaced, trigger an alarm, etc.
[0171] Further examples of sensors and methods for immersion cooling contamination monitoring may include probes for monitoring immersion cooling liquid parameters such as dielectric constant and dielectric loss tangent, and processors configured to identify trends in sensor data, model immersion cooling system behavior as a function of contamination, and alter operations of immersion cooling systems based on detected levels and / or states of contamination may be found in U.S. Provisional Patent Application 63 / 516,748, filed July 31, 2023 and entitled “Di-Electric Monitoring of Immersion Fluid During Cooling Operation,” the entirety of which is incorporated herein by reference.
[0172] The illustrated example of FIG. 28 is not intended to be to scale. The immersion cooling system 2800 may house and provide immersion cooling liquid 2864 to tens, hundreds, or even thousands of packages 2805. In some cases, the immersion cooling system 2800 can be small (e.g., the size of a floor unit air conditioner, approximately 1 meter high, 0.5 meter width, 0.5 meter depth or length). In some implementations, the immersion cooling system can be large (e.g., the size of a van or larger, approximately 2.5 meters high, 2.5 meters width, 4 meters depth or length).
[0173] The immersion cooling system 2800 can also include a controller 2802 (e.g., a microcontroller, programmable logic controller (PLC), microprocessor, field-programmable gate array, logic circuitry, memory, or some combination thereof) to manage system operation. Controller 2802 can perform various system functions such as monitoring temperatures of system components, cooling fluid level, tank access, chiller operation etc. The controller 2802 can further issue commands to control system operation such as executing a start-up sequence, executing a shut-down sequence, assigning workloads among the packages, changing cooling fluid level, changing the temperature of the heat-transfer fluid circulated by the chiller 2880, etc. In some implementations, controller 2802 can include (or itself be) a baseboard management controller (BMC) 2804. That is, the BMC 2804 may monitor and control allaspects of system operation for the immersion cooling system 2800 in addition to monitoring and controlling workloads of the semiconductor dies 2850 in the packages 2805 cooled by the system. The immersion cooling system 2800 can also include a network interface controller (NIC 2803) to allow the system to communicate over a network, such as a local area network or wide area network. The immersion cooling system 2800 can further include a fluid sensor array 2890 having a plurality of fluid sensors 2810. Fluid sensors 2810 may include one or more leak detection sensors at least partially submerged in immersion cooling liquid 2864.
[0174] The semiconductor die(s) 2850 and can be mounted on and attached to a printed circuit board (PCB) 2855 (sometimes referred to as a substrate) in device package 2805. The package 2805 can be made commercially available as an off-the-shelf (OTS) product. The package 2805 can be used for single-phase or two-phase immersion cooling of at least one semiconductor die 2850, such as a microprocessor (e.g., a central processing unit (CPU) and / or graphics processing unit (GPU)), voltage regulator (VR), high bandwidth memory (HBM), a digital signal processing (DSP) die, an artificial intelligence (Al) accelerator, an application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), and / or other densely patterned semiconductor die.
[0175] In the two-phase immersion cooling system 2800 of FIG. 28, heat flows from the semiconductor die 2850 where it is generated into the heat spreader 2852. The heat spreader 2852 is in thermal contact with an immersion cooling liquid 2864 that can flow over and extract heat from the heat spreader 2852. The amount of heat delivered by the heat spreader 2852 to the immersion cooling liquid 2864 is enough to boil the immersion cooling liquid 2864 that contacts the heat spreader 2852 (creating bubbles 2865 and potentially creating froth 2867 when bubbles 2865 reach the surface of immersion cooling liquid 2864). The vapor 2866 from the boiled immersion cooling liquid 2864 can be cooled and condensed back to liquid droplets 2868, for example, by the condenser tube 2870. The heat-transfer fluid, such as chilled water, from the chiller 2880 can be circulated through the condenser tube 2870 to lower the temperature of the condenser tube 2870 below the condensation point in the headspace 2808 of the tank 2820. As a result, vapor 2866 condenses on exterior surfaces of the condenser tube 2870 and liquid droplets 2868 from the condensed vapor can drip and / or flow back to the immersion cooling liquid 2864. There may be a plurality of condenser tubes 2870 in tank 2820 to condense the vapor 2866 into droplets. Some or all of the condenser tubes 2870 may or may not be located directly over the PCBs 2857. Instead, the condenser tube(s) 2870 can be located near one or more walls of the tank 2820, such that the condenser tube(s) 2870 are not directlyover the PCBs 2857 on which the packages 2805 are mounted.
[0176] To improve thermal performance in two-phase immersion cooling system 2800, the heat spreader 2852 can include a boiling enhancement coating (BEC) on at least one surface. The BEC can be formed from copper or a copper alloy and can be porous, for example, though BECs can take various forms. In some cases, the BEC is a micro porous copper coating having a thickness from approximately or exactly 50 microns to 500 microns thick (which may be produced by electroplating and / or etching). In some implementations, the BEC comprises a mesh copper layer bonded (e.g., via resistance heating) to at least an outer surface of the heat spreader 2852. In some cases, the BEC is applied as particulates to at least one smooth surface of the heat spreader 2852 and then subsequently sintered to adhere to one another and to the heat spreader 2852. The BEC provides an improved surface area to contact the immersion cooling liquid 2864 and can increase the heat transfer coefficient from the heat spreader 2852 to the immersion cooling liquid 2864 by up to a factor of 15 versus a smooth surface on the heat spreader 2852. Accordingly, BECs can increase thermal conductivity to, and accelerate the boiling of, the immersion cooling liquid 2864.
[0177] Further implementations of boiling enhancement coatings and enclosures are possible. Additional arrangements, applications, and methods of use of boiling enhancement coatings and enclosures, including with semiconductor dies and 3DIC stacks, are described in the below U.S. Patent Applications.
[0178] U.S. Patent Application No. 18 / 327,615, filed June 1, 2023 and entitled "Boiler Enhancement Coatings with Active Boiling Management,” discloses heat spreader and boiling enhancement enclosure architectures thermally and / or mechanically coupled to one or more semiconductor dies or logic ICs that may be used for passive and / or active management of immersion cooling fluid boiling, including through the use of valves to control pressure of boiling immersion cooling fluid within a boiling enhancement chamber, particularly in paragraphs
[0018] -
[0039] and FIGS. 3-5B. The entirety of U.S. Patent Application No. 18 / 327,615 is incorporated herein by reference.
[0179] U.S. Provisional Patent Application No. 63 / 500,167, filed May 4, 2023 and entitled “Direct to Chip Heat Spreader and Boiler Enhancement Coatings for Microelectronics,” discloses heat spreader and BECs thermally and / or mechanically coupled to one or more semiconductor dies, logic ICs, and / or 3DIC stacks, particularly in paragraphs
[0015] -
[0033] and FIGS. 2A-4. BEC form factors may include graphite heat spreader architectures, vaporchambers, heat pipes, copper plates, fins, and the like. BEC form factors may be thermally and / or mechanically coupled to the one or more semiconductor dies, logic ICs, and / or 3DIC stacks through a thermally conductive epoxy, and may have varying dimensions relative to a surface to which the semiconductor dies and / or logic ICs are mounted. The entirety of U.S. Provisional Patent Application No. 63 / 500,167 is incorporated herein by reference.
[0180] U.S. Patent Application No. 18 / 460,091, filed September 1, 2023 and entitled “Direct to Chip Application of Boiling Enhancement Coating,” discloses BECs and methods for applying BECs to semiconductor dies, logic ICs, and / or 3DIC stacks in accordance with the present technology. In particular, paragraphs
[0024] -
[0046] and FIGS. 2A-5 disclose embodiments of BEC layers, adhesives, solders, sintering, laser ablation, meshes, and other BECs and BEC application methods. The entirety of U.S. Patent Application No. 18 / 460,091 is incorporated herein by reference.
[0181] U.S. Provisional Patent Application No. 63 / 506,945, filed June 8, 2023 and entitled “Vapor- Shedding Structures for Boiler Plates in Two-Phase Immersion Cooling Systems,” discloses structures that may be thermally and / or mechanically coupled to computing hardware such as one or more semiconductor dies, logic ICs, and / or 3DIC stacks to enable the shedding of immersion cooling vapors generated from the boiling of immersion cooling fluid during operation of the computing hardware. In particular, paragraphs
[0021] -
[0039] and FIGS. 3A-5 disclose vapor-shedding structures including varying porosities, constituent materials, and geometries relative to the computing hardware on which they are mounted. The entirety of U.S. Provisional Patent Application No. 63 / 506,945 is incorporated herein by reference.
[0182] U.S. Provisional Application No. 63 / 513,828, filed July 14, 2023 and entitled “Grinding Apparatuses and Methods for Mechanically Modifying Surfaces of Processors to Promote Boiling of a Coolant Liquid,” discloses methods for creating boiling enhancement modifications to surfaces such as the surfaces of computing hardware such as one or more semiconductor dies, logic ICs, and / or 3DIC stacks, particularly in paragraphs
[0036] -
[0095] and FIGS. 2A-8. For example, grooves, patterns, gouges, trenches, or other structures may be added to a surface or lid of a processor, semiconductor die, logic IC, 3DIC stack component, and / or BEC to encourage nucleation sites for bubbles of immersion cooling vapor to form during a cooling process, thus decreasing the thermal resistance between the processor, semiconductor die, logic IC, and / or 3DIC stack component and the surrounding immersion cooling fluid. The entirety of U.S. Provisional Application No. 63 / 513,828 is incorporated herein by reference.
[0183] U.S. Provisional Patent Application No. 63 / 513,829, filed July 14, 2023 and entitled “Electrical Connector Having a Heater to Facilitate Boiling of a Coolant Liquid to Improve Signal Integrity in Immersion Cooling Environment,” discloses heaters for promoting boiling of immersion cooling fluid near electrical connectors such as connections between components of a 3DIC stack and enable improved impedances at those connectors, particularly in paragraphs
[0019] -
[0052] and FIGS. 1A-3B. The entirety of U.S. Provisional Patent Application No. 63 / 513,829 is incorporated herein by reference.
[0184] U.S. Provisional Patent Application No. 63 / 603,242, filed November 28, 2023 and entitled “Woven Boiler Enhancement Coatings,” provides additional examples of BECs including woven BECs with variable weave patterns, densities, attachment mechanisms, and materials (including copper and tungsten) that may be attached to computing hardware such as one or more semiconductor dies, logic ICs, and / or 3DIC stacks in order to promote more efficient heat transfer and immersion cooling vapor nucleation, particularly in paragraphs
[0031] -
[0055] and FIGS. 3-7. The entirety of U.S. Provisional Patent Application No. 63 / 603,242 is incorporated herein by reference.
[0185] Clause 1. A system-on-a-wafer (SoW), the SoW comprising: a substrate; a first computing component disposed on the substrate and a second computing component disposed on the substrate; and a plurality of interconnects disposed in the substrate in a regular pattern and communicatively coupling the first computing component and the second computing component.
[0186] Clause 2. The system of clause 1, further comprising: a plurality of third computing components, analogous to the first computing component, disposed on the substrate and communicatively coupled to at least one another through at least some of the plurality of interconnects; a plurality of fourth computing components, analogous to the second computing component, disposed on the substrate and communicatively coupled to at least one another through at least some of the plurality of interconnects; wherein: the first computing component and the plurality of third computing components are disposed on the substrate according to the regular pattern; and the second computing component and the plurality of fourth computing components are disposed on the substrate according to the regular pattern.
[0187] Clause 3. The system of clause 2, wherein the first computing component and the plurality of third computing components comprise graphics processing units (GPUs).
[0188] Clause 4. The system of clause 1, further comprising: a chip connector communicatively coupled to the first computing component and the second computing component and configured to transfer data between at least the first computing component, the second computing component, and a device external to the SoW.
[0189] Clause 5. The system of clause 4, wherein the device external to the SoW comprises a second SoW.
[0190] Clause 6. The system of clause 1, wherein the first computing component and the second computing component are analogous.
[0191] Clause 7. The system of clause 1, further comprising: a power module electrically coupled to the first computing component and the second computing component; wherein the power module is configured to provide power to the first computing component and the second computing component.
[0192] Clause 8. The system of clause 1, wherein the second computing component comprises at least one of a network interface card (NIC), a central processing unit (CPU), a switch, an optical engine, or a memory module.
[0193] Clause 9. The system of clause 1, wherein the substrate comprises a semiconductor wafer.
[0194] Clause 10. The system of clause 9, wherein the semiconductor wafer has a substantially circular shape.
[0195] Clause 11. The system of clause 9, wherein the semiconductor wafer has a substantially rectilinear shape.
[0196] Clause 12. A server comprising: a plurality of systems-on-a-wafer (SoWs), each SoW comprising: a semiconductor wafer; a plurality of chips mounted on the semiconductor wafer; a power module; one or more chip connectors; and a wiring layer that electrically couples each chip to the power module and to at least one of the one or more chip connectors; and a plurality of inter-wafer connectors, each inter-wafer connector electrically coupling at least one of the one or more chip connectors of one SoW to at least one of the one or more chip connectors of another SoW.
[0197] Clause 13. The server of clause 12, further comprising an inter-SoW connector that electrically connects at least a first chip connector of a first SoW of the plurality of SoWs to at least a second chip connector of a second SoW of the plurality of SoWs.
[0198] Clause 14. The server of clause 12, further comprising a plurality of inter-SoW connectors, each inter-SoW connector electrically connecting a respective chip connector of a respective SoW to another respective chip connector of another respective SoW.
[0199] Clause 15. The server of any of the preceding clauses, wherein each SoW has the same configuration as each other SoW.
[0200] Clause 16. The server of any of the preceding clauses, wherein each SoW of the plurality of SoWs has a different configuration than each other SoW of the plurality of SoWs.
[0201] Clause 17. The server of any of the preceding clauses, wherein the plurality of chips in each SoW includes at least one of: logic chips, switches, memory modules, network interface controllers (NICs), or input / output (I / O) controllers.
[0202] Clause 18. The server of clause 17, wherein the plurality of chips in each SoW includes the logic chips, the switches, the memory modules, the NICs, and the I / O controllers.
[0203] Clause 19. The server of clause 17 or 18, wherein the logic chips include at least one of: graphics processing units, central processing units, data processing units, or tensor processing units.
[0204] Clause 20. The server of any of the preceding clauses, wherein the plurality of chips in each SoW are selected and arranged to form an artificial intelligence training server.
[0205] Clause 21. The server of any of clauses 12-18, wherein the plurality of chips in each SoW are selected and arranged to form an artificial intelligence inference server.
[0206] Clause 22. The server of any of the preceding clauses, wherein the wiring layer comprises a redistribution layer that includes electrical wiring electrically coupled to each of: the chips, to the power module, and to the chip connectors.
[0207] Clause 23. The server of clause 22, wherein the wiring layer forms an integrated fanout packaging connection.
[0208] Clause 24. The server of any of clauses 12-22, wherein the wiring layer comprises: a first redistribution layer that includes first electrical wiring that is electrically coupled to the power module and to the chip connectors; a second redistribution layer that includes first electrical wiring electrically coupled to the chips; and a plurality of local silicon interconnects that is electrically coupled to the first and second electrical wiring.
[0209] Clause 25. The server of any of the preceding clauses, further comprising at least one thermal module in direct physical contact with and / or in thermal communication with the SoWs.
[0210] Clause 26. The server of clause 25, wherein the at least one thermal module includes an immersion cooler.
[0211] Clause 27. A high performance computing system on a wafer (SoW) comprising: a wafer-scale chip integration unit having two opposing faces and comprising a plurality of semiconductor computing chips and respective interconnects; a thermal module coupled to a first face of said wafer-scale chip integration unit; and an optical engine coupled to a second face of said wafer-scale chip integration unit.
[0212] Clause 28. The system of clause 27, wherein said wafer-scale chip integration unit comprises said plurality of semiconductor computing chips in a chip layer of said unit, and the respective interconnects comprise local silicon interconnects (LSI) in a redistribution layer (RDL) of said unit providing conduction pathways to said semiconductor computing chips.
[0213] Clause 29. The system of clause 27, wherein said semiconductor computing chips comprise any of: central processing units (CPU), general processing units (GPU), and input / output (IO) units.
[0214] Clause 30. The system of clause 27, wherein said thermal module comprises a two- phase immersion cooling system in thermal communication with said wafer-scale chip integration unit.
[0215] Clause 31. The system of clause 27, wherein said optical engine comprises any of: a linear drive pluggable optics (LPO) engine, a co-packaged optics (CPO) engine, and a near packaged optics (NPO) engine.
[0216] Clause 32. The system of clause 27, wherein said optical engine is coupled to said wafer-scale chip integration unit through any of: a hybrid bond layer, a bump bond layer, and a plug-in or push socket.
[0217] Clause 33. The system of clause 27, further comprising an integrated connector interface disposed on and in data communication with said wafer-scale chip integration unit.
[0218] Clause 34. The system of clause 27, further comprising a power module in electrical communication with said wafer-scale chip integration unit.
[0219] Clause 35. The system of clause 27, further comprising a data switch unit in data communication with said wafer-scale chip integration unit.
[0220] Clause 36. The system of clause 27, wherein said wafer-scale chip integration unit comprises a heterogeneous unit having a plurality of different types of semiconductor computing chips.
[0221] Clause 37. A high performance computing system on a wafer (SoW) comprising: a wafer substrate; a plurality of computing hardware chips (chips) disposed on a surface of said wafer substrate and defining a spatial arrangement between and among said chips; and a plurality of electrical interconnects placing said chips in electrical communication according to a layout of said chips on the surface of said wafer substrate; wherein : said electrical interconnects comprise respective positions with respect to one another and with respect to the surface of said wafer substrate; said electrical interconnects are disposed in a regular geometric arrangement and have a spatial density on a portion of the surface of said wafer substrate; and said spatial density is greater than a minimum required density to support said electrical communication according to said layout of said chips.
[0222] Clause 38. The system of clause 37, wherein said regular geometric arrangement comprises a regularly spaced grid of rows and columns on the surface of said wafer.
[0223] Clause 39. The system of clause 37, wherein said spatial density is a first spatial density and said portion of the surface is a first portion of the surface, and the system further comprises a second portion of said surface on which a second spatial density of interconnects is disposed.
[0224] Clause 40. The system of clause 37, wherein said plurality of computing hardware chips comprises a plurality of processing chips including any of: central processing units (CPU) and general processing units (GPU).
[0225] Clause 41. The system of clause 37, wherein said plurality of computing hardware chips comprises a plurality of data storage units including any of: DRAM, SRAM or other memory storage hardware chips.
[0226] Clause 42. The system of clause 37, wherein said electrical interconnects comprise redistribution layer (RDL) and local silicon interconnect (LSI) interconnects.
[0227] Clause 43. The system of clause 37, wherein said wafer substrate comprises a carrier silicon layer.
[0228] Clause 44. The system of clause 37, wherein said chips are configured and arranged in a layer substantially parallel to said wafer substrate and wherein said chips lie in spatial separation separated by a molding material therebetween.
[0229] Clause 45. The system of clause 37, wherein said electrical interconnects comprise a conducting metal material disposed in a dielectric layer of said SoW substantially parallel to said wafer substrate.
[0230] Clause 46. The system of clause 37, further comprising a plurality of stacked integrated circuits (3DIC) on said computing hardware chips.
[0231] Clause 47. The system of clause 37, further comprising an immersion cooling assembly in thermal communication with said computing hardware chips and operational to remove thermal energy generated by said chips.
[0232] Clause 48. The system of clause 39, wherein said first spatial density in said first portion of said surface is greater than said second spatial density in said second portion of said surface.
[0233] Clause 49. The system of clause 37, wherein said arrangement and said spatial density are such that said wafer substrate can support a plurality of computing hardware chips thereon and wherein only a subset of said plurality of electrical interconnects are required for interconnection of said chips while other electrical interconnects are not required.Conclusion
[0234] While various inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the inventive teachings is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specificallydescribed and claimed. Inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the inventive scope of the present disclosure.
[0235] Also, various inventive concepts may be embodied as one or more methods, of which an example has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.
[0236] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0237] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”
[0238] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0239] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly oneelement of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e. “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.
[0240] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0241] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.
Claims
CLAIMS1. A system-on-a-wafer (SoW), the SoW comprising: a substrate; a first computing component disposed on the substrate and a second computing component disposed on the substrate; and a plurality of interconnects disposed in the substrate in a regular pattern and communicatively coupling the first computing component and the second computing component.
2. The SoW of claim 1, further comprising: a plurality of third computing components, analogous to the first computing component, disposed on the substrate and communicatively coupled to at least one another through at least some of the plurality of interconnects; and a plurality of fourth computing components, analogous to the second computing component, disposed on the substrate and communicatively coupled to at least one another through at least some of the plurality of interconnects; wherein: the first computing component and the plurality of third computing components are disposed on the substrate according to the regular pattern; and the second computing component and the plurality of fourth computing components are disposed on the substrate according to the regular pattern.
3. The SoW of claim 2, wherein the first computing component and the plurality of third computing components comprise graphics processing units (GPUs).
4. The SoW of claim 1, further comprising: a chip connector communicatively coupled to the first computing component and the second computing component and configured to transfer data between at least the first computing component, the second computing component, and a device external to the SoW.
5. The SoW of claim 4, wherein the device external to the SoW comprises a second SoW.
6. The SoW of claim 1, wherein the first computing component and the second computingcomponent are analogous.
7. The SoW of claim 1, further comprising: a power module electrically coupled to the first computing component and the second computing component; wherein the power module is configured to provide power to the first computing component and the second computing component.
8. The SoW of claim 1, wherein the second computing component comprises at least one of a network interface card (NIC), a central processing unit (CPU), a switch, an optical engine, or a memory module.
9. The SoW of claim 1, wherein the substrate comprises a semiconductor wafer.
10. The SoW of claim 9, wherein the semiconductor wafer has a substantially circular shape.
11. The SoW of claim 9, wherein the semiconductor wafer has a substantially rectilinear shape.
12. A server compri sing : a plurality of systems-on-a-wafer (SoWs), each SoW comprising: a semiconductor wafer; a plurality of chips mounted on the semiconductor wafer; a power module; one or more chip connectors; and a wiring layer that electrically couples each chip to the power module and to at least one of the one or more chip connectors; and a plurality of inter-wafer connectors, each inter-wafer connector electrically coupling at least one of the one or more chip connectors of one SoW to at least one of the one or more chip connectors of another SoW.
13. The server of claim 12, further comprising an inter-SoW connector that electrically connects at least a first chip connector of a first SoW of the plurality of SoWs to at least asecond chip connector of a second SoW of the plurality of SoWs.
14. The server of claim 12, further comprising a plurality of inter-SoW connectors, each inter-SoW connector electrically connecting a respective chip connector of a respective SoW to another respective chip connector of another respective SoW.
15. The server of any one of claims 12-14, wherein each SoW has a same configuration as each other SoW.
16. The server of any one of claims 12-15, wherein each SoW of the plurality of SoWs has a different configuration than each other SoW of the plurality of SoWs.
17. The server of any one of claims 12-16, wherein the plurality of chips in each SoW includes at least one of: logic chips, switches, memory modules, network interface controllers (NICs), or input / output (I / O) controllers.
18. The server of claim 17, wherein the plurality of chips in each SoW includes the logic chips, the switches, the memory modules, the NICs, and the I / O controllers.
19. The server of claim 17 or 18, wherein the logic chips include at least one of: graphics processing units, central processing units, data processing units, or tensor processing units.
20. The server of any one of claims 12-19, wherein the plurality of chips in each SoW are selected and arranged to form an artificial intelligence training server.
21. The server of any of claims 12-18, wherein the plurality of chips in each SoW are selected and arranged to form an artificial intelligence inference server.
22. The server of any one of claims 12-21, wherein the wiring layer comprises a redistribution layer that includes electrical wiring electrically coupled to each of: the chips, the power module, and the chip connectors.
23. The server of claim 22, wherein the wiring layer forms an integrated fan-out packaging connection.
24. The server of any of claims 12-22, wherein the wiring layer comprises: a first redistribution layer that includes first electrical wiring that is electrically coupled to the power module and to the chip connectors; a second redistribution layer that includes first electrical wiring electrically coupled to the chips; and a plurality of local silicon interconnects that is electrically coupled to the first and second electrical wiring.
25. The server of any one of claims 12-24, further comprising at least one thermal module in direct physical contact with and / or in thermal communication with the SoWs.
26. The server of claim 25, wherein the at least one thermal module includes an immersion cooler.
27. A high performance computing system on a wafer (SoW) comprising: a wafer-scale chip integration unit having two opposing faces and comprising a plurality of semiconductor computing chips and respective interconnects; a thermal module coupled to a first face of said wafer-scale chip integration unit; and an optical engine coupled to a second face of said wafer-scale chip integration unit.
28. The system of claim 27, wherein said wafer-scale chip integration unit comprises said plurality of semiconductor computing chips in a chip layer of said unit, and the respective interconnects comprise local silicon interconnects (LSI) in a redistribution layer (RDL) of said unit providing conduction pathways to said semiconductor computing chips.
29. The system of claim 27, wherein said semiconductor computing chips comprise any of: central processing units (CPU), general processing units (GPU), and input / output (IO) units.
30. The system of claim 27, wherein said thermal module comprises a two-phase immersion cooling system in thermal communication with said wafer-scale chip integration unit.
31. The system of claim 27, wherein said optical engine comprises any of: a linear drive pluggable optics (LPO) engine, a co-packaged optics (CPO) engine, and a near packaged optics (NPO) engine.
32. The system of claim 27, wherein said optical engine is coupled to said wafer-scale chip integration unit through any of: a hybrid bond layer, a bump bond layer, and a plug-in or push socket.
33. The system of claim 27, further comprising an integrated connector interface disposed on and in data communication with said wafer-scale chip integration unit.
34. The system of claim 27, further comprising a power module in electrical communication with said wafer-scale chip integration unit.
35. The system of claim 27, further comprising a data switch unit in data communication with said wafer-scale chip integration unit.
36. The system of claim 27, wherein said wafer-scale chip integration unit comprises a heterogeneous unit having a plurality of different types of semiconductor computing chips.
37. A high performance computing system on a wafer (SoW) comprising: a wafer substrate; a plurality of computing hardware chips (chips) disposed on a surface of said wafer substrate and defining a spatial arrangement between and among said chips; and a plurality of electrical interconnects placing said chips in electrical communication according to a layout of said chips on the surface of said wafer substrate; wherein: said electrical interconnects comprise respective positions with respect to one another and with respect to the surface of said wafer substrate; said electrical interconnects are disposed in a regular geometric arrangement and have a spatial density on a portion of the surface of said wafer substrate; and said spatial density is greater than a minimum required density to support said electrical communication according to said layout of said chips.
38. The system of claim 37, wherein said regular geometric arrangement comprises a regularly spaced grid of rows and columns on the surface of said wafer.
39. The system of claim 37, wherein said spatial density is a first spatial density and said portion of the surface is a first portion of the surface, and the system further comprises a second portion of said surface on which a second spatial density of interconnects is disposed.
40. The system of claim 37, wherein said chips comprise any of: central processing units (CPU) and general processing units (GPU).
41. The system of claim 37, wherein said chips comprise any of: DRAM, SRAM or other memory storage hardware chips.
42. The system of claim 37, wherein said electrical interconnects comprise redistribution layer (RDL) and local silicon interconnect (LSI) interconnects.
43. The system of claim 37, wherein said wafer substrate comprises a carrier silicon layer.
44. The system of claim 37, wherein said chips are configured and arranged in a layer substantially parallel to said wafer substrate and wherein said chips lie in spatial separation separated by a molding material therebetween.
45. The system of claim 37, wherein said electrical interconnects comprise a conducting metal material disposed in a dielectric layer of said SoW substantially parallel to said wafer substrate.
46. The system of claim 37, further comprising a plurality of stacked integrated circuits (3DIC) on said chips.
47. The system of claim 37, further comprising an immersion cooling assembly in thermal communication with said chips and operational to remove thermal energy generated by said chips.
48. The system of claim 39, wherein said first spatial density in said first portion of said surface is greater than said second spatial density in said second portion of said surface.
49. The system of claim 37, wherein said arrangement and said spatial density are such that said wafer substrate can support the plurality of computing hardware chips thereon and wherein only a subset of said plurality of electrical interconnects are required for interconnection of said chips while other electrical interconnects are not required.
Citation Information
Patent Citations
Wafer-level system and generation method thereof, data processing method and storage medium
CN114743956A
4d device process and structure
US20110170266A1
Integrated Circuit Package and Method
US20200212018A1
Cooled system-on-wafer with means for reducing the effects of electrostatic discharge and / or electromagnetic interference
WO2022192034A1
Sandwiched multi-layer structure for cooling high power electronics
WO2023023090A1