Disaggregated ai server design for immersion

A disaggregated AI server design with flexible node connectivity over PCIe links addresses inefficiencies in existing systems by separating compute nodes and GPUs, enhancing resource utilization, scalability, and maintaining system functionality.

WO2025250380A1PCT designated stage Publication Date: 2025-12-04MTS IP HLDG LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/029618
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-28
Filing Date
2025-05-15
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing AI server designs with fixed compute nodes paired with system components, such as GPUs, are inefficient in resource utilization, maintenance, and scalability, leading to potential signal distortion and error rates in immersion cooling systems.

Method used

Implementing a disaggregated AI server design using flexible node connectivity over PCIe links, allowing compute nodes and GPUs to be separated across different tanks, enabling independent management and scalability, with optical switches facilitating communication between tanks.

Benefits of technology

The design enhances resource utilization and scalability, reduces power consumption, and maintains system functionality even if one device fails, while optimizing infrastructure costs and improving communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025029618_04122025_PF_FP_ABST
    Figure US2025029618_04122025_PF_FP_ABST
Patent Text Reader

Abstract

A device may include a compute tank, including at least one CPU, at least one Electrical-to-Optical (EO) converter, and a first set of computing nodes configured to execute compute functions. A device may include a GPU tank comprising a second set of computing nodes configured to execute at least one accelerator function. A device may include an optical switch coupled to the compute tank and to the GPU tank, wherein the at least one EO converter is coupled to the optical switch, the at least one CPU is coupled to the at least one EO converter, and the second set of computing nodes are configured to execute the at least one accelerator function, based at least in part on a request from an application and a state of the first set of computing nodes.
Need to check novelty before this filing date? Find Prior Art

Description

DISAGGREGATED Al SERVER DESIGN FOR IMMERSIONCROSS-REFERENCE TO RELATED APPLICATION

[0001] This Application claims the benefit of provisional U.S. Application No. 63 / 652,556, filed on May 28, 2024, and entitled “DISAGGREGATED Al SERVER DESIGN FOR IMMERSION” which is hereby incorporated by reference in its entirety.

[0002] In cases where the present application conflicts with a document incorporated by reference, the present application controls.BACKGROUND

[0003] Implementing computing systems and computations can require cooling, for example, immersion cooling of system components using a tank containing a liquid (i.e., coolant liquid) used to maintain appropriate temperatures for system components. These computing environments can include cloud computing, systems implementing machine learning tasks, or systems implementing high-power computations, including run-time computations for processes associated with artificial intelligence and inference. System components include, but are not limited to, CPUs, GPUs and / or integrated circuits (IC). These tanks can function as cloud computing systems providing high-power computations for processes including artificial intelligence (Al) training and inference. As such, immersion fluid systems often rely on Al systems to implement high-level computations.

[0004] Existing Al server designs are fixed compute nodes, paired physically with system components, such as GPUs. The fixed assembly combination of system components can be inefficient in resource utilization, and maintenance. For scalability, elasticity and cost efficiency, GPUs can be rented to accelerate Al workload requirements and / or other computation requirements associated with the immersion cooling system and / or tanks. Under fixed coupling scheme, server and system component functionalities can affect the availability of other resources and / or nodes. System components such as GPU and CPUs can be managed, scaled and upgraded independently, optimizing infrastructure expenses. In systems, including immersion fluid systems, once a threshold data rate has been reached, signal distortion is probable, leading to error rates impacting communication performance.

[0005] Disaggregated systems using a PCIe interface provide solutions for flexible node connectivity in an efficient manner. PCIe interface can further serve as an efficient solution in systems implementing artificial intelligence tasks, machine learning models, including but not limited to, training Large Language Models (LLMs). A disaggregated infrastructure with PCIeinterface can perform high-level computations with end-to-end reliability, and efficiency through latency and bandwidth. Utilizing high-performance network communications interfaces provides the link between system components using disaggregated set of nodes. The requirements associated with systems incorporating Al and cloud computing services rely on large scale computing in immersion cooling systems, including CPUs, GPUs and / or tanks. The present invention provides a disaggregated Al server design for immersion by using a flexible node connectivity over PCIe link.SUMMARY

[0006] The present technology relates generally to using flexible node connectivity over PCIe link for using compute nodes through different components of a system, such as an immersion cooling system.

[0007] Aspects of the present technology provide flexible node connectivity over PCIe links, with compute nodes and GPU nodes disaggregated across system components such as immersion tanks. A disaggregated infrastructure can be implemented at the software level (user-mode library, kernel-mode driver) and hardware level (PCIe level). Additionally, the design of system components, including but not limited to Al servers, allows for efficient “mix and match” of components such as, GPUs communicating with the other system components. Flexible node connectivity over PCIe links serves well to implement Al server designs which are disaggregated. Such disaggregated systems can accelerate Al operations when using components which are immersed.

[0008] Aspects of the present technology are directed to disaggregating compute nodes in different tanks and making use of flexible node connectivity over PCIe link for efficient computing in an immersed cooling environment.

[0009] Aspects of the present technology are directed to disaggregating compute nodes and GPUs separately in different tanks in an immersed cooling environment. The disaggregated compute nodes rely on a robust and efficient PCIe interface to interface, communicating in a flexible and efficient manner. This overcomes the disadvantages of existing Al server designs in which fixed compute nodes are paired (i.e. paired physically) with computing components, for example, GPUs.

[0010] In some aspects, the compute nodes are separated with for example, computer CPUs in one tank of an immersed cooling environment and peripheral devices, including but not limited to accelerators, GPUs, storage devices in another tank of the immersed cooling environment. In some aspects, the compute nodes can include but are not limited to compute servers (i.e. CPU) and accelerator servers, disaggregated for efficient computing in an immersed environment.

[0011] In some aspects, the immersed cooling environment includes a compute node tank, including compute nodes which can then be coupled to GPUs in another tank for implementing functions, including but not limited to, accelerating the implementation of complex Al related tasks. Based on the needs of an application, and the requirements related to the devices used, an efficient communication path and usage of the nodes is configured.

[0012] In some aspects, the disaggregated design takes into consideration that if one device fails, the remaining devices connected to the system can continue to function. For example, if one of the GPUs in the GPU tank fails, the server or CPU nodes in the compute tank can continue to function and may not require a reboot.

[0013] In some aspects, the immersed cooling environment includes multiple (i.e., more than two) tanks containing disaggregating nodes.

[0014] In some aspects, the disaggregated nodes, including but not limited to, CPU and GPUs nodes can function flexibly with the GPU functioning to carry out a separate task, leaving the CPUs to implement a different task. Additionally, in some aspects, an optimal interface for communication between these nodes is necessary, in the interface, including but not limited to, a PCIe interface. In some aspects, there is a compute tank which includes at least two CPU nodes, connected to storage devices, and a network interface controller (NIC).

[0015] In some aspects, the CPU nodes are coupled to an optical switch for connecting to the GPU nodes in the GPU tank. In some aspects, the disaggregated Al server design carries out Al operations including but not limited to, training of Al models and predicting using trained Al models. Additionally, in some aspects, the disaggregated Al server design includes CPUs, and GPUs to serve as accelerators to carry out various Al processes.

[0016] In some aspects, the techniques described herein relate to an apparatus for monitoring immersion cooling fluid. The system includes a controller and a sensor operably connected to the controller. In some aspects, the compute nodes, including but not limited to, the CPU nodes and the GPU nodes may function independently of each other as needed by a particular Al operation, and / or based on the state of the immersed cooling environment.

[0017] In some aspects, the techniques described include the CPU nodes receiving one or more Al operations. These inputs for the Al operations include inputs to one or more elements of the disaggregated GPU nodes and can include but are not limited to, one or more trained Al model, one or more Al model training parameters, Al model training data, inputs to be run through a trained Al model.

[0018] In some aspects, the techniques described herein applies to computing environments, including server systems and general computing environments not related to cooling systems. As such, the disaggregated design, described herein, can be implemented in generic computing systems with or without server components.

[0019] In some aspects, the techniques described herein relate to a system, including: a compute tank, including at least one CPU, at least one Electrical-to-Optical (EO) converter, and a first set of computing nodes configured to execute compute functions; a GPU tank including a second set of computing nodes configured to execute at least one accelerator function; and an optical switch coupled to the compute tank and to the GPU tank; wherein: the at least one EO converter is coupled to the optical switch; the at least one CPU is coupled to the at least one EO converter; and the second set of computing nodes are configured to execute the at least one accelerator function, based at least in part on a request from an application and a state of the first set of computing nodes.

[0020] In some aspects, the techniques described herein relate to a method for performing tasks with disaggregating computing nodes in an Al server, the Al server including: a compute tank including a second CPU and at least one EO converter; a first CPU; a GPU tank including an OE converter; and an optical switch coupled to the compute tank, the first CPU, and the GPU tank; the method including: receiving, by at least the first CPU, a request to execute a command from a user application; receiving, by the optical switch, a signal from the compute tank, the signal generated in response to the request; and completing a task associated with the request, based at least in part on the signal generated from the compute tank.

[0021] All combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are part of the inventive subject matter disclosed herein. The terminology used herein that also may appear in any disclosure incorporated by reference should be accorded a meaning most consistent with the particular concepts disclosed herein.BRIEF DESCRIPTIONS OF THE DRAWINGS

[0022] The skilled artisan will understand that the drawings primarily are for illustrative purposes and are not intended to limit the scope of the inventive subject matter described herein. The drawings are not necessarily to scale; in some instances, various aspects of the inventive subject matter disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like reference characters generally refer to like features (e.g., functionally similar and / or structurally similar elements).

[0023] FIG. 1 illustrates an immersion-cooled system for computation in accordance with the present technology.

[0024] FIG. 2 illustrates an example of an immersion cooling environment in accordance with the present technology.

[0025] FIG. 3 depicts aspects of an immersion cooling system.DETAILED DESCRIPTION

[0026] The present technology is directed toward a disaggregation technique to decouple computing components within an immersion cooling environment. Computing components, for example, can be aggregated into pools of devices and allocated based on system demands (i.e., user application requests). This present technology relies on flexibility and cost efficiency by using disaggregated nodes through a flexible node connectivity mechanism. Clusters of nodes are assessed to determine how they may be used to satisfy varying system requirements.

[0027] FIG. 1 illustrates an immersion-cooled system 100 for computation in accordance with the present technology. System 100 may be a disaggregated server architecture within an immersion- cooled computing system, where some or all of the components of system 100 are immersed in immersion cooling fluid. System 100 includes but is not limited to a compute tank 112 and a GPU tank 113 connected by at least one optical switch 114.

[0028] The CPU tank 112 can include storage 101, booting device 102, CPUs 103 and 104, NIC 106 (e.g., Network Card). The storage 101 and booting devices 102 are directly connected to the processing circuits (e.g., CPUs 103 and 104). The aggregated system design includes a GPU tank 113 with GPUs 108-111. The CPUs 103 and 104 are connected to a network card 106. The server may include additional processing resources (e.g., processors) and memory devices (e.g., system memory). In some cases, system components within the server may access system components of another server, and an aggregated system may be advantageous for this access to occur while minimizing the bandwidth or access.

[0029] The description sets forth the features of the present disclosure in connection with the illustrated embodiments. It is to be understood, that the same or equivalent functions and structures may be accomplished by different embodiments that are also intended to be encompassed within the scope of the disclosure. As denoted elsewhere herein, like element numbers are intended to indicate like elements or features.

[0030] System 100 discloses an example of an aggregated system architecture with a compute tank 112 connected to the GPU tank 113 through optical switch 114. Each of the compute tank 112 and GPU tank 113 contain Electrical-to-Optical (EO) converters 105 which are connected tothe optical switch 114. The EO converters 105 receive electrical inputs from CPUs 103 and 104, and convert these inputs into optical signals.

[0031] System 100 may also be described as a package (e.g., a package including a printed circuit board and components connected to it, or an enclosure including a printed circuit board) including the memory dies (e.g., storage 101), each storage 101 including a plurality of memory locations or cells. Storage 101 may be in a package connected to the printed circuit board. The components or nodes within the compute tank 112 may communicate with the nodes of the GPU tank 113 through the optical switch 114. The CPUs 103 and 104 communicate with the other computing components of the compute tank 112 through PCIe (Peripheral Component Interconnect Express) connections. Such a system design allows for the computing components of the compute tank 112 to be aggregated from the GPU tank computing components. The two tanks (computing components of the tank) communicate through the optical switch 114. The PCIe links of the compute tank 112 may be based on x4, x8 and xl6 slots.

[0032] In some embodiments, the disaggregated server architecture can include computing components such as, microprocessor (e.g., a central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), data processing unit (DPU), and the like), a digital signal processing (DSP) die, an artificial intelligence (Al) accelerator, an application- specific integrated circuit (ASIC), field-programmable gate array (FPGA), programmable logic controller (PEC), or any suitable processing circuitry. Examples of one or more memory modules that may be included in the tanks 112 and 113 include a dynamic random access memory (DRAM) module, a dual inline memory module (DIMM), a static random access memory (SRAM) module, a flash memory module, a solid-state drive (SSD), a non-volatile random access memory (NVRAM) module, a read-only memory (ROM) module (such as a floating-gate ROM, erasable programmable readonly memory (EPROM), electrically erasable programmable read-only memory (EEPROM), onetime programmable ROM (OTPROM), or the like), or any suitable type of memory module.

[0033] The aggregated architecture may include additional computing components, including multiple servers. Additionally, each of the EO converters 105 from the compute tank 112 can communicate with the Optical-to-Electrical (OE) converters 107 and GPUs 108-111 of the GPU tank 113. The OE converters 107 accept optical signals generated by the EO converters 105 and conducts optical-to-electrical conversion. This design schema allows for flexibility with using at least the compute tank node or computing component as standalone nodes for carrying out additional or separate functionality from those of the tasks carried out with the computing nodes in the GPU tank 113.

[0034] In some embodiments, the compute tank 112 can serve to carry out functions associated with a CPU architecture, with the compute tank 112 and GPU tank 113 together serving as an accelerator to implement and execute computation heavy tasks (i.e., machine learning tasks). The disaggregated design provides additional options for compute tank nodes and GPU tank nodes to communicate with each other to carry out tasks. For example, in one configuration, CPU 103 may work with the GPUs 108-111 to carry out some functionality. The aggregated nodes allows for flexibility of different combination of nodes to work with each other while also enabling independent functionality of the nodes to execute functions and work with other aggregated nodes in different tanks in the immersed cooling environment.

[0035] The system 100 allows for the GPU tank computing nodes to be disaggregated from the compute tank nodes and allocates the GPU nodes based on user, application and / or task specific functions. Such flexible allocation, in particular for Al specific workloads, can contribute to decreasing the power consumption for the system. Thus, the disaggregated tanks and / or specific nodes can be rented for tasks and returned to carry out its main functionality. For example, the GPU tank 113 serves a cluster of GPUs 108-111, which can be accessed to meet certain user and or system requirements. Meanwhile, the same cluster of GPUs 108-111, can also work with the compute tank 113 to carry out computation heavy tasks.

[0036] Although FIG 1. includes 4 GPUs, the number of GPUs can vary and can be adjusted based on how many can fit into any specific tank. Additionally, various embodiments disclose a system architecture include a plurality of tanks, including a hierarchy of tanks with at least more than two tanks. Additionally, each of the tanks can include a number of nodes, for example, in some scenarios, one GPU tank may include 500 GPU nodes.

[0037] User requests can vary and require a diverse set of resource requirements. Different workloads, in particular with reference to Al-related tasks, can result in large overheads when working with a static design of host server and GPU accelerators connected to each other. For example, in some cases, failure to the host server (compute tank 112) at any single point (at least one computing node failure) can result in the entire server being shut down, with the computing components becoming unavailable and the server requiring a reboot. In addition to failures, system updates also result in requiring a shut-down of the server in the fixed coupling scheme. As such, the costs can be high when maintaining the availability and functionality of these system components in a fixed or static coupling.

[0038] To address issues as discussed above, a disaggregated server design has been described herein which aims to physically decouple compute nodes (i.e., host devices, compute tank 112) from accelerator devices (GPUs 108-111), grouping them into a tank 114. In this example, theGPU workloads based on requests from applications on the host server are directed to allocated node(s) in the GPU tank 113. The allocation process is dynamic and allows for reclamation of available GPUs or available compute tank nodes for additional tasks. In some cases, only the compute tank nodes are used to meet a workload requirement. In other cases, the GPU tank nodes are relied upon to meet a workload requirement. Thus, the compute tank nodes and the GPU tank nodes can be managed, scaled and upgraded independently, optimizing infrastructure expenses.

[0039] PCIe level configurations can be used efficiently with the dynamic allocation of the nodes, where determining the most efficient path through a PCIe link can provide a dynamic performance forward solution to configuring the nodes in a disaggregated system. The optical switch 114 facilitates the interaction between the compute tank 112 and the GPU tank 113. This design of resource disaggregation involves decoupling of the resources and grouping into at least one additional tank (the architecture can include more than two tanks). GPU workloads are redirected to disaggregated GPUs with the scope of the disaggregation allowing for computing nodes (GPU nodes) to be allocated to compute tank nodes and other tanks within an immersion cooling environment (i.e., data center). Although the GPU tank of FIG. 1 discloses four GPU nodes, each of the tanks in the data center of the present technology can accommodate at least hundreds of GPU nodes (e.g., 128, 256, or 512).

[0040] The example of a disaggregated architecture design, as shown in FIG. 1 can be implemented with interconnected sleds, including but not limited to, compute sleds of processors, memories, storages, network interfaces, and other computing components. High speed interconnects can be used such as PCIe, and / or optical interconnects (or a combination thereof). The sled configuration within the racks (i.e., as depicted in FIG. 2) can be designed according to a disaggregated design policy which separates the main architectural components into pluggable components designated by sleds or racks. For examples, the nodes of a sled (or plurality of sleds) or a rack (or a plurality of racks) can be based on functionality, including but not limited to a processing component, a memory component, a storage component and / or an accelerator.

[0041] FIG. 2 depicts an example of an immersion cooling environment. Various embodiments, and configurations of a disaggregated architecture design may be used with this immersion cooling environment. As shown in FIG. 2, the immersion cooling environment 200 may include an optical switch 203. The optical switch 203 may generally include a combination of optical signaling media, including but not limited to optical cabling and optical switching devices used by at least one particular sled in the immersion cooling environment 200 to send or receive signals to another sled in the immersion cooling environment 200. Although the configuration of FIG. 2 displays three racks, the immersion cooling environment can include additional or less racks or sleds withinthese racks. In this example, the immersion cooling environment 200 includes three racks 201, 202 and 206, with each of these racks 201, 202 and 206 respectively housing sleds (204-207 and 209-210). In some examples, the design architecture of FIG 1. may be implemented with at least two of the racks from 202, 206 and 201 stored in tanks 112 and 113. For example, the compute tank 112 may store rack 202 and the GPU tank 113 may store rack 201 with the respective sleds within these racks storing the nodes (101-106 and 107-111).

[0042] FIG. 3 depicts aspects of an immersion cooling system 300 for dissipating heat from one or more heat-generating components such as semiconductor die packages 305 via immersion cooling. Each package 305 can include one or more semiconductor dies that produce heat when the system is in operation. The immersion cooling system 300 in the illustrated example of FIG. 3 is a two-phase immersion cooling system, though the invention may also be implemented in a single-phase immersion cooling system. Immersion cooling system 300 may include one or more computing systems such as system 100.

[0043] Immersion cooling system 300 includes a container such as tank 320 filled, at least in part, with immersion cooling liquid 364. The immersion cooling system 300 can further include at least one chiller 380 that flows a heat-transfer fluid through at least one condenser coil 370 that is disposed in the tank 320 and headspace 308. Condenser coil 370 and chiller 380 may be part of a heat exchanger. The packages 305 can be mounted on one or more printed circuit boards (PCBs) 357 that are immersed, at least in part, in the immersion cooling liquid 364. Immersion-cooling system 300 may further include a filter 375 disposed adjacent to the tank 320.

[0044] Filter 375 may include a filtration media, a housing, and a pump configured to force immersion cooling liquid 364 through filter 375 to remove contaminants, particulates, or other impurities that may be added to immersion cooling liquid 364 during use. Filter 375 may be housed outside of tank 320 while being in fluidic communication with immersion cooling liquid 364 in tank 320. Alternatively, filter 375 may be submerged within immersion cooling liquid 364 inside of tank 320.

[0045] Immersion cooling liquid 364 may be a hydrocarbon, a fluoroketone, an oil, or a similar dielectric liquid that will act as an insulator while simultaneously transferring heat from package 305 more efficiently than air. An example of immersion cooling liquid 364 is Novec™ 649 produced by 3M™. An exemplary immersion cooling liquid 364 used in accordance with embodiments of the present invention may have a dielectric constant baseline value of about 1.8- 2 at a frequency of about 1 kHz.

[0046] In an embodiment of the invention, immersion cooling liquid 364 may be considered unacceptably contaminated if the dielectric constant and / or dielectric loss tangent of immersioncooling fluid being used in an immersion cooling system 300 differs by a threshold amount as compared to unused or pure immersion cooling liquid 364. For example, immersion cooling liquid 364 may be considered unacceptably contaminated or degraded if the dielectric constant and / or dielectric loss tangent differs by a threshold of 10% or more as compared to unused or pure immersion cooling liquid 364. In an embodiment, a dielectric constant and / or dielectric loss tangent variation threshold may be 20%, 15%, 5%, 3%, 1%, or any suitable threshold.

[0047] Contamination of the immersion cooling liquid 364 and resulting changes to dielectric constant and / or dielectric loss tangent may alter or negatively impact operation of components within immersion cooling liquid 364 including semiconductor die(s) 350. An altered dielectric constant and / or dielectric loss tangent may result in undesirable cross-talk between components on a PCB, additional noise or reduction in signal strength transmitted along exposed wires of a PCB or semiconductor die(s) 350 submerged in immersion fluid, and / or signal dissipation through the immersion cooling liquid 364. Signal loss may be severe enough that two elements may be effectively represented as being separated by an open circuit despite being physically connected. In an embodiment, a dielectric constant and / or dielectric loss tangent variation threshold may be selected based on an observed or inferred effect on one or more submerged semiconductor die(s) 350. For example, an increase in PCIe bit error rate above an error rate baseline may be correlated with an increase in dielectric constant and / or dielectric loss tangent above a dielectric constant and / or dielectric loss tangent baseline. Accordingly, operation of semiconductor die(s) 350 may be throttled or suspended when a dielectric constant and / or dielectric loss tangent of immersion cooling liquid 364 exceeds a predetermined threshold.

[0048] Changes to dielectric constant and / or dielectric loss tangent may be caused by contaminants within immersion cooling liquid 364. In some cases, changes to dielectric constant and / or dielectric loss tangent may be reversed by filtering the contaminants from immersion cooling liquid 364. In some embodiments, upon detecting an increase in dielectric constant and / or dielectric loss tangent of immersion cooling liquid 364, controller 302 may instruct filter 375 to increase filtration throughput or notify a user that an immersion cooling liquid 364 filtration media may need to be replaced. If a dielectric constant and / or dielectric loss tangent exceeds a predetermined threshold, controller 302 may throttle or shut down one or more semiconductor die(s) 350, generate a notification that immersion cooling liquid 364 should be replaced, trigger an alarm, etc.

[0049] The illustrated example of FIG. 3 is not intended to be to scale. The immersion cooling system 300 may house and provide immersion cooling liquid 364 to tens, hundreds, or even thousands of packages 305. In some cases, the immersion cooling system 300 can be small (e.g.,the size of a floor unit air conditioner, approximately 1 meter high, 0.5 meter width, 0.5 meter depth or length). In some implementations, the immersion cooling system can be large (e.g., the size of a van or larger, approximately 2.5 meters high, 2.5 meters width, 4 meters depth or length).

[0050] The immersion cooling system 300 can also include a controller 302 (e.g., a microcontroller, programmable logic controller (PLC), microprocessor, field-programmable gate array, logic circuitry, memory, or some combination thereof) to manage system operation. Controller 302 can perform various system functions such as monitoring temperatures of system components, cooling fluid level, tank access, chiller operation etc. The controller 302 can further issue commands to control system operation such as executing a start-up sequence, executing a shut-down sequence, assigning workloads among the packages, changing cooling fluid level, changing the temperature of the heat-transfer fluid circulated by the chiller 380, etc. In some implementations, controller 302 can include (or itself be) a baseboard management controller (BMC) 304. That is, the BMC 304 may monitor and control all aspects of system operation for the immersion cooling system 300 in addition to monitoring and controlling workloads of the semiconductor dies 350 in the packages 305 cooled by the system. The immersion cooling system 300 can also include a network interface controller (NIC 303) to allow the system to communicate over a network, such as a local area network or wide area network. The immersion cooling system 300 can further include a fluid sensor array 390 having a plurality of fluid sensors 310. Fluid sensors 310 may include one or more leak detection sensors at least partially submerged in immersion cooling liquid 364.

[0051] The semiconductor die(s) 350 and can be mounted on and attached to a printed circuit board (PCB) 355 (sometimes referred to as a substrate) in device package 305. The package 305 can be made commercially available as an off-the-shelf (OTS) product. The package 305 can be used for single-phase or two-phase immersion cooling of at least one semiconductor die 350, such as a microprocessor (e.g., a central processing unit (CPU) and / or graphics processing unit (GPU)), voltage regulator (VR), high bandwidth memory (HBM), a digital signal processing (DSP) die, an artificial intelligence (Al) accelerator, an application- specific integrated circuit (ASIC), field- programmable gate array (FPGA), and / or other densely patterned semiconductor die.

[0052] In the two-phase immersion cooling system 300 of FIG. 3, heat flows from the semiconductor die 350 where it is generated into the heat spreader 352. The heat spreader 352 is in thermal contact with an immersion cooling liquid 364 that can flow over and extract heat from the heat spreader 352. The amount of heat delivered by the heat spreader 352 to the immersion cooling liquid 364 is enough to boil the immersion cooling liquid 364 that contacts the heat spreader 352 (creating bubbles 365 and potentially creating froth 367 when bubbles 365 reach thesurface of immersion cooling liquid 364). The vapor 366 from the boiled immersion cooling liquid 364 can be cooled and condensed back to liquid droplets 368, for example, by the condenser coil 370. The heat-transfer fluid, such as chilled water, from the chiller 380 can be circulated through the condenser coil 370 to lower the temperature of the condenser coil 370 below the condensation point in the headspace 308 of the tank 320. As a result, vapor 366 condenses on exterior surfaces of the condenser coil 370 and liquid droplets 368 from the condensed vapor can drip and / or flow back to the immersion cooling liquid 364. Although a single condenser coil 370 is depicted in FIG. 3, there can be a plurality of condenser coils 370 in tank 320 to condense the vapor 366 into droplets. Some or all of the condenser coils 370 may or may not be located directly over the PCBs 357. Instead, the condenser coil(s) 370 can be located near one or more walls of the tank 320, such that the condenser coil(s) 370 are not directly over the PCBs 357 on which the packages 305 are mounted.

[0053] To improve thermal performance in two-phase immersion cooling system 300, the heat spreader 352 can include a boiling enhancement coating (BEC) on at least one surface. The BEC can be formed from copper or a copper alloy and can be porous, for example, though BECs can take various forms. In some cases, the BEC is a micro porous copper coating having a thickness from approximately or exactly 50 microns to 500 microns thick (which may be produced by electroplating and / or etching). In some implementations, the BEC comprises a mesh copper layer bonded (e.g., via resistance heating) to at least an outer surface of the heat spreader 352. In some cases, the BEC is applied as particulates to at least one smooth surface of the heat spreader 352 and then subsequently sintered to adhere to one another and to the heat spreader 352. The BEC provides an improved surface area to contact the immersion cooling liquid 364 and can increase the heat transfer coefficient from the heat spreader 352 to the immersion cooling liquid 364 by up to a factor of 15 versus a smooth surface on the heat spreader 352. Accordingly, BECs can increase thermal conductivity to, and accelerate the boiling of, the immersion cooling liquid 364.

[0054] Further implementations of boiling enhancement coatings and enclosures are possible. Additional arrangements, applications, and methods of use of boiling enhancement coatings and enclosures, including with semiconductor dies and 3DIC stacks, are described in the below U.S. Patent Applications.

[0055] U.S. Patent Application No. 18 / 327,615, filed June 1, 2023 and entitled "Boiler Enhancement Coatings with Active Boiling Management,” discloses heat spreader and boiling enhancement enclosure architectures thermally and / or mechanically coupled to one or more semiconductor dies or logic ICs that may be used for passive and / or active management of immersion cooling fluid boiling, including through the use of valves to control pressure of boilingimmersion cooling fluid within a boiling enhancement chamber, particularly in paragraphs

[0018] -

[0039] and FIGS. 3-5B. The entirety of U.S. Patent Application No. 18 / 327,615 is incorporated herein by reference.

[0056] U.S. Provisional Patent Application No. 63 / 500,167, filed May 4, 2023 and entitled “Direct to Chip Heat Spreader and Boiler Enhancement Coatings for Microelectronics,” discloses heat spreader and BECs thermally and / or mechanically coupled to one or more semiconductor dies, logic ICs, and / or 3DIC stacks, particularly in paragraphs

[0015] -

[0033] and FIGS. 2A-4. BEC form factors may include graphite heat spreader architectures, vapor chambers, heat pipes, copper plates, fins, and the like. BEC form factors may be thermally and / or mechanically coupled to the one or more semiconductor dies, logic ICs, and / or 3DIC stacks through a thermally conductive epoxy, and may have varying dimensions relative to a surface to which the semiconductor dies and / or logic ICs are mounted. The entirety of U.S. Provisional Patent Application No. 63 / 500,167 is incorporated herein by reference.

[0057] U.S. Patent Application No. 18 / 460,091, filed September 1, 2023 and entitled “Direct to Chip Application of Boiling Enhancement Coating,” discloses BECs and methods for applying BECs to semiconductor dies, logic ICs, and / or 3DIC stacks in accordance with the present technology. In particular, paragraphs

[0024] -

[0046] and FIGS. 2A-5 disclose embodiments of BEC layers, adhesives, solders, sintering, laser ablation, meshes, and other BECs and BEC application methods. The entirety of U.S. Patent Application No. 18 / 460,091 is incorporated herein by reference.

[0058] U.S. Provisional Patent Application No. 63 / 506,945, filed June 8, 2023 and entitled “Vapor-Shedding Structures for Boiler Plates in Two-Phase Immersion Cooling Systems,” discloses structures that may be thermally and / or mechanically coupled to computing hardware such as one or more semiconductor dies, logic ICs, and / or 3DIC stacks to enable the shedding of immersion cooling vapors generated from the boiling of immersion cooling fluid during operation of the computing hardware. In particular, paragraphs

[0021] -

[0039] and FIGS. 3A-5 disclose vapor- shedding structures including varying porosities, constituent materials, and geometries relative to the computing hardware on which they are mounted. The entirety of U.S. Provisional Patent Application No. 63 / 506,945 is incorporated herein by reference.

[0059] U.S. Provisional Application No. 63 / 513,828, filed July 14, 2023 and entitled “Grinding Apparatuses and Methods for Mechanically Modifying Surfaces of Processors to Promote Boiling of a Coolant Liquid,” discloses methods for creating boiling enhancement modifications to surfaces such as the surfaces of computing hardware such as one or more semiconductor dies, logic ICs, and / or 3DIC stacks, particularly in paragraphs

[0036] -

[0095] and FIGS. 2A-8. For example,grooves, patterns, gouges, trenches, or other structures may be added to a surface or lid of a processor, semiconductor die, logic IC, 3DIC stack component, and / or BEC to encourage nucleation sites for bubbles of immersion cooling vapor to form during a cooling process, thus decreasing the thermal resistance between the processor, semiconductor die, logic IC, and / or 3DIC stack component and the surrounding immersion cooling fluid. The entirety of U.S. Provisional Application No. 63 / 513,828 is incorporated herein by reference.

[0060] U.S. Provisional Patent Application No. 63 / 513,829, filed July 14, 2023 and entitled “Electrical Connector Having a Heater to Facilitate Boiling of a Coolant Liquid to Improve Signal Integrity in Immersion Cooling Environment,” discloses heaters for promoting boiling of immersion cooling fluid near electrical connectors such as connections between components of a 3DIC stack and enable improved impedances at those connectors, particularly in paragraphs

[0019] -

[0052] and FIGS. 1A-3B. The entirety of U.S. Provisional Patent Application No. 63 / 513,829 is incorporated herein by reference.

[0061] U.S. Provisional Patent Application No. 63 / 603,242, filed November 28, 2023 and entitled “Woven Boiler Enhancement Coatings,” provides additional examples of BECs including woven BECs with variable weave patterns, densities, attachment mechanisms, and materials (including copper and tungsten) that may be attached to computing hardware such as one or more semiconductor dies, logic ICs, and / or 3DIC stacks in order to promote more efficient heat transfer and immersion cooling vapor nucleation, particularly in paragraphs

[0031] -

[0055] and FIGS. 3-7. The entirety of U.S. Provisional Patent Application No. 63 / 603,242 is incorporated herein by reference.Conclusion

[0062] While various inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the inventive teachings is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodimentsmay be practiced otherwise than as specifically described and claimed. Inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the inventive scope of the present disclosure.

[0063] Also, various inventive concepts may be embodied as one or more methods, of which an example has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.

[0064] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0065] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”

[0066] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.

[0067] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicatingexclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.

[0068] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.

[0069] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.

Claims

CLAIMS1. A system, comprising: a compute tank, comprising at least one CPU, at least one Electrical-to-Optical (EO) converter, and a first set of computing nodes configured to execute compute functions; a GPU tank comprising a second set of computing nodes configured to execute at least one accelerator function; and an optical switch coupled to the compute tank and to the GPU tank; wherein: the at least one EO converter is coupled to the optical switch; the at least one CPU is coupled to the at least one EO converter; and the second set of computing nodes are configured to execute the at least one accelerator function, based at least in part on a request from an application and a state of the first set of computing nodes.

2. The system of claim 1, further comprising: a PCIe link used for communicating between the first set of computing nodes or the second set of computing nodes.

3. The system of claim 1, wherein the GPU tank is configured to perform the at least one accelerator function using the compute tank, based at least in part on a signal received from the at least one EO converter; and the at least one CPU is configured to, using the optical switch, execute the at least one accelerator function using at least one GPU in the GPU tank.

4. The system of claim 1 , wherein at least one CPU is configured to perform Al tasks using the GPU tank.

5. The system of claim 1, wherein the second set of computing nodes comprises at least one Optical-to-Electrical (OE) converter and at least one GPU, and wherein the OE converter is coupled to the optical switch and the at least one GPU.

6. The system of claim 1, further comprising: receiving by the optical switch a signal from the compute tank, the signal generated in response to the request from the application; andcompleting a task associated with the request, based at least in part on the signal generated from the compute tank.

7. The system of claim 1, wherein the at least one CPU generates a signal for transmission to the at least one EO converter.

8. The system of claim 1, wherein the compute tank, the GPU tank and the optical switch are located in an immersion cooling environment.

9. The system of claim 1, wherein at least one of the compute tank, the GPU tank and the optical switch are immersed in immersion cooling liquid.

10. A method for performing tasks with disaggregating computing nodes in an Al server, the Al server comprising: a compute tank comprising a second CPU and at least one EO converter; a first CPU; a GPU tank comprising an OE converter; and an optical switch coupled to the compute tank, the first CPU, and the GPU tank; the method comprising: receiving, by at least the first CPU, a request to execute a command from a user application; receiving, by the optical switch, a signal from the compute tank, the signal generated in response to the request; and completing a task associated with the request, based at least in part on the signal generated from the compute tank.

11. The method of claim 10, further comprising: generating, by at least the second CPU, a signal for transmission to the at least one EO converter.

12. The method of claim 10, further comprising: determining by the optical switch, a status of the GPU tank.

Citation Information

Patent Citations

  • Server water cooling system

    CN113985979A

  • Systems and methods for interconnecting GPU accelerated compute nodes of an information handling system

    US20190042512A1

  • Device and method for accelerating graphics processor units, and computer readable storage medium

    US20200242724A1

  • Immersion cooling system with coolant boiling point reduction for increased cooling capacity

    US20210410320A1

  • Computer system and method for sharing computer memory

    US9582462B2