Cooling device around a central region

CN122776948APending Publication Date: 2026-09-18MELLANOX TECHNOLOGIES LTD(IL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610318811.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-16
Filing Date
2026-03-16
Publication Date
2026-09-18

Smart Images

  • Figure CN122776948A_ABST
    Figure CN122776948A_ABST
Patent Text Reader

Abstract

This disclosure relates to a cooling device surrounding a central region. A system for circulating coolant along an annular path surrounding the central region of the device is also provided. The device can receive cooling fluid through a manifold, circulate the cooling fluid along an annular path surrounding the central region, introduce the cooling fluid into the central region, and finally discharge the cooling fluid.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Cooling systems are commonly used in hardware environments to dissipate heat from one or more heat-generating components, such as CPUs, GPUs, and other processors. Some cooling systems use cooling loops to circulate coolant close to these heat-generating components. However, due to space constraints, some of these components can be difficult to access. For example, placing cooling components (such as cold plates) and cooling loops near heat-generating components may require multiple setups, which can lead to inefficiencies, difficult maintenance, and time-consuming operation. Attached Figure Description

[0002] Various embodiments according to this disclosure will be described with reference to the accompanying drawings, in which: Figure 1 The illustration depicts a system according to an example embodiment for dissipating heat from one or more heating elements via a cooling circuit; Figure 2A The illustration shows an apparatus according to an example embodiment including one or more cooling circuits for dissipating heat from one or more heating elements; Figure 2B The illustration shows an apparatus according to an example embodiment including one or more cooling circuits for dissipating heat from one or more heating elements; Figure 2C The illustration shows one or more insertion elements for engaging the device with a heating element according to an example embodiment; Figure 2D The illustration shows a bottom perspective view of an apparatus including one or more cooling circuits according to an example embodiment, the cooling circuits being used to dissipate heat from one or more heating elements; Figure 3A The illustration shows a top cross-sectional perspective view of an apparatus including a first cooling circuit according to an example embodiment, the first cooling circuit being used to dissipate heat from one or more heating elements; Figure 3B The illustration shows a top cross-sectional perspective view of an apparatus according to an example embodiment, including a first cooling circuit and a second cooling circuit, the first cooling circuit and the second cooling circuit being used to dissipate heat from one or more heating elements; Figure 3C The illustration shows a top cross-sectional perspective view of an apparatus including a cooling circuit according to an example embodiment, the cooling circuit being used to dissipate heat from one or more heat-generating elements; Figure 4 The illustration shows a system according to an example embodiment for generating temperature information from one or more heating elements and adjusting the flow rate of a cooling circuit; Figure 5AThe illustration depicts a process for dissipating heat from one or more heating elements according to an example embodiment; Figure 5B The illustration shows a process for generating temperature information from one or more heating elements and adjusting the flow rate of a cooling circuit according to an example embodiment. Figure 6 The illustration shows components of a distributed system that can be used in a data center according to an example embodiment; Figure 7 An example data center system according to an example embodiment is illustrated; Figure 8 The illustration depicts an example computing environment according to an example embodiment; Figure 9 The illustration depicts a computer system according to an example embodiment; Figure 10 The illustration depicts a computing system according to an example embodiment; Figure 11 The illustration shows an optoelectronic component according to an example embodiment; Figure 12 The illustration shows an optoelectronic component according to an example embodiment; Figure 13 The illustration shows an optoelectronic component having a substrate and an integrated circuit according to an example embodiment; Figure 14 The illustration shows an optoelectronic component in a transceiver according to an example embodiment; and Figure 15 The illustration shows a network device according to an example embodiment. Detailed Implementation

[0003] In the following description, various embodiments will be described. Specific configurations and details are set forth for ease of explanation, to provide a thorough understanding of these embodiments. However, it will be apparent to those skilled in the art that these embodiments can be practiced without these specific details. Furthermore, some well-known features may be omitted or simplified to avoid obscuring the described embodiments.

[0004] In traditional cooling environments, cooling loops and cold plates are typically used to dissipate heat from heat-generating components. However, due to the physical constraints of hardware and data center environments, the space around these components is often limited. Because of these limitations, some cooling systems do not allow cold plates with multiple thermally coupled components to be thermally connected to heat-generating components. As a non-limiting example, such an environment would not allow a cold plate with multiple thermally coupled components to be located around heat-generating components such as processors.

[0005] This disclosure relates to an apparatus having one or more cooling circuits and cooling elements, the cooling circuits being used to circulate coolant close to one or more heat-generating elements. In some embodiments, the apparatus is a cold plate, with annular flow of liquid surrounding the heat-generating elements for cooling, and then moving into the interior of the cold plate. The liquid path is annular and enters the interior of the main silicon. The annular flow increases the liquid volume, thereby allowing for more efficient cooling of the components. The annular flow ensures that all components maintain a relatively uniform temperature difference, thereby improving cooling efficiency. This configuration overcomes space constraints, particularly the limited space where separate cold plates are not permitted. Space constraints can be overcome by using different thermal interface materials, such as phase change materials for the main silicon and thermal pads for other components.

[0006] The device may include two regions: a central region and an annular region surrounding the central region. The annular region may be adjacent to one or more heating elements (e.g., directly above or otherwise physically close to the heating elements) or a thermal pad in thermal contact with the heating elements. The central region may also be adjacent to one or more heating elements and may include cooling elements, such as one or more cold plates. The annular region and the central region may be separated by walls or other partitions. A first cooling circuit may circulate coolant through the annular region to dissipate heat from the heating elements. One or more heating elements may be positioned around the periphery of the central region.

[0007] In some embodiments, the second cooling circuit may receive coolant from the first cooling circuit and circulate the coolant into the central region and near the cooling elements (e.g., cold plates). The coolant can be discharged from the device into a manifold or cooling distribution unit. By incorporating two distinct regions and cooling circuits into a single device, the device simplifies the cooling environment and reduces downtime for assembling cooling systems. The second cooling circuit may be located within the central region.

[0008] In other embodiments, the annular region may also be referred to as a surrounding region, circumferential region, ring-shaped region, peripheral region, and enclosing region, or any similar description, indicating that the annular region surrounds the central region of the device. Furthermore, the shape of the region surrounding the central region may be non-circular, such as a general square, rectangle, ellipse, triangle, or other similar shape, to allow coolant to circulate around the central region of the device.

[0009] Figure 1The illustration depicts a system 100 for circulating coolant into a device 190, which includes at least a first cooling circuit 144 located in an annular region 140 surrounding a central region 150 of the device 190. The system 100 can be used in various cooling environments, such as data centers, server racks, high-performance computing environments, or any other environment that may store or use hardware or heat-generating components. The system 100 may include a coolant distribution unit (CDU) 110, a cooling manifold 120, an inlet 130, the device 190, an outlet 160, and management equipment 170. The CDU 110 can circulate coolant 115 into the cooling manifold 120 and into the inlet 130, which is operatively connected or removably attached to the device 190. The device 190 may include at least a first cooling circuit 144 and a second cooling circuit 152 for circulating coolant 115 through the annular region 140 and the central region 150. The coolant 115 may circulate in a circulating or continuous flow pattern near one or more heat-generating components. The coolant 115 can then be discharged through outlet 160 and optionally returned to CDU 110. Device 190 can be operatively connected to management device 170 via a wired or wireless connection. Device 190 may include one or more temperature sensors ( Figure 1 Not shown in the image, but will be further referenced. Figure 4 and Figure 5B (To be discussed), the sensor can generate and send temperature information to the management device 170. The management device 170 can generate commands or queries based at least on this temperature information to adjust or change the flow rate of coolant 115 through cooling manifold 120 and device 190.

[0010] Typically, CDU 110 may include machinery and hardware designed to circulate coolant 115 through cooling manifold 120. Coolant 115 may include water, a water-glycol mixture, a dielectric fluid (e.g., mineral oil, synthetic fluid, ester), or other suitable fluid. CDU 110 may circulate coolant 115 to absorb or dissipate heat from the heat-generating elements 142A-142F, thereby stabilizing the temperature and preventing overheating of the heat-generating elements 142A-142F. In some example embodiments, CDU 110 may operate using a refrigeration cycle involving a compressor that compresses refrigerant gas to absorb heat from the coolant, thereby cooling the coolant 115 before it circulates back to cooling manifold 120 and device 190. Furthermore, CDU 110 may include a control system that monitors the temperature of coolant 115 and adjusts the cooling output as needed to maintain a desired temperature level.

[0011] Cooling manifold 120 may include a manifold system designed to facilitate the distribution and guidance of coolant 115 from CDU 110 to various components within system 100, including device 190. Cooling manifold 120 may serve as a central hub connecting cooling circuits to device 190, thereby enabling efficient flow and temperature regulation of coolant 115 across different areas of system 100. Although Figure 1 The illustration shows cooling manifold 120 connected to only one device 190, but it will be understood that in other example embodiments, cooling manifold 120 may be connected to multiple devices, or in fact connected to other cooling or heat-generating elements in system 100. In example embodiments, cooling manifold 120 may be made of materials such as durable plastics, aluminum or other metals or composites.

[0012] Device 190 receives coolant 115 from cooling manifold 120 through inlet 130 and discharges coolant 115 through outlet 160. Inlet 130 and outlet 160 can be operatively connected to or removably attached to device 190 via one or more connecting elements (e.g., screws, fasteners, or clips) and corresponding connection holes, thereby securing device 190 to one or more heating elements. Inlet 130 and outlet 160 may be designed to connect one or more pipes, conduits, loops, hoses, or other devices for circulating coolant 115. In some example embodiments, inlet 130 and outlet 160 may employ a quick-disconnect (QD) design, allowing pipes to be quickly attached to and removed from inlet 130 and outlet 160. In other embodiments, pipes may be secured to appropriate locations on inlet 130 and outlet 160 via screws or clips.

[0013] Device 190 can circulate coolant 115 from inlet 130 into a first cooling circuit 144, which is located in an annular region 140 surrounding a central region 150. Typically, device 190 has an annular or ring-shaped structure. Device 190 may have an outer annular region (i.e., annular region 140) surrounding a region within the outer annular region (i.e., central region 150). In other embodiments, device 190 may be concentric, having an outer concentric region (i.e., annular region 140) and an inner concentric region (i.e., central region 150). The central region 150 is separated from the annular region 140 by one or more separating elements (e.g., walls or partitions between the central region 150 and the annular region 140) (as described elsewhere herein, a second cooling circuit 152 connects the central region 150 to the annular region by passing over these walls or separating elements).

[0014] The annular region 140 allows coolant 115 to circulate through a first cooling circuit 144 near one or more heat-generating elements 142A-142F. The heat-generating elements 142A-142F may be located near the first cooling circuit 144; for example, they may be positioned directly below the device 190. In some example embodiments, the heat-generating elements 142A-142F may include thermal pads integrated into the device 190 that are in thermal contact with the CPU, GPU, or other heat-generating elements. That is, each of the heat-generating elements 142A-142F may correspond to one or more individual heat-generating elements or one or more portions of a heat-generating element. As the first cooling circuit 144 circulates coolant 115 near the heat-generating elements 142A-142F, heat is dissipated from them. The heat-generating elements 142A-142F may include, but are not limited to, processors (e.g., central processing unit (CPU) or graphics processing unit (GPU)), data processing units (DPU), quantum processing units (QPU), multiple parallel processing units (PPU), and application-specific integrated circuit (ASIC) memory modules or power supplies.

[0015] The second cooling circuit 152 can receive coolant 115 from the first cooling circuit 144 and circulate the coolant 115 into the central region 150. The second cooling circuit 152 can rise above the annular region 140 and pass over the partition element or wall between the annular region 140 and the central region 150. The central region 150 may include cooling elements 155, such as a cold plate in thermal contact with a heat-generating element. After circulating into the central region 150, the coolant 115 can be discharged through outlet 160 and optionally returned to CDU 110.

[0016] In at least one example embodiment, the cold plate may include adjustable fins that form microchannels through which fluid flows. In at least one embodiment, the fins in the cold plate allow heat to be transferred from at least one associated computing device to fluid flowing through the microchannels formed between the multiple fins. In at least one embodiment, the fins of the cold plate can be dynamically adjusted in real time to allow more heat to be transferred from at least one computing device to the fluid flowing through the finned cold plate. In at least one embodiment, such fins may be adjusted by a processor or processorless system in part based on a temperature determined (e.g., sensed) for the cold plate. In at least one embodiment, the temperature may be associated with at least one computing device, the workload of at least one computing device, or the cold plate at different time periods, and with the fluid at the inlet and outlet. In at least one embodiment, the processorless system may rely on the thermal properties of at least two materials used to form the fins of the cold plate, such that these fins can react without a processor, thereby exposing more surface area to the fluid. In at least one embodiment, such fins may include overlapping portions that can be exposed by the action of a control mechanism or by the properties of the at least two materials associated together for forming the fins.

[0017] Although Figure 1 Not shown, but device 190 may also include one or more temperature sensors configured to determine or generate temperature information regarding heating elements 142A-142F. This temperature information can be transmitted to management device 170 via a wired or wireless connection. Management device 170 can receive this temperature information and send one or more commands or queries to CDU 110 to change, alter, or adjust the flow rate of coolant 115 through cooling manifold 120 and device 190. Management device 170 may include a processor (e.g., a central processing unit (CPU) or graphics processing unit (GPU)), a data processing unit (DPU), a quantum processing unit (QPU), multiple parallel processing units (PPUs), and an application-specific integrated circuit (ASIC), a memory module, or a power supply. The QPUs are configured to perform one or more operations related to quantum algorithms. In some embodiments, each of the one or more QPUs may include multiple qubits, and the one or more QPUs may communicate with each other via quantum channels. In some embodiments, each of the multiple qubits may include local qubits, global qubits, and / or synchronization qubits. In some embodiments, the local qubits of each QPU can be configured to perform one or more operations related to quantum algorithms on the QPU associated with the local qubit.

[0018] Figure 2A The illustration shows a device 240 for circulating coolant. Device 240 may include a top portion 210 and a bottom portion 230. Although Figure 2A The illustration shows a device of a specific size, but it will be understood that the elements of the device may vary in other example embodiments. As described elsewhere in this document, in other example embodiments, the shape of device 240 may be generally square, rectangular, circular, elliptical, triangular, or any other shape suitable for the annular region 232 and the central region 233.

[0019] The bottom portion 230 of the device may include an annular region 232 surrounding a central region 233. In some embodiments, the annular region 232 itself may include a cooling circuit for circulating coolant along a path (e.g., a circulation path) surrounding the central region 233. In other embodiments, the cooling circuit may include a separate conduit or hose. The annular region 232 may be a recessed path surrounded by an inner wall 238 and an outer wall 239. The inner wall 238 separates the annular region 232 from the central region 233. The outer side of the inner wall 238 and the inner side of the outer wall 239 may define the annular region 232. The inner side of the inner wall 238 may define the central region 233. The outer side of the outer wall 239 may define the exterior of the bottom portion 230.

[0020] The circular path 232 can begin at entry point 236 and end at termination point 237. See further reference. Figure 1 , Figure 2D and Figures 3A-3B As discussed, the annular path 232 may be located close to one or more heating elements or heating pads. In some example embodiments, when the device 240 is removably attached to these heating elements, the heating elements may be located below the bottom portion 230. Within the central region 233, there may be one or more cooling elements 235, such as one or more cold plates. The cooling region 233 may also have one or more plunger connection holes 234 for connecting the device to one or more plunger elements 216, as further referenced. Figure 2C The bottom portion 230 may also include one or more mounting tabs 231A, each containing a mounting hole 231B. The mounting holes 231B allow the device to be connected to a receiving device via one or more connecting elements (e.g., screws, fasteners, or cable ties). Although... Figure 2C Only a specific number of mounting lugs are illustrated, but it will be understood that in other example embodiments, the device may include more or fewer mounting lugs 231A.

[0021] The top portion 210 may include a cap 211 for covering or spanning an annular region 232 defined by inner walls 238 and outer walls 239. The cap 211 may be removably attached to the bottom portion via one or more connecting elements. In other example embodiments, the cap 211 may snap into place or be securely mounted in a receiving groove on the bottom portion 230. The cap 211 may be annular or ring-shaped, similar in shape to the annular region 232. In example embodiments, the cap 211 may be made of metal, plastic, or a composite material. A manifold 213 may be present on or near the cap 211 for receiving coolant through an inlet and receiving element 212. The manifold 213 may be removably attached to the cap 211 via one or more connecting elements, such as screws, fasteners, snaps, or other suitable connecting elements. As shown, manifold 213 may be located near one or more sections of cover 211, such as one of the corners of cover 211; however, in other example embodiments, manifold 213 may also be located near other areas of cover 211, such as different corners of cover 211, or a portion of cover 211 that is off-corner.

[0022] Coolant can circulate from receiving element 212 into manifold 213 and enter annular region 232 at inlet point 236. Receiving element 212 can be removably attached to a cooling pipe, hose, or manifold. Once the coolant has circulated through annular region 232 and reached termination point 237, the coolant can circulate back into manifold 213 and into second cooling circuit 214, which circulates the coolant back into central region 233. Figure 2A As shown, the second cooling circuit 214 can pass over and / or bypass the inner wall 238 and outer wall 239 of the bottom 230 to reach the central region 233. Once the coolant has circulated through the central region 233, the coolant can be discharged through the outlet element 217 and optionally returned to the CDU or manifold. The outlet element 217 can be removably attached to a cooling pipe, hose, or manifold.

[0023] Figure 2B The illustration shows a device 240 comprising a top portion 210 and a bottom portion 230 operably connected together. In some example embodiments, the top portion 210 and the bottom portion 230 can be detached after being operably connected. In other example embodiments, the device 240 can be attached to a heating element or a housing near the heating element via a connection hole 231B. That is, screws or connecting elements can be passed through the connection hole 231B and through corresponding holes in the housing of the heating element. Further integration... Figure 2C-2DAs described, since the device 240 is connected to the receiving device via the connection hole 231B, the elastic mechanism 245 allows for conformability between the device 240 and the receiving device (allowing for different planes and tolerance ranges), such that when the cooling element 256 is pressed against the receiving device, the cooling element 256 can achieve conformability through the elastic mechanism 245.

[0024] Figure 2C The illustration depicts a plunger element 216 according to one or more example embodiments. Each plunger element may include a cap 221, a resilient mechanism 245, a base plate 243, a plunger element (e.g., a plunger head) 242, and a plunger housing 241. In the device 240, the plunger element 216 may be positioned between the cap 211 and the cooling element 256 such that when the device 240 is attached to a receiving device, the plunger element 216 can be compressed, allowing the cooling element 256 to move along the axis of the plunger element 216. In some example embodiments, the plunger element 216 may provide a clearance sufficient to flush with the cooling element 258.

[0025] Each plunger element 216 may include one or more caps 221. Each cap may include one or more receiving holes 222 through which one or more connecting elements 223 (e.g., screws or bolts) may pass. The connecting elements 223 may be screwed into a base plate 243, which has one or more receiving holes 244 aligned with the receiving holes 222. A plunger head 242 may engage in a plunger housing 241. Each resilient mechanism 245 may be operatively connected to the plunger head 242 such that the plunger head 242 may move within a predetermined gap within the plunger housing 241. The plunger housing 241 may be operatively connected to a cooling element 256 such that when the cooling element 256 is pushed against a receiving object (e.g., a receiving device), the resilient mechanism 245 is loaded, and the cooling element 256 conforms to the receiving object or receiving device. In some example embodiments, the resilient mechanism 245 applies a force to the cooling element 256, thereby applying a compressive force to the heating element. In some example embodiments, the heating element (or the housing of the heating element) may thermally expand due to high temperatures. When such expansion occurs, the elastic mechanism 245 can cause the cooling element 256 to conform to this expansion.

[0026] Figure 2DThe illustration shows a bottom perspective view of device 240. The bottom of device 240 may include a cooling element 256 and one or more heating elements 258. In an example embodiment, the cooling element 256 and heating element 258 will make thermal contact with one or more heat-generating elements on a receiving device. As a non-limiting example, the heating element 258 may include a thermal pad that makes thermal contact with a CPU, GPU, or other processing unit. The thermal pad may be made of a material with high thermal conductivity, such as silicone or graphite. The thermal pad can work by filling the gap between the heat-generating element and the heat sink, thereby improving heat transfer efficiency. The cooling element 256 may include a cold plate. Figure 1-2B As shown, the heating element 256 is located below or physically close to the annular region 232 having the first cooling circuit. The cooling element is located below or physically close to the central region 233. The device 240 can be operatively connected to or removably attached to one or more heating elements via mounting holes 231B.

[0027] Figure 3A The illustration shows a top cross-sectional perspective view of device 310, which has an annular region 312 and a central region 315. The annular region 312 may surround the central region 315 and circulate coolant close to one or more heating elements 313 (six in this example), which, in some example embodiments, may be located below the annular region 312. As described elsewhere herein, the annular region 312 may be defined by an outer wall 318 and an inner wall 319. The annular region 312 may serve as a first cooling circuit for cooling the coolant circulating into the device through a first inlet 311. The coolant may circulate around the annular region 312 through the first cooling circuit. At the end of the first cooling circuit in the annular region 312, the coolant may enter a second inlet 314, which, as discussed further in conjunction with manifold 213, circulates the coolant into a second cooling circuit 322, which may circulate the coolant over or around the outer wall 318 and the inner wall 319. Reference Figure 3B The coolant can circulate through the second inlet 314 into the second cooling circuit 322. The second circuit 322 can circulate the coolant into a central region 315 containing cooling elements (e.g., cold plate 316). The coolant can flow close to the cold plate 316 and be discharged through the outlet 328. Figure 3A and Figure 3B The illustration also shows a processing unit 350, which, in some example embodiments, generates heat during operation. A device 310 may be placed near or on top of the processing unit 350 to dissipate heat from it. Although Figures 3A-3BOnly one processing unit 350 of a specific size is illustrated. In other example embodiments, the dimensions of the device 310 and the processing unit 350 may be designed such that multiple processing units 350 can make thermal contact with the device 310. The processing unit 350 may include, but is not limited to, one or more CPUs, GPUs, DPUs, QPUs, multiple PPUs, one or more ASICs, or power supplies.

[0028] In some example embodiments, device 360 ​​may contain only a single cooling circuit. Figure 3C The illustration shows a top cross-sectional perspective view of device 360, which has an annular region 312 and a central region 315. The annular region 312 may surround the central region 315 and circulate coolant close to one or more heating elements 313 (six in this example), which, in some example embodiments, are located below the annular region 312. As described elsewhere herein, the annular region 312 may be defined by an outer wall 318 and an inner wall 319. The annular region 312 may serve as a cooling circuit for cooling the coolant circulating into the device through inlet 370. The coolant may circulate around the annular region 312 via the cooling circuit through inlet 370. At the end of the cooling circuit in the annular region 312, the coolant may be discharged through outlet 380. In some example embodiments, the coolant may be discharged from the outlet to a CDU.

[0029] Figure 4 The illustration shows a system 400 for collecting temperature information from devices. System 400 may include a CDU 410, a device 420, and a management device 450. CDU 410 may include a machine designed to circulate coolant through a cooling manifold connecting CDU 410 and device 420. The coolant may include water, a water-glycol mixture, a dielectric fluid, or other suitable fluid. CDU 410 may circulate the coolant to absorb or dissipate heat from these heat-generating elements, thereby stabilizing the temperature and preventing overheating of the heat-generating elements. In some example embodiments, CDU 410 may operate using a refrigeration cycle involving a compressor that compresses refrigerant gas, enabling it to absorb heat from water, thereby cooling the fluid before the coolant circulates back to the cooling manifold and device 420. Furthermore, CDU 410 may include a control system that monitors the coolant temperature and adjusts the cooling output as needed to maintain the desired temperature level. Additionally, CDU 410 may adjust or change the coolant flow rate based on one or more commands, prompts, or queries from management device 450.

[0030] Device 420 may include an annular region 421 surrounding the central region 430, such as in further combination Figures 1 to 3BAs discussed, device 420 may also include one or more temperature sensors 422A, 422B, 422C, and 422D. Temperature sensors 422A-422D may include and transmit one or more temperature information to management device 450. Device 420 and its associated temperature sensors 422A-422D may be connected to management device 450 via one or more wired or wireless connections. In an example embodiment, these temperature sensors 422A-422D may determine that one or more heat-generating elements associated with device 420 may be approaching a threshold, such as an overheating threshold. In response to this data, management device 450 may communicate with CDU 410 to circulate coolant into device 420. Once the coolant has circulated in device 420, temperature sensors 422A-422D may send updated temperature information to management device 450, which may determine that the temperature of the heat-generating element is below a threshold level (e.g., 70 degrees Celsius) to ensure proper operation of the heat-generating element. In some example embodiments, the management device 450 generates and sends a flow regulation 460, which can change the coolant flow rate based on temperature information 440. As a non-limiting example, the coolant flow rate can be defined as the volume of coolant passing through a specific point in a first or second cooling loop per unit time, for example, liters per minute (L / min). If temperature sensors 422A-422D detect a temperature rise in the heating element exceeding a desired threshold, the management device 450 can increase the coolant flow rate by adjusting the speed of the pump circulating the coolant. Conversely, if the temperature is below the threshold, the management device can reduce the flow rate to save energy and maintain optimal operating conditions.

[0031] In the example embodiment, temperature sensors 422A-422D can generate temperature information and send it to management device 450 and / or user equipment, which can then use the temperature data to adjust the coolant flow rate in the cooling loop or to determine the overall cooling strategy of system 400. Temperature sensors 422A-422D may include thermocouples or thermistors, which operate by measuring changes in resistance or voltage in response to temperature changes. Temperature sensors 422A-422D can collect temperature information by being in direct contact with or near the heat-generating element of device 420, allowing them to accurately collect temperature information about the heat-generating element.

[0032] Figure 5AThe illustration depicts a process 500 for dissipating heat from one or more heat-generating elements. Process 500 can be implemented in various environments discussed elsewhere in this document, including but not limited to data centers, server racks, or any other similar environment where heat-generating elements are cooled. It should be understood that, unless explicitly stated otherwise, the steps of this method can be performed in any order or in parallel. Furthermore, the method may contain more or fewer steps.

[0033] A cooling device 502 is provided, having a first cooling circuit for circulating coolant along a ring path around a central region. The coolant can circulate 504 through the first cooling circuit near one or more heat-generating elements. The cooling circuit can deliver coolant throughout the device or near one or more heat-generating elements to dissipate heat 506 from the heat-generating elements in thermal contact with the device.

[0034] Figure 5B The illustration depicts a process 520 for transmitting temperature sensor information to a management device. The device may include one or more temperature sensors that generate temperature information about one or more heating elements. For example, the temperature sensors may generate temperature information indicating that the temperature of one or more areas near the heating elements is 70 degrees Celsius or 160 degrees Fahrenheit. It should be understood that, unless otherwise explicitly stated, the steps of this method can be performed in any order or in parallel. Furthermore, the method may include more or fewer steps.

[0035] A cooling device 522 is provided with a first cooling circuit for circulating coolant along an annular path around a central region. The coolant can be circulated 524 through the first cooling circuit near one or more heat-generating elements. The cooling circuit can deliver coolant throughout the device or near one or more heat-generating elements to dissipate heat from the heat-generating elements in thermal contact with the device.

[0036] One or more temperature sensors on the device can generate temperature information and send 526 information to a controller, user equipment, management device, or other suitable device. The information can be transmitted via a wired or wireless connection. The management device can receive the temperature information and, when it determines that the temperature information indicates that the temperature corresponding to one or more heat-generating elements has reached a threshold, command the CDU to circulate coolant to the cold plate. In response to the command from the controller, coolant can be circulated 528 through a first cooling loop and a second cooling loop. These cooling loops can deliver coolant throughout the device or close to one or more heat-generating elements to dissipate heat from these elements.

[0037] Figure 6An example network configuration 600 is illustrated, comprising components that can be used to implement various aspects of the embodiments, such as components for providing, generating, modifying, encoding, processing, fusing, and / or transmitting generated image data, calculated measurements, or other such content. In at least one embodiment, client device 602 may use components of content application 604 on client device 602, as well as data locally stored on that client device, to generate or receive data for a session. In at least one embodiment, content application 624 running on computer or processor 620 (e.g., cloud server or control system) may initiate a session associated with at least one client device 602 (e.g., vehicle or robot), such as using a session manager and user data stored in user database 636, and may allow content such as liquid coolant or server thermal data to be selected and / or retrieved from repository 634 for use by test module 632, thereby calculating one or more performance metrics for monitoring module 628, which may provide flow or thermal data to control module 630 in an environment where the data is used to determine appropriate operating conditions for controlling flow or temperature. Content manager 626 can work with at least these various modules to perform testing and analysis, and may instruct any actions to be taken in response to performance metrics failing to meet operational requirements. At least a portion of the data or instruction content can be transmitted to client device 602 and / or physical device 670 using a suitable transmission manager 622 for sending via download, streaming, or other such transmission channels. At least a portion of the data can be encoded and / or compressed using an encoder before being transmitted to client device 602. In at least one embodiment, client device 602 receiving such content can provide it to a corresponding content application 604, which may also or optionally include a graphical user interface 610, a traffic monitoring module 612, and a control module 614 for providing, compositing, rendering, combining, modifying, or using the content on or through client device 602 for presentation, navigation, control (or other purposes), such as transmission to physical device 670. In some embodiments, the computer or processor 620 and the client device 602 may be able to communicate directly without transmitting data over the network 640 to avoid issues such as latency and availability. A decoder may also be used to decode data received over the network 640 for presentation by the client device 602, such as displaying image content or performance metrics via a display device 606, and playing audio, such as corresponding sound or synthesized speech, via at least one audio playback device 608 (e.g., a speaker or headphones).In at least one embodiment, at least a portion of the content may have been stored on, rendered on, or accessible by the client device 602, such that at least that portion of the content does not need to be transmitted over the network 640; for example, the content (e.g., hot data) may have been previously downloaded or stored locally on a hard drive or optical disc. In at least one embodiment, this content may be transmitted from the computer or processor 620 or the user database 636 to the client device 602 using a transmission mechanism such as data streaming. In at least one embodiment, at least a portion of the content may be acquired, enhanced, and / or streamed from other sources (e.g., third-party services 660 or other client devices 650), which may also include content applications for generating, updating, enhancing, or providing map content. In at least one embodiment, portions of this functionality may be executed using multiple computing devices, or multiple processors within one or more computing devices, such as a combination of a CPU and a graphics processing unit (GPU), a DPU, a QPU, or multiple parallel processing units (PPUs).

[0038] In this example, these client devices can include any suitable computing device, such as desktop computers, laptops, set-top boxes, streaming media devices, game consoles, smartphones, tablets, VR headsets, AR goggles, wearable computers, or smart TVs. Each client device can submit requests via at least one wired or wireless network, such as the Internet, Ethernet, a local area network (LAN), or a cellular network, and other such options. In this example, these requests can be submitted to an address associated with a cloud provider that can operate or control one or more electronic resources within the cloud provider's environment, such as data centers or server farms. In at least one embodiment, the request can be received or processed by at least one edge server located at the network edge and outside at least one security layer associated with the cloud provider's environment. In this way, latency can be reduced by enabling client devices to interact with servers closer to the network, while also improving the security of resources within the cloud provider's environment.

[0039] In at least one embodiment, such a system can be used to perform graphics rendering operations. In other embodiments, such a system can be used for other purposes, such as providing image or video content to test or validate autonomous machine applications, or for performing deep learning operations. In at least one embodiment, such a system can be implemented using an edge device, or may include one or more virtual machines (VMs). In at least one embodiment, such a system can be implemented at least partially in a data center, or at least partially using cloud computing resources.

[0040] Data Center Figure 7 An example data center 700 is illustrated, in which at least one embodiment can be used. In at least one embodiment, data center 700 includes a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.

[0041] In at least one embodiment, such as Figure 7 As shown, the data center infrastructure layer 710 may include a resource coordinator 712, packet computing resources 714, and node computing resources (“nodes CR”) 716(1)-716(N), where “N” represents any positive integer. In at least one embodiment, the nodes CR 716(1)-716(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more of the nodes CR 716(1)-716(N) may be servers having one or more of the aforementioned computing resources.

[0042] In at least one embodiment, the grouped computing resource 714 may include individual groups of node CRs housed in one or more racks (not shown), or multiple racks housed in data centers (also not shown) located in different geographical locations. The individual groups of node CRs in the grouped computing resource 714 may include grouped computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, a plurality of node CRs (including CPUs or processors) may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.

[0043] In at least one embodiment, resource coordinator 712 may be configured or otherwise control one or more nodes CR 716(1)-716(N) and / or group computing resources 714. In at least one embodiment, resource coordinator 712 may include a Software Design Infrastructure (“SDI”) management entity for data center 700. In at least one embodiment, resource coordinator may include hardware, software, or some combination thereof.

[0044] In at least one embodiment, such as Figure 7As shown, framework layer 720 includes job scheduler 722, configuration manager 724, resource manager 726, and distributed file system 728. In at least one embodiment, framework layer 720 may include a framework for supporting software 732 of software layer 730 and / or one or more applications 742 of application layer 740. In at least one embodiment, software 732 or application 742 may respectively include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 720 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark"), which can use distributed file system 728 for large-scale data processing (e.g., "big data"). In at least one embodiment, job scheduler 722 may include Spark drivers to facilitate the scheduling of workloads supported by the various layers of data center 700. In at least one embodiment, configuration manager 724 is capable of configuring different layers, such as software layer 730 and framework layer 720, including Spark and distributed file system 728, to support large-scale data processing. In at least one embodiment, resource manager 726 may be capable of managing clustered or grouped computing resources mapped or allocated to support distributed file system 728 and job scheduler 722. In at least one embodiment, clustered or grouped computing resources may include grouped computing resources 714 at data center infrastructure layer 710. In at least one embodiment, resource manager 726 may coordinate with resource coordinator 712 to manage these mapped or allocated computing resources.

[0045] In at least one embodiment, the software 732 included in the software layer 730 may include at least a portion of the nodes CR 716(1)-716(N), the software used by the grouped computing resources 714 and / or the distributed file system 728 of the framework layer 720. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.

[0046] In at least one embodiment, one or more applications 742 included in application layer 740 may include at least a portion of nodes CR 716(1)-716(N), grouped computing resources 714, and / or the distributed file system 728 of framework layer 720, one or more types of applications. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0047] In at least one embodiment, any of the configuration manager 724, resource manager 726, and resource coordinator 712 can perform any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. In at least one embodiment, the self-modification actions can protect the data center operator of the data center 700 from making potentially undesirable configuration decisions and may prevent underutilized and / or poorly performing portions of the data center.

[0048] In at least one embodiment, according to one or more embodiments described herein, data center 700 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information. For example, in at least one embodiment, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described above for data center 700. In at least one embodiment, the trained machine learning model corresponding to one or more neural networks can be used to infer or predict information using the resources described above for data center 700 by utilizing weight parameters calculated using one or more training techniques described herein.

[0049] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, DPU, QPU, multiple parallel processing units (PPU), or other hardware to perform training and / or inference using the resources described above. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as services to allow users to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0050] The inference and / or training logic 715 can be used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, the inference and / or training logic 715 can be used to... Figure 7In the system shown, inference or prediction operations are performed based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0051] Such components can be used in data centers that utilize the cooling loop described herein.

[0052] Figure 8 An example computing environment 800 according to at least one embodiment is illustrated, in which forward pass offload to available memory can be performed. It should be understood that embodiments of this disclosure can also be used with reference to alternative environments, and specific discussions of components are provided by way of non-limiting examples and may include equivalents. Furthermore, various features have been omitted for clarity and brevity. Additionally, the systems and methods can be used with a variety of different architectures. Example computing environment 800 may include server 802, which can be used to perform HPC workloads, such as AI training or machine learning model training. In one embodiment, server 802 may be an application instance or a compute node. Server 802 may include CPU 810 associated with a switch 820 (e.g., a Peripheral Component Interconnect High Speed ​​(PCIe) switch) that can control at least some data transfers on communication paths interconnecting various components. In one embodiment, CPU 810 may include a root complex processor.

[0053] The PCIe switch 820 may also be associated with GPU 830 and DPU 840, and may transfer data between at least some of the CPU 810, GPU 830, DPU 840, and other components. In one embodiment, the PCIe switch 820 may be associated with more than one GPU or more than one DPU. In another embodiment, the PCIe switch 820 may be located inside the DPU 840. The PCIe switch 820 may manage the transfer of at least a portion of the data between the CPU 810, GPU 830, and DPU 840. In another embodiment, the number of GPUs associated with the PCIe switch 820 may be equal to the number of DPUs associated with the PCIe switch 820. In at least one embodiment, the server 802 may include, but is not limited to, any number of CPUs 810, PCIe switches 820, GPUs 830, and / or DPUs 840, in any combination. For example, in at least one embodiment, the server 802 may include eight, sixteen, thirty-two, or more GPUs 830. In at least one embodiment, Figure 8The communication paths for interconnecting various components (including but not limited to CPU 810, PCIe switch 820, GPU 830 and DPU 840) can be implemented using any suitable protocol, such as peripheral component interconnect (PCI) based protocols (e.g. PCIe), or other bus or point-to-point communication interfaces and / or protocols, such as NV-Link high-speed interconnect or interconnect protocols.

[0054] DPU 840 may include a network interface controller (NIC) 842, DDR memory 844, and a non-volatile memory fast (NVMe) device 846. NIC 842 may interface with network 804, which may also interface with other NVMe devices available to DPU 840 (e.g., via a structure). In one embodiment, DPU 840 may not include NVMe device 846. In another embodiment, NVMe device 846 may reside on server 802 instead of DPU 840. In yet another embodiment, computing environment 800 may include more than one NVMe device 846, such as a first NVMe device in DPU 840 and a second NVMe device on server 802 and directly associated with PCIe switch 820. In one embodiment, DPU 840 may not include DDR memory 844 and may include compute storage service (CSS) as an alternative to or supplement to DDR memory 844. For example, computing environment 800 may include DPU compute storage (CS) memory 806, which is available to DPU 840 as part of CSS. Network 804 can interface with DPU CS memory 806 via NIC 842 according to any suitable interface protocol, such as Remote Direct Memory Access (RDMA) via Ethernet, InfiniBand, Fibre Channel, etc.

[0055] By using DPU 840 on a node of the system, the total memory available for data storage in the computing environment 800 can be expanded. DPU 840 can access memory pool 850 already available on server 802, such as Double Data Rate (DDR) memory, onboard NVMe devices, NVMe devices via the architecture, and CS. Memory pool 850 may include at least one of DDR memory 844, NVMe 846, and DPU CS memory 806. DPU 840 can also access the available memory of other DPUs as part of pool 850, and other DPUs can also access the available memory of DPU 840, such as pool 850. This available memory can be accessed and utilized for data storage without adding computing resources (such as compute nodes), whereas other solutions require additional computing resources. The available pool 850 accessible to DPU 840 can be configured for server 802, thereby expanding the total memory available for data storage, for example, reducing the data storage load on CPU 810 or GPU 830, which in turn can improve the utilization of the memory they use for processing. For example, during AI training, model states, residual states, activation functions, and checkpoints can be stored or unloaded into pool 850 accessible by DPU 840.

[0056] Figure 9 A computer system 900 according to at least one embodiment is illustrated. In at least one embodiment, the computer system 900 is configured to implement the various processes and methods described throughout this disclosure.

[0057] In at least one embodiment, the computer system 900 includes, but is not limited to, at least one central processing unit (“CPU”) 902 connected to a communication bus 910, which is implemented using any suitable protocol, such as PCI (“Peripheral Component Interconnect”), Peripheral Component Interconnect High Speed ​​(“PCI-Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 900 includes, but is not limited to, main memory 904, in which control logic (e.g., implemented in hardware, software, or a combination thereof) and data are stored, and main memory 904 may be in the form of random access memory (“RAM”). In at least one embodiment, a network interface subsystem (“network interface”) 922 provides an interface with other computing devices and networks to receive data from the computer system 900 and to send data from the computer system 900 to other systems.

[0058] In at least one embodiment, the computer system 900 includes, but is not limited to, an input device 908, a parallel processing system 912, and a display device 906, which may employ conventional cathode ray tube (“CRT”), liquid crystal display (“LCD”), light-emitting diode (“LED”), plasma display, or other suitable display technologies. In at least one embodiment, user input is received from the input device 908, such as a keyboard, mouse, touchpad, microphone, etc. In at least one embodiment, each of the above modules may reside on a single semiconductor platform, thereby forming the processing system.

[0059] In at least one embodiment, a computer program in the form of machine-readable executable code or computer control logic algorithms is stored in main memory 904 and / or auxiliary storage devices. According to at least one embodiment, if the computer program is executed by one or more processors, it enables system 900 to perform various functions. Memory 904, storage devices, and / or any other storage devices are possible examples of computer-readable media. In at least one embodiment, auxiliary storage devices can refer to any suitable storage device or system, such as hard disk drives and / or removable storage drives, representing floppy disk drives, magnetic tape drives, optical disk drives, digital versatile optical disc (“DVD”) drives, recording devices, Universal Serial Bus (“USB”) flash memory, etc. In at least one embodiment, the architecture and / or functionality in the various preceding figures are implemented in the context of: CPU 902; parallel processing system 912; integrated circuits capable of implementing at least a portion of the functions of both CPU 902 and parallel processing system 912; chipsets (e.g., a set of integrated circuits designed to operate and be sold as units for performing related functions); and any suitable combination of integrated circuits.

[0060] In at least one embodiment, the architecture and / or functionality of the various preceding figures are implemented within the context of general-purpose computer systems, circuit board systems, game console systems for entertainment purposes, dedicated systems, etc. In at least one embodiment, the computer system 900 may take the form of a desktop computer, laptop computer, tablet computer, server, supercomputer, smartphone (e.g., wireless handheld device), personal digital assistant (“PDA”), digital camera, vehicle, head-mounted display, handheld electronic device, mobile phone, television, workstation, game console, embedded system, and / or any other type of logic.

[0061] In at least one embodiment, the parallel processing system 912 includes, but is not limited to, multiple parallel processing units (“PPUs”) 914 and associated memory 916. In at least one embodiment, the PPUs 914 are connected to a host processor or other peripheral device via interconnects 918 and switches 920 or multiplexers. In at least one embodiment, the parallel processing system 912 distributes computational tasks across the parallelizable PPUs 914, for example, as part of distributing computational tasks across multiple graphics processing units (“GPUs”) thread blocks. In at least one embodiment, shared and accessible (e.g., for read and / or write access) memory is present on some or all of the PPUs 914, although such shared memory may result in performance penalties compared to using local memory and registers residing on the PPUs 914. In at least one embodiment, the operation of the PPUs 914 is synchronized using commands such as _syncthreads(), where all threads in a block (e.g., executing on multiple PPUs 914) must reach a certain point of execution in the code before proceeding.

[0062] Such components can be used in data centers that utilize the cooling loop described herein.

[0063] Figure 10 This is a block diagram schematically illustrating a computing system 1000, such as a data center or high-performance computing (HPC) cluster, according to embodiments described herein. According to at least one embodiment, system 1000 includes multiple subsystems, such as multiple processing devices, multiple network devices, and multiple networks coupled to each other. The computing system 1000 is designed to have multiple integrated circuits (referred to as processing devices), wherein each integrated circuit may include one or more CPUs and GPUs, thereby forming a powerful and flexible architecture.

[0064] The various processing devices are interconnected via NVLink or other high-speed interconnects to enable high-speed communication between subsystems, and are also connected via NICs or DPUs to ensure efficient data transmission within the computing system 1000 and with one or more external networks 1030, 1036. In this example, system 1000 includes: a switch 1048 that connects NIC / DPU 1028 to network 1030; and a packet switch 1050 that connects NIC / DPU 1032 to network 1036.

[0065] NVLink-coupled processing devices enable seamless data exchange and parallel processing, thereby improving overall computing performance. Processing devices connect to multiple networks via one or more Network Interface Controllers (NICs) or Data Processing Units (DPUs), enabling the system to handle complex, multi-network tasks with high bandwidth and low latency. This configuration is ideal for demanding applications requiring massive processing power, such as artificial intelligence (AI), machine learning (ML), and data-intensive computing, while ensuring robust connectivity and scalability across diverse network environments. The integrated circuits of the Computing System 1000 may include one or more CPUs and one or more GPUs.

[0066] Figure 10 An example architecture of a multi-GPU architecture is also shown. As shown in the figure, computing system 1000 includes a processing device 1002 with a multi-GPU architecture. Specifically, processing device 1002 may be a system-on-a-chip and includes multiple subsystems, such as CPU 1006, GPU 1008, and GPU 1010. CPU 1006 may be coupled to GPU 1008 via die-to-die (D2D) or chip-to-chip (C2C) interconnect 1012 (e.g., ground reference signaling interconnect (GRS interconnect)). CPU 1006 may be coupled to GPU 1010 via D2D or C2C interconnect 1014. CPU 1006 may also be coupled to GPU 1008 and GPU 1010 via PCIe interconnect.

[0067] The CPU 1006 can be coupled to one or more NICs or DPUs, which in turn can be coupled to one or more networks. For example, as Figure 10 As shown, CPU 1006 is coupled to a first NIC / DPU 1026, which is coupled to network 1030. CPU 1006 is also coupled to a second NIC / DPU 1028, which is coupled to network 1030 via switch 1048. For example, NIC / DPU 1026 and NIC / DPU 1028 can be coupled to network 1030 via Ethernet (ETH), NVLINK, or InfiniBand (IB) connections.

[0068] The computing system 1000 also includes a processing device 1004 with a multi-GPU architecture. Specifically, the processing device 1004 includes multiple subsystems, including a CPU 1016, a GPU 1018, and a GPU 1020. The CPU 1016 can be coupled to the GPU 1018 via a D2D or C2C interconnect 1022. The CPU 1016 can be coupled to the GPU 1020 via a D2D or C2C interconnect 1024. The CPU 1016 can also be coupled to the GPU 1018 and GPU 1020 via a PCIe interconnect. The CPU 1016 can be coupled to one or more NICs or DPUs, and the NICs or DPUs can be coupled to one or more networks. For example, as... Figure 10 As shown, CPU 1016 is coupled to a first NIC / DPU 1034, which is coupled to network 1036. CPU 1016 is also coupled to a second NIC / DPU 1032, which is coupled to network 1036 via switch 1050. NIC / DPU 1032 and NIC / DPU 1034 can be coupled to network 1036 via Ethernet (ETH), NVLINK, or InfiniBand (IB) connections.

[0069] In at least one embodiment, processing device 1002 and processing device 1004 can communicate with each other via NIC / DPU 1038, for example via PCIe interconnect. Processing device 1002 and processing device 1004 can also communicate with each other via high-bandwidth communication interconnect 1040, such as NVLink interconnect or other high-speed interconnect. Figure 10 The packet switches in the diagram can include, for example, Nvidia Quantum-2 switches. The NIC / DPU in the diagram can include, for example, Nvidia Bluefield DPUs.

[0070] In various embodiments, any of the network devices of system 1000, such as any of NIC / DPU 1026, 1028, 1032, 1034, and 1038, and / or any of switches 1048 and 1050, can use ILI packets according to the techniques described herein. These components can be used in data centers employing liquid cooling systems.

[0071] In an example embodiment, the heating element may be a silicon photonic element (e.g., a SiPh die) or a photonic IC. These elements may be co-packaged to form a multi-chip module (MCM) assembly.

[0072] refer to Figure 11 and Figure 12The figures illustrate a cross-sectional view and a top plan view of the optoelectronic component 1100. In some embodiments, the optoelectronic component 1100 may include a substrate 1102. For example, the substrate 1102 may be a printed circuit board, a metal carrier, an organic carrier, and / or a ceramic carrier. In some embodiments, the height of the substrate 1102 may vary. In this regard, for example, the height of a first portion 1102A of the substrate 1102 may be h1, and the height of a second portion 1102B of the substrate 1102 may be h2. In some embodiments, an electronic integrated circuit 1104 may be supported by the substrate 1102. The electronic integrated circuit 1104 may be any type of electronic integrated circuit. For example, the electronic integrated circuit 1104 may be a digital signal processor, a modulator driver, and / or a transimpedance amplifier. In some embodiments, the substrate 1102 may support more than one electronic integrated circuit. In some embodiments, the height of the electronic integrated circuit 1104 may be h3. In some embodiments, the optoelectronic component 1100 may support more than one electronic integrated circuit. In some embodiments, a photonic integrated circuit 1106 may be supported by the substrate 1102. The photonic integrated circuit 1106 may be any type of photonic integrated circuit. For example, the photonic integrated circuit 1106 may be an electro-optic modulator, a photodiode, a transmitter optical sub-assembly, and / or a receiver optical sub-assembly. In some embodiments, the photonic integrated circuit 1106 may comprise graphene. In some embodiments, the substrate 1102 may support more than one photonic integrated circuit. In some embodiments, the height of the photonic integrated circuit 1106 may be h4. In some embodiments, heights h1, h2, h3, and h4 may be different. For example, depending on the electronic integrated circuit and the photonic integrated circuit used, height h3 may be greater than height h4, and vice versa.

[0073] In some embodiments, the optoelectronic component 1100 may include one or more optical fibers 1118 connected to the photonic integrated circuit 1106. The one or more optical fibers 1118 may be configured to connect the optoelectronic component 1100 to other optical components and / or devices. In some embodiments, port 1116 may be connected to substrate 1102. Port 1116 may be configured to connect the optoelectronic component 1100 to other electronic components and / or devices. In some embodiments, the optoelectronic component 1100 may be configured to operate at a speed higher than 25 Gb / s.

[0074] The optoelectronic component 1100 may include a plurality of substrate interconnect connectors 1110 disposed on a substrate 1102, a plurality of electronic integrated circuit interconnect connectors 1112 disposed on an electronic integrated circuit 1104, and a plurality of photonic integrated circuit interconnect connectors 1114 disposed on a photonic integrated circuit 1106. The plurality of substrate interconnect connectors 1110, electronic integrated circuit interconnect connectors 1112, and photonic integrated circuit interconnect connectors 1114 may be made of any conductive material (e.g., conductive adhesive and / or solder). In some embodiments, the plurality of substrate interconnect connectors 1110, electronic integrated circuit interconnect connectors 1112, and photonic integrated circuit interconnect connectors 1114 may be flexible. In other words, in some embodiments, the plurality of substrate interconnect connectors 1110, electronic integrated circuit interconnect connectors 1112, and photonic integrated circuit interconnect connectors 1114 may be manipulated such that each connector can take on various shapes. In some embodiments, the spacing between the plurality of substrate interconnect connectors 1110 may be p1, the spacing between the plurality of electronic integrated circuit interconnect connectors 1112 may be p2, and the spacing between the plurality of photonic integrated circuit interconnect connectors 1114 may be p3. Spacing refers to the distance between each interconnect connector in a plurality of interconnect connectors. In some embodiments, the spacings p1, p2, and p3 can be different. For example, the spacing p2 of the plurality of electronic integrated circuit interconnect connectors 1112 can be 1.25 mm, while the spacing p3 of the plurality of photonic integrated circuits can be 1.5 mm.

[0075] In some embodiments, the optoelectronic component 1100 may include a first plurality of cable connectors 1108. In some embodiments, each of the first plurality of cable connectors 1108 can be connected to and communicate with the substrate 1102, the electronic integrated circuit 1104, and the photonic integrated circuit 1106 via a corresponding interconnect connector. In other words, the first plurality of cable connectors 1108 can be connected to and communicate with the substrate 1102 via a plurality of substrate interconnect connectors 1110, connected to and communicate with the electronic integrated circuit 1104 via a plurality of electronic integrated circuit interconnect connectors 1112, and connected to and communicate with the photonic integrated circuit 1106 via a plurality of photonic integrated circuit interconnect connectors 1114. Therefore, the first plurality of cable connectors 1108 can be used to facilitate communication between the substrate 1102, the electronic integrated circuit 1104, and the photonic integrated circuit 1106.

[0076] In some embodiments, the first plurality of cable connectors 1108 may define a first layout. In some embodiments, the first layout may define the overall connectivity of the optoelectronic component 1100.

[0077] In some embodiments, the optoelectronic component 1100 may include a substrate 1102. For example, the substrate 1102 may be a printed circuit board, a metal carrier, an organic carrier, and / or a ceramic carrier.

[0078] The electronic integrated circuit 1104 can be any type of electronic integrated circuit. For example, the electronic integrated circuit 1104 can be a digital signal processor, a modulator driver, and / or a transimpedance amplifier. In some embodiments, the substrate 1102 can support more than one electronic integrated circuit.

[0079] In some embodiments, the photonic integrated circuit 1106 may be supported by a substrate 1102. The photonic integrated circuit 1106 may be any type of photonic integrated circuit. For example, the photonic integrated circuit 1106 may be an electro-optic modulator, a photodiode, a transmitter optical sub-assembly, and / or a receiver optical sub-assembly.

[0080] For example, refer to Figure 13 In the illustrated example, the connectivity defined by the first layout allows the electronic integrated circuit 1304 to be connected via cable connector 1308 to the first photonic integrated circuit 1306A and the second photonic integrated circuit 1306B. In some embodiments, the first plurality of cable connectors 1108 can be interchanged with other plurality of cable connectors defining different layouts. Different layouts can alter the overall connectivity of the optoelectronic component 1100. For example, the first plurality of cable connectors 1108 can be interchanged with a second plurality of cable connectors defining a second layout, thereby altering the overall connectivity of the optoelectronic component 1100. In this way, the optoelectronic component 1100 can be easily modified to obtain the desired functionality by interchangeing the cable connectors.

[0081] In some embodiments, the first plurality of cable connectors 1108 may be flexible. This helps ensure that the first plurality of cable connectors 1108 can be used with various substrates, electronic integrated circuits, and photonic integrated circuits. For example, the substrates, electronic integrated circuits, and / or photonic integrated circuits may come from different manufacturers, may be different types of integrated circuits or substrates, and / or may have different functions. For example, substrate 1102, electronic integrated circuit 1104, and photonic integrated circuit 1106 may have different heights (e.g., the height h3 of electronic integrated circuit 1104 may be greater than the height h4 of photonic integrated circuit 1106). The flexibility of the first plurality of cable connectors 1108 allows them to bend as needed, thereby accommodating components of optoelectronic components 1100 with different heights, and enabling connection without any modification to the configuration of the optoelectronic components 1100 themselves. Furthermore, the flexibility of the first plurality of cable connectors 1108 allows them to be used with various substrates, electronic integrated circuits, and photonic integrated circuits having interconnect connectors with different pitches. For example, if the spacing p2 of the multiple electronic integrated circuit interconnect connectors 1112 is smaller than the spacing p3 of the multiple photonic integrated circuit interconnect connectors 1114, the first multiple cable connectors 1108 can be bent to accommodate the spacing difference and connect the electronic integrated circuit 1104 to the photonic integrated circuit 1106.

[0082] refer to Figure 13The illustration shows a portion of an example optoelectronic component 1300. For example, the example optoelectronic component 1300 may be part of a 1.6Tb / s demonstrator. The example optoelectronic component 1300 includes a substrate 1302, an electronic integrated circuit 1304 supported by the substrate 1302, a first photonic integrated circuit 1306A supported by the substrate 1302, and a second photonic integrated circuit 1306B supported by the substrate 1302. The example optoelectronic component 1300 may include a plurality of electronic integrated circuit interconnect connectors 1312 disposed on the electronic integrated circuit 1304 and a plurality of photonic integrated circuit interconnect connectors 1314 disposed on the first photonic integrated circuit 1306A and the second photonic integrated circuit 1306B. The electronic integrated circuit 1304 may be connected to and communicate with the first photonic integrated circuit 1306A and the second photonic integrated circuit 1306B via a plurality of cable connectors 1308. In the example optoelectronic component 1300, an electronic integrated circuit 1304 and a first photonic integrated circuit 1306A are located on a substrate 1302, such that a plurality of electronic integrated circuit interconnect connectors 1312 and a plurality of photonic integrated circuit interconnect connectors 1314 disposed on the first photonic integrated circuit 1306A are misaligned with each other (e.g., one is not positioned opposite the other). In this case, the flexibility of the plurality of cable connectors 1308 facilitating communication between the electronic integrated circuit 1304 and the first photonic integrated circuit 1306A allows the electronic integrated circuit 1304 and the first photonic integrated circuit 1306A to be connected by manipulating the cable connectors to accommodate the misaligned positions.

[0083] refer to Figure 14 The illustration shows another example optoelectronic component 1400. For example, this example optoelectronic component 1400 may be part of an eight-channel OSFP transceiver. The example optoelectronic component 1400 includes a substrate 1402, an electronic integrated circuit 1404 supported by the substrate 1402, and a photonic integrated circuit 1406 supported by the substrate 1402. The example optoelectronic component 1400 may include a plurality of electronic integrated circuit interconnect connectors 1412 disposed on the electronic integrated circuit 1404 and a plurality of photonic integrated circuit interconnect connectors 1414 disposed on the photonic integrated circuit 1406. The electronic integrated circuit 1404 may be connected to and communicate with the photonic integrated circuit 1406 via a plurality of cable connectors 1408. In the example optoelectronic component 1400, the spacing between the plurality of electronic integrated circuit interconnect connectors 1412 and the plurality of photonic integrated circuit interconnect connectors 1414 is different. In this case, the flexibility of the multiple cable connectors 1408 that facilitate communication between the electronic integrated circuit 1404 and the photonic integrated circuit 1406 makes it possible to connect the electronic integrated circuit 1404 and the photonic integrated circuit 1406, even if the spacing is different, by bending or other means to shape the cable connectors to accommodate the spacing difference.

[0084] Silicon photonics (SiP) is a technology that enables optical systems to be fabricated using silicon processes, with silicon serving as the optical medium. Various optical components, such as interconnects and signal processing components, can be fabricated and integrated into a single SiP device. Some SiP devices are fabricated on a silicon dioxide substrate or a silicon dioxide layer on a silicon substrate; this technique is often referred to as silicon-on-insulator (SOI). In some optical systems, SiP devices are attached to external devices to facilitate optical communication. However, it is often difficult to precisely align the optical signal on the SiP with the external device receiving the light.

[0085] In some optical systems, SiP devices are attached to external equipment to facilitate optical communication. However, it is often difficult to precisely align the optical signal on the SiP device with the external device receiving the light. For example, long-distance transmission of optical signals typically takes place in optical fibers. When generating or processing optical signals in a SiP device for transmission over an optical fiber, optical coupling is required between the SiP device and the fiber. Since waveguides within SiP devices typically have diameters smaller than the fiber diameter, this coupling between the SiP device and the fiber is often challenging. Therefore, a "world-to-chip" interface problem frequently arises in SiP technology, where optical coupling between silicon wire waveguides and optical fibers (and vice versa) is often inefficient.

[0086] Traditionally, fiber-to-chip (FTCH) coupling employs fiber coupling techniques using spot size converters (SSCs) or grating couplers. However, grating couplers used for FTCH typically offer narrow bandwidth and / or exhibit undesirable polarization sensitivity for certain optical applications. Furthermore, SSCs and grating couplers for FTCH are typically attached to the chip using adhesive techniques, resulting in fiber bundles attached to the silicon communication chip, increasing the complexity of handling and / or assembling the chip in other optical systems. Additionally, the wafer of a traditional SiP device is often diced (e.g., completely cut through) to form the wafer's edge, exposing the waveguide endface and / or facilitating the interfacing of the SiP device with external devices. Co-packaging can refer to the tight integration of different electrical and / or optoelectronic chips within the same package.

[0087] Figure 15A block diagram of a co-packaged network device 1500 is schematically illustrated according to embodiments disclosed herein. Different chips constituting the co-packaged network device are assembled on a single substrate; this assembly configuration is commonly referred to as an MCM assembly 1512. The MCM assembly 1512 may include a switching circuitry system 1516 surrounded by peripheral or satellite chips 1520. In some embodiments, the switching circuitry system 1516 and the surrounding satellite chips 1520 are mounted on a common substrate, although this configuration is not required. The MCM assembly 1512 may be disposed within a larger housing of the co-packaged network device 1500, behind the front panel 1504. The switching circuitry system 1516 may include one or more core digital application-specific integrated circuits (ASICs), CPUs, GPUs, microprocessors, FPGAs, combinations thereof, etc. The switching circuitry system 1516 may include several input ports and / or output ports 1524. The input / output (I / O) ports 1524 may include electrical ports and / or optical ports. Furthermore, the switching circuitry system 1516 may include a combination of electrical and optical modules. The electrical module of the switching circuit system 1516 may include a plurality of electrical switches configured to route signals in the electrical domain. The optical module of the switching circuit system 1516 may include a plurality of optical components configured to generate, detect, and route signals in the optical domain. In some embodiments, the MCM assembly 1512 may involve or include a plurality of satellite chips 1520 assembled on the same substrate as the switching circuit system 1516. In some embodiments, the configuration of the optical module and the electrical module depends on (e.g., based on) the number of optical ports in the I / O ports 1524.

[0088] As described above, optical I / O port 1508 (also referred to as an optical connector) is located on front panel 1504. As described above, the connection between MCM assembly 1512 and optical I / O port 1508 can be transmitted to front panel 1504 via optical fiber. This connection can be implemented directly through optical I / O port 1524 of the switching circuitry system, or through one or more satellite chips 1520. Connection is typically made through one or more satellite chips 1520 because satellite chips 1520 may include electro-optic converters and may contain SERDES to locally support the connection. Satellite chip 1520 may include one or more DSP processors, drivers, transimpedance amplifiers, lasers, modulators, photodiodes, serializers-deserializers, etc.

[0089] Some embodiments of this disclosure relate to a multi-chip module (MCM) having a central main chip and a plurality of peripheral MCM sockets configured to mechanically receive and electrically connect mezzanine packages, which may include co-packaged optics (CPO) packages and co-packaged copper (CPC) packages. Each mezzanine package may include a package substrate comprising a connector portion and a main portion configured to engage with an MCM socket, the main portion extending beyond the periphery of the MCM substrate. The main portion of the mezzanine package may be configured to receive optics and / or integrated circuits, for example via the mezzanine sockets, to allow connection between the optics and / or integrated circuits / RF copper connector and the main chip of the MCM. Because the mezzanine packages extend beyond the periphery of the MCM substrate, the physical size of the MCM substrate can be kept small, thereby reducing costs and avoiding the manufacturing challenges discussed above, while allowing connection of several optics and integrated circuits via the mezzanine packages, which occupy relatively inexpensive space around the periphery of the MCM substrate. As used herein, the terms "co-packaged optics" (or "CPO") and "co-packaged copper" (or "CPC") can refer to advanced heterogeneous integration of optics with silicon or copper with silicon, either of which can be achieved on a single package substrate. CPO can utilize pluggable optical modules containing optical engines (OEs) for converting optical signals to electrical signals and vice versa. CPO can also include optical components on a photonic die and electrical components on an electrical die.

[0090] As used in this article, a ball grid array (BGA) is a surface mount package type for integrated circuits. A BGA package uses an array of metal conductor balls arranged in a grid to permanently mount devices such as microprocessors onto a PCB. These metal conductor balls can then undergo the reflow process described above, where the metal conductor balls are preheated and then melted to bond the IC to the substrate, thereby forming the IC package.

[0091] As used herein, the term "flip chip (FC)" can refer to a method for interconnecting a die (e.g., a semiconductor device, IC chip, integrated passive device, and microelectromechanical system (MEMS)) to an external circuit system using solder bumps deposited on chip pads. Solder bumps can be deposited on chip pads on the top surface of the wafer during the final wafer fabrication process. The chip can be mounted onto an external circuit system (e.g., a circuit board or other chip or wafer) by "flipping" the chip so that its top surface is facing down and it is positioned so that the chip's pads are aligned with matching pads on the external circuitry. Solder is reflowed to complete the interconnect.

[0092] Various packaging schemes are employed in integrated circuit (IC) chip packaging, including traditional two-dimensional (2D) IC packaging and the recently introduced 2.5D and 3D IC packaging. In 2D IC packaging, multiple chips are mounted on a printed circuit board, where high-performance logic, low-performance logic, memory, analog / RF functions, and other functional components exist as discrete devices within separate chip packages. In contrast, in 2.5D and 3D IC packaging, multiple IC chips are mounted on a silicon interposer, rather than on a traditional packaging substrate. The silicon interposer is typically a silicon wafer, and because the manufacturing process used to form the conductive traces is the same as the process used to form the metal interconnects in the metallization layer of the silicon chip, very small and high-density conductive traces can be formed between multiple IC chips.

[0093] Compared to 2.5D and 3D IC packages, circuit boards using individually packaged chips (such as 2D IC packages) have several disadvantages. For example, 2D IC packages are typically larger, heavier, and consume more power, and are also slower than equivalent 2.5D or 3D IC packages because signals travel relatively slowly from one chip to another on the circuit board. Furthermore, 2D IC packages have more potential points of failure because the solder joints on the circuit board are more prone to failure than the electrical connections formed within the interposers. Nevertheless, troubleshooting 2D IC packages is relatively straightforward after the different chips have been mounted onto the circuit board. In particular, the conductive traces that transmit I / O signals between the individual chips on the circuit board are easily accessible and can therefore be used to measure specific I / O signals during troubleshooting.

[0094] In contrast, troubleshooting 2.5D or 3D IC packages is far more difficult because the I / O signals transmitted between different chips are typically embedded in a silicon interposer, making them physically inaccessible. Furthermore, due to the high bandwidth and high density characteristics of 2.5D and 3D IC packages, implementations often involve thousands of conductive traces routing between different chips. An example of such implementations is a memory bus existing between a processor and a high-bandwidth memory chip. In this implementation, even if a probe could be used to physically approach the traces traversing the silicon interposer, accurately and reliably selecting specific conductive traces or combinations of conductive traces for troubleshooting the IC package is extremely difficult, if not impossible.

[0095] In at least one embodiment, one or more parallel processors include circuitry optimized for graphics and video processing, including, for example, video output circuitry, and constitute a graphics processing unit (“GPU”). In at least one embodiment, one or more parallel processors include circuitry optimized for general-purpose processing. In at least one embodiment, components of the computing system may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors, memory hubs, processors, and I / O hubs may be integrated into a System-on-a-Chip (SoC) integrated circuit. In at least one embodiment, components of the computing system may be integrated into a single package to form a System-in-a-Package (“SIP”) configuration. In at least one embodiment, at least a portion of the components of the computing system 300 may be integrated into a multi-chip module (“MCM”) that can interconnect with other MCMs to form a modular computing system. In at least one embodiment, an I / O subsystem and a display device are omitted from the computing system.

[0096] When the light-transmitting medium is silicon, suitable insulators include, but are not limited to, silicon dioxide, and suitable substrates include silicon substrates. Silicon-on-insulator wafers are suitable platforms for optical devices having a silicon light-transmitting medium situated on a substrate having a silicon dioxide insulator and a silicon substrate.

[0097] The device includes one or more waveguides for transmitting optical signals to and / or from optical components. Examples of optical components that may be included on the device include, but are not limited to, one or more components selected from the group consisting of: end faces through which optical signals can enter and / or exit the waveguide; entry / exit ports through which optical signals can enter and / or exit the waveguide from above or below the device; multiplexers for combining multiple optical signals onto a single waveguide; demultiplexers for separating multiple optical signals so that different optical signals are received on different waveguides; optical couplers; optical switches; lasers serving as sources of optical signals; amplifiers for amplifying the intensity of optical signals; attenuators for attenuating the intensity of optical signals; modulators for modulating signals onto optical signals; modulators for converting optical signals into electrical signals; and vias that provide an optical path for optical signals to pass through the device from the bottom to the top of the device. Additionally, the device may optionally include electrical components. For example, the device may include electrical connections for applying a potential or current to the waveguide and / or for controlling other components on the optical device.

[0098] Various embodiments may be described by the following terms: 1. An apparatus comprising: The inlet is used to receive cooling fluid; A cooling circuit for circulating the cooling fluid along a ring-shaped path around a central region of the device, the cooling circuit being close to one or more heat-generating elements; and An outlet is used to discharge the cooling fluid.

[0099] 2. The device according to Clause 1 further includes one or more connecting elements for securing the device to the one or more heating elements.

[0100] 3. The apparatus according to Clause 2, wherein the one or more connecting elements comprise at least one of screws, bolts, or fasteners.

[0101] 4. The device according to Clause 1 further includes a resilient mechanism for achieving compliance between the device and the one or more heating elements.

[0102] 5. The device according to Clause 4, wherein the elastic mechanism is used to compensate for the thermal expansion of one or more heating elements.

[0103] 6. The apparatus according to Clause 1 further includes a second cooling circuit for receiving the cooling fluid from the cooling circuit, the second cooling circuit being adjacent to the cold plate located in the central region.

[0104] 7. The apparatus according to Clause 1 further includes one or more temperature sensors for: Generate one or more temperature data corresponding to the one or more heating elements; and The temperature data is sent to the controller.

[0105] 8. The apparatus according to Clause 1, wherein the inlet is for receiving cooling fluid from at least a cooling distribution unit (CDU), and the outlet is for discharging the cooling fluid to the CDU.

[0106] 9. The apparatus according to Clause 1, wherein the one or more heating elements comprise at least one of a central processing unit (CPU), a graphics processing unit (GPU), a quantum processing unit (QPU), an application-specific integrated circuit (ASIC), or a printed circuit board (PCB).

[0107] 10. A method comprising: A cooling device is provided, which has a cooling circuit for circulating cooling fluid along an annular path around a central region of the cooling device, close to one or more heating elements; The cooling fluid is circulated back into the cooling circuit to dissipate heat from the one or more heating elements; and The cooling fluid is discharged.

[0108] 11. The method according to Clause 10 further comprises: removably attaching the cooling device to the one or more heating elements.

[0109] 12. The method according to Clause 10 further comprises: circulating the cooling fluid from the cooling circuit to a second cooling circuit near the cold plate in the central region.

[0110] 13. The method according to Clause 10, comprising: applying a compressive force to the cooling device via an elastic mechanism to achieve thermal contact between the cooling device and the one or more heating elements.

[0111] 14. The method according to Clause 13, wherein the elastic mechanism compensates for the thermal expansion of the one or more heating elements.

[0112] 15. The method according to Clause 10 further comprises: circulating the cooling fluid at a variable flow rate based at least on the temperature of the one or more heating elements and the cold plate.

[0113] 16. An apparatus comprising a cooling circuit for circulating cooling fluid along a path surrounding a central region of the apparatus, close to a heat-generating element.

[0114] 17. The apparatus according to Clause 16 is further configured to circulate the cooling fluid at a variable flow rate, based at least on the heat load of the heating element and the cold plate.

[0115] 18. The apparatus according to Clause 16 further includes a second cooling circuit for receiving the cooling fluid from the cooling circuit and circulating the cooling fluid close to the cold plate in the central region.

[0116] 19. The device according to Clause 16, wherein the heating element comprises at least one of a central processing unit (CPU), a graphics processing unit (GPU), a quantum processing unit (QPU), an application-specific integrated circuit (ASIC), or a printed circuit board (PCB).

[0117] 20. The device according to Clause 16 further includes a resilient mechanism to achieve compliance between the device and the heating element.

[0118] 21. The apparatus according to Clause 16, wherein the cooling circuit circulates the cooling fluid on top of the heating element.

[0119] Other variations also fall within the spirit and scope of this disclosure. Therefore, while the disclosed technology can be modified and alternatively constructed in various ways, certain exemplary embodiments have been shown in the accompanying drawings and described in detail above. However, it should be understood that this disclosure is not intended to limit its scope to the specific forms or multiple specific forms disclosed, but rather to cover all modifications, alternative constructions, and equivalents that fall within the spirit and scope of this disclosure as defined in the appended claims.

[0120] In the context of describing the disclosed embodiments (especially in the context of the appended claims), the terms “a,” “an,” “the,” and similar pronouns should be interpreted as encompassing both singular and plural unless otherwise stated herein or explicitly denied by the context, and should not be considered as limiting the terms. The terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (meaning “including but not limited to”) unless otherwise stated. “Connection,” when unmodified and referring to a physical connection, should be interpreted as partially or wholly contained, attached to, or linked together, even if something intervenes. The enumeration of numerical ranges herein is intended only as a shorthand method to refer to each individual value falling within a range, unless otherwise stated herein, and each individual value is incorporated into the specification as if it had been individually enumerated herein. In at least one embodiment, unless otherwise stated or denied by the context, the terms “set” (e.g., “set of items”) or “subset” should be interpreted as a non-empty set containing one or more members. Furthermore, unless otherwise stated or denied by the context, a “subset” of a corresponding set does not necessarily mean an appropriate subset of the corresponding set, but rather the subset and the corresponding set can be equal.

[0121] Conjunctive phrases such as “at least one of A, B, and C” or “at least one of A, B, and C”, unless explicitly stated otherwise or explicitly denied by the context, should be understood in conjunction with the context and are generally used to indicate that an item, term, etc., can be A, B, or C, or any non-empty subset of the set consisting of A, B, and C. For example, in an illustrative example of a set containing three members, the conjunctive phrases “at least one of A, B, and C” and “at least one of A, B, and C” refer to any of the following sets: {A}, {B}, {C}, {A,B}, {A,C}, {B,C}, {A,B,C}. Therefore, such conjunctive phrases are generally not intended to imply that some embodiments require at least one A, at least one B, and at least one C to be present. Furthermore, unless explicitly stated otherwise or explicitly denied by the context, the term “multiple” indicates a plural state (e.g., “multiple items” means multiple items). In at least one embodiment, the number of multiple items is at least two, but the number can be more when explicitly stated or determined by the context. Furthermore, unless otherwise stated or the context clearly indicates otherwise, the phrase “based on” means “at least partially based on”, not “based on only”.

[0122] The operations of the processes described herein can be performed in any suitable order unless otherwise stated herein or explicitly denied by the context. In at least one embodiment, processes such as those described herein (or variations thereof and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions and implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that executes jointly on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program containing multiple instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-volatile computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes a non-volatile data storage circuitry system (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, code (e.g., executable code or source code) is stored in a collection of one or more non-volatile computer-readable storage media (or other memory for storing executable instructions), which, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described herein. In at least one embodiment, the collection of non-volatile computer-readable storage media includes multiple non-volatile computer-readable storage media, and one or more individual non-volatile storage media do not contain all the code, while the multiple non-volatile computer-readable storage media collectively store all the code. In at least one embodiment, the executable instructions are executed in such a manner that different instructions are executed by different processors; for example, the non-volatile computer-readable storage media stores the instructions, the main central processing unit (“CPU”) executes some of the instructions, and the graphics processing unit (“GPU”) executes the remaining instructions. In at least one embodiment, different components of the computer system have independent processors, and different processors execute different subsets of instructions.

[0123] In at least one embodiment, the arithmetic logic unit is a set of combinational logic circuits that receive one or more inputs to produce a result. In at least one embodiment, the processor uses the arithmetic logic unit to perform mathematical operations such as addition, subtraction, or multiplication. In at least one embodiment, the arithmetic logic unit is used to perform logical operations such as logical AND / OR or XOR. In at least one embodiment, the arithmetic logic unit is stateless and constitutes physical switching components such as semiconductor transistors arranged as logic gates. In at least one embodiment, the arithmetic logic unit can internally operate as a stateful logic circuit with an associated clock. In at least one embodiment, the arithmetic logic unit can be constructed as an asynchronous logic circuit whose internal state is not stored in an associated register set. In at least one embodiment, the processor uses the arithmetic logic unit to combine operands stored in one or more registers of the processor and produce an output, which can be stored by the processor in other registers or memory locations.

[0124] In at least one embodiment, as a result of processing instructions retrieved by the processor, the processor provides one or more inputs or operands to the arithmetic logic unit (ALU), causing the ALU to produce a result at least in part based on instruction code provided to the inputs. In at least one embodiment, the instruction code provided by the processor to the ALU is at least in part based on instructions executed by the processor. In at least one embodiment, combinational logic in the ALU processes the inputs and produces an output, which is placed on a bus within the processor. In at least one embodiment, the processor selects a destination register, memory location, output device, or output storage location on the output bus to time the processor so that the result produced by the ALU is sent to the desired location.

[0125] Within the scope of this application, the term "arithmetic logic unit" or "ALU" is used to refer to any computational logic circuit that processes operands to produce a result. For example, in this document, the term ALU may refer to a floating-point unit, a DSP, a tensor core, a shader core, a coprocessor, or a CPU.

[0126] Therefore, in at least one embodiment, the computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and the computer system is configured with corresponding hardware and / or software to enable the execution of the operations. Furthermore, the computer system implementing at least one embodiment of this disclosure may be a single device; in another embodiment, it may be a distributed computer system comprising multiple devices that operate in different ways such that the distributed computer system performs the operations described herein, and that a single device does not perform all operations.

[0127] The use of any and all examples or exemplary language provided herein (e.g., "for example") is intended solely to better illustrate embodiments of this disclosure and, unless otherwise stated, does not constitute a limitation on the scope of this disclosure. No language in the specification should be construed as indicating that any unstated element is essential to the practice of this disclosure.

[0128] The terms “coupled” and “connected” and their derivatives may be used in the specification and claims. It should be understood that these terms may not be intended to be synonyms with each other. More precisely, in certain examples, “connected” or “coupled” may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. “Coupled” can also indicate that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0129] Unless otherwise expressly stated, it will be understood that throughout this specification, terms such as “processing,” “calculation,” “operation,” and “determine” refer to the operations and / or processes of a computer or computing system or similar electronic computing device that process and / or transform data represented as physical quantities (e.g., electronic quantities) in the registers and / or memory of the computing system into other data represented as physical quantities in the memory, registers, or other such information storage, transmission, or display devices of the computing system.

[0130] Similarly, the term "processor" can refer to any device or part of a device that processes electronic data from registers and / or memory and transforms that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Furthermore, each process can refer to multiple processes for executing instructions sequentially or in parallel, continuously or intermittently. In at least one embodiment, the terms "system" and "method" are used interchangeably herein, provided that a system can embody one or more methods, and a method can be considered a system.

[0131] In this document, the acquisition, reception, or input of analog or digital data into a subsystem, computer system, or computer-implemented machine may be referred to. In at least one embodiment, the process of acquiring, receiving, or inputting analog and digital data can be implemented in various ways, for example, by receiving data as a parameter to a function call or a call to an application programming interface. In at least one embodiment, the process of acquiring, receiving, or inputting analog or digital data can be implemented by transmitting data via a serial or parallel interface. In at least one embodiment, the process of acquiring, receiving, or inputting analog or digital data can be implemented by transmitting data from a providing entity to an acquiring entity via a computer network. In at least one embodiment, the provision, output, transmission, sending, or presentation of analog or digital data may also be referred to. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter to a function call, an application programming interface, or an inter-process communication mechanism.

[0132] While the description herein illustrates example implementations of the described technologies, other architectures can be used to implement the described functionality, and all such architectures are within the scope of this disclosure. Furthermore, although specific assignments of responsibilities may have been defined above for ease of description, various functions and responsibilities can be assigned and divided in different ways depending on the specific circumstances.

[0133] Furthermore, although the subject matter has been described using language specific to structural features and / or method steps, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or steps described. Rather, the specific features and steps disclosed are exemplary forms for implementing the claims.

Claims

1. An apparatus comprising: The inlet is used to receive cooling fluid; A cooling circuit for circulating the cooling fluid along a ring-shaped path around a central region of the device, the cooling circuit being close to one or more heat-generating elements; and An outlet is used to discharge the cooling fluid.

2. The device according to claim 1, further comprising one or more connecting elements for securing the device to the one or more heating elements.

3. The apparatus of claim 2, wherein the one or more connecting elements comprise at least one of screws, bolts, or fasteners.

4. The device according to claim 1 further includes an elastic mechanism for achieving compliance between the device and the one or more heating elements.

5. The apparatus of claim 4, wherein the elastic mechanism is used to compensate for the thermal expansion of the one or more heating elements.

6. The apparatus of claim 1, further comprising a second cooling circuit for receiving the cooling fluid from the cooling circuit, the second cooling circuit being adjacent to the cold plate located in the central region.

7. The apparatus of claim 1, further comprising one or more temperature sensors for: Generate one or more temperature data corresponding to the one or more heating elements; and The temperature data is sent to the controller.

8. The apparatus of claim 1, wherein the inlet is for receiving cooling fluid from at least the cooling distribution unit (CDU), and the outlet is for discharging the cooling fluid into the CDU.

9. The apparatus of claim 1, wherein the one or more heating elements comprise at least one of a central processing unit (CPU), a graphics processing unit (GPU), a quantum processing unit (QPU), an application-specific integrated circuit (ASIC), or a printed circuit board (PCB).

10. A method comprising: A cooling device is provided, the cooling device having a cooling circuit for circulating cooling fluid along an annular path around a central region of the cooling device, close to one or more heating elements; The cooling fluid is circulated back into the cooling circuit to dissipate heat from the one or more heating elements; and The cooling fluid is discharged.

11. The method of claim 10, further comprising: The cooling device is removably attached to one or more heating elements.

12. The method of claim 10, further comprising: The cooling fluid is circulated from the cooling circuit to a second cooling circuit near the cold plate in the central region.

13. The method of claim 10, comprising: A compressive force is applied to the cooling device through an elastic mechanism to achieve thermal contact between the cooling device and the one or more heating elements.

14. The method of claim 13, wherein the elastic mechanism compensates for the thermal expansion of the one or more heating elements.

15. The method of claim 10, further comprising: The cooling fluid is circulated at a variable flow rate, based at least on the temperature of one or more heating elements and the cold plate.

16. An apparatus comprising a cooling circuit for circulating cooling fluid along a path surrounding a central region of the apparatus, close to a heat-generating element.

17. The apparatus of claim 16 is further configured to circulate the cooling fluid at a variable flow rate, based at least on the heat load of the heating element and the cold plate.

18. The apparatus of claim 16, further comprising a second cooling circuit for receiving the cooling fluid from the cooling circuit and circulating the cooling fluid close to the cold plate in the central region.

19. The apparatus of claim 16, wherein the heating element comprises at least one of a central processing unit (CPU), a graphics processing unit (GPU), a quantum processing unit (QPU), an application-specific integrated circuit (ASIC), or a printed circuit board (PCB).

20. The apparatus of claim 16 further includes an elastic mechanism for achieving compliance between the apparatus and the heating element.

21. The apparatus of claim 16, wherein the cooling circuit circulates the cooling fluid on top of the heating element.