Two-position cooling system for interconnect modules
Patent Information
- Application Number
- CN202610347284.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-20
- Filing Date
- 2026-03-20
- Publication Date
- 2026-09-22
Smart Images

Figure CN122803216A_ABST
Abstract
Description
Technical Field
[0001] At least one embodiment relates to cooling for one or more circuit components. For example, at least one embodiment relates to a two-position cooling system for cooling an interconnect module. Background Technology
[0002] Circuit components such as CPUs, DPUs, and GPUs are typically cooled using air cooling and / or liquid cooling. Some interconnect components used to send and / or receive signals between computing circuits generate heat and are cooled using air cooling and / or liquid cooling. Providing adequate heat transfer from interconnect devices to cooling devices can be challenging. Attached Figure Description
[0003] Various embodiments according to this disclosure will be described with reference to the accompanying drawings, in which: Figure 1A and 1B A simplified side view of a two-position cooling system for interconnecting modules according to at least some embodiments is illustrated.
[0004] Figures 2A-2E The illustration shows a perspective view of components of a two-position cooling system for an interconnect module according to at least some embodiments.
[0005] Figures 3A-3D It is a schematic side view illustrating forces in a two-position cooling system for an interconnect module according to at least some embodiments.
[0006] Figure 3E This is a simplified side view of a ramp feature used in a two-position cooling system for interconnecting modules according to at least some embodiments.
[0007] Figure 4A This is a top view of a computing tray incorporating a two-position cooling system for interconnecting modules, according to at least some embodiments.
[0008] Figure 4B This is a top view of a two-position cooling system for interconnecting modules using a cold plate, according to at least some embodiments.
[0009] Figure 5 This is a flowchart of an example method for using a two-bit cooling system for interconnecting modules, according to at least some embodiments.
[0010] Figures 6A-6B The diagram illustrates a network architecture according to at least some embodiments.
[0011] Figure 7 An example data center cooling system according to at least some embodiments is illustrated.
[0012] Figure 8The illustration shows a schematic diagram of a data center cooling system according to at least some embodiments.
[0013] Figure 9 The illustration depicts a computer system according to at least some embodiments.
[0014] Figure 10 This is a schematic diagram illustrating a computing system according to at least some embodiments.
[0015] Figure 11 An example computing environment according to at least one embodiment is illustrated.
[0016] Figure 12 The illustration shows an example network configuration of components that can be used to implement various aspects of different embodiments.
[0017] Figure 13 An example data center cooling system according to at least some embodiments is illustrated.
[0018] Figures 14A-14B A view is illustrated of a transceiver module operatively coupled to a network adapter according to at least some embodiments.
[0019] Figure 15 Exemplary use cases of transceivers according to at least some embodiments are described. Detailed Implementation
[0020] High-performance computing circuits typically include powerful computing components, such as processing devices (e.g., which may include graphics processing units (GPUs), central processing units (CPUs), data processing units (DPUs), memory, etc.). These high-performance computing components are typically implemented on printed circuit boards (PCBs) that have many other computing components and / or other circuit components. In at least one embodiment, an artificial intelligence (AI) data center infrastructure platform is provided. Examples of AI data center infrastructure platforms include Nvidia® DGX™ SuperPOD™ and DGX™ Foundry. In at least one embodiment, the AI data center infrastructure platform provides accelerated infrastructure and / or scalable performance tailored for AI, such as machine learning (ML) and other high-performance computing (HPC) workloads.
[0021] Transceiver modules can be used to transmit data between computing devices, such as servers, switch units, switch trays, etc. Interconnects between switches at different layers can be accomplished using active optical cables and optical links implemented with a pluggable form factor (also known as "pluggable devices"). Optical interconnects can be configured to connect between chips or between different communication systems. For example, it can provide optical interconnects from a network interface controller (NIC) to a switch, from a switch to a switch, and / or from a chip to a chip. Optical interconnects can be used in a variety of applications, such as switches, processing units (e.g., graphics processing units (GPUs), etc.). Optical interconnects can include optical links (e.g., optical fibers) to transmit optical signals. The bandwidth of an optical interconnect can be scaled by using the same optical link to transmit optical signals comprising multiple wavelengths. When doing so, the transmitter is tuned to generate an optical signal comprising multiple carrier wavelengths. Moreover, each modulator of the modular array can be tuned to receive and modulate the corresponding carrier frequency. Sequencing the multiple wavelengths is crucial to ensuring that the transmitter and receiver correctly convey the data contained in the optical signal.
[0022] In at least one example embodiment, the optical interconnect is part of a data center corresponding to a set of network devices, such as network switches (e.g., Ethernet switches, IP routers, multi-service platforms, various transport network elements, traditional communication equipment, or any other suitable communication system) connected to a set of servers or computing nodes. The switch configuration is used to transmit data between switch ports. The switch configuration includes one or more interconnect circuits that can be arranged in various switch configuration architectures, such as m*m crossbar, Banyan, Benes, Omega, Clos, multiplane, STS, TSI, shared memory, buffered crossbar, any other suitable blocking or non-blocking architecture, or any applicable hybrid architecture. In typical embodiments, the switch configuration is implemented in hardware, which may include field-programmable gate arrays (FPGAs) and / or application-specific integrated circuits (ASICs), and in some implementations, bus interconnects are also included. The data center may follow a networking topology (e.g., a hierarchical networking topology), such as a fat tree topology, a thin fly topology, a dragonfly topology, etc. The data center routes traffic between its network switches and servers, and at least one layer of the data center's topology is coupled to the communication network to allow network traffic to flow between the data center and one or more network devices.
[0023] Optical interconnects may include a substrate and an electro-optic component (VCSEL, photodiode, etc.) supported by the substrate and configured to convert between electrical and optical signals. The optical interconnect may also include a transport block defining a receiving surface configured to receive an optical fiber; and a waveguide configured to transmit optical signals between the electro-optic component and the receiving surface, such that in an operational configuration where the optical fiber is received at the receiving surface, the electro-optic component and the optical fiber are in optical communication.
[0024] While this disclosure illustrates and describes optical interconnects without housings or other protective enclosures, as will be understood by those skilled in the art based on this disclosure, some or all of the optical interconnects may be supported or enclosed by any housing used in the communication system to protect the components supported therein (e.g., as part of a quad small form factor pluggable (QSFP) connector, a small form factor pluggable (SFP) connector, etc.). Furthermore, the substrate housing the electro-optic components may be generally rectangular in shape and / or may be fabricated to fit the dimensions (e.g., size and shape) of any communication system regardless of geometric constraints (e.g., L-shape, square, etc.).
[0025] In various embodiments, the optical interconnect for receiving optical fibers can be implemented in a flip-chip configuration. In some embodiments, the optical interconnect is configured as a flip-chip assembly such that the longitudinal axis of a first thermally adiabatic transition profile of the optical interconnect and the longitudinal axis of a second thermally adiabatic transition profile of the optical interconnect can be collinear. In other words, the orientation of the optical interconnect in this embodiment does not require a mirror or other reflective surface to redirect the optical signal between the electro-optic component and the receiving surface. However, according to this disclosure, it will be apparent to those skilled in the art that the optical interconnect in a flip-chip configuration can also include one or more mirrors (e.g., reflective surfaces) to accommodate optical fibers received at different angles. In some optical interconnect embodiments, assuming that light can remain confined to the waveguide at some bending radii, only the mirror can be operatively configured to redirect the light. Thus, embodiments of field-replaceable modular optical interconnect units configured for reception by a main switch system box are described. The field-replaceable modular optical interconnect unit includes a housing comprising at least a front panel, a rear panel, and a side panel extending between the front and rear panels; a printed circuit board assembly supported within the housing; an optical module supported on the printed circuit board assembly and configured to convert between optical signals and corresponding electrical signals for transmitting or receiving optical signals via optical fiber, respectively; a board-to-board connector disposed on the rear panel of the housing and configured to enable the transmission of electrical signals between the printed circuit board assembly and a main switch system box; and an external connector disposed on the front panel of the housing and configured to engage external optical fibers for transmitting optical signals between the optical module and external components. The field-replaceable modular optical interconnect unit can be configured to be electrically connected to the main switch system box via engagement of the board-to-board connector with a corresponding connector of the main switch system box when the housing is received by the main switch system box.
[0026] In some embodiments, the optical module may be a mid-plate optical module (MBOM), and / or the field-replaceable modular optical interconnect unit may include multiple external connectors. For example, the external connector may be a first external connector, and the field-replaceable modular optical interconnect unit may also include a second external connector disposed on the front panel of the housing and configured to enable the transmission of electrical signals between the printed circuit board assembly and external components connected to the printed circuit board assembly.
[0027] Optical interconnects typically include driver circuitry to drive electro-optical elements, such as optical transmitters (typically with binary signals), waveguides (typically optical fibers), and receivers. In this setup, the optical transmitter typically consumes a large portion of the optical interconnect's power requirements. An optical interconnect typically consists of a transceiver module at each end, adapted to transmit optical information along one or two optical fibers. Each transceiver's transmitter typically includes driver circuitry coupled to a light source and receiver circuitry coupled to a photodetector. Typically, optical fiber is used as the transmission medium, in which case the light source and photodetector are coupled to the fiber. The driver circuitry (typically located on a driver chip) is circuitry customized to generate waveforms suitable for driving the light-emitting device in response to an input signal (typically a binary data stream). The combination of driver circuitry and the light source is called the transmitter. The receiver circuitry (typically located on a receiver chip) is circuitry customized to receive the output from the photodetector and generate a corresponding binary data stream. The combination of receiver circuitry and the photodetector is called the receiver. Typically, the receiver and transmitter provide multiple channels, i.e., the ability to transmit or receive via multiple light sources or photodetectors. Sometimes, the driver and receiver circuits are combined on the same chip, known as a transceiver chip. In addition to driver, receiver, and / or transceiver chips, the optical module may also include other chips and electronics, such as, for example, a microcontroller. Typically, the binary signal used in this optical link is an amplitude-modulated NRZ signal, but other signal types are possible in principle.
[0028] In typical optical interconnects, vertical-cavity surface-emitting laser (VCSEL) diodes are used as optical transmitters to transmit binary data over optical fibers. However, the light source can, in principle, be any suitable light source, and the transmitted waveform can be any suitable waveform used to transmit information. Most optical transmitters have a threshold current above which they essentially begin to emit light. Increasing the current driving through the transmitter from zero to above this threshold can be time-consuming; therefore, typically, a bias current is driven through the light source. Typically, the bias current is set just below the threshold, at the threshold, or above the threshold, but it can also be set much higher. This bias current is usually programmable, so the same circuit design can be used to drive different optical transmitters and / or for different applications. The additional time-varying current modulating the emission from the optical transmitter is called the modulation current.
[0029] In some embodiments, the disclosed technology provides an EO interconnect assembly including a pair of pluggable EO transceivers connected to respective ends of an optical fiber. EO transceivers are typically used to connect network connectivity devices (e.g., remote client switches, network adapters such as network interface controllers (NICs) and host channel adapters (HCAs), smart NICs (NICs with embedded CPUs), network-enabled graphics processing units (GPUs), etc.). The terms "network connectivity device" and "network device" are used interchangeably herein. In some embodiments, the optical interconnect includes a substrate, one or more optical waveguides, one or more first microlenses, one or more second microlenses, and a first and a second mechanical fastening device. Furthermore, when designing optical interconnect modules, it is highly desirable to place the EO component drive circuitry system close to the EO transducer to maintain high signal integrity. However, the heat generated in the drive circuitry system can increase the junction temperature of the transducer, thereby degrading its performance. To address the aforementioned heat dissipation problem, the disclosed optical interconnect module includes a cooling element highly integrated with other components of the optical interconnect module. Specifically, the optical interconnect module typically includes an optical coupling module for coupling optical signals between an optical fiber and a photoelectric transducer. The optical coupling module includes optical coupling elements, such as microlenses or prisms. In the disclosed embodiments, the optical coupling module additionally serves as a substrate for cooling elements. The resulting mechanical design is extremely compact yet highly efficient in heat dissipation. The optical coupling module is also referred to herein as an integrated optical cooling core.
[0030] Transceiver modules are inserted into sockets in computing device chassis to connect to computing components and / or switches within the chassis. Transceiver modules typically generate significant heat and can be cooled via a cold plate using liquid cooling. To improve heat transfer rates, a thermal interface pad can be included between the transceiver module and the cold plate. Spring pressure pushes the cold plate against the surface of the transceiver module, effectively clamping the thermal interface pad between the cold plate and the transceiver module. The thermal interface pad can be made of a fragile thermal interface material (TIM) such as graphene. In conventional arrangements, the top surface of the transceiver module can press against the thermal interface pad on the bottom surface of the cold plate during insertion into the socket. Shear and frictional forces can be generated in the thermal interface pad (when inserting or removing the transceiver module from the socket) and may damage the pad. To prevent damage to the thermal interface pad, a protective layer can be included. However, the protective layer reduces the heat transfer capacity of the thermal interface pad, and because the protective layer increases thermal resistance, it reduces thermal performance, thus impairing overall cooling efficiency. Moreover, the protective layer deteriorates after repeated module insertion, thereby impairing the effectiveness of the cooling system.
[0031] The aspects of this disclosure address the aforementioned deficiencies and other challenges by providing a system for coupling and decoupling transceiver modules in sockets (e.g., switch and / or server chassis) during module insertion in a high-performance system without damaging the thermal interface pads and without including a protective layer (that appropriately manages thermal resistance and mechanical reliability during multiple insertions).
[0032] In some embodiments, the cold plate is supported and / or attached to a sliding member. In some embodiments, the sliding member may be in one of two positions. For example, the sliding member may be in a first position when the transceiver module is not coupled into the socket, and in a second position when the transceiver module is coupled into the socket. When the transceiver module is inserted into the socket, the transceiver module pushes the sliding member, thereby moving the sliding member from the first position to the second position. In the first position, the sliding member (e.g., by applying force to the cold plate using one or more springs) supports the cold plate above and away from the transceiver module. In the second position, the sliding member lowers the cold plate to the transceiver module, such that the top surface of the transceiver module contacts a thermal interface pad on the bottom of the cold plate, thereby achieving effective thermal contact and improving heat dissipation. Heat can then be transferred from the transceiver module to the cold plate via the thermal interface pad. When the transceiver module is fully inserted into the socket, the sliding member moves to the second position (and lowers the cold plate). For example, inserting the transceiver module into the socket may include: a spring-driven push of the transceiver module, which allows the sliding member to transition from a first position to a second position. By supporting the cold plate away from the surface of the transceiver module during insertion (and allowing the cold plate to descend onto the transceiver module only after it is fully inserted), contact between the surface of the transceiver module and the thermal interface pad during insertion can be minimized and / or avoided, thereby eliminating or reducing shear forces on the thermal interface pad, regardless of vertical pressure. Therefore, damage to the thermal interface pad can also be minimized and / or avoided without a protective layer. Consequently, heat can be efficiently transferred from the transceiver module to the cold plate via the thermal interface pad.
[0033] The advantages of this disclosure include, but are not limited to, improved cooling of the transceiver module. Since no protective layer is required on the TIM pad, heat can flow more efficiently from the transceiver module to the cold plate via the TIM pad. Furthermore, damage to the TIM pad can be reduced by supporting the cold plate away from the transceiver module when it is being inserted into or removed from the socket. Because damage to the TIM pad can be reduced, cost savings can be achieved due to a decrease in the frequency of TIM pad maintenance and / or replacement. Additionally, unlike existing solutions, embodiments of this disclosure provide a system that can be repeatedly assembled / disassembled without damage (e.g., the TIM pad).
[0034] Figure 1A and Figure 1B A simplified side sectional view of a two-position cooling system for interconnecting modules according to at least some embodiments is illustrated. Figure 1A The first configuration of the system, 100A, is shown. Figure 1B The second configuration 100B of the system is shown.
[0035] refer to Figure 1A Interconnect module 104 is partially inserted into socket 102. Socket 102 may form a cage in which interconnect module 104 can be inserted. Socket 102 may include features that interact with corresponding features on interconnect module 104 to guide interconnect module 104 when it is inserted into socket 102. In some embodiments, interconnect module 104 is an optical transceiver. Interconnect module 104 may transmit electrical and / or optical signals to and / or receive electrical and / or optical signals from another remote computing unit, such as a server or switch tray. In some embodiments, sliding member 120 associated with socket 102 is configured to slide translationally between a first translational position and a second translational position. Figure 1A The image shows the sliding member 120 in a first translational position. Insertion of the interconnect module 104 into the socket 102 can cause the sliding member 120 to slide from the first translational position to a second translational position. The following references at least to... Figure 1B Further details are discussed. When the interconnect module 104 is inserted into the socket 102 and before the interconnect module 104 engages with the contact area 124 of the sliding member 120, the sliding member 120 supports the cold plate 110 in a first vertical position (e.g., Figure 1A (As shown). Additionally, removing the interconnect module 104 from the socket 102 causes the sliding member 120 to raise the cold plate 110 to a first vertical position. The cold plate 110 can be configured such that when the interconnect module 104 is fully inserted into the socket 102 (as shown). Figure 1B (As shown) and the cold plate 110 cools the interconnect module 104 when it is in its second vertical position in contact with the interconnect module 104. In an alternative arrangement, such as when the socket 102 is oriented vertically or when the cold plate 110 is positioned next to rather than above the interconnect module 104, the first vertical position of the cold plate 110 may correspond to a first horizontal position and the second vertical position may correspond to a second horizontal position. In some embodiments, the cold plate 110 moves along an axis orthogonal to the movement of the interconnect module 104 and / or the movement of the sliding member 120.
[0036] Sliding the sliding member 120 can adjust the height of the cold plate 110. In some embodiments, the sliding member 120 adjusts the distance from the cold plate 110 to the interconnect module 104. In a first vertical position (e.g., a first height, a first distance from the interconnect module, etc.), the cold plate 110 is supported away from the top surface of the interconnect module 104. In some embodiments, when the sliding member 120 is in a first translational position, a gap 132 exists between the top surface of the interconnect module 104 and the bottom surface of the TIM pad 130 attached to the bottom surface of the cold plate 110. The bottom surface of the cold plate 110 may be a cooling surface of the cold plate. In some embodiments, the width of the gap 132 is between approximately 0.2 mm and approximately 0.7 mm. In some embodiments, the gap 132 is approximately 0.6 mm wide. The gap 132 allows the interconnect module 104 to be at least partially inserted (e.g., mostly inserted) into the socket 102 without the top surface of the interconnect module 104 contacting the TIM pad 130. In some embodiments, the TIM pad 130 is made of a thermally conductive material such as graphene. The TIM pad 130 may be fragile. If the interconnect module 104 comes into contact with the TIM pad 130 during insertion (into the socket 102), the interconnect module 104 may damage the TIM pad 130. By supporting the cold plate 110 in a first vertical position (e.g., at a first height, a first distance, etc.), the TIM pad does not come into contact with the interconnect module, and therefore (e.g., due to gap 132) damage to the TIM pad 130 can be avoided.
[0037] In some embodiments, the cold plate 110 may be supported by skids (e.g., wheels) 112 that interact with ramps 122 formed in the sliding member 120. In some embodiments, the skids 112 move along the ramps 122 (e.g., up and down along the ramps 122) during the insertion and removal of the interconnect module 104 from the socket 102. The skids 112 may be at least partially circular protrusions extending from the body of the cold plate 110. The skids 112 may be metal or plastic. In some embodiments, the skids 112 are made of the same material as the body of the cold plate 110. The skids 112 may be formed on both sides of the cold plate 110. In some embodiments, the cold plate 110 includes four skids 112 (e.g., two skids 112 on each side of the cold plate). In some embodiments, the sliding member 120 includes four ramps 122 corresponding to the four skids 112.
[0038] In some embodiments, the cold plate 110 may be locked in a translational manner, allowing it to move only vertically relative to the socket 102 without lateral movement. The sliding member 120 may be substantially restricted to side-to-side motion (as illustrated), corresponding to the direction in which the interconnect module 104 is inserted into and removed from the socket 102. The sliding member 120 translates laterally along a first axis while the cold plate 110 remains stationary along that first axis (e.g., a left-right axis, as illustrated). In some embodiments, the left-right translation of the sliding member 120 along the first axis while the cold plate 110 remains stationary along the first axis may allow the slide rail 112 to optionally move vertically along a ramp 122 along a second axis that may be orthogonal to the first axis. The vertical movement of the slide rail 112 along the ramp 122 may cause the cold plate 110 to move vertically relative to the sliding member 120. In some embodiments, the cold plate 110 moves along a second axis orthogonal to the first axis, and the sliding member 120 may move along the first axis. Figure 1A As shown, the slide rail 112 is provided at the upper part of the ramp 122, so that the cold plate 110 (e.g., by means of the sliding member 120) is supported in a first vertical position.
[0039] As described herein, the first vertical position and the second vertical position refer to the reference frame shown in the figure. However, other orientations are also possible. For example, in some embodiments, the first vertical position and the second vertical position of the cold plate 110 may be changed to a first horizontal position and a second horizontal position (e.g., horizontal relative to the sliding member 120). In some embodiments, the sliding member 120 may move along a first axis, and the cold plate 110 may move along a second axis orthogonal to the first axis. In some embodiments, the second axis is arranged vertically. In some embodiments, the second axis is arranged horizontally. Regardless of the orientation of the second axis (e.g., whether vertical or horizontal), in some embodiments, the second axis is orthogonal to the first axis.
[0040] In some embodiments, the push-back assembly 140 pushes the sliding member 120 to a first translational position. Therefore, the default position of the cold plate 110 can be a first vertical position. The push-back assembly 140 includes a spring 142 that abuts against the push member 144 and a spring mount 146. The spring mount 146 can be rigidly coupled to the socket 102. In some embodiments, the spring (not shown) pushes against the cold plate 110 in a direction toward the interconnect module 104. The spring 142 can apply a spring force to the sliding member 120 to overcome the spring force applied to the cold plate 110.
[0041] refer to Figure 1BWhen the interconnect module 104 reaches a threshold position in the socket 102 (e.g., while the interconnect module 104 is being inserted into the socket 102), the interconnect module 104 may push against the contact area 124 of the sliding member 120. In some embodiments, the threshold position corresponds to the last approximately 2-3 mm of travel before the interconnect module 104 is fully inserted into the socket. In some embodiments, the threshold position is a distance corresponding to full insertion, which corresponds to the difference between a first translational position and a second translational position of the sliding member 120. When the interconnect module 104 is fully inserted into the socket 102 (e.g., past the threshold position), the sliding member 120 may move to the second translational position (…). Figure 1B (As shown in the diagram). In some embodiments, the interconnect module 104 is used to push the sliding member 120 from a first translational position to a second translational position. The spring 142 can be compressed. The spring 142 can move along the sliding member 120 from the second translational position (as shown in the diagram). Figure 1B As shown) to the first translation position ( Figure 1A The direction corresponding to the movement of the sliding member 120 (as shown in the diagram), i.e., the opposite of the movement of the sliding member 120 from the first translational position to the second translational position, applies a spring force to the sliding member 120. If the interconnect module 104 is removed from the socket 102, the spring 142 can push the sliding member back to... Figure 1A The first translation position is shown in the figure.
[0042] The sliding member 120 can be pushed by inserting the interconnect module 104 into the socket 102. In some embodiments, as the sliding member slides translatably from a first translational position to a second translational position, the slide rail 112 can move downwards (e.g., slide) along the ramp 122. Springs (e.g., one or more springs, not shown) can push the cold plate 110 in a direction toward the interconnect module 104 (e.g., downwards) and make thermal contact with the interconnect module 104. Figure 1BAs shown, a slide rail 112 is positioned at the lower portion of the ramp 122, allowing a spring force to push the cold plate 110 downwards to a second vertical position (e.g., a second height). In some embodiments, the spring force pushes the cold plate 110 along an axis perpendicular to the movement of the sliding member 120. In the second vertical position, the TIM pad 130 contacts the top surface of the interconnect module 104, allowing heat to be transferred (e.g., via the TIM pad 130) from the interconnect module 104 to the cold plate. The shear force in the TIM pad 130 caused by the movement of the interconnect module 104 relative to the TIM pad 130 can be distributed across substantially the entire surface of the TIM pad 130. In some embodiments, the thermal interface between the interconnect module 104 and the cold plate 110 has a thermal resistance of less than approximately 0.5 Kelvin per watt (K / W). In some embodiments, the thermal interface between the interconnect module 104 and the cold plate 110 has a thermal resistance of less than approximately 0.2 K / W. In some embodiments, the TIM pad 130 has a thermal conductivity between approximately 15 W / m Kelvin and approximately 45 W / m Kelvin. In some embodiments, the TIM pad 130 has a thermal conductivity of approximately 30 W / m Kelvin. When the cold plate 110 is in Figure 1B In the second vertical position shown, heat can be transferred from the interconnect module 104 to the cold plate 110 (e.g., via the TIM pad 130).
[0043] When the interconnect module 104 is removed from the socket 102, the TIM pad 130 may initially contact the top surface of the interconnect module 104. As the interconnect module 104 is removed from the socket 102, the spring 142 may push the sliding member 120 from a second translational position to a first translational position. As the spring 142 pushes the sliding member, the slide rail 112 may move upward along the ramp 122 (e.g., slide), causing the cold plate 110 to move upward from a second vertical position toward a first vertical position, thereby creating a gap between the top surface of the interconnect module 104 and the TIM pad 130. In some embodiments, the cold plate 110 moves away from the interconnect module 104 along an axis perpendicular to the movement of the sliding member 120. The interconnect module 104 can then be completely removed from the socket without contacting the TIM pad 130. Adjusting the height of the cold plate 110 (e.g., between a first height and / or a second height, between a first distance and / or a second distance from the interconnect module 104, etc.) can minimize the shear force on the TIM pad 130 during the insertion and / or removal of the interconnect module. For example, when the sliding member 120 is in a first translational position, the cold plate 110 can be adjusted to a first height (e.g., a first vertical position, a first distance from the interconnect module 104, etc.), wherein there is a gap between the TIM pad 130 and the interconnect module 104. The interconnect module 104 can be moved (e.g., inserted or removed) without contact with the TIM pad 130. When the sliding member 120 is in a second translational position, the cold plate 110 can be adjusted to a second height (e.g., a second vertical position, a second distance from the interconnect module 104, etc.), wherein there is a gap between the TIM pad 130 and the interconnect module 104. When the interconnect module 104 is fully inserted into the socket 102, the TIM pad 130 can contact the interconnect module 104, thereby minimizing the shear force applied to the TIM pad 130 by inserting the interconnect module 104.
[0044] In some embodiments, Figure 1A and 1B The illustrated system is configured to receive multiple insertions (e.g., into receptacle 102) and removals of interconnect module 104 without damaging TIM pad 130 and / or causing mechanical failure of sliding member 120. For example, interconnect module 104 can be inserted into and / or removed from receptacle 102 at least one hundred times without substantially damaging TIM pad 130.
[0045] Figure 2AA perspective view of the cold plate 110 is illustrated. In some embodiments, the cold plate 110 is used to cool the interconnect module 104. A TIM pad 130 may be attached to the bottom surface of the cold plate 110. As discussed above, the TIM pad 130 may be made of a thermal interface material such as graphene. As discussed above, in the last 2-3 mm of insertion, the module push mechanism and the cold plate 110 move downwards to contact the interconnect module 104. This minimizes shear movement, thereby allowing the use of soft, high-conductivity TIMs, such as graphene-based TIMs, which can significantly reduce thermal resistance by 60-70%. The bottom surface of the cold plate 110 may be a cooling surface. In some embodiments, the cold plate 110 receives a coolant flow (e.g., via a coolant inlet). Heat received at the cooling surface can be transferred to the coolant. The heated coolant can be discharged from the cold plate 110 (e.g., via a coolant outlet). The discharged coolant can be cooled (e.g., via a data center cooling system).
[0046] The cold plate 110 may be formed with multiple members to support it. In some embodiments, a slide rail 112 protrudes from the side of the cold plate 110. The slide rail may be at least partially circular to facilitate sliding along the surface of the ramp 122. In some embodiments, the cold plate 110 may be vertically movable between a first vertical position and a second vertical position. The second vertical position may be lower than the first vertical position. In some embodiments, one or more springs (not shown) apply a downward spring force to the cold plate 110. In some embodiments, the spring force on the cold plate is in a direction orthogonal to the movement of the sliding member 120. The cold plate 110 may be formed with a groove 114 to interface with a spring. For example, the spring may be at least partially located within the groove 114 and may push against the top surface of the cold plate 110. The spring may apply a spring force to the cold plate 110 in a direction toward the interconnect module.
[0047] In at least one embodiment, the cold plate is a metal plate that can be thermally coupled to electronic devices (e.g., computing components) such as CPUs or GPUs, or interconnect modules. In at least one embodiment, the cold plate can achieve localized cooling of electro-electronic devices by transferring heat from the electronic devices to a liquid coolant flowing to a remote heat exchanger. In at least one embodiment, the cold plate comprises a thick metal plate having one or more internal passages (e.g., finger-like structures) through which liquid coolant can flow to dissipate heat from the module. Liquid circulates through the internal passages (e.g., interconnect fingers), thereby serially cooling multiple internal passages. In at least one embodiment, the cold plate can be made of a material such as aluminum, steel, stainless steel, or copper. In at least one embodiment, electronic devices (such as interconnect modules) in contact with the cold plate are cooled by conduction. Heat from the electronic devices can be conducted from the devices to the attached cold plate. Heat can be carried away by the liquid coolant flowing through the cold plate.
[0048] Figure 2B A perspective view of the sliding member 120 is illustrated. In some embodiments, the sliding member 120 can slide translationally between a first translational position and a second translational position. When the interconnecting module 104 is inserted into the socket 102, the interconnecting module 104 can push against the contact area 124 to move the sliding member 120 from the first translational position to the second translational position. In some embodiments, the contact area 124 is formed on the surface of the bridging member 128B. The sliding member 120 includes bridging members 128A and 128B to couple the parallel member 126. The bridging members can provide structural rigidity to the sliding member 120. The parallel members 126 can be generally parallel to each other. The bridging members 128A and 128B can be generally orthogonal to the parallel members 126. In some embodiments, the bridging member 128B forms a cut to provide clearance to coolant lines coupled to the coolant inlet and outlet of the cold plate 110. In some embodiments, the parallel members 126 each form a ramp 122. When the sliding member 120 moves from the first translational position to the second translational position, the slide rail 112 of the cold plate 110 can move along the surface of the ramp 122 to move the cold plate 110 up and down (e.g., between the first vertical position and the second vertical position). In some embodiments, the sliding member 120 is made of a polymer (such as polyetheretherketone (PEEK) plastic, polytetrafluoroethylene (PTFE) plastic, or other suitable plastic). Alternatively, the sliding member 120 may be made of a metal such as aluminum.
[0049] Figure 2C-2EThe illustration shows a perspective view of the push-back assembly 140. When the interconnect device 104 is removed from the receptacle 102, the push-back assembly 140 can push against the sliding member 120 to return the sliding member 120 from a second translational position to a first translational position. In some embodiments, the spring mount 146 is rigidly coupled to the receptacle. The push member 144 can be translationally movable relative to the spring mount 146. The push member 144 may include a guide pin 145 adapted to a hole formed in the spring mount 146 to guide the movement of the push member 144. In some embodiments, a spring 142 is coupled to the spring mount 146 and pushes against the push member 144 to apply a spring force to the sliding member 120. In some embodiments, the push member 144 forms a cutout to provide clearance to coolant lines coupled to the coolant inlet and outlet of the cold plate 110. The spring mount 146 may form a similar cutout / feature. In some embodiments, the actuating member 144 and / or the spring mount 146 are made of a polymer such as PEEK plastic, PTFE plastic, or other suitable plastic. Alternatively, the actuating member 144 and / or the spring mount 146 may be made of a metal such as aluminum. The spring 142 may be made of metal.
[0050] Figures 3A-3D It is a schematic side view illustrating forces in a two-position cooling system for an interconnect module according to at least some embodiments.
[0051] Figure 3A This is a schematic side view 300A illustrating the forces in the two-position cooling system of the interconnect module. The forces can be shown when the cold plate 110 is moved away from the interconnect module 104 (e.g., in a first vertical position, at a first distance from the interconnect module 104, etc.). A first spring force F1 can act on the cold plate 110 in the negative Y direction. The cold plate 110 can transmit the spring force F1 to the sliding member 120. A second spring force F2 can act on the sliding member 120 in the negative X direction. As shown, the X-axis can be horizontal and the Y-axis can be vertical, or the X-axis can be vertical and the Y-axis can be horizontal. The sliding member 120 can apply a reaction force. In some embodiments, the sliding member 120 applies a first reaction force R1 opposite to the first spring force F1 and a second reaction force R2 opposite to the second spring force F2. In some embodiments, the first spring force F1 is provided by a spring pushing against the cold plate 110, and the second spring force F2 is provided by one or more springs 142.
[0052] The forces along the Y-axis and along the X-axis can be modeled as follows: Figure 3BThis is a schematic side view 300B illustrating forces in a two-position cooling system for an interconnect module. The forces are shown when the interconnect module 104 is inserted into the socket 102. In some embodiments, the interconnect module 104 can be pushed into the socket 102 by a force R2. Force R2 can be transmitted to the sliding member 120 when the interconnect module 104 contacts the sliding member 120. One or more springs 142 can push the sliding member 120 back using a spring force F2. Springs can push the cold plate 110 downwards using a spring force F1. In some embodiments, the spring force F1 is between approximately 10 Newtons and approximately 40 Newtons. In some embodiments, the spring force F1 is approximately 20 Newtons. In some embodiments, the spring force F2 is between approximately 20 Newtons and approximately 40 Newtons. In some embodiments, the spring force F2 is greater than about 25 Newtons. In some embodiments, the spring force F2 is approximately 30 Newtons.
[0053] As one or more slide rails 112 move downward along the ramp 122, a force can be applied between the sliding member 120 and the cold plate (e.g., via one or more slide rails 112 and the ramp 122). In some embodiments, the ramp 122 is at an angle α relative to the X-axis. Normal force F n The normal force F can be applied to the ramp 122 by the slide rail 112. In some embodiments, the normal force F n It can be normal to the surface of slope 122. Parallel force F μ The parallel force F can be applied to the ramp 122 by the slide rail 112. In some embodiments, the parallel force F μ The sliding member 120 can apply a reaction force to the cold plate 110 in a direction along the surface of the ramp 122 at an angle α relative to the X-axis. In some embodiments, the sliding member 120 (via the surface of the ramp 122) applies a reaction force F in the negative X direction. rμ and the reaction force F in the positive Y direction r .
[0054] The forces along the Y-axis and along the X-axis can be modeled as follows: Figure 3C This is a schematic side view 300C illustrating the forces in a two-position cooling system for an interconnect module. The forces can be shown when the interconnect module 104 is fully inserted into the socket 102. In some embodiments, a locking mechanism (not shown) of the socket 102 locks the force R. lock The force is applied to the interconnect module 104 (e.g., to lock the interconnect module 104 in a socket). In some embodiments, the interconnect device 104 applies a reaction force R1 to the cold plate 110. The reaction force R1 can be applied from the top surface of the interconnect device 104 to the cooling surface (e.g., the bottom surface) of the cold plate 110.
[0055] The forces along the Y-axis and along the X-axis can be modeled as follows: Figure 3D This is a schematic side view 300D illustrating forces in a two-position cooling system for an interconnect module. The forces can be shown when the interconnect module 104 is removed from the socket 102. Forces can be applied between the sliding member 120 and the cold plate (e.g., via one or more slide rails 112 and the ramp 122) as one or more slide rails 112 move upward along the ramp 122. In some embodiments, the ramp 122 is at an angle α relative to the X-axis. Normal force F n The normal force F can be applied to the ramp 122 by the slide rail 112. In some embodiments, the normal force F n It can be normal to the surface of slope 122. Parallel force F μ The parallel force F can be applied to the ramp 122 by the slide rail 112. In some embodiments, the parallel force F μ The sliding member 120 can apply a reaction force to the cold plate 110 in a direction along the surface of the ramp 122 at an angle α relative to the X-axis. In some embodiments, the sliding member 120 (via the surface of the ramp 122) applies a reaction force F in the positive X-direction. rμ and the reaction force F in the positive Y direction r .
[0056] The forces along the Y-axis and along the X-axis can be modeled as follows: Figure 3EThis is a simplified side view 300E of a ramp feature used in a two-position cooling system for an interconnect module according to at least some embodiments. In some embodiments, a slide rail 112 moves along the surface of a ramp 122 as the sliding member moves between a first translational position and a second translational position. For example, when the sliding member 120 moves to the right, as shown, the slide rail can move downwards along the ramp 122, thereby lowering the cold plate 110. When the sliding member 120 moves to the left, the slide rail can move upwards along the ramp 122, thereby raising the cold plate 110. In some embodiments, the surface of the ramp 122 is at an angle 123 relative to a horizontal reference system (e.g., angle 123 relative to the socket 102). Angle 123 can determine how quickly the cold plate 110 moves along the travel range of the sliding member 120 to contact the interconnect module 104. If angle 123 has a higher value, the cold plate moves up and down (e.g., away from or toward the interconnect module 104) more quickly (relative to the movement of the sliding member 120) to contact the interconnect module 104. If angle 123 has a smaller value, the cold plate moves up and down more slowly (e.g., away from or toward interconnect module 104). In some embodiments, the surface of ramp 122 has two phases, each with its own angle. The first phase may be steep to quickly move the cold plate away from the interconnect module, and the second phase may be gentle to balance the spring force without causing excessive stress on the component parts.
[0057] In some embodiments, the sliding member 120 can slide in a translational manner between approximately 1 mm and approximately 5 mm (e.g., along a movement axis, whether in a vertical or horizontal orientation as shown). The difference between a first translational position and a second translational position of the sliding member 120 can be between approximately 1 mm and approximately 5 mm. In some embodiments, the cold plate can move between approximately 1.5 mm and approximately 2 mm (e.g., along a movement axis orthogonal to the sliding member 120). The difference between a first vertical position and a second vertical position of the cold plate 110 can be between approximately 1.5 mm and approximately 2 mm. The ratio of the distances traveled by the sliding member 120 and the cold plate 110 can depend on angle 123. In some embodiments, angle 123 is between approximately 5 degrees and approximately 45 degrees.
[0058] Figure 4AThis is a top view of a computing tray 400A incorporating a two-position cooling system for interconnect modules, according to at least some embodiments. In some embodiments, the computing tray 400 includes multiple computing components (e.g., GPUs, DPUs, CPUs, switches, and / or interconnect components, etc.) within a chassis 450. The chassis 450 may include side walls, a bottom wall, and / or a top wall. Components for cooling the computing components (such as cold plates, coolant lines, coolant manifolds, etc.) may also be housed within the chassis 450. To connect the computing components to external computing components (e.g., other servers or computing devices), the chassis 450 includes receptacles 402, each receptacle 402 for receiving an interconnect module (e.g., interconnect module 104). The receptacles 402 may be positioned adjacent to each other, and so on. The interconnect modules may be cooled using a cold plate 410.
[0059] Figure 4B This is a top view of a two-position cooling system for interconnecting modules using a cold plate 410, according to at least some embodiments. (See again...) Figure 4A In some embodiments, cold plate 410 receives liquid flow and / or two-phase coolant. Cold coolant may be received into chassis 450 via coolant inlet 444. One or more conduits (e.g., coolant lines) may deliver cold coolant from the inlet to one or more manifolds 442. Coolant may be distributed from manifold 442 to cold plate 410. In some embodiments, coolant flows from manifold 442 into cold plate 410 via inlet 414. Cold coolant received in cold plate 410 may receive heat from interconnect modules. Heated coolant may be discharged from cold plate 410 via outlet 416. Inlet 414 and outlet 416 may keep cold plate 410 substantially stationary within chassis 450. In some embodiments, discharged heated coolant may flow to return manifold 442. In some embodiments, cold plates 410 are connected in series such that coolant flows from one cold plate 410 to adjacent cold plates 410, etc., before being collected into return manifold 442. In some embodiments, the cold plates 410 are connected in parallel such that coolant flows out of the cold plates and is collected in a return manifold 442 without flowing to another cold plate 410. The heated coolant can be collected in the return manifold 442 and (e.g., via one or more conduits) directed to a coolant outlet 446. In some embodiments, the heated coolant exits the casing 450 via the outlet 446 for cooling (e.g., in a cooling tower or refrigeration unit).
[0060] Figure 5This is a flowchart of an example method 500 using a two-position cooling system for interconnecting modules, according to at least some embodiments. Although shown in a specific order or sequence, the order of processes may be modified unless otherwise stated. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes may be performed in different orders, and some processes may be performed in parallel. Additionally, one or more processes may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are also possible.
[0061] At operation 510, the interconnect module is inserted into the socket. In some embodiments, the interconnect module is an optical transceiver. The interconnect module can be used to send / receive electrical and / or optical signals between computing servers or switch trays, etc. In some embodiments, the socket is formed in the side of a computer enclosure (e.g., a server chassis, a switch tray chassis, etc.). The switch within each layer can be a 1U switch, where "1U" refers to the industry-standard size of rack-mount switches and servers. The switch can be an electrical switch, an optical switch, a hybrid electro-optical switch, or any combination thereof. The switch can be implemented using suitable hardware and / or software that enables signal routing in the appropriate domain. For example, an electrical switch may include receivers that receive optical signals and convert the optical signals into electrical signals for routing within the electrical switch. The receivers of an electrical switch may include a transimpedance amplifier (TIA), a photodetector, and a controller, all of which are used to convert optical signals into electrical signals. Each electrical switch may also include a transmitter that converts electrical signals routed within the electrical switch into optical signals for output to another (optical or electrical) switch within the system. For example, the transmitter of an electrical switch may include a light source, a modulator, and a controller that controls the modulator and the light source. In some embodiments, the receiver / transmitter pair may be integrated into a single transceiver. Each electrical switch may also include an internal switching circuitry for routing electrical signals within the electrical switch.
[0062] At operation 520, the sliding member is pushed from a first translational position to a second translational position. In some embodiments, the sliding member is pushed when the interconnect module is inserted into the socket. The interconnect module can be inserted into the socket by a threshold distance before the interconnect module contacts the sliding member and then the sliding member is pushed.
[0063] At operation 530, the cold plate moves (e.g., lowers) from a first distance (e.g., a first vertical position) from the interconnect module to a second distance (e.g., a second vertical position) from the interconnect module, wherein when the cold plate is at the second distance from the interconnect module, the TIM pad on the cold plate contacts the interconnect module. In some embodiments, when the sliding member is in the first translational position, the sliding member supports the cold plate away from the surface of the interconnect module by a first distance. A gap can separate the top surface of the interconnect module from the TIM pad on the cooling surface of the cold plate. In some embodiments, the interconnect module can be at least partially inserted into the socket without contacting the TIM pad. When the sliding member moves to the second translational position, the sliding member can allow the cold plate to move to the second distance. When the cold plate is at the second distance, the TIM pad can contact the surface of the interconnect module. In some embodiments, the sliding member forms one or more ramp features, and a slide rail protruding from the cold plate can slide on one or more ramp features. When the sliding member moves from the first translational position to the second translational position, the slide rail can slide along the ramp features, thereby moving the cold plate from the first distance to the second distance.
[0064] At operation 540, the interconnect module is cooled via a cold plate. In some embodiments, heat from the interconnect module is supplied to the cold plate via a TIM pad on the cooling surface of the cold plate. The heat from the interconnect module can be carried away by coolant flowing through the cold plate.
[0065] In some embodiments, when removing the interconnect module from the receptacle, the sliding member can move from a second translational position to a first translational position. The slide rail of the cold plate can slide along one or more ramp features of the sliding member, thereby moving the cold plate from a second distance to a first distance and separating the cold plate from the interconnect module. At the first distance, a gap may exist between the TIM pad on the cooling surface of the cold plate and the surface of the interconnect module. The interconnect module can be removed from the receptacle without contacting the TIM pad. In some embodiments, the interconnect module can be inserted into and removed from the receptacle multiple times without damaging the TIM pad.
[0066] In some embodiments, when the outlet is horizontally oriented, a first distance (e.g., of the cold plate) corresponds to a first vertical position and a second distance corresponds to a second vertical position, as described above. Other orientations of the outlet are also possible.
[0067] Servers and data centers The following figures illustrate, but are not limited to, exemplary network servers and data center-based systems that can be used to implement at least one embodiment.
[0068] Data centers may include multiple network switches in a specific topology, such as a fat tree topology, a thin fly topology, or a dragonfly topology. The specifications and composition of the network switches in the topology affect the overall network performance of the data center (e.g., bandwidth capacity). In at least one embodiment, an artificial intelligence (AI) data center infrastructure platform is provided. Examples of AI data center infrastructure platforms include Nvidia® DGX. TM SuperPOD TM and DGX TM Foundry. In at least one embodiment, the AI data center infrastructure platform provides accelerated infrastructure and / or scalable performance tailored for AI, such as machine learning (ML) and other high-performance computing (HPC) workloads.
[0069] Data center environment example Data centers and high-performance computing clusters, as described above, are typically formed by various computing components or networked devices, and communication networks formed by electrical and / or optical equipment can be used to enable communication between these networked devices. For example, see reference... Figures 6A-6B Network architecture 600 may include data center 602, communication network 604, and one or more network devices 606. Network architecture 600 may illustrate a general computing architecture, within which more specific systems and / or subsystems may operate. Although network architecture 600 and / or data center 602, in which embodiments of the present disclosure may be implemented, are described below with reference to such embodiments, the present disclosure contemplates that the transceiver resilient devices and techniques described herein may be applicable to any communication implementation without limitation.
[0070] For example, data center 602 may be a centralized facility designed to house computing resources and related components. Data center 602 can operate to support the infrastructure required for advanced computing tasks to achieve efficient, secure, and reliable operation. Data center 602 may include building and structural components, including power, cooling systems, fire suppression systems, and physical security measures configured to maintain optimal operating conditions and / or protect equipment from environmental hazards and unauthorized access. Example data center 602 may include high-performance servers or compute nodes typically arranged in racks and connected via high-speed networks as described herein, such as... Figure 6BThe illustration depicts a high-performance server or computing node. These servers may include processors (e.g., central processing units (CPUs), graphics processing units (GPUs), data processing units (DPUs), etc.), memory (e.g., RAM), and storage solutions (e.g., hard disk drives (HDDs), solid-state drives (SSDs), etc.). The hardware configuration can be designed for parallel processing and high throughput to meet the demands of high-performance computing (HPC) applications.
[0071] Data center 602 may include high-speed network devices (such as network switches, routers, firewalls, etc.) to facilitate fast and secure data transfer between the data center 602 (e.g., between servers or compute nodes) and external networks. Data center 602 can facilitate communication between servers or compute nodes by ensuring efficient data exchange, minimizing latency, and maximizing bandwidth through a network topology. The network topology specifies how various network devices (such as switches and routers) interconnect to enable data flow. By implementing an efficient network topology, data center 602 can support high-performance computing tasks. Examples of various network topologies may include hierarchical networking topologies, such as fat-tree topologies, thin-fly topologies, dragonfly topologies, etc.
[0072] Communication network 604 can communicatively couple data center 602 to one or more network devices 606 and other external devices for data exchange and connectivity. Examples of communication network 604 may include Internet Protocol (IP) networks, Ethernet, InfiniBand (IB) networks, Fibre Channel networks, the Internet, cellular communication networks, wireless communication networks, combinations thereof (e.g., Ethernet Fibre Channel), and variations thereof. The ability of communication network 604 to combine many network types and configurations allows data center 602 to adapt to a variety of application requirements, from general data communication to dedicated HPC tasks. As described herein, communication network 604 can utilize various optical components to establish communication links between components in architecture 600 (e.g., communicatively coupled between components in architecture 700). Thus, communication network 604 may include various optical devices, transceivers, modules, etc., configured to generate optical signals (e.g., provide optical transmitter functionality) and / or receive optical signals (e.g., provide optical receiver functionality).
[0073] One or more network devices 606 may include various computing devices capable of transmitting and receiving signals via communication network 604. The range of network devices 606 can range from personal computing devices to complex server configurations. Examples include personal computers (PCs), laptops, tablets, smartphones, and servers. One or more network devices 606 can facilitate user interaction with data center 602, allowing data to be entered, retrieved, and processed from remote locations. In addition to individual computing devices, one or more network devices 606 may also include collections of servers or collections of additional data centers. For example, these could be other data centers similar to or identical to data center 602. This interconnection can allow the formation of distributed computing environments to improve redundancy, load balancing, and disaster recovery capabilities. By linking multiple data centers, network architecture 600 can leverage geographically dispersed resources to optimize performance and ensure high availability.
[0074] As described herein, data center 602 and / or one or more network devices 606 may include storage devices and processing circuitry for performing computational tasks, such as controlling data flow within and on communication network 604. The processing circuitry may include software, hardware, or a combination thereof. For example, the processing circuitry may include a memory containing executable instructions and a processor (e.g., a microprocessor) that executes those instructions. The memory may correspond to any suitable type of memory device or collection of memory devices configured to store instructions. Non-limiting examples of suitable memory devices include flash memory, random access memory (RAM), read-only memory (ROM), variations thereof, combinations thereof, or similar technologies. In certain embodiments, the memory and processor may be integrated into a common device, such as a microprocessor with integrated memory. Additionally or alternatively, the processing circuitry may include hardware components such as application-specific integrated circuits (ASICs). Other non-limiting examples of processing circuitry systems include integrated circuit (IC) chips, CPUs, GPUs, quantum processing units (QPUs), multiple parallel processing units (PPUs), microprocessors, field-programmable gate arrays (FPGAs), collections of logic gates or transistors, resistors, capacitors, inductors, and diodes. Some or all of the processing circuitry systems may be mounted on a printed circuit board (PCB) or an assembly of PCBs. It should be understood that any suitable type of electrical component or collection of electrical components may be suitable for inclusion in the processing circuitry system. A QPU is configured to perform one or more operations associated with a quantum algorithm. In some embodiments, each of the one or more QPUs may include a plurality of qubits, and the one or more QPUs may communicate with each other via a quantum channel. In some embodiments, each of the plurality of qubits may include local qubits, global qubits, and / or synchronization qubits. In some embodiments, the local qubits of each QPU may be configured to perform one or more operations associated with a quantum algorithm on the QPU associated with the local qubit.
[0075] Additionally, although not explicitly shown, this disclosure envisions that data center 602 and one or more network devices 606 may include one or more communication interfaces for facilitating wired and / or wireless communication between each other and other unillustrated elements of network architecture 600. These communication interfaces may include a variety of technologies, including but not limited to Ethernet ports, fiber optic connections, Wi-Fi® transceivers, Bluetooth® modules, and cellular communication modules for integration and interoperability among various components within network architecture 600.
[0076] Furthermore, this disclosure envisions that network architecture 600 may include additional components and functionalities. For example, the network architecture may include, but is not limited to, additional processing units, dedicated accelerators (such as tensor processing units or TPUs), enhanced security modules, and redundant power supplies. Including these elements may be intended to ensure that network architecture 600 is robust, scalable, and capable of meeting various operational requirements. Any changes, modifications, or adaptations of the elements falling within the spirit and scope of this disclosure are considered to be covered by this disclosure. This includes any combination, sub-combination, or enhancement of the various elements to achieve improvements in performance, reliability, and efficiency in network architecture 600.
[0077] Figure 7 An example data center cooling system 700 according to at least some embodiments is illustrated. In at least one embodiment, system 700 includes a data center 708 having one or more servers 712. In at least one embodiment, the servers 712 are rack-based servers. For example, servers 712 are disposed in one or more racks of data center 708. In at least one embodiment, as referenced above... Figure 1A-1B Each server 712 discussed includes multiple computing components and / or interconnect modules. In at least one embodiment, server 712 includes one or more computing components having a power density greater than a threshold. Power density can be used to describe the relationship between component power and component size. Computing components having a power density greater than the threshold can be referred to as high-power computing components. In at least one embodiment, a high-power computing component can be a processing unit, such as a central processing unit (CPU) or a graphics processing unit (GPU). In at least one embodiment, a high-power computing component can include dedicated or general-purpose processing devices, such as the aforementioned GPUs and CPUs, field-programmable gate arrays (FPGAs), data processing units (DPUs), etc. In at least one embodiment, the high-power computing components of server 712 can release heat exceeding a threshold amount. Similarly, in at least one embodiment, server 712 includes one or more computing components having a power density less than a threshold. Computing components having a power density less than the threshold can be referred to as low-power computing components. In at least one embodiment, the low-power computing components of server 712 can output heat less than a threshold amount. In at least one embodiment, the low-power computing components of the server can include a power supply, motherboard, memory, network interface controller (NIC), solid-state drive or hard disk drive, sound card, etc.
[0078] In at least one embodiment, the computing components and / or interconnect modules of server 712 are cooled via one or more cooling loops. In at least one embodiment, a first cooling loop 714 allows a first coolant to flow to server 712 to cool one or more computing components and / or interconnect modules of server 712. In at least one embodiment, cooling loop 714 includes conduits such as piping and / or tubing systems to allow coolant to flow between cooling distribution unit (CDU) 724 and server 712. In at least one embodiment, first cooling loop 714 allows coolant to flow along pipes, tubing systems, and / or one or more manifolds from first CDU 724 to server 712 and back to first CDU 724. In at least one embodiment, the first coolant can transport heat from server 712 to CDU 724. In at least one embodiment, the first coolant is provided to cooling devices (e.g., cold plates) in server 712 for cooling interconnect modules as described herein.
[0079] In at least one embodiment, the cold plate is a metal plate that can be attached to an electronic device (e.g., a computing component) such as a CPU or GPU. In at least one embodiment, the cold plate is attached to the electronic device by an adhesive such as a thermally applied epoxy resin. In at least one embodiment, the cold plate is attached to the electronic device by one or more mechanical fasteners. In at least one embodiment, the cold plate is supported by a sliding member and / or lowered by a sliding member to contact an interconnect module as described herein. In at least one embodiment, the cold plate can achieve localized cooling of the electro-electronic device by transferring heat from the electronic device to a liquid coolant flowing to a remote heat exchanger. In at least one embodiment, the cold plate comprises a thick metal plate having one or more internal passages through which a liquid coolant can flow. In at least one embodiment, the cold plate may be made of a material such as aluminum, steel, stainless steel, or copper. In at least one embodiment, the electronic device in contact with the cold plate is cooled by conduction. Heat from the electronic device can be conducted from the device to the attached cold plate. Heat can be carried away by the liquid coolant flowing through the cold plate.
[0080] In at least one embodiment, the first coolant is a conductive coolant. In at least one embodiment, the first coolant may include water, deionized water, or refrigerants such as R-134a, R-1234YF, 515B, or any low global warming potential (GWP) coolant or any perfluoroalkyl and polyfluoroalkyl (PFA) compatible coolant. In at least one embodiment, the first coolant comprises a mixture of water and additives, such as a mixture of water and ethylene glycol or a mixture of water and propylene glycol. In at least one embodiment, the first coolant comprises a 25% aqueous solution of deionized propylene glycol. Heat from the high-power computing component is transported by the first coolant to the first CDU 724. In at least one embodiment, the first cooling circuit 714 includes one or more supply conduits (indicated by solid lines) and one or more return conduits (indicated by dashed lines). In at least one embodiment, the first coolant is a single-phase coolant. In at least one embodiment, the first coolant is a two-phase coolant. In at least one embodiment, the first coolant does not evaporate when heated by the first computing component.
[0081] In at least one embodiment, CDU 724 includes a heat exchanger for exchanging heat between a first cooling circuit 714 and a second cooling circuit 732. In at least one embodiment, the second cooling circuit 732 allows another coolant (such as water) to flow from CDU 724 to cooling tower 730 to exchange heat from the first cooling circuit 714 with the surrounding environment. In at least one embodiment, the second cooling circuit 732 allows coolant to flow from CDU 724 to one or more coolers to exchange heat with a cold source such as the surrounding environment. In at least one embodiment, the surrounding environment includes an air environment or a liquid environment. In at least one embodiment, CDU 724 includes one or more pumps for pumping the first coolant and / or another coolant into the second cooling circuit 732. In at least one embodiment, CDU 724 includes a controller for controlling the flow and / or distribution of coolant along the first cooling circuit 714 and / or the second cooling circuit 732. In at least one embodiment, CDU 724 includes one or more valves to implement this control. In at least one embodiment, the second cooling circuit 732 may be referred to as the “main” cooling circuit, while the first cooling circuit 714 may be referred to as the “auxiliary” cooling circuit.
[0082] Figure 8 A schematic diagram of a data center cooling system 800 according to at least some embodiments is shown. Figure 8 The diagram shown in the figure is the same as Figure 7Features illustrated in the diagram with similar numbers may have similar functions. In at least one embodiment, system 800 includes a plurality of servers 812 disposed in a data center rack 810. In at least one embodiment, the servers 812 are part of a data center having multiple racks 810, each rack supporting multiple servers 812. Although Figure 8 Only one rack 810 is shown, but system 800 can provide cooling for servers 812 in multiple racks 810.
[0083] In at least one embodiment, a first coolant flows from CDU 824 to server 812 along one or more flow paths of one or more first cooling loops. In at least one embodiment, the first coolant flows through one or more manifolds. In at least one embodiment, the first coolant is supplied from CDU 824 to supply manifold 844. In at least one embodiment, supply manifold 844 can distribute the first coolant to multiple servers 812 supported in rack 810. In at least one embodiment, the first coolant can flow from supply manifold 844 into server 812 to cool one or more high-power computing components and / or one or more interconnect modules of server 812. The first coolant can flow to one or more cold plates in server 812 and / or to one or more cooling devices for cooling interconnect modules, as described herein.
[0084] In at least one embodiment, the high-power computing components and / or interconnect devices within server 812 may each be coupled to one or more cold plates to receive a first coolant. In at least one embodiment, one or more cold plates may transfer heat from the high-power computing components or interconnect modules to the first coolant. In at least one embodiment, the first coolant carries heat away from the high-power computing components and / or interconnect modules of server 812. In at least one embodiment, a return manifold 846 collects the heated first coolant output from each of the servers 812. In at least one embodiment, the first coolant flows from the return manifold 846 to CDU 824, where the first coolant is cooled by chilled water or other coolant flowing between cooler 832 and CDU 824. In at least one embodiment, heat may be exchanged between the first coolant and cooling water or other coolant in a liquid-liquid heat exchanger within CDU 824. In at least one embodiment, the cooled first coolant again flows from CDU 824 to server 812 via supply manifold 844.
[0085] In at least one embodiment, water or another coolant flows between the chiller 832 and the CDU 824 via a second cooling circuit. In at least one embodiment, cold air 831 is drawn into the chiller 832 by one or more fans. In at least one embodiment, the water or other coolant carrying heat transferred from the first cooling circuit (e.g., via a heat exchanger in the CDU 824) is cooled by the cold air 831. In at least one embodiment, heat from the water or other coolant is transferred to the air, and the heated air 833 (by one or more fans) is exhausted from the chiller 832. In at least one embodiment, the water or other coolant is thus cooled by the air. In at least one embodiment, the cooling water flows back to the CDU 824 along the flow path of the second cooling circuit.
[0086] Figure 9 A computer system 900 according to at least one embodiment is illustrated. In at least one embodiment, the computer system 900 is configured to implement the various processes and methods described throughout this disclosure.
[0087] In at least one embodiment, the computer system 900 includes, but is not limited to, at least one central processing unit (“CPU”) 902 connected to a communication bus 910 implemented using any suitable protocol, such as PCI (“Peripheral Component Interconnect”), PCI-Express (“PCI-Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or one or more point-to-point communication protocols. In at least one embodiment, the computer system 900 includes, but is not limited to, main memory 904 and control logic (e.g., implemented in hardware, software, or a combination thereof), and data is stored in main memory 904, which may be in the form of random access memory (“RAM”). In at least one embodiment, a network interface subsystem (“network interface”) 922 provides an interface to other computing devices and networks for receiving data from the computer system 900 and transferring data to other systems.
[0088] In at least one embodiment, the computer system 900 includes, but is not limited to, an input device 908, a parallel processing system 912, and a display device 906, which may be implemented using conventional cathode ray tube (“CRT”), liquid crystal display (“LCD”), light-emitting diode (“LED”), plasma display, or other suitable display technologies. In at least one embodiment, user input is received from the input device 908, such as a keyboard, mouse, touchpad, microphone, etc. In at least one embodiment, each of the foregoing modules may reside on a single semiconductor platform to form the processing system.
[0089] In at least one embodiment, a computer program in the form of machine-readable executable code or computer control logic algorithms is stored in main memory 904 and / or secondary memory. If executed by one or more processors, the computer program enables system 900 to perform various functions according to at least one embodiment. Memory 904, storage devices, and / or any other storage devices are possible examples of computer-readable media. In at least one embodiment, secondary storage devices can refer to any suitable storage device or system, such as hard disk drives and / or removable storage drives, representing floppy disk drives, magnetic tape drives, optical disk drives, digital versatile disc (“DVD”) drives, recording devices, Universal Serial Bus (“USB”) flash memory, etc. In at least one embodiment, the architecture and / or functionality of the various preceding figures are implemented in the context of: CPU 902; parallel processing system 912; integrated circuits having at least a portion of the capabilities of two CPUs 902; chipsets (e.g., groups of integrated circuits designed to operate and be marketed as units for performing related functions); and any suitable combination of one or more integrated circuits.
[0090] In at least one embodiment, the architecture and / or functionality of the various preceding figures are implemented within the context of general-purpose computer systems, circuit board systems, game console systems dedicated to entertainment purposes, special-purpose systems, etc. In at least one embodiment, computer system 900 may take the form of a desktop computer, laptop computer, tablet computer, server, supercomputer, smartphone (e.g., wireless handheld device), personal digital assistant (“PDA”), digital camera, vehicle, head-mounted display, handheld electronic device, mobile phone device, television, workstation, game console, embedded system, and / or any other type of logic.
[0091] In at least one embodiment, the parallel processing system 912 includes, but is not limited to, multiple parallel processing units (“PPUs”) 914 and associated memory 916. In at least one embodiment, the PPUs 914 are connected to a host processor or other peripheral devices via interconnects 918 and switches 920 or multiplexers. In at least one embodiment, the parallel processing system 912 distributes computational tasks across the parallelizable PPUs 914—for example, as part of distributing computational tasks across multiple graphics processing units (“GPUs”) thread blocks. In at least one embodiment, although such shared memory may incur a performance penalty compared to using memory and registers residing locally on the PPUs 914, memory can be shared and accessed (e.g., for read and / or write access) across some or all of the PPUs 914. In at least one embodiment, the operation of the PPUs 914 is synchronized using commands such as syncthreads(), where all threads in a block (e.g., threads executing across multiple PPUs 914) must reach a certain point in code execution before continuing execution.
[0092] Figure 10 This is a schematic block diagram of a computing system 1000 (e.g., a data center or high-performance computing (HPC) cluster) according to embodiments described herein. According to at least one embodiment, system 1000 includes multiple subsystems, such as multiple processing devices, multiple network devices, and multiple networks coupled to each other. The computing system 1000 is designed with multiple integrated circuits (referred to as processing devices), each of which may include one or more CPUs and GPUs, thereby forming a powerful and flexible architecture.
[0093] Various processing devices are interconnected via NVLink or other high-speed interconnects to enable high-speed communication between subsystems; and are also connected via NICs or DPUs to ensure efficient data transfer across computing system 1000 and one or more external networks 1030, 1036. In this example, system 1000 includes a packet switch 1048 that connects NIC / DPU 1028 to network 1030 and a packet switch 1050 that connects NIC / DPU 1032 to network 1036.
[0094] Seamless data exchange and parallel processing are enabled through NVLink-coupled processing devices, thereby improving overall computing performance. The processing devices connect to multiple networks via one or more Network Interface Controllers (NICs) or Data Processing Units (DPUs), allowing the system to handle complex multi-network tasks with high bandwidth and low latency. This configuration is ideal for demanding applications requiring significant processing power, such as artificial intelligence (AI), machine learning (ML), and data-intensive computing, while ensuring robust connectivity and scalability across diverse networking environments. The integrated circuits of the Computing System 1000 may include one or more CPUs and one or more GPUs.
[0095] Figure 10 An example architecture of a multi-GPU architecture is also demonstrated. As illustrated, the computing system 1000 includes a processing device 1002 with a multi-GPU architecture. Specifically, the processing device 1002 may be a system-on-a-chip and includes multiple subsystems such as a CPU 1006, a GPU 1008, and a GPU 1010. The CPU 1006 may be coupled to the GPU 1008 via die-to-die (D2D) or chip-to-chip (C2C) interconnects 1012 (such as a ground reference signaling interconnect (GRS interconnect)). The CPU 1006 may be coupled to the GPU 1010 via a D2D or C2C interconnect 1014. The CPU 1006 may also be coupled to the GPU 1008 and GPU 1010 via a PCIe interconnect.
[0096] The CPU 1006 can be coupled to one or more NICs or DPUs, which in turn are coupled to one or more networks. For example, as Figure 10 As illustrated, CPU 1006 is coupled to a first NIC / DPU 1026, which is coupled to network 1030. CPU 1006 is also coupled to a second NIC / DPU 1028, which is coupled to network 1030 via switch 1048. For example, NIC / DPU 1026 and NIC / DPU 1028 can be coupled to network 1030 via Ethernet, NVLink, or InfiniBand (IB) connections.
[0097] The computing system 1000 also includes a processing device 1004 with a multi-GPU architecture. Specifically, the processing device 1004 includes multiple subsystems, including a CPU 1016, a GPU 1018, and a GPU 1020. The CPU 1016 can be coupled to the GPU 1018 via a D2D or C2C interconnect 1022. The CPU 1016 can be coupled to the GPU 1020 via a D2D or C2C interconnect 1024. The CPU 1016 can also be coupled to the GPU 1018 and GPU 1020 via a PCIe interconnect. The CPU 1016 can be coupled to one or more NICs or DPUs, which in turn are coupled to one or more networks. For example, as... Figure 10 As illustrated, CPU 1016 is coupled to a first NIC / DPU 1032, which is coupled to network 1036. CPU 1016 is also coupled to a second NIC / DPU 1034, which is coupled to network 1036 via switch 1050. NIC / DPU 1032 and NIC / DPU 1034 can be coupled to network 1036 via Ethernet, NVLink, or InfiniBand (IB) connections.
[0098] In at least one embodiment, processing device 1002 and processing device 1004 can communicate with each other via NIC / DPU 1038 (such as via PCIe interconnect). Processing device 1002 and processing device 1004 can also communicate with each other via high-bandwidth communication interconnect 1040 (such as NVLink interconnect or other high-speed interconnect). Figure 10 The packet switches in the diagram can include, for example, Nvidia Quantum-2 switches. The NIC / DPU in the diagram can include, for example, Nvidia Bluefield DPUs.
[0099] Figure 11 An example computing environment according to at least one embodiment is illustrated.
[0100] Figure 11An example computing environment 1100 is illustrated, in which forward pass-off to available memory can be performed. It should be understood that embodiments of this disclosure can also be used in alternative environments, and specific discussions of components are provided by way of non-limiting examples and may include equivalents. Furthermore, various features are omitted for clarity and brevity. Additionally, the systems and methods can be used with a variety of different architectures. Example computing environment 1100 may include server 1102, which can be used to perform HPC workloads such as AI training or machine learning model training. In one embodiment, server 1102 may be an application instance or a compute node. Server 1102 may include CPU 1110 associated with a switch 1120, such as a Peripheral Component Interconnect Fast (PCIe) switch, which can control at least some data transfers via communication paths interconnecting various components. In one embodiment, CPU 1110 may include a root complex processor.
[0101] PCIe switch 1120 may also be associated with GPU 1130 and DPU 1140, and may transfer data between at least some of CPU 1110, GPU 1130, DPU 1140, and other components. In one embodiment, PCIe switch 1120 may be associated with more than one GPU or more than one DPU. In another embodiment, PCIe switch 1120 may be located within DPU 1140. PCIe switch 1120 may manage the transfer of at least some of the data between CPU 1110, GPU 1130, and DPU 1140. In another embodiment, the number of GPUs associated with PCIe switch 1120 may be equal to the number of DPUs associated with PCIe switch 1120. In at least one embodiment, server 1102 may include, but is not limited to, any number of CPUs 1110, PCIe switch 1120, GPUs 1130, and / or DPUs 1140 in any combination. For example, in at least one embodiment, server 1102 may include eight, sixteen, thirty-two, and / or more GPUs 1130. In at least one embodiment, various components (including but not limited to) are interconnected. Figure 11 The communication paths of the CPU 1110, PCIe switch 1120, GPU 1130 and DPU 1140 can be implemented using any suitable protocol, such as peripheral component interconnect (PCI) based protocols (e.g., PCIe) or other bus or point-to-point communication interfaces and / or one or more protocols (e.g., NV-Link high-speed interconnect or interconnect protocols)).
[0102] DPU 1140 may include a network interface controller (NIC) 1142, DDR memory 1144, and a non-volatile memory fast (NVMe) device 1146. NIC 1142 is capable of interfaced with network 1104, which may also interface (e.g., via fabric) with additional NVMe devices available to DPU 1140. In one embodiment, DPU 1140 may not include NVMe device 1146. In another embodiment, NVMe device 1146 may reside on server 1102 instead of DPU 1140. In yet another embodiment, computing environment 1100 may include more than one NVMe device 1146, such as a first NVMe device in DPU 1140 and a second first NVMe device on server 1102 directly associated with PCIe switch 1120. In one embodiment, DPU 1140 may not include DDR memory 1144 and may include compute storage service (CSS) in place of DDR memory 1144, or may include compute storage service (CSS) in addition to DDR memory 1144. For example, computing environment 1100 may include DPU compute storage (CS) memory 1106 available to DPU 1140 as part of the CSS. Network 1104 is capable of interfaced with DPU CS memory 1106 via NIC 1142 according to any suitable interface protocol, such as Remote Direct Memory Access (RDMA) via Ethernet, InfiniBand, Fibre Channel, etc.
[0103] The total memory available for data storage in the computing environment 1100 can be expanded by using DPU 1140 on system nodes. DPU 1140 can access memory pools 1150 already available on server 1102, such as dual data rate (DDR) memory, onboard NVMe devices, architecture-based NVMe devices, and CS. Memory pool 1150 may include at least one of DDR memory 1144, NVMe devices 1146, and DPU CS memory 1106. DPU 1140 is also able to access available memory of other DPUs that are part of pool 1150, and other DPUs can access available memory of DPU 1140, such as pool 1150. This available memory can be accessed and used for data storage without adding computing resources (such as compute nodes), which would be required with other solutions. Server 1102 can be supplied with an available pool 1150 accessible to DPU 1140 to expand the total memory available for data storage, such as reducing the data storage load on CPU 1110 or GPU 1130, which can actually increase the utilization of the memory they use for processing. For example, during AI training, model states, residual states, activation functions, and checkpoints can be stored or offloaded to pool 1150 accessible to DPU 1140.
[0104] Figure 12An example network configuration 1200 is illustrated, comprising components that can be used to implement various aspects of different embodiments, such as providing, generating, modifying, encoding, processing, fusing, and / or transmitting generated image data, calculated measurement results, or other such content. In at least one embodiment, client device 1202 may use components of content application 1204 on client device 1202 and data locally stored on the client device to generate or receive data for a session. In at least one embodiment, content application 1224 executing on computer or processor 1220 (e.g., cloud server or control system) may initiate a session associated with at least one client device 1202 (e.g., vehicle or robot), and may cause content such as liquid coolant or server thermal data to be selected and / or retrieved from repository 1234 for use by test module 1232 to calculate one or more performance metrics for monitoring module 1228, which may provide flow or thermal data to control flow or temperature in an environment where the data is to be used to determine appropriate operation. Content manager 1226 can work with these various modules to perform tests and analyses, and may instruct any actions to be taken in response to performance metrics failing to meet operational requirements. At least a portion of the data or instruction content can be transmitted to client device 1202 / or physical device 1270 using appropriate delivery manager 1222 for transmission via download, streaming, or another such delivery channel. Encoders can be used to encode and / or compress at least some of the data before transmission to client device 1202. In at least one embodiment, client device 1202 receiving such content can provide the content to a corresponding content application 1204, which may also or alternatively include a graphical user interface 1210, a flow control module 1212, and a control module 1214 for providing, composing, rendering, combining, modifying, or using the content on or by client device 1202 for presentation, navigation, control (or other purposes), such as transmission to physical device 1270. In some embodiments, the computer / processor 1220 and the client device 1202 can communicate directly without transmitting data over the network 1240 to avoid issues such as latency and availability. The decoder can also be used to decode data received over the network 1240 for presentation via the client device 1202, such as presenting image content or performance metrics via a display device 1206 and presenting audio (such as corresponding voice or synthesized speech) via at least one audio playback device 1208 (such as a speaker or headphones).In at least one embodiment, at least some of the content may already be stored on, rendered on, or accessible to client device 1202, such that at least that portion of the content does not need to be transmitted over network 1240, for example, where the content (e.g., hot data) may have previously been downloaded or locally stored on a hard drive or optical disc. In at least one embodiment, a transmission mechanism such as data streaming may be used to transmit the content from computer / processor 1220 or user database 1236 to client device 1202. In at least one embodiment, at least a portion of the content may be obtained, enhanced, and / or streamed from another source (such as third-party service 1260 or other client device 1250), which may also include a content application for generating, updating, enhancing, or providing map content. In at least one embodiment, the various parts of the functionality may be executed using multiple computing devices or multiple processors (such as a combination of CPU and graphics processing unit (GPU)) within one or more computing devices.
[0105] In at least some of these examples, the client device can include any suitable computing device, such as a desktop computer, laptop computer, set-top box, streaming device, game console, smartphone, tablet computer, VR headset, AR goggles, wearable computer, or smart TV. Each client device can submit requests across at least one wired or wireless network (such as the Internet, Ethernet, Local Area Network (LAN), or cellular network, and other such options). In this example, these requests can be submitted to an address associated with a cloud provider that can operate or control one or more electronic resources in a cloud provider environment (such as a data center or server farm). In at least one embodiment, the request can be received or processed by at least one edge server located at the network edge and outside at least one security layer associated with the cloud provider environment. This reduces latency by allowing client devices to interact with a closer server, while also improving the security of resources in the cloud provider environment.
[0106] In at least one embodiment, such a system can be used to monitor or manage the thermal conditions of a server, including a cold plate serving as a liquid manifold. In other embodiments, such a system can be used for other purposes, such as providing control over the flow of liquid coolant or performing deep learning operations. In at least one embodiment, such a system can be implemented using an edge device or can be incorporated into one or more virtual machines (VMs). In at least one embodiment, such a system can be implemented at least partially in a data center or at least partially using cloud computing resources.
[0107] Cold plate Figure 13 An example data center cooling system 1300 according to at least some embodiments is illustrated. In at least one embodiment, the cold plate includes adjustable fins forming microchannels through which fluid flows. In at least one embodiment, the fins in the cold plate enable heat transfer from at least one associated computing device to fluid flowing through the microchannels formed between multiple fins. In at least one embodiment, the fins of the cold plate are dynamically adjustable in real time to allow more heat to be transferred from at least one computing device to the fluid flowing through the finned cold plate. In at least one embodiment, such fins may be adjusted by a processor system or a processorless system in part based on a temperature determined (e.g., sensed) for the cold plate. In at least one embodiment, the temperature may be associated with at least one computing device, the workload of at least one computing device, or different time periods, and the fluid at the inlet and outlet of the cold plate. In at least one embodiment, the processorless system may rely on the thermal properties of at least two materials used to form the fins for the cold plate, thereby enabling such fins to react without a processor, allowing more surface area to be exposed to the fluid. In at least one embodiment, such fins may include overlapping portions that can be exposed by the action of a control mechanism or by the properties of at least two materials associated together to form the fins.
[0108] Cooling system In at least one embodiment, it is possible to utilize, such as Figure 13 The illustrated exemplary data center 1300 has an improved cooling system subject to the description herein. In at least one embodiment, numerous specific details are set forth to provide a thorough understanding; however, the concepts herein can be practiced without one or more of these specific details. In at least one embodiment, the data center cooling system can respond to sudden high heat demands caused by changes in the computing load in today's computing components. In at least one embodiment, because these requirements change or tend to vary from minimum to maximum across different cooling requirements, an appropriate cooling system must be used to meet these requirements economically. In at least one embodiment, a liquid cooling system can be used for medium to high cooling requirements. In at least one embodiment, high cooling requirements are met economically through localized immersion cooling. In at least one embodiment, these different cooling requirements also reflect different thermal characteristics of the data center. In at least one embodiment, the heat generated from these components, servers, and racks is collectively referred to as thermal characteristics or cooling requirements, because cooling requirements must fully address thermal characteristics.
[0109] In at least one embodiment, a data center liquid cooling system is disclosed. In at least one embodiment, this data center cooling system addresses the thermal characteristics of associative computing or data center equipment, such as graphics processing units (GPUs), switches, dual in-line memory modules (DIMMs), or central processing units (CPUs). In at least one embodiment, these components may be referred to herein as high heat density computing components. Furthermore, in at least one embodiment, the associative computing or data center equipment may be a processing card having one or more GPUs, switches, or CPUs thereon. In at least one embodiment, each of the GPU, switch, and CPU may be a heat-generating characteristic of the computing device. In at least one embodiment, the GPU, CPU, or switch may have one or more cores, and each core may be a heat-generating characteristic.
[0110] In at least one embodiment, the cold plate includes adjustable fins forming microchannels through which fluid flows. In at least one embodiment, the fins in the cold plate enable heat transfer from at least one associated computing device to fluid flowing through the microchannels formed between multiple fins. In at least one embodiment, the fins of the cold plate are dynamically adjustable in real time to allow more heat to be transferred from at least one computing device to the fluid flowing through the finned cold plate. In at least one embodiment, such fins may be adjusted by a processor system or a processorless system, in part, based on a temperature determined (e.g., sensed) for the cold plate. In at least one embodiment, the temperature may be associated with at least one computing device, the workload of at least one computing device, or at different time periods, and with the fluid at the inlet and outlet of the cold plate. In at least one embodiment, the processorless system may rely on the thermal properties of at least two materials used to form the fins for the cold plate, allowing such fins to react without a processor to expose more surface area to the fluid. In at least one embodiment, such fins may include overlapping portions that may be exposed by the action of a control mechanism or by the properties of the at least two materials associated together to form the fins.
[0111] In at least one embodiment, the cold plate has a top plate, a bottom plate, and fins therebetween. In at least one embodiment, the bottom plate may be a base for the cold plate. In at least one embodiment, the top plate may be located between a cover plate of the cold plate and the bottom plate or substrate of the cold plate. In at least one embodiment, the fins may be coupled to the bottom plate or substrate and the top plate, such that the top plate can be moved to expose the overlapping portion of each fin and expose the overlapping portion of each fin to fluid flowing through the cold plate. In at least one embodiment, the exposure of the overlapping portion of each fin results in the previously covered surface area being exposed to the fluid and provides additional cooling to the fins of the cold plate, and provides additional cooling to the fins of the associated computing device through association.
[0112] In at least one embodiment, multiple fins form microchannels for fluid to flow therethrough. In at least one embodiment, such fins are capable of actively or passively responding to thermal feedback by modifying the microchannels, allowing the fluid to absorb more heat from at least one computing device. In at least one embodiment, the active response can be implemented by at least one processor that can expose more surface area of such fins by expanding the overlapping portions of such fins. In at least one embodiment, the passive response can be achieved by the thermal properties of the materials associated with each fin in such fins, which allow such fins to expand.
[0113] In at least one embodiment, the problem of the cold plate being a static device is addressed by the intelligent dynamic cold plate described herein. In at least one embodiment, the intelligent dynamic cold plate allows the cold plate (via its internal features) to respond to a temperature sensed or determined from at least one computing device. In at least one embodiment, the intelligent aspect of the intelligent dynamic cold plate, compared to a static cold plate, allows for modification of the fins of such a cold plate using sensor input. In at least one embodiment, it is possible to alter the microchannels formed by such fins to cause changes in fluid paths or increase the interaction surface between the fluid and each such fin. In at least one embodiment, such aspects allow more fluid to pass through certain areas, or through certain areas and allow heat to be removed from those areas where high heat density occurs from at least one computing device.
[0114] In at least one embodiment, a static cold plate for liquid cooling of GPUs, CPUs, switches, and other high heat density components may have static microchannels that allow fluid to flow therethrough to remove heat from such heat-dissipating components in a data center. In at least one embodiment, the static cold plate incorporates designs and methods for heat removal that are independent of or do not respond to the heat density or heat dissipation of the computing components. In at least one embodiment, some cold plates may have varying thermal behavior based on computing properties, environmental properties, and other properties that require dynamic behavior to achieve optimal heat removal capabilities for available resources in a liquid-cooled environment.
[0115] In at least one embodiment, the liquid-cooled indirect cooling plate can be composed of components therein. In at least one embodiment, two components can be provided, such that the bottom or base layer or component provides a rigid mechanical attachment and also serves as a highly thermally conductive medium, wherein heat is conducted from computing components (GPUs, switches, and CPUs) to multiple fins forming microchannels on the bottom or base plate or component. In at least one embodiment, such fins are built into a base metal and can be associated with a top or upper plate or component, which is an intermediate plate or component. In at least one embodiment, the top or upper plate or component (as an intermediate plate or component) is non-conductive and made of a non-conductive material. In at least one embodiment, a cover plate (along with side plates) encloses these components within the smart dynamic cooling plate.
[0116] In at least one embodiment, the top plate or upper plate or component can dynamically modify the microchannel fluid path of the substrate, partially based on the instantaneous thermal behavior of the heating assembly associated with the smart dynamic cold plate, through the function of its intermediate plate or component. In at least one embodiment, using smart sensing, inference, and adaptive modification, the microchannels forming the fluid path can be dynamically adjusted to direct more fluid (for heat removal) toward or surge toward high heat density areas within the smart dynamic cold plate. In at least one embodiment, such features also enable the simultaneous reduction of fluid flow to block microchannels in areas of lower heat removal demand determined or sensed from the substrate of the smart dynamic cold plate. In at least one embodiment, the overlapping portions of fins disposed within the smart dynamic cold plate can be used to block or constrain microchannels by making them thicker at the overlapping portions between the fins. In at least one embodiment, multiple top plates can be provided to act as intermediate plates, and each top plate can be associated with different fins. In at least one embodiment, movement of different top plates achieves different blocking, constraint, or flow redirection within the smart dynamic cold plate.
[0117] In at least one embodiment, it is possible to utilize, such as Figure 13The illustrated exemplary data center 1300 has a cooling system affected by the improvements described herein. In at least one embodiment, the data center 1300 may be one or more rooms 1302 having racks 1310 and auxiliary equipment to accommodate one or more servers on one or more server trays. In at least one embodiment, the data center 1300 is supported by a cooling tower 1304 located outside the data center 1300. In at least one embodiment, the cooling tower 1304 dissipates heat from the data center 1300 by acting on a main cooling circuit 1306. In at least one embodiment, a cooling distribution unit (CDU) 1312 is used between the main cooling circuit 1306 and a second or auxiliary cooling circuit 1308 to enable heat to be absorbed from the second or auxiliary cooling circuit 1308 to the main cooling circuit 1306. In at least one embodiment, in one aspect, the auxiliary cooling circuit 1308 may be connected to various plumbing systems within the server trays as needed. In at least one embodiment, loops 1306, 1308 are illustrated as line drawings, but those skilled in the art will recognize that one or more piping system features may be used. In at least one embodiment, flexible polyvinyl chloride (PVC) pipe may be used with the associated piping system to allow fluid movement in each configured loop 1306; 1308. In at least one embodiment, one or more coolant pumps may be used to maintain a pressure differential within the coolant loops 1306, 1308 to allow coolant to move in various locations (including in a room, in one or more racks 1310, and / or in server enclosures or server trays within one or more racks 1310) according to temperature sensors.
[0118] In at least one embodiment, the coolant in the main cooling circuit 1306 and the auxiliary cooling circuit 1308 may be at least water and an additive. In at least one embodiment, the additive may be ethylene glycol or propylene glycol. In operation, in at least one embodiment, each cooling circuit in the main cooling circuit and the auxiliary cooling circuit has its own coolant. In at least one embodiment, the coolant in the auxiliary cooling circuit may be dedicated to the requirements of components in the server tray or associated rack 1310. In at least one embodiment, the CDU 1312 is capable of complex control over the coolant within the configured coolant circuits 1306, 1308, either independently or concurrently. In at least one embodiment, the CDU may be adapted to control the flow rate of the coolant to properly distribute the coolant to absorb heat generated within the associated rack 1310. In at least one embodiment, a more flexible piping system 1314 is provided from the auxiliary cooling circuit 1308 to enter each server tray and supply coolant to the electrical and / or computing components therein.
[0119] Heat transfer fluids typically include water, aqueous solutions (e.g., propylene glycol-water), brine, antifreeze, mixtures of antifreeze and water, oil, alcohol, mercury, or any other suitable heat transfer fluid. The heat transfer fluid may be a conductive coolant and may comprise water, deionized water or a coolant (such as R-134a), a mixture of water and additives (such as a mixture of water and ethylene glycol or a mixture of water and propylene glycol (e.g., a 25% aqueous solution of deionized propylene glycol). The heat transfer fluid may also be a single dielectric fluid (e.g., water-free for the purposes of this disclosure) or a combination of water and additives comprising at least one dielectric fluid, such as one or more of deionized water, ethylene glycol, and propylene glycol. In at least one embodiment, the heat transfer fluid may be an absorption cooler whose working fluid is a mixed solution containing lithium bromide as an absorbent material and water as a carrier material. The heat transfer fluid may also be a two-phase coolant with a boiling point below the intended operating temperature of the electronic device. Exemplary two-phase coolants include 2,3,3,3-tetrafluoropropylene, 1,1,1,2-tetrafluoroethane, and water.
[0120] In at least one embodiment, the piping system 1318 forming part of the auxiliary cooling loop 1308 may be referred to as a room manifold. Additionally, in at least one embodiment, other piping systems 1316 may extend from the row manifold piping system 1318 and may also be part of the auxiliary cooling loop 1308, but may be referred to as a row manifold. In at least one embodiment, the coolant piping system 1314 enters the rack as part of the auxiliary cooling loop 1308, but may be referred to as a rack cooling manifold within one or more racks. In at least one embodiment, the row manifold 1316 extends along rows in the data center 1300 to all racks. In at least one embodiment, the piping system of the auxiliary cooling loop 1308, including coolant manifolds 1318, 1316, and 1314, may be improved by at least one embodiment described herein. In at least one embodiment, a cooler 1320 may be positioned in the main cooling loop within the data center 1302 to support cooling prior to the cooling tower. In at least one embodiment, for the purposes of this disclosure, an additional cooling circuit may exist within the main control circuit and provide cooling outside the rack and outside the auxiliary cooling circuit. The additional cooling circuit may be used together with the main cooling circuit and may be different from the auxiliary cooling circuit.
[0121] In at least one embodiment, during operation, heat generated within the server trays of the provided rack 1310 can be transferred via a flexible piping system of row manifolds 1314 of the second cooling circuit 1308 to coolant leaving one or more racks 1310. In at least one embodiment, a second coolant from CDU 1312 (in the auxiliary cooling circuit 1308) for cooling the provided racks 1310 moves toward one or more racks 1310 via the provided piping system. In at least one embodiment, the second coolant from CDU 1312 reaches one side of the rack 1310 from one side of the room manifold having piping system 1318 via row manifold 1316 and passes through one side of the server tray via different piping systems 1314. In at least one embodiment, used or returned second coolant (or departing second coolant carrying heat from computing components) exits from the other side of the server tray (e.g., after circulating through the server tray or components on the server tray, entering the left side of the rack of the server tray and exiting from the right side of the rack). In at least one embodiment, the used second coolant exiting the server tray or rack 1310 exits from a different side (e.g., the exit side) of the piping system 1314 and moves to the parallel exit side of the row manifold 1316. In at least one embodiment, the used second coolant enters from the row manifold 1316 into a parallel portion of the room manifold 1318 and is traveling in the opposite direction to the incoming second coolant (which may also be regenerated second coolant) and toward the CDU 1312.
[0122] In at least one embodiment, the used second coolant exchanges heat with the main coolant in the main cooling circuit 1306 via CDU 1312. In at least one embodiment, the used second coolant can be regenerated (e.g., after relative cooling compared to its temperature during the used second coolant phase) and prepared to be circulated back to one or more computing components via the second cooling circuit 1308. In at least one embodiment, various flow and temperature control features in CDU 1312 enable control of the heat exchanged from the used second coolant or the flow rate of the second coolant into and out of CDU 1312. In at least one embodiment, CDU 1312 is also capable of controlling the flow of the main coolant in the main cooling circuit 1306.
[0123] Figure 14A and 14B Top and perspective views are illustrated respectively of a transceiver module operatively coupled to a network adapter (in this example, a network interface controller (NIC) 1400) according to an embodiment of the present disclosure. Figure 14A and 14BAs shown, the transceiver module may include a first optical module 1401, a second optical module 1403, an adapter 1410, and a dual-port NIC 1420 for the server. Both the first optical module 1401 and the second optical module 1403 may be dual-fiber transceivers configured for full-duplex communication, allowing communication between a source (e.g., a server) and a target (e.g., a leaf switch) in both directions. The adapter 1410 may be a linked physical component configured to link the first optical module 1401 and the second optical module 1403 for sending data to and receiving data from the leaf switch.
[0124] In some embodiments, adapter 1410 can be configured to operate in two configurations, such as a first configuration and a second configuration. In one aspect, the first configuration can be a default operating configuration, wherein the first optical module 1401 can be operationally active. The second configuration can be an emergency configuration implemented when the first optical module 1401 experiences an operational failure. When such a failure is detected, the second optical module 1403, which was originally operationally inactive or idle, can become operationally active and handle all network traffic initially handled by the first optical module 1401.
[0125] In some embodiments, transceiver module 1400 can be configured to operate in a leaf-spine architecture. A leaf-spine architecture is a data center network topology that may include two switching layers (spine and leaf). The leaf layer may include access switches (leaf switches) that aggregate traffic from servers and are directly connected to the backbone or network core. Spine switches interconnect all leaf switches in a full mesh topology, and access switches are located in the leaf layer and aggregate traffic from servers. Thus, in one embodiment, to ensure reliable downlink operation, transceiver module 1400 can be configured to operate between the server and the leaf layer. Specifically, as... Figure 14A and Figure 14B As shown, adapter 1410 can be operatively coupled to first optical module 1401 and second optical module 1403, and first optical module 1401 and second optical module 1403 can be operatively coupled to dual-port NIC 1420 of the server.
[0126] In various embodiments, the NIC 1400 may include one or more processing circuits as detailed above; the processing circuits may include a firmware loaded according to the techniques described above.
[0127] Figure 15Exemplary use cases of the optical transceiver 1502 according to some embodiments are depicted. To name just a few examples, the optical transceiver 1502 can be used in computing systems 1504 (e.g., in server farms or within server computer systems), vehicles 1506 (e.g., cars, trucks, trains, or airplanes), and (or robots in factories) robots 1508. The optical transceiver 1502 is particularly useful for high-speed communication in environments subject to high levels of electromagnetic interference (EMI).
[0128] Other variations are within the spirit of this disclosure. Therefore, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments are shown in the accompanying drawings and have been described in detail above. However, it should be understood that this disclosure is not intended to be limited to the one or more specific forms disclosed, but rather is intended to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of this disclosure as defined in the appended claims.
[0129] Unless otherwise stated herein or obviously contradicted by the context, the use of the terms “a,” “an,” “the,” and similar pronouns in the context of describing the disclosed embodiments (particularly in the context of the appended claims) should be interpreted as encompassing both singular and plural forms, rather than as definitions of the terms. Unless otherwise indicated, the terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (meaning “including, but not limited to”). When unmodified and referring to a physical connection, “connection” should be interpreted as partially or completely contained in, attached to, or combined with, even with intervening elements. Unless otherwise indicated herein, statements of value ranges herein are intended only as a way of referring to each individual value falling within that range separately, and each individual value is incorporated into the specification as if it were individually stated herein. In at least one embodiment, unless otherwise specified or contradicted by the context, the use of the terms “set” (e.g., “item set”) or “subset” should be interpreted as a non-empty set comprising one or more members. Furthermore, unless otherwise specified or contradicted by the context, the term "subset" of the corresponding set does not necessarily refer to an appropriate subset of the corresponding set, but rather the subset and the corresponding set can be equal.
[0130] Unless explicitly stated otherwise or otherwise clearly contradicted by the context, connectives (such as phrases of the form "at least one of A, B, and C" or "at least one of A, B, and C") are otherwise understood, in conjunction with the context, to generally refer to items, terms, etc., and can be any non-empty subset of the set A or B or C, or A and B and C. For example, in an illustrative example of a set with three members, the connective phrases "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such connectives are generally not intended to imply that certain embodiments require the separate presence of at least one of A, at least one of B, and at least one of C. Additionally, unless explicitly stated or contradicted by the context, the term "multiple" indicates a plural state (e.g., "multiple items" indicates multiple items). In at least one embodiment, the number of items in the multiple is at least two, but may be more if explicitly indicated or indicated by the context. Furthermore, unless otherwise stated or clearly understood from the context, the phrase “based on” means “at least partially based on” rather than “based on only”.
[0131] Unless otherwise indicated herein or otherwise clearly contradicted by the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions and are implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that does not include transient signals (e.g., propagation of transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver that includes transient signals. In at least one embodiment, code (e.g., executable code or source code) is stored in a set of one or more non-transitory computer-readable storage media having executable instructions (or other memory for storing executable instructions) stored thereon. These executable instructions, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media comprises multiple non-transitory computer-readable storage media, and one or more individual non-transitory storage media lack the complete code; instead, the multiple non-transitory computer-readable storage media collectively store the complete code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors.
[0132] Therefore, in at least one embodiment, the computer system is configured to implement one or more services that perform the operations of the processes described herein, either individually or collectively, and such a computer system is configured with suitable hardware and / or software to enable the performance of the operations. Furthermore, the computer system implementing at least one embodiment of this disclosure is a single device, while in another embodiment it is a distributed computer system comprising multiple devices operating in different ways, such that the distributed computer system performs the operations described herein, and that a single device does not perform all the operations.
[0133] Unless otherwise required, the use of any and all examples or exemplary language (e.g., “such as”) provided herein is intended only to better illustrate embodiments of this disclosure and is not intended to limit the scope of this disclosure. The language in the specification should not be construed as indicating that any unclaimed element is essential to the practice of this disclosure.
[0134] In the specification and claims, the terms “coupled” and “connected” along with their derivatives may be used. It should be understood that these terms are not intended to be synonyms for each other. Rather, in specific examples, “connected” or “coupled” can be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. “Coupled” can also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0135] Unless otherwise specifically stated, it should be understood that throughout this specification, terms such as “processing,” “calculating,” “operating,” “determining,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that manipulate and / or transform data representing physical (such as electronic) quantities within the registers and / or memory of the computing system into other data representing physical quantities within the memory, registers, or other such information storage, transmission, or display devices of the computing system.
[0136] Similarly, the term "processor" can refer to any device or part of a device that processes electronic data from registers and / or memory and transforms that electronic data into other electronic data that can be stored in registers and / or memory. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Furthermore, each process can refer to multiple processes for executing instructions sequentially or in parallel, continuously or intermittently. In at least one embodiment, the terms "system" and "method" are used interchangeably herein, provided that the system can embody one or more methods, and the method can be considered a system.
[0137] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer implementation machine. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog and digital data can be accomplished in various ways, such as by receiving data as a parameter of a function call or a call to an application programming interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. In at least one embodiment, reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter of a function call, an application programming interface, or an inter-process communication mechanism.
[0138] While the description herein illustrates exemplary embodiments of the technology, other architectures may also be used to implement the functionality and are intended to fall within the scope of this disclosure. Furthermore, although a specific allocation of responsibilities has been defined above for descriptive purposes, various functions and responsibilities may be allocated and divided differently depending on the circumstances.
[0139] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological actions, it is to be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or actions described. Rather, the specific features and actions are disclosed as exemplary forms for implementing the claims.
Claims
1. A system comprising: A sliding member configured to move a cold plate between a first configuration and a second configuration, wherein in the first configuration the cold plate is positioned away from the surface of the interconnect module, and wherein in the second configuration the cold plate is positioned in thermal communication with the interconnect module.
2. The system of claim 1, further comprising: The cold plate, wherein the cold plate is configured to cool the interconnect module; as well as A thermal interface pad on the cooling surface of the cold plate, wherein the thermal interface pad does not contact the interconnect module when the cold plate is supported in the first configuration, and wherein the thermal interface pad contacts the interconnect module when the cold plate is in the second configuration.
3. The system of claim 2, wherein the thermal interface pad comprises graphene.
4. The system of claim 2, wherein the system is configured to receive multiple insertions and removals of the interconnect module without damaging the thermal interface pad and without mechanical failure of the sliding member.
5. The system of claim 1, wherein the thermal interface between the interconnect module and the cold plate has a thermal resistance of less than approximately 0.2 Kelvin per watt (K / W).
6. The system of claim 1, wherein the sliding member is configured to slide in a translational manner between a first translational position corresponding to the first configuration and a second translational position corresponding to the second configuration.
7. The system of claim 6, wherein during the process of inserting the interconnect module into the socket, the interconnect module is used to push the sliding member from the first translational position to the second translational position.
8. The system of claim 6, wherein in the first translational position, the sliding member positions the cold plate at a first vertical position away from the interconnect module, wherein in the second translational position, the sliding member positions the cold plate at a second vertical position, and wherein in the second vertical position, the cold plate is in thermal communication with the interconnect module.
9. The system of claim 8, wherein the sliding member is formed with a ramp feature configured to interact with a protrusion of the cold plate, and wherein as the sliding member slides in a translational manner between the first translational position and the second translational position, the protrusion is used to slide along the ramp feature to vertically move the cold plate between the first vertical position and the second vertical position.
10. The system of claim 9, wherein the sliding member comprises two generally parallel members coupled by at least one bridging member, and each of the two generally parallel members forms at least one ramp feature.
11. The system of claim 6, further comprising: One or more first springs are configured to apply a first spring force to the sliding member, wherein the first spring force is in a first direction corresponding to the movement of the sliding member from the second translational position to the first translational position.
12. The system of claim 11, further comprising: One or more second springs are configured to apply a second spring force to the cold plate, wherein the second spring force is in a second direction toward the interconnecting module coupled within the socket.
13. A system comprising: A sliding member configured to adjust the height of a cold plate between a first height and a second height during the process of inserting an interconnect module into a socket, wherein the adjustment of the height of the cold plate minimizes the shear force on the thermal interface pad of the cold plate during the process of inserting the interconnect module into the socket.
14. The system of claim 13, wherein the thermal interface pad does not contact the interconnect module when the cold plate is at the first height, and wherein the thermal interface pad contacts the interconnect module when the cold plate is at the second height.
15. The system of claim 13, wherein the thermal interface formed at least partially by the thermal interface pad between the interconnect module and the cold plate has a thermal resistance of less than approximately 0.2 Kelvin per watt (K / W).
16. The system of claim 13, wherein when the cold plate is at the first height, the cold plate is positioned vertically away from the interconnect module, and wherein when the cold plate is at the second height, the cold plate is in thermal communication with the interconnect module.
17. The system of claim 13, wherein the sliding member is configured to slide in a translational manner between a first translational position and a second translational position, wherein when the sliding member is in the first translational position, the cold plate is adjusted to the first height; and wherein when the sliding member is in the second translational position, the cold plate is adjusted to the second height.
18. The system of claim 17, wherein during the process of inserting the interconnect module into the socket, the interconnect module is used to push the sliding member from the first translational position to the second translational position.
19. The system of claim 17, wherein the sliding member is formed with a ramp feature configured to interact with a protrusion of the cold plate, and wherein as the sliding member slides in a translational manner between the first translational position and the second translational position, the protrusion is used to slide along the ramp feature to vertically move the cold plate between the first height and the second height.
20. The system of claim 17, further comprising: One or more first springs are configured to apply a first spring force to the sliding member, wherein the first spring force is in a first direction corresponding to the movement of the sliding member from the second translational position to the first translational position; as well as One or more second springs are configured to apply a second spring force to the cold plate, wherein the second spring force is in a second direction toward the interconnecting module coupled within the socket.
21. A computing server, comprising: A socket configured to receive interconnect modules; Cold plate; as well as A sliding member configured to move the cold plate between a first configuration and a second configuration, wherein in the first configuration the cold plate is positioned away from the surface of the interconnect module inserted into the socket, and wherein in the second configuration the cold plate is positioned in thermal communication with the interconnect module.
22. The computing server of claim 21, further comprising: A thermal interface pad on the cooling surface of the cold plate, wherein the thermal interface pad does not contact the interconnect module when the cold plate is supported in the first configuration, and wherein the thermal interface pad contacts the interconnect module when the cold plate is in the second configuration.
23. The computing server of claim 21, wherein the sliding member is configured to slide in a translational manner between a first translational position corresponding to the first configuration and a second translational position corresponding to the second configuration, and wherein during the process of inserting the interconnect module into the socket, the interconnect module is used to push the sliding member from the first translational position to the second translational position.
24. The computing server of claim 23, wherein in the first translational position, the sliding member positions the cold plate at a first vertical position away from the interconnect module, wherein in the second translational position, the sliding member positions the cold plate at a second vertical position, and wherein in the second vertical position, the cold plate is in thermal communication with the interconnect module.
25. The computing server of claim 24, wherein the sliding member is formed with a ramp feature configured to interact with a protrusion of the cold plate, and wherein as the sliding member slides in a translational manner between the first translational position and the second translational position, the protrusion is configured to slide along the ramp feature to vertically move the cold plate between the first vertical position and the second vertical position.
26. The computing server of claim 23, further comprising: One or more first springs are configured to apply a first spring force to the sliding member, wherein the first spring force is in a first direction corresponding to the movement of the sliding member from the second translational position to the first translational position; as well as One or more second springs are configured to apply a second spring force to the cold plate, wherein the second spring force is in a second direction toward the interconnecting module coupled within the socket.
27. A data center, comprising: One or more computing servers, wherein at least one of the one or more computing servers comprises: A socket configured to receive interconnect modules; Cold plate; and A sliding member configured to move the cold plate between a first configuration and a second configuration, wherein in the first configuration the cold plate is positioned away from the surface of the interconnect module inserted into the socket, and wherein in the second configuration the cold plate is positioned in thermal communication with the interconnect module.
28. The data center of claim 27, wherein the at least one computing server among the one or more computing servers further comprises: A thermal interface pad on the cooling surface of the cold plate, wherein the thermal interface pad does not contact the interconnect module when the cold plate is supported in the first configuration, and wherein the thermal interface pad contacts the interconnect module when the cold plate is in the second configuration.
29. The data center of claim 27, wherein the sliding member is configured to slide in a translational manner between a first translational position corresponding to the first configuration and a second translational position corresponding to the second configuration, and wherein during the process of inserting the interconnect module into the socket, the interconnect module is used to push the sliding member from the first translational position to the second translational position.