Liquid cooling plant room system

By designing a liquid-cooled server room system in the data center, and adopting cold plate liquid cooling technology and high-performance computing systems, the problems of low space utilization and high PUE value in traditional data centers have been solved, achieving more efficient energy utilization and lower PUE value, and supporting high-performance artificial intelligence computing.

CN118450668BActive Publication Date: 2025-11-21CHINA CONSTRUCTION BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410542365.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-11-21
Estimated Expiration
2044-04-30

AI Technical Summary

Technical Problem

Traditional data center server rooms and racks have low space utilization, a small number of modular servers deployed per server room, and a high PUE value, which limits the popularization of large-scale artificial intelligence model clusters.

Method used

Design a liquid-cooled data center system, including a training cluster, a general inference cluster, and a scalable inference cluster. Employ cold plate liquid cooling technology, deploy multiple GPU servers in liquid-cooled cabinets, and include network switches in air-cooled cabinets, to achieve flexible deployment and high availability design of the high-performance computing system.

Benefits of technology

It improves the utilization rate of data center and rack space and energy, achieves a lower PUE value, creates a green and low-carbon data center, and supports high-performance artificial intelligence computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118450668B_ABST
    Figure CN118450668B_ABST
Patent Text Reader

Abstract

The application provides a liquid-cooled machine room system, and relates to the technical field of liquid-cooled machine rooms.The liquid-cooled machine room system comprises a training cluster, a general reasoning cluster and an extensible reasoning cluster; the extensible reasoning cluster and the training cluster are arranged on the same floor of the same machine room building; the general reasoning cluster, the extensible reasoning cluster and the training cluster are arranged in different machine room buildings; the training cluster and the extensible reasoning cluster each comprise a first number of liquid-cooled cabinets and a second number of air-cooled cabinets; the general reasoning cluster comprises a third number of liquid-cooled cabinets and a fourth number of air-cooled cabinets; the liquid-cooled cabinet comprises a plurality of GPU servers; and the air-cooled cabinet comprises a single network switch, which can realize high availability and flexible deployment design of the liquid-cooled machine room system, improve the utilization rate of machine room and cabinet space and energy, and realize a lower PUE value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of liquid-cooled computer room technology, and more particularly to a liquid-cooled computer room system. Background Technology

[0002] This section is intended to provide background or context for embodiments of the present invention. The description herein is not intended to imply that it is prior art simply because it is included in this section.

[0003] ChatGPT (Chat Generative Pre-trained Transformer) has become a highly anticipated technological innovation, marking a significant breakthrough in the field of artificial intelligence. The construction of large-scale AI model clusters is also in full swing. To effectively utilize AI, powerful computing resources are essential; in addition to servers, data center infrastructure (power, water, and electricity) also plays a crucial role.

[0004] Traditional servers consume approximately 150-300W, GPU (Graphics Processing Unit) servers consume approximately 9-10kW, and liquid-cooled GPU servers consume approximately 8-9kW. As can be seen, the power consumption of servers used for artificial intelligence is significantly higher than that of traditional servers.

[0005] Regarding the deployment of AI GPU servers in data center server rooms, the current common solutions are as follows: in traditional air-cooled server rooms, most high-performance GPU servers are deployed in a rack of one, with a full load power of 10kW per rack; a few are deployed in a rack of 2-3, with a full load power of 20-25kW per rack, at the cost of requiring a large number of in-row air conditioners and reducing the server rack utilization rate; or in general liquid-cooled server rooms (single rack power consumption of 20kW and below), two GPU servers are deployed in a rack. Currently, there are very few high-performance computing liquid-cooled server rooms with single rack power consumption of 40kW or more that are in use.

[0006] Existing solutions suffer from problems such as low utilization of data center and rack space, limited number of modular servers deployed per data center, and high PUE (Power Usage Effectiveness) values ​​when deploying high-power GPU servers. These issues limit the widespread use of large-scale artificial intelligence model clusters. Summary of the Invention

[0007] To address the problems existing in the prior art, this invention proposes a liquid-cooled data center system. This invention enables high availability and flexible deployment design of the liquid-cooled data center system, improves data center and rack space and energy utilization, and achieves a lower PUE value.

[0008] This invention provides a liquid-cooled data center system, comprising: a training cluster, a general inference cluster, and a scalable inference cluster; the training cluster is a high-performance computing system for implementing the artificial intelligence training process; the general inference cluster is a high-performance computing system for implementing the artificial intelligence inference process; and the scalable inference cluster is a high-performance computing system that meets the requirements of a general inference cluster and has the conditions to be modified into a training cluster.

[0009] The scalable inference cluster and the training cluster are deployed on the same floor of the same data center building; the general inference cluster, the scalable inference cluster, and the training cluster are deployed in different data center buildings.

[0010] The training cluster and the scalable inference cluster each include a first number of liquid-cooled cabinets and a second number of air-cooled cabinets; the general inference cluster includes a third number of liquid-cooled cabinets and a fourth number of air-cooled cabinets; the liquid-cooled cabinets include multiple liquid-cooled graphics processing unit (GPU) servers; the air-cooled cabinets include a single network switch; the liquid-cooled GPU servers adopt cold plate liquid cooling technology.

[0011] Compared with the traditional air-cooled data center and deployment solutions in the prior art, the embodiments of this invention deploy training clusters, general inference clusters, and scalable inference clusters. The training cluster is a high-performance computing system for implementing the artificial intelligence training process; the general inference cluster is a high-performance computing system for implementing the artificial intelligence inference process; the scalable inference cluster is a high-performance computing system that meets the requirements of the general inference cluster and has the conditions to be transformed into a training cluster; the scalable inference cluster and the training cluster are deployed on the same floor of the same data center building; the general inference cluster, the scalable inference cluster, and the training cluster are deployed in different data center buildings, which can realize a high-availability design of dual availability zones for inference clusters deployed in different data center buildings, and a flexible deployment design where the scalable inference cluster can be converted into a training cluster; both the training cluster and the scalable inference cluster include a first number of liquid-cooled cabinets and a second number of air-cooled cabinets; the general inference cluster includes a third number of liquid-cooled cabinets and a fourth number of air-cooled cabinets; the liquid-cooled cabinet includes multiple liquid-cooled graphics processing unit (GPU) servers; the air-cooled cabinet includes a single network switch; the liquid-cooled GPU servers adopt cold plate liquid cooling technology, and deploying multiple GPU servers in a single liquid-cooled cabinet can improve the space and energy utilization of the data center and cabinets, and achieve a lower PUE value. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1This is a schematic diagram of a liquid-cooled computer room system according to an embodiment of the present invention;

[0014] Figure 2 This is a schematic diagram of the cold plate liquid cooling technology according to an embodiment of the present invention;

[0015] Figure 3 This is a floor plan of the floor where the training cluster and scalable inference cluster are located, according to an embodiment of the present invention.

[0016] Figure 4 This is a floor plan of the floor where the general inference cluster is located, according to an embodiment of the present invention.

[0017] Figure 5 This is a schematic diagram of the structure of the liquid-cooled machine room according to an embodiment of the present invention. Detailed Implementation

[0018] The principles and spirit of the invention will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement the invention, and are not intended to limit the scope of the invention in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0019] Those skilled in the art will recognize that embodiments of the present invention can be implemented as a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0020] First, let's introduce the technical terms used in this article:

[0021] Liquid cooling: A cooling method that uses liquid as a cooling medium to exchange heat with the heat-generating components of IT (Information Technology) equipment, removing the heat generated by these components. It is suitable for applications requiring increased computing power, energy efficiency, and deployment density.

[0022] Cold plate liquid cooling technology: This is a method of transferring heat from heat-generating components to cooling liquid enclosed in a circulation pipeline through a cold plate (usually a closed cavity made of thermally conductive metals such as copper or aluminum), and then carrying away the heat through the cooling liquid.

[0023] Liquid Cooling Distribution Unit (CDU): Also known as a "cooling distribution unit," it is a module used for heat exchange between the secondary side high-temperature liquid cooling medium and the primary side cold source, and for providing cooling capacity distribution and intelligent management for liquid-cooled IT equipment.

[0024] Primary side: This refers to the primary cooling loop, a cooling system within the liquid-cooled computer room system responsible for dissipating heat generated by components in the computer room from the secondary cooling loop to the outdoor atmosphere or recovering it through a heat recovery system, while simultaneously circulating the cooling medium. The primary cooling loop can consist of a liquid cooling distribution unit (primary circulation channel), cooling water circulation pipes, a water pump, and a cold source.

[0025] Secondary side: This refers to the secondary cooling loop, which is the cooling medium circulation system within the liquid-cooled computer room system responsible for removing the heat generated by the high heat flux density components of electronic information equipment from the computer room and delivering it to the cooling distribution unit for heat exchange with the external circulation system. It mainly consists of the liquid-cooled distribution unit (secondary side circulation channel), liquid-cooled equipment, cooling medium supply and return manifolds, circulation pipelines, and connecting pipelines.

[0026] PUE stands for Power Usage Effectiveness, a metric for evaluating the energy efficiency of data centers. It is the ratio of all energy consumed by the data center to the energy consumed by the IT load. PUE = Total Data Center Energy Consumption / IT Equipment Energy Consumption. The total data center energy consumption includes the energy consumption of IT equipment and systems such as cooling and power distribution. A value greater than 1 indicates that the non-IT equipment consumes less energy, and the better the energy efficiency.

[0027] The realization of artificial intelligence involves two stages: training and inference.

[0028] Training, also known as the learning process, refers to training a complex neural network model using large amounts of labeled data. This allows the system to adapt to specific functions. Training requires high computing power, the ability to handle massive amounts of data, and a degree of versatility to accomplish a wide variety of learning tasks.

[0029] Reasoning refers to using a trained model to deduce various conclusions from new data. It involves using a neural network model to perform calculations and obtain the correct conclusion in one go using new input data. This is also called prediction or inference.

[0030] Cluster: A cluster is a high-performance computing system that connects multiple computers on a local area network (LAN) or the internet, using software and hardware to share resources and allocate tasks to collaboratively complete a single task. Clusters are primarily used in large-scale data processing, scientific computing, and digital media processing, offering higher computing efficiency and larger data storage capacity.

[0031] Training cluster: A high-performance computing system capable of implementing the artificial intelligence training process. It trains the corresponding system using a large amount of labeled data, enabling it to adapt to specific functions and possessing a certain degree of versatility.

[0032] Inference clusters are high-performance computing systems that enable artificial intelligence inference processes. Inference clusters are actually designed for applications, using trained models to solve specific application problems.

[0033] Scalable inference cluster: The cluster itself meets the requirements of an inference cluster, while reserving some potential conditions, and has the ability to be transformed into a training cluster.

[0034] This invention provides a liquid-cooled data center system that addresses the problems of low utilization of data center and rack space, limited number of single-data center module servers, and high PUE values ​​in existing systems. It achieves high availability and flexible deployment design for the liquid-cooled data center system, improves the utilization of data center and rack space and energy, and achieves a lower PUE value.

[0035] The principles and spirit of the present invention will be explained in detail below with reference to several representative embodiments.

[0036] Figure 1 This is a schematic diagram of a liquid-cooled computer room system according to an embodiment of the present invention, as shown below. Figure 1 As shown, the system may include: a training cluster, a general inference cluster, and a scalable inference cluster; the training cluster is a high-performance computing system for implementing the artificial intelligence training process; the general inference cluster is a high-performance computing system for implementing the artificial intelligence inference process; and the scalable inference cluster is a high-performance computing system that meets the requirements of a general inference cluster and has the conditions to be modified into a training cluster.

[0037] The scalable inference cluster and the training cluster are deployed on the same floor of the same data center building; the general inference cluster, the scalable inference cluster, and the training cluster are deployed in different data center buildings.

[0038] The training cluster and the scalable inference cluster each include a first number of liquid-cooled cabinets and a second number of air-cooled cabinets; the general inference cluster includes a third number of liquid-cooled cabinets and a fourth number of air-cooled cabinets; the liquid-cooled cabinets include multiple liquid-cooled graphics processing unit (GPU) servers; the air-cooled cabinets include a single network switch; the liquid-cooled GPU servers adopt cold plate liquid cooling technology.

[0039] The embodiments of the present invention mainly solve the problems of low room and rack space utilization, small number of modular servers deployed per room, and high PUE value when deploying high-power liquid-cooled GPU servers in traditional air-cooled server rooms of data centers. It enables the deployment of more GPU servers per rack and server room, which can maximize the utilization of energy and room and rack space, achieve a lower PUE value, and thus create a green and low-carbon data center.

[0040] In one embodiment, the architectural design and floor plan layout of a liquid-cooled data center system can be derived by analyzing the deployment requirements of a bank's large-scale artificial intelligence model cluster. For example, the bank's requirements for liquid-cooled GPU servers and network switches are: a plan to deploy 1972 liquid-cooled GPU servers in the new data center, with a single unit power consumption of approximately 4.5–5 kW. Based on these requirements, the liquid-cooled racks in the liquid-cooled data center system of this invention can include 8 liquid-cooled GPU servers, meaning each liquid-cooled GPU server deploys 8 nodes in each rack, with a single rack power consumption of approximately 40 kW. The liquid-cooled GPU servers occupy approximately 250 liquid-cooled racks. Due to the high height and high power consumption of the accompanying network switches, they can be deployed as one network switch per air-cooled rack, with the network switches occupying approximately 190 air-cooled racks.

[0041] The 250 liquid-cooled racks and 190 air-cooled racks are divided into three data center modules according to their functions and requirements. The corresponding functions are to realize the artificial intelligence training process, to realize the artificial intelligence inference process, and to realize both the artificial intelligence inference process and, after modification, to realize the artificial intelligence training process. That is, the three data center modules correspond to training clusters, general inference clusters, and scalable inference clusters, respectively. The training cluster is a high-performance computing system for realizing the artificial intelligence training process; the general inference cluster is a high-performance computing system for realizing the artificial intelligence inference process; and the scalable inference cluster is a high-performance computing system that meets the requirements of the inference cluster and has the conditions to be modified into a training cluster.

[0042] The training cluster is planned to deploy 625 liquid-cooled GPU servers after the new data center goes into operation, requiring 80 liquid-cooled racks; the supporting network switches will require 80 units, requiring 80 air-cooled racks.

[0043] The inference cluster and other related needs will require approximately 1347 liquid-cooled GPU servers and 170 liquid-cooled racks over the next 5 years after the new data center goes into operation. However, considering that the current demand for the inference cluster is less than it will be in the next 5 years, it is planned to build two data center modules, namely:

[0044] (1) Scalable Inference Cluster: This refers to an inference cluster that can be expanded into a training cluster. Considering the future need to build a training cluster with 10,000 training chips, this inference cluster must be able to be transformed into a training cluster and form a unified training cluster with the original training cluster. Therefore, it must meet the networking requirements of network switches, i.e., it needs to deploy 80 network switches, 80 air-cooled cabinets, and 80 liquid-cooled cabinets. This cluster should be deployed on the same floor of the same building as the training cluster. In summary, the design scheme of this data center module and the data center module of the training cluster can be completely consistent.

[0045] (2) General Inference Cluster: This refers to a typical inference cluster. A general inference cluster only requires a small-scale RDMA (Remote Direct Memory Access) network. The number of network switches required is significantly reduced compared to the training cluster. Therefore, reserving 30 air-cooled racks is sufficient, and 90 liquid-cooled racks are also needed. However, the inference cluster has a high availability requirement of dual AZ (Availability Zones) deployment. Therefore, this cluster needs to be deployed in two separate data center buildings from the two clusters mentioned above.

[0046] In summary, based on the requirements, the liquid-cooled data center system needs to be set up with a total of 3 data center modules. Two data center modules each have 160 IT cabinets (80 liquid-cooled cabinets + 80 air-cooled cabinets), and one data center module has 120 IT cabinets (90 liquid-cooled cabinets + 30 air-cooled cabinets), for a total of 440 IT cabinets (250 liquid-cooled cabinets + 190 air-cooled cabinets). The total IT power consumption of a single data center module is about 4800kW, and the total IT power consumption of the three data center modules is 14400kW.

[0047] In one embodiment, the fourth quantity is less than the second quantity. Generally, inference clusters only need to build a small-scale RDMA (Remote Direct Memory Access) network, and the number of network switches required is greatly reduced compared to training clusters. Therefore, reserving 30 air-cooled racks is sufficient to meet the requirements. However, scalable inference clusters and training clusters require the deployment of 80 network switches and 80 air-cooled racks.

[0048] In one embodiment, both the training cluster and the scalable inference cluster include a primary side and a secondary side; the primary side is deployed outside the liquid-cooled server room and includes a cooling tower and primary side piping; the secondary side is deployed inside the liquid-cooled server room and includes a liquid-cooled distribution unit (CDU), a liquid-cooled cabinet, and secondary side piping; the liquid-cooled cabinet is connected to the CDU through the secondary side piping; the cooling tower is connected to the CDU through the primary side piping; the primary side and the secondary side exchange heat through the CDU.

[0049] Based on the contact method between the liquid refrigerant and the heat source, liquid cooling technology can be divided into three types: cold plate type (indirect contact), immersion type (direct contact), and spray type (direct contact). Through surveys of server manufacturers, a comprehensive comparison of air cooling and different liquid cooling technologies was conducted, concluding that cold plate type liquid cooling has higher commercial maturity, is compatible with existing networks, is superior overall, meets future evolution requirements, and aligns with the development trend of green and low-carbon data centers.

[0050] Figure 2 This is a schematic diagram of the cold plate liquid cooling technology according to an embodiment of the present invention, as shown below. Figure 2As shown, the liquid cooling system of the cold plate type liquid cooling technology consists of a primary side and a secondary side. The primary side can be located outside the liquid cooling room and mainly includes the cooling tower, heat exchange plate, and primary side piping. The secondary side is located inside the liquid cooling room and mainly includes the CDU, secondary side piping, and liquid cooling cabinet. The outdoor cold source enters the CDU through the primary side piping and then enters the liquid cooling cabinet through the secondary side piping.

[0051] According to research conducted with server and rack manufacturers, some liquid-cooled racks can currently achieve a single-rack power consumption of 40-50kW and are equipped with liquid cooling doors, enabling all heat dissipation of the cold plate liquid-cooled GPU server to be carried away by liquid cooling.

[0052] Figure 3 This is a floor plan of the floor where the training cluster and scalable inference cluster are located, according to an embodiment of the present invention. Figure 4 This is a floor plan of the floor where the general inference cluster is located, as shown in this embodiment of the invention. Figure 3 , Figure 4 As shown, the design of the liquid-cooled data center system is mainly based on data center design specifications. Each data center building adopts dual power supply, dual water supply, and dual network routing. Heating and drainage pipelines, as well as resource networks, are independent of each other. Each data center building has independent infrastructure such as ventilation, heating, water, electricity, and network to achieve high availability design between data center buildings. Data center buildings should be equipped with dedicated incoming line rooms for the production network, weak current shafts, and reserved equipment holes. Outdoor pipe shafts should be reserved between data center buildings and connected to the operator's pipe shafts to ensure that the production network of a Class A data center meets the 3-way redundancy requirements. The data center buildings where training clusters and scalable inference clusters are located should also be equipped with spare parts rooms, cylinder rooms, evacuation staircases, pipe shafts, hoisting ports, air conditioning water pipe shafts, freight elevator halls, cooling water pipe shafts, weak current rooms, connecting corridors, etc. Generally, the data center buildings where inference clusters are located should also be equipped with degaussing rooms and fresh air rooms in addition to the configurations of the data center buildings where training clusters and scalable inference clusters are located. To power the training cluster, scalable inference cluster, and general inference cluster, a power distribution room and a battery room should also be deployed on the same floor for the liquid-cooled server room and the air-cooled server room.

[0053] like Figure 3 As shown, in one embodiment, the liquid-cooled cabinets and air-cooled cabinets in the training cluster can be deployed in two separate liquid-cooled server rooms; similarly, the liquid-cooled cabinets and air-cooled cabinets in the scalable inference cluster can be deployed in two separate liquid-cooled server rooms. The CPUs (Computer-Aided Units) can be deployed in the air-conditioned area of ​​the liquid-cooled server room, while the liquid-cooled cabinets can be deployed in the space within the liquid-cooled server room excluding the air-conditioned area. The CPUs and liquid-cooled cabinets can be located in the same room or deployed separately in another room. Considering the high water pressure on the primary side, it is recommended that the CPUs and liquid-cooled cabinets be deployed separately to reduce the risk of leakage.

[0054] The main design requirements for data centers with liquid cooling are a smaller footprint, higher load requirements for the main server room, and the need to reserve space for liquid cooling piping. Because liquid cooling systems replace high-energy-consuming refrigeration equipment such as chillers, liquid-cooled server room systems can achieve lower PUE values, are more energy-efficient and carbon-lower. With the same external power capacity, liquid cooling allows for the configuration of more IT equipment, maximizing energy and space utilization.

[0055] Liquid cooling technology uses liquid as a heat exchange medium to exchange heat and cool near the heat source. Because liquid has a relatively high specific heat capacity, its heat dissipation capacity is much higher than that of air cooling. Therefore, the power density of a single rack in a liquid-cooled data center system is often several times or even dozens of times that of a traditional air-cooled data center. This makes the area occupied by a liquid-cooled data center system much smaller than that of a traditional air-cooled data center.

[0056] Furthermore, in liquid-cooled data center system projects, liquid cooling load accounts for a large proportion of the total cooling load. The liquid cooling load mainly dissipates heat directly to the outside through heat dissipation equipment such as cooling towers. Therefore, liquid-cooled data center systems have a smaller demand for mechanical refrigeration equipment such as chillers, and thus require less floor space for mechanical refrigeration systems such as refrigeration rooms and cold storage equipment. As a result, the building area of ​​a liquid-cooled data center system of the same IT scale is often smaller than that of a data center built using the traditional air-cooling mode.

[0057] In one embodiment, a raised floor is installed in the liquid cooling room, and secondary side pipelines are laid within the raised floor; the laying height of the raised floor is not less than a preset height; and the width of the channel under the raised floor for laying the secondary side pipelines is not less than a preset width.

[0058] There are several infrastructure requirements for the construction of liquid-cooled server rooms, such as floor load-bearing requirements, raised floor requirements, power supply requirements, and HVAC requirements. The equipment layout in the main server room is basically the same as that of a traditional air-cooled server room, especially for those using a cold-plate liquid cooling system. However, because the secondary side piping for the cold-plate liquid cooling system needs to be laid within the raised floor, it is generally recommended that the raised floor be at least 600mm high. When laying the secondary side piping for the cold-plate liquid cooling system under the raised floor, the width of the passageway should generally be at least 1200mm to provide sufficient space for the installation of the secondary side piping and future maintenance.

[0059] In one embodiment, a water collection tray is provided on the secondary side piping, and a leakage detection device is installed inside the water collection tray. It is advisable to add a water collection tray to the secondary side piping of the liquid cooling system, and install a leakage detection device inside the water collection tray to detect leaks.

[0060] In one embodiment, the load of the liquid-cooled server room exceeds the preset load. Higher power density per rack means greater equipment weight; as power density increases, the weight of liquid-cooled racks also increases significantly. Therefore, the main server room design load for liquid-cooled server room systems typically needs to reach 15 kN / m². 2 .

[0061] Figure 5 This is a schematic diagram of the structure of the liquid-cooled machine room according to an embodiment of the present invention, as shown below. Figure 5 As shown, the liquid-cooled server room uses an independent cold source. The closed-loop cooling tower is placed outside the liquid-cooled server room, while the circulating water pumps and hydraulic module equipment are arranged inside. The CDUs are centrally located in the air-conditioning room. The liquid-cooled server room has a raised floor with a height of 800mm. In this invention, the liquid-cooled GPU servers in the liquid-cooled server room utilize native liquid-cooled racks, so 100% of the heat generated by the racks can be removed by the liquid-cooling system, requiring only a small number of room-level air conditioners for comfort. The air-cooled racks housing the chassis network switches in the liquid-cooled server room are cooled using water-cooled in-row air conditioners, with four units per row, one of which is a backup. Each air conditioner has a cooling capacity of no less than 45kW.

[0062] The heat exchange in a liquid-cooled server room is mainly divided into two parts: primary-side heat exchange and secondary-side heat exchange. The secondary side acquires heat from heat sources such as liquid-cooled GPU servers through direct or indirect heat exchange and transfers it to the primary side. The primary side then transfers the heat to the outside through outdoor cooling equipment to complete the entire heat dissipation process. The primary and secondary sides exchange heat through CDUs (cooled CPU coolers).

[0063] Cooling units (CDUs) often have high requirements for water quality, therefore a closed-loop cooling water circulation system is typically used on the primary side. Closed-loop cooling towers are commonly used heat dissipation devices in liquid cooling technology. If an open cooling tower is used, an intermediate plate heat exchanger needs to be added between the cooling tower and the CDU to ensure the quality of the inlet and outlet water for the CDU. Liquid-cooled computer room systems have a wider range of applicable primary side water temperatures. CDUs generally support primary side inlet temperatures above 33°C and heat exchange temperature differences of 8-10°C, thus allowing for natural cooling and heat dissipation throughout the year over a wider range.

[0064] In one embodiment, a filter with a higher mesh size than a preset size is installed at the connection interface between the CDU and the primary side pipeline; the water treatment equipment in the liquid cooling room system uses a device with a filtration accuracy higher than the preset filtration accuracy.

[0065] In selecting primary-side water treatment equipment, since CDU has more stringent requirements for primary-side water quality than traditional air-cooled data centers, in addition to setting up necessary water treatment devices such as dosing devices and water softening devices, it is also necessary to install higher mesh filters at the CDU inlet, and the system's full-process or bypass water treatment equipment also needs to use equipment with higher filtration precision.

[0066] In terms of system design, if an A-level data center is required, in addition to selecting equipment with N+X redundancy and using a ring network, the uninterrupted cooling of the liquid-cooled computer room system is achieved by equipping the cooling tower with a water supply tank that meets the required water supply.

[0067] In one embodiment, the liquid-cooled data center system further includes a power supply system for providing high-voltage DC power to the training cluster, general inference cluster, and scalable inference cluster; the power supply system includes a distribution transformer, a low-voltage distribution cabinet, an uninterrupted power supply (UPS), and a high-voltage DC power supply; the power supply system also includes a backup power supply for working simultaneously with the distribution transformer, low-voltage distribution cabinet, UPS, and high-voltage DC power supply to form a dual power supply, or as a backup power supply for the distribution transformer, low-voltage distribution cabinet, UPS, and high-voltage DC power supply.

[0068] Because liquid-cooled data center systems have a much higher rack density than traditional air-cooled data centers, the cross-section of the terminal power distribution cables is larger, and the requirements for pipeline space are greater. Traditionally, data centers use AC power systems to power ICT (information and communications technology) servers, with UPS providing uninterrupted power. This power supply method involves 10kV power passing through a 10kV / 400V distribution transformer, a low-voltage distribution cabinet, and a UPS, ultimately providing 380V AC power to the servers. However, servers are DC-powered devices, so a switching power supply is still needed at the terminal to rectify and convert the AC power to DC / DC before it can power the various components on the server motherboard. From the perspective of system energy transfer efficiency, the entire system, from the grid power supply to the final server, undergoes four energy conversion stages: AC / DC, DC / AC, AC / DC, and DC / DC. Each stage has corresponding energy losses, so overall, the energy transfer efficiency of the entire AC power supply system is relatively low.

[0069] In recent years, DC-related technologies have developed rapidly in fields such as power (DC transmission) and communications (240V, 336V high-voltage DC power supplies). With the development of power electronics technology and the increasing proportion of DC devices in end-user electrical equipment, the technological advantages of DC power supply and distribution are gradually becoming apparent. In this invention, the liquid-cooled server room system adopts high-voltage DC power supply. Specifically, a 10kV power supply passes through a 10kV / 400V distribution transformer, a low-voltage distribution cabinet, a UPS, and a high-voltage DC power supply, ultimately providing 380V DC power to the liquid-cooled GPU server, which can significantly improve the energy transfer efficiency of the power supply system.

[0070] In this embodiment of the invention, the liquid-cooled server room system needs to meet the requirements of a Class A data center. It requires fault-tolerant configuration of the power distribution system, uninterruptible power supply (UPS), etc., with backup equipment located in separate physical compartments. The battery backup time of the UPS system is no less than 15 minutes. IT equipment is powered by a 2N UPS configuration. The harmonic current at the UPS input must meet the requirements of the financial industry.

[0071] The CDU, liquid-cooled cooling water pumps and cooling towers, and computer room air conditioning equipment must be equipped with an uninterruptible power supply (UPS) system according to the computer room's classification, with backup time sufficient to meet the power distribution requirements of the air conditioning system. The air conditioning system power distribution should use dual AC 380V / 220V power supplies with terminal switching.

[0072] Furthermore, in this embodiment of the invention, the liquid-cooled machine room system should be powered by dual power sources and should be equipped with a 10KV or 0.4KV backup power supply. The backup power supply can be a diesel generator system. The load-carrying characteristics of the backup diesel generator system should not be lower than a preset level, the insulation class of the low-voltage generator should not be lower than H class, and the insulation class of the medium-voltage generator should not be lower than F class. The backup diesel generator system uses continuous power and 70% of its basic power. The backup diesel generator system is configured with N+X redundancy, and the fuel storage capacity should preferably be not less than 12 hours. When the external fuel supply time is guaranteed, the fuel storage capacity only needs to be greater than the external fuel supply time.

[0073] In this embodiment of the invention, the location of the substation for the liquid-cooled machine room system should be close to the load center to facilitate power supply lines and equipment transportation and installation. A reserved bay (switch) is provided on the 110kV output side to meet the building's increased power supply capacity requirements. The terminal power distribution unit (PDU) adopts a 3-phase 32A design, with the A and B PDUs using a black and yellow color-coded design.

[0074] Currently, a complete set of standards for liquid-cooled data center systems has not yet been established, and there are significant differences in product forms among different equipment manufacturers. Furthermore, liquid cooling technology is developing rapidly, while data center construction cycles are relatively long. Therefore, it is necessary to study the division of labor in implementing high-performance computing liquid-cooled data centers. This invention proposes the following division of labor to decouple engineering construction from equipment installation: The project team is responsible for implementing the infrastructure of the liquid-cooled data center system, including the cooling towers and the openings, shafts, piping facilities, power distribution busbars, raised floors, and ceilings. The server manufacturers are responsible for implementing the CDUs, liquid-cooled cabinets and GPU servers, liquid-cooled piping, and power cables from the busbars to the liquid-cooled and air-cooled cabinets within the liquid-cooled data center.

[0075] This invention provides a high-performance computing liquid-cooled data center system. By analyzing the deployment requirements of a bank's large-scale artificial intelligence model cluster, it outlines the division of labor for the system's architectural design and floor plan, power supply and distribution design, air conditioning design, and implementation. This invention enables the scientific deployment of the large-scale artificial intelligence model training and inference clusters, improves the space utilization of the data center and server racks, and reduces the power usage effectiveness (PUE).

[0076] Compared with the traditional air-cooled data center and deployment solutions in the prior art, the embodiments of this invention deploy training clusters, general inference clusters, and scalable inference clusters. The training cluster is a high-performance computing system for implementing the artificial intelligence training process; the general inference cluster is a high-performance computing system for implementing the artificial intelligence inference process; the scalable inference cluster is a high-performance computing system that meets the requirements of the general inference cluster and has the conditions to be transformed into a training cluster; the scalable inference cluster and the training cluster are deployed on the same floor of the same data center building; the general inference cluster, the scalable inference cluster, and the training cluster are deployed in different data center buildings, which can realize a high-availability design of dual availability zones for inference clusters deployed in different data center buildings, and a flexible deployment design where the scalable inference cluster can be converted into a training cluster; both the training cluster and the scalable inference cluster include a first number of liquid-cooled cabinets and a second number of air-cooled cabinets; the general inference cluster includes a third number of liquid-cooled cabinets and a fourth number of air-cooled cabinets; the liquid-cooled cabinet includes multiple liquid-cooled graphics processing unit (GPU) servers; the air-cooled cabinet includes a single network switch; the liquid-cooled GPU servers adopt cold plate liquid cooling technology, and deploying multiple GPU servers in a single liquid-cooled cabinet can improve the space and energy utilization of the data center and cabinets, and achieve a lower PUE value.

[0077] The beneficial effects of the embodiments of the present invention are as follows:

[0078] (1) This invention provides a planar layout and electromechanical design scheme for a high-performance computing liquid-cooled computer room system, which can improve the utilization rate of computer room space for deploying large artificial intelligence model cluster equipment.

[0079] (2) The PUE value of the liquid-cooled computer room system of the present invention can be reduced to about 1.1, which significantly improves the energy utilization rate.

[0080] (3) The high availability design of the inference cluster with dual availability zones deployed in different computer room buildings, and the flexible deployment design of the scalable inference cluster that can be converted into a training cluster.

[0081] (4) The division of labor interface for the implementation of high-performance computing liquid-cooled computer room is defined, realizing the decoupling of engineering construction and equipment installation, which facilitates scientific implementation.

[0082] The terms "module" or "unit" used above can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the above embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0083] It should be noted that although several modules of the liquid-cooled server room system have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more modules described above can be embodied in a single module. Conversely, the features and functions of a single module described above can be further divided and embodied by multiple modules.

[0084] The acquisition, storage, use, and processing of data in this application comply with relevant laws and regulations.

[0085] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0086] This invention is described with reference to flowchart illustrations and / or block diagrams of methods and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0087] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0088] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0089] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A liquid-cooled computer room system, characterized in that, include: Training clusters, general inference clusters, and scalable inference clusters; The training cluster is a high-performance computing system that enables the artificial intelligence training process; The general inference cluster is a high-performance computing system that implements the artificial intelligence inference process; the scalable inference cluster is a high-performance computing system that meets the requirements of the general inference cluster and has the conditions to be transformed into a training cluster. The scalable inference cluster and the training cluster are deployed on the same floor of the same data center building; the general inference cluster, the scalable inference cluster, and the training cluster are deployed in different data center buildings. The training cluster and the scalable inference cluster each include a first number of liquid-cooled cabinets and a second number of air-cooled cabinets; the general inference cluster includes a third number of liquid-cooled cabinets and a fourth number of air-cooled cabinets; the liquid-cooled cabinets include multiple liquid-cooled graphics processing unit (GPU) servers; the air-cooled cabinets include a single network switch; the liquid-cooled GPU servers adopt cold plate liquid cooling technology.

2. The liquid-cooled computer room system according to claim 1, characterized in that, The liquid-cooled and air-cooled cabinets in the training cluster are deployed in two liquid-cooled server rooms, respectively; the liquid-cooled and air-cooled cabinets in the scalable inference cluster are deployed in two liquid-cooled server rooms, respectively.

3. The liquid-cooled computer room system according to claim 2, characterized in that, The load on the liquid cooling room is greater than the preset load.

4. The liquid-cooled computer room system according to claim 2, characterized in that, Both the training cluster and the scalable inference cluster include a primary side and a secondary side; the primary side is deployed outside the liquid-cooled server room and includes a cooling tower and primary side piping; the secondary side is deployed inside the liquid-cooled server room and includes a liquid-cooled distribution unit (CDU), a liquid-cooled cabinet, and secondary side piping. The liquid-cooled cabinet is connected to the CDU via secondary side piping; the cooling tower is connected to the CDU via primary side piping; the primary and secondary sides exchange heat through the CDU.

5. The liquid-cooled computer room system according to claim 4, characterized in that, The CDU is deployed in the air-conditioned area of ​​the liquid-cooled server room, while the liquid-cooled cabinet is deployed in the space outside the air-conditioned area of ​​the liquid-cooled server room.

6. The liquid-cooled computer room system according to claim 4, characterized in that, The liquid cooling room is equipped with a raised floor, and secondary side pipelines are laid within the raised floor; the laying height of the raised floor is not less than a preset height; the width of the channel under the raised floor for laying secondary side pipelines is not less than a preset width.

7. The liquid-cooled computer room system according to claim 4, characterized in that, The secondary side pipeline is equipped with a water receiving tray, and a leakage detection device is installed inside the water receiving tray.

8. The liquid-cooled computer room system according to claim 4, characterized in that, The connection interface between the CDU and the primary side pipeline is equipped with a filter with a higher mesh size than the preset value; the water treatment equipment in the liquid cooling room system uses a device with a filtration accuracy higher than the preset filtration accuracy.

9. The liquid-cooled computer room system according to claim 1, characterized in that, The liquid-cooled data center system also includes a power supply system for providing high-voltage DC power to the training cluster, general inference cluster, and scalable inference cluster. The power supply system includes a distribution transformer, a low-voltage distribution cabinet, an uninterruptible power supply (UPS), and a high-voltage DC power supply. The power supply system also includes a backup power supply for working simultaneously with the distribution transformer, low-voltage distribution cabinet, UPS, and high-voltage DC power supply to form a dual power supply, or as a backup power supply for the distribution transformer, low-voltage distribution cabinet, UPS, and high-voltage DC power supply.

10. The liquid-cooled computer room system according to claim 1, characterized in that, The fourth quantity is less than the second quantity.

Citation Information

Patent Citations

  • Multi-refrigeration system combined control method based on artificial intelligence

    CN117425317A

  • Cooling system

    CN117835647A