A data center

CN119342017BActive Publication Date: 2026-09-18BEIJING KUSHU INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411461814.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2026-09-18
Estimated Expiration
2044-10-18

AI Technical Summary

Technical Problem

[0005]本发明要解决的技术问题是:针对现有技术的上述缺陷,提供一种数据中心,解决传统数据中心布局中大量使用光电转换模块导致的可靠性、延时高、成本高、功耗高等问题

Benefits of technology

[0021]The present invention has the following beneficial effects: The innovative design of the present invention adopts a three-layer ring structure layout. Each cluster unit is a three-layer ring structure, consisting of a first ring layer, a second ring layer, and a third ring layer from the inside out. The cluster unit is radially divided into multiple modules. Each module's first ring layer is equipped with a network cabinet containing a Spine switch. Each module's second and third ring layers are equipped with server racks, each containing multiple GPU servers. The top layer of the server racks in each module's second ring layer is equipped with a Leaf switch. By rationally setting the distance between each ring layer and carefully arranging the positions of the Spine and Leaf switches, the distance between the GPU servers and Leaf switches within each module, the distance between Leaf switches and Spine switches, and the distance between Spine switches within each cluster unit are all less than 5 meters. The GPU servers and Leaf switches, Leaf switches and Spine switches, and Spine switches are all connected via copper cables. This solution improves network reliability, reduces network latency, networking costs, and power consumption by optimizing the network layout, reducing the distance between components, and using copper cables instead of fiber optic cables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119342017B_ABST
    Figure CN119342017B_ABST
Patent Text Reader

Abstract

The application discloses a kind of data centers, belong to data center process layout technical field.The cluster unit of the scheme data center is three-layer annular structure, and is divided into multiple modules along radial direction;First annular layer is deployed spine switch, second, third annular layer is deployed GPU server cabinet, and second annular layer is also deployed leaf switch;By setting the distance between each layer, the position of spine switch and leaf switch, the distance between GPU server in each module and leaf switch, leaf switch and spine switch, the spine switch in cluster unit is less than 5 meters, and GPU server and leaf switch, leaf switch and spine switch, between spine switch are all connected by copper cable.This scheme improves the reliability of network, reduces network delay, networking cost and power consumption by reasonable layout.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data center layout technology, and more particularly to a data center. Background Technology

[0002] With the rapid development of high-performance computing, artificial intelligence, especially generative artificial intelligence and large-scale pre-trained models, humanity is rapidly moving towards general artificial intelligence. The current demand for AI computing power is increasing daily, and traditional data center layouts are increasingly unable to meet the requirements of high-performance computing. The core technology of current AI computing power mimics the connection methods of the human brain's neural networks. This is achieved by stacking simple GPU (Graphics Processing Unit) cores using efficient packaging technologies (such as CoWoS packaging technology) and ultimately packaging them on a substrate to form a 2.5D or 3D form, enabling efficient interconnection between GPUs and HBM (High Bandwidth Memory). Furthermore, NVLink technology enables high-speed interconnection between GPUs within servers, and lossless networks based on RDMA (Remote Direct Memory Access) technology enable high-speed interconnection between each GPU within the server. This hardware environment requires that any GPU can achieve high-speed, lossless interconnection with all other GPUs, posing a significant challenge to network organization and design.

[0003] Currently, the most common layout for data centers is a linear column layout. This layout has been widely used for a long time, primarily due to its focus on maximizing space utilization. However, traditional data center racks are rectangular, resulting in low power density, typically only 3-5-7 kilowatts. Furthermore, the business model is dominated by north-south traffic, with very little east-west traffic that is not sensitive to latency, making it difficult to meet the high bandwidth and low latency communication requirements of current computing scenarios. Because of the significant distance between server racks and switches, traditional data centers are only suitable for fiber optic-based networking. Each GPU server's network interface requires a photoelectric conversion module to convert electrical signals within the GPU server into laser signals, which are then transmitted to the switch port. The photoelectric conversion module then converts the laser signals back into electrical signals for switching and distribution. Subsequently, the laser signals are again converted back into laser signals from the switch port to the network interface of another GPU server, where they are converted back into electrical signals for processing. The structure of the photoelectric conversion module is very complex, and the reliability of the link is affected by the number of photoelectric conversions. For example, in NVIDIA's networking solution based on 64-port InfiniBand switches, using leaf layer switches to network fewer than 32 H100 GPU servers requires four photoelectric conversion modules for continuous conversion (GPU-Leaf-GPU) to complete a single link. When there are more than 32 but less than 256 GPU servers, a maximum of eight photoelectric conversion modules are needed for continuous conversion in a single link (GPU-Leaf-Spine-Leaf-GPU), at which point the reliability of a single optical transmission link decreases by nearly 10 times. When there are more than 256 GPU servers, a third core layer network is required, at which point a single optical transmission link requires 12 photoelectric conversions (GPU-Leaf-Spine-Core-Spine-Leaf-GPU), further reducing the reliability of a single link. During large model training, any photoelectric conversion failure can lead to congestion of training data and an increased system failure rate, thereby affecting training efficiency and significantly increasing costs.

[0004] Furthermore, each photoelectric conversion module has a latency of 30-40 nanoseconds. Combined with long-distance fiber optic links and a large number of photoelectric conversion modules, this latency impacts the high-frequency, high-volume data throughput efficiency during large model training, extending training time and increasing costs. In addition, the large number of photoelectric conversion modules in the network leads to high networking costs and high power consumption. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a data center that addresses the above-mentioned deficiencies of the prior art and solves the problems of high reliability, high latency, high cost, and high power consumption caused by the extensive use of photoelectric conversion modules in traditional data center layouts.

[0006] To achieve the above objectives, the present invention provides a data center based on a networking scheme using NVIDIA's 64-port Infiniband switch, characterized in that it includes a single-story or multi-story building, and each floor of the building includes at least one cluster unit.

[0007] The cluster unit is a three-layer ring structure, consisting of a first ring layer RL1, a second ring layer RL2, and a third ring layer RL3 from the inside out.

[0008] The cluster unit is radially divided into A modules; each module has a network cabinet 1 in the first annular layer RL1, and B Spine switches in the network cabinet 1; each module has a server rack 2 in both the second annular layer RL2 and the third annular layer RL3, and D GPU servers in each server rack 2; and E Leaf switches are installed on the top layer of the server rack 2 in the second annular layer RL2 of each module.

[0009] Each of the GPU servers is equipped with multiple GPU cards, and the GPU cards of each GPU server support low-latency interconnection communication.

[0010] A first distance L1 exists between the outer edge of the first annular layer RL1 and the inner edge of the second annular layer RL2; a second distance L2 exists between the outer edge of the second annular layer RL2 and the inner edge of the third annular layer RL3; by reasonably setting the first distance L1 and the second distance L2, as well as the positions of the spine switch and the leaf switch, the distances between the GPU server and the leaf switch in each module, between the leaf switch and the spine switch, and between the spine switches in the cluster unit are all less than 5 meters; the GPU server in each module is connected to the leaf switch in the same module via a copper cable; the leaf switch in each module is connected to the spine switch in the same module via a copper cable; the spine switches in the cluster unit are connected to each other via copper cables.

[0011] Each of the network cabinets 1 is also equipped with a core switch. The spine switch and the core switch are connected by copper cables, and the cluster units are interconnected through the core switch.

[0012] Preferably, an installation and maintenance channel is reserved in the middle of the first annular layer RL1, the second annular layer RL2, and the third annular layer RL3.

[0013] Preferably, both the spine switch and the leaf switch have 64 ports, wherein 32 downlink ports of the leaf switch are used to connect to the GPU server and 32 uplink ports are used to connect to the spine switch; 32 downlink ports of the spine switch are used to connect to the leaf switch and 32 uplink ports are used to connect to other spine switches.

[0014] Preferably, A=8, B=8, C=4, D=4, E=2, and each GPU server is equipped with 8 GPU cards; the first distance L1 and the second distance L2 are both 1.2 meters.

[0015] Preferably, the core switches of multiple cluster units on the same floor are connected by optical fiber; when the distance between the core switches of two cluster units on adjacent floors is less than 5 meters, copper cables are used to connect the core switches of the two cluster units; otherwise, optical fiber cables are used to connect the core switches of the two cluster units.

[0016] Preferably, the GPU card in the GPU server is an NVIDIA H100.

[0017] Preferably, the GPU card in the GPU server is an NVIDIA H200.

[0018] Preferably, a wire mesh is provided between the top of the first annular layer RL1 and the top of the second annular layer RL2, and between the top of the second annular layer RL2 and the top of the third annular layer RL3, and the copper cable connecting the adjacent annular layers is fixed and supported by the wire mesh.

[0019] Preferably, the server racks 2 in the second annular layer RL2 and the third annular layer RL3 of each module are arranged in a row along the edge of the annular layer. A power distribution unit 3 is provided at the head of the row of server racks, and the power distribution unit 3 is connected to the server rack 2 by a cable.

[0020] Preferably, each of the third annular layer RL3 of the module is equipped with a cooling and dehumidifying unit 4 at the head or tail of the server rack queue. The GPU chip and CPU chip in the GPU server are fixedly attached to the water-cooled plate with thermally conductive adhesive. The water-cooled plate is connected to the cooling and dehumidifying unit 4 through stainless steel water-cooling pipes. The network cabinet 1, server rack 2, and power distribution unit 3 are all equipped with cabinet doors 5. A chilled water fan coil unit cabinet panel 6 is provided on the outer side of the cabinet opposite to the cabinet door 5.

[0021] The present invention has the following beneficial effects: The innovative design of the present invention adopts a three-layer ring structure layout. Each cluster unit is a three-layer ring structure, consisting of a first ring layer, a second ring layer, and a third ring layer from the inside out. The cluster unit is radially divided into multiple modules. Each module's first ring layer is equipped with a network cabinet containing a Spine switch. Each module's second and third ring layers are equipped with server racks, each containing multiple GPU servers. The top layer of the server racks in each module's second ring layer is equipped with a Leaf switch. By rationally setting the distance between each ring layer and carefully arranging the positions of the Spine and Leaf switches, the distance between the GPU servers and Leaf switches within each module, the distance between Leaf switches and Spine switches, and the distance between Spine switches within each cluster unit are all less than 5 meters. The GPU servers and Leaf switches, Leaf switches and Spine switches, and Spine switches are all connected via copper cables. This solution improves network reliability, reduces network latency, networking costs, and power consumption by optimizing the network layout, reducing the distance between components, and using copper cables instead of fiber optic cables. Attached Figure Description

[0022] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0023] Figure 1 This is a schematic diagram of the layout structure of a cluster unit consisting of 256 GPU servers.

[0024] Figure 2 This is a network topology diagram of a cluster unit consisting of 256 GPU servers.

[0025] Figure 3 A schematic diagram of the layout structure of a cluster unit consisting of 32 GPU servers.

[0026] Figure 4 A schematic diagram of the layout structure of a cluster unit consisting of 128 GPU servers. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] The general concept of this invention is as follows: each cluster unit has a three-layer ring structure, consisting of a first ring layer, a second ring layer, and a third ring layer from the inside out. The cluster unit is radially divided into multiple modules; each module's first ring layer deploys Spine switches, while the second and third ring layers each deploy GPU server racks. The second ring layer also deploys Leaf switches. By rationally setting the distances between each layer and carefully arranging the positions of the Spine and Leaf switches, the distances between GPU servers and Leaf switches within each module, between Leaf switches and Spine switches, and between Spine switches within each cluster unit are all less than 5 meters. All connections between GPU servers and Leaf switches, between Leaf switches and Spine switches, and between Spine switches are made via copper cables.

[0029] The embodiments of the present invention will now be described in further detail with reference to the accompanying drawings. It should be understood that the embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the scope of the invention.

[0030] like Figure 1 As shown, this embodiment of the invention provides a data center, which is networked based on NVIDIA's 64-port Infiniband switch. The data center includes a single-story or multi-story building, and each floor of the building includes at least one cluster unit. The cluster unit has a three-layer ring structure, consisting of a first ring layer RL1, a second ring layer RL2, and a third ring layer RL3 from the inside out.

[0031] The cluster unit is divided into A modules radially; each module has a network cabinet 1 in the first annular layer RL1, and B Spine switches in the network cabinet 1; each module has C server racks 2 in the second annular layer RL2 and the third annular layer RL3, and D GPU servers in each server rack 2; and E Leaf switches are installed on the top layer of the server racks 2 in the second annular layer RL2 of each module.

[0032] The structure of a single module is as follows Figure 3 As shown. Figure 1There are a total of 8 modules, corresponding to M1-M8 in the diagram.

[0033] Each of the GPU servers is equipped with multiple GPU cards, and the GPU cards of each GPU server support low-latency interconnection communication.

[0034] There is a first distance L1 between the outer edge of the first annular layer RL1 and the inner edge of the second annular layer RL2; there is a second distance L2 between the outer edge of the second annular layer RL2 and the inner edge of the third annular layer RL3; by reasonably setting the first distance L1 and the second distance L2, as well as the positions of the spine switch and the leaf switch, the distance between the GPU server and the leaf switch in each module, between the leaf switch and the spine switch, and between the spine switches in the cluster unit are all less than 5 meters; the GPU server in each module is connected to the leaf switch in the same module via a copper cable; the leaf switch in each module is connected to the spine switch in the same module via a copper cable; the spine switches in the cluster unit are connected to each other via copper cables. In practical applications, several fixed copper cable specifications can be standardized based on the distance from the GPU server to the leaf switch, the distance from the leaf switch to the spine switch, and the distance between spine switches. For example, copper cables of lengths of 2 meters, 3 meters, and 4 meters can be selected according to the actual distance to improve cable consistency and enhance the accuracy of topology awareness.

[0035] Each of the network cabinets 1 is also equipped with a core switch. The spine switch and the core switch are connected by copper cables, and the cluster units are interconnected through the core switch.

[0036] In this embodiment of the invention, installation and maintenance channels are reserved in the middle of the first annular layer RL1, the second annular layer RL2, and the third annular layer RL3. For example... Figure 1 As shown, the cluster unit has a left-right symmetrical structure with 4 modules on each side. There are reserved installation and maintenance channels between modules M1 and M8, and between modules M4 and M5. The installation and maintenance channels pass through the first annular layer RL1, the second annular layer RL2, and the third annular layer RL3.

[0037] In this embodiment of the invention, both the spine switch and the leaf switch contain 64 ports. The leaf switch has 32 downlink ports for connecting to the GPU server and 32 uplink ports for connecting to the spine switch. The spine switch has 32 downlink ports for connecting to the leaf switch and 32 uplink ports for connecting to other spine switches.

[0038] like Figure 1 As shown, in some embodiments of the present invention, A=8, B=8, C=4, D=4, E=2, and each GPU server is equipped with 8 GPU cards. Figure 1 In this cluster unit, there are eight modules. Each module forms a fan-shaped structure radially along the cluster unit. The second annular layer RL2 and the third annular layer RL3 of each module form an approximately octagonal structure. Each module's first annular layer RL1 is equipped with eight 64-port Spine switches. The second annular layer RL2 and the third annular layer RL3 each have four server racks 2, each containing four GPU servers. Each server rack 2 in the second annular layer RL2 has two 64-port Leaf switches. Each GPU server is equipped with eight GPU cards. Therefore, Figure 1 In the illustrated embodiment, each module contains 8x4 = 32 GPU servers, and a cluster unit has a total of 8x32 = 256 GPU servers. Each module contains 32x8 GPU cards, requiring 8 leaf switches and 8 spine switches with 64 ports each. The 32 downlink ports of the leaf switches are used to connect to the GPU servers, and the 32 uplink ports are used to connect to the spine switches; the 32 downlink ports of the spine switches are used to connect to the leaf switches, and the 32 uplink ports are used to connect to the other spine switches. Figure 2 for Figure 1 The network topology diagram shown is shown.

[0039] Figure 1 In the embodiment described, two network cabinets 1 are arranged in the first annular layer of each module, and four 64-port Spine switches are arranged on the top layer of each network cabinet 1. In practical applications, the number of network cabinets 1 and the number of Spine switches in each network cabinet 1 can be set as needed.

[0040] In practical applications, if a GPU server is equipped with 4 GPU cards, 16 GPU cards, or a combination of GPU servers with different numbers of GPU cards, the number of leaf switches, spine switches, server racks, and servers within the server racks can be adjusted according to the actual situation.

[0041] In this embodiment of the invention, when the number of GPU servers is less than 256, the layout is still based on 32 GPU servers per module. Figure 3 The diagram shows the layout of 32 GPU servers, which together form one module. Figure 4 The diagram shows the layout of 128 GPU servers. The 128 GPU servers form 4 modules, which are arranged consecutively for easy future expansion.

[0042] In this embodiment of the invention, the server rack has a length, width, and height of 1.2 meters, 0.6 meters, and 2-2.5 meters, respectively. Each server rack can accommodate four GPU servers with a height of 8U-10U, where U is the basic unit for measuring device height within a server rack, and 1U ≈ 44.45mm. The GPU cards in the GPU servers are either NVIDIA H100 or NVIDIA H200. The NVIDIA H100 is a high-performance GPU accelerator launched by NVIDIA, designed specifically for data centers and high-performance computing (HPC). The H100 is based on the NVIDIA Hopper architecture and supports NVLink high-speed interconnect technology, enabling high-speed communication between multiple GPUs. It is suitable for high-performance computing, deep learning training, deep learning inference, and data center acceleration. The NVIDIA H200 is the next-generation product of the NVIDIA H100, with the same size and power consumption. It should be noted that the GPU cards supported by this invention are not limited to NVIDIA H100 and NVIDIA H200. In practical applications, the size of the server rack can be adjusted according to the model of the GPU card and the size of the GPU server.

[0043] In this embodiment of the invention, both the first distance L1 and the second distance L2 are 1.2 meters. A leaf switch is installed on the top layer of the server rack 2 in the second ring layer RL2, and a spine switch is installed on the top layer of the network rack 1 in the first ring layer RL1. This ensures that the distance between the GPU server and the leaf switch in each module, the distance between the leaf switch and the spine switch, and the distance between the spine switches in the cluster unit are all less than 5 meters. The GPU server and the leaf switch, the leaf switch and the spine switch, and the spine switches can all be connected via copper cables.

[0044] Copper cables have a lower bit error rate compared to other materials, which helps improve the reliability of data transmission and reduces the error rate during data transmission. This is especially important in large-scale clusters, where stable data transmission is crucial for the stability of the entire system. In this embodiment of the invention, a direct connection of all electrical signals is achieved using passive direct-connect cables (DAC, <3m) and active cables (AEC, <5m). Compared to optical cables, DAC / AEC copper cables operate at lower temperatures, consume less power, are less expensive, and have higher reliability. Therefore, this design reduces intermittent network link failures and malfunctions, which are major problems faced by all high-speed interconnects using optical components. Figure 1 The layout shown minimizes the latency differences between downlink ports of both leaf and spine switches, and also reduces the distance differences between spine switches, significantly reducing the latency differences between links in the network. In this embodiment, cable length maintains a positive correlation with communication routing, avoiding the negative impact of communication latency on topology awareness and improving its accuracy. Compared to fiber optic connections, the total cable length is significantly reduced, signal quality is significantly improved, and signal round-trip time is significantly reduced, resulting in significantly improved interconnection efficiency between GPU servers and reduced networking costs.

[0045] In practical applications, the positions of the first distance L1, the second distance L2, and the leaf switches, spine switches, and server racks can be adjusted according to the specific conditions of the data center and industry standards. This ensures that the distance between the GPU server and the leaf switch, the distance between the leaf switch and the spine switch, and the distance between the spine switches in each module are all less than 5 meters or as far as possible less than 5 meters, in order to minimize the number of photoelectric conversion modules in the network.

[0046] In some embodiments of the present invention, a wire mesh is provided between the top of the first annular layer RL1 and the top of the second annular layer RL2, and between the top of the second annular layer RL2 and the top of the third annular layer RL3. The copper cables connecting adjacent annular layers are fixed and supported by the wire mesh. The wire mesh ensures a neat and orderly cable arrangement, facilitating subsequent maintenance. In other embodiments of the present invention, a cable tray corresponding to the positions of the first annular layer RL1, the second annular layer RL2, and the third annular layer RL3 is provided above the cluster unit to support and organize the copper cables.

[0047] In this embodiment of the invention, when the number of GPU servers is greater than 256, see [link to relevant documentation]. Figure 1Multiple clusters are constructed as shown, with the core switches of multiple cluster units on the same floor connected via fiber optic cables. When the distance between the core switches of two cluster units on adjacent floors is less than 5 meters, copper cables are used to connect the core switches of the two cluster units; otherwise, fiber optic cables are used. When deploying cluster units on adjacent floors, the cluster units on adjacent floors can be arranged vertically opposite each other. If the distance between the core switches of two cluster units is less than 5 meters, vertical cabling can be carried out through shafts, cable trays, or conduit systems, and copper cables are used to connect the core switches of the two cluster units. Therefore, the solution of this invention has good network scalability, easily enabling expansion from 256 servers to 512 and 1024 servers, and significantly reducing the distance between clusters. This minimizes system latency, improves reliability, provides flexibility and sustainability for future cluster growth, and ensures performance stability during the expansion process.

[0048] like Figure 1 , Figure 3 , Figure 4 As shown in this embodiment of the invention, the server racks 2 in the second annular layer RL2 and the third annular layer RL3 of each module are arranged in a row along the edge of the annular layer. A power distribution unit 3 is provided at the head of the row of server racks. The power distribution unit 3 is connected to the racks, switches and other equipment in the server racks 2 via cables. Each server rack 2 is equipped with two UPS power supplies.

[0049] In this embodiment of the invention, multiple GPU servers are concentrated in one server rack. To extend the lifespan of the GPU chips, precise heat dissipation is required for the GPU servers. Figure 1 , Figure 3 , Figure 4 As shown in this embodiment of the invention, each of the third annular layer RL3 of the module is equipped with a cooling and dehumidification unit 4 at the head or tail of the server rack queue. The GPU chip and CPU chip in the GPU server are fixedly attached to the water-cooled plate with thermally conductive adhesive, and the water-cooled plate is connected to the cooling and dehumidification unit 4 through stainless steel water-cooling pipes. The network cabinet 1, server rack 2, and power distribution unit 3 are all equipped with cabinet doors 5, and a chilled water fan coil unit cabinet panel 6 is provided on the outer side of the cabinet opposite to the cabinet door 5. An air conditioner for automatic dehumidification is also installed in the data center computer room. Through liquid cooling, chilled water fan coil units, and air conditioning, the long-term reliability of the GPU server is improved.

[0050] The present invention has the following beneficial effects: The innovative design of the present invention adopts a three-layer ring structure layout. Each cluster unit is a three-layer ring structure, consisting of a first ring layer, a second ring layer, and a third ring layer from the inside out. The cluster unit is radially divided into multiple modules. Each module's first ring layer is equipped with a network cabinet containing a Spine switch. Each module's second and third ring layers are equipped with server racks, each containing multiple GPU servers. The top layer of the server racks in each module's second ring layer is equipped with a Leaf switch. By rationally setting the distance between each ring layer and carefully arranging the positions of the Spine and Leaf switches, the distance between the GPU servers and Leaf switches within each module, the distance between Leaf switches and Spine switches, and the distance between Spine switches within each cluster unit are all less than 5 meters. The GPU servers and Leaf switches, Leaf switches and Spine switches, and Spine switches are all connected via copper cables. This solution improves network reliability, reduces network latency, networking costs, and power consumption by optimizing the network layout, reducing the distance between components, and using copper cables instead of fiber optic cables.

[0051] The above are merely specific embodiments of the present invention and should not be construed as limiting the scope of the present invention. Equivalent variations made by those skilled in the art based on this invention, as well as changes well-known to those skilled in the art, should still fall within the scope of the present invention.

Claims

1. A data center network based on a 64-port Infiniband switch from NVIDIA, suitable for large-scale GPU training cluster intelligent computing centers, characterized in that... It includes single-story or multi-story buildings, each of which contains at least one cluster unit; The cluster unit is a three-layer ring structure, consisting of a first ring layer (RL1), a second ring layer (RL2), and a third ring layer (RL3) from the inside out. The cluster unit is divided into A modules radially; each module has a network cabinet (1) in the first ring layer (RL1), and B Spine switches in the network cabinet (1); each module has a server rack (2) in both the second ring layer (RL2) and the third ring layer (RL3), and D GPU servers in each server rack (2); E Leaf switches are installed on the top layer of the server rack (2) in the second ring layer (RL2) of each module. Each of the GPU servers is equipped with multiple GPU cards, and the GPU cards of each GPU server support low-latency interconnection communication. There is a first distance (L1) between the outer edge of the first annular layer (RL1) and the inner edge of the second annular layer (RL2); there is a second distance (L2) between the outer edge of the second annular layer (RL2) and the inner edge of the third annular layer (RL3); by reasonably setting the first distance (L1) and the second distance (L2), as well as the positions of the spine switch and the leaf switch, the distance between the GPU server and the leaf switch in each module, between the leaf switch and the spine switch, and between the spine switches in the cluster unit are all less than 5 meters; the GPU server in each module is connected to the leaf switch in the same module via a copper cable; the leaf switch in each module is connected to the spine switch in the same module via a copper cable; the spine switches in the cluster unit are connected to each other via copper cables. Each of the network cabinets (1) is also equipped with a core switch. The spine switch and the core switch are connected by copper cables, and the cluster units are interconnected through the core switch.

2. The data center according to claim 1, characterized in that, Installation and maintenance channels are reserved in the middle of the first annular layer (RL1), the second annular layer (RL2), and the third annular layer (RL3).

3. The data center according to claim 1, characterized in that, Both the spine switch and the leaf switch contain 64 ports. The leaf switch has 32 downlink ports for connecting to the GPU server and 32 uplink ports for connecting to the spine switch. The spine switch has 32 downlink ports for connecting to the leaf switch and 32 uplink ports for connecting to other spine switches.

4. The data center according to claim 3, characterized in that, A=8, B=8, C=4, D=4, E=2, each of the GPU servers is equipped with 8 GPU cards; the first distance (L1) and the second distance (L2) are both 1.2 meters.

5. The data center according to claim 1, characterized in that, The core switches of multiple cluster units on the same floor are connected by optical fiber. When the distance between the core switches of two cluster units on adjacent floors is less than 5 meters, copper cables are used to connect the core switches of the two cluster units; otherwise, optical fibers are used to connect the core switches of the two cluster units.

6. The data center according to claim 1, characterized in that, The GPU card in the GPU server is an NVIDIA H100.

7. The data center according to claim 1, characterized in that, The GPU card in the GPU server is an NVIDIA H200.

8. The data center according to claim 1, characterized in that, A wire mesh is provided between the top of the first annular layer (RL1) and the top of the second annular layer (RL2), and between the top of the second annular layer (RL2) and the top of the third annular layer (RL3). The copper cable connecting the adjacent annular layers is fixed and supported by the wire mesh.

9. The data center according to claim 1, characterized in that, The second annular layer (RL2) of each module and the server rack (2) in the third annular layer (RL3) are arranged in a row of server racks along the edge of the annular layer. A power distribution unit (3) is provided at the head of the server rack row, and the power distribution unit (3) is connected to the server rack (2) by a cable.

10. The data center according to claim 9, characterized in that, Each of the modules is equipped with a cooling and dehumidifying unit (4) at the head or tail of the server rack queue of the third annular layer (RL3). The GPU chip and CPU chip in the GPU server are fixedly attached to the water-cooled plate by thermal adhesive. The water-cooled plate is connected to the cooling and dehumidifying unit (4) by stainless steel water-cooled pipes. The network cabinet (1), server rack (2), and power distribution unit (3) are all equipped with cabinet doors (5). A chilled water fan coil unit cabinet panel (6) is provided on the outer side of the cabinet opposite to the cabinet door (5).

Citation Information

Patent Citations

  • Networking method for datacenter network and datacenter network

    CN107211036A

  • Switch managed resource allocation and software execution

    CN115668886A