Processor, computing node, and computing cluster

By introducing universal protocol ports that support multiple communication protocols into processors and computing nodes, the problem of complex computing cluster network architecture is solved, and port utilization is improved and resources are saved.

WO2025237203A1PCT designated stage Publication Date: 2025-11-20HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2025/093947
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-14
Filing Date
2025-05-09
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

The existing network architecture of computing clusters is complex because scale-up and scale-out expansion use different communication protocols and require different network devices, resulting in redundant device setups and wasted resources.

Method used

A processor, computing node, and computing cluster are provided, which adopt a universal protocol port that supports both communication protocol 1 and communication protocol 2, thereby achieving port compatibility, reducing the setup of network devices, and simplifying the network architecture.

Benefits of technology

By sharing network devices, the number of network devices required is reduced, port utilization is improved, the network architecture of the computing cluster is simplified, and resources are saved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025093947_20112025_PF_FP_ABST
    Figure CN2025093947_20112025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the embodiments of the present application are a processor, a computing node, and a computing cluster, which are used for simplifying the network architecture of the computing cluster. The processor comprises: a general-purpose protocol port, which is used for supporting a first communication protocol and a second communication protocol, wherein the second communication protocol is different from the first communication protocol; the first communication protocol is a communication protocol of a first network, the second communication protocol is a communication protocol of a second network, the first network is used for supporting communication between a plurality of processors in one computing node including the processor, the second network is used for supporting communication between the processor and other processors comprised in other computing nodes, and the processor is any one of a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU) and a general-purpose graphics processing unit (GPGPU).
Need to check novelty before this filing date? Find Prior Art

Description

A processor, a computing node and a computing cluster TECHNICAL FIELD

[0001] The present application relates to the field of computers, and in particular to a processor, a computing node and a computing cluster. BACKGROUND

[0002] With the application of large AI models, the underlying computing power demand is increasing day by day. At present, the training parameters of AI large models have soared to the trillion level. Such a huge training task cannot be completed by a single GPU or a single server, and a large number of servers are needed as computing nodes to form a computing cluster through high-speed networks to complete tasks. At present, the computing cluster is usually formed by expanding in two directions, including Scale-Up and Scale-Out.

[0003] Scale-Up refers to vertical expansion, which can also be referred to as expansion in the vertical direction. That is, to enhance the computing power of a computing node, more GPUs, storage devices and memories are usually added in the computing node. Correspondingly, Scale-Out refers to horizontal expansion, which can also be referred to as expansion in the horizontal direction. That is, taking a computing node as a unit, more computing nodes are interconnected to form the above computing cluster to achieve the purpose of enhancing the computing power.

[0004] At present, the network formed by Scale-Up expansion and the network formed by Scale-Out expansion use different communication protocols. In order to support these two different communication protocols, different network devices are needed to support them, resulting in a relatively complex network architecture of the computing cluster. SUMMARY

[0005] Embodiments of the present application provide a processor, a computing node and a computing cluster to reduce the setting of network devices and simplify the network architecture of the computing cluster.

[0006] In a first aspect, the present application provides a processor, which comprises a general protocol port capable of supporting a communication protocol 1 and a communication protocol 2, wherein the communication protocol 1 is different from the communication protocol 2. The processor can be a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU), or a general purpose graphics processing unit (GPGPU), or other general processors such as a central processing unit (CPU), which is not limited herein. Taking the processor as an example, the network 1 is used to support the communication between a plurality of GPUs in a computing node including the GPU, the communication protocol is the communication protocol 1, the network 2 is used to support the communication between the GPU and other GPUs included in other computing nodes, and the communication protocol is the communication protocol 2. Since the port included in the GPU is capable of supporting the communication protocol 1 and the communication protocol 2 at the same time, the network 1 and the network 2 can be combined to reduce the setting of network devices, thereby simplifying the network architecture of the computing cluster. Further, compared with the prior art, since the port included in the GPU is capable of supporting the communication protocol 1 and the communication protocol 2 at the same time, the maximum number of available ports of a single network, such as the network 1 or the network 2, can be improved.

[0007] In a possible implementation, the processor further comprises a single port supporting the communication protocol 1; or the processor further comprises a single port supporting the communication protocol 2; or the processor further comprises a single port supporting the communication protocol 1 and a single port supporting the communication protocol 2, wherein the number of ports supporting the two communication protocols and the number of ports supporting a single communication protocol can be set according to actual needs.

[0008] In a possible implementation, the network 1 is a network formed by vertically expanding the processors, and the vertical expansion can be a Scale-Up expansion. The network 2 is a network formed by horizontally expanding the computing nodes, and the horizontal expansion can be a Scale-Out expansion. The Scale-Out expansion can form a larger computing cluster scale.

[0009] In a possible implementation, the general protocol port comprises a communication interface 1 and a communication interface 2. The communication interface 1 is used to support the communication protocol 1. If a data packet needs to be forwarded in the network 1, the data packet is directly processed by the communication interface 1. If a data packet needs to be forwarded in the network 2, the data packet is processed by the communication interface 2, so that the technical solution can realize that a port can support two different communication protocols at the same time.

[0010] In a possible implementation, the processor further comprises other general protocol ports, and the communication interface 2 is a shared communication interface of the other general protocol ports and the general protocol port, which is used to encapsulate and forward the data packet based on the communication protocol 2 when the data packet needs to be forwarded in the network 2.

[0011] In a second aspect, the present application provides a computing node comprising a plurality of processors as described in the first aspect, a switch chip for connecting the plurality of processors, wherein the switch chip comprises a general protocol forwarding port for supporting a communication protocol 1 and a communication protocol 2.

[0012] In a possible implementation, the switch chip further comprises a forwarding port for single supporting the communication protocol 1.

[0013] In a possible implementation, the switch chip further comprises a forwarding port for single supporting the communication protocol 2.

[0014] In a possible implementation, the switch chip further comprises a forwarding port for single supporting the communication protocol 1 and a forwarding port for single supporting the communication protocol 2.

[0015] In a third aspect, the present application further provides a computing cluster comprising a plurality of computing nodes as described in the second aspect, and a primary switch device for connecting the plurality of computing nodes, wherein the primary switch device comprises a general protocol switch port for supporting the communication protocol 1 and the communication protocol 2.

[0016] In a possible implementation, the primary switch device further comprises a primary switch port for single supporting the communication protocol 2.

[0017] In a possible implementation, the computing cluster further comprises a secondary switch device for connecting the plurality of primary switch devices, wherein the secondary switch device comprises a secondary switch port for single supporting the communication protocol 2, and through the secondary switch device, the cluster size can be further expanded.

[0018] In a fourth aspect, the present application further provides a switch device comprising a general protocol port for supporting a communication protocol 1 and a communication protocol 2, wherein the communication protocol 1 is different from the communication protocol 2, the communication protocol 1 is a communication protocol of a network 1, and the communication protocol 2 is a communication protocol of a network 2, wherein the network 1 is used for supporting communication between a plurality of processors included in a computing node, the network 2 is used for supporting communication between the processors included in the computing node and other computing nodes, and the switch device is used for connecting the computing node and the other computing nodes. BRIEF DESCRIPTION OF DRAWINGS

[0019] Fig. 1 is a network architecture of a computing cluster in the prior art;

[0020] Fig. 2 is a schematic diagram of an architecture of a computing node in Fig. 1;

[0021] Fig. 3 is a schematic diagram of another architecture of a computing node in the prior art;

[0022] Fig. 4 is a schematic diagram of ports included in a processor provided by the present application;

[0023] Figure 5 is a structural diagram of a communication protocol port of a processor provided by the present application;

[0024] Figure 6 is an architectural diagram of a computing node provided by the present application;

[0025] Figure 7 is a diagram of ports included in a switching chip provided by the present application;

[0026] Figure 8 is an architectural diagram of a computing cluster provided by the present application;

[0027] Figure 9 is a diagram of ports included in a primary switching device provided by the present application. DETAILED DESCRIPTION

[0028] To make the purpose, technical solutions and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below with reference to the drawings.

[0029] Currently, the following different networking methods exist in the prior art to implement Scale-Up networks and Scale-Out networks, which are introduced as follows:

[0030] The first method is shown in Figure 1, which includes 32 computing nodes, each of which includes 8 GPUs, GPU0-GPU7, and the 32 nodes have a total of 256 GPUs, which are connected through an NVLink network to form a super node (Super POD). The expansion from one GPU to 8 GPUs and the expansion from 1 computing node to 32 computing nodes can both be referred to as Scale-Up direction expansion. In the Scale-Up scenario, the number of computing nodes is usually limited, so to further expand the interconnection range, multiple Super PODs can be connected to form a larger scale GPU cluster, wherein the multiple Super PODs can be connected through an InfiniBand network, an Ethernet, a Unified Bus (UB)-Global. The expansion from one Super POD to multiple Super PODs can be referred to as Scale-Out direction expansion.

[0031] Please refer to FIG. 2, which is an architecture diagram of one of the 32 computing nodes, the architecture diagram of the computing node includes 8 GPUs, namely GPU0-GPU7, each GPU is interconnected within one computing node through NVLink, 4 NVSwitches (abbreviated as NVS in FIG. 2). For nodes outside, it is also connected to the ConnectX-7 (abbreviated as CX-7 in FIG. 2) network module through the Peripheral Component Interconnect Express (PCIe) port, one CX-7 network module includes 4 CX-7 network cards, and then accesses the IB network through the CX-7 network card. Although the export bandwidth of the GPU can be 900GB / s+128GB / s, because the ports supporting different protocols are physically separated, flexible time division multiplexing cannot be performed, that is, if it is desired to use the bandwidth of the NVLink protocol to be 1000GB / s, even if there is more than 1000GB / s of export bandwidth in the physical, it cannot be achieved.

[0032] Because the Scale-Up expansion and the Scale-Out expansion use different network protocols, in order to efficiently process these two protocols, different network devices need to be used, for example, in order to perform Scale-Up direction expansion, NVSwitch needs to be used, and in order to perform Scale-Out direction expansion, ConnectX-7 intelligent network card, IB / Eth switch and the like need to be used, which makes the network formed by Scale-Up expansion and the network formed by Scale-Out expansion two independent networks.

[0033] The second kind: please refer to FIG. 3, the server shown in FIG. 3 includes 8 GPUs, the GPUs are interconnected through a PCIe switch, each GPU includes 24 Ethernet interfaces with a bandwidth of 100G / s, 21 of which are used for interconnection between GPUs in the server, that is, each GPU has 2.1 TB / s of bandwidth resources for interconnection with the remaining 7 GPUs, and the remaining 3 Ethernet interfaces are used for interconnection between Scale-Out expanded machines, wherein the Scale-Out expanded interface can be a double-density quad small form factor pluggable (QSPF-DD), further, the server further includes 10 100G Ethernet ports and 16 PCIe 4.0 interfaces, wherein the Ethernet ports are used for interconnection between processors in the server, and the PCIe 4.0 interfaces are used to connect host CPUs. In the architecture shown in FIG. 3, the communication protocol used by Scale-Up expansion is the same as the communication protocol used by Scale-Out expansion, and in the case of using the same communication protocol, the cost required for high-performance transmission is large due to the requirements of transmission data volume, transmission distance, and transmission reliability.

[0034] To this end, the application provides a processor, which includes a general protocol port for supporting a communication protocol 1 and a communication protocol 2, the communication protocol 1 is a communication protocol of a network 1, and the communication protocol 2 is a communication protocol of a network 2, wherein the network 1 is a network formed by interconnecting a plurality of processors after Scale-Up expansion of one processor. If one processor is Scale-Up expanded to form a computing node, then the network 2 is a network formed by interconnecting a plurality of computing nodes after Scale-Out expansion of one computing node. Although the network 1 and the network 2 use different communication protocols, the general protocol port included in the processor can support the communication protocol 1 and the communication protocol 2 at the same time, so that the network 1 and the network 2 can be combined into one, without the need to separately set network devices for the network 1 and the network 2, that is, the network 1 and the network 2 can share network devices, thereby saving network device resources. At the same time, in the prior art, the ports on the processor for supporting the communication protocol 1 and the ports for supporting the communication protocol 2 are physically separated from each other, and the port setting mode of the processor in the application can maximize the utilization of each port. Further, in the embodiment of the application, the network 1 and the network 1 use appropriate communication protocols, that is, without the need to use the same communication protocol, thereby reducing protocol processing overhead.

[0035] In a first aspect, referring to FIG. 4, a processor 40 shown in FIG. 4 includes a general protocol port 401, which is configured to support a communication protocol 1 and a general protocol 2, the communication protocol 1 and the communication protocol 2 being different, and the number of the general protocol ports 401 can be M1, M1 being an integer greater than zero.

[0036] The processor 40 can be one of a GPU, a general-purpose graphics processing unit (GPGPU), a neural network processing unit (NPU), a tensor processing unit (TPU), or another special-purpose processor XPU, which is not limited herein. For ease of description, the following description takes the processor 40 as a GPU 40 as an example.

[0037] The GPU 40 is taken as a unit for Scale-Up expansion. After the communication connection between the multiple GPUs 40 after the expansion, a network formed is network 1. If a node including multiple GPUs is taken as a computing node, a computing node is taken as a unit for Scale-Out expansion. After the communication connection between the multiple computing nodes after the expansion, a network formed is network 2. It should be noted that the Scale-Up expansion is not limited to the expansion from one GPU to multiple GPUs. In the specific implementation process, a computing node can be taken as a unit for further Scale-Up expansion to form a super node. In this case, the network formed after the communication connection between the multiple computing nodes can also be referred to as network 1. On this basis, a super node can be taken as a unit for Scale-Out expansion to form a large AI cluster scale. In this case, the network formed after the communication connection between the multiple super nodes is referred to as network 2.

[0038] In the embodiments of the present application, if the plane to be extended by Scale-Up is referred to as a Scale-Up plane, then the network 1 can be referred to as a Scale-Up plane network. Of course, the plane to be extended by Scale-Up can also be referred to as a bus plane, a SuperPOD plane, a High-Bandwidth Domain, or an in-chassis interconnection. If the plane to be extended by Scale-Out is referred to as a Scale-Out plane, then the network 2 can be referred to as a Scale-Out plane network. The networking manner of the Scale-Up plane network can be any one or a combination of Torus topology, Mesh topology, or Clos topology. The Torus topology is a multi-dimensional ring structure. In the Torus, each node is connected to its adjacent nodes to form a ring structure. The Torus topology structure can achieve efficient communication and routing, because each node only needs to know the location of its neighbor nodes to communicate. The Mesh topology is a mesh topology structure similar to the Torus, but it is a planar structure. In the Mesh, each node is connected to its adjacent nodes above, below, left, and right to form a planar mesh structure. The Mesh topology structure can achieve efficient communication and routing, but it needs more links and switches to connect all nodes. The Clos is a three-layer structure topology, which is composed of multiple switches. In the Clos, each switch has multiple input ports and multiple output ports, and the input ports and output ports can be connected to form a three-layer structure. The Clos topology structure can achieve efficient communication and routing, because it can load balance and failover between multiple switches. The networking manner of the Scale-Out plane can be the Clos topology or the Optical cross-connect (OXC), or a combination of the two, or through other networking manners, which are not limited herein.

[0039] After the concepts of network 1 and network 2 are introduced, the communication protocol used by network 1 is communication protocol 1, and the communication protocol used by network 2 is communication protocol 2, that is, the GPUs in network 1 communicate with each other using communication protocol 1, and the GPUs in different network 1 communicate with each other using communication protocol 2. The communication protocol 1 includes but is not limited to NVLink protocol, Infinity protocol, Compute Express Link (CXL) protocol, PCIe protocol, Unified Bus (UB)-Clan protocol, or other communication protocols. The communication protocol 2 includes but is not limited to Ethernet protocol, InfiniBand protocol or UB-G protocol, or other communication protocols with good scalability and communication performance, which are not limited in the embodiment of the present application.

[0040] The specific implementation of the general protocol port 401 included in the GPU 40 will be introduced below, please refer to FIG. 5, the general protocol port 401 includes module 4011 and module 4012, wherein the module 4011 is used to support the communication protocol 1, and the module 4012 is used to support the communication protocol 2. In this implementation, the general protocol port 401 further includes a detection module 4013 and a switch module 4014, the detection module 4013 is in communication connection with the switch module 4014, and the switch module 4014 is selectively in communication connection with the module 4011 and the module 4012. The detection module 4013 is used to detect whether the data packet is a data packet to be sent and whether it needs to be forwarded in network 1 or network 2. In the specific implementation process, it is assumed that the GPU 40-1 needs to send a message to the GPU 40-2, the GPU 40-1 can know the destination network address of the GPU 40-2, and then can know the destination port number of the GPU 40-2 for receiving the message, so that the GPU 40-1 can determine whether the to-be-sent message needs to be forwarded in network 1 or network 2 according to the destination port number. In some possible implementation, the GPU 40 can also determine whether the to-be-sent data needs to be forwarded in network 1 or network 2 through other ways, specifically, an identification bit is set in the to-be-sent data packet, the identification bit is used to identify whether the data packet needs to be forwarded in network 1 or network 2, and then determine the processing mode of the data packet. Through this way, the forwarding network of the data packet can be quickly determined, so that the message forwarding efficiency can be improved. When the data packet needs to be forwarded in network 1, the switch module 4014 is controlled to be in communication connection with the module 4011, so as to encapsulate and forward the data packet based on the communication protocol 1, and when the data packet needs to be forwarded in network 2, the switch module 4014 is controlled to be in communication connection with the module 4012, so as to encapsulate and forward the data packet based on the communication protocol 2.

[0041] In the embodiments of the present application, when there are multiple general protocol ports 401, the module 4012 can be exclusive to one general protocol port 401, that is, one general protocol port 401 is configured with one module 4012. In some possible implementation manners, the module 4012 can also be shared by multiple general protocol ports 401 included in the GPU, but when the number of packets to be forwarded in the network 2 is large, the module 4012 should also include a data buffer for buffering data packets to be processed. The data buffer can be a memory for temporarily storing data, and the memory includes at least two types of memories, for example, the memory can be a random access memory, for example, a dynamic random access memory (DRAM). The memory can also include other random access memories, for example, a static random access memory (SRAM) and the like. In addition, when the number of packets to be forwarded is large, a certain queuing strategy can also be set, for example, first-in first-out, that is, the data packets that enter first are processed first, and the data packets that enter later are processed later.

[0042] In addition to supporting the communication protocol 1 and the communication protocol 2 through the hardware mode described above, the general protocol port 401 can also be implemented in the form of software. Specifically, the software-defined network (SDN) technology can be used to receive data packets through a hardware interface, that is, the SDN controller processes the data packets according to a flow table to implement the conversion of the protocols.

[0043] After introducing the general protocol port 401, in some possible implementation manners, the GPU 40 further includes: a first protocol port 402 configured to support the communication protocol 1, wherein the number of the first protocol ports 402 is M2, and M2 is an integer greater than zero; in some possible implementation manners, the GPU 40 further includes: a second protocol port 403 configured to support the communication protocol 2, wherein the number of the second protocol ports 403 is M3, and M3 is an integer greater than zero; in some possible implementation manners, the GPU 40 further includes: the first protocol port 402 and the second protocol port 403, wherein the first protocol port 402 is configured to support the communication protocol 1, the second protocol port 403 is configured to support the communication protocol 2, the number of the first protocol ports 402 is M2, the number of the second protocol ports 403 is M3, and M2 and M3 are integers greater than zero. That is, in the embodiment of the present application, in addition to the general protocol port 401, other protocol ports supporting only one communication protocol can also be set, so that the application scenarios of the GPU 40 can be expanded, and the number of the general protocol port 401 and the protocol port supporting only one communication protocol can be set according to actual needs, which is not limited here.

[0044] In the second aspect, referring to FIG. 6, the present application further provides a computing node 60, the computing node 60 shown in FIG. 6 includes a plurality of GPUs 40 shown in FIG. 4 and a switch chip 601, wherein the switch chip 601 includes a general protocol forwarding port 6011 configured to support the communication protocol 1 and the communication protocol 2 in the first aspect. The specific implementation manner of the general protocol forwarding port 6011 can be referred to the description of the general protocol port 401 in the first aspect, which is not repeated here. The number of the general protocol forwarding ports 6011 is N1, and N1 is an integer greater than zero.

[0045] In the embodiment of the present application, one GPU 40 shown in FIG. 4 is vertically expanded to form a computing node 60, so that the computing power of the computing node 60 can be enhanced, wherein the plurality of GPUs 40 can communicate based on the communication protocol 1.

[0046] In some possible implementation manners, the switch chip 601 further includes first protocol forwarding ports 6012 for supporting the communication protocol 1, the number of the first protocol forwarding ports 6012 is N2, and N2 is an integer greater than zero. In some possible implementation manners, the switch chip 601 further includes second protocol forwarding ports 6013 for supporting the communication protocol 2, the number of the second protocol forwarding ports 6013 is N3, and N3 is an integer greater than zero. In some possible implementation manners, the switch chip 601 further includes the first protocol forwarding ports 6012 and the second protocol forwarding ports 6013, wherein the first protocol forwarding ports 6012 are configured to support the communication protocol 1, and the second protocol forwarding ports 6013 are configured to support the communication protocol 2. A schematic diagram of ports included in the switch chip 601 can be referred to FIG. 7.

[0047] Further, in the embodiment of the present application, the computing node 60 further includes a CPU 602, which can be in communication connection with the switch chip 601 or directly connected with the GPU 40; in the case where the number of the CPU 602 is multiple, part of the CPUs 602 are in communication connection with the switch chip 601, and part of the CPUs 602 are in communication connection with the switch chip 601, and the above three implementation manners are all possible, wherein the CPU 602 and the switch chip 601 or the GPU 40 can be in communication connection through a PCIe bus or other buses, which is not limited in the embodiment of the present application.

[0048] Third aspect, please refer to FIG. 8, the present application further provides a computing cluster 80, the computing cluster 80 includes multiple computing nodes and a first switching device 801, the first switching device 801 is used for interconnecting multiple computing nodes, wherein the first switching device 801 includes a general protocol switching port 8011, which is used for supporting the communication protocol 1 and the communication protocol 2 in the first aspect, and the number of the general protocol switching port 8011 is P1, and P1 is an integer greater than zero. Wherein, the computing nodes included in the computing cluster 80 can be the computing node 60 shown in FIG. 6, or can be a supernode formed by Scale-Up expansion of the computing node 60 shown in FIG. 6, which is not limited in the embodiment of the present application. In the following introduction, the computing nodes included in the computing cluster 80 are taken as the computing node 60 shown in FIG. 6.

[0049] In the embodiment of the present application, the GPU 40 included in the computing node 60 exchanges data through the switch chip 601, and the GPUs 40 in different computing nodes 60 can exchange data through the first switching device 801, without the need of the CX-7 network card shown in FIG. 2, so that the setting of the network device can be reduced, and the network device resources can be saved.

[0050] Further, in the embodiment of the present application, the primary switching device 801 further includes a second protocol primary switching port 8012 for supporting a communication protocol 2, and the number of the second protocol primary switching ports 8012 is P2, P2 is an integer greater than zero. The port arrangement of the primary switching device 801 can be referred to FIG. 9. The primary switching device 801 can further include other components, such as linear-drive pluggable optics (LPO). The so-called “pluggable” means that there is a port of the optical module on the switch, and the corresponding optical module is inserted into the switch, and then the optical fiber can be connected. If the optical fiber or optical module is damaged, it can be repaired and replaced by pulling out the optical module. The interface of the 400G optical module package of the LPO can be divided into a multimode optical module interface, including SR8, S is the first letter of short, which means 100 meters of transmission distance, and “8” is 8 optical signal channels, each channel is 50G; SR4.2, SR is also short, which means 100 meters of transmission distance, “4” is 4 optical fiber channels, and “2” is 2 wavelengths bidirectional multiplexing per channel, each channel is 2x50G; the interface of the 400G optical module package also has a single-mode optical module interface, including FR8 and FR4, where FR represents a distance of 2 kilometers (km), “8” means 8 wavelength multiplexing on one optical fiber, and “4” means 4 wavelength multiplexing on one optical fiber. If the Scale-Out plane adopts the above-mentioned networking mode, the primary switching device 801 adopts the LPO optical module, and the bandwidth of the LPO optical module is 400G. Under normal circumstances, the primary switching device 801 includes 128 interfaces, so that the primary switching device 801 can provide a bandwidth of 50T.

[0051] Further, in the embodiment of the present application, the computing cluster 80 further includes a secondary switching device 802, and the secondary switching device 802 is used to interconnect the primary switching device 801. The secondary switching device 802 includes a second protocol secondary switching port for supporting a communication protocol 2, and can support further expansion of the computing cluster 80.

[0052] In the embodiments of the present application, the number of computing nodes 60 included in the computing cluster 80 and the number of GPUs 40 included in the computing node 60 can be set according to actual needs. In general, the number of GPUs 40 included in each computing node 60 is the same, and the number of GPUs 40 included in one computing node 60 can be 8, 16, 32, 64, 128, or 256. In actual design, the number of GPUs 40 included in the computing node 60 can be set according to the actual cooling scheme adopted. For example, when the cooling scheme adopted is liquid cooling, the number of GPUs 40 included in the computing node 60 can be 64, and when the cooling scheme adopted is air cooling, the number of GPUs 40 included in the computing node 60 can be 16. Wherein, liquid cooling refers to using liquid as a cooling medium to cool the computing devices of the computing node, and air cooling refers to using air as a cooling medium to cool. The number of computing nodes 60 can be 4, 8, 16, 64, or 128. In this way, when the cooling scheme is air cooling, 16 GPUs 40 can be included in one computing node 60, and 120 computing nodes can form a computing cluster of 1920 pieces; when the cooling scheme is liquid cooling, 64 GPUs 40 can be included in one computing node 60, and 128 computing nodes can form a computing cluster of 8912 pieces. Correspondingly, the number of primary switching devices 801 and the number of secondary switching devices 802 can be calculated according to the number of computing nodes 60 and the number of GPUs 40 included in the computing node 60 and the required bandwidth.

[0053] After introducing the GPUs 40, computing nodes 60, switching chips 601, primary switching devices 801, and secondary switching devices 802 included in the computing cluster 80, the port settings of the GPUs 40, switching chips 601, primary switching devices 801, and secondary switching devices 802 when building the computing cluster 80 are further introduced, which are introduced as follows.

[0054] Method one

[0055] The M1 general protocol ports 401 included in the GPU 40 all support communication protocol 1 and communication protocol 2, that is, all support data forwarding of network 1 and data forwarding of network 2. In this way, the number of available ports for data forwarding of network 1 is M1, and the number of available ports for data forwarding of network 2 is also M1. Compared with the prior art, in the case of a total of M1 ports, the number of available ports for data forwarding of network 1 is less than M1, and the number of available ports for data forwarding of network 2 is also less than M1, so that the technical scheme provided by the embodiments of the present application can maximize the number of available ports for data forwarding of network 1 and network 2.

[0056] Switch chip 600 includes N1 general protocol forwarding ports, each of which supports communication protocol 1 and communication protocol 2, that is, each of which supports receiving and forwarding data of network 1 and receiving and forwarding data of network 2.

[0057] It should be noted that the communication protocol 1 supported by GPU 40 needs to be the same as the communication protocol 1 supported by switch chip 601 and level-1 switching device 801, and as an example, the communication protocol 1 supported by GPU 40 is the NVLink protocol, and the communication protocol 1 supported by switch chip 601 and level-1 switching device 801 is also the NVLink protocol; similarly, the communication protocol 2 supported by GPU 40 needs to be the same as the communication protocol 2 supported by switch chip 601 and level-1 switching device 801, and as an example, the communication protocol 2 supported by GPU 40 is the Ethernet protocol, and the communication protocol 2 supported by switch chip 601 and level-1 switching device 801 is also the Ethernet protocol, so as to ensure normal transmission and reception of the switching device.

[0058] Method two

[0059] GPU 40 includes M1 general protocol ports 401, each of which supports communication protocol 1 and communication protocol 2; switch chip includes N1 general protocol forwarding ports 6011, each of which supports communication protocol 1 and communication protocol 2; level-1 switching device 801 includes general protocol switching port 8011 and second protocol level-1 switching port 8012, wherein general protocol switching port 8011 is used to support communication protocol 1 and communication protocol 2, and second protocol level-1 switching port 8012 is used to support communication protocol 2.

[0060] On the basis of the above-mentioned method one and method two, the computing cluster 80 can further include a level-2 switching device 802. In the case where the computing cluster 80 further includes the level-2 switching device 802, the number of level-2 switching devices 802 and the number of level-1 switches 801 satisfy a certain convergence ratio, and as an example, the convergence ratio of level-2 switching device 802 and level-1 switching device 801 is 7:1.

[0061] Method three

[0062] GPU 40 includes M1 general protocol ports 401, which are used to support communication protocol 1 and communication protocol 2, and M2 first protocol ports 402, which are used to support communication protocol 1, and the sum of M1 and M2 is M. In this way, the number of available ports for forwarding data of network 1 is M, and the number of available ports for forwarding data of network 2 is also M1. Compared with the prior art, the number of data forwarding ports of network 1 is also maximized.

[0063] The switch chip 601 includes N1 general protocol forwarding ports 6011 for supporting the communication protocol 1 and the communication protocol 2; or includes N1 general protocol forwarding ports 6011 and N2 first protocol forwarding ports 6012, wherein the N1 general protocol forwarding ports 6011 can support the communication protocol 1 and the communication protocol 2, and the N2 first protocol forwarding ports 6012 are used for supporting the communication protocol 1.

[0064] For the primary switching device 801, it includes P1 general protocol switching ports 8011 for supporting the communication protocol 1 or the general protocol 2; or includes P1 general protocol switching ports 8011 and P2 second protocol primary switching ports 8012, wherein the P1 general protocol switching ports 8011 can support the communication protocol 1 and the communication protocol 2, and the P2 second protocol primary switching ports 8012 are used for supporting the second communication protocol.

[0065] When the primary switching device 801 includes P1 general protocol ports 8011 and P2 second protocol primary switching ports 8012, the port for connecting the GPU 40 in the primary switching device 801 for supporting the communication protocol 1 and the communication protocol 2 also needs to be a port capable of supporting the two communication protocols.

[0066] Mode four

[0067] The GPU 40 includes M1 general protocol ports 401 for supporting the communication protocol 1 and the communication protocol 2, and M3 second protocol ports 403 for supporting the communication protocol 2. In this way, the number of available data forwarding ports of the network 1 is M1, and the number of available data forwarding ports of the network 2 is M1+M3. Compared with the prior art, at least the number of available data forwarding ports of the network 2 is maximized.

[0068] The switch chip 600 includes N1 general protocol forwarding ports 6011 for supporting the communication protocol 1 and the communication protocol 2; or includes N1 general protocol forwarding ports 6011 and N2 first protocol forwarding ports 6012, wherein the N1 general protocol forwarding ports can support the communication protocol 1 and the communication protocol 2, and the N2 protocol forwarding ports 6011 are used for supporting the communication protocol 1.

[0069] The primary switching device 801 includes P1 general protocol switching ports 8011 for supporting the communication protocol 1 or the general protocol 2; or includes P1 general protocol switching ports 8011 and P2 protocol second protocol primary switching ports 8012, wherein the P1 general protocol switching ports 8011 can support the communication protocol 1 and the communication protocol 2, and the P2 protocol second protocol primary switching ports 8012 are used for supporting the second communication protocol.

[0070] The first exchange device 801 includes P1 general protocol ports 8011, and P2 protocol first exchange ports. The port for connecting the GPU 40 in the first exchange device 801 to support the communication protocol 1 and the communication protocol 2 is also required to support the two communication protocols. The port for connecting the GPU 40 in the first exchange device 801 to support the communication protocol 2 can be a port capable of supporting the communication protocol 2 or a port capable of supporting the communication protocol 1 and the communication protocol 2.

[0071] Mode five

[0072] The GPU 40 includes M1 general protocol ports 401 for supporting the communication protocol 1 and the communication protocol 2, M2 first protocol ports 402 for supporting the communication protocol 1, and M3 second protocol ports 403 for supporting the communication protocol 2. Thus, the number of available data forwarding ports of the network 1 is M1+M2, and the number of available data forwarding ports of the network 2 is M1+M3. Compared with the prior art, the number of available data forwarding ports of the network 1 and the number of available data forwarding ports of the network 2 are maximized.

[0073] In this implementation mode, the exchange chip 601 includes N1 general protocol forwarding ports 6011 for supporting the communication protocol 1 and the communication protocol 2, N2 first protocol forwarding ports 6012 for supporting the communication protocol 1, and N3 second protocol forwarding ports 6013 for supporting the communication protocol 2. In this implementation mode, the port for connecting the GPU 40 in the first exchange device 801 to support the communication protocol 1 and the communication protocol 2 is also required to support the two communication protocols. The port for connecting the GPU 40 in the first exchange device 801 to support the communication protocol 1 can be a port capable of supporting the communication protocol 1 or a port capable of supporting the communication protocol 1 and the communication protocol 2. The port for connecting the GPU 40 in the first exchange device 801 to support the communication protocol 2 can be a port capable of supporting the communication protocol 2 or a port capable of supporting the communication protocol 1 and the communication protocol 2.

[0074] In the specific implementation process, the number of ports for supporting only the communication protocol 1, the number of ports for supporting only the communication protocol 2, and the number of ports for supporting both the communication protocol 1 and the communication protocol 2 on the GPU 40, the first exchange device 801, and the second exchange device 802 can be set according to actual needs. In actual application, a reasonable congestion control algorithm can be allocated according to the size of the data flow of the network 1 and the data flow of the network 2 to dynamically call these ports and improve the link utilization rate.

[0075] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any change or replacement within the technical scope disclosed by the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A processor, comprising: The processor comprises: a general protocol port for supporting a first communication protocol and a second communication protocol different from the first communication protocol; wherein the first communication protocol is a communication protocol of a first network for supporting communication among a plurality of processors including the processor within one computing node, and the second communication protocol is a communication protocol of a second network for supporting communication between the processor and other processors included in other computing nodes, and the processor is any one of a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU), or a general-purpose graphics processing unit (GPGPU).

2. The processor of claim 1, wherein, The processor further comprises: a first protocol port for supporting the first communication protocol.

3. The processor of claim 1 or 2, wherein, The processor further comprises: a second protocol port for supporting the second communication protocol.

4. The processor of claim 1 or 2, wherein, The first network is a network formed by vertically expanding the processor, and the second network is a network formed by horizontally expanding the computing node.

5. The processor of claim 1 or 2, wherein, The first communication protocol is any one of an NVlink protocol, an Infinity protocol, a CXL protocol, a peripheral component interconnect express (PCIe) protocol, or a unified bus (UB-C) protocol, and the second communication protocol is any one of an Ethernet protocol, an InfiniBand protocol, or a UB-G protocol.

6. The processor of claim 1 or 2, wherein, The general protocol port comprises a first communication interface and a second communication interface; The first communication interface is for supporting the first communication protocol; The second communication interface is for supporting the second communication protocol.

7. A computing node, characterized in that, The processor comprises: a plurality of processors as claimed in claim 1; a switch chip for connecting the plurality of processors; wherein the switch chip comprises a general protocol forwarding port for supporting the first communication protocol and the second communication protocol.

8. The computing node of claim 7, wherein, The switch chip further comprises: a first protocol forwarding port for supporting the first communication protocol.

9. The computing node according to claim 7 or 8, c h a r a c t e r i z e d b y The switch chip further comprises: a second protocol forwarding port for supporting the second communication protocol.

10. A computing cluster, characterized by, The computing cluster comprises: a plurality of computing nodes as claimed in claim 9; a primary switching device for connecting the plurality of computing nodes, wherein the primary switching device comprises a general protocol switching port for supporting the first communication protocol and the second communication protocol.

11. The computing cluster of claim 10, wherein, The primary switching device further comprises: a second protocol primary switching port for supporting the second communication protocol.

12. The computing cluster according to claim 10 or 11, c h a r a c t e r i z e d b y The computing cluster further comprises: a secondary switching device for connecting a plurality of the primary switching devices, wherein the secondary switching device comprises: a second protocol secondary switching port for supporting the second communication protocol.

13. A switching device, characterized by The computing cluster comprises: a general protocol switching port for supporting a first communication protocol and a second communication protocol different from the first communication protocol; The first communication protocol is a communication protocol of a first network, the second communication protocol is a communication protocol of a second network, the first network supports communication between a plurality of processors included in a computing node, the second network is used to support communication between a processor included in the computing node and another processor included in another computing node, and the switching device is used to connect the computing node and the another computing node.

14. The switching device of claim 13, wherein, The switching device further comprises: A second protocol exchange port used to support the second communication protocol.

Citation Information

Patent Citations

  • System and manufacture method of multi-node configuration of processor cards connected via processor fabrics

    CN101324877A

  • Processor topology switches

    CN102461088A

  • Systems, methods, and apparatus for storage controller with multiple heterogeneous network interface ports

    CN111367844A

  • Computing device, management controller and data processing method

    CN117873924A

  • Intelligent network processor and method of using intelligent network processor

    US20070266179A1

Cited By

  • Semantic exchange chip, computer system, semantic exchange method, network interface card, device, medium and program product

    CN122309441A