Processor, computing node and computing cluster

By introducing universal protocol ports that support multiple communication protocols into processors, computing nodes, and computing clusters, the problem of complex computing cluster network architecture is solved, achieving resource savings and improved transmission efficiency.

CN120950174APending Publication Date: 2025-11-14HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410599159.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-14
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

The network architecture of existing computing clusters is complex because scale-up and scale-out expansion use different communication protocols, which requires different network devices, increasing the complexity of device setup and wasting resources.

Method used

A processor, computing node, and computing cluster are provided, which adopt a universal protocol port that supports both communication protocol 1 and communication protocol 2, thereby achieving port compatibility, reducing the setup of network devices, and simplifying the network architecture.

Benefits of technology

By supporting ports that support multiple communication protocols, the network architecture of the computing cluster is simplified, network device resources are saved, and port utilization and data transmission efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950174A_ABST
    Figure CN120950174A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a processor, a computing node and a computing cluster, which are used for simplifying the network architecture of the computing cluster. The processor comprises a general protocol port used for supporting a first communication protocol and a second communication protocol, and the second communication protocol is different from the first communication protocol; wherein the first communication protocol is a communication protocol of a first network, the second communication protocol is a communication protocol of a second network, and the first network is used for supporting communication among a plurality of processors including the processor in a computing node. The second network is used for supporting communication between the processor and other processors included in other computing nodes, and the processor is any one of a graphics processing unit (GPU), a tensor processor (TPU), a neural processor (NPU) or a general purpose graphics processing unit (GPGPU).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more particularly to a processor, computing node, and computing cluster. Background Technology

[0002] With the application of large-scale artificial intelligence (AI) models, the demand for underlying computing power is increasing daily. Currently, the training parameters of large AI models have soared to the trillion level. Such a massive training task cannot be completed by a single graphics processing unit (GPU) or a single server. It requires a large number of servers as computing nodes, forming computing clusters through high-speed networks to complete the task together. Currently, these computing clusters are typically formed by scaling in two directions: vertical scaling (Scale-Up) and horizontal scaling (Scale-Out).

[0003] Scale-up refers to vertical scaling, which means expanding within a single compute node to enhance its computing power. This typically involves adding more GPUs, storage devices, and memory to that node. Conversely, scale-out refers to horizontal scaling, which means expanding from a single compute node by interconnecting more compute nodes to form a compute cluster, thereby increasing computing power.

[0004] Currently, networks formed by scale-up expansion and networks formed by scale-out expansion use different communication protocols. In order to support these two different communication protocols, different network devices are required, which makes the network architecture of computing clusters more complex. Summary of the Invention

[0005] This application provides a processor, computing node, and computing cluster to reduce the need for network devices and simplify the network architecture of the computing cluster.

[0006] Firstly, this application provides a processor including ports supporting both Communication Protocol 1 and Communication Protocol 2, where Communication Protocol 1 and Communication Protocol 2 are different. The processor can be a Graphics Processing Unit (GPU), Tensor Processing Unit (TPU), Neural Processing Unit (NPU), or a General Purpose Graphics Processing Unit (GPGPU), or other general-purpose processors such as a CPU; no limitation is made here. Taking a GPU as an example, Network 1 supports communication between multiple GPUs, including the GPU, within a computing node, using Communication Protocol 1. Network 2 supports communication between the GPU and other GPUs within other computing nodes, using Communication Protocol 2. Because the ports included in the GPU can simultaneously support both Communication Protocol 1 and Communication Protocol 2, Network 1 and Network 2 can be combined into one, reducing the number of network devices required and simplifying the network architecture of the computing cluster. Furthermore, compared to existing technologies, because the ports included in the GPU can simultaneously support both Communication Protocol 1 and Communication Protocol 2, the maximum number of available ports for a single network, such as Network 1 or Network 2, can be increased.

[0007] In one possible implementation, the processor may further include a single port supporting communication protocol 1; or the processor may further include a single port supporting communication protocol 2; or the processor may further include a single port supporting communication protocol 1 and a single port supporting communication protocol 2, wherein the number of ports supporting two communication protocols and the number of ports supporting only one communication protocol can be set according to actual needs.

[0008] In one possible implementation, network 1 is a network formed by vertical scaling on the basis of processors, which can be scale-up scaling. Network 2 is a network formed by horizontal scaling on the basis of computing nodes, which can be scale-out scaling. Scale-out scaling can form a larger computing cluster.

[0009] In one possible implementation, the general protocol port includes communication interface 1 and communication interface 2. Communication interface 1 is used to support communication protocol 1. If a data packet needs to be forwarded within network 1, the data packet is directly processed by communication interface 1. If a data packet needs to be forwarded within network 2, it is processed by communication interface 2. Thus, this technical solution can enable a port to support two different communication protocols simultaneously.

[0010] In one possible implementation, the processor further includes other general protocol ports, and the communication interface 2 is a shared communication interface between the other general protocol ports and the general protocol port, used to encapsulate the data packets based on the communication protocol 2 before forwarding them when the data packets need to be forwarded within the network 2.

[0011] In a second aspect, this application provides a computing node including a plurality of processors as described in the first aspect, and a switching chip for connecting the plurality of processors, wherein the switching chip includes a general protocol forwarding port for supporting communication protocol 1 and communication protocol 2.

[0012] In one possible implementation, the switching chip also includes a forwarding port for a single communication protocol 1.

[0013] In one possible implementation, the switching chip also includes a forwarding port for a single device supporting Communication Protocol 2.

[0014] In one possible implementation, the switching chip also includes a forwarding port for a single support of communication protocol 1 and a forwarding port for a single support of communication protocol 2.

[0015] Thirdly, this application also provides a computing cluster, including multiple computing nodes as described in the second aspect, and a primary switching device for connecting the multiple computing nodes. The primary switching device includes a general protocol switching port for supporting communication protocol 1 and communication protocol 2.

[0016] In one possible implementation, the primary switching device also includes a primary switching port for a single device supporting Communication Protocol 2.

[0017] In one possible implementation, the computing cluster also includes secondary switching devices for connecting multiple primary switching devices. The secondary switching devices include a secondary switching port for a single device supporting Communication Protocol 2. The cluster size can be further expanded through the secondary switching devices.

[0018] Fourthly, this application also provides a switching device, including: a general protocol port for supporting communication protocol 1 and communication protocol 2, wherein communication protocol 1 is different from communication protocol 2, communication protocol 1 is the communication protocol of network 1, and communication protocol 2 is the communication protocol of network 2, wherein network 1 is used to support communication between multiple processors included in a computing node, and network 2 is used to support communication between the computing node and other processors included in other computing nodes, and the switching device is used to connect the computing node and the other computing nodes. Attached Figure Description

[0019] Figure 1 This refers to the network architecture of computing clusters in existing technologies;

[0020] Figure 2 for Figure 1 A schematic diagram of the architecture of a computing node;

[0021] Figure 3 This is a schematic diagram of another architecture for a computing node in the prior art;

[0022] Figure 4 A schematic diagram of the ports included in a processor provided in this application;

[0023] Figure 5 A schematic diagram of the communication protocol port of the processor provided in this application;

[0024] Figure 6 A schematic diagram of a computing node architecture provided for this application;

[0025] Figure 7 A schematic diagram of the ports included in a switching chip provided in this application;

[0026] Figure 8 A schematic diagram of a computing cluster architecture provided for this application;

[0027] Figure 9 This is a schematic diagram of the ports included in a primary switching device provided in this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the specific embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0029] Currently, there are different networking methods in existing technologies to implement Scale-Up networks and Scale-Out networks, which are described below:

[0030] The first type: such as Figure 1 As shown, the system includes 32 compute nodes, each containing 8 GPUs (GPU0-GPU7), resulting in a total of 256 GPUs. These 256 GPUs are connected via an NVLink network to form a Super POD. Both scaling from one GPU to eight GPUs and scaling from one compute node to 32 compute nodes can be considered scale-up scaling. In scale-up scenarios, the number of compute nodes is typically limited. To further expand the interconnect range, connecting multiple Super PODs can form a larger-scale GPU cluster. These Super PODs can be connected via InfiniBand networks, Ethernet, or a Unified Bus (UB)-Global. Scaling from one Super POD to multiple Super PODs can be considered scale-out scaling.

[0031] Please see Figure 2This is an architecture diagram of one of the 32 compute nodes. This compute node's architecture includes 8 GPUs, namely GPU0-GPU7, each connected via NVLink and 4 NVSwitch interfaces. Figure 2 The above (abbreviated as NVS) enables interconnection within a compute node. For external nodes, it also connects to ConnectX-7 (Peripheral Component Interconnect Express, PCIe) ports. Figure 2 The above is an abbreviation for CX-7 network module. One CX-7 network module includes four CX-7 network cards, which are then connected to the IB network. Although the GPU's outbound bandwidth can be 900GB / s + 128GB / s, because the ports supporting different protocols are physically separate, flexible time-division multiplexing is not possible. That is, if you want to use the NVLink protocol with a bandwidth of 1000GB / s, even if there is a physical outbound bandwidth of more than 1000GB / s, it cannot be achieved.

[0032] Because Scale-Up and Scale-Out expansion use different network protocols, different network devices are needed to efficiently handle these two protocols. For example, an NVSwitch is needed for Scale-Up expansion, while a ConnectX-7 smart network card, IB / Eth switch, etc. are needed for Scale-Out expansion. This makes the network formed by Scale-Up expansion and the network formed by Scale-Out expansion two independent networks.

[0033] The second option: Please refer to [link / reference] Figure 3 , Figure 3 The server shown includes eight GPUs interconnected via a PCIe switch. Each GPU has 24 Ethernet interfaces with a bandwidth of 100G / s. 21 of these interfaces are used for interconnection between GPUs within the server, meaning each GPU has 2.1TB / s of bandwidth for interconnecting with the other seven GPUs. The remaining three Ethernet interfaces are used for scale-out expansion interconnection, where the scale-out interfaces can be Quad Small Form Factor Pluggable-Double Density (QSPF-DD). Furthermore, the server includes ten 100G Ethernet ports and sixteen PCIe 4.0 interfaces. The Ethernet ports are used for interconnecting processors within the server, while the PCIe 4.0 interfaces are used to connect to the host CPU. Figure 3In the architecture shown, the communication protocol used by the Scale-Up extension is the same as that used by the Scale-Out extension. However, when the same communication protocol is used for all of them, the cost of high-performance transmission will be relatively high due to the requirements of data volume, transmission distance and transmission reliability.

[0034] To address this, this application provides a processor including a general-purpose protocol port for supporting communication protocol 1 and communication protocol 2. Communication protocol 1 is the communication protocol for network 1, and communication protocol 2 is the communication protocol for network 2. Network 1 is a network formed by interconnecting multiple processors after scale-up expansion based on a single processor. If a processor scales up to form a computing node, then network 2 is a network formed by interconnecting multiple computing nodes after scale-out expansion based on a single computing node. Although network 1 and network 2 use different communication protocols, the general-purpose protocol port included in the processor can simultaneously support communication protocol 1 and communication protocol 2. This allows network 1 and network 2 to be combined into one, eliminating the need for separate network devices for network 1 and network 2. In other words, network 1 and network 2 can share network devices, thereby saving network device resources. Furthermore, in the prior art, the ports on processors supporting communication protocol 1 and communication protocol 2 are physically separate. The port configuration method of the processor in this application maximizes the utilization of each port. Furthermore, in this embodiment of the application, network 1 and network 2 each use suitable communication protocols, that is, they do not need to use the same communication protocol, thereby reducing protocol processing overhead.

[0035] Firstly, please see Figure 4 , Figure 4 The processor 40 shown includes a general protocol port 401, which is used to support communication protocol 1 and general protocol 2. Communication protocol 1 and communication protocol 2 are different. The number of general protocol ports 401 can be M1, where M1 is an integer greater than zero.

[0036] The processor 40 can be one of the following: a GPU, a general-purpose computing on graphics processing unit (GPGPU), a neural network processing unit (NPU), a tensor processing unit (TPU), or another dedicated processor (XPU); no limitation is made here. For ease of description, the following description will use GPU 40 as an example.

[0037] Network 1 is the network formed by scaling up with GPUs as units (40 GPUs) and connecting them. Network 2 is the network formed by scaling up with a single compute node as the unit (a compute node) and connecting them. It's important to note that scaling up doesn't strictly refer to expanding from one GPU to multiple GPUs. In practice, a single compute node can be further scaled up to form a supernode. In this case, the network formed by connecting multiple compute nodes can also be called Network 1. Furthermore, scaling up with a single supernode can create a large AI cluster. In this case, the network formed by connecting multiple supernodes is called Network 2.

[0038] In this embodiment, if the plane undergoing scale-up expansion is called the scale-up plane, then network 1 can be called a scale-up plane network. Of course, the plane undergoing scale-up expansion can also be called a bus plane, SuperPOD plane, high-bandwidth domain, or intra-machine interconnect. If the plane undergoing scale-out expansion is called the scale-out plane, then network 2 can be called a scale-out plane network. The networking method of the scale-up plane network can be any combination of torus topology, mesh topology, or Clos topology. A torus topology is a multi-dimensional ring structure. In a torus, each node is connected to its neighboring nodes, forming a ring structure. The torus topology allows for efficient communication and routing because each node only needs to know the location of its neighboring nodes to communicate. A mesh topology is a grid topology similar to a torus, but it is a planar structure. In a mesh, each node is connected to its adjacent nodes (up, down, left, and right), forming a planar mesh structure. Mesh topology enables efficient communication and routing, but it requires more links and switches to connect all nodes. Clos is a three-layer topology composed of multiple switches. In Clos, each switch has multiple input ports and multiple output ports, which can be connected to form a three-layer structure. Clos topology enables efficient communication and routing because it can perform load balancing and failover among multiple switches. The networking method for the scale-out plane can be Clos topology, optical cross-connect (OXC) or a combination of both, or other networking methods; no restrictions are placed here.

[0039] Having established the concepts of Network 1 and Network 2, Network 1 uses Communication Protocol 1, while Network 2 uses Communication Protocol 2. This means that GPUs within Network 1 communicate using Communication Protocol 1, while GPUs in different Networks 1 communicate using Communication Protocol 2. Communication Protocol 1 includes, but is not limited to, NVLink, Infinity, Compute Express Link (CXL), PCIe, Unified Bus (UB)-Clan, or other communication protocols. Communication Protocol 2 includes, but is not limited to, Ethernet, InfiniBand, or UB-G, or other communication protocols with good scalability and communication performance; specific limitations are not imposed in this embodiment.

[0040] The following section describes the specific implementation of the general protocol port 401 included in GPU40. Please refer to [link / reference]. Figure 5 The general protocol port 401 includes modules 4011 and 4012, where module 4011 supports communication protocol 1 and module 4012 supports communication protocol 2. In this implementation, the general protocol port 401 also includes a detection module 4013 and a switch module 4014. The detection module 4013 is communicatively connected to the switch module 4014, and the switch module 4014 selectively communicates with modules 4011 and 4012. The detection module 4013 detects whether a data packet needs to be forwarded in network 1 or network 2. In a specific implementation, assuming GPU 40-1 needs to send a message to GPU 40-2, GPU 40-1 knows the destination network address of GPU 40-2, and therefore knows the destination port number used by GPU 40-2 to receive messages. Thus, GPU 40-1 can determine whether the message to be sent needs to be forwarded in network 1 or network 2 based on the destination port number. In some possible implementations, GPU40 can also determine whether the data to be sent needs to be forwarded in network 1 or network 2 through other methods. Specifically, an identifier bit is set in the data packet to be sent to indicate whether the data packet needs to be forwarded in network 1 or network 2, thereby determining the processing method for the data packet. This method can quickly determine the forwarding network of the data packet, thereby improving the packet forwarding efficiency. When the data packet needs to be forwarded in network 1, the control switch module 4014 communicates with module 4011 to encapsulate the data packet according to communication protocol 1 before forwarding. When the data packet needs to be forwarded in network 2, the control switch module 4014 communicates with module 4012 to encapsulate the data packet according to communication protocol 2 before forwarding.

[0041] In this embodiment, when there are multiple general protocol ports 401, module 4012 can be dedicated to one general protocol port 401, meaning one module 4012 is configured for one general protocol port 401. In some possible implementations, module 4012 can also be shared by multiple general protocol ports 401 included in the GPU. However, when module 4012 is shared by multiple general protocol ports 401 and the number of packets to be forwarded in network 2 is large, module 4012 should also include a data buffer for caching data packets to be processed. This data buffer can be memory for temporary data storage. Memory includes at least two types of memory, such as random access memory (RAM) or dynamic random access memory (DRAM). Memory can also include other types of random access memory, such as static random access memory (SRAM). In addition, when the number of packets to be forwarded is large, a queuing strategy can be set, such as first-in-first-out (FIFO), meaning that the first data packet to arrive is processed first, and the data packets that arrive later are processed later.

[0042] Besides supporting Communication Protocol 1 and Communication Protocol 2 through the hardware methods described above, the General Protocol Port 401 can also be implemented in software. Specifically, Software-Defined Networking (SDN) technology can be used to receive data packets through a hardware interface. This means the SDN controller processes the data packets according to flow tables to achieve protocol conversion.

[0043] After introducing the general protocol port 401, in some possible implementations, GPU 40 also includes: a first protocol port 402 for supporting communication protocol 1, wherein the number of first protocol ports 402 is M2, and M2 is a positive integer; in some possible implementations, GPU 40 also includes: a second protocol port 403 for supporting communication protocol 2, wherein the number of second protocol ports 403 is M3, and M3 is a positive integer; in some possible implementations, GPU 40 also includes: a first protocol port 402 and a second protocol port 403, wherein the first protocol port 402 is used to support communication protocol 1, and the second protocol port 403 is used to support communication protocol 2, wherein the number of first protocol ports 402 is M2, and the number of second protocol ports 403 is M3, and both M2 and M3 are positive integers. In other words, in this embodiment of the application, in addition to the general protocol port 401, other protocol ports that only support one communication protocol can also be set, thereby expanding the application scenarios of GPU 40. The number of general protocol ports 401 and ports that only support one communication protocol can be set according to actual needs, and no specific limitation is made here.

[0044] Secondly, please see Figure 6 This application also provides a computing node 60. Figure 6 The computing node 60 shown includes multiple such nodes. Figure 4 The GPU 40 and switching chip 601 shown are illustrated. Switching chip 601 includes a general protocol forwarding port 6011 for supporting communication protocol 1 and communication protocol 2 as described in the first aspect. The specific implementation of the general protocol forwarding port 6011 can be found in the description of general protocol port 401 in the first aspect, and will not be repeated here. The number of general protocol forwarding ports 6011 is N1, where N1 is a positive integer.

[0045] In the embodiments of this application, a such Figure 4 The GPU 40 shown can be vertically expanded to form a computing node 60, thereby enhancing the computing power of a computing node 60. Multiple GPUs 40 can communicate with each other based on the aforementioned communication protocol 1.

[0046] In some possible implementations, the switching chip 601 further includes a first protocol forwarding port 6012 for supporting communication protocol 1, wherein the number of first protocol forwarding ports 6012 is N², where N² is a positive integer. In some possible implementations, the switching chip 601 further includes a second protocol forwarding port 6013 for supporting communication protocol 2, wherein the number of second protocol forwarding ports 6013 is N³, where N³ is a positive integer. In some possible implementations, the switching chip 601 further includes a first protocol forwarding port 6012 and a second protocol forwarding port 6013, wherein the first protocol forwarding port 6012 supports communication protocol 1 and the second protocol forwarding port 6013 supports communication protocol 2. A schematic diagram of the ports included in the switching chip 601 can be found [reference needed]. Figure 7 .

[0047] Furthermore, in this embodiment, the computing node 60 further includes a CPU 602. The CPU 602 can be communicatively connected to the switching chip 601 or directly connected to the GPU 40. When there are multiple CPUs 602, some CPUs 601 among the multiple CPUs 602 can be communicatively connected to the switching chip 601, and some CPUs 602 among the multiple CPUs 602 can be communicatively connected to the switching chip 601. All three implementation methods are acceptable. The CPU 602 can communicate with the switching chip 601 or the GPU 40 through the PCIe bus or other buses. This embodiment does not impose any restrictions.

[0048] Thirdly, please see Figure 8 This application also provides a computing cluster 80, which includes multiple computing nodes and a primary switching device 801. The primary switching device 801 is used to interconnect the multiple computing nodes. The primary switching device 801 includes general protocol switching ports 8011 for supporting communication protocol 1 and communication protocol 2 as described in the first aspect. The number of general protocol switching ports 8011 is P1, where P1 is a positive integer. The computing nodes included in the computing cluster 80 can be... Figure 6 The computation node 60 shown can also be... Figure 6 The supernode formed by scaling up the compute node 60 shown is not specifically limited in this embodiment. In the following description, the compute nodes included in the compute cluster 80 are... Figure 6 Take computing node 60 as an example.

[0049] In this embodiment, the GPU 40 included in the computing node 60 exchanges data through the switching chip 601. GPUs 40 in different computing nodes 60 can exchange data through a primary switching device 801, without requiring... Figure 2As shown, a CX-7 network card is also required, which can reduce the number of network devices required and thus save network device resources.

[0050] Furthermore, in this embodiment, the primary switching device 801 also includes a second protocol primary switching port 8012 for supporting communication protocol 2. The number of second protocol primary switching ports 8012 is P2, where P2 is a positive integer. The port configuration of the primary switching device 801 can be found in [reference needed]. Figure 9 The primary switching equipment 801 may also include other components, such as linear-drive pluggable optical modules (LPOs). "Pluggable" means that the switch has ports for optical modules; by inserting the corresponding optical module into the switch, an optical fiber can be connected. If the optical fiber or optical module is damaged, it can be repaired or replaced simply by removing the optical module. The 400G optical module encapsulation interface of LPO can be divided into multimode optical module interfaces, including SR8, where S stands for Short, indicating a transmission distance of 100 meters, and "8" represents 8 optical signal channels, each channel being 50G; SR4.2, where SR also means short distance, 100-meter transmission distance, "4" represents 4 fiber channels, and "2" represents 2 wavelengths bidirectionally multiplexed per channel, with each channel being 2×50G; the 400G optical module encapsulation interface also includes single-mode optical module interfaces, including FR8 and FR4, where FR indicates a distance of 2 kilometers (km), "8" refers to 8 wavelengths multiplexed on one fiber, and "4" refers to 4 wavelengths multiplexed on one fiber. If the scale-out plane adopts the above networking method, the primary switching device 801 uses LPO optical modules, and the bandwidth of the LPO optical modules is 400G. Typically, the primary switching device 801 includes 128 interfaces, thus providing a bandwidth of 50T.

[0051] Furthermore, in this embodiment of the application, the computing cluster 80 also includes a secondary switching device 802, which is used to interconnect the primary switching device 801. The secondary switching device 802 includes a second protocol secondary switching port to support communication protocol 2, which can support further expansion of the computing cluster 80.

[0052] In this embodiment, the number of computing nodes 60 and the number of GPUs 40 included in a computing cluster 80 can be set according to actual needs. Typically, each computing node 60 includes the same number of GPUs 40; a single computing node 60 can include 8, 16, 32, 64, 128, or 256 GPUs 40. In actual design, the number of GPUs 40 included in a computing node 60 can be set according to the actual cooling solution used. For example, if liquid cooling is used, the number of GPUs 40 included in a computing node 60 can be 64; if air cooling is used, the number of GPUs 40 included in a computing node 60 can be 16. Liquid cooling refers to using liquid as the cooling medium to dissipate heat from the computing devices of the computing nodes, while air cooling refers to using air as the cooling medium. The number of computing nodes 60 can be 4, 8, 16, 64, or 128. Thus, when the cooling solution is air cooling, a single compute node 60 can include 16 GPUs 40, and 120 compute nodes can form a compute cluster of 1920 GPUs. When the cooling solution is liquid cooling, a single compute node 60 can include 64 GPUs 40, and 128 compute nodes can form a compute cluster of 8912 GPUs. Accordingly, the number of primary switching devices 801 and the number of secondary switching devices 802 can be calculated based on the number of compute nodes 60, the number of GPUs 40 included in the compute nodes 60, and the required bandwidth.

[0053] After introducing the GPU40, compute nodes60, switching chip601, primary switching device801, and secondary switching device802 included in the computing cluster 80, we will further introduce the port settings of GPU40, switching chip601, primary switching device801, and secondary switching device802 when building the computing cluster 80. These will be introduced separately below.

[0054] Method 1

[0055] The GPU 40 includes M1 general-purpose protocol ports 401, all of which support communication protocol 1 and communication protocol 2, meaning they both support data forwarding from network 1 and network 2. Thus, the number of available ports for data forwarding from network 1 is M1, and the number of available ports for data forwarding from network 2 is also M1. Compared to existing technologies, where the total number of ports is M1, the number of available ports for data forwarding from network 1 and network 2 is less than M1. Therefore, the technical solution provided in this application maximizes the number of available ports for data forwarding from both network 1 and network 2.

[0056] The switching chip 600 includes N1 general protocol forwarding ports, all of which support communication protocol 1 and communication protocol 2, meaning they all support receiving and forwarding data from network 1 as well as receiving and forwarding data from network 2.

[0057] It should be noted here that the communication protocol 1 supported by GPU40 must be the same as the communication protocol 1 supported by switching chip 601 and primary switching device 801. For example, the communication protocol 1 supported by GPU40 is the NVLink protocol, and the communication protocol 1 supported by switching chip 601 and primary switching device 801 is also the NVLink protocol. Similarly, the communication protocol 2 supported by GPU40 must be the same as the communication protocol 2 supported by switching chip 601 and primary switching device 801. For example, the communication protocol 2 supported by GPU40 is the Ethernet protocol, and the communication protocol 2 supported by switching chip 601 and primary switching device 801 is also the Ethernet protocol. This is to ensure the normal transmission and reception of the switching device.

[0058] Method 2

[0059] The GPU40 includes M1 general-purpose protocol ports 401, all of which support communication protocol 1 and communication protocol 2; the switching chip includes N1 general-purpose protocol forwarding ports 6011, all of which support communication protocol 1 and communication protocol 2; the primary switching device 801 includes a general-purpose protocol switching port 8011 and a secondary protocol primary switching port 8012, wherein the general-purpose protocol switching port 8011 is used to support communication protocol 1 and communication protocol 2, and the secondary protocol primary switching port 8012 is used to support communication protocol 2.

[0060] Based on the above methods one and two, the computing cluster 80 may also include secondary switching devices 802. When the computing cluster 80 also includes secondary switching devices 802, the number of secondary switching devices 802 and the number of primary switches 801 meet a certain convergence ratio. For example, the convergence ratio of secondary switching devices 802 to primary switching devices 801 is 7:1.

[0061] Method 3

[0062] GPU40 includes M1 general protocol ports 401 for supporting communication protocol 1 and communication protocol 2, and M2 primary protocol ports 402 for supporting communication protocol 1. The sum of M1 and M2 is M. Thus, network 1 has M available ports for data forwarding, and network 2 also has M1 available ports for data forwarding. Compared to existing technologies, this maximizes the number of data forwarding ports for network 1.

[0063] The switching chip 601 includes N1 general protocol forwarding ports 6011 for supporting communication protocol 1 and communication protocol 2; or it includes N1 general protocol forwarding ports 6011 and N2 first protocol forwarding ports 6012, wherein the N1 general protocol forwarding ports 6011 can all support communication protocol 1 and communication protocol 2, and the N2 first protocol forwarding ports 6012 are used to support communication protocol 1.

[0064] For a primary switching device 801, it includes P1 general protocol switching ports 8011 for supporting communication protocol 1 or general protocol 2; or it includes P1 general protocol switching ports 8011 and P2 secondary protocol primary switching ports 8012, wherein the P1 general protocol switching ports can support both communication protocol 1 and communication protocol 2, and the P2 secondary protocol primary switching ports 8012 are used to support the second communication protocol.

[0065] When the primary switching device 801 includes P1 general protocol ports 8011 and P2 secondary protocol primary switching ports 8012, the ports in the primary switching device 801 used to connect to the GPU 40 and to support communication protocol 1 and communication protocol 2 also need to be ports that can support both communication protocols.

[0066] Method 4

[0067] GPU40 includes M1 general protocol ports 401 for supporting communication protocol 1 and communication protocol 2, and M3 secondary protocol ports 403 for supporting communication protocol 2. Thus, network 1 has M1 available ports for data forwarding, and network 2 has M1 + M3 available ports for data forwarding. Compared to existing technologies, this maximizes the number of available ports for data forwarding in network 2.

[0068] The switching chip 600 includes N1 general protocol forwarding ports 6011 for supporting communication protocol 1 and communication protocol 2; or it includes N1 general protocol forwarding ports 6011 and N2 first protocol forwarding ports 6012, wherein the N1 general protocol forwarding ports can all support communication protocol 1 and communication protocol 2, and the N2 protocol forwarding ports 6011 are used to support communication protocol 1.

[0069] The primary switching device 801 includes P1 general protocol switching ports 8011 for supporting communication protocol 1 or general protocol 2; or it includes P1 general protocol switching ports 8011 and P2 primary switching ports 8012 for supporting the second protocol. The P1 general protocol switching ports 8011 can support both communication protocol 1 and communication protocol 2, and the P2 primary switching ports 8012 are used to support the second communication protocol.

[0070] When the primary switching device 801 includes P1 general protocol ports 8011 and P2 protocol primary switching ports, the port in the primary switching device 801 used to connect to the GPU 40 and supporting communication protocol 1 and communication protocol 2 also needs to be able to support both communication protocols. The port in the primary switching device 801 used to connect to the GPU 40 and supporting communication protocol 2 can be a port that supports communication protocol 2, or it can be a port that supports both communication protocol 1 and communication protocol 2 simultaneously.

[0071] Method 5

[0072] GPU40 includes M1 general-purpose protocol ports 401 for supporting communication protocol 1 and communication protocol 2, M2 first protocol ports 402 for supporting communication protocol 1, and M3 second protocol ports 403 for supporting communication protocol 2. Thus, the number of available data forwarding ports for network 1 is M1 + M2, and the number of available data forwarding ports for network 2 is also M1 + M3. Compared to existing technologies, this maximizes the number of available data forwarding ports for both network 1 and network 2.

[0073] In this implementation, the switching chip 601 includes N1 general-purpose protocol forwarding ports 6011 for supporting communication protocol 1 and communication protocol 2, N2 first protocol forwarding ports 6012 for supporting communication protocol 1, and N3 second protocol forwarding ports 6013 for supporting communication protocol 2. In this implementation, the ports of the primary switching device 801 used to connect to the GPU 40 for supporting communication protocol 1 and communication protocol 2 also need to be ports capable of supporting both communication protocols; the ports of the primary switching device 801 used to connect to the GPU 40 for supporting communication protocol 1 can be ports capable of supporting communication protocol 1, or ports capable of supporting both communication protocol 1 and communication protocol 2 simultaneously; the ports of the primary switching device 801 used to connect to the GPU 40 for supporting communication protocol 2 can be ports capable of supporting communication protocol 2, or ports capable of supporting both communication protocol 1 and communication protocol 2 simultaneously.

[0074] In the specific implementation, the number of ports on GPU40, primary switching device 801, and secondary switching device 802 that support only communication protocol 1, only communication protocol 2, and both communication protocols 1 and 2 can be configured according to actual needs. In practical applications, appropriate congestion control algorithms can be allocated based on the data traffic of network 1 and network 2 to dynamically utilize these ports and improve link utilization.

[0075] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A processor, characterized in that, include: A general protocol port is used to support a first communication protocol and a second communication protocol, wherein the second communication protocol is different from the first communication protocol; Wherein, the first communication protocol is the communication protocol of the first network, the second communication protocol is the communication protocol of the second network, the first network is used to support communication between multiple processors, including the processor, within a computing node, and the second network is used to support communication between the processor and other processors in other computing nodes, wherein the processor is any one of a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU), or a general-purpose graphics processing unit (GPGPU).

2. The processor according to claim 1, characterized in that, The processor also includes: The first protocol port is used to support the first communication protocol.

3. The processor according to claim 1 or 2, characterized in that, The processor also includes: The second protocol port is used to support the second communication protocol.

4. The processor according to claim 1 or 2, characterized in that, The first network is a network formed by vertically expanding the network based on the processor, and the second network is a network formed by horizontally expanding the network based on the computing nodes.

5. The processor according to claim 1 or 2, characterized in that, The first communication protocol is any one of NVlink protocol, Infinity protocol, CXL protocol, PCIe protocol for peripheral devices, and UB-C protocol; the second communication protocol is any one of Ethernet protocol, InfiniBand protocol, or UB-G protocol.

6. The processor according to claim 1 or 2, characterized in that, The general protocol port includes a first communication interface and a second communication interface; The first communication interface is used to support the first communication protocol; The second communication interface is used to support the second communication protocol.

7. A computing node, characterized in that, include: Multiple processors as described in claim 1; A switching chip for connecting the plurality of processors; The switching chip includes a general protocol forwarding port, which is used to support the first communication protocol and the second communication protocol.

8. The computing node according to claim 7, characterized in that, The switching chip also includes: The first protocol forwarding port is used to support the first communication protocol.

9. The computing node according to claim 7 or 8, characterized in that, The switching chip also includes: The second protocol forwarding port is used to support the second communication protocol.

10. A computing cluster, characterized in that, include: Multiple computing nodes as described in claim 9; A primary switching device is used to connect the plurality of computing nodes, wherein the primary switching device includes a general protocol switching port, the general protocol switching port being used to support the first communication protocol and the second communication protocol.

11. The computing cluster according to claim 10, characterized in that, The primary switching equipment also includes: The second protocol level 1 switch port is used to support the second communication protocol.

12. The computing cluster according to claim 10 or 11, characterized in that, The computing cluster also includes: A secondary switching device is used to connect multiple primary switching devices, wherein the secondary switching device includes: The second protocol level 2 switch port is used to support the second communication protocol.

13. A switching device, characterized in that, include: A general protocol switching port is used to support a first communication protocol and a second communication protocol, wherein the second communication protocol is different from the first communication protocol; Wherein, the first communication protocol is the communication protocol of the first network, the second communication protocol is the communication protocol of the second network, the first network supports communication between multiple processors included in a computing node, the second network is used to support communication between the processors included in the computing node and other processors included in other computing nodes, and the switching device is used to connect the computing node and the other computing nodes.

14. The switching device according to claim 13, characterized in that, The switching equipment also includes: The second protocol switching port is used to support the second communication protocol.