Computing system and communication method
Patent Information
- Application Number
- PCT/CN2025/139791
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-14
- Filing Date
- 2025-12-03
- Publication Date
- 2026-09-17
Smart Images

Figure CN2025139791_17092026_PF_FP_ABST
Abstract
Description
Computing systems and communication methods
[0001] This application claims priority to Chinese Patent Application No. 202510307688.6, filed on March 14, 2025, entitled "Computing System and Communication Method", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of communication technology, and in particular to a computing system and a communication method. Background Technology
[0003] To enhance the computing power of a computing system, it typically comprises multiple computing nodes that can execute computational tasks in parallel. These nodes communicate with each other via interconnect links, forming a fully connected network (full mesh) structure.
[0004] Although direct communication links between nodes have lower latency, the number of network ports on each node in a fully connected network architecture is linearly related to the scale of the computing system (e.g., the number of computing nodes) because any two nodes in the fully connected network architecture have direct communication links. For example, if the computing system includes n computing nodes, where n is an integer greater than 1, then to build a fully connected network, each computing node needs to have n-1 network ports.
[0005] However, if the number of network ports on a computing node is less than the number of computing nodes in the computing system, the computing nodes in the computing system cannot build a fully connected network. Summary of the Invention
[0006] The purpose of this application is to provide a computing system and a communication method.
[0007] The first aspect of this application provides a computing system comprising: M computing nodes and X switching nodes, wherein the number of ports of each switching node is greater than 0 and less than M, where M is an integer greater than 2 and X is an integer greater than 1; any two computing nodes among the M computing nodes are connected to the same switching node through a direct communication link, wherein a first computing node among the M computing nodes is connected to a first switching node among the X switching nodes through a first direct communication link, and a second computing node among the M computing nodes is connected to the first switching node through a second direct communication link; and the first computing node and the second computing node communicate with each other through the first direct communication link, the first switching node, and the second direct communication link.
[0008] In this embodiment, computing nodes in the computing system establish communication connections through multiple low-port switching nodes, and any two computing nodes can communicate through at most one switching node. Since the internal structure of low-port switching nodes is relatively simple (e.g., consisting of only a crossbar switch matrix), the communication latency is also low. Establishing communication connections between multiple computing nodes in the computing system through multiple low-port switching nodes can reduce the communication latency between any two computing nodes.
[0009] In this embodiment of the application, any two computing nodes among the M computing nodes are connected to the same switching node through one or more direct communication links.
[0010] In this embodiment of the application, any two computing nodes among the M computing nodes are connected to one or more identical switching nodes.
[0011] In this embodiment of the application, the first computing node sends the first data to the first switching node through the first direct communication link, and the first switching node sends the received first data to the second computing node through the second direct communication link, so as to realize the communication between the first computing node and the second computing node.
[0012] In one possible implementation of the first aspect described above, K of the M computing nodes are connected pairwise via direct communication links, where K is an integer greater than 1 and less than or equal to M; the third computing node among the K computing nodes is connected to the second switching node among the X switching nodes via a third direct communication link, the fourth computing node among the K computing nodes is connected to the second switching node via a fourth direct communication link, and the third computing node and the fourth computing node are connected via a fifth direct communication link; furthermore, the third computing node and the fourth computing node communicate via the third direct communication link, the second switching node and the fourth direct communication link, and / or, the third computing node and the fourth computing node communicate via the fifth direct communication link.
[0013] In this embodiment of the application, K of the M computing nodes are connected to each other through one or more direct communication links.
[0014] In this embodiment, the third computing node sends the second data to the second switching node via a third direct communication link, and the second switching node sends the received second data to the fourth computing node via a fourth direct communication link, thereby enabling communication between the third computing node and the fourth computing node; alternatively, the third computing node sends the second data to the fourth computing node via a fifth direct communication link, thereby enabling communication between the third computing node and the fourth computing node; alternatively, the third computing node sends a portion of the second data (e.g., the first sub-data) to the second switching node via the third direct communication link, and the second switching node sends a portion of the received second data to the fourth computing node via the fourth direct communication link; the third computing node then sends another portion of the second data (e.g., the second sub-data) to the fourth computing node via the fifth direct communication link, thereby enabling communication between the third computing node and the fourth computing node.
[0015] In one possible implementation of the first aspect described above, the computing nodes in the first and second computing node groups of the N computing node groups are connected to at least two switching nodes via direct communication links, and the X switching nodes include at least two switching nodes.
[0016] In one possible implementation of the first aspect described above, the M computing nodes correspond to N computing node groups, each computing node group includes at least one computing node from the M computing nodes, wherein computing nodes from any two computing node groups in the N computing node groups are connected to the same switching node.
[0017] In this embodiment of the application, computing nodes from any two computing node groups in the N computing node groups are connected to the same switching node through one or more direct communication links. Computing nodes from any two computing node groups in the N computing node groups are connected to one or more identical switching nodes.
[0018] In this embodiment, the first computing node and the second computing node are computing nodes in the same computing node group, or the first computing node and the second computing node are computing nodes in different computing node groups; and the first computing node and the second computing node communicate with each other through a first direct communication link, a first switching node, and a second direct communication link. For example, the first computing node sends first data to the first switching node through the first direct communication link, and the first switching node sends the received first data to the second computing node through the second direct communication link, thereby realizing communication between the first computing node and the second computing node.
[0019] In one possible implementation of the first aspect described above, computing nodes in each of the Y computing node groups are connected to each other via direct communication links, where Y is an integer greater than 0 and less than or equal to N; and / or, in the N computing node groups, computing nodes with the same identification information are connected to each other via direct communication links.
[0020] In the embodiments of this application, the identification information includes at least one of the following: number, position. For example, computing nodes with the same number in N computing node groups (e.g., numbered 1-D1, 2-D1, 3-D1, and 4-D1) have the same identification information. As another example, computing nodes with the same position in N computing node groups (e.g., the first one) have the same identification information.
[0021] In the embodiments of this application, computing nodes in each of the Y computing node groups are connected to each other through one or more direct communication links; and / or, in the N computing node groups, computing nodes with the same identification information are connected to each other through one or more direct communication links.
[0022] In this embodiment of the application, each computing node group includes M / N computing nodes, wherein the i-th computing node in each computing node group is connected to each other through a direct communication link, and i is any integer from 1 to M / N.
[0023] In one possible implementation of the first aspect described above, the fifth computing node among the M computing nodes is connected to the third switching node among the X switching nodes via a sixth direct communication link, the sixth computing node among the M computing nodes is connected to the third switching node via a seventh direct communication link, and the fifth computing node and the sixth computing node are connected via an eighth direct communication link; wherein, the fifth computing node and the sixth computing node are computing nodes within one of the Y computing node groups, or, the fifth computing node and the sixth computing node are computing nodes with the same identification information in N computing node groups; and, the fifth computing node and the sixth computing node communicate via the sixth direct communication link, the third switching node and the seventh direct communication link, and / or, the fifth computing node and the sixth computing node communicate via the eighth direct communication link.
[0024] In this embodiment, the fifth computing node sends the third data to the third switching node via the sixth direct communication link, and the third switching node sends the received third data to the sixth computing node via the seventh direct communication link, thereby realizing communication between the fifth computing node and the sixth computing node; or, the fifth computing node sends the third data to the sixth computing node via the eighth direct communication link, thereby realizing communication between the fifth computing node and the sixth computing node; or, the fifth computing node sends a portion of the third data (e.g., the first sub-data) to the third switching node via the sixth direct communication link, and the third switching node sends a portion of the received third data to the sixth computing node via the seventh direct communication link; the fifth computing node then sends another portion of the third data (e.g., the second sub-data) to the sixth computing node via the eighth direct communication link, thereby realizing communication between the fifth computing node and the sixth computing node.
[0025] In one possible implementation of the first aspect described above, in N computing node groups, the computing nodes within each computing node group are connected to the same one or more control nodes.
[0026] A second aspect of this application provides a communication method applied to a computing system, the computing system comprising: M computing nodes and X switching nodes, each switching node having a number of ports greater than 0 and less than M, where M is an integer greater than 2 and X is an integer greater than 1; any two computing nodes among the M computing nodes are connected to the same switching node via a direct communication link, wherein a first computing node among the M computing nodes is connected to a first switching node among the X switching nodes via a first direct communication link, and a second computing node among the M computing nodes is connected to the first switching node via a second direct communication link; and the method comprises: a first computing node detecting a first request, the first request instructing the first computing node to send first data to a second computing node; the first computing node sending the first data to the second computing node via the first direct communication link, the first switching node, and the second direct communication link.
[0027] In this embodiment of the application, the first computing node sends the first data to the first switching node through the first direct communication link, and the first switching node sends the first data to the second computing node through the second direct communication link.
[0028] It is understood that data transmission between the first computing node and the second computing node can be forwarded once through the first switching node connected to the first computing node and the second computing node. That is, data transmission between the first computing node and the second computing node only needs to be forwarded once through the switching node connected to the data transmission between the first computing node and the second computing node, resulting in low communication latency.
[0029] In one possible implementation of the second aspect described above, K of the M computing nodes are connected pairwise via direct communication links, where K is an integer greater than 1 and less than or equal to M. A third computing node among the K nodes is connected to a second switching node among the X switching nodes via a third direct communication link, and a fourth computing node among the K nodes is connected to the second switching node via a fourth direct communication link. The third and fourth computing nodes are connected via a fifth direct communication link. Furthermore, the method further includes: the third computing node detecting a second request, the second request instructing the third computing node to send second data to the fourth computing node; the third computing node sending the second data to the fourth computing node via the third direct communication link, the second switching node, and the fourth direct communication link; or, the third computing node sending the second data to the fourth computing node via the fifth direct communication link; or, the third computing node sending a portion of the second data to the fourth computing node via the third direct communication link, the second switching node, and the fourth direct communication link, and sending another portion of the second data to the fourth computing node via the fifth direct communication link.
[0030] In this embodiment, the third computing node sends the second data to the second switching node via a third direct communication link, and the second switching node sends the second data to the second computing node via a fourth direct communication link. Alternatively, the third computing node sends the second data to the fourth computing node via a fifth direct communication link. Alternatively, the third computing node splits the second data and sends a portion of the second data (e.g., the first sub-data) to the second switching node via the third direct communication link; the second switching node sends a portion of the second data to the fourth computing node via the fourth direct communication link; and sends the remaining portion of the second data (e.g., the second sub-data) to the fourth computing node via the fifth direct communication link.
[0031] It is understandable that data transmission between the third and fourth computing nodes can be forwarded once through a second switching node connected to both the third and fourth computing nodes, and / or directly transmitted through the direct communication link between the third and fourth computing nodes. The direct communication link between computing nodes has low latency; therefore, data transmission between the third and fourth computing nodes requires at most one forwarding through a switching node connected to the data transmission link between them, resulting in low communication latency.
[0032] In one possible implementation of the second aspect described above, the M computing nodes correspond to N computing node groups, each computing node group including at least one computing node from the M computing nodes, wherein computing nodes in any two computing node groups in the N computing node groups are connected to the same switching node; in the N computing node groups, computing nodes in each of the Y computing node groups are connected pairwise via direct communication links, where Y is an integer greater than 0 and less than or equal to N; and / or, in the N computing node groups, computing nodes with the same identification information are connected pairwise via direct communication links, the fifth computing node in the M computing nodes is connected to the third switching node in the X switching nodes via a sixth direct communication link, the sixth computing node in the M computing nodes is connected to the third switching node via a seventh direct communication link, and the fifth computing node and the sixth computing node are connected via an eighth direct communication link; furthermore, the method also includes The fifth computing node out of M computing nodes is connected to the third switching node out of X switching nodes via a sixth direct communication link; the sixth computing node out of M computing nodes is connected to the third switching node via a seventh direct communication link; and the fifth computing node and the sixth computing node are connected via an eighth direct communication link. The fifth computing node detects a third request, which instructs the fifth computing node to send third data to the sixth computing node. The fifth computing node sends the third data to the sixth computing node via the sixth direct communication link, the third switching node, and the seventh direct communication link; or, the fifth computing node sends the third data to the sixth computing node via the eighth direct communication link; or, the fifth computing node sends a portion of the third data to the sixth computing node via the sixth direct communication link, the third switching node, and the seventh direct communication link, and sends another portion of the third data to the sixth computing node via the eighth direct communication link.
[0033] In this embodiment, the fifth computing node sends the third data to the third switching node via the sixth direct communication link, and the third switching node sends the third data to the sixth computing node via the seventh direct communication link. Alternatively, the fifth computing node sends the third data to the sixth computing node via the eighth direct communication link. Alternatively, the fifth computing node splits the third data and sends a portion of the third data (e.g., the first sub-data) to the third switching node via the sixth direct communication link; the third switching node sends a portion of the third data to the sixth computing node via the seventh direct communication link; and sends another portion of the third data (e.g., the second sub-data) to the sixth computing node via the eighth direct communication link.
[0034] It is understandable that data transmission between the fifth and sixth computing nodes can be forwarded once through a third switching node connected to both nodes, and / or transmitted directly through the direct communication link between them. The direct communication link between computing nodes has low latency; therefore, data transmission between the fifth and sixth computing nodes requires at most one forwarding through a switching node connected to their data transmission link, resulting in low communication latency.
[0035] In one possible implementation of the second aspect described above, the fifth computing node and the sixth computing node are computing nodes within one of the Y computing node groups; or, the fifth computing node and the sixth computing node are computing nodes with the same identification information in N computing node groups. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0037] Figure 1 shows a schematic diagram of a computing system according to an embodiment of this application;
[0038] Figure 2 shows a schematic diagram of a single-layer non-blocking network structure according to an embodiment of this application;
[0039] Figure 3 shows a schematic diagram of another layer-1 non-blocking network structure according to an embodiment of this application;
[0040] Figure 4 shows a schematic diagram of a cross switch matrix architecture according to an embodiment of this application;
[0041] Figure 5 illustrates a schematic diagram of a switching node architecture based on a combination of multiple cross switch matrix modules according to an embodiment of this application;
[0042] Figure 6 shows a schematic diagram of a two-layer fat tree network structure according to an embodiment of this application;
[0043] Figure 7 shows a schematic diagram of a non-blocking network structure according to an embodiment of this application;
[0044] Figure 8 shows a schematic diagram of another non-blocking network structure according to an embodiment of this application;
[0045] Figure 9 illustrates a schematic diagram of a network structure among computing nodes within a computing node group according to an embodiment of this application;
[0046] Figure 10 shows a schematic diagram of a network structure between computing nodes in a computing node group according to an embodiment of this application;
[0047] Figure 11 shows a schematic diagram of a network structure between a control node and a computing node according to an embodiment of this application;
[0048] Figure 12 shows a schematic diagram of a network structure of a computing system according to an embodiment of this application;
[0049] Figure 13 illustrates a flowchart of a communication method according to an embodiment of this application;
[0050] Figure 14 illustrates a flowchart of another communication method according to an embodiment of this application;
[0051] Figure 15 shows a logic diagram of a communication method according to an embodiment of this application;
[0052] Figure 16 illustrates a schematic diagram of the software framework of a computing system according to an embodiment of this application;
[0053] Figure 17 shows a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0054] The illustrative embodiments of this application include, but are not limited to, computing systems and communication methods.
[0055] Before introducing the technical solutions involved in the embodiments of this application, some of the terms included in the embodiments of this application will be explained.
[0056] (1) Node
[0057] A node is a basic unit in a computing system, used to perform specific functions and tasks. For example, a node that performs computational tasks in a computing system can be called a computing node; a node that performs management tasks such as task scheduling and resource allocation in a computing system can be called a control node (or management node); and a node that performs communication tasks between computing nodes and between computing nodes and control nodes can be called a switching node.
[0058] It should be noted that a node is a logical unit in a computing system. A node can be one or more electronic devices at the physical layer, or it can be a hardware circuit in one or more electronic devices used to perform a specific function.
[0059] (2) Calculate nodes
[0060] A compute node is a node in a computing system used to perform computational tasks.
[0061] A computing node typically includes one or more processors for performing computational operations. Processors include, but are not limited to, graphics processing units (GPUs), neural network processing units (NPUs), data processing units (DPUs), tensor processing units (TPUs), field-programmable gate arrays (FPGAs), or dedicated artificial intelligence (AI) processors.
[0062] (3) Control Node
[0063] A control node, also known as a management node, is a node in a computing system used to perform management tasks such as task scheduling and resource allocation. For example, control nodes are used to schedule computing tasks to the corresponding computing nodes for execution, and to control data interaction between computing nodes.
[0064] A control node typically includes one or more processors, including but not limited to a central processing unit (CPU), a digital processing unit (DPU), an FPGA, or a microprocessor.
[0065] (4) Exchange Node
[0066] A switching node is a node used for forwarding data between nodes in a computing system (e.g., between computing nodes, or between computing nodes and control nodes). For example, a switching node receives data (e.g., data frames, data packets) from a node and forwards the data to the target node based on the data's destination node information (e.g., address information).
[0067] A switching node typically includes one or more switching chips. These chips support store-and-forward switching, cut-through switching, and high-speed network interconnect technologies. High-speed network interconnect technologies include, but are not limited to, NVLink, Compute Express Link (CXL), Peripheral Component Interconnect Express (PCIE), Universal Chiplet Interconnect Express (UCIE), Huawei Cache Coherent System (HCCS), and Cache Coherent Interconnect for Accelerators (CCIX).
[0068] (5) Direct communication link
[0069] A direct link, also known as a direct connection, refers to a physical link that connects two nodes directly. This link can be wired (e.g., cables, network cables, fiber optics; wireless communication ports such as Bluetooth). Data can be transmitted directly between the two nodes via this direct link without needing any intermediate devices.
[0070] The computing system involved in this application is described below with reference to the accompanying drawings.
[0071] Figure 1 shows a schematic diagram of a computing system. The computing system includes one or more control nodes 10, multiple computing nodes (computing node 201, computing node 202, ..., computing node 232), and one or more switching nodes 30. All computing nodes (computing node 201, computing node 202, ..., computing node 232) are connected to the control node 10.
[0072] Multiple computing nodes in a computing system can be connected via network interconnection technology or via at least one switching node.
[0073] It is understood that a computing system may include one or more computing devices.
[0074] If the computing system comprises only one computing device, that is, the single computing device constitutes the computing system, then the control node 10, computing nodes 201 to 232, and switching node 30 in the computing system shown in Figure 1 can be hardware circuits with corresponding functions. The computing nodes, control nodes, and switching nodes can be connected via high-speed network interconnection technologies, such as NVLink, CXL, PCIE, UCIE, HCCS, CCIX, etc., and this application does not impose any restrictions on this.
[0075] If the computing system includes multiple computing devices, then the control node 10, computing nodes 201 to 232, and switching node 30 in the computing system shown in Figure 1 can be independent electronic devices (such as independent hardware devices such as servers and switches) or hardware circuits with corresponding functions.
[0076] If the control node 10, computing nodes 201 to 232, and switching node 30 in the computing system are independent electronic devices, the computing nodes, control nodes, and switching nodes can be connected through network interconnection technologies, such as Ethernet, Fibre Channel, remote direct memory access (RDMA), RDMA over converged ethernet (RoCE) network, infinite bandwidth (IB) network, etc., and this application does not impose any restrictions on this.
[0077] If control node 10, computing nodes 201 to 232, and switching node 30 are hardware circuits with corresponding functions, then the multiple computing nodes shown in Figure 1 can be computing nodes located on the same physical device or computing nodes located on different physical devices. For example, computing nodes 201 to 232 shown in Figure 1 are located on the same server. Alternatively, computing nodes 201 to 216 are located on the same server, while computing nodes 217 to 232 are located on another server.
[0078] For ease of description, the following uses the computing system shown in Figure 1 as an example to illustrate the technical solution of this application.
[0079] It is understood that the computing system in the embodiments of this application can be used in fields such as high performance computing (HPC), artificial intelligence and machine learning, cloud data and cloud computing.
[0080] For example, taking the training and application of AI inference models as an example, the computing system processes data through multiple computing nodes, and then completes the final calculation through the interaction of data between multiple computing nodes.
[0081] For example, when the amount of data to be processed (such as training data during model training or input data during model application) is large, the control node of the computing system can distribute the data to be processed to multiple computing nodes for parallel processing. After multiple computing nodes have finished processing the corresponding data, they synchronize the calculation results obtained by themselves to other computing nodes, that is, data parallelism (DP).
[0082] For example, when the model is large, the control node of the computing system can distribute different parts of the model to be run (such as different operations) to multiple computing nodes for parallel processing. After the multiple computing nodes have finished processing the distributed computing tasks, they synchronize the computing results they have processed to other computing nodes, that is, model parallelism (MP).
[0083] In other words, multiple computing nodes in a computing system need to send data to each other during the parallel processing of data (such as the aforementioned data parallelism or model parallelism) in order to synchronize and aggregate the data.
[0084] Computing nodes typically communicate via direct communication links between themselves (e.g., sending and receiving data) or via shared switching nodes.
[0085] As mentioned earlier, although the communication latency of direct communication links between nodes is low, since there is a direct communication link between any two nodes in a fully connected network architecture, the number of network ports of each node in the fully connected network architecture is linearly related to the number of computing nodes in the computing system.
[0086] Currently, the number of network ports on computing nodes is usually small. If the number of computing nodes in a computing system is large, it is difficult to build a fully connected network if the number of network ports on the computing nodes is less than the number of computing nodes in the computing system.
[0087] For example, referring to the architecture shown in Figure 1, if compute nodes 201 to 232 only have 18 or 24 network ports, then to build a fully connected network, compute nodes 201 to 232 need to have at least 31 network ports, so a fully connected network cannot be built.
[0088] To address the aforementioned issues, in some cases, a non-blocking network structure (such as a CLOS network) can be constructed using switching nodes to facilitate data transmission between multiple nodes in the computing system.
[0089] For example, if the number of computing nodes in the computing system is large, a large-port switch node with a large number of ports (such as a switch chip with more than 24 ports, like 72 ports) can be used to connect all the computing nodes in the computing system. For example, taking the architecture shown in Figure 1 as an example, 32 computing nodes in the computing system shown in Figure 1 are connected through a large-port switch node.
[0090] For example, refer to the schematic diagram of a single-layer CLOS network structure shown in Figure 2. As shown in Figure 2, compute nodes 201 to 232 are connected to switch node 001, and switch node 001 has at least 32 ports.
[0091] For example, if a computing node requires a large bandwidth for data transmission, it can be achieved by establishing multiple communication links with a single switching node or multiple communication links with multiple switching nodes.
[0092] For example, refer to Figure 3, which illustrates another layer CLOS network structure. Computation nodes 201 to 232 are connected to switching nodes 0-1 to 0-9, respectively, and there are two communication links between each computation node and each switching node. In the network structure shown in Figure 3, each computation node requires 18 network ports, and every two network ports connect to one switching node, forming two communication links. Each switching node requires 72 network ports, of which 64 ports connect to 32 computation nodes (every two ports connect to one computation node), and 8 ports connect to 8 switching nodes.
[0093] It's understandable that if a CLOS network layer is constructed using switching nodes, the number of ports required by the switching nodes is linearly related to the number of compute nodes in the computing system. For example, the number of ports required by the switching nodes needs to be an integer multiple greater than or equal to the number of compute nodes in the computing system. If the number of compute nodes in the computing system is large, then the number of ports required by the switching nodes used to construct a CLOS network layer will also be large; that is, large-port switching nodes (e.g., switching nodes with more than 24 ports) are required.
[0094] However, current switching nodes (such as switches or switching chips) are typically implemented based on a crossbar architecture. Each input line and each output line of a switching node has a crossbar, connected by a semiconductor switch. When an input line of one port needs to send data to the output line of another port, the switch at the crossbar between that input and output line is connected (or turned on), allowing data to be sent from one port's input line to the other's output line.
[0095] For example, refer to the schematic diagram of the cross switch matrix architecture shown in Figure 4. As shown in Figure 4, assuming that the switching node includes L ports, where L is an integer greater than 0, then the L ports correspond to L input lines (e.g., input line IN1, input line IN2, input line IN3, input line IN4, input line IN5, ..., input line INL) and L output lines (e.g., output line OUT1, output line OUT2, output line OUT3, output line OUT4, output line OUT5, ..., output line OUTL).
[0096] It can be understood that the number of crossbars in the crossbar switch matrix architecture corresponding to a switching node is proportional to the square of the number of ports in the switching node. For example, if the number of ports in a switching node is L, then the number of crossbars in the crossbar switch matrix architecture corresponding to that switching node is L. 2 As the number of ports on a switching node increases, the hardware cost and complexity of the corresponding crossbar switch matrix architecture will increase dramatically.
[0097] For example, as the number of ports in a switching node increases, the crossbar switch matrix architecture corresponding to the switching node needs to be laid out with a greater number of connection lines (input lines and output lines), and the complexity and difficulty of cabling will increase significantly with the increase in the number of ports.
[0098] Therefore, current large-port switching nodes typically design the crossbar switch matrix architecture as a modular structure and build large-port switching nodes based on the combination of multiple crossbar switch matrix modules.
[0099] For example, refer to Figure 5, which shows a schematic diagram of a switching node architecture based on a combination of multiple cross switch matrix modules.
[0100] As shown in Figure 5, assuming that each cross switch matrix includes 16 cross points, that is, each cross switch matrix includes 4 input lines and 4 output lines, a 16-port switching node (e.g., port 01 to port 16 in the figure) can be constructed based on 8 cross switch matrices (e.g., cross switch matrix 1 to cross switch matrix 8 in the figure).
[0101] It's understandable that large-port switching nodes employ an architecture combining multiple crossbar switch matrix modules, requiring data to be forwarded through these modules within the node. Therefore, large-port switching nodes typically experience higher communication latency. The more ports a switching node has, the more crossbar switch matrix modules it may contain, potentially leading to even higher communication latency.
[0102] Furthermore, because large-port switching nodes employ an architecture combining multiple crossbar switch matrix modules, while low-port switching nodes use a single crossbar switch matrix module or a smaller number of crossbar switch matrix modules, the manufacturing process and cost of large-port switching nodes are significantly higher. Therefore, large-port switching nodes are considerably more expensive than low-port switching nodes; for example, their cost can be tens of times higher.
[0103] Furthermore, large-port switching nodes employ an architecture combining multiple crossbar switch matrix modules, while low-port switching nodes use a single crossbar switch matrix module or a smaller number of crossbar switch matrix modules. Therefore, large-port switching nodes consume significantly more power than low-port switching nodes; for example, the power consumption of a large-port switching node is tens of times that of a low-port switching node.
[0104] In other cases, a fat-tree network structure can be constructed using swap nodes, enabling data transmission between multiple nodes in the computing system through a multi-layer (e.g., two-layer) swap node network.
[0105] For example, taking a two-layer fat-tree network structure as an example, if the number of computing nodes in the computing system is large, multiple low-port switching nodes (e.g., with fewer than or equal to 24 ports) can be used as the access layer of the fat-tree network structure to connect all computing nodes in the computing system, and one switching node can be used as the aggregation layer of the fat-tree network structure to connect multiple switching nodes in the access layer. For example, 32 computing nodes in the computing system shown in Figure 1 can be connected through four 16-port low-port switching nodes, and then the four 16-port low-port switching nodes can be connected through one 16-port low-port switching node.
[0106] For example, refer to the schematic diagram of the two-layer fat tree network structure shown in Figure 6. As shown in Figure 6, computing nodes 201 to 208 are connected to switching node 11, computing nodes 209 to 216 are connected to switching node 12, computing nodes 217 to 224 are connected to switching node 13, computing nodes 225 to 232 are connected to switching node 14, and switching nodes 11 to 14 are connected to switching node 15.
[0107] For example, if compute node 201 needs to send data to compute node 232, compute node 201 sends the data to be sent to exchange node 11, exchange node 11 then sends the received data to exchange node 15, exchange node 15 then sends the received data to exchange node 14, and exchange node 14 then sends the received data to compute node 232.
[0108] It is understandable that data transmission between computing nodes requires forwarding through multiple switching nodes (such as the three forwardings shown in Figure 6) to complete, resulting in high communication latency. In scenarios with high timeliness requirements, such as running large models for querying, this can lead to longer execution times for large models, thus impacting the user experience.
[0109] In summary, communication between computing nodes in a computing system can be achieved by connecting a single large-port switching node to all computing nodes in the computing system to form a single CLOS network, or by using multiple low-port switching nodes in a multi-layer fat tree structure. However, both methods result in high communication latency between nodes.
[0110] In view of this, embodiments of this application provide a computing system in which computing nodes establish communication connections through multiple low-port switching nodes, and any two computing nodes can communicate (e.g., send and receive data) by forwarding through at most one switching node. Since the internal structure of low-port switching nodes is relatively simple (e.g., consisting of only a crossbar switch matrix) and the communication latency is also low, establishing communication connections between multiple computing nodes through multiple low-port switching nodes in the computing system can reduce the communication latency between any two computing nodes.
[0111] Specifically, the computing system includes X switching nodes and M computing nodes. Each switching node has fewer than M ports. Any two nodes in the system are connected to the same switching node via a direct communication link. Each switching node has more than 0 ports and less than M ports. Here, M is an integer greater than 2, and X is an integer greater than 1. Furthermore, any two computing nodes among the M nodes communicate (e.g., for data synchronization) through the same switching node, with each node having a direct communication link to the other two nodes.
[0112] For example, a first computing node out of M computing nodes is connected to a first switching node out of X switching nodes via a first direct communication link, and a second computing node out of M computing nodes is connected to the first switching node via a second direct communication link. Furthermore, the first computing node and the second computing node communicate with each other via the first direct communication link, the first switching node, and the second direct communication link. For example, the first computing node sends first data to the first switching node via the first direct communication link, and the first switching node sends the received first data to the second computing node via the second direct communication link.
[0113] It is understandable that the direct communication links between each computing node and the switching node can be implemented based on high-speed interconnect technology, such as NVLink, CXL or other types of technology.
[0114] It is understood that the switching node in this application embodiment includes a cross switch matrix, and data will not be forwarded multiple times within the switching node.
[0115] Specifically, since the low-port switching node with fewer ports has the aforementioned cross-switch matrix architecture, and the large-port switching node with more ports has the aforementioned architecture of multiple cross-switch matrix modules, the communication latency of the low-port switching node with fewer ports is lower than that of the large-port switching node with more ports. Compared with the above-mentioned method of using large-port switching nodes with more ports, the communication latency of data transmission between computing nodes in the computing system of this application embodiment is lower.
[0116] Furthermore, compared to the above-mentioned method of constructing a fat tree network structure using switching nodes, the data transmission between computing nodes in the computing system of this application embodiment only needs to be forwarded once by the switching node, resulting in lower communication latency for data transmission.
[0117] Furthermore, since the cost and power consumption of low-port switching nodes are much lower than those of high-port switching nodes (e.g., by tens of times), the computing system of this application embodiment can also reduce network costs and network power consumption.
[0118] In some embodiments, K of the M computing nodes are connected to each other via direct communication links, where K is an integer greater than 1 and less than or equal to M. Furthermore, any two of the K computing nodes communicate through the same switching node, each having a direct communication link with both of them, and / or, any two of the K computing nodes communicate through a direct communication link between any two computing nodes.
[0119] For example, the third computing node out of M computing nodes is connected to the second switching node out of X switching nodes via a third direct communication link, and the fourth computing node out of M computing nodes is connected to the second switching node via a fourth direct communication link. There is a fifth direct communication link between the third computing node and the fourth computing node (i.e., they are connected via a fifth direct communication link).
[0120] Furthermore, the third computing node and the fourth computing node communicate through a third direct communication link, the second switching node and the fourth direct communication link, and / or, the third computing node and the fourth computing node communicate through a fifth direct communication link.
[0121] For example, the third computing node sends the second data to the fourth computing node through the third direct communication link, the second switching node, and the fourth direct communication link.
[0122] Alternatively, the third computing node can send the second data to the fourth computing node via the fifth direct communication link.
[0123] Alternatively, the third computing node may send a portion of the second data to the fourth computing node via a third direct communication link, the second switching node, and a fourth direct communication link, and send another portion of the second data to the fourth computing node via a fifth direct communication link. For example, the third computing node may split the second data into a first sub-data and a second sub-data. The third computing node sends the first sub-data to the second switching node via the third direct communication link, the second switching node sends the received first sub-data to the fourth computing node via the fourth direct communication link, and the third computing node sends the second sub-data to the fourth computing node via the fifth direct communication link.
[0124] In some embodiments, the amount of data in the first sub-data and the second sub-data can be determined based on the bandwidth of the third communication link (and / or the fourth communication link) and the bandwidth of the fifth communication link.
[0125] In other embodiments, the data volume of the first sub-data and the second sub-data can be determined based on the bandwidth of the third communication link (and / or the fourth communication link) and the available bandwidth of the fifth communication link. This application does not limit this.
[0126] It is understandable that, since the communication latency of direct communication links between computing nodes is low, when latency requirements are high, if there is a direct communication link between two computing nodes (for example, there is a fifth direct communication link between the third and fourth computing nodes mentioned above), communication can be carried out solely through the direct communication link between the computing nodes, thus reducing communication latency.
[0127] It is understandable that when the amount of data to be transmitted between two computing nodes is large, the data can be split into two parts, and the two parts can be transmitted separately through the direct communication link between the computing nodes and the switching node that is jointly linked by the computing nodes. This can improve the data transmission efficiency.
[0128] It is understandable that when the amount of data to be transmitted between two computing nodes is large, if there is a direct communication link between the two computing nodes and the available bandwidth of the direct communication link is limited, the data can be split into two parts and transmitted separately through the direct communication link between the computing nodes and the switching node jointly linked by the computing nodes. This can improve the data transmission efficiency.
[0129] In some embodiments, the M computing nodes in the computing system are divided into N computing node groups, and each computing node group includes at least one computing node. Then, any two computing nodes in the computing node group can be connected to the same switching node, thus enabling any two computing nodes in the computing system to be connected to the same switching node.
[0130] In this configuration, the first computing node and the second computing node are located in the same computing node group, or the first computing node and the second computing node are located in different computing node groups. Similarly, the third computing node and the fourth computing node are located in the same computing node group, or the third computing node and the fourth computing node are located in different computing node groups.
[0131] It is understood that the number of computing nodes in each computing node group can be the same or different. Computing nodes in any two computing node groups can be connected through one or more switching nodes. This application does not limit this connection. For example, one computing node and another computing node can be simultaneously connected to both switching node A and switching node B. This not only increases the bandwidth of the communication links between computing nodes but also improves the security and stability of the computing system. For instance, if switching node A fails, communication can still be maintained through switching node B.
[0132] It is understandable that if the number of computing nodes in the computing system is large, by dividing the computing nodes into computing node groups, each switching node only needs to connect to computing nodes in at least two computing node groups, and low-port switching nodes with fewer ports can be used.
[0133] In some embodiments, the computing nodes in the first computing node group and the second computing node group of the N computing node groups are all connected to at least two switching nodes via direct communication links, and the X switching nodes include at least two switching nodes.
[0134] For example, consider computing nodes 201 to 232 in the computing system shown in Figure 1. Assume that computing nodes 201 to 232 are divided into 4 computing node groups, and each computing node group includes 8 computing nodes. The computing nodes in any two computing node groups are connected through a switching node.
[0135] Refer to Figure 7 for a schematic diagram of the CLOS network structure.
[0136] As shown in Figure 7, computing nodes 201 to 208 form computing node group 1, computing nodes 209 to 216 form computing node group 2, computing nodes 217 to 224 form computing node group 3, and computing nodes 225 to 232 form computing node group 4.
[0137] Switching node 1-2 connects the data transmission between each computing node in computing node group 1 and each computing node in computing node group 2; that is, each computing node in computing node group 1 and each computing node in computing node group 2 are connected to switching node 1-2.
[0138] Switching nodes 1-3 are used for data transmission between the computing nodes in computing node group 1 and the computing nodes in computing node group 3. That is, the computing nodes in computing node group 1 and the computing nodes in computing node group 3 are connected to switching nodes 1-3.
[0139] Switching nodes 1-4 are used for data transmission between the computing nodes in computing node group 1 and the computing nodes in computing node group 4. That is, the computing nodes in computing node group 1 and the computing nodes in computing node group 4 are connected to switching nodes 1-4.
[0140] Switching nodes 2-3 are used for data transmission between the computing nodes in computing node group 2 and the computing nodes in computing node group 3. That is, the computing nodes in computing node group 2 and the computing nodes in computing node group 3 are connected to switching nodes 2-3.
[0141] Switching nodes 2-4 are used for data transmission between the computing nodes in computing node group 2 and the computing nodes in computing node group 4. That is, the computing nodes in computing node group 2 and the computing nodes in computing node group 4 are connected to switching nodes 2-4.
[0142] Switching nodes 3-4 are used for data transmission between the computing nodes in computing node group 3 and the computing nodes in computing node group 4. That is, the computing nodes in computing node group 3 and the computing nodes in computing node group 4 are connected to switching nodes 3-4.
[0143] Each switching node has at least 16 ports, and each computing node has at least 3 ports.
[0144] It is understood that Figure 7 only shows the connection relationship between exchange node 1-2 and each computing node in computing node group 1 and each computing node in computing node group 2. The connection relationship between exchange node 1-3, exchange node 1-4, exchange node 2-3, exchange node 2-4 and exchange node 3-4 and the computing units in the corresponding computing node groups can be referred to the connection relationship between exchange node 1-2 and each computing node in computing node group 1 and each computing node in computing node group 2.
[0145] For example, consider computing nodes 201 to 232 in the computing system shown in Figure 1. Assume that computing nodes 201 to 232 are divided into four computing node groups, and each computing node group includes eight computing nodes. The computing node groups in any two of the four computing node groups are connected through six switching nodes.
[0146] Refer to Figure 8 for a schematic diagram of another CLOS network structure.
[0147] As shown in Figure 8, computing nodes 201 to 208 form computing node group 1, computing nodes 209 to 216 form computing node group 2, computing nodes 217 to 224 form computing node group 3, and computing nodes 225 to 232 form computing node group 4.
[0148] Switching nodes 11-12 are used for data transmission between the computing nodes in computing node group 1; that is, switching nodes 11-12 are connected to each computing node in computing node group 1. Switching nodes 21-22 are connected to each computing node in computing node group 2. Switching nodes 31-32 are connected to each computing node in computing node group 3. Switching nodes 41-42 are connected to each computing node in computing node group 4.
[0149] Switching nodes 11-21, 11-22, 12-21, and 21-22 are used for data transmission between the computing nodes in computing node group 1 and the computing nodes in computing node group 2.
[0150] Referring to Figure 8, for example, switching nodes 11-21 are connected to computing nodes 201, 202, 203, and 204 in computing node group 1, and computing nodes 209, 210, 211, and 212 in computing node group 2. Switching nodes 11-22 are connected to computing nodes 201, 202, 203, and 204 in computing node group 1, and computing nodes 213, 214, 215, and 216 in computing node group 2. Switching nodes 12-21 are connected to computing nodes 205, 206, 207, and 208 in computing node group 1, and computing nodes 209, 210, 211, and 212 in computing node group 2. Switching nodes 12-22 are connected to compute nodes 205, 206, 207, and 208 in compute node group 1 and compute nodes 213, 214, 215, and 216 in compute node group 2.
[0151] Switching nodes 11-31, 11-32, 12-31, and 12-32 are used for data transmission between the computing nodes in computing node group 1 and the computing nodes in computing node group 3.
[0152] Switching nodes 11-41, 11-42, 12-41, and 12-42 are used for data transmission between the computing nodes in computing node group 1 and the computing nodes in computing node group 4.
[0153] Switching nodes 21-31, 21-32, 22-31, and 22-32 are used for data transmission between the computing nodes in computing node group 2 and the computing nodes in computing node group 3.
[0154] Switching nodes 21-41, 21-42, 22-41, and 22-42 are used for data transmission between the computing nodes in computing node group 2 and the computing nodes in computing node group 4.
[0155] Switching nodes 31-41, 31-42, 32-41, and 32-42 are used for data transmission between the computing nodes in computing node group 3 and the computing nodes in computing node group 4.
[0156] Each switching node has at least 8 ports, and each computing node has at least 7 ports.
[0157] It is understood that Figure 8 only shows the connection relationships between each computing node in computing node group 1 and each computing node in computing node group 2 and their corresponding exchange nodes. The connection relationships between computing nodes in other computing node groups and their corresponding exchange nodes can be referenced from the connection relationships between each computing node in computing node group 1 and each computing node in computing node group 2 and their corresponding exchange nodes.
[0158] It is understood that the connection schemes shown in Figures 7 and 8 above can also be combined. For example, each computing node in computing node group 1 and each computing node in computing node group 2 can be connected through 6 switching nodes as shown in Figure 8. The connection relationship between computing nodes in other computing node groups and their corresponding switching nodes can be connected through 1 switching node as shown in Figure 7. This application does not impose any restrictions on this.
[0159] It is understood that if the number of computing nodes in each computing node group is different, for example, in the computing system shown in Figure 7 or Figure 8 above, computing nodes 201 to 232 are divided into 5 computing node groups, and computing node groups 1 to 3 each include 8 computing nodes, while computing node groups 4 and 5 each include 4 computing nodes. Then, the computing nodes in every two computing node groups 1 to 5 are connected to a switching node. Alternatively, the computing nodes in every two computing node groups 1 to 3 are connected to a switching node, and the computing nodes in computing node groups 1 to 3 are connected to the same switching node as the computing nodes in computing node groups 4 and 5, respectively. This application does not impose specific restrictions on this, as long as any two computing nodes in the computing system are connected to the same switching node.
[0160] It is understood that the connection schemes shown in Figures 7 and 8 above are only one example of the connection between computing nodes and switching nodes. In other embodiments, the number of each switching node in Figures 7 and 8 can also be multiple, thereby improving the security and stability of the computing system. For example, assuming that the number of each switching node in Figures 7 and 8 is one, the constructed CLOS network is a dual-track network. When one of the two switching nodes fails, the other can be used for communication, thereby ensuring the normal operation of the computing system and improving the security and stability of the computing system.
[0161] It is understood that in some embodiments, any two computing nodes in the computing system are connected via direct communication links. That is, all computing nodes in the computing system are interconnected, forming a complete network structure. In other words, K computing nodes are connected to each other via direct communication links, where K equals M. Thus, in addition to data transmission through switching nodes, data can also be transmitted between computing nodes via direct communication links.
[0162] In other words, any two computing nodes among the K computing nodes communicate through the same switching node that has a direct communication link with each of the two computing nodes, and / or, any two computing nodes among the K computing nodes communicate through a direct communication link between any two computing nodes.
[0163] However, if there are many computing nodes in the computing system, each computing node needs to be configured with a lot of communication ports to connect the computing nodes to each other and form a complete network structure.
[0164] Therefore, in some other embodiments, in the N computing node groups of the computing system, the computing nodes in each of the Y computing node groups are connected to each other through direct communication links, where Y is an integer greater than 0 and less than or equal to N.
[0165] For example, for a computing node group with more than one computing node, any two computing nodes in the computing node group are connected via a direct communication link. That is, all computing nodes in the computing node group are interconnected, forming a complete network structure (hereinafter referred to as a one-dimensional complete network structure). In other words, Y is the number of computing node groups with more than one computing node in N computing node groups.
[0166] For example, consider computing nodes 201 to 232 in the computing system shown in Figure 1. Refer to Figure 9 for a schematic diagram of the network structure between computing nodes within the computing node group.
[0167] As shown in Figure 9, compute nodes 201 to 208 form compute node group 1, compute nodes 209 to 216 form compute node group 2, compute nodes 217 to 224 form compute node group 3, and compute nodes 225 to 232 form compute node group 4. The connection relationships between compute nodes and switching nodes can be found in the relevant description in Figure 7.
[0168] Furthermore, any two computing nodes within each computing node group are connected to form a one-dimensional full network structure.
[0169] For example, continuing to refer to Figure 11, computing node 225 is connected to computing nodes 226, 227, 228, 229, 230, 231 and 232 respectively, forming 7 direct communication links.
[0170] Computing node 226 is connected to computing nodes 225, 227, 228, 229, 230, 231 and 232 respectively, forming 7 direct communication links.
[0171] Computing node 227 is connected to computing nodes 225, 226, 228, 229, 230, 231 and 232 respectively, forming 7 direct communication links.
[0172] Computing node 228 is connected to computing nodes 225, 226, 227, 229, 230, 231 and 232 respectively, forming 7 direct communication links.
[0173] Computing node 228 is connected to computing nodes 225, 226, 227, 229, 230, 231 and 232 respectively, forming 7 direct communication links.
[0174] Computing node 229 is connected to computing nodes 225, 226, 227, 228, 230, 231 and 232 respectively, forming 7 direct communication links.
[0175] Computing node 230 is connected to computing nodes 225, 226, 227, 228, 229, 231 and 232 respectively, forming 7 direct communication links.
[0176] Computing node 231 is connected to computing nodes 225, 226, 227, 228, 229, 230 and 232 respectively, forming 7 direct communication links.
[0177] Computing node 232 is connected to computing nodes 225, 226, 227, 228, 229, 230 and 231 respectively, forming 7 direct communication links.
[0178] There are a total of 28 direct communication links within computing node group 4.
[0179] It is understood that Figure 9 only shows a schematic diagram of the connection relationships between the computing nodes in computing node group 4. For details on the connection relationships between the computing nodes in computing node group 1, computing node group 2, and computing node group 3, please refer to the connection relationships between the computing nodes in computing node group 4.
[0180] In some embodiments, the direct communication links between computing nodes can be implemented based on high-speed interconnect technologies, such as NVLink, CXL, or other types of technologies.
[0181] In some embodiments, computing nodes with the same identification information in N computing node groups are connected to each other via direct communication links. The same identification information includes, but is not limited to, number and location.
[0182] For example, N computing nodes with the same identification information in an N-node group are connected to each other through direct communication links to form a full network structure (hereinafter referred to as a two-dimensional full network structure).
[0183] For example, assuming that each computing node group includes M / N computing nodes, the i-th computing node in each computing node group is connected through a direct communication link, where i is any integer from 1 to M / N.
[0184] In this way, in addition to data transmission through exchange nodes, computing nodes with the same identification information in different computing node groups can also transmit data through direct communication links between computing nodes.
[0185] For example, consider computing nodes 201 to 232 in the computing system shown in Figure 1. Refer to Figure 10 for a schematic diagram of the network structure between computing nodes in the computing node group.
[0186] As shown in Figure 10, compute nodes 201 to 208 form compute node group 1, compute nodes 209 to 216 form compute node group 2, compute nodes 217 to 224 form compute node group 3, and compute nodes 225 to 232 form compute node group 4. The connection relationships between compute nodes and switching nodes can be found in the relevant description in Figure 7.
[0187] For example, when the identification information is the position of a computing node within a computing node group, computing nodes 201, 209, 217, and 225 have the same identification information because they are the first computing nodes in their respective computing node groups. Similarly, computing nodes 202, 210, 218, and 226 have the same identification information. Computing nodes 203, 211, 219, and 227 have the same identification information. Computing nodes 204, 212, 220, and 228 have the same identification information. Computing nodes 205, 213, 221, and 229 have the same identification information. Computing nodes 206, 214, 222, and 230 have the same identification information. Computational nodes 207, 215, 223, and 231 have the same identification information. Computational nodes 208, 216, 224, and 232 have the same identification information.
[0188] Furthermore, any two computing nodes with the same positional relationship are connected to form a two-dimensional full network structure. For example, continuing to refer to Figure 12, computing node 201 is connected to computing nodes 209, 217, and 225, forming three direct communication links. Computing node 209 is connected to computing nodes 201, 217, and 225, forming three direct communication links. Computing node 217 is connected to computing nodes 201, 209, and 225, forming three direct communication links. Computing node 225 is connected to computing nodes 201, 209, and 217, forming three direct communication links. A total of six direct communication links exist between multiple computing nodes with the same identification information in computing node groups 1, 2, 3, and 4.
[0189] It is understood that Figure 10 only shows a schematic diagram of the connection relationships between the computing nodes in computing node group 4. For details on the connection relationships between the computing nodes in computing node group 1, computing node group 2, and computing node group 3, please refer to the connection relationships between the computing nodes in computing node group 4.
[0190] In some embodiments, the direct communication links between computing nodes can be implemented based on high-speed interconnect technologies, such as NVLink, CXL, or other types of technologies.
[0191] In some embodiments, computing nodes connected via direct communication links can achieve memory pooling and share memory resources through memory sharing technology.
[0192] In other embodiments, the computing node groups in the computing system may also have independent memory resources. That is, computing nodes within each computing node group share the same memory resources, while computing nodes in different computing node groups do not share memory resources, thereby improving security. For example, different computing node groups can be used to process data from different users. In other words, user data from different users is stored in different memories and cannot be accessed by computing nodes in other computing node groups.
[0193] It is understood that the computing system provided in the embodiments of this application may include only the structure shown in FIG9 above, or only the structure shown in FIG10 above, or both the structure shown in FIG9 above and the structure shown in FIG10 above. This application does not limit this.
[0194] In other words, in the N computing node groups of the computing system, the computing nodes in each of the Y computing node groups are connected to each other through direct communication links, where Y is an integer greater than 0 and less than or equal to N; and / or, computing nodes with the same identification information in the N computing node groups are connected to each other through direct communication links.
[0195] In some embodiments, the fifth computing node among the M computing nodes is connected to the third switching node among the X switching nodes via a sixth direct communication link, the sixth computing node among the M computing nodes is connected to the third switching node via a seventh direct communication link, and the fifth computing node and the sixth computing node are connected via an eighth direct communication link (i.e., connected via an eighth direct communication link); wherein, the fifth computing node and the sixth computing node are computing nodes within one of the Y computing node groups, or the fifth computing node and the sixth computing node are computing nodes with the same identification information in N computing node groups.
[0196] Furthermore, the fifth computing node and the sixth computing node communicate through the sixth direct communication link, the third switching node and the seventh direct communication link, and / or, the fifth computing node and the sixth computing node communicate through the eighth direct communication link.
[0197] For example, the fifth computing node sends the third data to the sixth computing node through the sixth direct communication link, the third switching node, and the seventh direct communication link.
[0198] Alternatively, the fifth computing node can send the third data to the sixth computing node via the eighth direct communication link.
[0199] Alternatively, the fifth computing node can send the third data to the sixth computing node via a sixth direct communication link, a third switching node, a seventh direct communication link, and an eighth direct communication link. For example, the fifth computing node splits the third data into a first sub-data and a second sub-data. The fifth computing node sends the first sub-data to the third switching node via the sixth direct communication link, and the third switching node sends the received first sub-data to the sixth computing node via the seventh direct communication link. The fifth computing node then sends the second sub-data to the sixth computing node via the eighth direct communication link.
[0200] In some embodiments, the computing system includes N control node groups, each control node group including at least one control node, and a control node in one control node group controls at least one computing node in a computing node group. That is, in the N computing node groups, the computing nodes in each computing node group are connected to the same one or more control nodes.
[0201] In this process, the control node sends control signals to the connected computing node through the communication link between the control node and the computing node to control the computing node to process the corresponding computing tasks.
[0202] For example, taking computing nodes 201 to 232 in the computing system shown in Figure 1 as an example, refer to the schematic diagram of the network structure between the control node and the computing node shown in Figure 11.
[0203] As shown in Figure 11, compute nodes 201 to 208 form compute node group 1, compute nodes 209 to 216 form compute node group 2, compute nodes 217 to 224 form compute node group 3, and compute nodes 225 to 232 form compute node group 4. The connection relationships between compute nodes and switching nodes can be found in the relevant description in Figure 7.
[0204] The control node 301 in the control node group 300 is connected to each computing node in the computing node group 1, forming 8 communication links.
[0205] The control node 311 in the control node group 310 is connected to each computing node in the computing node group 2, forming 8 communication links.
[0206] The control node 321 in the control node group 320 is connected to each computing node in the computing node group 3, forming 8 communication links.
[0207] The control node 331 in the control node group 330 is connected to each computing node in the computing node group 4, forming 8 communication links.
[0208] It is understood that Figure 11 only shows a schematic diagram of the connection relationship between control node 301 and each computing node in computing node group 1. The connection relationship between control node 311, control node 321 and control node 331 and the computing units in the corresponding computing node groups can be referred to the connection relationship between control node 301 and each computing node in computing node group 1.
[0209] In some embodiments, the communication link between the computing node and the control node can be based on the PCIE protocol or other types of protocols.
[0210] It is understood that in other embodiments, the control node may also connect to the corresponding computing nodes through a switching node. For example, one control node and multiple corresponding computing nodes may be connected to the same switching node. The control node communicates with the corresponding computing nodes through communication links between the control node and the switching node, and between the switching node and the computing nodes. The communication links between the control node and the switching node, and between the switching node and the computing nodes, can be based on the PCIe protocol or other types of protocols.
[0211] Thus, a control node only needs to connect to a switching node through a single communication interface, instead of connecting to at least one computing node in the computing node group through multiple communication interfaces. When the number of computing nodes in the computing node group is large, the control node does not need to be configured with a large number of communication interfaces.
[0212] The following section will take computing nodes 201 to 232 in the computing system shown in Figure 1 as an example, and describe in detail the structure of the computing system provided in the embodiments of this application with reference to the accompanying drawings.
[0213] For example, Figure 12 shows a schematic diagram of the network structure of a computing system according to an embodiment of this application.
[0214] As shown in Figure 12, the computing system includes computing node group 1, computing node group 2, computing node group 3 and computing node group 4. Each computing node group includes 8 computing nodes (shown as 1-D1 to 1-D8, 2-D1 to -D8, 3-D1 to 3-D8 or 4-D1 to 4-D8).
[0215] Switch node SW-12 connects the compute nodes in compute node group 1 and compute node group 2. Switch node SW-13 connects the compute nodes in compute node group 1 and compute node group 3. Switch node SW-14 connects the compute nodes in compute node group 1 and compute node group 4. Switch node SW-23 connects the compute nodes in compute node group 2 and compute node group 3. Switch node SW-24 connects the compute nodes in compute node group 2 and compute node group 4. Switch node SW-34 connects the compute nodes in compute node group 3 and compute node group 4.
[0216] The CLOS network consists of two switching nodes: SW-12, SW-13, SW-14, SW-23, SW-24, and SW-34. This dual-track network enhances the security of the computing system. For example, if one switching node, SW-12, fails, the compute nodes in compute node group 1 and compute node group 2 can communicate through the other connected switching node, SW-12.
[0217] Referring again to Figure 12, the computing nodes within each computing node group (shown as 1-D1 to 1-D8, 2-D1 to 1-D8, 3-D1 to 3-D8, or 4-D1 to 4-D8) form a one-dimensional fully connected network. A direct communication link exists between any two computing nodes within each computing node group; the specific connection relationships can be found in the network structure diagram shown in Figure 11.
[0218] A two-dimensional fully connected network is constructed by any two computing nodes in a computing node group that have the same identification information (e.g., D1 (1-D1, 2-D1, 3-D1, and 4-D1) in a four-node computing node group, D2 (1-D2, 2-D2, 3-D2, and 4-D2) in a four-node computing node group, ..., D8 (1-D8, 2-D8, 3-D8, and 4-D8) in a four-node computing node group). A direct communication link exists between any two computing nodes with the same identification information (e.g., number). The specific connection relationship can be referred to in the schematic diagram of the network structure shown in Figure 12.
[0219] Referring again to Figure 12, control node C1 connects to all computing nodes in computing node group 1, control node C2 connects to all computing nodes in computing node group 2, control node C3 connects to all computing nodes in computing node group 3, and control node C4 connects to all computing nodes in computing node group 4. The number of control nodes C1, C2, C3, and C4 is two.
[0220] It is understood that in the network structure shown in Figure 12, each computing node includes at least 18 network ports. Among them, 6 network ports are used to connect to the switching node, 7 network ports are used to connect to the 7 computing nodes in the computing node group corresponding to this computing node, 3 network ports are used to connect to 3 computing nodes with the same identification information as this computing node, and 2 network ports are used to connect to the corresponding control node.
[0221] Each switching node includes at least 16 ports for connecting compute nodes in two compute node groups.
[0222] It is understandable that if the number of ports on a switching node is greater than the number of connected compute nodes, for example, if each switching node shown in Figure 12 includes 18 ports, then two of these ports can be used to connect backup nodes. Backup nodes are used to perform the corresponding computational tasks in place of the failed compute nodes. If the forwarding latency of a switching node is 100 nanoseconds (ns), and the communication latency of a direct communication link between any two compute nodes is 300 ns, then the communication latency for any two compute nodes communicating through a switching node is 400 ns.
[0223] It is understood that the network structure shown in Figure 12 is only an example. In other embodiments, the number of computing nodes in each computing node group may be different, and computing nodes in any two computing node groups may communicate through more or fewer switching nodes. This application does not limit this.
[0224] It is understandable that, compared to the network structure shown in Figure 3, which uses 9 large-port switching nodes (including 72 ports), the network structure shown in Figure 12 only uses 12 low-port switching nodes (including at least 16 ports). Since the cost and power consumption of large-port switching nodes are usually tens of times that of low-port switching nodes, and the number of low-port switching nodes used in the network structure shown in Figure 12 is only a few more than the number of large-port switching nodes used in the network structure shown in Figure 3, the network cost and power consumption of the network structure shown in Figure 3 are much higher than those of the network structure shown in Figure 12.
[0225] It is understood that this application also provides a communication method for the above-described computing system.
[0226] The following description uses a computing system comprising computing node P (e.g., the aforementioned first computing node, third computing node, or fifth computing node) and computing node Q (e.g., the aforementioned second computing node, fourth computing node, or sixth computing node) as an example to illustrate the communication method provided in this application.
[0227] In this computing system, computing node P and computing node Q are connected to the same switching node H (e.g., the aforementioned first switching node, second switching node, or third switching node), computing node P and switching node H have a direct communication link (e.g., the aforementioned first direct communication link, third direct communication link, or sixth direct communication link), and computing node Q and switching node H have a direct communication link (e.g., the aforementioned second direct communication link, fourth direct communication link, or seventh direct communication link).
[0228] Computing node P sends the data to be sent (such as the aforementioned first data, second data, or third data) to switching node H through a direct communication link with switching node H. Switching node H then sends the data to be sent to computing node Q through a direct communication link with computing node Q.
[0229] In some embodiments, the M computing nodes in the computing system are divided into N computing node groups, and each computing node group includes at least one computing node. Then, computing node P and computing node Q may be located in the same computing node group or in different computing node groups.
[0230] In some embodiments, a direct communication link exists between any two computing nodes in the computing system. For example, a direct communication link exists between computing node P and computing node Q (e.g., the aforementioned fifth or eighth direct communication link). Computing node P sends data to be sent to computing node Q through this direct communication link.
[0231] In other embodiments, for a computing node group with more than one computing node, any two computing nodes in the computing node group are connected via a direct communication link. If computing node P and computing node Q are located in the same computing node group, then there is a direct communication link between computing node P and computing node Q (e.g., the aforementioned eighth direct communication link). Computing node P sends data to be sent to computing node Q via the direct communication link with switching node H, switching node H, and the direct communication link between switching node H and computing node Q; and / or, computing node P sends data to be sent to computing node Q via the direct communication link with computing node Q.
[0232] In other embodiments, there are direct communication links between computing nodes with the same identification information in N computing node groups. The same identification information includes, but is not limited to, number and location. If computing node P and computing node Q are located in different computing node groups and have the same identification information, then there is a direct communication link between computing node P and computing node Q (e.g., the aforementioned eighth direct communication link). Computing node P sends data to be sent to computing node Q through the direct communication link with switching node H, switching node H, and the direct communication link between switching node H and computing node Q; and / or, computing node P sends data to be sent to computing node Q through the direct communication link with computing node Q.
[0233] It is understood that the aforementioned data to be sent (such as the first data, second data, or third data mentioned above) is data that computing node P needs to send to computing node Q. The data to be sent can be data generated by computing node P during the process of the computing system running the neural network model, or it can be data received by computing node P during the process of the computing system running the neural network model. This application does not impose any restrictions on this.
[0234] In some embodiments, the direct communication link described above is implemented based on high-speed interconnect technology, such as NVLink, CXL, or other types of technology.
[0235] In some embodiments, the collective communication method for synchronizing data between computing node P and computing node Q may specifically include various methods such as reduce-scatter, all-gather, and all-reduce.
[0236] It is understood that the above description only illustrates the connection between any two computing nodes through a single direct communication link. In other embodiments, any two computing nodes may be connected through multiple direct communication links, and this application does not limit this. Similarly, the above description only illustrates the connection between a computing node and a switch through a single direct communication link. In other embodiments, a computing node and a switch may be connected through multiple direct communication links, and this application does not limit this. Finally, the above description only illustrates the connection between any two computing nodes and the same switch node. In other embodiments, any two computing nodes may also be connected to multiple identical switch nodes, and this application does not limit this.
[0237] For example, Figure 13 illustrates a flowchart of a communication method according to an embodiment of this application. It can be understood that the communication method shown in Figure 13 can be applied to a computing system having the network structure provided in the above embodiments.
[0238] As shown in Figure 13, the method includes:
[0239] S1301: Computing node P detects a send request, which instructs computing node P to send the data to be sent to computing node Q.
[0240] In some embodiments, the data to be sent is data generated by computing node P during the data processing (running an AI inference model) of the computing system (e.g., gradient data of computing node P during data parallel processing).
[0241] In other embodiments, the data to be sent may also be data received by computing node P during data processing by the computing system (for example, data received by computing node P from other computing nodes besides computing node Q). This application does not impose any limitations on this.
[0242] It is understood that the data to be sent is the data that computing node P needs to send to computing node Q, as indicated by a send request detected by computing node P (as an instance of a first request, second request, or third request). For example, the data to be sent may specifically be the data that computing node P needs to send to computing node Q during stages such as reduce-scatter, all-gather, and all-reduce. This application does not limit this.
[0243] It is understandable that before sending data to computing node Q, computing node P can first determine whether there is a direct communication link between computing node P and computing node Q.
[0244] Because the communication latency of direct communication links between computing nodes is low, if there is a direct communication link between computing node P and computing node Q (such as the aforementioned fifth or eighth direct communication link), computing node P can communicate with computing node Q through the direct communication link between them.
[0245] If there is no direct communication link between computing node P and computing node Q, computing node P can communicate with computing node Q through the switching node H, which is connected to computing node Q.
[0246] S1302: This corresponds to a direct communication link between computing node P and computing node Q. Computing node P sends the data to be sent to computing node Q through the direct communication link between computing node P and computing node Q.
[0247] In some embodiments, a direct communication link exists between any two computing nodes in the computing system. Computing node P sends data to be sent to computing node Q via the direct communication link with computing node Q.
[0248] In some embodiments, any two computing nodes in a computing node group with more than one computing node in the computing system are connected via a direct communication link. Furthermore, if computing node P and computing node Q are located in the same computing node group, then a direct communication link exists between computing node P and computing node Q. Computing node P sends data to be sent to computing node Q via this direct communication link.
[0249] In other embodiments, there are direct communication links between computing nodes with the same identification information in N computing node groups of the computing system. Furthermore, if computing node P and computing node Q are located in different computing node groups and have the same identification information, then there is a direct communication link between computing node P and computing node Q. Computing node P sends data to be sent to computing node Q through the direct communication link with computing node Q.
[0250] S1303: This corresponds to the absence of a direct communication link between computing node P and computing node Q. Computing node P sends the data to be sent to computing node Q through the same switching node H connected to computing node Q.
[0251] In some embodiments, if there is no direct communication link between computing node P and computing node Q in the computing system, computing node P sends the data to be sent to computing node Q through the same switching node H connected to computing node Q.
[0252] Specifically, computing node P sends the data to be sent to switching node H via a direct communication link with switching node H. After receiving the data to be sent, switching node H sends the received data to computing node Q via a direct communication link with computing node Q.
[0253] It is understood that the communication method shown in Figure 13 is only illustrated using computing node P and computing node Q as an example. Any two computing nodes in the computing system provided in this application embodiment can communicate (e.g., data synchronization) using the communication method shown in Figure 13.
[0254] It is understandable that data transmission between computing node P and computing node Q can be carried out through a direct communication link between computing nodes P and Q, or through a single forwarding via a switching node H connected to computing nodes P and Q. That is, data transmission between computing nodes P and Q requires at most one forwarding via a switching node connected to computing nodes P and Q, resulting in low communication latency.
[0255] It is understandable that, due to the low communication latency of direct communication links between computing nodes, in the communication method shown in Figure 13 above, when there is a direct communication link between computing node P and computing node Q, the direct communication link between computing nodes is preferred for communication.
[0256] In other embodiments, if the amount of data to be transmitted between computing node P and computing node Q is large, and computing node P and computing node Q have both a direct communication link and are jointly connected to switching node H, in order to improve data transmission efficiency, data transmission between computing node P and computing node Q can be carried out through both the direct communication link between computing node P and computing node Q and the switching node H to which computing node P and computing node Q are jointly connected.
[0257] For example, Figure 14 illustrates a flowchart of another communication method according to an embodiment of this application. It can be understood that the communication method shown in Figure 14 can be applied to computing systems having the network structures provided in the above embodiments.
[0258] As shown in Figure 14, the method includes:
[0259] S1401: Computing node P detects a send request, which instructs computing node P to send the data to be sent to computing node Q.
[0260] Specifically, referring to the relevant description of S1301 in Figure 13 above, this application will not repeat it here.
[0261] It is understandable that before sending data to computing node Q, computing node P can first determine whether there is a direct communication link between computing node P and computing node Q.
[0262] If the amount of data to be transmitted between computing node P and computing node Q is large, and there is a direct communication link between computing node P and computing node Q (such as the aforementioned fifth or eighth direct communication link), then computing node P can communicate with computing node Q through the direct communication link with computing node Q and the switching node H.
[0263] If the amount of data to be transmitted between computing node P and computing node Q is large, and there is a direct communication link between computing node P and computing node Q (such as the aforementioned fifth or eighth direct communication link), and the available bandwidth of the direct communication link between computing node P and computing node Q is small, then computing node P can communicate with computing node Q through the direct communication link between computing node P and computing node Q and the switching node H.
[0264] If there is no direct communication link between computing node P and computing node Q, computing node P can communicate with computing node Q through the switching node H, which is connected to computing node Q.
[0265] S1402: Computing node P sends the data to be sent to computing node Q through the same switching node H connected to computing node Q.
[0266] In some embodiments, if there is no direct communication link between computing node P and computing node Q in the computing system, computing node P sends the data to be sent to computing node Q through the same switching node H connected to computing node Q.
[0267] Specifically, computing node P sends the data to be sent to switching node H via a direct communication link with switching node H. After receiving the data to be sent, switching node H sends the received data to computing node Q via a direct communication link with computing node Q.
[0268] S1403: Computing node P splits the data to be sent into a first sub-data and a second sub-data.
[0269] It is understandable that since computing nodes P and Q in the computing system are connected to the same switching node H, if there is a direct communication link between computing nodes P and Q, then the first node and the second node can communicate either through the direct communication link or through the switching node H that computing nodes P and Q are connected to.
[0270] Then, computing node P can split the data to be sent into a first sub-data and a second sub-data. The first sub-data is sent to computing node Q via the switching node H, which is jointly connected to computing node Q. The second sub-data is sent to computing node Q via the direct communication link between computing node P and computing node Q.
[0271] In some embodiments, the splitting of the data to be sent by computing node P into a first sub-data and a second sub-data can be determined based on the bandwidth of the direct communication link between computing nodes and the bandwidth of the direct communication link between computing node P and switching node H. For example, the ratio of the data size of the first sub-data to the data size of the second sub-data can be the ratio of the bandwidth of the direct communication link between computing nodes and the bandwidth of the direct communication link between computing node P and switching node H. For example, if the bandwidth of the direct communication link between computing node P and computing node Q is 400 gigabits per second (Gbps), and the bandwidth of the direct communication link between computing node P and switching node H, as well as the bandwidth of the communication link between computing node Q and switching node H, are also 400 Gbps, then the data size of the first sub-data and the data size of the second sub-data can be equal.
[0272] In other embodiments, the splitting of the data to be sent by computing node P into a first sub-data and a second sub-data can be determined based on the available bandwidth of the direct communication links between computing nodes and the available bandwidth of the direct communication link between computing node P and switching node H. For example, the ratio of the data size of the first sub-data to the data size of the second sub-data can be the ratio of the available bandwidth of the direct communication links between computing nodes and the available bandwidth of the direct communication link between computing node P and switching node H. For example, if the bandwidth of the direct communication link between computing node P and computing node Q is 100 gigabits per second (Gbps), and the bandwidth of the direct communication link between computing node P and switching node H, and the bandwidth of the communication link between computing node Q and switching node H are 400 Gbps, then the ratio of the data size of the first sub-data to the data size of the second sub-data can be 1:4.
[0273] It is understood that, due to the low communication latency of direct communication links between computing nodes, in some embodiments, the splitting of the data to be sent by computing node P into a first sub-data and a second sub-data can also be determined based on the available bandwidth of the direct communication link between computing node P and computing node Q. For example, if the available bandwidth of the direct communication link between computing node P and computing node Q is equal to the bandwidth of the direct communication link between computing node P and computing node Q, then the data to be sent is not split, and the data to be sent is sent directly through the direct communication link between computing node P and computing node Q. If the available bandwidth of the direct communication link between computing node P and computing node Q is greater than 60% of the bandwidth of the direct communication link between computing node P and computing node Q, and less than or equal to 80% of the bandwidth of the direct communication link between computing node P and computing node Q, then 80% of the data to be sent is sent through the direct communication link between computing node P and computing node Q. This application does not impose any limitations on this.
[0274] S1404: Computing node P sends the first sub-data to computing node Q through the direct communication link between computing node P and computing node Q, and sends the second sub-data to computing node Q through the same switching node H connected to computing node Q.
[0275] In some embodiments, a direct communication link exists between any two computing nodes in the computing system. Computing node P sends first sub-data to computing node Q via the direct communication link with computing node Q, and sends second sub-data to computing node Q via the same switching node H connected to computing node Q.
[0276] In some embodiments, any two computing nodes in a computing node group with more than one computing node in the computing system are connected via a direct communication link. Furthermore, if computing node P and computing node Q are located in the same computing node group, then a direct communication link exists between computing node P and computing node Q. Computing node P sends first sub-data to computing node Q via the direct communication link with computing node Q, and sends second sub-data to computing node Q via the same switching node H connected to computing node Q.
[0277] In other embodiments, there are direct communication links between computing nodes with the same identification information in N computing node groups of the computing system. Furthermore, if computing node P and computing node Q are located in different computing node groups and have the same identification information, then there is a direct communication link between computing node P and computing node Q. Computing node P sends first sub-data to computing node Q through the direct communication link with computing node Q, and sends second sub-data to computing node Q through the same switching node H connected to computing node Q.
[0278] The method by which computing node P sends the second sub-data to computing node Q through the same exchange node H connected to computing node Q can be referred to the description in S1403 above, and will not be repeated here.
[0279] It is understandable that when the amount of data to be transmitted between two computing nodes is large, if there is a direct communication link between computing nodes P and Q, the data can be split into two parts and transmitted separately through the direct communication link between computing nodes P and Q and the switching node H that is jointly linked by computing nodes P and Q. This can improve the data transmission efficiency.
[0280] It is understandable that when the amount of data to be transmitted between computing node P and computing node Q is large, if there is a direct communication link between computing node P and computing node Q and the available bandwidth of the direct communication link is limited, the data can be split into two parts and transmitted through the direct communication link between computing node P and computing node Q, and through the switching node H that is jointly connected to computing node P and computing node Q, respectively. This can improve the data transmission efficiency.
[0281] If there is sufficient available bandwidth in the direct communication link between computing node P and computing node Q, data can also be transmitted solely through the direct communication link between computing node P and computing node Q. This application does not impose any limitations on this.
[0282] It is understandable that data transmission between computing node P and computing node Q can be carried out through a direct communication link between computing nodes P and Q, or through a forwarding process via a switching node H connected to computing nodes P and Q. That is, data transmission between computing nodes P and Q requires at most one forwarding process via a switching node connected to computing nodes P and Q, resulting in low communication latency.
[0283] It can be understood that if M computing nodes in a computing system are divided into N computing node groups, and each computing node group includes at least one computing node, then the M computing nodes in the computing system can be encoded based on the computing node group to which the computing node belongs and its position within that group. For example, taking the computing node group to which a computing node belongs as the first dimension and the position of the computing node within its group as the second dimension, each computing node can obtain a unique position code (a, b). Here, a represents the computing node group to which the computing node belongs, and b represents the position of the computing node within its group.
[0284] For example, referring to the network architecture shown in Figure 12 above, the position code of computing node 1-D1 in computing node group 1 can be (1, 1), the position code of computing node 1-D2 in computing node group 1 can be (1, 2), and so on, the position code of computing node 4-D8 in computing node group 4 can be (4, 8).
[0285] If there is a direct communication link between any two computing nodes in the computing node group, it means that there is a direct communication link between any two computing nodes with the same first dimension.
[0286] If there is a direct communication link between computing nodes with the same identification information in each computing node group, it means that there is a direct communication link between any two computing nodes with the same second dimension.
[0287] Therefore, in the communication methods shown in Figures 13 and 14 above, it can be determined whether there is a direct communication link between the first computing node and the second computing node based on the location code of the first computing node and the location code of the second computing node.
[0288] For example, if the location code of the first computing node and the location code of the second computing node have the same first or second dimension, it indicates that there is a direct communication link between the first computing node and the second computing node.
[0289] If the location codes of the first computing node and the second computing node are different in both the first and second dimensions, it means that there is no direct communication link between the first computing node and the second computing node.
[0290] For example, taking the computing system shown in Figure 12 as an example, the communication method applied to the computing system shown in Figure 12 can refer to the logic diagram of the communication method shown in Figure 15.
[0291] As shown in Figure 15, the communication methods include:
[0292] S1501: Determine whether the location code of the target computing node and the location code of the local computing node have the same dimension.
[0293] It can be understood that the local computing node can be computing node P mentioned above, and the target computing node can be computing node Q mentioned above. The data packet to be sent includes the location code of the target computing node.
[0294] In some embodiments, if the determination result is yes, it means that the location code of the target computing node and the location code of the local computing node have the same value in the first dimension and / or the same value in the second dimension. This indicates that there is a direct communication link between the local computing node and the target computing node, or that the local computing node is the target computing node. Then, S1502 is executed, and the data to be sent is sent to the target computing node using a fully connected routing algorithm.
[0295] In other embodiments, if the determination result is negative, it indicates that the values of the first dimension and the second dimension of the location code of the target computing node and the location code of the local computing node are different, which means that there is no direct communication link between the local computing node and the target computing node. Then, step S1503 is executed, and the data to be sent is sent to the target computing node using the CLOS routing algorithm.
[0296] S1502: Use a fully connected routing algorithm to send the data to be sent to the target computing node.
[0297] In some embodiments, if the location code of the target computing node and the location code of the local computing node have the same value in the first dimension and / or the same value in the second dimension, then based on the location code of the target computing node and the location code of the local computing node, a direct communication link between the local computing node and the target computing node is determined, and the data to be sent is sent to the target computing node.
[0298] Specifically, including:
[0299] S1502A: If the location code of the target computing node and the location code of the local computing node have the same value in the first dimension but different values in the second dimension, the data to be sent will be sent to the target computing node through the corresponding direct communication link.
[0300] It is understandable that if the location code of the target computing node and the location code of the local computing node have the same value in the first dimension but different value in the second dimension, it means that the target computing node and the local computing node are located in the same computing node group. In this case, the data to be sent will be sent to the target computing node through the direct communication link between the local computing node and the target computing node.
[0301] S1502B: If the location code of the target computing node and the location code of the local computing node have different values in the first dimension but the same value in the second dimension, the data to be sent will be sent to the target computing node through the corresponding direct communication link.
[0302] It is understandable that if the location code of the target computing node and the location code of the local computing node have different values in the first dimension but the same value in the second dimension, it means that the target computing node and the local computing node are located in different computing node groups and have the same identification information. In this case, the data to be sent will be sent to the target computing node through the direct communication link between the local computing node and the target computing node.
[0303] S1502C: If the location code of the target computing node and the location code of the local computing node have the same values in the first and second dimensions, receive the data to be sent.
[0304] It is understandable that if the location code of the target computing node and the location code of the local computing node have the same values in the first and second dimensions, it means that the local computing node is the target computing node, and the local computing node receives the target data.
[0305] S1503: Use the CLOS routing algorithm to send the data to be sent to the target computing node.
[0306] In some embodiments, if the values of the first dimension and the second dimension of the location code of the target computing node and the location code of the local computing node are different, then based on the values of the first dimension and the second dimension of the location code of the target computing node and the location code of the local computing node, the exchange node that the local computing node and the target computing node are connected to is determined, and the data to be sent is sent to the target computing node through the exchange node.
[0307] Specifically, including:
[0308] S1503A: The local computing node determines the corresponding exchange node based on the value of the first dimension of the location code of the target computing node of the data to be sent, and sends the data to be sent to the corresponding exchange node.
[0309] S1503B: The switching node determines the corresponding target computing node based on the value of the second dimension of the location code of the target computing node of the data to be sent, and sends the data to be sent to the corresponding target computing node.
[0310] In summary, since low-port switching nodes with fewer ports employ the aforementioned cross-connect matrix architecture, while large-port switching nodes with more ports employ an architecture combining multiple cross-connect matrix modules, the communication latency of low-port switching nodes with fewer ports is lower than that of large-port switching nodes with more ports. Compared to the aforementioned approach using large-port switching nodes, the data transmission latency between computing nodes in the computing system of this embodiment is lower. Compared to the aforementioned approach using switching nodes to construct a fat-tree network structure, data transmission between computing nodes in the computing system of this embodiment only requires one forwarding by the switching node, resulting in lower data transmission latency. Furthermore, since low-port switching nodes are less expensive than large-port switching nodes, the computing system of this embodiment can reduce network costs to a certain extent.
[0311] Furthermore, the communication latency between computing nodes in the communication method applied to the computing system provided in the embodiments of this application is low.
[0312] The methods of the embodiments of this application have been described in detail above. In order to facilitate better implementation of the above-described solutions of the embodiments of this application, relevant equipment for cooperating in implementing the above solutions is also provided below.
[0313] The software framework of the computing system is described below with reference to Figure 16. The layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. For example, as shown in Figure 16, this includes a distributed parallel framework, a training framework, and a library of ensemble communication operators.
[0314] Distributed parallel frameworks are used to enable computing systems to distribute computational tasks to multiple computing nodes for parallel processing. Distributed parallel frameworks can include the large-scale distributed training framework Megatron, which achieves efficient model training by combining data parallelism, tensor parallelism, and pipelined parallelism. Distributed parallel frameworks can also include the large model acceleration library Mindspeed, which supports various parallel strategies such as model parallelism, expert parallelism, and sequence parallelism.
[0315] Training frameworks are used to complete computational tasks, and can include PyTorch, TensorFlow, and MindSpore, among others.
[0316] Collective communication operator libraries implement collective communication between computing nodes. Specifically, these may include high-performance collective communication libraries such as HCCL (Huawei collective communication library), NCCL (NVIDIA collective communications library), and TCCL (Tencent collective communication library). Any of these collective communication libraries can be used to implement the communication methods in the embodiments of Figures 13, 14, and 15, and will not be elaborated further here.
[0317] Figure 17 is a schematic diagram of the structure of an electronic device 100 provided in this application. As shown in Figure 17, the electronic device 100 includes a processor 110, a communication interface 120, and a memory 130. The processor 110, communication interface 120, and memory 130 can be interconnected via an internal bus 140, or they can communicate via wireless transmission or other means. This embodiment of the application takes the connection via bus 140 as an example. Bus 140 can be a PCIe bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. Bus 140 can be divided into an address bus, a data bus, a control bus, etc. In addition to the data bus, bus 140 can also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus 140 in the figure.
[0318] Processor 110 may consist of one or more processors, such as a CPU, a combination of a CPU and hardware chips. The hardware chips may be application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. The PLDs may be complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), generic array logic (GALs), or any combination thereof. Processor 110 executes various types of digital storage instructions, such as digital storage instructions in software or firmware stored in memory 130, enabling electronic device 100 to provide a variety of services.
[0319] The memory 130 stores program code, which is executed under the control of the processor 110 to perform the processing steps of the communication method in the above embodiments. The program code may include one or more software modules that can execute the processing steps of the communication method in the above embodiments.
[0320] The memory 130 may include volatile memory, such as random access memory (RAM); the memory 130 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); the memory 130 may also include combinations of the above types. The memory 130 may store program code, specifically used to execute the processing steps of the communication method in the above embodiments, which will not be described in detail here.
[0321] The communication interface 120 can be an internal interface (such as a high-speed serial computer expansion bus), a wired interface (such as an Ethernet interface), or a wireless interface (such as a cellular network interface or a wireless LAN interface) for communicating with other devices or modules.
[0322] It should be noted that Figure 17 is only one possible implementation of the embodiment of this application. In actual applications, the electronic device 100 may include more or fewer components, which is not limited here.
[0323] It should be understood that the electronic device 100 shown in Figure 17 can be a standalone physical device, and the electronic device 100 includes one or more of the aforementioned computing nodes, switching nodes or control nodes.
[0324] The electronic device 100 shown in Figure 17 can also be a computer cluster consisting of at least one physical device (e.g., a server). The aforementioned computing nodes, switching nodes, and control nodes can be independent physical devices or logical units within one or more physical devices. This application does not impose any specific limitations on these nodes.
[0325] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more sets of available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs), or semiconductor media. Semiconductor media can be SSDs.
[0326] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A computing system, characterized in that, The computing system includes M computing nodes and X switching nodes, wherein the number of ports of each switching node is greater than 0 and less than M, where M is an integer greater than 2 and X is an integer greater than 1. Any two computing nodes among the M computing nodes are connected to the same switching node via a direct communication link, wherein the first computing node among the M computing nodes is connected to the first switching node among the X switching nodes via a first direct communication link, and the second computing node among the M computing nodes is connected to the first switching node via a second direct communication link; and... The first computing node and the second computing node communicate with each other through the first direct communication link, the first switching node, and the second direct communication link.
2. The computing system according to claim 1, characterized in that, K of the M computing nodes are connected to each other via direct communication links, where K is an integer greater than 1 and less than or equal to M. The third computing node among the K computing nodes is connected to the second switching node among the X switching nodes via a third direct communication link; the fourth computing node among the K computing nodes is connected to the second switching node via a fourth direct communication link; and the third computing node and the fourth computing node are connected via a fifth direct communication link. The third computing node and the fourth computing node communicate via the third direct communication link, the second switching node, and the fourth direct communication link, and / or... The third computing node and the fourth computing node communicate with each other through the fifth direct communication link.
3. The computing system according to claim 1, characterized in that, The M computing nodes correspond to N computing node groups, and each computing node group includes at least one computing node from the M computing nodes. Computing nodes in any two of the N computing node groups are connected to the same exchange node.
4. The computing system according to claim 3, characterized in that, The computing nodes in the first and second computing node groups of the N computing node groups are all connected to at least two switching nodes via direct communication links, and the X switching nodes include the at least two switching nodes.
5. The computing system according to claim 3 or 4, characterized in that, The computing nodes in each of the Y computing node groups are connected to each other through direct communication links, where Y is an integer greater than 0 and less than or equal to N; And / or, In the N computing node groups, computing nodes with the same identification information are connected to each other through direct communication links.
6. The computing system according to claim 5, characterized in that, The fifth computing node among the M computing nodes is connected to the third switching node among the X switching nodes via a sixth direct communication link; the sixth computing node among the M computing nodes is connected to the third switching node via a seventh direct communication link; and the fifth computing node and the sixth computing node are connected via an eighth direct communication link. The fifth and sixth computing nodes are computing nodes within one of the N computing node groups, or the fifth and sixth computing nodes are computing nodes with the same identification information in the N computing node groups; and... The fifth computing node and the sixth computing node communicate via the sixth direct communication link, the third switching node, and the seventh direct communication link, and / or, The fifth computing node and the sixth computing node communicate through the eighth direct communication link.
7. The computing system according to any one of claims 3 to 6, characterized in that, In the N computing node groups, the computing nodes in each computing node group are connected to the same one or more control nodes.
8. A communication method, characterized in that, Applied to computing systems The computing system includes M computing nodes and X switching nodes, wherein the number of ports of each switching node is greater than 0 and less than M, where M is an integer greater than 2 and X is an integer greater than 1. Any two computing nodes among the M computing nodes are connected to the same switching node via a direct communication link, wherein the first computing node among the M computing nodes is connected to the first switching node among the X switching nodes via a first direct communication link, and the second computing node among the M computing nodes is connected to the first switching node via a second direct communication link; and... The method includes: The first computing node detects a first request, which instructs the first computing node to send first data to the second computing node; The first computing node sends the first data to the second computing node through the first direct communication link, the first switching node, and the second direct communication link.
9. The method according to claim 8, characterized in that, K of the M computing nodes are connected in pairs via direct communication links, where K is an integer greater than 1 and less than or equal to M. The third computing node among the K computing nodes is connected to the second switching node among the X switching nodes via a third direct communication link, and the fourth computing node among the K computing nodes is connected to the second switching node via a fourth direct communication link. The third computing node and the fourth computing node are connected via a fifth direct communication link. and, The method further includes: The third computing node detects the second request, which instructs the third computing node to send the second data to the fourth computing node. The third computing node sends the second data to the fourth computing node via the third direct communication link, the second switching node, and the fourth direct communication link; or... The third computing node sends the second data to the fourth computing node via the fifth direct communication link; or... The third computing node sends a portion of the second data to the fourth computing node through the third direct communication link, the second switching node, and the fourth direct communication link, and sends another portion of the second data to the fourth computing node through the fifth direct communication link.
10. The method according to claim 8, characterized in that, The M computing nodes correspond to N computing node groups, and each computing node group includes at least one computing node from the M computing nodes. The computing nodes in any two computing node groups in the N computing node groups are connected to the same switching node. In the N computing node groups, the computing nodes within each of the Y computing node groups are connected to each other via direct communication links, where Y is an integer greater than 0 and less than or equal to N; and / or, In the N computing node groups, computing nodes with the same identification information are connected to each other via direct communication links. The fifth computing node among the M computing nodes is connected to the third switching node among the X switching nodes via a sixth direct communication link; the sixth computing node among the M computing nodes is connected to the third switching node via a seventh direct communication link; and the fifth computing node and the sixth computing node are connected via an eighth direct communication link. The method further includes: The fifth computing node among the M computing nodes is connected to the third switching node among the X switching nodes via a sixth direct communication link, the sixth computing node among the M computing nodes is connected to the third switching node via a seventh direct communication link, and the fifth computing node and the sixth computing node are connected via an eighth direct communication link. The fifth computing node detects a third request, which instructs the fifth computing node to send third data to the sixth computing node. The fifth computing node sends the third data to the sixth computing node via the sixth direct communication link, the third switching node, and the seventh direct communication link; or... The fifth computing node sends the third data to the sixth computing node via the eighth direct communication link; or... The fifth computing node sends a portion of the third data to the sixth computing node via the sixth direct communication link, the third switching node, and the seventh direct communication link, and sends another portion of the third data to the sixth computing node via the eighth direct communication link.
11. The method according to claim 10, characterized in that, The fifth computing node and the sixth computing node are computing nodes within one of the Y computing node groups; or... The fifth computing node and the sixth computing node are computing nodes with the same identification information in the N computing node group.