Collective communication method and apparatus, and electronic device
Patent Information
- Application Number
- PCT/CN2025/138666
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2025-11-28
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025138666_01102026_PF_FP_ABST
Abstract
Description
A collection of communication methods, devices and electronic equipment
[0001] This application claims priority to patent application No. 202510388105.7 filed in China on March 28, 2025, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of communications, and more particularly to a combined communication method, apparatus, and electronic device. Background Technology
[0003] Collective communication refers to the collaborative operation among multiple processes or nodes in a distributed system to exchange global data. It is commonly used in high-performance computing and distributed training. The core operations of collective communication include allreduce, reduce-scatter, and all-gather. Allreduce aggregates the input data from each node (e.g., summing or finding the maximum value) and then distributes the result to all nodes. Reduce-scatter aggregates the input data from each node and distributes the result to each node according to its node number (each node receives only a portion of the result). All-gather concatenates the input data from each node in node number order and then broadcasts the complete data to all nodes.
[0004] In large-scale distributed computing scenarios, the allreduce operation is crucial. For a one-dimensional fully-mesh topology, the allreduce operation can be accomplished by combining reduce-scatter and all-gather operations. However, as the scale of computing clusters increases and application scenarios become more complex, multi-dimensional fully-mesh topologies are more widely used due to their lower cost and higher performance.
[0005] Multidimensional fully connected topologies are more complex, with more diverse communication paths between nodes. Therefore, for multidimensional fully connected topologies, how to fully utilize the bandwidth of each dimension to achieve allreduce operations with shorter communication times and higher efficiency has become an urgent problem to be solved. Summary of the Invention
[0006] This application provides a collection communication method, device, and electronic device that can effectively utilize the bandwidth of each dimension in a multidimensional fully connected topology to achieve the allreduce operation of the multidimensional fully connected topology with shorter communication time and higher efficiency.
[0007] Firstly, a collection communication method is provided, comprising: dividing the input data of each node in a first network into n data blocks. The topology of the first network is an n-dimensional fully connected topology, where n is greater than or equal to 2. After performing reduction and distribution operations on the i-th data block of each node by traversing each dimension in the i-th order, a global collection operation is then performed by traversing each dimension in the reverse order of the i-th order. Here, the value of i ranges from 1 to n. The i-th order indicates the order between the dimensions.
[0008] Based on this scheme, the global reduction operation of a multidimensional fully connected topology (i.e., an n-dimensional fully connected topology) is decomposed into a reduction distribution operation and a global collection operation performed separately in different dimensions. This reduces the number of hops in cross-dimensional communication and lowers communication latency. Furthermore, since the reduction distribution operation and the global collection operation are performed separately in different dimensions, it helps to shorten communication time and improve bandwidth utilization and communication efficiency.
[0009] Combining the aggregated communication method provided in the first aspect, in some possible implementations, the reduction and distribution operation of the p-th data block of each node in the q-th dimension is executed concurrently with the reduction and distribution operation of the k-th data block of each node in the l-th dimension. p, q, k, and l are all arbitrary numbers from 1 to n, where p is different from k, and q is different from l. Based on this scheme, by executing the reduction and distribution operations of multiple data blocks in parallel in different dimensions, the physical link resources of the multidimensional topology can be fully utilized, bandwidth utilization can be improved, and overall latency can be reduced.
[0010] Combining the aggregated communication method provided in the first aspect, in some possible implementations, the global collection operation of the p-th data block of each node in the q-th dimension is executed concurrently with the global collection operation of the k-th data block of each node in the l-th dimension. Based on this scheme, by executing the global collection operations of multiple data blocks in parallel in different dimensions, the physical link resources of the multidimensional topology can be fully utilized, bandwidth utilization can be improved, and overall latency can be reduced.
[0011] Combining the set communication method provided in the first aspect, in some possible implementations, the first order is from the first dimension to the nth dimension. The second order is from the second dimension to the nth dimension, and then back to the first dimension. When i is greater than or equal to 3, the i-th order is from the i-th dimension to the nth dimension, and then from the first dimension to the (i-1)-th dimension. Based on this scheme, the load distribution across dimensions can be balanced through a cyclic offset dimension traversal strategy, avoiding link traffic conflicts, and adapting to the path diversity of high-dimensional topologies.
[0012] Combining the set communication method provided in the first aspect, in some possible implementations, the n data blocks have the same data size. Based on this scheme, it is beneficial to avoid the synchronization waiting overhead caused by data skew, and to achieve efficient communication scheduling.
[0013] Combining the collection communication method provided in the first aspect, in some possible implementations, the segmentation method determines the time required for each reduction distribution operation and each global collection operation based on the input data of each node. Based on this scheme, it is beneficial to reduce the overall time consumption of the global reduction operation and improve communication efficiency.
[0014] Combining the aggregated communication method provided in the first aspect, in some possible implementations, the traversal of each dimension is performed under the premise of no link traffic conflict. Based on this scheme, it is beneficial to improve bandwidth utilization.
[0015] Combining the set communication method provided in the first aspect, in some possible implementations, n is 2, and the n-dimensional fully connected topology is a 2-dimensional fully connected topology, where the 2 dimensions include the first and second dimensions. After performing reduction and distribution operations on the i-th data block of each node in the i-th order across each dimension, a global collection operation is performed in the reverse order across each dimension. This includes: performing reduction and distribution operations on the first data block of each node sequentially in the first and second dimensions, followed by a global collection operation sequentially in the second and first dimensions. Similarly, performing reduction and distribution operations on the second data block of each node sequentially in the second and first dimensions, followed by a global collection operation sequentially in the first and second dimensions.
[0016] Secondly, a collective communication device is provided, comprising: a first module, used to divide the input data of each node in a first network into n data blocks. The topology of the first network is an n-dimensional fully connected topology. n is greater than or equal to 2. A second module is used to perform reduction and distribution operations on the i-th data block of each node by traversing each dimension in the i-th order, and then perform a global collection operation by traversing each dimension in the reverse order of the i-th order. Here, the value of i ranges from 1 to n. The i-th order indicates the order between the dimensions.
[0017] Thirdly, a collective communication device is provided, comprising multiple interacting modules for implementing the method as described in any of the first aspects.
[0018] Fourthly, an electronic device is provided, comprising one or more processors. The one or more processors are configured to execute computer programs or instructions to implement the method of any of the first aspects.
[0019] Fifthly, an electronic device is provided, characterized in that it includes a memory and one or more processors. The memory is used to store computer programs or instructions. The one or more processors are used to execute the computer programs or instructions in the memory, causing the electronic device to perform the methods as described in any of the first aspects.
[0020] In a sixth aspect, a computer-readable storage medium is provided, comprising a computer program or instructions that, when executed, cause the method of any one of the first aspects to be implemented.
[0021] In a seventh aspect, a computer program product is provided, comprising a computer program or instructions that, when executed, cause the method of any one of the first aspects to be implemented.
[0022] Eighthly, a chip device is provided, including a processor and a memory. The processor is used to invoke a computer program or computer instructions stored in the memory to cause the processor to execute any of the implementations described in the first aspect. Optionally, the processor is coupled to the memory via an interface.
[0023] It should be understood that the second to eighth aspects of this application are consistent with or correspond to the technical solutions of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description
[0024] Figure 1 is a schematic diagram of a one-dimensional fully connected topology;
[0025] Figure 2 is a schematic diagram of a two-dimensional fully connected topology;
[0026] Figure 3 is a schematic diagram of a three-dimensional fully connected topology;
[0027] Figure 4 is a schematic diagram of a global reduction operation;
[0028] Figure 5 is a schematic diagram of a protocol distribution operation;
[0029] Figure 6 is a schematic diagram of a global collection operation;
[0030] Figure 7 is a schematic diagram of another global reduction operation;
[0031] Figure 8 is a schematic diagram of a computing cluster architecture provided in an embodiment of this application;
[0032] Figure 9 is a flowchart illustrating a collection communication method provided in an embodiment of this application;
[0033] Figure 10 is a schematic diagram of another two-dimensional fully connected topology provided in an embodiment of this application;
[0034] Figure 11 is a schematic diagram of a communication calculation process for data A provided in an embodiment of this application;
[0035] Figure 12 is a schematic diagram of a communication calculation process for B data provided in an embodiment of this application;
[0036] Figure 13 is a schematic diagram illustrating the time consumption of each step in an embodiment of this application;
[0037] Figure 14 is a schematic diagram of a data change provided in an embodiment of this application;
[0038] Figure 15 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0039] Figure 16 is a schematic diagram of a collection communication module provided in an embodiment of this application. Detailed Implementation
[0040] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0041] References to "one embodiment" or "some embodiments" as described in this application mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0042] Furthermore, in the embodiments of this application, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as an "example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the term "example" is intended to present concepts in a concrete manner. In the embodiments of this application, "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably, and it should be noted that their intended meanings are consistent unless their distinction is emphasized.
[0043] The application scenarios described in this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.
[0044] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0045] To facilitate understanding, some technical terms involved in the embodiments of this application will be introduced below.
[0046] One-dimensional fully connected topology: Also known as one-dimensional full-mesh topology, one-dimensional fully interconnected topology, etc., this refers to a network topology where every node is directly connected to every other node. The connections between nodes can be transmission links such as cables or fiber optics, and can be one or multiple physical cables. The bandwidth of the connections between nodes can be equal or unequal. It should be understood that in a one-dimensional fully connected topology with N nodes, with only one cable connecting each node, each node is connected to other nodes through N-1 cables.
[0047] For example, as shown in Figure 1, in a one-dimensional fully connected topology containing 5 nodes (represented by node in Figure 1), each node is directly connected to any other node in the network, and the topology includes a total of 10 cables.
[0048] Two-dimensional fully connected topology: This consists of multiple one-dimensional fully connected topologies. Nodes are grouped according to a two-dimensional matrix, with each dimension corresponding to an independent direction (e.g., the X-axis is horizontal, the Y-axis is vertical). Within each dimension, groups of nodes in the same direction (nodes in the same row, nodes in the same column) constitute a one-dimensional fully connected topology. Two-dimensional fully connected topologies have the following characteristics: In the first dimension (e.g., the X-axis direction), there are M N*1 one-dimensional fully connected topologies (M and N can be equal or unequal); in the second dimension (e.g., the Y-axis direction), there are N M*1 one-dimensional fully connected topologies. In other words, M*N nodes are divided into N groups, forming N M*1 one-dimensional fully connected topologies in the Y-axis direction and M N*1 one-dimensional fully connected topologies in the X-axis direction.
[0049] For example, as shown in Figure 2, a two-dimensional fully connected topology containing 4*4 nodes (including 16 nodes, node1 to node16) includes four 4*1 one-dimensional fully connected topologies in the X-axis direction and four 4*1 one-dimensional fully connected topologies in the Y-axis direction. It should be understood that in a two-dimensional fully connected topology, any two nodes in different dimensions are not necessarily directly connected. For example, the node in the first row and first column of Figure 2 may not be directly connected to the node in the second row and second column.
[0050] A 3D fully connected topology is composed of multiple 1D fully connected topologies. Nodes are grouped according to a 3D cube, with each dimension corresponding to an independent direction. Within each dimension, groups of nodes in the same direction constitute a 1D fully connected topology. A 3D fully connected topology has the following characteristics: M*N*K nodes distributed along the direction of N nodes (called the X-axis) comprise M*K N*1 1D fully connected topologies; the direction of M nodes distributed along the direction of M nodes (called the Y-axis) comprises N*K M*1 1D fully connected topologies; and the direction of K nodes distributed along the direction of K nodes (called the Z-axis) comprises N*M K*1 1D fully connected topologies (N, M, and K can be equal or unequal).
[0051] For example, as shown in Figure 3, a 3D fully connected topology containing 3*3*3 nodes includes 3*3 three-dimensional fully connected topologies along the X-axis, 3*3 three-dimensional fully connected topologies along the Y-axis, and 3*3 three-dimensional fully connected topologies along the Z-axis (for clarity, only two three-dimensional fully connected topologies along the Z-axis are shown in Figure 3, omitting the other seven). It should be understood that, similar to 2D fully connected topologies, any two nodes in different dimensions of a 3D fully connected topology are not necessarily directly connected.
[0052] An n-dimensional fully connected topology consists of multiple one-dimensional fully connected topologies. Nodes are grouped according to n dimensions, each corresponding to an independent direction. Within each dimension, groups of nodes in the same direction form a one-dimensional fully connected topology. An n-dimensional fully connected topology has the following characteristics: N*M*K*...*S nodes form an n-dimensional full-mesh. Each node provides (N-1), (M-1), (K-1), ..., (S-1) logical ports (each logical port can be one or more physical ports) in each of the n dimensions, forming multiple groups of one-dimensional full-mesh consisting of N nodes, M nodes, K nodes, ..., S nodes. The n numbers N, M, K, ..., S can be equal or unequal. Similarly, any two nodes in different dimensions of an n-dimensional fully connected topology are not necessarily directly connected, which will not be elaborated further here.
[0053] Global reduction refers to aggregating (e.g., adding) the input data of all process IDs (ranks, also known as process identifiers) and then distributing the aggregation result to the outputs of all ranks. Here, a rank identifies an independent execution unit in a distributed system, such as a process, graphics processing unit (GPU), or node. Each rank has independent input and output data and collaborates with other ranks through communication operations (such as global reduction operations).
[0054] For example, please refer to Figure 4, which is a schematic diagram of a global reduction operation. As shown in Figure 4, the one-dimensional fully connected topology includes nodes 0, 1, 2, and 3. The input of node 0 is in0, the input of node 1 is in1, the input of node 2 is in2, and the input of node 3 is in3. After performing a global reduction operation on this one-dimensional fully connected topology (assuming the above aggregation is addition), the output of each node is in0 + in1 + in2 + in3.
[0055] Reduced distribution refers to aggregating all rank input data and then distributing the aggregation results to the output data of each rank according to their rank number. Each rank receives 1 / ranksize of the data from the other processes and performs reduction. Here, ranksize refers to the total number of ranks involved in the operation, i.e., the total number of processes.
[0056] For example, please refer to Figure 5, which is a schematic diagram of a reduction distribution operation. As shown in Figure 5, the one-dimensional fully connected topology includes nodes 0, 1, 2, and 3. The input of node 0 is in0 (including in01, in02, in03, in04), the input of node 1 is in1 (including in11, in12, in13, in14), the input of node 2 is in2 (including in21, in22, in23, in24), and the input of node 3 is in3 (including in31, in32, in33, in34). After performing a reduction distribution operation on the one-dimensional fully connected topology (assuming the above aggregation is addition), the output of node 0 is in01+in11+in21+in31, the output of node 1 is in02+in12+in22+in32, the output of node 2 is in03+in13+in23+in33, and the output of node 3 is in04+in14+in24+in34.
[0057] Global collection: This refers to concatenating all the input data of all ranks according to the rank number, and then sending the concatenated result to the output of all ranks.
[0058] For example, please refer to Figure 6, which is a schematic diagram of a global collection operation. As shown in Figure 6, the one-dimensional fully connected topology includes nodes 0, 1, 2, and 3. The input of node 0 is in0, the input of node 1 is in1, the input of node 2 is in2, and the input of node 3 is in3. After performing a global collection operation on this one-dimensional fully connected topology, the output of each node is the concatenation result of in1, in2, in3, and in4, which can be written as [in0, in1, in2, in3].
[0059] Based on the above description of the terminology, the background of the proposed embodiments of this application will be explained below.
[0060] In a one-dimensional fully connected topology, the global reduction operation can be implemented step-by-step by a reduction distribution operation and a global collection operation. Taking a one-dimensional fully connected topology with N nodes as an example, the reduction distribution operation can be performed first, which involves unidirectionally transmitting 1 / N of the data on each link, transmitting different data blocks to different nodes. In this way, each node receives a total of (N-1) / N data blocks sent by the other N-1 nodes, plus its local 1 / N data blocks, reducing it to a data block of size 1 / N. Then, the global collection operation is performed, which involves unidirectionally transmitting the reduced 1 / N data blocks on each link, with each node transmitting the same data to different nodes. In this way, each node collects the total of (N-1) / N data blocks from the other N-1 nodes, plus its local 1 / N data blocks, obtaining the data block after the global reduction operation.
[0061] To facilitate understanding, a specific example is provided below. Please refer to Figure 7, which illustrates another global reduction operation. As shown in Figure 7, the one-dimensional fully connected topology includes nodes a, b, c, and d. The input data for node a includes a0, a1, a2, and a3; the input data for node b includes b0, b1, b2, and b3; the input data for node c includes c0, c1, c2, and c3; and the input data for node d includes d0, d1, d2, and d3. The global reduction operation performed on the one-dimensional fully connected topology shown in Figure 7 can be divided into steps S1, S2, S3, and S4. Steps S1 and S2 are reduction distribution operations, and steps S3 and S4 are global collection operations. As shown in Figure 1, in step S1, node a sends a1 to node b, a2 to node c, and a3 to node d. Node b sends b0 to node a, b2 to node c, and b3 to node d. Node c sends c0 to node a, c1 to node b, and c3 to node d. Node d sends d0 to node a, d1 to node b, and d2 to node c. In step S2, each node reduces the received data and its local data to obtain the data after the reduction and distribution operation. In step S3, each node sends the data after the reduction and distribution operation to other nodes. In step S4, each node concatenates the received data with its local data to obtain the data after the global reduction operation.
[0062] The decomposition of the global reduction operation into reduction-distribution and global collection operations is primarily based on considerations of communication complexity. For example, attempting to complete global data aggregation and distribution in a single phase (e.g., each node directly broadcasts its own data and synchronously receives data from others) would result in exponentially increasing communication complexity. Specifically, each node would need to send a complete data block to all other nodes and synchronously receive and aggregate all input data, leading to a total communication complexity of O(N^2). 2 (D) (N is the number of nodes, D is the amount of data). This will not only cause network congestion, but also result in the actual effective throughput being far lower than the theoretical value due to multiple nodes simultaneously competing for link bandwidth.
[0063] The decomposition strategy, employing a reduction and distribution operation followed by a global collection operation, breaks down the problem into two optimization phases using a divide-and-conquer approach. In the reduction and distribution phase, data is fragmented and only undergoes local reduction between nodes, reducing the communication load to O(ND). In the global collection phase, the fragmentation results are broadcast via multicast trees or ring topologies, maintaining a total communication load of O(ND). This phased approach reduces the bandwidth requirement of each phase to within the maximum capacity of a single link, while chunk pipelining overlaps computation and communication. Furthermore, the decomposition strategy naturally adapts to the parallelization of multi-dimensional topologies; different data fragments can be mapped to links of different dimensions for transmission. Leveraging the spatial multiplexing characteristics of multi-dimensional physical links, the time complexity is ultimately reduced from O(ND). 2 The optimization is from D) to O(D / K+logN)O (where K is the number of fragments). This design has been mathematically proven to be the bandwidth-optimal global reduction implementation paradigm.
[0064] The above describes the implementation of global reduction operations in one-dimensional fully connected topologies. With the expansion of networks and the diversification of application scenarios, networking methods are diverse. For example, computer systems such as central processing units (CPUs), graphics processing units (GPUs), neural processing units (NPUs), and extended processing units (XPUs) are interconnected by networks to form large-scale network clusters for AI training and inference clusters, or for scientific computing or massive data analysis and processing. In these application scenarios, multi-dimensional fully connected topologies are widely used due to their significant advantages such as low cost and high performance.
[0065] It should be understood that multidimensional fully connected topologies are more complex, with more diverse communication paths between nodes. Therefore, for multidimensional fully connected topologies, how to fully utilize the bandwidth of each dimension to achieve allreduce operations with shorter communication times and higher efficiency becomes an urgent problem to be solved.
[0066] To address the aforementioned issues, embodiments of this application provide a collective communication method, apparatus, and electronic device capable of performing global reduction operations on multidimensional fully connected topologies (i.e., the aforementioned n-dimensional fully connected topologies), exhibiting high bandwidth utilization, short communication time, and high communication efficiency. The collective communication method, apparatus, and electronic device provided in the embodiments of this application will be described in detail below.
[0067] The aggregation communication method, apparatus, and electronic device provided in this application are applied to a computing cluster architecture. The computing cluster architecture is a multi-layered system designed to provide robust support for high-performance computing and distributed tasks through the collaborative work of each layer. For ease of understanding, the computing cluster architecture is described herein.
[0068] Please refer to Figure 8, which is a schematic diagram of a computing cluster architecture provided in an embodiment of this application. As shown in Figure 8, the computing cluster architecture may include, from bottom to top, a hardware layer 801, an operating system layer 802, a communication layer 803, a software layer 804, and an application layer 805.
[0069] The hardware layer 801 is the physical foundation of the computing cluster architecture, which can include computing nodes (such as CPUs, GPUs, NPUs, XPUs, etc.), network devices (such as switches, routers, etc.), and storage devices (such as hard drives, solid-state storage, etc.). Among them, computing nodes are used to handle the actual computing tasks, network devices are used to handle communication between nodes, and storage devices are used to handle data storage.
[0070] The operating system layer 802 is responsible for managing hardware resources and providing basic system services, such as resource management (managing resources like CPU and memory on computing nodes), process scheduling (scheduling processes on nodes to ensure efficient resource utilization), and file system (providing file storage and access functions). The operating system layer 802 ensures efficient utilization of hardware resources and concurrent execution of multiple tasks.
[0071] The communication layer 803 is responsible for data transmission between nodes, including network interface cards (for implementing the physical connection between nodes and the network), drivers (for providing driver support for network interface cards and ensuring their proper functioning), and communication protocol stacks (such as Transmission Control Protocol (TCP) / Internet Protocol (IP), Remote Direct Memory Access (RDMA), etc., for data transmission and reception). The communication layer 803 ensures efficient and reliable data transmission within the cluster.
[0072] The software layer 804 contains various software components to support cluster management and application development. These components include cluster management software, communication libraries, control plane software, and middleware. The cluster management software is responsible for resource scheduling, task allocation, and monitoring of the entire cluster, ensuring its efficient operation. The communication library provides efficient communication primitives (such as message passing interfaces (MPI)) to support efficient communication between nodes. The control plane software is responsible for topology configuration and management, providing network topology information to the communication library. Middleware provides additional services and support, such as load balancing and fault recovery.
[0073] The application layer 805 contains user applications and business logic processing, directly facing users and providing services such as high-performance computing, distributed training, scientific computing, and data analysis.
[0074] In this embodiment, the aforementioned multidimensional fully connected topology can be implemented in hardware layer 801. The aggregated communication method provided in this embodiment can be applied to (e.g., configured in) a communication library in software layer 804. The specific structure of the multidimensional fully connected topology can be configured by the control plane software to the communication library, or it can be obtained by the communication library itself through topology awareness; no limitation is made here.
[0075] In one possible implementation, the aggregation communication method, apparatus, and electronic device provided in this application embodiment can be applied to distributed parallel computing scenarios in artificial intelligence (AI) and high-performance computing (HPC). In this computing cluster architecture, nodes consist of one or more CPUs / GPUs / NPUs / XPUs or other accelerators or computing units, and all nodes are interconnected through a multi-dimensional fully connected network topology to form a computing cluster. A cluster communication library distributed software (i.e., the aforementioned communication library) is deployed on the nodes. This communication library distributed software manages the node communication tasks in the cluster for data exchange between nodes in distributed computing.
[0076] The specific implementation of the embodiments of this application will be described below.
[0077] Please refer to Figure 9, which is a flowchart illustrating a collective communication method provided in an embodiment of this application. As described in the foregoing embodiments, this method is for networks with an n-dimensional fully connected topology.
[0078] As shown in Figure 9, the method includes the following steps.
[0079] S901. Divide the input data of each node in the first network into n data blocks.
[0080] The first network is an n-dimensional fully connected topology. For an introduction to the n-dimensional fully connected topology, please refer to the aforementioned embodiment; it will not be repeated here.
[0081] In some possible implementations, the input data of each node can be divided into n data blocks. For example, when n is 3 and the input data of the node is [1,2,3,4,5,6], the input data can be divided into the following 3 data blocks: [1,2], [3,4], [5,6].
[0082] In some other possible implementations, the segmentation method may not be equal segmentation. For example, when n is 3 and the input data of the node is [1,2,3,4,5,6], the input data can be segmented into the following 3 data blocks: [1], [2,3,4], [5,6]. As a possible example, when the switching method is not equal segmentation, the segmentation ratio can be determined based on the principle of minimizing the total completion time of the communication operator, without any restrictions here.
[0083] For ease of explanation, the dimension is represented by D, for example, D1 represents the first dimension and D2 represents the second dimension. For instance, in the two-dimensional fully connected topology of the aforementioned embodiments, D1 represents the direction of the X-axis and D2 represents the direction of the Y-axis. In the three-dimensional fully connected topology of the aforementioned embodiments, D1 represents the direction of the X-axis, D2 represents the direction of the Y-axis, and D3 represents the direction of the Z-axis.
[0084] It should be understood that for an n-dimensional fully connected topology network, Di represents the i-th dimension, and the value of i ranges from 1 to n.
[0085] Furthermore, Sj represents the j-th data block after each node has divided the data into n data blocks. It should be noted that in this embodiment, Sj represents a data block, not the specific data content. For example, for a certain n-dimensional fully connected topology, Sj includes the j-th data block of a certain node. The data content of this j-th data block is 1. After reduction, distribution, global collection, and other operations, the data content of the j-th data block of this node becomes 2. In this case, Sj represents the data content 2 in the j-th data block of this node, not 1. Similarly, the value of j ranges from 1 to n.
[0086] It should be noted that the naming order of each dimension and the naming order of each data block after segmentation are not limited in the embodiments of this application. For example, for the two-dimensional fully connected topology in the aforementioned embodiments, D1 can be used to represent the direction of the X-axis and D2 to represent the direction of the Y-axis, or D1 can be used to represent the direction of the Y-axis and D2 to represent the direction of the X-axis. For another example, after the data [1,2,3,4,5,6] is segmented into [1], [2,3,4], and [5,6], [1] can be called the first data block, [2,3,4] can be called the second data block, and [5,6] can be called the third data block, or [1] can be called the second data block, [2,3,4] can be called the third data block, and [5,6] can be called the first data block, etc., without specific limitations.
[0087] Performing an operation on Sj in Di (taking reduction and distribution as an example) means reducing and distributing the data content of the j-th data block of each node in the i-th dimension. For example, taking the two-dimensional fully connected topology shown in Figure 2, where D1 is the X-axis direction and D2 is the Y-axis direction, performing a reduction and distribution operation on S1 in D1 means reducing and distributing the first data block of each node in the X-axis direction. Specifically, this includes reducing and distributing the first data blocks of nodes 1 to 4, nodes 5 to 8, nodes 9 to 12, and nodes 13 to 16.
[0088] S902. After performing reduction and distribution operations on each dimension by traversing the i-th data block of each node in the i-th order, perform global collection operations by traversing each dimension in the reverse order of the i-th order.
[0089] The value of i ranges from 1 to n. The i-th order is used to indicate the order between the various dimensions.
[0090] S902 can also be expressed as follows: after performing reduction and distribution operations on Si by traversing D1 to Dn in the i-th order, global collection operations are then performed by traversing D1 to Dn in the reverse order of the i-th order.
[0091] The first order, the second order, ..., the nth order (hereinafter referred to as the first to the nth order in the embodiments) can be partially the same or completely different, without specific limitations. For example, when the first to the nth order are all different, the first order can be from the first dimension to the nth dimension sequentially. The second order can be from the second dimension to the nth dimension sequentially, then back to the first dimension. The third order can be from the third dimension to the nth dimension sequentially, then back to the second dimension from the first dimension. And so on. When i is greater than or equal to 3, the i-th order can be from the i-th dimension to the nth dimension, then back to the (i-1)-th dimension.
[0092] After the value of i is traversed from 1 to n to complete the global collection operation, the global reduction operation of the n-dimensional fully connected topology can be realized.
[0093] Based on this scheme, the global reduction operation of the multidimensional fully connected topology (i.e., the aforementioned n-dimensional fully connected topology) is decomposed into reduction distribution operations and global collection operations performed separately in different dimensions. This helps reduce the number of hops in cross-dimensional communication and lowers communication latency. In addition, since the reduction distribution operations and global collection operations are performed separately in different dimensions, it helps improve bandwidth utilization and communication efficiency.
[0094] In some possible implementations, the reduction and distribution operations of different data blocks in various dimensions can be executed concurrently, as can the global collection operations in various dimensions. The reduction and distribution operation of the i-th data block in one dimension and the global collection operation of the j-th data block in another dimension can also be executed concurrently (i and j are different), which is not limited here.
[0095] This section provides an illustrative example of the concurrent execution of reduction and distribution operations across various dimensions. For instance, the first order is from dimension 1 to dimension n. The second order is from dimension 2 to dimension n, then back to dimension 1. When i is greater than or equal to 3, the i-th order is from dimension i to dimension n, then from dimension 1 to dimension (i-1). Thus, when reducing and distributing the first data block of each node in dimension 1, the second data block of each node can be reduced and distributed in parallel in dimension 2, the third data block in dimension 3, and the nth data block in dimension n. Then, the first data block of each node can be reduced and distributed in parallel in dimension 2, the second data block in dimension 3, the third data block in dimension 4, and so on, until the nth data block of each node is reduced and distributed in dimension 1. This process continues until the reduction and distribution operations for each data block in each dimension are completed. This allows the value of i to iterate from 1 to n, performing reduction and distribution operations on the i-th data block of each node in the i-th order across each dimension. The concurrent execution of other operations is similar and will not be elaborated upon here.
[0096] It should be understood that concurrent execution can significantly improve bandwidth utilization, shorten communication time, and improve communication efficiency.
[0097] In this embodiment, the dimension switching can be performed only when there is no data communication in the dimension to be switched (referred to as the destination dimension). For example, after the first block of data from each node is reduced and distributed in the first dimension, it can be determined whether there is data communication in the dimension to be switched (such as the second dimension), such as whether the reduction and distribution operation of the second block of data from each node in the second dimension has been completed. If there is no data communication in the second dimension, then the second block of data from each node is reduced and distributed in the second dimension. In this way, traffic conflicts can be avoided and communication efficiency can be improved.
[0098] The collection communication method provided in the embodiments of this application will be described below using another descriptive method.
[0099] Performing a global reduction operation on an S-dimensional fully connected topology consisting of N*M*·.·*K nodes can include the following steps.
[0100] First, the input data that needs to be allreduce on each node is divided into S data parts (also called S data blocks). The optimal split ratio is evaluated based on the principle of minimizing the total time, and the data is split equally or unequally.
[0101] Then, the first portion of data undergoes a reduce-scatter operation on the first axis, followed by another reduce-scatter operation on the second axis, and so on. Finally, a reduce-scatter operation is performed on the S-th axis, followed by an all-gather operation on the S-th axis, then an all-gather operation on the S-1-th axis, and finally an all-gather operation on the first axis. When switching axes, it is ensured that there is no other data communicating on the destination axis to avoid traffic conflicts.
[0102] The second set of data is first reduced and scattered on the second axis, then moved to the third axis for further reduction and scattering, and so on, until the S-th axis is used for reduction and scattering, then back to the first axis, followed by an allgather operation on the first axis, and so on, until the S-th axis is used for allgathering, then back to the first axis, and finally the second axis is used for allgathering. When switching axes, it must be confirmed that there is no other data communicating on the destination axis to avoid traffic conflicts.
[0103] The same applies to the other data. Choose a starting axis, perform reduce scatter operations by changing axes in turn, and then perform allgather operations by changing axes in reverse order.
[0104] In this way, the allreduce operation for the entire cluster can be completed.
[0105] The following uses a two-dimensional fully connected topology as an example to exemplify the set communication method provided in the embodiments of this application.
[0106] Please refer to Figure 10, which is a schematic diagram of another two-dimensional fully connected topology provided in an embodiment of this application. As shown in Figure 10, this two-dimensional fully connected topology includes M*N nodes (the connection relationships between the nodes are not shown for clarity). Specifically, it includes M N*1 one-dimensional fully connected topologies in the X-axis direction (i.e., the first dimension) and N M*1 one-dimensional fully connected topologies in the Y-axis direction (i.e., the second dimension).
[0107] First, the input data from the M*N nodes can be divided into two data blocks. For ease of explanation, the first data block can be called data A (assuming data size a), and the second data block can be called data B (assuming data size b).
[0108] The communication calculation process for data A can be divided into four steps, referred to as step A1, step A2, step A3, and step A4. As shown in Figure 11, step A1 involves reducing and distributing data A along the X-axis (transmitting a*1 / N data per link). Step A2 involves reducing and distributing data A along the Y-axis (transmitting a*1 / (N*M) data per link). Step A3 involves globally collecting data A along the Y-axis (transmitting a*1 / (N*M) data per link). Step A4 involves globally collecting data A along the X-axis (transmitting a*1 / N data per link).
[0109] Similarly, the communication calculation process for data B can also be divided into four steps, referred to as step B1, step B2, step B3, and step B4. As shown in Figure 12, step B1 is the reduction and distribution operation of data B along the Y-axis (transmitting b*1 / N data per link). Step B2 is the reduction and distribution operation of data B along the X-axis (transmitting b*1 / (N*M) data per link). Step B3 is the global collection operation of data B along the X-axis (transmitting b*1 / (N*M) data per link). Step B4 is the global collection operation of data B along the Y-axis (transmitting b*1 / N data per link).
[0110] In some possible implementations, the communication computation process for data A and the communication computation process for data B can be performed concurrently. For example, step A1 and step B1 can be executed concurrently. After step A1 and step B1 are completed, step A2 and step B2 can be executed concurrently, and so on, until step A4 and step B4 are completed, thus achieving the global reduction operation of this two-dimensional fully connected topology.
[0111] It should be understood that during the aforementioned communication calculation process, when switching from the X-axis to the Y-axis after completing the calculation, or vice versa, it is advisable to first confirm whether there is data communication on the destination axis. Switching should only proceed if no data communication is available. This avoids traffic conflicts and improves communication efficiency.
[0112] Furthermore, in some possible implementations, the above-mentioned method of dividing the input data into two data blocks can be equal partitioning, meaning that the data sizes of data A and data B can be the same. For example, for topologies with equal dimensional bandwidth and number of nodes, equal partitioning can be used.
[0113] In other possible implementations, the data partitioning method (i.e., partitioning ratio) can be determined based on the principle of minimizing the total completion time. For example, the time required to complete the aggregate communication method provided in this application embodiment under different partitioning methods can be calculated separately, and the partitioning ratio with the shortest time or a time shorter than a certain preset threshold can be selected as the final partitioning ratio. For instance, for topologies with unequal dimensional bandwidth and number of nodes, the data partitioning method can be determined based on the principle of minimizing the total completion time.
[0114] For example, referring to Figure 13, the time taken for step A1 is tA1, the time taken for step A2 is tA2, the time taken for step A3 is tA3, and the time taken for step A4 is tA4. The time taken for step B1 is tB1, the time taken for step B2 is tB2, the time taken for step B3 is tB3, and the time taken for step B4 is tB4.
[0115] It should be understood that after step A1, data A needs to wait for step B1 to complete (ensuring no data communication along the Y-axis) before proceeding to step A2, and so on for other steps. Therefore, the total completion time (i.e., completion time minus start time) is max(tA1, tB1) + max(tA2, tB2) + max(tA3, tB3) + max(tA4, tB4). Thus, the total completion time for different segmentation methods can be calculated separately, and the segmentation method with the shortest total completion time (or shorter than a certain preset threshold) can be used as the final determined segmentation method.
[0116] To better understand the data change process, the following section uses node1 in the 4*4 two-dimensional fully connected topology shown in Figure 2 as an example to introduce the data change process of data A in node1 in each step.
[0117] Please refer to Figure 14. Initially, the data size of data A is 16MB, consisting of 16 data blocks, namely a1-1, a1-2, ..., a1-16. Each data block is 1MB in size. The naming rules for each data block of data A in other nodes are similar to those in node1. For example, data A in node2 includes a2-1, a2-2, ..., a2-16, and data A in node16 includes a16-1, a16-2, ..., a16-16.
[0118] As shown in Figure 14, after step A1, the first four data blocks held by node1 are the result of reducing the first 4MB of data from the four nodes (node1 / 2 / 3 / 4) along the X-axis. After step A2, the first data block held by node1 is the result of reducing the first 1MB of data from all 16 nodes (node1 / 2…15 / 16). After step A3, the first four data blocks held by node1 are the result of reducing the first 4MB of data from all 16 nodes (node1 / 2…15 / 16). After step A4, the 16 data blocks held by node1 are the result of reducing the total 16MB of data from all 16 nodes (node1 / 2…15 / 16).
[0119] It should be understood that during the computation process, the A data of other nodes is similar to that of node1, the difference being the different offset positions of the reduce data obtained in steps A1 / A2 / A3. After steps A1-A4 are completed, all 16 nodes obtain 16MB of data from the allreduce operation.
[0120] Based on the above description, it should be understood that the collection communication method provided in this application can realize global reduction operation of multi-dimensional fully connected topologies. In this method, traffic on each physical link of the multi-dimensional fully connected topology is concurrent and there are no traffic conflicts, achieving maximum bandwidth utilization and minimizing the time required to complete the global reduction operation. After data blocks are continuously distributed through cascaded reduction operations in all dimensions, a global collection operation is performed in reverse order across all dimensions, resulting in a small total amount of transmitted data. Furthermore, the method has few implementation steps, and the number of steps does not increase with the size of the topology given a fixed number of dimensions. Additionally, this method perfectly matches multi-dimensional Full-mesh physical topologies, does not require supporting nodes to have routing and forwarding capabilities, and has a low implementation cost.
[0121] This application also provides an electronic device including one or more processors. The one or more processors are used to execute computer programs or instructions to implement any of the methods described in the above method embodiments.
[0122] In some possible implementations, the electronic device may be a node in the foregoing embodiments, or it may be the execution subject of one, some, or all of the steps in the foregoing method embodiments, without limitation here.
[0123] Electronic devices can be devices or apparatuses with chips, or devices or apparatuses with integrated circuits, or chips, chip systems, modules, or control units in the devices or apparatuses shown above; this application does not specifically limit their scope. It should be noted that, in this application, the term "electronic device" can refer to the electronic device itself, or to chips, functional modules, or integrated circuits within the electronic device that perform the methods provided in this application; this application does not specifically limit their scope in this regard.
[0124] For example, please refer to FIG15, which is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. As shown in FIG15, the electronic device 1500 may include one or more processors 1501 (FIG. 15 uses one processor as an example). Optionally, the electronic device 1500 may also include one or more memories 1502 coupled to the processor 1501 (FIG. 15 uses one memory as an example). The memory 1502 is used to store computer programs or instructions and / or data, and the processor 1501 is used to execute the computer programs or instructions and / or data stored in the memory 1502, so that the methods or steps in the embodiments of this application are executed.
[0125] Alternatively, the memory 1502 may be integrated with the processor 1501, or it may be set separately.
[0126] Optionally, the electronic device 1500 may further include a communication module 1503, which is used to implement communication functions. For example, the processor 1501 is used to control the communication module 1503 to implement communication with other nodes.
[0127] In the embodiments of this application, the processor may include one or more of the following: central processing unit (CPU), digital signal processor (DSP), microprocessor unit (MPU), microcontroller unit (MCU), graphics processing unit (GPU), field programmable gate array (FPGA), artificial intelligence processor (AI processor), or neural processing unit (NPU).
[0128] The memory may include one or more combinations of random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), phase-change memory (PCM), resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), cache, register, read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), and hard disk. Embodiments of this application also provide a collective communication device, which may include multiple interacting modules for implementing the methods of any of the above method embodiments.
[0129] This application also provides a collective communication device. Please refer to Figure 16, which is a schematic diagram of a collective communication device provided in this application embodiment. As shown in Figure 16, the collective communication device 1600 includes: a first module 1601, used to divide the input data of each node in a first network into n data blocks. The topology of the first network is an n-dimensional fully connected topology. n is greater than or equal to 2. A second module 1602, used to perform reduction and distribution operations on the i-th data block of each node by traversing each dimension in the i-th order, and then perform global collection operations by traversing each dimension in the reverse order of the i-th order. Here, the value of i ranges from 1 to n. The i-th order indicates the order between the dimensions.
[0130] Further functions of the first module 1601 and the second module 1602 can be found in the aforementioned method embodiments, and will not be elaborated here.
[0131] This application also provides a computer-readable storage medium, which includes a computer program or instructions that, when executed, enable the implementation of any of the methods described in the above embodiments.
[0132] This application also provides a computer program product, which includes a computer program or instructions, such that when the computer program or instructions are run, the method of any one of the above method embodiments is implemented.
[0133] This application also provides a chip device including a processor and a memory. The processor is used to invoke a computer program or computer instructions stored in the memory to cause the processor to execute any implementation of the above-described method embodiments. Optionally, the processor is coupled to the memory via an interface.
[0134] It should be noted that some optional features in the various embodiments of this application may not depend on other features in certain scenarios, or may be combined with other features in certain scenarios, without limitation.
[0135] The solutions in the various embodiments of this application can be used in reasonable combinations, and the explanations or descriptions of various terms, similar operations, or steps appearing in the embodiments can be referenced or explained to each other in the various embodiments, without limitation.
[0136] This application also provides a system that includes one or more of the above-described devices, apparatuses, computer-readable storage media, computer program products, chips, or chip systems.
[0137] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the explanation of the relevant content and beneficial effects of any of the devices, equipment, and media provided above can be referred to the corresponding method embodiments provided above, and will not be repeated here.
[0138] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0139] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0140] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0141] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the essential contribution of the technical solution of this application, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0142] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A collective communication method, characterized in that, include: The input data of each node in the first network is divided into n data blocks; The topology of the first network is an n-dimensional fully connected topology; n is greater than or equal to 2; After performing reduction and distribution operations on the i-th data block of each node by traversing each dimension in the i-th order, global collection operations are then performed by traversing each dimension in the reverse order of the i-th order; where i ranges from 1 to n; the i-th order is used to indicate the order between each dimension.
2. The collective communication method according to claim 1, characterized in that, The reduction and distribution operation of the p-th data block of each node in the q-th dimension is executed concurrently with the reduction and distribution operation of the k-th data block of each node in the l-th dimension; p, q, k, and l are all arbitrary numbers from 1 to n, where p is different from k and q is different from l.
3. The collective communication method according to claim 2, characterized in that, The global collection operation of the p-th data block of each node in the q-th dimension is executed concurrently with the global collection operation of the k-th data block of each node in the l-th dimension.
4. The collective communication method according to any one of claims 1-3, characterized in that, The first order is from the first dimension to the nth dimension; the second order is from the second dimension to the nth dimension, and then to the first dimension; when i is greater than or equal to 3, the i-th order is from the i-th dimension to the nth dimension, and then from the first dimension to the (i-1)-th dimension.
5. The collective communication method according to any one of claims 1-4, characterized in that, The n data blocks have the same data size.
6. The collective communication method according to any one of claims 1-4, characterized in that, The segmentation method is based on the time determination of each node's input data to complete each protocol distribution operation and each global collection operation.
7. The collective communication method according to any one of claims 1-6, characterized in that, The traversal of each dimension is performed under the premise that there is no conflict in the link traffic.
8. The collective communication method according to any one of claims 1-7, characterized in that, The n is 2, and the n-dimensional fully connected topology is a 2-dimensional fully connected topology, where the 2 dimensions include the first dimension and the second dimension; after performing reduction and distribution operations on the i-th data block of each node according to the i-th order across each dimension, and then performing global collection operations on each dimension according to the reverse order of the i-th order, the process includes: After performing reduction and distribution operations on the first data block of each node in the first and second dimensions respectively, a global collection operation is performed in the second and first dimensions respectively. After performing reduction and distribution operations on the second data block of each node in the second and first dimensions respectively, a global collection operation is performed in the first and second dimensions respectively.
9. A collective communication device, characterized in that, include: The first module is used to divide the input data of each node in the first network into n data blocks respectively; The topology of the first network is an n-dimensional fully connected topology; n is greater than or equal to 2; The second module is used to perform reduction and distribution operations on the i-th data block of each node by traversing each dimension in the i-th order, and then perform global collection operations on each dimension by traversing each dimension in the reverse order of the i-th order; wherein the value of i traverses from 1 to n; the i-th order is used to indicate the order between each dimension.
10. A collective communication device, characterized in that, It includes multiple interacting modules for implementing the method as described in any one of claims 1 to 8.
11. An electronic device, characterized in that, It includes one or more processors; the one or more processors are configured to execute computer programs or instructions to implement the method of any one of claims 1-8.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program or instructions that, when executed, cause the method of any one of claims 1-8 to be implemented.