All-to-all communication method for multi-dimensional full-connection network topology

By determining the data block segmentation and communication methods between processes based on the cardinality n and dimension d in the multi-dimensional fully connected network topology, the problems of link congestion and low hardware utilization in Alltoall communication in the multi-dimensional fully connected network topology are solved, and efficient and redundant communication is achieved.

CN120223541APending Publication Date: 2025-06-27SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510384305.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing Alltoall communication algorithm has problems of link congestion and low hardware utilization in the multi-dimensional fully connected network topology, and cannot fully utilize all links of the network.

Method used

By determining the data block segmentation and communication method between processes based on the cardinality n and dimension d of the multi-dimensional fully connected network topology, each process communicates only with adjacent processes and exchanges data blocks in data block groups of different dimensions until the data block exchange of all dimension groups is completed.

Benefits of technology

Alltoall communication without link congestion in a multi-dimensional fully connected network topology is realized, with hardware utilization reaching 100%, the number of communication steps is the dimension d of the network topology, the data block transmission path is the shortest path, the total traffic volume is the best, and there is no communication redundancy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223541A_ABST
    Figure CN120223541A_ABST
Patent Text Reader

Abstract

The invention discloses a full-to-full communication method for a multi-dimensional full-connection network topology, and belongs to the field of full-to-full communication, and the method comprises the steps: constructing the multi-dimensional full-connection network topology; determining the total number of processes participating in the Allall communication task according to the cardinal number and the dimension number; equally dividing a plurality of parts of data, participating in the Allall communication task, of each process into a plurality of data blocks according to dimensions; numbering and grouping each data block of each process based on the source process, the target process and the dimension of the data block to obtain a plurality of data block groups; each process communicates with an adjacent process of each dimension at the same time, and exchanges data blocks in the data block groups of different dimensions; repeating the exchange of the data blocks until the data blocks of all dimension groups are exchanged between each process and the adjacent process of each dimension; and recombining the data blocks into final complete data by each process according to the serial numbers of the data blocks. According to the invention, the problems of link congestion and low hardware utilization rate in the multi-dimensional fully-connected network topology are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of all-to-all communication, and particularly relates to an all-to-all communication method for a multi-dimensional fully connected network topology. Background Art

[0002] The multi-dimensional fully connected network topology is a new type of highly scalable and cost-effective topology, mainly used in the supercomputer interconnection network system. A typical multi-dimensional fully connected network topology is constructed by two parameters: the radix n and the dimension d. The n processes within one dimension are fully connected, and the processes at the same corresponding positions in each dimension are also fully connected. In the case of the same total number of processes, the multi-dimensional fully connected network has a lower cost than the fully interconnected network. Under the uniform traffic pattern, the multi-dimensional fully connected network has excellent performance and can meet the requirements of high bandwidth and low latency for the supercomputer network.

[0003] The collective communication algorithm plays a crucial role in the fields of high-performance computing (HPC) and artificial intelligence (AI). Among them, the Alltoall operation is the key to achieving efficient parallel computing, and it plays an important role in various parallel algorithms and applications, such as model parallelism, matrix transpose, fast Fourier transform and other scenarios. Traditional Alltoall algorithms rely on algorithm design based on process indices to reduce communication latency. By optimizing the Alltoall operation, the performance and efficiency of parallel computing can be significantly improved.

[0004] The Message Passing Interface (MPI) is the most important and mainstream parallel programming framework in the field of high-performance computing today. MPI not only includes basic one-to-one communication interfaces such as sending and receiving, but also includes many collective communication interfaces, such as: MPI_Alltoall, MPI_Allgather, MPI_Alltoall, etc. Programmers only need to call the function interfaces to complete the communication, without having to care about the implementation of the communication algorithm, let alone understand the use of the hardware responsible for communication. The parallel applications written also have good portability. In common open-source implementations, the algorithms used for the Alltoall operation are the Pairwise algorithm and the Bruck algorithm. However, these algorithms do not fully utilize the characteristics of the multi-dimensional fully connected network topology, resulting in a large amount of link congestion and low hardware utilization, greatly reducing the communication speed. This is caused by the characteristics of the multi-dimensional fully connected network topology. Therefore, it is very necessary to optimize the implementation method of the Alltoall communication operation for the multi-dimensional fully connected network topology.

[0005] One existing method is the Bruck algorithm: an index-based Alltoall algorithm that reduces the number of communication steps from P to log2P by using a radix r = 2. The algorithm consists of three phases: an initial local rotation phase, a communication phase containing multiple rounds of point-to-point data exchange, and a final local rotation phase. It is a store-and-forward algorithm and thus has good performance for short messages that are sensitive to latency. The algorithm first performs local copying and upward shifting of the data blocks of each process from the input buffer to the output buffer, such that the data blocks that each process is to send to itself are located at the top of the output buffer. To achieve this, process i' must rotate its data upward by i' blocks. At each communication step k , process i' sends all data blocks whose k-th bit is 1 to process (i' + 2k) % P, receives data from process (i' - 2k) % P, and stores it in the positions where the k-th bit is 1. After all communication steps are executed, all data is routed to the correct destination processes, but the order of the data blocks in the output buffer is incorrect. In the last step, each process performs a local inverse shift (memory copy) on the data blocks to place the data in the correct order.

[0006] For example Figure 1 , when the number of processes is 4, with process labels P0, P1, P2, and P3 respectively, each process divides its data into 4 equal data blocks. The data blocks are represented by serial numbers, where the 0th bit represents the destination process and the 1st bit represents the source process. For example, the data block serial numbers of P0 are 00, 01, 02, and 03, representing the final receiving processes P0, P1, P2, and P3 respectively. The positions of the data blocks are represented by binary serial numbers. In the first step, each process performs an initial rotation on the data blocks, moving the data blocks with the same source and destination process serial numbers upward to position 00. For example Figure 1in the "1. Initial Rotation" section; then in the "2. Communication Phase", step 0, each process sends data to the process with the next sequence number and receives data from the previous process, sending and receiving all data blocks with the least significant bit being 1. That is, P0 sends data blocks "01" and "03" to P1, P1 sends data blocks "12" and "10" to P2, P2 sends data blocks "23" and "21" to P3, and P3 sends data blocks "30" and "32" to P0. The received data blocks are placed in the corresponding positions. In step 1, P0 and P2 exchange data with each other, and P1 and P3 exchange data with each other, sending and receiving all data blocks with the first bit of the position sequence number being 1. That is, P0 sends data blocks "02" and "32" to P2, P1 sends data blocks "13" and "03" to P3, P2 sends data blocks "20" and "10" to P0, and P3 sends data blocks "31" and "21" to P1. The received data blocks are placed in the corresponding positions. After the two-step communication, each process has obtained the required data blocks, but the order of the data blocks is incorrect. After going through the "3. Final Rotation" section, the order is corrected and the algorithm ends.

[0007] The disadvantages are as follows:

[0008] 1. In the Brook algorithm, some data blocks of each process are forwarded by other processes to the target process, resulting in redundant communication.

[0009] 2. The Brook algorithm is based on the index of the processes and does not consider the actual connection relationship between the processes, that is, the actual network topology. When data must be forwarded through relay processes between processes, problems such as link congestion will occur, and not all links of the network can be fully utilized.

[0010] Another existing method is the Dimension Order Algorithm (DO): The DO algorithm is an Alltoall algorithm for two-dimensional fully connected networks. This algorithm arranges the sending and receiving of data for each process according to the number of processes in each dimension and the process sequence numbers of the two-dimensional fully connected network. The DO algorithm is divided into two similar steps. In the first step, the algorithm exchanges data in the first dimension. In the second step, the algorithm sends data in the second dimension and forwards the data from the first step to complete the algorithm.

[0011] For example Figure 2 , a 2 2Two-dimensional fully connected network, with processes P0 and P1 interconnected, processes P1 and P3 interconnected, processes P2 and P3 interconnected, and processes P2 and P0 interconnected. In the first step, each process communicates with the processes on the x-axis, i.e., P0 communicates with P1, and P2 communicates with P3. P0 sends data D[0,1] and D[0,3], P1 sends data D[1,0] and D[1,2], P2 sends data D[2,1] and D[2,3], and P3 sends data D[3,0] and D[3,2]. One of the two pieces of data sent by each process is what the other process needs, and the other is forwarded by the other process in the next step. In the second step, each process communicates with the processes on the y-axis, i.e., P0 communicates with P2, and P1 communicates with P3. P0 sends data D[0,2] and D[1,2], P1 sends data D[0,3] and D[1,3], P2 sends data D[2,0] and D[3,0], and P3 sends data D[2,1] and D[3,1], thus completing the Alltoall task of the network.

[0012] Disadvantages are as follows:

[0013] 1. The dimension-order algorithm only uses the links in one dimension of the network for each communication, which results in insufficient hardware utilization and does not fully utilize all the links of the network.

[0014] 2. The dimension-order algorithm is only applicable to two-dimensional fully connected networks and does not consider higher-dimensional fully connected networks, and cannot handle the Alltoall task of multi-dimensional fully connected network topologies under arbitrary parameters. Summary of the Invention

[0015] In view of the above deficiencies in the prior art, a method for all-to-all communication for multi-dimensional fully connected network topologies provided by the present invention solves the problems of link congestion and low hardware utilization in multi-dimensional fully connected network topologies.

[0016] To achieve the above invention objective, the technical solution adopted by the present invention is: A method for all-to-all communication for multi-dimensional fully connected network topologies, including:

[0017] According to the number of computing nodes P, determine the radix n and the dimension d, and construct a multi-dimensional fully connected network topology according to the radix and the dimension;

[0018] Determine the total number of processes participating in the Alltoall communication task according to the radix n and the dimension d, and represent each process with a coordinate number;

[0019] According to the dimension, evenly divide the n d pieces of data of each process participating in the Alltoall communication task into d data blocks;

[0020] Number each data block of each process based on the source process, destination process, and dimension of the data block, and group them according to the dimension to obtain several data block groups;

[0021] Each process communicates with adjacent processes in each dimension simultaneously to exchange data blocks in the data block groups of different dimensions; the exchange of data blocks is repeated until each process has exchanged data blocks of all dimension groups with adjacent processes in each dimension;

[0022] Each process recombines the data blocks into the final complete data according to the data block numbers.

[0023] Further, the number of computing nodes satisfies the following relationship with the base and the dimension:

[0024] P = n d

[0025] n ≥ 2, n ∈ Z +

[0026] d ≥ 1, d ∈ Z +

[0027] where P is the number of computing nodes; Z + is the set of positive integers.

[0028] Further, the coordinate numbers of each process are:

[0029] (x d-1 x d-2 … x i x i-1 … x1x0)

[0030] 0 ≤ x i ≤ n - 1

[0031] 0 ≤ i ≤ d - 1

[0032] where x i is the value of the i-th bit in the coordinate number, which is an n-based number; n is the base of the multi-dimensional fully connected network topology; d is the dimension of the multi-dimensional fully connected network topology; i is the dimension.

[0033] Further, the n d data sizes of each process participating in the Alltoall communication task are equal.

[0034] Further, the data blocks in the data block group are arranged in the order of the destination processes.

[0035] Further, there is a communication link between the adjacent process and the current process; when the Hamming distance between the coordinate numbers of two processes is 1, there is a communication link between the two processes.

[0036] Further, each process communicates with adjacent processes in each dimension simultaneously to exchange data blocks in the data block groups of different dimensions, specifically:

[0037] In the exchange of the s-th step data block, each process communicates with the adjacent process of dimension (s + j - 1) % d, and exchanges a part of the data blocks in the data block group of dimension j; where 0 ≤ j ≤ d - 1, 1 ≤ s ≤ d; The data blocks received by each process are filled in sequence into the positions of the data block groups of the same dimension locally where the data blocks are sent.

[0038] The data blocks received by each process are filled in sequence into the positions of the data block groups of the same dimension locally where the data blocks are sent.

[0039] Furthermore, the part of the data blocks to be exchanged are the data blocks in the data block group of dimension j where the (s + j - 1) % d-th bit of the destination process number is the same as the (s + j - 1) % d-th bit of the adjacent process coordinate number.

[0040] Furthermore, the dimension of the adjacent process is the i value with different numerical values in the coordinate numbers of the adjacent process and the current process.

[0041] The beneficial effects of the present invention are as follows: The Alltoall method provided by the present invention for the multi-dimensional fully connected network topology determines how to arrange data segmentation and inter-process communication according to two parameters, the cardinality n and the dimension d of the multi-dimensional fully connected network topology, and can flexibly adapt to the multi-dimensional fully connected network topology under all parameters; in the network topology, each process in this method only communicates and exchanges data with adjacent processes, and there is no problem of link congestion; at the same time, all links of the network topology are utilized during the entire communication process, and the hardware utilization rate is 100%; the number of communication steps of this method is the dimension d of the network topology, that is, the minimum number of steps to complete the Alltoall communication; the number of data blocks sent in each step of the Alltoall method is the same, so the link load is the same, and the transmission path of each data block is the shortest path, and the total communication volume is optimal, without communication redundancy. Description of the Drawings

[0042] Figure 1 It is a schematic execution diagram of a Brook algorithm in the related technology of the present invention;

[0043] Figure 2 It is a schematic execution diagram of a two-dimensional order algorithm in the related technology of the present invention;

[0044] Figure 3 It is a flowchart of the method of the present invention.

[0045] Figure 4 It is a schematic diagram of the multi-dimensional fully connected network topology in an embodiment of the present invention;

[0046] Figure 5 It is a schematic diagram of data segmentation in an embodiment of the present invention.

[0047] Figure 6 It is a schematic diagram of the first-step communication in an embodiment of the present invention.

[0048] Figure 7 This is the communication schematic diagram for the second step in the embodiments of the present invention.

[0049] Figure 8 This is the data integration schematic diagram for the embodiments of the present invention. Specific Embodiments

[0050] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0051] As Figure 3 shown, in an embodiment of the present invention, a full-to-full communication method for a multi-dimensional fully connected network topology includes:

[0052] Determine the radix n and the dimension d according to the number of computing nodes P, and construct a multi-dimensional fully connected network topology based on the radix and the dimension;

[0053] Determine the total number of processes participating in the Alltoall communication task according to the radix n and the dimension d, and represent each process with a coordinate number;

[0054] According to the dimension, evenly divide the n d pieces of data for each process participating in the Alltoall communication task into d data blocks;

[0055] Number each data block of each process based on the source process, destination process, and dimension of the data block, and group them according to the dimension to obtain several data block groups;

[0056] Each process communicates with the adjacent processes in each dimension simultaneously to exchange the data blocks in the data block groups of different dimensions; repeat the exchange of the data blocks until each process has exchanged the data blocks of all dimension groups with the adjacent processes in each dimension;

[0057] Each process recombines the data blocks into the final complete data according to the data block numbers.

[0058] The number of computing nodes satisfies the following relationship with the radix and the dimension:

[0059] P = n d

[0060] n ≥ 2, n ∈ Z +

[0061] d ≥ 1, d ∈ Z +

[0062] Where P is the number of computing nodes; Z + is a set of positive integers.

[0063] The coordinate numbers of the processes are as follows:

[0064] (x d-1 x d-2 …x i x i-1 …x1x0)

[0065] 0 ≤ x i ≤ n - 1

[0066] 0 ≤ i ≤ d - 1

[0067] Where xi is the value of the i-th bit in the coordinate number, which is an n-ary number; n is the base of the multi-dimensional fully connected network topology; d is the dimension of the multi-dimensional fully connected network topology; and i is the dimension.

[0068] The n d data sizes of the processes participating in the Alltoall communication task are equal.

[0069] In this embodiment, step 1: Determine that the total number of processes participating in the Alltoall communication task is n according to the input multi-dimensional fully connected network parameters n and d passed in by the user. d , and each process is uniquely represented by a d-bit n-ary coordinate number, in the form of (x d-1 x d-2 …x i x i-1 …x1x0), where 0 ≤ x i ≤ n - 1, 0 ≤ i ≤ d - 1. If the Hamming distance between the coordinates of two processes is 1, there is a link connection between the processes. Each process has n d data pieces of the same size participating in the Alltoall communication.

[0070] In this embodiment, Figure 4 shows a schematic diagram of the multi-dimensional fully connected network topology of 2 2 . There are 4 processes in the network, represented by 2-bit binary numbers, namely 00, 01, 10, 11. Two processes with a Hamming distance of 1 in the numbers are connected to each other. In the 2D fully connected network topology, if only the least significant bit of the two process numbers is different, the other is called an adjacent process in dimension 0; if only the most significant bit is different, the other is called an adjacent process in dimension 1.

[0071] The data blocks in the data block group are arranged in the order of the destination processes.

[0072] The adjacent process and the current process have a communication link; when the Hamming distance between the coordinate numbers of two processes is 1, there is a communication link between the two processes.

[0073] In this embodiment, step 2: According to the dimension parameter d in the network parameters, for the n d pieces of data for each process participating in the Alltoall communication task, perform a data splitting operation on each piece of data. Each piece of data is split into d data blocks of equal size, and they are grouped and numbered according to the dimensions of the multi-dimensional fully connected network topology. The numbering contains information such as the source process, destination process, and dimension group of the data block. The data blocks in the same dimension group are arranged in the order of the destination process.

[0074] In this embodiment, Figures 5 - 8 shows an execution schematic diagram of the all-reduce algorithm based on the 2 2 -dimensional fully connected network topology according to an embodiment of the present invention. For the Alltoall communication operation, as Figure 5 shown, each process has 4 pieces of data to be processed of the same size. The numbering of the data indicates its source process and destination process. For example, for the data "02" of process 00, the first digit of the number indicates that it comes from process 00, and the "2" in the 0th digit indicates that it will finally be sent to destination process 10. After the Alltoall communication is executed, each process will obtain the data in all processes whose destination process is itself and arrange them in order. The data splitting and grouping operations are as Figure 5 shown. Each piece of data of each process is evenly split into 2 data blocks and re-numbered and grouped. For example, the data 03 of process 00 is evenly divided into data blocks [00, 11, 0] and [00, 11, 1]. Among them, "00" indicates that the data block originates from process 00, "11" indicates that the data block is required by process 11, and the final "0" and "1" indicate the dimension grouping, belonging to dimension 0 group and dimension 1 group respectively. After grouping, the Figure 5 data block arrangement on the right is obtained.

[0075] Each process communicates with adjacent processes in each dimension simultaneously, and exchanges data blocks in the data block groups of different dimensions. Specifically:

[0076] In the exchange of data blocks in the s-th step, each process communicates with adjacent processes in dimension (s + j - 1) % d, and exchanges partial data blocks in the data block group of dimension j; where, 0 ≤ j ≤ d - 1, 1 ≤ s ≤ d;

[0077] The data blocks received by each process are filled in the positions of the data block groups of the same dimension locally where the data blocks are sent in order.

[0078] The partial data blocks to be exchanged are the data blocks in the data block group of dimension j where the (s + j - 1) % d-th digit of the destination process number is the same as the (s + j - 1) % d-th digit of the adjacent process coordinate number.

[0079] The dimension of the adjacent process is the i value where the numerical values in the coordinate numbers of the adjacent process and the current process are different.

[0080] In this embodiment, step 3: All processes in the multi-dimensional fully connected network topology perform communication operations. In the entire communication process, the total number of communication rounds is d. In each communication round, each process simultaneously communicates with the processes adjacent to it in each dimension, exchanges data blocks of different dimension groups, and the number of communication data blocks is The communication strategy is that in the s-th communication round, each process communicates with the adjacent process of dimension (s + j - 1) % d, exchanges the data blocks in the number of dimension j groups, and the exchanged data blocks are the data blocks where the (s + j - 1) % d-th bit of the destination process number in dimension j group is the same as the (s + j - 1) % d-th bit of the coordinate number of this adjacent process. The data blocks received by each process are filled in sequence into the positions of the data blocks sent in the local same dimension group. After D communication rounds, each node has exchanged all the data blocks of all dimension groups with its adjacent nodes in each dimension. Assuming the size of each data block is Then the link load of the Alltoall method is The total communication volume is d(n - 1)×n d-1 .

[0081] In this embodiment, according to the exclusive OR operation of the process numbers, it is known that process 00 and processes 01, 10, and 11 are adjacent processes on dimension 0, and process 00 and processes 10, 01, and 11 are adjacent processes on dimension 1. The cardinality of the 2D fully connected network topology is n = 2. Therefore, data blocks are transmitted on each link in each communication step. The communication operations of each process in the first communication round are as Figure 6 shown. Process 00 and process 01, process 10 and process 11 exchange the data of dimension 0 group, and process 00 and process 10, process 01 and process 11 exchange the data blocks of dimension 1 group. The positions of the exchanged data blocks are determined according to the target process number. For example, process 00 and process 01 exchange the data blocks of dimension 0 group. The 0-th bit of process 01's number is 1. Therefore, process 00 sends all the data blocks in its dimension 0 group whose 0-th bit of the destination process number is 1 to process 01, that is Figure 6 the data blocks [00, 01, 0] and [00, 11, 0] pointed by the brackets. Each process determines which data blocks to send according to the target process number. After receiving the data blocks, the process fills them in sequence into the positions of the data blocks sent in the local corresponding dimension group. For example, after process 00 receives the data blocks [01, 00, 0] and [01, 10, 0] from process 01, it fills them into the positions of the data blocks [00, 01, 0] and [00, 11, 0] sent in dimension 0 group in this round of communication respectively. The data block situation of each process after the first communication step is asFigure 6 Shown on the right.

[0082] Furthermore, the communication operation of each process in the second communication round is as follows: Figure 7 As shown, process 00 exchanges data of dimension 1 group with process 01, process 10 exchanges data blocks of dimension 0 group with process 00 and process 10, process 01 and process 11. Similarly, data block sending is arranged according to the target process number. For example, process 00 sends all data blocks whose first bit of the destination process number is 1 in the local dimension 0 group to process 10, that is, Figure 7 The data blocks [00, 10, 0] and [01, 10, 0] pointed to by the brackets on process 00 are filled in the sent data block positions with the data blocks [10, 00, 0] and [11, 00, 0] received from process 10 respectively. The other processes communicate in the same way. After the communication operation is completed, each process has obtained all the data blocks it needs to receive and arranged them in the order of the process number of the dimension group.

[0083] Step 4: After all processes complete the communication operation, they obtain all the required data blocks and reassemble the data blocks into the final complete data according to the data block numbers.

[0084] Data block adjustment operations are as follows Figure 8 As shown, for example, the data blocks [11, 00, 0] and [11, 00, 1] of process 00 are spliced ​​into the data "30", and each process adjusts and splices all local data blocks according to the initial data segmentation method into the final complete data, and the Alltoall communication operation is completed.

[0085] The Alltoall method for multi-dimensional fully connected network topology provided by the present invention determines how to arrange data segmentation and communication between processes according to two parameters of the cardinality n and dimension d of the multi-dimensional fully connected network topology. Compared with the related technology 2 which can only be applied to 2-dimensional fully connected networks, it can flexibly adapt to the multi-dimensional fully connected network topology under all parameters; in the network topology, each process of the Alltoall method only communicates and exchanges data with adjacent processes, and there is no problem of link congestion; at the same time, the Alltoall method uses all links of the network topology in the entire communication process, and the hardware utilization rate is 100%; the number of communication steps of the Alltoall method is the dimension d of the network topology, that is, the minimum number of steps to complete the Alltoall communication, which is better than the number of steps in the related technology 1 In the 2D fully connected network topology, the number of communication steps is the same as that of the related technology 2; the Alltoall method sends the same number of data blocks in each communication step, so the link load is the same, and the transmission path of each data block is the shortest path, the total communication volume is optimal, and there is no communication redundancy.

Claims

1. An all-to-all communication method for a multi-dimensional fully connected network topology, characterized in that: include: According to the number of computing nodes P, the cardinality n and the dimension d are determined, and a multi-dimensional fully connected network topology is constructed based on the cardinality and the dimension; Determine the total number of processes participating in the Alltoall communication task according to the cardinality n and the dimension d, and represent each process with a coordinate number; According to the dimension, each process participates in the Alltoall communication task n d The data is equally divided into d data blocks; Each data block of each process is numbered based on the source process, destination process and dimension of the data block, and grouped according to the dimension to obtain a number of data block groups; Each process communicates with the adjacent processes of each dimension at the same time to exchange data blocks in data block groups of different dimensions; the exchange of data blocks is repeated until each process exchanges data blocks of all dimension groups with the adjacent processes of each dimension; Each process reassembles the data blocks into the final complete data according to the data block numbers.

2. The all-to-all communication method for a multi-dimensional fully connected network topology according to claim 1, characterized in that: The number of computing nodes, cardinality and dimension satisfy: P=n d n≥2,n∈Z + d≥1,d∈Z + Where P is the number of computing nodes; Z + is a set of positive integers.

3. The all-to-all communication method for a multi-dimensional fully connected network topology according to claim 1, characterized in that: The coordinate numbers of the processes are: (x d-1 x d-2 …x i x i-1 …x1x0) 0≤x i ≤n-1 0≤i≤d-1 Among them, x i is the value of the i-th position in the coordinate number, which is an n-ary number; n is the cardinality of the multi-dimensional fully connected network topology; d is the dimension of the multi-dimensional fully connected network topology; i is the dimension.

4. The all-to-all communication method for a multi-dimensional fully connected network topology according to claim 1, characterized in that: Each process participates in the Alltoall communication task d The data sizes are all equal.

5. The all-to-all communication method for a multi-dimensional fully connected network topology according to claim 1, characterized in that: The data blocks in the data block group are arranged in order of the destination process.

6. The all-to-all communication method for a multi-dimensional fully connected network topology according to claim 1, characterized in that: The adjacent process has a communication link with the current process; when the Hamming distance between the coordinate numbers of the two processes is 1, the two processes have a communication link.

7. The all-to-all communication method for a multi-dimensional fully connected network topology according to claim 1, characterized in that: Each process communicates with adjacent processes of each dimension at the same time to exchange data blocks in data block groups of different dimensions, specifically: In the exchange of data blocks in step s, each process communicates with adjacent processes of dimension (s+j-1)%d to exchange the data blocks of dimension j. Partial data block; where 0≤j≤d-1, 1≤s≤d; The data blocks received by each process are sequentially filled into the position of the sent data blocks in the local data block group of the same dimension.

8. The all-to-all communication method for a multi-dimensional fully connected network topology according to claim 7, characterized in that: For exchange Some data blocks are data blocks whose destination process number (s+j-1)%d is the same as the adjacent process coordinate number (s+j-1)%d in the data block group with dimension j.

9. The all-to-all communication method for a multi-dimensional fully connected network topology according to claim 7, characterized in that: The dimension of the adjacent process is the i value that is different from the coordinate number of the adjacent process and the current process.

Citation Information

Cited By

  • Method for performing multi-configuration simultaneous control on large-scale full switching matrix switch by using All-to-All broadcast algorithm

    CN120949629A