A method, device, terminal device and medium for sparse gradient reduction
In the distributed training of large-scale neural network models, the devices and sparse tensors are divided into two-dimensional grids according to the network topology structure, and the gradient specification is optimized by using Reduce-Scatter and AllGather algorithms, which solves the low bandwidth bottleneck in data parallel training and improves training efficiency and performance.
Patent Information
- Application Number
- CN202310144049.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-02-21
AI Technical Summary
In distributed training of large-scale neural network models, standardized communication of data parallelism methods becomes a bottleneck, resulting in inefficient training, and the existing methods of sparse gradient tensors fail to make full use of the underlying bandwidth of heterogeneous networks.
By obtaining the network topology, the global device and sparse tensor are divided into two-dimensional device grids and two-dimensional sparse tensors. The Reduce-Scatter algorithm and the AllGather algorithm are used to perform communication regulations, the gradient specification process is optimized, and an efficient gradient specification mode is formulated based on the underlying bandwidth.
The gradient specification performance of distributed data parallel training under heterogeneous networks is optimized, and network bandwidth is fully utilized to improve training efficiency and performance.
Smart Images

Figure CN116405398B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural network model training, and particularly relates to a method, device, terminal device and medium for sparse gradient reduction. Background Art
[0002] At present, with the popularization and commercialization of large-scale neural network models, it has become a development trend to further improve the training speed of the models through distributed training. Currently, in large-scale model training, most use the data parallel method for training: splitting the data onto different devices, while having the same copy of the model on different devices, and performing reduction communication (AllReduce) on the reverse gradient tensors during the training process. AllReduce is the most commonly used collective communication primitive in the field of deep learning, mainly used for gradient synchronization between multiple machines / multiple cards. AllReduce was born in the HPC field, and MPI (Message Passing Interface, a communication library in the HPC field) implemented the AllReduce primitive in multiple ways, which is the main reference basis in the field of deep learning. The Message Passing Interface (MPI) is the most important and mainstream parallel programming framework in the current high-performance computing field. Allgather is an important collective communication interface in MPI. It is responsible for aggregating the data in the send buffers of all processes in a communication domain, arranging the data in the order of the send process numbers, and then broadcasting all the data to all processes and placing it in their receive buffers. In the case of an increasing number of devices in data parallelism, reduction communication becomes the bottleneck of the overall training, suppressing the efficiency of model training and significantly increasing the machine occupancy and time overhead. One of the key technologies to alleviate the performance bottleneck is to reduce the amount of data in reduction communication. Through the method of sparsification, some values in the reverse gradient are clipped, and reduction communication is performed on the sparsified gradient tensors. The sparsified gradient tensors reduce the amount of data transmitted in reduction communication, thus greatly alleviating the communication bottleneck. Summary of the Invention
[0003] The present invention aims to at least partly solve one of the technical problems in the above technologies. For this reason, the first object of the present invention is to propose a method for network bandwidth-aware sparse gradient reduction, which formulates an efficient gradient reduction mode according to the underlying bandwidth to achieve efficient utilization of network bandwidth, alleviate the low-bandwidth bottleneck, and improve training performance.
[0004] The second object of the present invention is to propose a device for network bandwidth-aware sparse gradient reduction.
[0005] The third object of the present invention is to propose a terminal device.
[0006] The fourth object of the present invention is to propose a computer medium.
[0007] To achieve the above object, an embodiment of the first aspect of the present invention proposes a method for network bandwidth-aware sparse gradient reduction, including:
[0008] Obtain the network topology structure;
[0009] According to the network topology structure, divide the global devices into a two-dimensional device grid (n, d), and at the same time divide the sparse tensor into a two-dimensional sparse tensor (d, n), where n represents the number of nodes and d represents the number of devices in each node;
[0010] Based on a preset Reduce-Scatter algorithm, perform communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) to determine the reduction result;
[0011] Based on the AllGather algorithm, process the reduction result to obtain the final processing result.
[0012] In one embodiment, dividing the global devices into a two-dimensional device grid (n, d) according to the network topology structure, and at the same time dividing the sparse tensor into a two-dimensional sparse tensor (d, n), includes:
[0013] Based on the network topology structure, determine the global device index. The global device index is a total of N devices from 0,..., N - 1; abstract it into a two-dimensional matrix (n, d) according to the network connection conditions and determine it as the two-dimensional device grid (n, d); where n * d = N;
[0014] The sparse tensor includes sparse values and sparse indexes, both of which are divided into N parts and mapped to a two-dimensional structure of (d, n) to be determined as a two-dimensional sparse tensor (d, n).
[0015] In one embodiment, the preset Reduce-Scatter algorithm is the first Reduce-Scatter algorithm;
[0016] Based on the first Reduce-Scatter algorithm, perform communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) to determine the reduction result, including:
[0017] Determine that the id of each device gid in the two-dimensional device grid (n, d) is [r1, r2] = [gid / d, gid % d];
[0018] There are a total of N - 1 rounds of communication. In the r-th round, distribute its own sparse tensor block [(r2 + r % d) % d, (r1 + r / d) % n] to the device numbered [(r1 + r / d) % n, (r2 + r % d) % d];
[0019] After completing N-1 rounds of communication, each device sums up the received sparse tensor data blocks by reduction to determine the reduction result.
[0020] In one embodiment, the preset Reduce-Scatter algorithm is the second Reduce-Scatter algorithm;
[0021] Performing communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) based on the second Reduce-Scatter algorithm to determine the reduction result, including:
[0022] Determine that the id of each device gid in the two-dimensional device grid (n, d) is [r1, r2] = [gid / d, gid % d];
[0023] There are a total of n - 1 + d - 1 rounds of communication. In the first n - 1 rounds, in the r-th round, distribute its own [*, (r1 - r + n) % n] sparse tensor data block to the device numbered [(r1 - r + n) % n, r2]; where * represents all data in this dimension;
[0024] After completing n - 1 rounds of communication, each device sums up the received sparse tensor data blocks by reduction to determine the first intermediate reduction result;
[0025] In the subsequent d - 1 rounds, in the r-th round, distribute the [(r2 - r + d) % d, r1] sparse tensor data block in the first intermediate reduction result to the device numbered [r1, (r2 - r + d) % d]; after completing d - 1 rounds of communication, each device sums up the received sparse tensor data blocks by reduction to determine the reduction result.
[0026] In one embodiment, the preset Reduce-Scatter algorithm is the third Reduce-Scatter algorithm;
[0027] Performing communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) based on the third Reduce-Scatter algorithm to determine the reduction result, including:
[0028] Determine that the id of each device gid in the two-dimensional device grid (n, d) is [r1, r2] = [gid / d, gid % d];
[0029] There are a total of d - 1 + n - 1 rounds of communication. In the first d - 1 rounds, in the r-th round, distribute its own [(r2 + r) % d, *] sparse tensor data block to the device numbered [r1, (r2 + r) % d], where * represents all data in this dimension;
[0030] After completing d - 1 rounds of communication, each device sums up the received sparse tensor data blocks by reduction to determine the second intermediate reduction result;
[0031] In the last n - 1 rounds, in the r-th round, the sparse tensor data block in the second intermediate reduction result at positions [r2, (r1 + r) % n] is distributed to the device numbered [(r1 + r) % n, r2].
[0032] After completing n - 1 rounds of communication, each device performs a reduction sum on the received sparse tensor data blocks to determine the reduction result.
[0033] In one embodiment, the reduction result is processed based on the AllGather algorithm to obtain the final processing result, including:
[0034] For any device coordinates [r1, r2], first, the reduction results on devices with the same second - dimension coordinate r2 are aggregated to obtain an intermediate aggregated sparse tensor of [r2, *], where * represents all data in this dimension;
[0035] Devices with the same first - dimension coordinate r1 aggregate the intermediate aggregated sparse tensor [r2, *] to obtain the final processing result.
[0036] In one embodiment, it further includes performing VGG16, LSTM, and BERT task training on 2 - node 4 - card, 2 - node 8 - card, and 4 - node 8 - card hardware devices respectively.
[0037] To achieve the above object, an embodiment of the second aspect of the present invention proposes a network - bandwidth - aware sparse gradient reduction device, including:
[0038] An acquisition module, configured to acquire the network topology structure;
[0039] A partitioning module, configured to divide the global devices into a two - dimensional device grid (n, d) according to the network topology structure, and at the same time divide the sparse tensor into a two - dimensional sparse tensor (d, n), where n represents the number of nodes and d represents the number of devices in each node;
[0040] A determination module, configured to perform communication reduction on the two - dimensional device grid (n, d) and the two - dimensional sparse tensor (d, n) based on a preset Reduce - Scatter algorithm to determine the reduction result;
[0041] A processing module, configured to process the reduction result based on the AllGather algorithm to obtain the final processing result.
[0042] To achieve the above object, an embodiment of the third aspect of the present invention provides a terminal device, which includes a memory, a processor, and a program for network bandwidth-aware sparse gradient reduction stored in the memory and executable on the processor. When the program for network bandwidth-aware sparse gradient reduction is executed by the processor, it implements the steps of the method for network bandwidth-aware sparse gradient reduction as described above.
[0043] To achieve the above object, an embodiment of the fourth aspect of the present invention provides a computer medium, which is a computer-readable storage medium. A program for network bandwidth-aware sparse gradient reduction is stored on the computer-readable storage medium. When the program for network bandwidth-aware sparse gradient reduction is executed by a processor, it implements the steps of the method for network bandwidth-aware sparse gradient reduction as described above.
[0044] The present invention provides a method, device, terminal device, and computer medium for network bandwidth-aware sparse gradient reduction. The method includes: obtaining a network topology structure; dividing global devices into a two-dimensional device grid (n, d) according to the network topology structure, and at the same time dividing a sparse tensor into a two-dimensional sparse tensor (d, n), where n represents the number of nodes and d represents the number of devices within each node; performing communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) based on a preset Reduce-Scatter algorithm to determine a reduction result; and processing the reduction result based on an AllGather algorithm to obtain a final processing result. This method can greatly optimize the performance of gradient reduction in distributed data parallel training under a heterogeneous network and make full use of network bandwidth.
[0045] Other features and advantages of the present invention will be described in the following specification, and will, in part, become apparent from the specification or be learned by practicing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structure specifically pointed out in the written specification and the drawings.
[0046] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0047] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0048] Figure 1 is a schematic diagram of the implementation of dense Ring Reduce-Scatter in the prior art;
[0049] Figure 2A flowchart of a method for network bandwidth-aware sparse gradient reduction according to an embodiment of the present invention;
[0050] Figure 3 A schematic diagram of the Reduce-Scatter algorithm in the prior art and the preset Reduce-Scatter algorithm proposed by the present invention;
[0051] Figure 4 A schematic diagram of the first Reduce-Scatter algorithm according to an embodiment of the present invention;
[0052] Figure 5 A schematic diagram of the second and third Reduce-Scatter algorithms according to an embodiment of the present invention;
[0053] Figure 6 A block diagram of a device for network bandwidth-aware sparse gradient reduction according to an embodiment of the present invention. Detailed implementation manners
[0054] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0055] According to the attached Figure 1-6 A method for network bandwidth-aware sparse gradient reduction proposed by the present invention will be described.
[0056] In the prior art, as Figure 1 shown, it is the implementation of dense ring Reduce-Scatter, with a total of N - 1 rounds on N devices. Each device i distributes its (i + j) % N-th data block to the (i - 1 + N) % N-th device at Round j, and at the same time reduces the received (i + j + 1) % N-th data block. For example, including devices 0, 1, 2, 3, that is, N is 4. At round 1, device 0 distributes its (0 + 1) % 4, that is, the 1st data block, to the (0 - 1 + 4) % 4-th device, that is, device 3. In the prior art, communication optimization has been achieved from the perspective of communication volume based on gradient tensor sparsification, but the impact of the underlying heterogeneous network topology on communication performance has not been considered. Due to the pruning of dense tensors, each sparse gradient tensor has a different non-zero element distribution, and the general AllReduce for dense tensors cannot support the reduction of sparse gradient tensors. At the same time, most of the manually implemented AllReduce operators do not consider the distribution of the dynamically changing sparse tensor data volume during the reduction process on heterogeneous networks, and cannot fully utilize the underlying hardware to improve performance. As Figure 3The implementation of serial number 0 is an existing sparse tensor reduction algorithm process with excellent performance: Each device i distributes its data block numbered (i - j + N) % N to the device numbered (i - j + N) % N in Round j. Then each device reduces and sums the received data blocks, but there are problems of low network bandwidth utilization and inconvenient communication.
[0057] As Figure 2 shown, the embodiment of the present invention proposes a method for network bandwidth-aware sparse gradient reduction, including steps S1 - S4:
[0058] S1. Obtain the network topology structure;
[0059] S2. Divide the global devices into a two-dimensional device grid (n, d) according to the network topology structure, and at the same time divide the sparse tensor into a two-dimensional sparse tensor (d, n), where n represents the number of nodes and d represents the number of devices in each node;
[0060] S3. Perform communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) based on a preset Reduce-Scatter algorithm to determine the reduction result;
[0061] S4. Process the reduction result based on the AllGather algorithm to obtain the final processing result.
[0062] The working principle and beneficial effects of the above technical solution: Obtain the network topology structure; divide the global devices into a two-dimensional device grid (n, d) according to the network topology structure. The two-dimensional device network is a two-dimensional network obtained by dividing the device numbers of the global devices. The two-dimensional device grid includes two dimensions, namely the first dimension and the second dimension. Devices with the same subscript in the first dimension are in the same node and have high device communication bandwidth. Otherwise, they are cross-node devices with lower network bandwidth. At the same time, divide the sparse tensor into a two-dimensional sparse tensor (d, n), where n represents the number of nodes and d represents the number of devices in each node; the global devices are all processing devices; perform communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) based on a preset Reduce-Scatter algorithm to determine the reduction result; process the reduction result based on the AllGather algorithm to obtain the final processing result. Develop an efficient gradient reduction mode according to the underlying bandwidth to achieve efficient utilization of network bandwidth, relieve the low-bandwidth bottleneck, and improve training performance. This method can greatly optimize the performance of gradient reduction in distributed data parallel training under heterogeneous networks and make full use of network bandwidth.
[0063] In one embodiment, dividing the global devices into a two-dimensional device grid (n, d) according to the network topology structure and at the same time dividing the sparse tensor into a two-dimensional sparse tensor (d, n) includes:
[0064] Based on the network topology, determine the global device index. The global device index is a total of N devices from 0 to N-1. Abstract the network connection conditions into a two-dimensional matrix (n, d) and determine it as a two-dimensional device grid (n, d). Among them, n*d = N.
[0065] The sparse tensor includes sparse values and sparse indices, which are evenly divided into N parts and mapped to a two-dimensional structure of (d, n), and determined as a two-dimensional sparse tensor (d, n).
[0066] The beneficial effects of the above technical solution: It is convenient to accurately allocate the data processing process and improve the processing efficiency.
[0067] The optimization process of the Reduce-Scatter algorithm for sparse tensors is as Figure 3 shown. There are two devices on node A and node B respectively, each with the subscript and value of its own sparse tensor. We propose three algorithms: the first Reduce-Scatter algorithm, the second Reduce-Scatter algorithm, and the third Reduce-Scatter algorithm.
[0068] In one embodiment, the preset Reduce-Scatter algorithm is the first Reduce-Scatter algorithm.
[0069] Based on the first Reduce-Scatter algorithm, perform communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n), and determine the reduction result, including:
[0070] Determine that the id of each device gid in the two-dimensional device grid (n, d) is [r1, r2] = [gid / d, gid%d].
[0071] There are a total of N-1 rounds of communication. In the r-th round, distribute its own sparse tensor block [(r2 + r%d)%d, (r1 + r / d)%n] to the device numbered [(r1 + r / d)%n, (r2 + r%d)%d].
[0072] After completing N-1 rounds of communication, each device reduces and sums the received sparse tensor data blocks to determine the reduction result.
[0073] The working principle and beneficial effects of the above technical solution: The first Reduce-Scatter algorithm adjusts the execution order of the Reduce-Scatter algorithm, and adjusts the communication within the node to precede the cross-node communication. For example, device [0, 0] distributes its own sparse tensor block [1, 0] to device [0, 1] in the first round. This method can balance the use of the underlying heterogeneous network topology bandwidth. As Figure 3As shown, the figure corresponding to serial number 1 is the execution allocation schematic diagram of the first Reduce-Scatter algorithm, such as Figure 4 is the schematic diagram of the first Reduce-Scatter algorithm.
[0074] In one embodiment, the preset Reduce-Scatter algorithm is the second Reduce-Scatter algorithm;
[0075] Based on the second Reduce-Scatter algorithm, perform communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n), and determine the reduction result, including:
[0076] Determine that the id of each device gid in the two-dimensional device grid (n, d) is [r1, r2] = [gid / d, gid % d];
[0077] There are a total of n - 1 + d - 1 rounds of communication. In the first n - 1 rounds, in the r-th round, distribute its own sparse tensor data block of [*, (r1 - r + n) % n] to the device numbered [(r1 - r + n) % n, r2]; where * represents all the data in this dimension;
[0078] After completing n - 1 rounds of communication, each device reduces and sums the received sparse tensor data blocks to determine the first intermediate reduction result;
[0079] In the next d - 1 rounds, in the r-th round, distribute the sparse tensor data block of [(r2 - r + d) % d, r1] in the first intermediate reduction result to the device numbered [r1, (r2 - r + d) % d]; after completing d - 1 rounds of communication, each device reduces and sums the received sparse tensor data blocks to determine the reduction result.
[0080] The working principle and beneficial effects of the above technical solution: The second Reduce-Scatter algorithm first performs communication reduction of the second-dimensional sparse tensor of the sparse tensor within the devices with the same second dimension subscript in the two-dimensional device grid. Then, according to the intermediate reduction sparse tensor result, perform reduction of Reduce-Scatter within the devices with the same first dimension subscript in the two-dimensional device grid for the first dimension of the intermediate reduction sparse tensor. That is:
[0081] Determine that the id of each device gid in the two-dimensional device grid (n, d) is [r1, r2] = [gid / d, gid % d]; there are a total of n - 1 + d - 1 rounds of communication. In the first n - 1 rounds, in the r-th round, distribute its sparse tensor data block of [*, (r1 - r + n) % n] to the device numbered [(r1 - r + n) % n, r2]; where, * represents all the data in this dimension, specifically, the tensor formed by aggregating all the sparse tensor data blocks with the second coordinate being (r1 - r + n) % n. For example, device [0, 0] distributes its sparse tensor block [*, n - 1] to device [n - 1, 0] in the first round.
[0082] After completing n - 1 rounds of communication, each device performs a reduction sum on the received sparse tensor data blocks to determine the first intermediate reduction result; in the subsequent d - 1 rounds, in the r-th round, distribute the sparse tensor data block of [(r2 - r + d) % d, r1] in the first intermediate reduction result to the device numbered [r1, (r2 - r + d) % d]; after completing d - 1 rounds of communication, each device performs a reduction sum on the received sparse tensor data blocks to determine the reduction result. For example, device [0, 0] distributes its sparse tensor block [d - 1, 0] to device [0, d - 1] in the first round. This method can reduce the number of communication times across nodes, thereby reducing communication latency. As Figure 3 shown, the one corresponding to serial number 2 is the execution allocation schematic diagram of the second Reduce-Scatter algorithm. Figure 3 In it, Index is the sparse index and value is the sparse value.
[0083] In one embodiment, the preset Reduce-Scatter algorithm is the third Reduce-Scatter algorithm;
[0084] Based on the third Reduce-Scatter algorithm, perform communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) to determine the reduction result, including:
[0085] Determine that the id of each device gid in the two-dimensional device grid (n, d) is [r1, r2] = [gid / d, gid % d];
[0086] There are a total of d - 1 + n - 1 rounds of communication. In the first d - 1 rounds, in the r-th round, distribute its sparse tensor data block of [(r2 + r) % d, *] to the device numbered [r1, (r2 + r) % d], where, * represents all the data in this dimension;
[0087] After completing d - 1 rounds of communication, each device performs a reduction sum on the received sparse tensor data blocks to determine the second intermediate reduction result;
[0088] In the last n - 1 rounds, in the r-th round, the sparse tensor data block in the [r2, (r1 + r) % n] of the second intermediate reduction result is distributed to the device numbered [(r1 + r) % n, r2];
[0089] After completing n - 1 rounds of communication, each device reduces and sums the received sparse tensor data blocks to determine the reduction result.
[0090] The working principle and beneficial effects of the above technical solution: The third Reduce-Scatter algorithm first performs communication reduction of the first-dimensional sparse tensor of the sparse tensor within the devices with the same first-dimensional subscript in the two-dimensional device grid. Then, according to the intermediate reduction sparse tensor result, reduction of the second dimension of the intermediate reduction sparse tensor is performed within the devices with the same second-dimensional subscript in the two-dimensional device grid. That is:
[0091] Determine that the id of each device gid in the two-dimensional device grid (n, d) is [r1, r2] = [gid / d, gid % d]; There are a total of d - 1 + n - 1 rounds of communication. In the first d - 1 rounds, in the r-th round, the [(r2 + r) % d, *] sparse tensor data block of itself is distributed to the device numbered [r1, (r2 + r) % d], where * represents all the data in this dimension, specifically the tensor formed by aggregating all the sparse tensor data blocks with the first coordinate of (r2 + r) % d; For example, the device [0, 0] distributes its sparse tensor block [1, *] to the device [0, 1] in the first round.
[0092] After completing d - 1 rounds of communication, each device reduces and sums the received sparse tensor data blocks to determine the second intermediate reduction result;
[0093] In the last n - 1 rounds, in the r-th round, the sparse tensor data block in the [r2, (r1 + r) % n] of the second intermediate reduction result is distributed to the device numbered [(r1 + r) % n, r2];
[0094] After completing n - 1 rounds of communication, each device reduces and sums the received sparse tensor data blocks to determine the reduction result. For example, the device [0, 0] distributes its sparse tensor block [0, 1] to the device [1, 0] in the first round. This method advances the order of cross-node communication to after in-node communication, thereby reducing the bottleneck of cross-node communication. As Figure 3 shown, the one corresponding to serial number 3 is the execution allocation schematic diagram of the third Reduce-Scatter algorithm. As Figure 5 is the schematic diagram of the second Reduce-Scatter algorithm and the third Reduce-Scatter algorithm.
[0095] In one embodiment, the reduction result is processed based on the AllGather algorithm to obtain the final processing result, including:
[0096] For any device coordinates [r1, r2], first aggregate the reduction results on devices with the same second-dimensional coordinate r2 to obtain an intermediate aggregated sparse tensor of [r2, *]; where * represents all data in this dimension;
[0097] Devices with the same first-dimensional coordinate r1 aggregate the intermediate aggregated sparse tensor [r2, *] to obtain the final processing result.
[0098] The working principle and beneficial effects of the above technical solution: The reverse determination of the third Reduce-Scatter algorithm is the communication order of the AllGather algorithm; for any device coordinates [r1, r2], first aggregate the reduction results on devices with the same second-dimensional coordinate r2 to obtain an intermediate aggregated sparse tensor of [r2, *]; where * represents all data in this dimension; specifically, * represents the tensor formed by aggregating all sparse tensor data blocks with the first coordinate r2.
[0099] Devices with the same first-dimensional coordinate r1 aggregate the intermediate aggregated sparse tensor [r2, *] to obtain the final processing result. For example, device [0, 0] first aggregates with device [*, 0] to obtain the intermediate sparse tensor aggregation result of [0, *]. Then it aggregates with device [0, *] to obtain the aggregation result of all sparse tensors. This is convenient for improving the aggregation efficiency and processing efficiency, and can greatly optimize the performance of gradient reduction in distributed data parallel training under heterogeneous networks, making full use of the network bandwidth. Compared with the existing Allgather implementation, the hierarchical topology-aware reduction order is adopted, which improves the performance.
[0100] In one embodiment, it further includes performing VGG16, LSTM, and BERT task training on 2-node 4-card, 2-node 8-card, and 4-node 8-card hardware devices respectively.
[0101] Using the bandwidth-aware sparse gradient reduction method of the present invention, under the above cross-bandwidth network topology, perform VGG16, LSTM, and BERT task training on 2-node 4-card, 2-node 8-card, and 4-node 8-card hardware devices respectively. The results show that it has little impact on the accuracy of the model reaching the convergence state. From the test results, after adopting this method, under the above different configurations, the throughput of model training has a large increase. It basically meets the requirements of efficient data parallel model training. The present invention is simple to implement, effectively accelerates, and meets the application requirements.
[0102] Such as Figure 6As shown in the figure, an embodiment of the second aspect of the present invention provides an apparatus for network bandwidth-aware sparse gradient reduction, including:
[0103] An acquisition module, configured to acquire a network topology structure;
[0104] A partitioning module, configured to divide global devices into a two-dimensional device grid (n, d) according to the network topology structure, and at the same time divide a sparse tensor into a two-dimensional sparse tensor (d, n), where n represents the number of nodes and d represents the number of devices in each node;
[0105] A determination module, configured to perform communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) based on a preset Reduce-Scatter algorithm, and determine a reduction result;
[0106] A processing module, configured to process the reduction result based on an AllGather algorithm to obtain a final processing result.
[0107] To achieve the above object, an embodiment of the third aspect of the present invention provides a terminal device, where the terminal device includes a memory, a processor, and a network bandwidth-aware sparse gradient reduction program stored in the memory and executable on the processor. When the network bandwidth-aware sparse gradient reduction program is executed by the processor, the steps of the network bandwidth-aware sparse gradient reduction method as described above are implemented.
[0108] To achieve the above object, an embodiment of the fourth aspect of the present invention provides a computer medium, where the medium is a computer-readable storage medium, and a network bandwidth-aware sparse gradient reduction program is stored on the computer-readable storage medium. When the network bandwidth-aware sparse gradient reduction program is executed by a processor, the steps of the network bandwidth-aware sparse gradient reduction method as described above are implemented.
[0109] The beneficial effects of the apparatus, terminal device, and computer medium for network bandwidth-aware sparse gradient reduction proposed by the present invention are consistent with those of the network bandwidth-aware sparse gradient reduction method, and will not be elaborated here.
[0110] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A method for network bandwidth-aware sparse gradient reduction, characterized in that Including: Obtain the network topology structure; According to the network topology structure, divide the global devices into a two-dimensional device grid (n, d), and at the same time divide the sparse tensor into a two-dimensional sparse tensor (d, n), where n represents the number of nodes and d represents the number of devices in each node; Based on the preset Reduce-Scatter algorithm, perform communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) to determine the reduction result; Based on the AllGather algorithm, process the reduction result to obtain the final processing result; According to the network topology structure, divide the global devices into a two-dimensional device grid (n, d), and at the same time divide the sparse tensor into a two-dimensional sparse tensor (d, n), including: Based on the network topology structure, determine the global device index. The global device index is a total of N devices from 0,..., N - 1; abstract it into a two-dimensional matrix (n, d) according to the network connection condition, and determine it as the two-dimensional device grid (n, d); where, n * d = N; The sparse tensor includes sparse values and sparse indices, both of which are divided into N parts and mapped to a two-dimensional structure of (d, n) to be determined as the two-dimensional sparse tensor (d, n).
2. The method for network bandwidth-aware sparse gradient reduction according to claim 1, wherein, The preset Reduce-Scatter algorithm is the first Reduce-Scatter algorithm; Based on the first Reduce-Scatter algorithm, perform communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) to determine the reduction result, including: Determine that the id of each device gid in the two-dimensional device grid (n, d) is [r1, r2] = [gid / d, gid % d]; There are a total of N - 1 rounds of communication. In the r-th round, distribute its own sparse tensor block [(r2 + r % d) % d, (r1 + r / d) % n] to the device numbered [(r1 + r / d) % n, (r2 + r % d) % d]; After completing N - 1 rounds of communication, each device reduces and sums the received sparse tensor data blocks to determine the reduction result.
3. The method for network bandwidth-aware sparse gradient reduction according to claim 1, characterized in that, The preset Reduce-Scatter algorithm is the second Reduce-Scatter algorithm; Based on the second Reduce-Scatter algorithm, perform communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) to determine the reduction result, including: Determine that the id of each device gid in the two-dimensional device grid (n, d) is [r1, r2] = [gid / d, gid % d]; There are a total of n - 1 + d - 1 rounds of communication. In the first n - 1 rounds, in the r-th round, distribute its own [*, (r1 - r + n) % n] sparse tensor data block to the device numbered [(r1 - r + n) % n, r2]; where, * represents all data in this dimension; After completing n - 1 rounds of communication, each device reduces and sums the received sparse tensor data blocks to determine the first intermediate reduction result; In the (d - 1) rounds after that, in the r-th round, the sparse tensor data block of [(r2 - r + d) % d, r1] in the first intermediate reduction result is distributed to the device numbered [r1, (r2 - r + d) % d]; after completing the (d - 1) rounds of communication, each device reduces and sums the received sparse tensor data blocks to determine the reduction result.
4. The method for network bandwidth-aware sparse gradient reduction according to claim 1, wherein The preset Reduce-Scatter algorithm is the third Reduce-Scatter algorithm; Based on the third Reduce-Scatter algorithm, performing communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) to determine the reduction result, including: Determine that the id of each device gid in the two-dimensional device grid (n, d) is [r1, r2] = [gid / d, gid % d]; There are a total of (d - 1) + (n - 1) rounds of communication. In the first (d - 1) rounds, in the r-th round, the sparse tensor data block of [(r2 + r) % d, *] of itself is distributed to the device numbered [r1, (r2 + r) % d], where * represents all the data in this dimension; After completing the (d - 1) rounds of communication, each device reduces and sums the received sparse tensor data blocks to determine the second intermediate reduction result; In the (n - 1) rounds after that, in the r-th round, the sparse tensor data block of [r2, (r1 + r) % n] in the second intermediate reduction result is distributed to the device numbered [(r1 + r) % n, r2]; After completing the (n - 1) rounds of communication, each device reduces and sums the received sparse tensor data blocks to determine the reduction result.
5. The method for network bandwidth-aware sparse gradient reduction according to claim 1, characterized in that Based on the AllGather algorithm, processing the reduction result to obtain the final processing result, including: For any device coordinates [r1, r2], first gather the reduction results on the devices with the same second-dimensional coordinate r2 to obtain the intermediate gathered sparse tensor of [r2, *]; where * represents all the data in this dimension; The devices with the same first-dimensional coordinate r1 gather the intermediate gathered sparse tensor [r2, *] to obtain the final processing result.
6. The method for network bandwidth-aware sparse gradient reduction according to claim 1, wherein It also includes training VGG16, LSTM, and BERT tasks on 2-node 4-card, 2-node 8-card, and 4-node 8-card hardware devices respectively.
7. An apparatus for network bandwidth-aware sparse gradient reduction, characterized in that, Including: An acquisition module for acquiring the network topology structure; A partitioning module for dividing the global devices into a two-dimensional device grid (n, d) according to the network topology structure, and at the same time dividing the sparse tensor into a two-dimensional sparse tensor (d, n), where n represents the number of nodes and d represents the number of devices in each node; A determination module for performing communication reduction on the two-dimensional device grid (n, d) and the two-dimensional sparse tensor (d, n) based on the preset Reduce-Scatter algorithm to determine the reduction result; A processing module for processing the reduction result based on the AllGather algorithm to obtain the final processing result; The method by which the partitioning module divides the global devices into a two-dimensional device grid (n, d) according to the network topology structure and at the same time divides the sparse tensor into a two-dimensional sparse tensor (d, n) includes: Dividing the global devices into a two-dimensional device grid (n, d) according to the network topology structure and at the same time dividing the sparse tensor into a two-dimensional sparse tensor (d, n), including: Based on the network topology, determine the global device index. The global device index is a total of N devices numbered from 0 to N-1. Abstract it into a two-dimensional matrix (n, d) according to the network connection conditions, and determine it as a two-dimensional device grid (n, d), where n*d = N. The sparse tensor includes sparse values and sparse indices, both of which are divided into N parts and mapped to a two-dimensional structure of (d, n), and determined as a two-dimensional sparse tensor (d, n).
8. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a program for network bandwidth-aware sparse gradient reduction stored in the memory and executable on the processor. When the program for network bandwidth-aware sparse gradient reduction is executed by the processor, it implements the steps of the method for network bandwidth-aware sparse gradient reduction according to any one of claims 1-6.
9. A computer medium, the medium being a computer-readable storage medium, characterized in that, A program for network bandwidth-aware sparse gradient reduction is stored on the computer-readable storage medium. When the program for network bandwidth-aware sparse gradient reduction is executed by the processor, it implements the steps of the method for network bandwidth-aware sparse gradient reduction according to any one of claims 1-6.