Data processing method and related equipment

By performing backpropagation and gradient compression operations in parallel in distributed training of neural networks, combining selective gradient compression and multi-threaded management, the gradient synchronization bottleneck problem is solved, and training efficiency and communication efficiency are improved.

CN120297367APending Publication Date: 2025-07-11HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410050300.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In distributed training of neural networks, gradient synchronization becomes a bottleneck that restricts training speed, and the computational overhead brought by existing gradient compression algorithms leads to less significant improvement in training efficiency.

Method used

By performing backpropagation and gradient compression operations of neural networks in parallel on different threads, using different hardware unit resources, reducing computing resource preemption, and selectively compressing gradients, combining multi-threaded management and gradient packaging transmission, optimizing communication efficiency.

Benefits of technology

It improves the distributed training efficiency of neural networks, reduces the overall data processing time, and improves training speed and network bandwidth utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297367A_ABST
    Figure CN120297367A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method, and the method comprises the steps: enabling a first calculation node to execute a calculation operation in a training process of a neural network of the first calculation node through a first thread, and obtaining a gradient; and compressing the gradient through the second thread to obtain a compressed gradient. Visibly, in the method, through the first thread and the second thread, processing such as compression of the generated gradient can be hidden in calculation processes such as back propagation of the neural network in time in the training process of the neural network, so that the distributed training efficiency of the neural network is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a data processing method and related devices. Background Art

[0002] Currently, as the number of parameters of neural networks increases, in order to meet the training requirements of neural networks, a distributed method can be used to train neural networks.

[0003] As the number of parameters of neural networks continues to rise, the amount of gradient data that needs to be synchronized between multiple distributed devices during the iterative process of neural networks becomes larger and larger, resulting in gradient synchronization becoming a bottleneck restricting the training speed in the distributed training of neural networks. For this reason, a gradient compression algorithm has been proposed to reduce the amount of gradient data during the gradient synchronization process, thereby alleviating the pressure of network communication and improving the training speed.

[0004] However, when using the gradient compression algorithm to compress the gradients of neural networks to reduce network communication pressure, operations such as gradient compression and decompression will bring additional computational overhead and introduce additional computational time consumption. Therefore, it is still possible that the duration of the overall training process cannot be effectively reduced, resulting in the benefits brought by introducing the gradient compression algorithm being not obvious, and the distributed training efficiency of neural networks not being significantly improved. Summary of the Invention

[0005] Embodiments of this application provide a data processing method, which can improve the distributed training efficiency of neural networks during the actual distributed training process. This application also provides corresponding devices, equipment, computer-readable storage media, computer program products, etc.

[0006] The first aspect of this application provides a data processing method. This method is applied to a first computing node and includes: through a first thread, performing a computing operation during the training process of the neural network of the first computing node to obtain a gradient; through a second thread, compressing the gradient to obtain a compressed gradient.

[0007] In the existing solutions, when using the gradient compression algorithm to compress the gradients of neural networks to reduce network communication pressure, researchers believe that both the backpropagation of neural networks and the compression and decompression of gradients require computational resources to be implemented. If the backpropagation of neural networks and the compression and decompression of gradients are executed in parallel, there will be a situation of preempting computational resources. Therefore, the backpropagation of neural networks and the compression and decompression of gradients are usually executed serially. For example, when performing backpropagation on a neural network through the main thread to obtain the gradients of the neural network, the computational resources occupied during backpropagation are released; then the main thread obtains computational resources to compress the gradients of the neural network.

[0008] This causes the traditional gradient compression algorithm-based solution to introduce additional computational time consumption, resulting in the inability to effectively reduce the duration of the overall training process, making the benefits brought by introducing the gradient compression algorithm not obvious, and the distributed training efficiency of the neural network not being significantly improved.

[0009] In the first aspect, the inventors believe that operations such as compression belong to memory access-intensive operators, while neural network computations such as backpropagation are computation-intensive operators, and the two rely on different hardware units in chips such as AI intelligent chips. Therefore, it is considered to implement different tasks through different threads (such as the first thread and the second thread).

[0010] For example, the first thread can be used to implement computational operations such as the backpropagation of the neural network. And during backpropagation, after obtaining the gradient of the nth layer of the neural network, the second thread can perform operations such as compressing the gradient of the nth layer, so that the resources of the first thread can continuously execute backpropagation of different layers such as the (n - 1)th layer of the neural network.

[0011] It can be seen that through the second thread and the first thread, mutual hiding (that is, parallel execution) of different types of tasks for different layers of the neural network (such as the nth layer and the (n - 1)th layer in the above example) can be achieved. Compared with the traditional solution where the computational operations of the neural network and the compression of gradients and other operations are serially executed, in the first aspect, through the first thread and the second thread, the compression and other processing of the gradients generated during the backpropagation process can be hidden in time from the computational processes such as the backpropagation of the neural network, thereby improving the distributed training efficiency of the neural network.

[0012] For example, during the backpropagation process of the neural network, after the gradient of the nth layer of the neural network is obtained through backpropagation, during the process of performing backpropagation on the (n - 1)th layer of the neural network, it can be considered to perform operations such as compressing the already obtained gradient of the nth layer of the neural network. That is to say, the compression and other operations of the gradient of the nth layer of the neural network by the second thread are executed in parallel with the backpropagation computational operation of the (n - 1)th layer of the neural network by the first thread, so that the compression and other operations of the already obtained gradient of the nth layer of the neural network are hidden in the backpropagation and other computational operations of the (n - 1)th layer of the neural network, thereby reducing the overall data processing time consumption and improving the training efficiency.

[0013] In a possible implementation of the first aspect, before compressing the gradients, it further includes: determining whether to compress the gradients, where the gradients that need to be compressed are the first gradients, and the gradients that do not need to be compressed are the second gradients; compressing the gradients includes: in the case of compressing the gradients, compressing the first gradients to obtain the compressed first gradients.

[0014] In practical applications, the gradients of a neural network can be compressed through a gradient compression algorithm to reduce the data volume of the gradients during the gradient synchronization process, thereby alleviating the pressure of network communication.

[0015] However, when using a gradient compression algorithm to compress the gradients of a neural network, operations such as gradient compression and decompression will bring additional computational overhead and introduce additional computational time consumption. Therefore, it is still possible that the overall training process duration cannot be effectively reduced, resulting in the benefits brought by introducing the gradient compression algorithm being not obvious and the distributed training efficiency of the neural network not being significantly improved.

[0016] Specifically, in some practical application scenarios, the magnitudes of the gradients of a neural network vary from a few bytes to hundreds of megabytes. Compressing small gradients may not only fail to reduce the network communication latency but may instead increase the additional computational overhead of compression and decompression, thereby having a negative impact on the actual processing efficiency.

[0017] In this possible implementation, it is possible to determine whether to compress the gradients to identify the gradients that need to be compressed as the first gradients and the gradients that do not need to be compressed as the second gradients. Thus, through a selective gradient compression scheme, the negative impact on the actual processing efficiency caused by unreasonable gradient compression (such as compressing small gradients) can be reduced or even avoided, ensuring the benefits brought by gradient compression, and thereby more effectively improving the distributed training efficiency of the neural network.

[0018] In a possible implementation of the first aspect, determining whether to compress the gradients includes: determining whether to compress the gradients according to a first preset condition, where the first preset condition is determined based on the communication overheads of the gradients before and after compression.

[0019] In this possible implementation, when calculating the communication overheads of the gradients before and after compression, it may not be necessary to actually perform compression operations and / or transmission operations on the gradients. Instead, based on pre-collected relevant parameter information, the communication overheads based on the compressed gradients and the communication overheads based on the gradients before compression can be estimated to determine whether to compress the gradients.

[0020] In a possible implementation of the first aspect, determining whether to compress the gradient according to the first preset condition includes: determining whether to compress the gradient according to the first preset condition through a third thread.

[0021] During one iteration of the neural network, it is usually to perform backpropagation layer by layer on multiple layers of the neural network. Therefore, the gradients of the layers of the neural network are also generated layer by layer, rather than being generated all at once.

[0022] In this possible implementation, after each new gradient is generated through calculation operations such as backpropagation, the third thread can obtain the gradient from the first thread and determine whether to compress the gradient according to the first preset condition. For example, during one iteration, after the gradient of the nth layer of the neural network is obtained through the calculation operation of the first thread, the gradient of the nth layer of the neural network can be passed from the first thread to the third thread, and during this iteration, the gradient of the (n - 1)th layer of the neural network has usually not been obtained through calculation operations yet.

[0023] In this way, through the first thread, the second thread, and the third thread, the calculation operations of the neural network and subsequent processing operations such as compressing the generated gradients can be managed and processed separately, providing a basis for parallel processing of operations such as calculation and compression, and enabling operations such as calculation and compression to be hidden from each other.

[0024] In a possible implementation of the first aspect, after determining whether to compress the gradient according to the first preset condition through the third thread, it further includes: in the case of compressing the gradient, storing the first gradient in the compression task queue through the third thread; compressing the gradient through the second thread to obtain the compressed gradient, including: obtaining the first gradient from the compression task queue through the second thread for compression.

[0025] In this possible implementation, the first gradient can be passed to the compression task queue through the third thread, so that the second thread that manages the compression task queue triggers the calculation unit to perform the compression operation on the first gradient to obtain the compressed first gradient.

[0026] In a possible implementation of the first aspect, the communication overhead is determined based on one or more of the following information: the network transmission method of the gradient, the network transmission speed, the compression speed of the gradient, and the decompression speed of the compressed gradient.

[0027] In this possible implementation, a method for selectively compressing gradients is proposed, taking into account aspects such as the magnitude of the gradients, compression and decompression, and data transmission overhead. Among them, the communication overhead corresponding to the compressed and uncompressed gradients can be theoretically analyzed, and a first preset condition can be pre-constructed according to performance parameters such as transmission and compression and decompression, so as to determine whether any gradient needs to be compressed according to the first preset condition, reducing or even avoiding the negative impact on the actual processing efficiency caused by inappropriate compression operations such as small gradient compression, ensuring the benefits brought by gradient compression, and thus more effectively improving the distributed training efficiency of the neural network.

[0028] In a possible implementation of the first aspect, after compressing the first gradient to obtain the compressed first gradient, it further includes: sending the compressed first gradient through the fourth thread.

[0029] In this possible implementation, the computing operations of the neural network and subsequent processing operations such as compression and communication of the generated gradients can be managed and processed respectively through the first thread, the second thread, the third thread, and the fourth thread, providing a basis for parallel processing of operations such as computing, compression, and communication, and enabling operations such as computing, compression, and communication to be hidden from each other.

[0030] In a possible implementation of the first aspect, the second gradient and the compressed first gradient are stored in the communication task queue; sending the compressed first gradient through the fourth thread includes: when the second gradient and the compressed first gradient in the communication task queue meet the second preset condition, packing and sending the second gradient and the compressed first gradient through the fourth thread.

[0031] In the traditional solution, each time gradient synchronization is performed, each single small gradient calls a communication primitive to implement network transmission, which will lead to a reduction in network bandwidth utilization and an increase in communication time consumption.

[0032] To solve the above problems, in this possible implementation, it is considered to pack the compressed gradient and the uncompressed gradient together and then perform network transmission to improve network bandwidth utilization. Specifically, when it is detected that the second gradient and the compressed first gradient stored in the communication task queue meet the second preset condition, the second gradient and the compressed first gradient can be packed to obtain the packed data, so as to send the packed data to the second computing node among multiple computing nodes. For example, when using the allreduce communication primitive for transmission, the packed data can achieve data synchronization among each computing node through one allreduce communication primitive.

[0033] In a possible implementation of the first aspect, after determining whether to compress the gradient, it further includes: when not compressing the gradient, sending the second gradient through the fourth thread.

[0034] In a possible implementation of the first aspect, the method further includes: receiving the gradient of the second computing node, where the second computing node is a computing node other than the first computing node, and the gradient of the second computing node includes the compressed gradient; decompressing the gradient of the second computing node through the fifth thread.

[0035] In this possible implementation, the gradient of the second computing node can be decompressed through the fifth thread. The fifth thread is different from the above-mentioned first thread, second thread, third thread, and fourth thread. Therefore, this decompression operation can be executed in parallel with computing operations, compression operations, communication operations, etc., improving the data processing efficiency.

[0036] In a possible implementation of the first aspect, the received gradient of the second computing node is stored in the decompression task queue; decompressing the gradient of the second computing node through the fifth thread includes: obtaining the gradient of the second computing node from the decompression task queue through the fifth thread for decompression.

[0037] In this possible implementation, the fifth thread managing the decompression task queue can determine the execution timing of the decompression task for the gradient of the second computing node according to the execution situation of the decompression task and the resource situation of the computing unit, etc., and instruct the computing unit to execute the decompression task for the gradient of the second computing node to obtain the decompressed gradient of the second computing node.

[0038] The second aspect of the present application provides a data processing device, which has the function of implementing the method of the first aspect or any possible implementation of the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, such as a computing module and a processing module.

[0039] The third aspect of the present application provides a computing node, which includes at least one processor, a memory, and computer execution instructions stored in the memory and executable on the processor. When the computer execution instructions are executed by the processor, the processor executes the method of the first aspect or any possible implementation of the first aspect as described above.

[0040] The fourth aspect of the present application provides a computer-readable storage medium storing one or more computer execution instructions. When the computer execution instructions are executed by the processor, the processor executes the method of the first aspect or any possible implementation of the first aspect as described above.

[0041] The fifth aspect of this application provides a computer program product for storing one or more computer-executable instructions. The computer program product includes computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor executes the method according to the first aspect or any possible implementation manner of the first aspect as described above.

[0042] The sixth aspect of this application provides a chip system. The chip system includes a processor for supporting a computing node to implement the functions involved in the first aspect or any possible implementation manner of the first aspect as described above. In a possible design, the chip system may further include a memory for storing necessary program instructions and data. The chip system may be composed of chips or may include chips and other discrete devices.

[0043] Among them, for the technical effects brought by the second aspect to the sixth aspect or any possible implementation manner thereof, reference may be made to the technical effects brought by the first aspect or the related possible implementation manners of the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is an exemplary schematic diagram of the distributed training architecture of the neural network provided by an embodiment of this application;

[0045] Figure 2a is an exemplary schematic diagram of the traditional data processing flow provided by an embodiment of this application;

[0046] Figure 2b is an exemplary schematic diagram of the improved data processing flow provided by an embodiment of this application;

[0047] Figure 3 is an embodiment schematic diagram of the data processing method provided by an embodiment of this application;

[0048] Figure 4 is an embodiment schematic diagram of the data processing method provided by an embodiment of this application;

[0049] Figure 5 is an exemplary schematic diagram of a low-rank gradient compression algorithm provided by an embodiment of this application;

[0050] Figure 6 is an exemplary schematic diagram of multi-thread management provided by an embodiment of this application;

[0051] Figure 7 is an exemplary schematic diagram of packing and sending data provided by an embodiment of this application;

[0052] Figure 8a is an exemplary schematic diagram of a test result of the training speed provided by an embodiment of this application;

[0053] Figure 8b It is an exemplary schematic diagram of another test result of the training speed provided by an embodiment of the present application;

[0054] Figure 9 It is a schematic diagram of an embodiment of a data processing device provided by an embodiment of the present application;

[0055] Figure 10 It is a schematic diagram of a structure of a computing node provided by an embodiment of the present application. Detailed implementation manners

[0056] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the implementation manners part of the present application are only used to explain the specific embodiments of the present application, rather than to limit the present application.

[0057] As can be known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0058] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. The terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these process, method, product or device.

[0059] In the field of artificial intelligence, due to the strong learning and understanding capabilities of large artificial intelligence (AI) models, they have shown outstanding performance in fields such as natural language processing, dialogue systems, and image processing. With the continuous development of technology, the number of parameters of AI large models is increasing, showing an exponential growth trend. Currently, the number of parameters has exceeded the trillion level and is moving towards the trillion scale. At the same time, the training datasets are also getting larger and larger. This has made it difficult for a single machine to meet the needs of AI large models in terms of both storage capacity and computing power. To address this, one solution is to use a distributed approach to train neural networks.

[0060] There are various ways to use a distributed approach to train neural networks.

[0061] In one example, neural network training can be carried out through distributed devices in the dimension of data parallelism.

[0062] Data parallelism is a distributed approach to accelerate neural network training. In a data parallel scenario, the local data of different computing nodes can be regarded as different parts of the training dataset. Each computing node independently trains the neural network using its own local data, and the gradients of the neural network generated during the training process are synchronized across the network among the nodes. For example, a commonly used synchronization framework is to use the allreduce primitive in the field of high performance computing (HPC) for communication.

[0063] Among them, the gradient of the neural network refers to the gradient of the weights of the neural network, which is generated through the backpropagation of the neural network. This gradient can be in the form of a tensor, reflecting the deviation of the corresponding weights. During the iteration process, the weights of the neural network can be updated according to the gradients obtained from the backpropagation of the neural network to obtain the updated weights.

[0064] Such as Figure 1 shown in a schematic diagram of traditional data parallel training.

[0065] In Figure 1In the illustrated example, a neural network is trained by 4 computing nodes (Node0, Node1, Node2, and Node3 respectively). During the training process, a complete neural network is deployed on each computing node, and the training data in the training dataset is sliced into 4 parts and distributed to the 4 computing nodes respectively. Each computing node can load a part of the training data locally and train the neural network separately. Any computing node can obtain the gradients of the neural network through forward propagation and backward propagation based on the locally loaded training data. In this way, a total of 4 gradient information is generated by the 4 computing nodes. Through cross-node information transmission via the network, gradient synchronization among the 4 computing nodes can be achieved, so that each computing node can obtain the 4 gradient information, and each computing node can update the neural network according to the 4 gradient information respectively.

[0066] As the number of parameters of the neural network continues to increase, the amount of data of the gradients to be synchronized becomes larger and larger, resulting in gradient synchronization becoming a bottleneck restricting the training speed in the distributed training of the neural network. Moreover, in recent years, with the rapid development of computing power and the rise of compilation technology, computing operations such as forward propagation and backward propagation during the training process are more efficient, which means that gradient synchronization will be more frequent, resulting in the bottleneck problem of gradient synchronization becoming more serious, and even the time-consuming proportion in end-to-end training is more than 90%.

[0067] Based on this, the embodiments of the present application provide a data processing method, which can improve the training efficiency in the actual distributed training process.

[0068] The embodiments of the present application can be applied to a computing device cluster.

[0069] The computing device cluster may include multiple computing nodes, and the multiple computing nodes are used to implement distributed neural network training. The multiple computing nodes involved in the embodiments of the present application can be considered as the computing nodes that need to perform gradient synchronization during the distributed training of the neural network. The neural networks to be trained in different computing nodes can be the same or partially the same.

[0070] Among them, the type of the neural network is not limited here. Exemplarily, the neural network can be one or more of a recurrent neural network (RNN), a long short term memory network (LSTM), a convolutional neural network (CNN), a Transformer model, etc.

[0071] The specific application scenarios of this neural network are not limited here. Exemplarily, this neural network can be one or more of a natural language processing model, an image processing model, a speech processing model, and a graph neural network, etc.

[0072] The specific type of each computing node is not limited here. Exemplarily, each computing node can be a terminal device, a single server, or a server cluster, and can also be a virtual machine (VM) or a container, etc.

[0073] The specific types and functions of each computing node can be the same or different, which are not limited here. Each computing node includes but is not limited to a computing function. Moreover, multiple computing nodes in the computing node cluster can be located in the same physical location, for example, in the same data center; or, they can also be located in different physical locations. For example, any two computing nodes are located in different data centers.

[0074] Different computing nodes are connected through a network, and the specific type of this network and the communication method adopted are not limited here.

[0075] Exemplarily, different computing nodes can adopt communication methods such as allreduce communication method, parameter server communication method, etc. for cross-node data synchronization. For example, to achieve parameter synchronization of the neural network of each computing node.

[0076] The specific form of the distributed training of the neural network implemented by multiple computing nodes is not limited here. Exemplarily, in this distributed training, distributed training can be carried out based on the dimension of data parallelism. In addition, a three-dimensional (3D) parallel technology can also be used for the distributed training of the neural network. In this three-dimensional parallel technology, it includes data parallelism, tensor parallelism, and pipelining parallelism.

[0077] In addition, the specific form of the system software stack architecture in each computing node is also not limited here.

[0078] In some examples, in any computing node, it can include a bottom-layer hardware acceleration chip and network link, a chip operator library in the layer above the bottom layer, a computing framework in the layer above that, and a communication library in the even higher layer. In this example, the data processing method implemented can involve the communication library level. For example, it can be implemented using the allreduce communication method without modifying the logic at the bottom layer of the software stack.

[0079] Some traditional theories hold that when using gradient compression algorithms to compress the gradients of neural networks to reduce network communication pressure, both the calculations such as backpropagation of neural networks and the compression and decompression of gradients require the use of computing resources to be implemented. If the backpropagation of neural networks and the compression and decompression of gradients are executed in parallel, there will be a situation of competing for computing resources. Therefore, the backpropagation of neural networks and the compression and decompression of gradients are usually executed serially. For example, after the backpropagation of the neural network is performed through the main thread to obtain the gradients of the neural network, the computing resources occupied during backpropagation are released; then the main thread obtains the computing resources to compress the gradients of the neural network.

[0080] As Figure 2a shown, during the training process of the neural network, through computational operations such as backpropagation T11, the gradients of the (n + 1)-th layer can be obtained. After the computational operation T11 ends, the computing resources are released, and then through the compression operation T12, the gradients of the (n + 1)-th layer can be compressed. After compression, through the communication operation T13, the compressed gradients of the (n + 1)-th layer are sent to other computing nodes. Then, referring to the relevant operations of the (n + 1)-th layer, the computational operation T21, compression operation T22, and communication operation T23 for the n-th layer are sequentially executed to send the compressed gradients of the n-th layer to other computing nodes, and then the computational operation T31, compression operation T32, communication operation T33, and subsequent operations for the (n - 1)-th layer are sequentially executed.

[0081] This will cause the traditional gradient compression algorithm-based scheme to introduce additional computational time consumption, resulting in the fact that the overall duration of the training process cannot be effectively reduced, making the benefits brought by introducing the gradient compression algorithm not obvious, and the distributed training efficiency of the neural network has not been significantly improved.

[0082] However, the inventors of the embodiments of the present application believe that the calculations such as backpropagation of neural networks are computationally intensive and have a relatively high occupancy rate of computing units such as computing cores. However, compression and decompression are memory access intensive, with usually few computational operations, but often need to repeatedly scan data, having a relatively high demand for memory bandwidth but less demand for computing units.

[0083] Therefore, their dependencies on hardware are different, one depends on computing units and the other depends on memory units. Therefore, resource conflicts usually do not occur and they can be executed concurrently.

[0084] For example, as Figure 2bIn the illustrated example, during the backpropagation process of a neural network, after the gradient of the neural network of the (n + 1)-th layer is obtained through computational operations such as backpropagation (operation T11'), during the process of performing computational operations such as backpropagation (operation T21') on the neural network of the n-th layer, operations such as compression operation T12' and communication operation T13' on the obtained gradient of the neural network of the (n + 1)-th layer can be considered. That is to say, operations such as compression operation T12' and communication operation T13' on the gradient of the neural network of the (n + 1)-th layer are executed in parallel with computational operations such as backpropagation (operation T21') on the neural network of the n-th layer, so that operations such as compression operation T12' and communication operation T13' on the obtained gradient of the neural network of the (n + 1)-th layer are hidden in computational operations such as backpropagation (operation T21') on the neural network of the n-th layer. Similarly, operations such as compression operation T22' and communication operation T23' on the gradient of the neural network of the n-th layer are executed in parallel with computational operations such as backpropagation (operation T31') on the neural network of the (n - 1)-th layer, so that operations such as compression operation T22' and communication operation T23' on the obtained gradient of the neural network of the n-th layer are hidden in computational operations such as backpropagation (operation T31') on the neural network of the (n - 1)-th layer, thereby reducing the overall data processing time and improving the training efficiency.

[0085] Next, any one of the multiple computing nodes in the computing node cluster is referred to as a first computing node, and taking this first computing node as an example, the data processing method in the embodiments of the present application is introduced.

[0086] As Figure 3 shown, the data processing method may include steps 301 - 302.

[0087] Step 301, through a first thread, execute the computational operations during the training process of the neural network of the first computing node to obtain a gradient.

[0088] In the embodiments of the present application, during the computational process of the neural network of the first computing node, the computational operations of the neural network of the first computing node can be executed through a first thread. The computational operations may include backpropagation of the neural network and may also include forward propagation of the neural network, etc.

[0089] The first thread may be one or more threads in the first process of the first computing node. For example, the first thread may be the main process in the first process.

[0090] In step 301, the obtained gradient may be a partial gradient of the neural network. For example, when performing computational operations such as backpropagation on a certain layer of the neural network, one or more gradients are generated. The number of gradients in step 301 is not limited herein.

[0091] Step 302: Compress the gradients through the second thread to obtain the compressed gradients.

[0092] In the embodiments of the present application, the second thread is a thread different from the first thread. For example, the second thread may be a thread in the first process, or the second thread may be a thread split from the first process, or the second thread may also be a thread in another process other than the first process.

[0093] In the embodiments of the present application, part or all of the content in the gradients may be compressed. For example, if the gradients include a first gradient that needs to be compressed and a second gradient that does not need to be compressed, the first gradient in the gradients may be compressed.

[0094] It can be seen that in the embodiments of the present application, considering that operations such as compression are memory access-intensive operators, while neural network calculations such as backpropagation are compute-intensive operators, and the two rely on different hardware units in chips such as AI intelligent chips. Therefore, in the embodiments of the present application, different threads (such as the first thread and the second thread) are considered to implement different tasks.

[0095] For example, the first thread can be used to implement calculation operations such as backpropagation of the neural network. And after obtaining the gradients of the nth layer of the neural network during backpropagation, the second thread can perform operations such as compression on the gradients of the nth layer, so that the resources of the first thread can continuously execute backpropagation of different layers such as the (n - 1)th layer of the neural network.

[0096] It can be seen that through the second thread and the first thread, mutual hiding (that is, parallel execution) of different types of tasks for different layers of the neural network (such as the subsequent processing such as compression of the nth layer and the backpropagation of the (n - 1)th layer in the above example) can be achieved. Compared with the solution in the traditional technology where the calculation operations of the neural network and the operations such as compression of the gradients are serially executed, in the solution of the embodiments of the present application, through the first thread and the second thread, the processing such as compression of the gradients generated during backpropagation can be hidden in time in the calculation process such as backpropagation of the neural network, thereby improving the distributed training efficiency of the neural network.

[0097] For example, during the backpropagation of a neural network, after the gradient of the neural network in the n-th layer is obtained through backpropagation, during the backpropagation of the neural network in the (n - 1)-th layer, operations such as compression can be considered for the obtained gradient of the neural network in the n-th layer. That is to say, the operations such as compression of the gradient of the neural network in the n-th layer by the second thread are executed in parallel with the backpropagation calculation operation of the neural network in the (n - 1)-th layer by the first thread, so that the operations such as compression of the obtained gradient of the neural network in the n-th layer are hidden in the calculation operations such as backpropagation of the neural network in the (n - 1)-th layer, thereby reducing the overall data processing time and improving the training efficiency.

[0098] In some embodiments, after performing the calculation operations in the training process of the neural network of the first computing node through the first thread and obtaining the gradient, it can be determined whether to compress the gradient.

[0099] Specifically, as Figure 4 shown, in some embodiments, after step 301, step 303 can be executed, and based on the execution result of step 303, the subsequent execution steps 304 or 305 can be determined.

[0100] Among them, step 303 includes:

[0101] Determine whether to compress the gradient.

[0102] The gradient that needs to be compressed is the first gradient, and the gradient that does not need to be compressed is the second gradient.

[0103] In practical applications, the gradient of the neural network can be compressed through a gradient compression algorithm to reduce the amount of gradient data during the gradient synchronization process, thereby alleviating the pressure of network communication and improving the training speed.

[0104] However, when using a gradient compression algorithm to compress the gradient of the neural network, operations such as gradient compression and decompression will bring additional computational overhead and introduce additional computational time. Therefore, it is still possible that the overall training process duration cannot be effectively reduced, resulting in the benefits brought by introducing the gradient compression algorithm not being obvious and the distributed training efficiency of the neural network not being significantly improved.

[0105] Specifically, in some practical application scenarios, the size of the gradient of the neural network ranges from a few bytes to hundreds of megabytes. Compressing small gradients may not only not reduce the network communication delay but may instead increase the computational overhead of additional compression and decompression, thereby having a negative impact on the actual processing efficiency.

[0106] Based on this, in the embodiments of the present application, it is possible to determine whether to compress the gradients, so as to determine that the gradients that need to be compressed are the first gradients, and the gradients that do not need to be compressed are the second gradients, thereby reducing the impact of inappropriate compression and decompression operations on the processing efficiency.

[0107] Specifically, in the embodiments of the present application, information in multiple dimensions such as transmission overhead and additional overhead brought by compression and decompression operations can be comprehensively considered to determine whether to compress the gradients. Thus, through a selective gradient compression scheme, the negative impact on the actual processing efficiency caused by unreasonable gradient compression (such as compressing small gradients) can be reduced or even avoided, ensuring the benefits brought by gradient compression, and thus more effectively improving the distributed training efficiency of the neural network.

[0108] In some embodiments, step 303 includes:

[0109] Determine whether to compress the gradients according to a first preset condition.

[0110] The first preset condition is determined based on the communication overhead of the gradients before and after compression.

[0111] Among them, when the number of gradients is multiple, it is possible to determine whether to compress any gradient according to the first preset condition.

[0112] When calculating the communication overhead of the gradients before and after compression, it is not necessary to actually perform compression operations and / or transmission operations on the gradients. Instead, based on pre-collected relevant parameter information, the communication overhead of the gradients after compression and the communication overhead of the gradients before compression can be estimated to determine whether to compress the gradients.

[0113] When calculating the above communication overhead, there can be various types of relevant parameter information involved.

[0114] In some embodiments, the communication overhead is determined based on one or more of the following information:

[0115] The network transmission method of the gradients, network transmission speed, compression speed of the gradients, decompression speed of the compressed gradients.

[0116] The above parameter information can be pre-collected through tests and other means.

[0117] Among them, the communication overhead based on the gradients after compression can include the time taken to compress the gradients, the time taken to transmit the compressed gradients between multiple computing nodes so that each computing node obtains the compressed gradients, and the time taken to decompress the compressed gradients. And the communication overhead based on the gradients before compression can include the time taken to transmit the gradients between multiple computing nodes so that each computing node obtains the gradients.

[0118] Taking the data synchronization method between multiple computing nodes as the synchronization method using the allreduce algorithm and the gradient compression algorithm as the low-rank gradient compression algorithm as an example, the process of determining whether to compress the gradient according to the first preset condition will be introduced exemplarily.

[0119] Currently, a commonly used gradient compression algorithm is the low-rank compression algorithm. During the training process, the gradient matrix generated by the neural network is relatively sparse and has a small rank. Using the concept of linear algebra, a low-rank gradient matrix can be decomposed into two matrices through QR decomposition. For example, as Figure 5 shown, an M*N matrix can be decomposed into L k = m*k and R K T = k*n matrices, where k is the rank of the matrix, so as to achieve the purpose of compressing the gradient and reducing the communication volume. At this time, the compression ratio is ((m + n)*k) / (m*k).

[0120] When the data synchronization method between multiple computing nodes is the synchronization method using the allreduce algorithm, the gradient synchronization between multiple computing nodes can be achieved through a set of allreduce communication primitives. Among them, communication primitives refer to a set of communication templates obtained by combining one or more basic operations such as send, receive, and copy.

[0121] Based on the characteristics of the low-rank gradient compression algorithm, when the data synchronization method between multiple computing nodes is the allreduce algorithm synchronization method, the gradients compressed by the low-rank gradient compression algorithm can be directly accumulated. Therefore, the gradients of each computing node (which can have uncompressed gradients and compressed gradients) can use the same set of allreduce communication primitives for accumulation to achieve gradient synchronization between multiple computing nodes. After each computing node obtains the accumulated data, it can obtain the gradient to be synchronized after performing a decompression operation once, without performing multiple compression and decompression operations at each computing node during the communication process.

[0122] It can be seen that if the data synchronization method between multiple computing nodes is the synchronization method using the allreduce algorithm and the gradient compression algorithm is the low-rank gradient compression algorithm, usually only the gradient needs to be compressed before communication, and the accumulated data needs to be decompressed after communication, while a large number of compression and decompression operations can be reduced during the communication process, greatly improving the data synchronization efficiency between multiple computing nodes.

[0123] In this example, if the data synchronization method between computing nodes is the synchronization method using the allreduce algorithm and the gradient compression algorithm is the low-rank compression algorithm, in an exemplary first preset condition, the calculation formulas for the communication overhead based on the compressed gradient and the communication overhead based on the uncompressed gradient are generated based on the following method.

[0124] Before generating the first preset condition, exemplarily, the parameters in Table 1 can be defined.

[0125] Table 1: Parameters related to the first preset condition

[0126] Parameter Explanation m Gradient magnitude, in Byte K Number of segments into which the gradient is split during communication N Number of computing nodes participating in training r Gradient compression ratio <![CDATA[T enc (m)]]> Time taken to compress m Bytes of gradient <![CDATA[T dec (m)]]> Time taken to decompress m Bytes of gradient <![CDATA[T send (m)]]> Time taken to send m Bytes of gradient point-to-point

[0127] Using the allreduce communication primitive to implement the gradient synchronization operation with a gradient size of m among N computing nodes, the gradient needs to be split into K parts for network transmission (where K is not greater than N), and after 2(N - 1) steps of communication, all computing nodes obtain the synchronized gradient. Then, the communication overhead based on the uncompressed gradient can be calculated by the following formula:

[0128]

[0129] While using the low-rank gradient compression algorithm to compress the gradient once before communication and decompress it once after 2(N - 1) steps of communication, so that all computing nodes obtain the synchronized gradient. Then, in the scenario where the gradient needs to be compressed, the communication overhead based on the compressed gradient can be calculated by the following formula:

[0130]

[0131] Among them, the values of performance parameters such as T enc (m), T dec (m), T send (m), etc. can be obtained through pre-tests and / or transmission performance queries and analyses, etc. Since and their respective calculation formulas are both convex functions, therefore, the K values that minimize the two formulas can be determined. That is to say, in order to minimize the communication overhead, the K value in the communication overhead based on the uncompressed gradient and the K value in the communication overhead based on the compressed gradient can be different.

[0132] In this way, when constructing and After determining the K values in their respective corresponding formulas, the variable in these two formulas is the magnitude m of the gradient to be synchronized. Therefore, before training the neural network, and their respective corresponding formulas can be configured on the first computing node, and the first preset condition can be that if the communication overhead based on the gradient before compression is greater than the communication overhead based on the compressed gradient then it is determined that the gradient needs to be compressed; otherwise, it is determined that the gradient does not need to be compressed.

[0133] In this way, in this example, when performing the step of determining whether to compress the gradient according to the first preset condition, the magnitude of the gradient can be obtained as the value of m, and through and their respective corresponding formulas, calculate the difference between the communication overhead based on the compressed gradient and the communication overhead based on the gradient before compression, so as to determine whether to compress the gradient according to this difference.

[0134] Among them, if the communication overhead based on the compressed gradient is greater than the communication overhead based on the gradient before compression, then it is determined that this gradient needs to be compressed, and this gradient that needs to be compressed is used as the first gradient; if the communication overhead based on the compressed gradient is not greater than the communication overhead based on the gradient before compression, then it is determined that this gradient does not need to be compressed, and this gradient that does not need to be compressed is used as the second gradient.

[0135] It can be seen that in the embodiments of the present application, a method for selectively compressing gradients is proposed by considering aspects such as the magnitude of the gradient, compression and decompression, and data transmission overhead. Among them, the communication overheads corresponding to the compressed and uncompressed gradients can be theoretically analyzed, and the first preset condition can be pre-constructed according to performance parameters such as transmission and compression and decompression, so as to determine whether any gradient needs to be compressed according to the first preset condition, reduce or even avoid the negative impact on the actual processing efficiency caused by inappropriate compression operations such as small gradient compression, ensure the benefits brought by gradient compression, and thus more effectively improve the distributed training efficiency of the neural network.

[0136] In the actual neural network training process, the iteration process of the neural network usually has multiple times. Therefore, the same weight of the neural network needs to be updated multiple times. That is to say, gradients will be generated multiple times for this weight and gradient synchronization will be performed among multiple computing nodes. Since the data volume of the gradient corresponding to this weight and the corresponding network transmission method, etc. are often fixed, therefore, in each iteration process, the processing method of the generated gradient (that is, judging whether to compress this gradient) can follow the processing method determined for the corresponding gradient in the first iteration. That is to say, in some examples, step 303 can be executed only once during the training process of the neural network, and it may not be necessary to re-execute this step every time an iteration occurs. In other examples, step 303 can be executed multiple times during the training process of the neural network. For example, if the data volume of the gradient corresponding to the same weight changes in different iteration processes of the neural network, then this step can be executed separately in different iteration processes.

[0137] In addition, in some embodiments, the above step 303 can be executed by a third thread.

[0138] Specifically, in some embodiments, the above step 303 includes:

[0139] Using the third thread, determine whether to compress the gradient according to a first preset condition.

[0140] In the embodiments of the present application, this third thread is a thread different from the first thread and the second thread. For example, this third thread can be a thread in the main process, or the third thread can be a thread split from the main process, or the third thread can also be a thread in other processes other than the main process.

[0141] In one iteration process of the neural network, usually the multiple layers of the neural network are propagated backward layer by layer. Therefore, the gradients of the layers of the neural network are also generated layer by layer, rather than being generated all at once.

[0142] In the embodiments of the present application, it can be that after each new gradient is generated through calculation operations such as backward propagation, the third thread can obtain this gradient from the first thread and determine whether to compress the gradient according to the first preset condition. For example, in one iteration process, after the gradient of the nth layer of the neural network is obtained through the calculation operation of the first thread, the gradient of the nth layer of this neural network can be passed from the first thread to the third thread, and in this iteration process, the gradient of the (n - 1)th layer of the neural network usually has not been obtained through calculation operations yet.

[0143] In this way, through the first thread, the second thread, and the third thread, the computational operations of the neural network and subsequent processing operations such as the compression of the generated gradients can be managed and processed separately, providing a basis for the parallel processing of operations such as computing and compression, and enabling operations such as computing and compression to be hidden from each other.

[0144] Among them, through the third thread, subsequent operations such as compression and / or communication can be managed.

[0145] Next, in combination with multi-thread management, subsequent operations of the first gradient that needs to be compressed (such as step 304) and subsequent operations of the second gradient that does not need to be compressed (such as step 305) will be introduced exemplarily.

[0146] 1. Multi-thread management of the first gradient that needs to be compressed.

[0147] In some embodiments, after step 303 is executed, step 304 includes:

[0148] In the case of compressing the gradient, compress the first gradient to obtain the compressed first gradient.

[0149] In the embodiments of the present application, the first gradient may be the gradient that needs to be compressed among the gradients obtained in step 301.

[0150] When step 303 is executed by the third thread, the compression operation of the first gradient can be managed through the third thread.

[0151] (1) Compression operation of the first gradient.

[0152] Specifically, in some embodiments, after determining whether to compress the gradient according to the first preset condition through the third thread, it further includes:

[0153] In the case of compressing the gradient, store the first gradient in the compression task queue through the third thread;

[0154] Step 302 includes:

[0155] Obtain the first gradient from the compression task queue through the second thread for compression.

[0156] Refer to Figure 6 As shown in the example, in the embodiments of the present application, the first gradient can be passed to the compression task queue through the third thread, so that the second thread that manages the compression task queue triggers the computing unit to execute the compression operation on the first gradient to obtain the compressed first gradient.

[0157] In some embodiments, it can be through a task identifier (such as Figure 6The tasks to be executed are described by step id) in. For example, in Figure 6 In the example shown, the task identification list can be used to describe the task identification and the tasks to be executed corresponding to the task identification.

[0158] For example, in this task identification list, the compression task identification is step0, indicating that the task to be executed is the compression task; the communication task identification is step1, indicating that the task to be executed is the communication task; the decompression task identification is step2, indicating that the task to be executed is the decompression task.

[0159] In this way, the next task operation can be indicated by the task identification, and the dependency relationship between tasks can be maintained, facilitating the asynchronous management and task scheduling of different task queues and corresponding threads, and providing a basis for the parallel execution of tasks such as computing, compression, decompression, and communication.

[0160] For the compression task, when it is determined to compress the first gradient through the first preset condition, the first gradient compression task identification can be assigned, thereby triggering the third thread to store the first gradient in the compression task queue. In some examples, the first gradient stored by the third thread in the compression task queue can carry this compression task identification.

[0161] In Figure 6 the example, this compression task identification can be step0, and the compression task queue is Qcomp.

[0162] The second thread that manages the compression task queue Qcomp can determine the execution timing of the compression task for the first gradient based on the execution status of the compression tasks of other gradients before the first gradient and the idle resource status of the computing unit, etc., and at this execution timing, obtain the first gradient from the compression task queue, thereby instructing the computing unit to compress the first gradient to obtain the compressed first gradient. This computing unit can be executed in the second thread or in other threads outside the second thread.

[0163] (2) Communication operation for the compressed first gradient.

[0164] In some embodiments, after compressing the first gradient to obtain the compressed first gradient, it further includes:

[0165] Sending the compressed first gradient through the fourth thread.

[0166] In this way, through the first thread, the second thread, the third thread, and the fourth thread, the computing operation of the neural network and subsequent processing operations such as compression and communication of the generated gradients can be managed and processed separately, providing a basis for the parallel processing of operations such as computing, compression, and communication, and enabling operations such as computing, compression, and communication to be hidden from each other.

[0167] In some embodiments, the second thread may store the compressed first gradient in a communication task queue; then, the fourth thread may obtain the compressed first gradient from the communication task queue to send the compressed first gradient to a second computing node.

[0168] Exemplarily, the second thread may assign a communication task identifier (e.g., step1 in Figure 6 ) to the obtained compressed first gradient, so that the second thread stores the compressed first gradient in a communication task array Qcomm according to the communication task identifier. Then, the fourth thread managing the communication task queue Qcomm may determine the sending timing of the compressed first gradient according to communication conditions, etc., to read the compressed first gradient from the communication task queue Qcomm and instruct a communication unit to send the compressed first gradient to the second computing node among multiple computing nodes.

[0169] Among them, the communication unit may be considered as a software module for implementing communication. The communication unit may provide communication primitives through a underlying communication library and call communication interfaces such as an allreduce interface to implement the transmission of a second gradient. The communication unit may be located in the fourth thread or in other threads different from the fourth thread.

[0170] 2. Multi-thread management of a second gradient that does not need to be compressed.

[0171] In some embodiments, after determining whether to compress a gradient, the method further includes step 305:

[0172] In the case of not compressing the gradient, send the second gradient through the fourth thread.

[0173] In the embodiments of the present application, in the case of not compressing the gradient, the third thread may store the second gradient in a communication task queue. Then, the fourth thread managing the communication task queue Qcomm may send the second gradient to the second computing node.

[0174] For example, in the example shown in Figure 6 , in a first thread, forward propagation and backward propagation and other computing operations of a neural network may be performed. And, during the backward propagation, as the backward propagation is executed, gradients of the n-th layer of the neural network, gradients of the (n - 1)-th layer of the neural network, etc. are generated in sequence.

[0175] When the first thread performs computing operations such as backward propagation and generates a new gradient (e.g., the gradient of the n-th layer) each time, the gradient is passed to the third thread.

[0176] The third thread determines a second gradient that does not need to be compressed through a first preset condition, and then can store the second gradient in the communication task queue Qcomm to obtain a communication task regarding the second gradient.

[0177] Then, a fourth thread that manages the communication task queue Qcomm can determine the sending timing of the second gradient according to communication conditions, etc., so as to read the second gradient from the communication task queue Qcomm and instruct the communication unit to send the second gradient to a second computing node among multiple computing nodes.

[0178] Among them, the communication unit can be regarded as a software module for implementing communication. The communication unit can provide communication primitives through a underlying communication library and call communication interfaces such as the allreduce interface to implement the transmission of the second gradient. The communication unit can be located in the fourth thread or in other threads different from the fourth thread.

[0179] In addition, in some embodiments, the network transmission efficiency can also be improved by reasonably executing communication tasks.

[0180] For example, in some embodiments, the second gradient and the compressed first gradient are stored in the communication task queue;

[0181] Sending the compressed first gradient through the fourth thread includes:

[0182] When the second gradient and the compressed first gradient in the communication task queue meet a second preset condition, the second gradient and the compressed first gradient are packaged and sent through the fourth thread.

[0183] In the traditional solution, each single gradient calls a communication primitive once to implement network transmission every time gradient synchronization is performed, which will lead to a reduction in network bandwidth utilization and an increase in communication time consumption.

[0184] To solve the above problems, in the embodiments of the present application, it is considered to package the compressed gradient and the uncompressed gradient together and then perform network transmission to improve network bandwidth utilization.

[0185] Specifically, when it is detected that the second gradient and the compressed first gradient stored in the communication task queue meet the second preset condition, the second gradient and the compressed first gradient that meet the second preset condition are obtained from the communication task queue and packaged and sent.

[0186] Among them, there can be various situations for the second preset condition.

[0187] For example, the second preset condition can be that the time interval between the current time and the time when the previous communication task was executed is not less than a preset time threshold, and the time threshold is a parameter that can be preconfigured. For example, the time threshold can be configured to 5 ms, etc.

[0188] For another example, the second preset condition may refer to that the sum of the data volumes of the second gradient and the compressed first gradient stored in the communication task queue is not less than the data volume threshold. The data volume threshold is a parameter that can be pre-configured. For example, the data volume threshold may be 64 MB or the like.

[0189] When it is detected that the second gradient and the compressed first gradient stored in the communication task queue satisfy the second preset condition, the second gradient and the compressed first gradient can be packed to obtain the packed data, so as to send the packed data to the second computing node among multiple computing nodes. For example, when using the allreduce communication primitive for transmission, the packed data can achieve data synchronization among each computing node through one allreduce communication primitive.

[0190] For example, in the example shown in Figure 7 In the traditional solution, in the example shown, 3 compressed gradients and 2 uncompressed small gradients respectively call one allreduce communication primitive to achieve data synchronization among each computing node. That is to say, the 3 compressed gradients and the 2 uncompressed small gradients need to call the allreduce communication primitive 5 times.

[0191] In this example, when the combined data volume of the 3 compressed gradients and the 2 uncompressed small gradients satisfies the second preset condition (for example, less than 64 MB), the 3 compressed gradients and the 2 uncompressed small gradients can be packed and then call one allreduce communication primitive to achieve data synchronization among each computing node, greatly improving the network bandwidth utilization rate.

[0192] In addition, in some embodiments, the decompression task of the data can also be executed through multi-thread management.

[0193] The following is an exemplary introduction to the execution manner of the decompression task.

[0194] In some embodiments, the method further includes:

[0195] Receiving the gradient of the second computing node, where the second computing node is a computing node other than the first computing node, and the gradient of the second computing node includes the compressed gradient;

[0196] Decompressing the gradient of the second computing node through the fifth thread.

[0197] In the embodiments of the present application, in scenarios such as distributed training, the first computing node may receive the compressed gradient from the second computing node, so a decompression operation is required.

[0198] For example, if multiple computing nodes use the allreduce algorithm for gradient synchronization, the gradients of the neural network after synchronization can be data obtained by accumulation. And, in this example, the gradients of the neural network after synchronization are sent from the second computing node to the first computing node.

[0199] Among them, in the gradients of the neural network after synchronization, all gradients can be compressed, or some gradients can be compressed while some gradients are not compressed. Then, in this example, among the gradients of the neural network after synchronization, the compressed gradients can be used as the gradients of the second computing node to be decompressed.

[0200] In this way, the gradients of the second computing node can be decompressed through the fifth thread. This fifth thread is different from the above-mentioned first thread, second thread, third thread, and fourth thread. Therefore, this decompression operation can be executed in parallel with computing operations, compression operations, communication operations, etc., improving the data processing efficiency.

[0201] In some embodiments, the received gradients of the second computing node are stored in the decompression task queue;

[0202] Decompressing the gradients of the second computing node through the fifth thread includes:

[0203] Through the fifth thread, obtain the gradients of the second computing node from the decompression task queue for decompression.

[0204] In the embodiments of the present application, the fifth thread that manages the decompression task queue can determine the execution timing of the decompression task for the gradients of the second computing node according to the execution situation of the decompression task and the resource situation of the computing unit, etc., and instruct the computing unit to execute the decompression task for the gradients of the second computing node to obtain the decompressed gradients of the second computing node.

[0205] In addition, in some examples, after completing the communication primitive task, the first computing node can add a decompression task identifier (such as Figure 6 step2 shown) to the received gradients of the second computing node to indicate that the gradients of the second computing node are stored in the decompression task queue (such as Figure 6 Qdecomp in) so as to obtain the gradients of the second computing node from the compression task queue Qdecomp through the fifth thread and perform decompression through the computing unit.

[0206] It can be seen that the dependency relationship between tasks can be maintained through task identifiers such as communication task identifiers, compression task identifiers, and decompression task identifiers, facilitating asynchronous management and task scheduling of different task queues and corresponding threads, and providing a basis for the parallel execution of tasks such as computing, compression, decompression, and communication.

[0207] In the embodiments of the present application, it is possible to achieve the parallel execution of operations such as computing, compression, decompression, and communication in terms of timing through the asynchronous management of multiple task arrays and corresponding threads. For example, in Figure 5 the example shown, operations such as compression and decompression of the gradients of the nth layer are executed in parallel with operations such as computing of the gradients of the (n - 1)th layer and other layers; and again, operations such as communication of the gradients of the nth layer are executed in parallel with operations such as compression of the gradients of the (n - 1)th layer and other layers. It can be seen that through the embodiments of the present application, it is possible to hide operations such as computing, compression, decompression, and communication from each other in terms of timing, thereby improving the overall data processing efficiency.

[0208] In one example, the training speed of distributed neural network training using a traditional solution and the training speed of distributed neural network training using any of the above embodiments were tested.

[0209] Specifically, in the open-source communication framework HiPress, 16 servers (these 16 servers can be used as multiple computing nodes in any of the above embodiments) can be used for data-parallel neural network training. Each server has 8 GPUs, and the servers are interconnected using a 100 Gbps remote direct memory access (RDMA) network. In addition, the PyTorch computing framework is used as the backend of the communication framework, the neural network to be trained is the computer vision processing model UGATIT and the model LSTM selected from language models, and the gradient compression algorithm is the low-rank gradient compression algorithm PowerSGD.

[0210] In the traditional solution without gradient compression, the communication frameworks used are BytePS and Ring; while in the framework for gradient compression, the PowerSGD method integrated in the distributed training framework distributed data parallel (DDP) under the PyTorch framework is selected, denoted as PyTorch(OSS-PowerSGD); and the method using any of the above embodiments is integrated in the HiPress open-source framework, denoted as HiPress-CaSync-Ring(CompLL-PowerSGD); the HiPress open-source framework also integrates a quantized gradient compression algorithm TernGrad, denoted as HiPress-CaSync-PS(CompLL-TernGrad).

[0211] The aggregated training speeds corresponding to the above solutions are as Figure 8a and Figure 8b shown. Among them, Figure 8a and Figure 8bThe "linear scaling" in

[0212] Figure 8a In Figure 8a it is known that when training the neural network UGATIT using the data parallel method in the PyTorch framework, the abscissa represents the number of GPUs participating in the training, and the ordinate represents the number of images processed per second. This ordinate can reflect the end-to-end aggregated training speed. It can be seen that the higher the value of this ordinate, the higher the training speed and the higher the training performance. And from

[0213] Figure 8b In Figure 8b it is known that when training the neural network LSTM using the data parallel method in the PyTorch framework, the abscissa represents the number of GPUs participating in the training, and the ordinate represents the number of statements processed per second. This ordinate can reflect the end-to-end aggregated training speed. And from

[0214] it can be seen that when using the method of any of the above embodiments of the present application to train the neural network LSTM integrated in the HiPress open source framework, the training speed is the fastest.

[0215] As described above, the data processing method provided by the embodiments of the present application has been introduced from multiple aspects. Next, in conjunction with the accompanying drawings, the data processing device provided by the embodiments of the present application will be introduced.

[0216] As Figure 9 shown, the embodiments of the present application provide a data processing device 90. This data processing device 90 is applied to the first computing node, and this data processing device 90 includes:

[0217] A computing module 901, configured to execute a computing operation in the training process of the neural network of the first computing node through a first thread to obtain a gradient;

[0218] A processing module 902, configured to compress the gradient through a second thread to obtain a compressed gradient.

[0219] Optionally, the processing module 902 is configured to:

[0220] Determine whether to compress the gradient. The gradient that needs to be compressed is the first gradient, and the gradient that does not need to be compressed is the second gradient;

[0221] In the case of compressing the gradient, compress the first gradient to obtain the compressed first gradient.

[0222] Optionally, the processing module 902 is configured to: determine whether to compress the gradient according to a first preset condition, where the first preset condition is determined based on the communication overhead of the gradient before compression and the compressed gradient.

[0223] Optionally, the processing module 902 is configured to: through a third thread, determine whether to compress the gradient according to the first preset condition.

[0224] Optionally, the processing module 902 is configured to:

[0225] In the case of compressing the gradient, through a third thread, store the first gradient in the compression task queue;

[0226] Through a second thread, obtain the first gradient from the compression task queue for compression.

[0227] Optionally, the communication overhead is determined based on one or more of the following information:

[0228] The network transmission method of the gradient, the network transmission speed, the compression speed of the gradient, and the decompression speed of the compressed gradient.

[0229] Optionally, the processing module 902 is configured to: through a fourth thread, send the compressed first gradient.

[0230] Optionally, the second gradient and the compressed first gradient are stored in the communication task queue;

[0231] The processing module 902 is configured to: when the second gradient and the compressed first gradient in the communication task queue meet a second preset condition, through a fourth thread, package and send the second gradient and the compressed first gradient.

[0232] Optionally, the processing module 902 is configured to: in the case of not compressing the gradient, through a fourth thread, send the second gradient.

[0233] Optionally, the processing module 902 is configured to:

[0234] Receive the gradient of the second computing node, where the second computing node is a computing node other than the first computing node, and the gradient of the second computing node includes the compressed gradient;

[0235] Through a fifth thread, decompress the gradient of the second computing node.

[0236] Optionally, the received gradient of the second computing node is stored in the decompression task queue;

[0237] The processing module 902 is configured to: obtain the gradients of the second computing node from the decompression task queue through the fifth thread for decompression.

[0238] Figure 10 As shown in the figure, it is a possible schematic diagram of the logical structure of the computing node 100 provided by an embodiment of the present application. The computing node 100 is used to implement the functions of the electronic device involved in any of the above embodiments. The computing node 100 includes: a memory 1001, a processor 1002, a communication interface 1003, and a bus 1004. Among them, the memory 1001, the processor 1002, and the communication interface 1003 are communicatively connected to each other through the bus 1004.

[0239] The memory 1001 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1001 may store a program. When the program stored in the memory 1001 is executed by the processor 1002, the processor 1002 and the communication interface 1003 are used to execute one or more steps in the above data processing method embodiments.

[0240] The processor 1002 may be a central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any combination thereof, and is used to execute relevant programs to implement the functions required by the computing module and the processing module in the data processing device in the above embodiments, or execute one or more steps in the method embodiments of the present application. The steps of the method disclosed in combination with the embodiments of the present application may be completed by a hardware decoding processor, or may be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1001, and the processor 1002 reads the information in the memory 1001 and combines its hardware to execute one or more steps in the above data processing method embodiments.

[0241] The communication interface 1003 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the computing node 100 and other devices or communication networks.

[0242] The bus 1004 can implement a path for transmitting information between various components of the computing node 100 (for example, the memory 1001, the processor 1002, and the communication interface 1003). The bus 1004 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in representation, Figure 10 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0243] In another embodiment of the present application, a computer-readable storage medium is further provided. Computer-executable instructions are stored in the computer-readable storage medium. When the processor of the device executes the computer-executable instructions, the device executes the above-mentioned Figure 10 steps executed by the processor.

[0244] In another embodiment of the present application, a computer program product is further provided. The computer program product includes computer-executable instructions, and the computer-executable instructions are stored in a computer-readable storage medium; when the processor of the device executes the computer-executable instructions, the device executes the above-mentioned Figure 10 steps executed by the processor.

[0245] In another embodiment of the present application, a chip system is further provided. The chip system includes a processor, and the processor is used to implement the above-mentioned Figure 10 steps executed by the processor. In a possible design, the chip system may further include a memory for storing the necessary program instructions and data for the device to write data. The chip system can be composed of chips or can include chips and other discrete devices.

[0246] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0247] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0248] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0249] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0250] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs and other various media that can store program codes.

Claims

1. A data processing method, characterized in that, The method is applied to a first computing node, and the method includes: Executing a computing operation in the training process of the neural network of the first computing node through a first thread to obtain a gradient; Compressing the gradient through a second thread to obtain a compressed gradient.

2. The method according to claim 1, wherein Before compressing the gradient, it further includes: Judging whether to compress the gradient, the gradient that needs to be compressed is the first gradient, and the gradient that does not need to be compressed is the second gradient; The compressing of the gradient includes: In the case of compressing the gradient, compressing the first gradient to obtain a compressed first gradient.

3. The method according to claim 2, wherein The judging whether to compress the gradient includes: Judging whether to compress the gradient according to a first preset condition, and the first preset condition is determined based on the communication overhead of the gradient before compression and the gradient after compression.

4. The method according to claim 3, characterized in that, The judging whether to compress the gradient according to the first preset condition includes: Judging whether to compress the gradient according to the first preset condition through a third thread.

5. The method according to claim 4, wherein After judging whether to compress the gradient through a third thread, it further includes: In the case of compressing the gradient, storing the first gradient into a compression task queue through the third thread; The compressing of the gradient through the second thread to obtain a compressed gradient includes: Obtaining the first gradient from the compression task queue through the second thread for compression.

6. The method according to any one of claims 3-5, characterized in that, The communication overhead is determined based on one or more of the following information: The network transmission mode of the gradient, the network transmission speed, the compression speed of the gradient, and the decompression speed of the compressed gradient.

7. The method according to any one of claims 2-6, characterized in that, After compressing the first gradient to obtain a compressed first gradient, it further includes: Sending the compressed first gradient through a fourth thread.

8. The method according to claim 7, wherein The second gradient and the compressed first gradient are stored in a communication task queue; The sending of the compressed first gradient through the fourth thread includes: When the second gradient and the compressed first gradient in the communication task queue meet a second preset condition, packing and sending the second gradient and the compressed first gradient through the fourth thread.

9. The method according to any one of claims 2-8, characterized in that, After judging whether to compress the gradient, it further includes: In the case of not compressing the gradient, sending the second gradient through the fourth thread.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: Receiving the gradient of a second computing node, where the second computing node is a computing node other than the first computing node, and the gradient of the second computing node includes a compressed gradient; Decompressing the gradient of the second computing node through a fifth thread.

11. The method according to claim 10, wherein The received gradient of the second computing node is stored in a decompression task queue; The decompressing of the gradient of the second computing node through the fifth thread includes: Obtaining the gradient of the second computing node from the decompression task queue through the fifth thread for decompression.

12. A data processing device, characterized in that, Applied to a first computing node, the device includes: A computing module, configured to perform computing operations during the training process of the neural network of the first computing node through a first thread to obtain gradients; A processing module, configured to compress the gradients through a second thread to obtain compressed gradients.

13. The apparatus according to claim 12, wherein: The processing module is configured to: Determine whether to compress the gradients. The gradients that need to be compressed are the first gradients, and the gradients that do not need to be compressed are the second gradients; In the case of compressing the gradients, compress the first gradients to obtain compressed first gradients.

14. The apparatus according to claim 13, wherein: The processing module is configured to: Determine whether to compress the gradients according to a first preset condition, and the first preset condition is determined based on the communication overheads of the gradients before and after compression.

15. The apparatus according to claim 14, wherein: The processing module is configured to: Determine whether to compress the gradients according to the first preset condition through a third thread.

16. The apparatus according to claim 15, wherein: The processing module is configured to: In the case of compressing the gradients, store the first gradients in a compression task queue through the third thread; Obtain the first gradients from the compression task queue through the second thread for compression.

17. The device according to any one of claims 14 to 16, characterized in that The communication overhead is determined based on one or more of the following information: The network transmission mode of the gradients, the network transmission speed, the compression speed of the gradients, and the decompression speed of the compressed gradients.

18. The apparatus according to any one of claims 13-17, wherein: The processing module is configured to: Send the compressed first gradients through a fourth thread.

19. The device according to claim 18, characterized in that, The second gradients and the compressed first gradients are stored in a communication task queue; The processing module is configured to: When the second gradients and the compressed first gradients in the communication task queue meet a second preset condition, pack and send the second gradients and the compressed first gradients through the fourth thread.

20. The apparatus according to any one of claims 13-19, wherein: The processing module is configured to: In the case of not compressing the gradients, send the second gradients through a fourth thread.

21. The apparatus according to any one of claims 12-20, wherein: The processing module is configured to: Receive the gradients of a second computing node, where the second computing node is a computing node other than the first computing node, and the gradients of the second computing node include compressed gradients; Decompress the gradients of the second computing node through a fifth thread.

22. The device according to claim 21, characterized in that, The received gradients of the second computing node are stored in a decompression task queue; The processing module is configured to: Obtain the gradients of the second computing node from the decompression task queue through the fifth thread for decompression.

23. A computing node, characterized in that, The computing node includes at least one processor, a memory, and instructions stored on the memory and executable by the at least one processor, and the at least one processor executes the instructions to implement the steps of the method according to any one of claims 1-11.

24. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-11.

25. A computer program product comprising instructions, characterized in that, When the instructions are executed by a processor, it implements the method according to any one of claims 1-11.