Data processing method and related device
By performing backpropagation and gradient compression operations in parallel in distributed training of neural networks, combined with selective gradient compression and optimized communication, the gradient synchronization bottleneck problem is solved and the training efficiency of neural networks is improved.
Patent Information
- Application Number
- PCT/CN2024/134122
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2024-11-25
- Publication Date
- 2025-07-17
AI Technical Summary
In the distributed training of neural networks, gradient synchronization has become a bottleneck that restricts training speed, and the introduction of existing gradient compression algorithms takes time to calculate, resulting in less significant improvement in training efficiency.
By performing backpropagation and gradient compression operations of neural networks in parallel on different threads, using different hardware unit resources, parallel processing of gradient compression and backpropagation is realized, and compression is selectively performed according to the gradient size, optimizing communication methods to improve network bandwidth utilization.
It effectively reduces the overall data processing time, improves the distributed training efficiency of neural networks, and significantly improves the training speed.
Smart Images

Figure CN2024134122_17072025_PF_FP_ABST
Abstract
Description
A data processing method and related equipment
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 11, 2024, with Chinese application number 202410050300.4 and invention name “A data processing method and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence technology, and in particular to a data processing method and related equipment. Background Art
[0003] At present, as the number of parameters of neural networks becomes larger and larger, in order to meet the training needs of neural networks, neural networks can be trained in a distributed manner.
[0004] As the number of neural network parameters continues to increase, the amount of gradient data that needs to be synchronized across multiple distributed devices during neural network iterations is increasing. This has caused gradient synchronization to become a bottleneck restricting the training speed of distributed neural network training. To this end, a gradient compression algorithm was proposed to reduce the amount of gradient data during gradient synchronization, thereby alleviating network communication pressure and improving training speed.
[0005] However, when using the gradient compression algorithm to compress the gradient of the neural network to reduce network communication pressure, operations such as gradient compression and decompression will bring additional computational overhead and introduce additional computational time. Therefore, it is still possible that the duration of the overall training process cannot be effectively reduced, resulting in the benefits brought by the introduction of the gradient compression algorithm being not obvious, and the distributed training efficiency of the neural network not being significantly improved. Summary of the Invention
[0006] The present application provides a data processing method that can improve the distributed training efficiency of neural networks during actual distributed training. The present application also provides corresponding devices, equipment, computer-readable storage media, and computer program products.
[0007] A first aspect of the present application provides a data processing method, which is applied to a first computing node. The method includes: executing computing operations in a training process of a neural network of the first computing node through a first thread to obtain a gradient; and compressing the gradient through a second thread to obtain a compressed gradient.
[0008] In existing solutions, when using gradient compression algorithms to compress neural network gradients to reduce network communication pressure, researchers believe that both the backpropagation of the neural network and the compression and decompression of the gradients require computing resources to implement. If these processes are performed in parallel, computing resources will be preempted. Therefore, backpropagation and gradient compression and decompression are usually performed serially. For example, after backpropagating the neural network through the main thread to obtain the gradient, the computing resources occupied by the backpropagation are released; then the main thread obtains computing resources to compress the gradient of the neural network.
[0009] This will cause the traditional solution using the gradient compression algorithm to introduce additional computational time, resulting in the overall training process duration not being effectively reduced, making the benefits of introducing the gradient compression algorithm not obvious, and the distributed training efficiency of the neural network not significantly improved.
[0010] In the first aspect, the inventors believe that compression and other processing are memory-intensive operators, while neural network calculations such as back propagation are compute-intensive operators. Both rely on different hardware units in chips such as AI smart chips. Therefore, it is considered to use different threads (such as the first thread and the second thread) to implement different tasks.
[0011] For example, computing operations such as back propagation of a neural network can be implemented through the first thread. Moreover, in back propagation, after obtaining the gradient of the nth layer of the neural network, operations such as compression can be performed on the gradient of the nth layer through the second thread, so that the resources of the first thread can continuously perform back propagation of different layers such as the n-1th layer of the neural network.
[0012] It can be seen that through the second thread and the first thread, different types of tasks (such as compression and other subsequent processing of the nth layer and back propagation of the n-1th layer) of different layers of the neural network (such as the nth layer and the n-1th layer in the above example) can be hidden from each other (that is, they can be executed in parallel). Compared to the solution in which the calculation operations of the neural network and operations such as gradient compression are performed serially in traditional technologies, in the first aspect, through the first thread and the second thread, the compression and other processing of the gradients generated during the back propagation process can be hidden in time during the calculation process such as the back propagation of the neural network, thereby improving the distributed training efficiency of the neural network.
[0013] For example, in the back propagation process of the neural network, after the gradient of the neural network of the nth layer is obtained through back propagation, it is possible to consider performing operations such as compression on the obtained gradient of the neural network of the nth layer during the back propagation of the neural network of the n-1th layer. That is, the compression and other operations of the gradient of the neural network of the nth layer by the second thread are executed in parallel with the back propagation calculation operation of the neural network of the n-1th layer by the first thread, so that the compression and other operations of the gradient of the neural network of the nth layer are hidden in the back propagation and other calculation operations of the neural network of the n-1th layer, thereby reducing the overall data processing time and improving training efficiency.
[0014] In a possible implementation of the first aspect, before compressing the gradient, the method further includes: determining whether to compress the gradient, where the gradient to be compressed is a first gradient, and the gradient not to be compressed is a second gradient; and compressing the gradient includes: if compressing the gradient, compressing the first gradient to obtain a compressed first gradient.
[0015] In practical applications, the gradient of the neural network can be compressed using a gradient compression algorithm to reduce the amount of gradient data in the gradient synchronization process, thereby alleviating the pressure on network communication.
[0016] However, when using the gradient compression algorithm to compress the gradient of the neural network, operations such as gradient compression and decompression will bring additional computational overhead and introduce additional computational time. Therefore, it is still possible that the duration of the overall training process cannot be effectively reduced, resulting in the benefits brought by the introduction of the gradient compression algorithm being not obvious, and the distributed training efficiency of the neural network not being significantly improved.
[0017] Specifically, in some practical application scenarios, the size of the gradient of the neural network ranges from a few bytes to hundreds of MB. Compressing small gradients may not only fail to reduce network communication latency, but may increase the computational overhead of additional compression and decompression, thereby negatively affecting actual processing efficiency.
[0018] In this possible implementation, it is possible to determine whether to compress the gradient, so as to determine that the gradient that needs to be compressed is the first gradient, and the gradient that does not need to be compressed is the second gradient. Thus, through a selective gradient compression scheme, the negative impact of unreasonable gradient compression (such as compressing small gradients) on the actual processing efficiency is reduced or even avoided, and the benefits brought by gradient compression are guaranteed, thereby more effectively improving the distributed training efficiency of the neural network.
[0019] In a possible implementation of the first aspect, determining whether to compress the gradient includes determining whether to compress the gradient according to a first preset condition, where the first preset condition is determined based on communication overhead of the gradient before compression and the gradient after compression.
[0020] In this possible implementation, when calculating the communication overhead based on the pre-compression gradient and the post-compression gradient, it is not necessary to actually perform a compression operation and / or a transmission operation on the gradient. Instead, based on pre-collected relevant parameter information, the communication overhead based on the post-compression gradient and the communication overhead based on the pre-compression gradient can be estimated to determine whether to compress the gradient.
[0021] In a possible implementation of the first aspect, determining whether to compress the gradient according to the first preset condition includes: determining, by a third thread, whether to compress the gradient according to the first preset condition.
[0022] During one iteration of a neural network, backpropagation is usually performed on multiple layers of the neural network layer by layer. Therefore, the gradients of the layers of the neural network are also generated layer by layer, rather than all at the same time.
[0023] In this possible implementation, each time a new gradient is generated through a computational operation such as backpropagation, the third thread may obtain the gradient from the first thread and determine whether to compress the gradient based on a first preset condition. For example, during an iteration, after the gradient of the nth layer of the neural network is obtained through computational operations by the first thread, the gradient of the nth layer of the neural network may be transferred from the first thread to the third thread. However, during this iteration, the gradient of the n-1th layer of the neural network has not yet been obtained through computational operations.
[0024] In this way, through the first thread, the second thread and the third thread, the calculation operations of the neural network and subsequent processing operations such as compression of the generated gradients can be managed and processed separately, providing a basis for the parallel processing of operations such as calculation and compression, so that operations such as calculation and compression can be hidden from each other.
[0025] In a possible implementation of the first aspect, after determining, by a third thread, according to a first preset condition, whether to compress the gradient, the method further includes: if the gradient is to be compressed, storing, by the third thread, the first gradient in a compression task queue; and compressing, by the second thread, the gradient to obtain a compressed gradient, including: obtaining, by the second thread, the first gradient from the compression task queue for compression.
[0026] In this possible implementation, the first gradient may be transferred to the compression task queue through a third thread, so that the computing unit may be triggered to perform a compression operation on the first gradient through a second thread that manages the compression task queue, thereby obtaining a compressed first gradient.
[0027] In a possible implementation manner of the first aspect, the communication overhead is determined based on one or more of the following information: a network transmission mode of the gradient, a network transmission speed, a compression speed of the gradient, and a decompression speed of the compressed gradient.
[0028] This possible implementation proposes a method for selectively compressing gradients, taking into account factors such as gradient size, compression and decompression, and data transmission overhead. This method theoretically analyzes the communication overhead associated with compressed and uncompressed gradients, and pre-establishes a first pre-condition based on performance parameters such as transmission and compression and decompression. Based on this pre-condition, it determines whether any gradient requires compression. This reduces or even avoids the negative impact of inappropriate compression operations, such as small gradient compression, on actual processing efficiency, ensuring the benefits of gradient compression and thus more effectively improving the efficiency of distributed neural network training.
[0029] In a possible implementation manner of the first aspect, after compressing the first gradient to obtain the compressed first gradient, the method further includes: sending the compressed first gradient through a fourth thread.
[0030] In this possible implementation, the computational operations of the neural network and subsequent processing operations such as compression and communication of the generated gradients can be managed and processed separately through the first thread, the second thread, the third thread, and the fourth thread, providing a basis for parallel processing of operations such as computation, compression, and communication, so that the operations such as computation, compression, and communication can be hidden from each other.
[0031] In a possible implementation of the first aspect, the second gradient and the compressed first gradient are stored in a communication task queue; and sending the compressed first gradient through a fourth thread includes: when the second gradient and the compressed first gradient in the communication task queue meet a second preset condition, packaging and sending the second gradient and the compressed first gradient through the fourth thread.
[0032] In traditional solutions, each time gradient synchronization is performed, each single small gradient calls a communication primitive to achieve network transmission, which will reduce network bandwidth utilization and increase communication time.
[0033] To address the above issues, this possible implementation considers packaging the compressed gradient and the uncompressed gradient together before network transmission to improve network bandwidth utilization. Specifically, when it is detected that the second gradient stored in the communication task queue and the compressed first gradient meet the second preset condition, the second gradient and the compressed first gradient can be packaged to obtain packaged data, which can then be sent to the second computing node among the multiple computing nodes. For example, when the allreduce communication primitive is used for transmission, the packaged data can be synchronized between the computing nodes through a single allreduce communication primitive.
[0034] In a possible implementation manner of the first aspect, after determining whether to compress the gradient, the method further includes: sending the second gradient through a fourth thread without compressing the gradient.
[0035] In a possible implementation of the first aspect, the method further includes: receiving a gradient of a second computing node, where the second computing node is a computing node other than the first computing node, and the gradient of the second computing node includes a compressed gradient; and decompressing the gradient of the second computing node through a fifth thread.
[0036] In this possible implementation, the gradient of the second computing node can be decompressed by a fifth thread. This fifth thread is distinct from the first, second, third, and fourth threads. Therefore, the decompression operation can be performed in parallel with the computing, compression, and communication operations, thereby improving data processing efficiency.
[0037] In a possible implementation of the first aspect, the received gradient of the second computing node is stored in a decompression task queue; and decompressing the gradient of the second computing node by a fifth thread includes: obtaining, by the fifth thread, the gradient of the second computing node from the decompression task queue for decompression.
[0038] In this possible implementation, the fifth thread that manages the decompression task queue can determine the execution timing of the decompression task regarding the gradient of the second computing node based on the execution status of the decompression task and the resource status of the computing unit, and instruct the computing unit to execute the decompression task regarding the gradient of the second computing node to obtain the decompressed gradient of the second computing node.
[0039] A second aspect of the present application provides a data processing device that has the functionality to implement the method of the first aspect or any possible implementation of the first aspect. This functionality can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functionality, such as a computing module and a processing module.
[0040] The third aspect of the present application provides a computing node, which includes at least one processor, a memory, and computer execution instructions stored in the memory and executable by the processor. When the computer execution instructions are executed by the processor, the processor executes the method as described in the first aspect or any possible implementation of the first aspect.
[0041] The fourth aspect of the present application provides a computer-readable storage medium storing one or more computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor executes the method as described in the first aspect or any possible implementation of the first aspect.
[0042] The fifth aspect of the present application provides a computer program product that stores one or more computer-executable instructions. The computer program product includes computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor executes a method as described in the first aspect or any possible implementation of the first aspect.
[0043] A sixth aspect of the present application provides a chip system, which includes a processor for supporting a computing node to implement the functions involved in the first aspect or any possible implementation of the first aspect. In one possible design, the chip system may also include a memory for storing necessary program instructions and data. The chip system may be composed of a chip or may include a chip and other discrete devices.
[0044] Among them, the technical effects brought about by the second to sixth aspects or any possible implementation methods thereof can refer to the technical effects brought about by the first aspect or the relevant possible implementation methods of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] FIG1 is an exemplary schematic diagram of a distributed training architecture of a neural network provided in an embodiment of the present application;
[0046] FIG2a is an exemplary schematic diagram of a traditional data processing flow provided by an embodiment of the present application;
[0047] FIG2 b is an exemplary schematic diagram of an improved data processing flow provided by an embodiment of the present application;
[0048] FIG3 is a schematic diagram of an embodiment of a data processing method provided in an embodiment of the present application;
[0049] FIG4 is a schematic diagram of an embodiment of a data processing method provided in an embodiment of the present application;
[0050] FIG5 is an exemplary schematic diagram of a low-rank gradient compression algorithm provided in an embodiment of the present application;
[0051] FIG6 is an exemplary schematic diagram of multi-thread management provided by an embodiment of the present application;
[0052] FIG7 is an exemplary schematic diagram of data packet transmission according to an embodiment of the present application;
[0053] FIG8 a is an exemplary schematic diagram of a test result of training speed provided in an embodiment of the present application;
[0054] FIG8 b is an exemplary schematic diagram of another test result of the training speed provided in an embodiment of the present application;
[0055] FIG9 is a schematic diagram of an embodiment of a data processing device provided in an embodiment of the present application;
[0056] FIG10 is a schematic diagram of the structure of a computing node provided in an embodiment of the present application. DETAILED DESCRIPTION
[0057] The following describes the embodiments of the present application in conjunction with the accompanying drawings. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.
[0058] Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0059] In this application, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of singular or plural items. The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable where appropriate. This is merely a way of distinguishing objects with the same properties when describing them in the embodiments of this application. Furthermore, the terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover a non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a list of elements is not necessarily limited to those elements but may include other elements not expressly listed or inherent to such process, method, product, or apparatus.
[0060] In the field of artificial intelligence (AI), large AI models, due to their strong learning and comprehension capabilities, have demonstrated remarkable performance in areas such as natural language processing, conversational systems, and image processing. However, with the continuous advancement of technology, the number of parameters in large AI models is increasing exponentially, currently exceeding hundreds of billions and heading towards trillions. Simultaneously, training datasets are also growing larger. As a result, single machines are unable to meet the storage and computing power requirements of large AI models. One solution is to adopt a distributed approach to neural network training.
[0061] There are many ways to train neural networks in a distributed manner.
[0062] In one example, neural network training can be performed on distributed devices in a data-parallel dimension.
[0063] Data parallelism is a distributed approach to accelerate neural network training. In data parallel scenarios, the local data of different compute nodes can serve as different parts of the training dataset. Each compute node independently trains the neural network using its own local data. The gradients generated during training are synchronized across the network between nodes. For example, a common synchronization framework uses the allreduce primitive from the high-performance computing (HPC) field for communication.
[0064] The gradient of a neural network refers to the gradient of the neural network's weights, generated by backpropagation through the neural network. This gradient can be in the form of a tensor, reflecting the deviation of the corresponding weights. During the iterative process, the gradient obtained by backpropagation through the neural network can be used to update the neural network weights to obtain updated weights.
[0065] Figure 1 shows a schematic diagram of traditional data parallel training.
[0066] In the example shown in Figure 1, the neural network is trained through four computing nodes (Node0, Node1, Node2 and Node3). During the training process, each computing node is deployed with a complete neural network, and the training data in the training data set is divided into four parts and distributed to four computing nodes respectively. Each computing node can load a piece of training data locally and train the neural network separately. Any computing node can obtain the gradient of the neural network through forward propagation and back propagation based on the training data loaded locally. In this way, a total of four gradient information is generated by the four computing nodes. By transmitting information across nodes through the network, gradient synchronization between the four computing nodes can be achieved, so that each computing node can obtain the four gradient information, and each computing node can update the neural network according to the four gradient information.
[0067] As the number of neural network parameters continues to increase, the amount of gradient data that needs to be synchronized is growing, causing gradient synchronization to become a bottleneck that restricts the training speed of distributed neural network training. Furthermore, with the rapid development of computing power in recent years and the rise of compilation technology, computational operations such as forward propagation and backpropagation during training have become more efficient, meaning that gradient synchronization becomes more frequent, making the gradient synchronization bottleneck even more severe. It can even account for over 90% of the time spent in end-to-end training.
[0068] Based on this, an embodiment of the present application provides a data processing method that can improve training efficiency in an actual distributed training process.
[0069] The embodiments of the present application can be applied to a computing device cluster.
[0070] The computing device cluster may include multiple computing nodes, which are used to implement distributed neural network training. The multiple computing nodes involved in the embodiments of the present application can be considered as computing nodes that need to perform gradient synchronization during the distributed training of the neural network. The neural networks to be trained in different computing nodes can be the same or partially the same.
[0071] The type of neural network is not limited herein. For example, the neural network may be one or more of a recurrent neural network (RNN), a long short term memory network (LSTM), a convolutional neural network (CNN), a Transformer model, and the like.
[0072] The specific application scenarios of the neural network are not limited here. For example, the neural network can be one or more of a natural language processing model, an image processing model, a speech processing model, and a graph neural network.
[0073] The specific type of each computing node is not limited herein. For example, each computing node can be a terminal device, a single server or a server cluster, or a virtual machine (VM) or a container.
[0074] The specific type and functionality of each computing node can be the same or different, and are not limited here. Each computing node includes, but is not limited to, computing functionality. Furthermore, multiple computing nodes in a computing node cluster can be located in the same physical location, such as in the same data center; or they can be located in different physical locations, such as where any two computing nodes are located in different data centers.
[0075] Different computing nodes are connected via a network, and the specific type of the network and the communication method used are not limited here.
[0076] Exemplarily, different computing nodes can use allreduce communication methods, parameter server communication methods, etc. to synchronize data across nodes, for example, to achieve parameter synchronization of the neural network of each computing node.
[0077] The specific form of distributed training of neural networks implemented by multiple computing nodes is not limited here. For example, the distributed training can be performed based on the dimension of data parallelism. In addition, the distributed training of neural networks can also be performed using three-dimensional (3D) parallel technology, which includes data parallelism, tensor parallelism, and pipeline parallelism.
[0078] In addition, the specific form of the system software stack architecture in each computing node is not limited here.
[0079] In some examples, any computing node may include a bottom-level hardware acceleration chip and network link, a chip operator library in the layer above the bottom level, a computing framework in the layer above that, and a communication library in the layer above that. In this example, the data processing method implemented can involve the communication library layer, for example, using the allreduce communication method, without modifying the underlying logic of the software stack.
[0080] Some traditional theories suggest that when using gradient compression algorithms to compress neural network gradients to reduce network communication pressure, computations such as backpropagation and gradient compression and decompression require computing resources. If these operations are performed in parallel, computing resources will be preempted. Therefore, backpropagation and gradient compression and decompression are typically performed serially. For example, after backpropagating the neural network through the main thread to obtain the gradient, the computing resources occupied by backpropagation are released; the main thread then acquires computing resources to compress the gradient.
[0081] As shown in Figure 2a, during the training of a neural network, the gradient of the n+1th layer can be obtained through a computational operation T11 such as backpropagation. After computational operation T11 is completed, the computing resources are released, and the gradient of the n+1th layer can be compressed through a compression operation T12. After compression, the compressed gradient of the n+1th layer is sent to other computing nodes through a communication operation T13. Afterwards, referring to the relevant operations of the n+1th layer, computational operation T21, compression operation T22, and communication operation T23 are sequentially performed on the nth layer to send the compressed gradient of the nth layer to other computing nodes. Then, computational operation T31, compression operation T32, communication operation T33, and subsequent operations are sequentially performed on the n-1th layer.
[0082] This will cause the traditional solution using the gradient compression algorithm to introduce additional computational time, resulting in the overall training process duration not being effectively reduced, making the benefits of introducing the gradient compression algorithm not obvious, and the distributed training efficiency of the neural network not significantly improved.
[0083] The inventors of the present embodiments believe that computations such as backpropagation in neural networks are computationally intensive, requiring high utilization of computing units such as computing cores. However, compression and decompression are memory-intensive, requiring relatively few computational operations but often requiring repeated data scans. This places high demands on memory bandwidth but low demands on computing units.
[0084] Therefore, the two have different dependencies on hardware: one relies on the computing unit and the other relies on the memory unit. Therefore, resource conflicts usually do not occur and they can be executed concurrently.
[0085] For example, as shown in the example of FIG2b , in the back propagation process of the neural network, after the gradient of the neural network of the n+1th layer is obtained through the back propagation and other calculation operations T11', it is possible to consider performing operations such as compression operation T12' and communication operation T13' on the obtained gradient of the neural network of the n+1th layer in the process of performing the back propagation and other calculation operations T21' on the neural network of the nth layer. That is, the compression operation T12' and the communication operation T13' of the gradient of the neural network of the n+1th layer are performed in parallel with the back propagation and other calculation operations T21' of the neural network of the nth layer, so that the compression operation T12' and the communication operation T13' of the gradient of the neural network of the n+1th layer are hidden in the back propagation and other calculation operations T21' of the neural network of the nth layer. Similarly, the compression operation T22' and the communication operation T23' of the gradient of the neural network of the nth layer are performed in parallel with the calculation operation T31' such as the back propagation of the neural network of the n-1th layer, so that the obtained compression operation T22' and the communication operation T23' of the gradient of the neural network of the nth layer are hidden in the calculation operation T31' such as the back propagation of the neural network of the n-1th layer, thereby reducing the overall data processing time and improving training efficiency.
[0086] In the following, any computing node among the multiple computing nodes of the computing node cluster is referred to as a first computing node, and the data processing method in the embodiment of the present application is introduced by taking the first computing node as an example.
[0087] As shown in FIG3 , the data processing method may include steps 301 - 302 .
[0088] Step 301: Execute the computational operations in the training process of the neural network of the first computing node through the first thread to obtain a gradient.
[0089] In an embodiment of the present application, during the computation of the neural network of the first computing node, the computational operation of the neural network of the first computing node may be performed by the first thread. The computational operation may include back propagation of the neural network, and may also include forward propagation of the neural network.
[0090] The first thread may be one or more threads in a first process of the first computing node. For example, the first thread may be a main process in the first process.
[0091] In step 301, the gradient obtained may be a partial gradient of the neural network, for example, one or more gradients generated when performing back propagation calculations on a layer of the neural network. The number of gradients in step 301 is not limited here.
[0092] Step 302: Compress the gradient through the second thread to obtain a compressed gradient.
[0093] In the embodiment of the present application, the second thread is a thread different from the first thread. For example, the second thread can be a thread in the first process, or the second thread can be a thread separated from the first process, or the second thread can be a thread in a process other than the first process.
[0094] In the embodiment of the present application, part or all of the content in the gradient may be compressed. For example, if the gradient includes a first gradient that needs to be compressed and a second gradient that does not need to be compressed, the first gradient in the gradient may be compressed.
[0095] It can be seen that in the embodiments of the present application, considering that compression and other processing are memory-intensive operators, and neural network calculations such as back propagation are compute-intensive operators, both rely on different hardware units in chips such as AI smart chips. Therefore, in the embodiments of the present application, it is considered to implement different tasks through different threads (for example, the first thread and the second thread).
[0096] For example, computing operations such as back propagation of a neural network can be implemented through the first thread. Moreover, in back propagation, after obtaining the gradient of the nth layer of the neural network, operations such as compression can be performed on the gradient of the nth layer through the second thread, so that the resources of the first thread can continuously perform back propagation of different layers such as the n-1th layer of the neural network.
[0097] It can be seen that, through the second thread and the first thread, different types of tasks (such as compression and other subsequent processing of the nth layer and back propagation of the n-1th layer) of different layers of the neural network (such as the nth layer and the n-1th layer in the above example) can be hidden from each other (that is, they can be executed in parallel). Compared with the solution in which the calculation operation of the neural network and the compression of the gradient and other operations are performed serially in the traditional technology, in the solution of the embodiment of the present application, through the first thread and the second thread, the compression and other processing of the gradient generated during the back propagation process can be hidden in time in the calculation process such as the back propagation of the neural network, thereby improving the distributed training efficiency of the neural network.
[0098] For example, in the back propagation process of the neural network, after the gradient of the neural network of the nth layer is obtained through back propagation, it is possible to consider performing operations such as compression on the obtained gradient of the neural network of the nth layer during the back propagation of the neural network of the n-1th layer. That is, the compression and other operations of the gradient of the neural network of the nth layer by the second thread are executed in parallel with the back propagation calculation operation of the neural network of the n-1th layer by the first thread, so that the compression and other operations of the gradient of the neural network of the nth layer are hidden in the back propagation and other calculation operations of the neural network of the n-1th layer, thereby reducing the overall data processing time and improving training efficiency.
[0099] In some embodiments, after the computing operations in the training process of the neural network of the first computing node are executed by the first thread and the gradient is obtained, it can be determined whether to compress the gradient.
[0100] Specifically, as shown in FIG4 , in some embodiments, after step 301 , step 303 may be executed, and the subsequent execution of step 304 or 305 may be determined based on the execution result of step 303 .
[0101] Wherein, step 303 includes:
[0102] Determine whether to compress the gradient.
[0103] The gradient that needs to be compressed is the first gradient, and the gradient that does not need to be compressed is the second gradient.
[0104] In practical applications, the gradient compression algorithm can be used to compress the gradient of the neural network to reduce the amount of gradient data in the gradient synchronization process, thereby alleviating the pressure on network communication and improving the training speed.
[0105] However, when using the gradient compression algorithm to compress the gradient of the neural network, operations such as gradient compression and decompression will bring additional computational overhead and introduce additional computational time. Therefore, it is still possible that the duration of the overall training process cannot be effectively reduced, resulting in the benefits brought by the introduction of the gradient compression algorithm being not obvious, and the distributed training efficiency of the neural network not being significantly improved.
[0106] Specifically, in some practical application scenarios, the size of the gradient of the neural network ranges from a few bytes to hundreds of MB. Compressing small gradients may not only fail to reduce network communication latency, but may increase the computational overhead of additional compression and decompression, thereby negatively affecting actual processing efficiency.
[0107] Based on this, in an embodiment of the present application, it can be determined whether to compress the gradient to determine that the gradient that needs to be compressed is the first gradient, and the gradient that does not need to be compressed is the second gradient, thereby reducing the impact of inappropriate compression and decompression operations on processing efficiency.
[0108] Specifically, in an embodiment of the present application, information from multiple dimensions, such as transmission overhead and additional overhead caused by compression and decompression operations, can be comprehensively considered to determine whether to compress the gradient. Thus, through a selective gradient compression scheme, the negative impact of unreasonable gradient compression (such as compressing small gradients) on actual processing efficiency can be reduced or even avoided, thereby ensuring the benefits brought by gradient compression, thereby more effectively improving the distributed training efficiency of neural networks.
[0109] In some embodiments, step 303 includes:
[0110] Whether to compress the gradient is determined according to a first preset condition.
[0111] The first preset condition is determined based on the communication overhead of the gradient before compression and the gradient after compression.
[0112] When there are multiple gradients, it can be determined whether to compress any gradient according to a first preset condition.
[0113] When calculating the communication overhead based on the gradient before compression and the gradient after compression, it is not necessary to actually perform a compression operation and / or a transmission operation on the gradient. Instead, the communication overhead based on the gradient after compression and the communication overhead based on the gradient before compression can be estimated based on pre-collected relevant parameter information to determine whether to compress the gradient.
[0114] When calculating the above communication overhead, there may be multiple types of related parameter information involved.
[0115] In some embodiments, the communication overhead is determined based on one or more of the following information:
[0116] The network transmission method of the gradient, the network transmission speed, the compression speed of the gradient, and the decompression speed of the compressed gradient.
[0117] The above parameter information may be collected in advance through testing or other methods.
[0118] The communication overhead based on compressed gradients can include the time it takes to compress the gradients, the time it takes to transmit the compressed gradients between multiple computing nodes so that all computing nodes obtain the compressed gradients, and the time it takes to decompress the compressed gradients. The communication overhead based on uncompressed gradients can include the time it takes to transmit the gradients between multiple computing nodes so that all computing nodes obtain the gradients.
[0119] The following uses the data synchronization method between multiple computing nodes using the allreduce algorithm and the gradient compression algorithm using a low-rank gradient compression algorithm as an example to exemplify the process of determining whether to compress the gradient according to the first preset condition.
[0120] Currently, a commonly used gradient compression algorithm is the low-rank compression algorithm. During the training process, the gradient matrix generated by the neural network is relatively sparse and has a relatively small rank. Using the concept of linear algebra, a low-rank gradient matrix can be decomposed into two matrices through QR. For example, as shown in Figure 5, an M*N matrix can be decomposed into L k =m*k and R K T = two matrices of k*n, where k is the rank of the matrix, in order to compress the gradient and reduce the amount of communication. At this time, the compression ratio is ((m+n)*k) / (m*k).
[0121] When data synchronization between multiple compute nodes uses the allreduce algorithm, gradient synchronization can be achieved between the compute nodes using a set of allreduce communication primitives. A communication primitive is a set of communication templates formed by combining one or more basic operations, such as send, receive, and copy.
[0122] Due to the characteristics of the low-rank gradient compression algorithm, when data synchronization between multiple compute nodes is performed using the allreduce algorithm, the compressed gradients can be directly accumulated. Therefore, the gradients of each compute node (which can include both uncompressed and compressed gradients) can be accumulated using the same set of allreduce communication primitives, achieving gradient synchronization between multiple compute nodes. After obtaining the accumulated data, each compute node performs a decompression operation to obtain the gradient to be synchronized, eliminating the need for multiple compression and decompression operations at each compute node during the communication process.
[0123] It can be seen that if the data synchronization method between multiple computing nodes is the synchronization method using the allreduce algorithm and the gradient compression algorithm is a low-rank gradient compression algorithm, it is usually only necessary to compress the gradient before communication and decompress the accumulated data after communication. During the communication process, a large number of compression and decompression operations can be reduced, greatly improving the data synchronization efficiency between multiple computing nodes.
[0124] In this example, if the data synchronization method between computing nodes is a synchronization method using the allreduce algorithm and the gradient compression algorithm is a low-rank compression algorithm, in an exemplary first preset condition, the calculation formula for the communication overhead based on the compressed gradient and the communication overhead based on the gradient before compression is generated based on the following method.
[0125] Before generating the first preset condition, illustratively, the parameters in Table 1 may be defined.
[0126] Table 1: Parameters related to the first preset condition
[0127] The allreduce communication primitive is used to synchronize the gradients of size m between N computing nodes. The gradient needs to be split into K parts for network transmission (where K is not greater than N). After 2(N-1) steps of communication, all computing nodes get the synchronized gradients. Then, the communication overhead based on the gradient before compression is It can be calculated by the following formula:
[0128] The low-rank gradient compression algorithm compresses the gradient before communication and decompresses it after 2(N-1) steps of communication, so that all computing nodes can obtain synchronized gradients. It can be calculated by the following formula:
[0129] Among them, such as T enc (m), T dec (m), T send The values of performance parameters such as (m) can be obtained through pre-testing and / or transmission performance query and analysis. as well as The corresponding calculation formulas are all convex functions, so we can determine the K value that minimizes the two formulas. In other words, in order to minimize the communication overhead, the communication overhead based on the gradient before compression is The K value in and the communication overhead based on the compressed gradient The value of K in can be different.
[0130] So, when building as well as After the corresponding formulas are determined and the K value is determined, the variables in the two formulas are the size m of the gradient to be synchronized. Therefore, before training the neural network, as well as The corresponding formulas are configured in the first computing node, and the first preset condition can be based on the communication overhead of the gradient before compression. Greater than the communication overhead of compressed gradients If , it is determined that the gradient needs to be compressed, otherwise, it is determined that the gradient does not need to be compressed.
[0131] Thus, in this example, when executing the step of determining whether to compress the gradient according to the first preset condition, the size of the gradient can be obtained as the value of m, and as well as The corresponding formulas calculate the communication overhead based on the compressed gradient Communication overhead compared to pre-compression gradients The difference between them is used to determine whether to compress the gradient.
[0132] Among them, if the communication overhead based on the compressed gradient Greater than the communication overhead based on the gradient before compression It is determined that the gradient needs to be compressed, and the gradient that needs to be compressed is used as the first gradient; and if the communication overhead based on the compressed gradient No more than the communication overhead of the gradient before compression It is determined that the gradient does not need to be compressed, and the gradient that does not need to be compressed is used as the second gradient.
[0133] As can be seen, in the embodiments of the present application, a method for selectively compressing gradients is proposed, taking into account aspects such as gradient size, compression and decompression, and data transmission overhead. Specifically, a theoretical analysis can be conducted on the communication overhead corresponding to compressed and uncompressed gradients, and a first preset condition can be pre-established based on performance parameters such as transmission and compression and decompression. Based on the first preset condition, it can be determined whether any gradient needs to be compressed, thereby reducing or even avoiding the negative impact of inappropriate compression operations such as small gradient compression on actual processing efficiency, ensuring the benefits of gradient compression, and thus more effectively improving the distributed training efficiency of neural networks.
[0134] In the actual neural network training process, the neural network usually has multiple iterations. Therefore, the same weight of the neural network needs to be updated multiple times. In other words, the gradient will be generated multiple times for the weight and the gradient synchronization will be performed between multiple computing nodes. Since the amount of data of the gradient corresponding to the weight and the corresponding network transmission method are often fixed, the processing method of the generated gradient in each iteration (that is, whether to compress the gradient) can continue to use the processing method determined for the corresponding gradient in the first iteration. That is to say, in some examples, step 303 can be performed only once during the training process of the neural network, and it is not necessary to re-execute this step in each iteration. In other examples, step 303 can be performed multiple times during the training process of the neural network; for example, if the amount of data of the gradient corresponding to the same weight changes in different iterations of the neural network, this step can be performed separately in different iterations.
[0135] Furthermore, in some embodiments, step 303 may be performed via a third thread.
[0136] Specifically, in some embodiments, step 303 includes:
[0137] Through the third thread, it is determined whether to compress the gradient according to the first preset condition.
[0138] In the embodiment of the present application, the third thread is a thread different from the first thread and the second thread. For example, the third thread can be a thread in the main process, or the third thread can be a thread separated from the main process, or the third thread can be a thread in a process other than the main process.
[0139] During one iteration of a neural network, backpropagation is usually performed on multiple layers of the neural network layer by layer. Therefore, the gradients of the layers of the neural network are also generated layer by layer, rather than all at the same time.
[0140] In an embodiment of the present application, each time a new gradient is generated through a computational operation such as backpropagation, the third thread may obtain the gradient from the first thread and determine whether to compress the gradient based on a first preset condition. For example, during an iteration, after the gradient of the nth layer of the neural network is obtained through a computational operation by the first thread, the gradient of the nth layer of the neural network may be transferred from the first thread to the third thread. However, during this iteration, the gradient of the n-1th layer of the neural network has not yet been obtained through a computational operation.
[0141] In this way, through the first thread, the second thread and the third thread, the calculation operations of the neural network and subsequent processing operations such as compression of the generated gradients can be managed and processed separately, providing a basis for the parallel processing of operations such as calculation and compression, so that operations such as calculation and compression can be hidden from each other.
[0142] The third thread can be used to manage subsequent compression and / or communication operations.
[0143] In the following, in combination with multi-thread management, the subsequent operation of the first gradient that needs to be compressed (eg, step 304) and the subsequent operation of the second gradient that does not need to be compressed (eg, step 305) are respectively described by way of example.
[0144] 1. Multi-threaded management of the first gradient that needs to be compressed.
[0145] In some embodiments, after executing step 303, step 304 includes:
[0146] When the gradient is compressed, the first gradient is compressed to obtain a compressed first gradient.
[0147] In the embodiment of the present application, the first gradient may be the gradient that needs to be compressed among the gradients obtained in step 301 .
[0148] When step 303 is executed by the third thread, the compression operation of the first gradient can be managed by the third thread.
[0149] (1) Compression operation on the first gradient.
[0150] Specifically, in some embodiments, after determining, by the third thread, whether to compress the gradient according to the first preset condition, the method further includes:
[0151] When the gradient is compressed, the first gradient is stored in the compression task queue through the third thread;
[0152] Step 302 includes:
[0153] The second thread obtains the first gradient from the compression task queue for compression.
[0154] Referring to the example shown in FIG6 , in an embodiment of the present application, the first gradient can be passed to the compression task queue through a third thread, so that the computing unit can be triggered to perform a compression operation on the first gradient through a second thread that manages the compression task queue to obtain a compressed first gradient.
[0155] In some embodiments, the task to be performed may be described by a task identifier (eg, step id in FIG6 ). For example, in the example shown in FIG6 , a task identifier list may be used to describe the task identifier and the task to be performed corresponding to the task identifier.
[0156] For example, in the task identification list, the compression task identification is step0, indicating that the task to be executed is a compression task; the communication task identification is step1, indicating that the task to be executed is a communication task; and the decompression task identification is step2, indicating that the task to be executed is a decompression task.
[0157] In this way, the next task operation can be indicated by the task identifier, and the dependency relationship between tasks can be maintained, which facilitates asynchronous management and task scheduling of different task queues and corresponding threads, and provides a basis for the parallel execution of tasks such as calculation, compression, decompression, and communication.
[0158] For the compression task, if the first gradient is determined to be compressed according to the first preset condition, a compression task identifier can be assigned to the first gradient, thereby triggering the third thread to store the first gradient in the compression task queue. In some examples, the first gradient stored by the third thread in the compression task queue can carry the compression task identifier.
[0159] In the example of FIG6 , the compression task identifier may be step0 , and the compression task queue may be Qcomp .
[0160] The second thread managing the compression task queue Qcomp can determine the execution timing of the compression task for the first gradient based on the execution status of compression tasks for other gradients preceding the first gradient and the idle resources of the computing unit. At this execution timing, the second thread retrieves the first gradient from the compression task queue, thereby instructing the computing unit to compress the first gradient and obtain the compressed first gradient. This computing unit can be executed in the second thread or in a thread other than the second thread.
[0161] (2) Communication operation on the compressed first gradient.
[0162] In some embodiments, after compressing the first gradient to obtain the compressed first gradient, the method further includes:
[0163] The compressed first gradient is sent through the fourth thread.
[0164] In this way, the computational operations of the neural network and subsequent processing operations such as compression and communication of the generated gradients can be managed and processed separately through the first thread, the second thread, the third thread, and the fourth thread, providing a basis for the parallel processing of computation, compression, communication and other operations, so that computation, compression, communication and other operations can be hidden from each other.
[0165] In some embodiments, the second thread may store the compressed first gradient in a communication task queue; then, the fourth thread may obtain the compressed first gradient from the communication task queue to send the compressed first gradient to the second computing node.
[0166] For example, the second thread can assign a communication task identifier to the obtained compressed first gradient (e.g., step 1 in FIG6 ), so that the second thread can store the compressed first gradient in the communication task array Qcomm based on the communication task identifier. Then, a fourth thread managing the communication task queue Qcomm can determine the timing for sending the compressed first gradient based on communication conditions, read the compressed first gradient from the communication task queue Qcomm, and instruct the communication unit to send the compressed first gradient to a second computing node among the multiple computing nodes.
[0167] The communication unit can be considered a software module that implements communication. It can provide communication primitives through the underlying communication library and call communication interfaces such as the allreduce interface to implement the transmission of the second gradient. The communication unit can be located in the fourth thread or in another thread different from the fourth thread.
[0168] 2. Multi-threaded management of the second gradient that does not require compression.
[0169] In some embodiments, after determining whether to compress the gradient, the method further includes step 305:
[0170] Without compressing the gradient, the second gradient is sent through the fourth thread.
[0171] In the embodiment of the present application, when the gradient is not compressed, the third thread can store the second gradient in the communication task queue. Then, the second gradient can be sent to the second computing node through the fourth thread that manages the communication task queue Qcomm.
[0172] For example, as shown in the example of FIG6 , the first thread can execute computational operations such as forward propagation and back propagation of the neural network, and in the process of back propagation, as the back propagation is executed, the gradient of the neural network of the nth layer, the gradient of the neural network of the n-1th layer, and so on are generated in sequence.
[0173] When the first thread performs a computational operation such as back propagation, each time a new gradient (such as the gradient of the nth layer) is generated, the gradient is passed to the third thread.
[0174] The third thread determines, through the first preset condition, that the second gradient does not need to be compressed, and then stores the second gradient in the communication task queue Qcomm to obtain a communication task related to the second gradient.
[0175] Then, the fourth thread managing the communication task queue Qcomm can determine the sending timing of the second gradient according to the communication situation, etc., to read the second gradient from the communication task queue Qcomm, and instruct the communication unit to send the second gradient to the second computing node among the multiple computing nodes.
[0176] The communication unit can be considered a software module that implements communication. It can provide communication primitives through the underlying communication library and call communication interfaces such as the allreduce interface to implement the transmission of the second gradient. The communication unit can be located in the fourth thread or in another thread different from the fourth thread.
[0177] In addition, in some embodiments, network transmission efficiency can be improved by reasonably executing communication tasks.
[0178] For example, in some embodiments, the second gradient and the compressed first gradient are stored in a communication task queue;
[0179] Sending the compressed first gradient through the fourth thread includes:
[0180] When the second gradient and the compressed first gradient in the communication task queue meet a second preset condition, the second gradient and the compressed first gradient are packaged and sent through a fourth thread.
[0181] In traditional solutions, each time gradient synchronization is performed, each single gradient calls a communication primitive to implement network transmission, which will reduce network bandwidth utilization and increase communication time.
[0182] In order to solve the above problem, in an embodiment of the present application, it is considered to package the compressed gradient and the uncompressed gradient together and then transmit them over the network to improve the network bandwidth utilization.
[0183] Specifically, when it is detected that the second gradient and the compressed first gradient stored in the communication task queue meet the second preset condition, the second gradient and the compressed first gradient are obtained from the communication task queue and packaged for transmission.
[0184] There are multiple situations for the second preset condition.
[0185] For example, the second preset condition may be that the time interval between the current time and the execution of the last communication task is not less than a preset time threshold. The time threshold is a preconfigurable parameter. For example, the time threshold may be configured to be 5ms.
[0186] For another example, the second preset condition may be that the sum of the data volume of the second gradient and the compressed first gradient stored in the communication task queue is not less than a data volume threshold. The data volume threshold is a preconfigurable parameter, for example, the data volume threshold may be 64MB.
[0187] When it is detected that the second gradient and the compressed first gradient stored in the communication task queue meet a second preset condition, the second gradient and the compressed first gradient can be packaged to obtain packaged data, and the packaged data can be sent to a second computing node among the multiple computing nodes. For example, when the allreduce communication primitive is used for transmission, the packaged data can be synchronized between the computing nodes through a single allreduce communication primitive.
[0188] For example, in the example shown in Figure 7, in the traditional solution, the three compressed gradients and the two uncompressed small gradients each require one call to the allreduce communication primitive to synchronize data between the compute nodes. This means that the three compressed gradients and the two uncompressed small gradients require five calls to the allreduce communication primitive.
[0189] In this example, the data size of the three compressed gradients and two uncompressed small gradients meets the second preset condition (for example, less than 64 MB). After the three compressed gradients and the two uncompressed small gradients are packaged, the allreduce communication primitive can be called once to synchronize data between the computing nodes, greatly improving network bandwidth utilization.
[0190] Furthermore, in some embodiments, the data decompression task may be performed through multi-thread management.
[0191] The following is an example of how to execute the decompression task.
[0192] In some embodiments, the method further comprises:
[0193] receiving a gradient of a second computing node, where the second computing node is a computing node other than the first computing node, and the gradient of the second computing node includes a compressed gradient;
[0194] The fifth thread decompresses the gradient of the second computing node.
[0195] In an embodiment of the present application, in scenarios such as distributed training, the first computing node may receive compressed gradients from the second computing node, and therefore a decompression operation is required.
[0196] For example, if multiple computing nodes use the allreduce algorithm to synchronize gradients, the gradient of the synchronized neural network can be accumulated data. In this example, the gradient of the synchronized neural network is sent from the second computing node to the first computing node.
[0197] Among the gradients of the synchronized neural network, all gradients may be compressed, or some gradients may be compressed while others are not. In this example, among the gradients of the synchronized neural network, the compressed gradients may serve as the gradients of the second computation node to be decompressed.
[0198] In this way, the gradient of the second computing node can be decompressed by the fifth thread. The fifth thread is different from the first, second, third, and fourth threads. Therefore, the decompression operation can be performed in parallel with the computing operation, compression operation, communication operation, etc., thereby improving data processing efficiency.
[0199] In some embodiments, the received gradient of the second computing node is stored in a decompression task queue;
[0200] The fifth thread decompresses the gradient of the second computing node, including:
[0201] The fifth thread obtains the gradient of the second computing node from the decompression task queue for decompression.
[0202] In an embodiment of the present application, the fifth thread that manages the decompression task queue can determine the execution timing of the decompression task regarding the gradient of the second computing node based on the execution status of the decompression task and the resource status of the computing unit, and instruct the computing unit to execute the decompression task regarding the gradient of the second computing node to obtain the decompressed gradient of the second computing node.
[0203] In addition, in some examples, after completing the communication primitive task, the first computing node may add a decompression task identifier to the received gradient of the second computing node (e.g., step 2 shown in FIG6 ) to indicate that the gradient of the second computing node is stored in a decompression task queue (e.g., Qdecomp in FIG6 ), so that the gradient of the second computing node can be obtained from the compression task queue Qdecomp through a fifth thread and decompressed through the computing unit.
[0204] It can be seen that the dependencies between tasks can be maintained through task identifiers such as communication task identifiers, compression task identifiers, and decompression task identifiers, which facilitates asynchronous management and task scheduling of different task queues and corresponding threads, and provides a basis for the parallel execution of tasks such as computing, compression, decompression, and communication.
[0205] In the embodiment of the present application, the possibility of parallel execution of operations such as calculation, compression, decompression, and communication in time sequence can be achieved through asynchronous management of multiple task arrays and corresponding threads. For example, in the example shown in Figure 5, operations such as compression and decompression of the gradient of the nth layer are executed in parallel with operations such as calculation of the gradient of the n-1th layer and other layers; for another example, operations such as communication of the gradient of the nth layer are executed in parallel with operations such as compression of the gradient of the n-1th layer and other layers. It can be seen that through the embodiment of the present application, operations such as calculation, compression, decompression, and communication can be hidden from each other in time sequence, thereby improving the overall data processing efficiency.
[0206] In one example, the training speed of a distributed neural network training using a traditional solution and the training speed of a distributed neural network training using any of the above embodiments are tested.
[0207] Specifically, data-parallel neural network training can be performed using 16 servers (which can serve as multiple computing nodes in any of the above embodiments) within the open-source communication framework HiPress. Each server has eight GPUs, and the servers are interconnected using a 100Gbps remote direct memory access (RDMA) network. Furthermore, the PyTorch computing framework serves as the backend for the communication framework. The neural networks to be trained are the computer vision processing model UGATIT and the language selection model LSTM, and the gradient compression algorithm is the low-rank gradient compression algorithm PowerSGD.
[0208] In traditional solutions that do not compress gradients, the communication frameworks used are BytePS and Ring. In the framework that compresses gradients, the PowerSGD method integrated with the distributed data parallel (DDP) distributed training framework under the PyTorch framework is selected, denoted as PyTorch (OSS-PowerSGD). The method using any of the above embodiments is integrated into the HiPress open source framework, denoted as HiPress-CaSync-Ring (CompLL-PowerSGD). The HiPress open source framework also integrates a quantized gradient compression algorithm TernGrad, denoted as HiPress-CaSync-PS (CompLL-TernGrad).
[0209] The corresponding aggregation training speed of the above scheme is shown in Figures 8a and 8b. Among them, linear sealing in Figures 8a and 8b represents linear scalability, which is an ideal result. The closer the measured results are to linear-scaling, the better the performance.
[0210] In Figure 8a, the neural network UGATIT is trained using a data parallel approach in the PyTorch framework. The horizontal axis represents the number of GPUs involved in training, and the vertical axis represents the number of images processed per second. This vertical axis can reflect the end-to-end aggregate training speed. As can be seen, the higher the vertical axis value, the higher the training speed and the higher the training performance. As can be seen from Figure 8a, the neural network UGATIT trained using the method of any of the above embodiments of the present application, integrated into the HiPress open source framework, has the fastest training speed.
[0211] In Figure 8b, the LSTM neural network is trained using data parallelism in the PyTorch framework. The horizontal axis represents the number of GPUs involved in training, and the vertical axis represents the number of statements processed per second. This vertical axis can reflect the end-to-end aggregate training speed. As can be seen from Figure 8b, the LSTM neural network trained using the method of any of the above embodiments of this application, integrated into the HiPress open source framework, achieves the fastest training speed.
[0212] It can be seen that the method of any of the above embodiments of the present application can significantly improve the training efficiency when the gradient compression algorithm is used to compress the gradient during the actual distributed training process.
[0213] The data processing method provided in the embodiments of the present application has been described above from multiple aspects. The data processing device provided in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0214] As shown in FIG9 , an embodiment of the present application provides a data processing device 90 , which is applied to a first computing node and includes:
[0215] A computing module 901 is configured to execute, through a first thread, a computing operation in a training process of a neural network of a first computing node to obtain a gradient;
[0216] The processing module 902 is configured to compress the gradient through a second thread to obtain a compressed gradient.
[0217] Optionally, the processing module 902 is configured to:
[0218] Determine whether to compress the gradient, the gradient that needs to be compressed is the first gradient, and the gradient that does not need to be compressed is the second gradient;
[0219] When the gradient is compressed, the first gradient is compressed to obtain a compressed first gradient.
[0220] Optionally, the processing module 902 is configured to determine whether to compress the gradient according to a first preset condition, where the first preset condition is determined based on communication overheads of the gradient before compression and the gradient after compression.
[0221] Optionally, the processing module 902 is configured to: determine, through a third thread, according to a first preset condition, whether to compress the gradient.
[0222] Optionally, the processing module 902 is configured to:
[0223] When the gradient is compressed, the first gradient is stored in the compression task queue through the third thread;
[0224] The second thread obtains the first gradient from the compression task queue for compression.
[0225] Optionally, the communication overhead is determined based on one or more of the following information:
[0226] The network transmission method of the gradient, the network transmission speed, the compression speed of the gradient, and the decompression speed of the compressed gradient.
[0227] Optionally, the processing module 902 is configured to send the compressed first gradient through a fourth thread.
[0228] Optionally, the second gradient and the compressed first gradient are stored in a communication task queue;
[0229] The processing module 902 is configured to package and send the second gradient and the compressed first gradient in the communication task queue through a fourth thread when the second gradient and the compressed first gradient meet a second preset condition.
[0230] Optionally, the processing module 902 is configured to send the second gradient through a fourth thread without compressing the gradient.
[0231] Optionally, the processing module 902 is configured to:
[0232] receiving a gradient of a second computing node, where the second computing node is a computing node other than the first computing node, and the gradient of the second computing node includes a compressed gradient;
[0233] The fifth thread decompresses the gradient of the second computing node.
[0234] Optionally, the received gradient of the second computing node is stored in a decompression task queue;
[0235] The processing module 902 is configured to obtain the gradient of the second computing node from the decompression task queue through a fifth thread for decompression.
[0236] Figure 10 shows a possible logical structure diagram of a computing node 100 provided in an embodiment of the present application. The computing node 100 is used to implement the functions of the electronic device involved in any of the above embodiments. The computing node 100 includes: a memory 1001, a processor 1002, a communication interface 1003, and a bus 1004. The memory 1001, processor 1002, and communication interface 1003 are connected to each other via bus 1004.
[0237] The memory 1001 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1001 may store programs. When the program stored in the memory 1001 is executed by the processor 1002, the processor 1002 and the communication interface 1003 are used to perform one or more steps in the above-described data processing method embodiment.
[0238] The processor 1002 can be a central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component or any combination thereof, used to execute relevant programs to implement the functions required to be executed by the calculation module and the processing module in the data processing device in the above embodiment, or to perform one or more steps in the embodiment of the method of the present application. The steps of the method disclosed in conjunction with the embodiment of the present application can be performed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 1001, and the processor 1002 reads the information in the memory 1001 and performs one or more steps in the above-mentioned data processing method embodiment in combination with its hardware.
[0239] The communication interface 1003 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the computing node 100 and other devices or a communication network.
[0240] Bus 1004 provides a pathway for transmitting information between the various components of computing node 100 (e.g., memory 1001, processor 1002, and communication interface 1003). Bus 1004 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, control buses, and so on. For ease of illustration, FIG10 shows a single thick line, but this does not imply that there is only one bus or only one type of bus.
[0241] In another embodiment of the present application, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When the processor of the device executes the computer-executable instructions, the device executes the steps executed by the processor in Figure 10 above.
[0242] In another embodiment of the present application, a computer program product is also provided, which includes computer execution instructions stored in a computer-readable storage medium; when the processor of the device executes the computer execution instructions, the device executes the steps performed by the processor in Figure 10 above.
[0243] In another embodiment of the present application, a chip system is provided, comprising a processor configured to implement the steps performed by the processor in FIG. 10 . In one possible design, the chip system may further comprise a memory configured to store program instructions and data necessary for data writing. The chip system may be comprised of a chip alone or may include a chip and other discrete components.
[0244] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0245] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0246] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0247] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0248] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
Claims
1. A data processing method, characterized in that, The method is applied to a first computing node, and the method includes: Executing, by a first thread, a computing operation in the training process of a neural network of the first computing node to obtain gradients; Compressing, by a second thread, the gradients to obtain compressed gradients.
2. The method according to claim 1, characterized in that, Before compressing the gradients, it further includes: Judging whether to compress the gradients, the gradients that need to be compressed are first gradients, and the gradients that do not need to be compressed are second gradients; The compressing of the gradients includes: In the case of compressing the gradients, compressing the first gradients to obtain compressed first gradients.
3. The method according to claim 2, wherein The judging whether to compress the gradients includes: Judging whether to compress the gradients according to a first preset condition, where the first preset condition is determined based on the communication overheads of the gradients before and after compression.
4. The method according to claim 3, characterized in that, The judging whether to compress the gradients according to the first preset condition includes: Judging, by a third thread, whether to compress the gradients according to the first preset condition.
5. The method according to claim 4, wherein After judging, by a third thread, whether to compress the gradients according to the first preset condition, it further includes: In the case of compressing the gradients, storing, by the third thread, the first gradients into a compression task queue; The compressing, by a second thread, the gradients to obtain compressed gradients includes: Obtaining, by the second thread, the first gradients from the compression task queue for compression.
6. The method according to any one of claims 3 to 5, characterized in that, The communication overheads are determined based on one or more of the following information: The network transmission mode of the gradients, the network transmission speed, the compression speed of the gradients, the decompression speed of the compressed gradients.
7. The method according to any one of claims 2-6, characterized in that, After compressing the first gradients to obtain compressed first gradients, it further includes: Sending, by a fourth thread, the compressed first gradients.
8. The method according to claim 7, wherein The second gradients and the compressed first gradients are stored in a communication task queue; The sending, by a fourth thread, the compressed first gradients includes: When the second gradients and the compressed first gradients in the communication task queue meet a second preset condition, packing and sending, by a fourth thread, the second gradients and the compressed first gradients.
9. The method according to any one of claims 2-8, characterized in that, After judging whether to compress the gradients, it further includes: In the case of not compressing the gradients, sending, by a fourth thread, the second gradients.
10. The method according to any one of claims 1-9, characterized in that, The method further includes: Receiving the gradients of a second computing node, where the second computing node is a computing node other than the first computing node, and the gradients of the second computing node include compressed gradients; Decompressing, by a fifth thread, the gradients of the second computing node.
11. The method according to claim 10, wherein The received gradients of the second computing node are stored in a decompression task queue; The decompressing, by a fifth thread, the gradients of the second computing node includes: Obtaining, by the fifth thread, the gradients of the second computing node from the decompression task queue for decompression.
12. A data processing device, characterized in that, Applied to a first computing node, the device includes: A computing module, configured to perform computing operations during the training process of the neural network of the first computing node through a first thread to obtain gradients; A processing module, configured to compress the gradients through a second thread to obtain compressed gradients.
13. The apparatus according to claim 12, wherein The processing module is configured to: Determine whether to compress the gradients, the gradients that need to be compressed are the first gradients, and the gradients that do not need to be compressed are the second gradients; In the case of compressing the gradients, compress the first gradients to obtain compressed first gradients.
14. The apparatus according to claim 13, wherein The processing module is configured to: determine whether to compress the gradients according to a first preset condition, and the first preset condition is determined based on the communication overheads of the gradients before and after compression.
15. The apparatus according to claim 14, wherein The processing module is configured to: determine whether to compress the gradients according to the first preset condition through a third thread.
16. The apparatus according to claim 15, wherein The processing module is configured to: In the case of compressing the gradients, store the first gradients in a compression task queue through the third thread; Obtain the first gradients from the compression task queue through the second thread for compression.
17. The device according to any one of claims 14-16, characterized in that, The communication overhead is determined based on one or more of the following information: The network transmission mode of the gradients, the network transmission speed, the compression speed of the gradients, and the decompression speed of the compressed gradients.
18. The apparatus according to any one of claims 13-17, wherein The processing module is configured to: send the compressed first gradients through a fourth thread.
19. The device according to claim 18, wherein The second gradients and the compressed first gradients are stored in a communication task queue; The processing module is configured to: when the second gradients and the compressed first gradients in the communication task queue meet a second preset condition, package and send the second gradients and the compressed first gradients through the fourth thread.
20. The apparatus according to any one of claims 13-19, wherein The processing module is configured to: send the second gradients through a fourth thread in the case of not compressing the gradients.
21. The apparatus according to any one of claims 12-20, wherein The processing module is configured to: Receive the gradients of a second computing node, the second computing node is a computing node other than the first computing node, and the gradients of the second computing node include compressed gradients; Decompress the gradients of the second computing node through a fifth thread.
22. The device according to claim 21, characterized in that, The received gradients of the second computing node are stored in a decompression task queue; The processing module is configured to: obtain the gradients of the second computing node from the decompression task queue through the fifth thread for decompression.
23. A computing node, characterized in that, The computing node includes at least one processor, a memory, and instructions stored on the memory and executable by the at least one processor, and the at least one processor executes the instructions to implement the steps of the method according to any one of claims 1-11.
24. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1-11.
25. A computer program product comprising instructions, characterized in that, When the instructions are executed by a processor, the method according to any one of claims 1-11 is implemented.
Citation Information
Patent Citations
Edge-end collaborative gradient compression polymerization method and device
CN112418440A
Gradient synchronization method for compressed sensing in distributed deep learning training scene
CN113592089A
Gradient compression method and device, equipment and storage medium
CN114386622A
Federated learning gradient quantification method, efficient communication federated learning method and related devices
CN115392348A
Cloud computing data compression for allreduce in deep learning
US20200311539A1