Distributed Training System and Method for Layer Sparse Neural Network Based on Supernet

By designing a hypernet on each computer node and selecting gradient transmission of important layers through sparse operation, the problem of large communication overhead in distributed neural network training is solved, training efficiency and accuracy is improved, and it is suitable for distributed training of image classification and natural language processing neural networks.

CN119272840BActive Publication Date: 2025-07-04SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411286352.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2025-07-04
Estimated Expiration
2044-09-13

AI Technical Summary

Technical Problem

The existing distributed neural network training methods consume too much communication overhead, resulting in low training efficiency. The existing gradient compression technology fails to effectively reduce communication overhead and increase computing costs.

Method used

A supernet is designed on each computer node, and the gradient of the important layer is selected for transmission through sparse operations. The layer's importance probability is smoothed by using the exponential moving average value, and the k-layer gradient with the greatest probability is selected for transmission, reducing the gradient transmission of non-important layers.

Benefits of technology

It improves the transmission efficiency of distributed training, reduces communication overhead, reduces training time, and maintains the training accuracy of neural networks. It is suitable for distributed training of image classification and natural language processing neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119272840B_ABST
    Figure CN119272840B_ABST
Patent Text Reader

Abstract

The present invention discloses a distributed training system and method for a layer-sparsified neural network based on a supernetwork. The distributed training system includes a global node and multiple computer nodes, and a supernetwork is set on each computer node. The computer node is used to obtain a copy of the neural network to be trained, and based on the local training dataset, train the neural network to be trained and the supernetwork to obtain the complete gradient of the neural network to be trained for the current training. Perform a layer-sparsification operation according to the output of the supernetwork, select the gradients of multiple layers in the neural network to be trained for transmission. The global node is used to receive the gradients uploaded by all computer nodes and perform aggregation, update the global neural network weights, and determine whether the global neural network converges or the number of training times reaches a preset number. If so, determine the trained neural network model based on the network models trained by multiple computer nodes. Otherwise, distribute the global neural network weights to all computer nodes for re-training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to neural network training technology, and in particular to a distributed training system and method for layer-sparse neural network based on supernet. Background Art

[0002] As the complexity of deep neural networks (DNNs) continues to increase, the scale of their parameters continues to grow, resulting in extremely time-consuming training. Distributed training strategies have been introduced into the neural network training process to accelerate the training speed and reduce the training time. For example, in the training process of convolutional networks for image classification or neural networks for natural language processing, the training time of a single computer node is 1 hour. Theoretically, adding 1 node can achieve 2-fold acceleration and reduce 50% of the training time. However, in practice, we can only reduce the training time by 20-30%. The reason is that the efficiency of distributed training is reduced by the overhead introduced by frequent communication between different computer nodes, resulting in the training efficiency not reaching the expected effect.

[0003] To solve the above problems, related research work has introduced various gradient compression techniques (such as sparsification, quantization, low rank) to reduce communication overhead and reduce the overall training time. However, these methods mainly focus on numerical characteristics. For example, reducing the data transmission scale or changing the gradient stored in floating-point format to a data structure with fewer bits, ignoring the convergence characteristics inherent in the neural network training process. In addition, these techniques need to perform operations on all tensors and perform compression operations on each element, which consumes a large amount of computing time. Although the above methods can reduce the communication time, they increase the data compression time, and the overall time is not significantly reduced, that is, the improvement degree of the training efficiency of the neural network is not ideal. Summary of the Invention

[0004] Aiming at the above deficiencies in the prior art, the distributed training system and method for layer-sparse neural network based on supernet provided by the present invention solve the problem of low efficiency in distributed training of neural networks by existing methods.

[0005] To achieve the above invention purpose, the technical solution adopted by the present invention is as follows:

[0006] In the first aspect, a distributed training system for layer-sparse neural network based on supernet is provided, which includes a global node and multiple computer nodes, and a supernet is set on each computer node;

[0007] The computer node is used to obtain a copy of the neural network to be trained, and based on the local training data set, train the neural network to be trained and the supernet to obtain the complete gradient of the neural network to be trained for the current training; perform layer sparsification operation according to the output of the supernet, and select the gradients of multiple layers in the neural network to be trained for transmission;

[0008] The global node is used to receive the gradients uploaded by all computer nodes, aggregate them, update the global neural network weights, and determine whether the global neural network converges or the number of training times reaches a preset number. If so, a trained neural network model is determined based on the network models trained by multiple computer nodes; otherwise, the global neural network weights are distributed to all computer nodes for retraining.

[0009] The neural network to be trained is an image classification ResNet or a natural language processing neural network.

[0010] Furthermore, the method for performing layer sparsification operations based on the output of the supernet includes:

[0011] Calculate the difference between the outputs of the supernet in two adjacent trainings, and use the exponential moving average to smooth the difference to obtain the probability of each layer of the neural network to be trained generated by the supernet being selected.

[0012] According to the probability of each layer of the neural network to be trained generated by the supernet being selected, select the gradients of the k layers with the largest probability for transmission.

[0013] Furthermore, the expression of the output of the supernet is:

[0014] α i = HN i (v i , ψ i ) where HN i () is the supernet of the i-th computer node, υ i and ψ i are the embedding vector of the i-th computer node and the network weights of the supernet respectively; α i is the output of the supernet of the i-th computer node, that is, the probability of each layer of the neural network to be trained generated by the supernet of the i-th computer node being selected.

[0015] The expression for smoothing the difference using the exponential moving average is:

[0016]

[0017] where ∈ is a hyperparameter; are the probabilities of each layer of the neural network to be trained generated by the supernet of the i-th computer node being selected at the (t - 1)-th and t-th iterations of the supernet respectively; is the difference between the output of the supernet of the i-th computer node at the t-th iteration and the output at the (t - 1)-th iteration.

[0018] Furthermore, the expression of the objective function of the supernet is:

[0019]

[0020] Among them, f i (θ i ; α i ) is the objective function value of the supernetwork of the i-th computer node; θ i is the neural network weight to be trained on the i-th computer node; and are the weights of the l1-th and ln-th layers in θ i respectively; L i (·) is the loss function; and are the probabilities of the l1-th and ln-th layers of the neural network to be trained being selected in α i respectively; ⊙ is the Hadamard product.

[0021] Further, the expressions for updating the gradients of the embedding vector and the supernetwork weights during each iteration are:

[0022]

[0023] Among them, is the gradient of υ i ; is the gradient of ψ i ; θ i is the neural network weight to be trained on the i-th computer node; is the intermediate parameter; and are the weights of the l1-th and ln-th layers in θ i respectively; L i (·) is the loss function; and are the probabilities of the l1-th and ln-th layers of the neural network to be trained being selected in α i respectively; ⊙ is the Hadamard product.

[0024] Further, the supernetwork of each computer node includes an input layer, two fully connected layers, and an output layer, and each fully connected layer includes multiple neurons.

[0025] Second, a distributed training method for a layer-sparse neural network based on a supernetwork is provided, which includes the steps of:

[0026] S1. Arrange a supernetwork on each computer node, and each computer node obtains a copy of the neural network to be trained;

[0027] S2. Train the neural network to be trained and the supernetwork based on the local training dataset to obtain the complete gradient of the neural network to be trained for the current training;

[0028] S3. Perform layer sparsification operations based on the output of the supernet, and select the gradients of multiple layers in the neural network to be trained for transmission;

[0029] S4. Receive the gradients uploaded by all computer nodes and aggregate them to update the global neural network weights;

[0030] S5. Determine whether the global neural network converges or whether the number of training times reaches the preset number. If so, go to step S6; otherwise, distribute the global neural network weights to all computer nodes and return to step S2;

[0031] S6. Based on the network models trained by multiple computer nodes, determine the trained neural network model;

[0032] The neural network to be trained is an image classification ResNet or a natural language processing neural network.

[0033] Furthermore, the method for performing layer sparsification operations based on the output of the supernet includes:

[0034] Calculate the difference between the outputs of the supernet in two adjacent trainings, and use the exponential moving average to smooth the difference to obtain the probability of each layer of the neural network to be trained generated by the supernet being selected;

[0035] According to the probability of each layer of the neural network to be trained generated by the supernet being selected, select the gradients of the k layers with the highest probability for transmission.

[0036] Furthermore, the expression of the output of the supernet is:

[0037] α i = HN i (v i , ψ i )

[0038] where α i is the output of the i-th computer node; HN i () is the supernet of the i-th computer node, v i and ψ i are the embedding vector of the i-th computer node and the network weights of the supernet respectively; α i is the output of the supernet of the i-th computer node, that is, the probability of each layer of the neural network to be trained being selected by the supernet of the i-th computer node;

[0039] The expression for smoothing the difference using the exponential moving average is:

[0040]

[0041] where ∈ is a hyperparameter; are the probabilities that each layer of the neural network to be trained generated by the supernet of the i-th computer node is selected at the (t-1)-th and t-th iterations, respectively; is the difference between the output of the supernet of the i-th computer node at the t-th iteration and the output at the (t-1)-th iteration.

[0042] The beneficial effects of the present invention are as follows: In this solution, an independent supernet is designed on each computer node and trained to generate the weights of each layer of the neural network to be trained to represent the importance of that layer. That is, the supernet is used to identify the most "important" layers in each iteration, and the gradients of these "important" neural network layers are transmitted, while the other layers are discarded and gradient accumulation is performed for subsequent transmission. By this means, this solution improves the transmission efficiency in the distributed training process, avoids the additional computational cost brought by frequent operations on tensor elements in the current mainstream gradient compression methods, and thus improves the training efficiency of the neural network.

[0043] Through comprehensive experiments on Resnet-18 and VGG-16, the results show that this solution can reduce the communication overhead with a small loss of accuracy. In addition, the method provided by this solution can be used in overlapping with other compression methods to further reduce the communication volume, especially suitable for the distributed training of current mainstream graph classification and natural language processing neural networks, which can greatly reduce the communication overhead and shorten the training time. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a schematic diagram of distributed training of a layer-sparse neural network based on a supernet.

[0045] Figure 2 is a flowchart of a method for distributed training of a layer-sparse neural network based on a supernet. DETAILED DESCRIPTION OF THE INVENTION

[0046] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

[0047] This solution provides a distributed training system for a layer-sparse neural network based on a supernet, which includes a global node and multiple computer nodes. A supernet is set on each computer node. The supernet includes an input layer, two fully connected layers, and an output layer, and each fully connected layer includes multiple neurons.

[0048] The computer node is used to obtain a copy of the neural network to be trained, and based on the local training dataset, train the neural network to be trained and the hypernetwork to obtain the complete gradient of the current training of the neural network to be trained; perform layer sparsification operations according to the output of the hypernetwork, and select the gradients of multiple layers in the neural network to be trained for transmission.

[0049] The global node is used to receive the gradients uploaded by all computer nodes and aggregate them, update the global neural network weights, and determine whether the global neural network converges or the number of training times reaches the preset number of times. If so, determine the trained neural network model based on the network models trained by multiple computer nodes. Otherwise, distribute the global neural network weights to all computer nodes for re-training.

[0050] The neural network to be trained is an image classification ResNet or a natural language processing neural network.

[0051] The hypernetworks on the computer node is a special neural network architecture. Its core idea is to train the weights of another larger network in a smaller network. As Figure 1 shown, the hypernetwork realizes the mapping of model parameters associated with the target task to the corresponding model parameters by mapping the embedding vector. In this solution, the hypernetwork is used on each computer node and is trained to generate the weights of each layer of the network to represent the importance of that layer. Then, after calculating the gradients of each layer, the target network is pruned according to the weights generated by the hypernetwork. Different from manual pruning, this solution allows the pruning structures of different node networks to be different. This means that our method is more in line with the dynamic characteristics of the training process.

[0052] In an embodiment of the present invention, the method for performing layer sparsification operations according to the output of the hypernetwork includes:

[0053] Calculate the difference between the outputs of the hypernetwork in two adjacent trainings. The expression for the output of the hypernetwork is:

[0054] α i = HN i (v i , ψ i )

[0055] where NH i () is the hypernetwork of the i-th computer node, υ i and ψ i are the embedding vector and the network weights of the hypernetwork of the i-th computer node respectively; α i is the output of the hypernetwork of the i-th computer node, that is, the probability that the hypernetwork of the i-th computer node selects each layer of the neural network to be trained;

[0056] The exponential moving average is used to smooth the difference, and the probability of each layer of the neural network to be trained generated by the supernet is obtained; the expression for smoothing the difference using the exponential moving average is:

[0057]

[0058] where ∈ is a hyperparameter; are the probabilities of each layer of the neural network to be trained generated by the supernet of the i-th computer node at the (t - 1)-th and t-th iterations respectively; is the difference between the output of the supernet of the i-th computer node at the t-th iteration and the output at the (t - 1)-th iteration.

[0059] According to the probabilities of each layer of the neural network to be trained generated by the supernet, the gradients of the k layers with the largest probabilities are selected for transmission.

[0060] In actual use, α i is directly generated by the supernet, and this scheme uses the change Δα i between two times of α i to represent the importance of this layer. However, in some extreme cases, due to problems with training samples, the change in this value is relatively large and it is not very suitable to directly represent the importance of this layer. This scheme can well solve the problem of large change in Δα i caused by training sample problems by using the exponential moving average to smooth the difference.

[0061] In order to make the weights generated for each layer approximately reflect their importance in each training iteration, in implementation, the expression of the objective function of the supernet is preferably:

[0062]

[0063] where f i (θ i ; α i ) is the objective function value of the supernet of the i-th computer node; θ i is the weight of the neural network to be trained on the i-th computer node; and are the weights of the l1-th and ln-th layers in θ i respectively; L i (·) is the loss function; and are the probabilities of the l1-th and ln-th layers of the neural network to be trained being selected in α i respectively; ⊙ is the Hadamard product.

[0064] The purpose of the loss function is to make α iThe range is restricted between 0 and 1. The objective function essentially represents the loss of the original network, with each layer multiplied by a shrinkage factor. When all α = 1, the function value will reach the minimum. Obviously, the network without any modification will produce the minimum loss. To minimize the function value, the expressions for updating the gradients of the embedding vector and the hypernetwork weights at each iteration are as follows:

[0065]

[0066] Among them, is the gradient of υ i ; is the gradient of ψ i ; θ i is the neural network weight to be trained on the i-th computer node; is an intermediate parameter; and are the weights of the l1-th and ln-th layers in θ i respectively; L i (·) is the loss function; and are the probabilities of the l1-th and ln-th layers of the neural network to be trained being selected in α i respectively; ⊙ is the Hadamard product.

[0067] As Figure 2 shown, this solution also provides a distributed training method for a layer-sparse neural network based on a hypernetwork. This method S includes steps S1 to S6:

[0068] In step S1, a hypernetwork is arranged on each computer node, and each computer node obtains a copy of the neural network to be trained;

[0069] In step S2, the neural network to be trained and the hypernetwork are trained based on the local training dataset to obtain the complete gradient of the current training of the neural network to be trained;

[0070] In step S3, layer sparsification operations are performed according to the output of the hypernetwork, and the gradients of multiple layers in the neural network to be trained are selected for transmission; the methods of the sparsification operations include:

[0071] Calculate the difference between the outputs of the hypernetwork in two adjacent trainings, and use the exponential moving average to smooth the difference to obtain the probability of each layer of the neural network to be trained generated by the hypernetwork being selected;

[0072] According to the probability of each layer of the neural network to be trained generated by the hypernetwork being selected, select the gradients of the k layers with the largest probability for transmission.

[0073] In step S4, receive the gradients uploaded by all computer nodes and aggregate them to update the global neural network weights;

[0074] In step S5, it is determined whether the global neural network converges or the number of training times reaches a preset number. If so, go to step S6; otherwise, distribute the global neural network weights to all computer nodes and return to step S2;

[0075] In step S6, based on the network models completed by training on multiple computer nodes, a neural network model that has completed training is determined.

[0076] To verify the effectiveness of the distributed training method of this solution, the effect of this solution is described below in combination with the training of two neural networks:

[0077] In this embodiment, two image classification neural networks (ResNet-18 and VGG-16) are selected for training on two different image classification data sets, and the time consumed by data compression and the accuracy of the neural network are compared. For the convenience of the experiment, the distributed training process is simulated on an NVIDIA Tesla V100 card. A total of 2 groups of experiments are set up, and the number of nodes participating in the training (the number of workers) is set to 4 nodes and 8 nodes respectively. For the CIFAR10 data set used for image classification, when performing classification training, the number of training epochs is set to 100, while for the CIFAR100 data set, the number of training epochs is set to 150.

[0078] All hyperparameters remain consistent throughout the experiment. Compared with four classical methods: Baseline is the original method without compression, the method of the largest K elements (Top-k), the quantization method (QSGD), two-way compression (TernGrad), and the distributed training method of the layer sparse neural network based on the super network of this solution (LSH). The transmission compression ratio is 1 / 3, that is, the original neural network gradient has a total of 30 layers, and 10 layers are selected for transmission.

[0079] The accuracy comparison of ResNet-18 based on 5 different gradient compression methods on two data sets (CIFAR10 and CIFAR100) is shown in Table 1. The accuracy comparison of VGG-16 based on 5 different gradient compression methods on two data sets (CIFAR10 and CIFAR100) is shown in Table 2.

[0080] Table 1

[0081]

[0082]

[0083] ResNet-18 was trained on two datasets (CIFAR10 and CIFAR100). Comparing with mainstream methods (data complete transmission, the method of the largest K elements (Top-k), quantization method (QSGD), two-way compression (TernGrad)), the accuracy before and after the method LSH of this solution is shown in Table 3.

[0084] Table 3

[0085]

[0086]

[0087] VGG-16 was trained on two datasets (CIFAR10 and CIFAR100). Comparing with four other methods, the accuracy before and after using the hypernetwork-based layer sparsification method (LSH) is shown in Table 4.

[0088] Table 4

[0089]

[0090]

[0091] It can be seen from Table 1 and Table 2 that the ResNet-18 and VGG-16 networks were tested on two datasets. Under different cluster scales, the method LSH of this solution can be basically consistent with the accuracy of the baseline (uncompressed). In addition, compared with other classic methods in the industry combined with the method LSH of this solution (data complete transmission, the method of the largest K elements (Top-k), quantization method (QSGD), two-way compression (TernGrad)), the accuracy is higher. As shown in Tables 3 and 4, comparing the compression rate, the method LSH of this solution also has obvious advantages.

[0092] To sum up, in the neural network distributed training environment, this solution uses the hypernetwork to help screen the optimal layer for gradient transmission in each iteration, which can reduce the communication overhead and improve the training efficiency, especially suitable for the distributed training of the image classification network ResNet and the natural language neural network model.

Claims

1. A distributed training system for a layer-sparsified neural network based on a supernet, characterized in that It includes a global node and multiple computer nodes, and a supernet is set on each computer node. The computer node is used to obtain a copy of the neural network to be trained, and based on the local training dataset, train the neural network to be trained and the supernet to obtain the complete gradient of the current training of the neural network to be trained; perform layer sparsification operations according to the output of the supernet, and select the gradients of multiple layers in the neural network to be trained for transmission. The method for performing layer sparsification operations according to the output of the supernet includes: Calculate the difference between the outputs of the supernet in two adjacent trainings, and use the exponential moving average to smooth the difference to obtain the probability of each layer of the neural network to be trained generated by the supernet being selected. According to the probability of each layer of the neural network to be trained generated by the super network, select the layer with the largest probability k and transmit the gradient of this layer; The global node is used to receive the gradients uploaded by all computer nodes and aggregate them, update the global neural network weights, and determine whether the global neural network converges or the number of training times reaches the preset number. If so, determine the trained neural network model based on the network models trained by multiple computer nodes. Otherwise, distribute the global neural network weights to all computer nodes for retraining. The neural network to be trained is an image classification ResNet or a natural language processing neural network.

2. The distributed training system of the layer sparsification neural network based on the super network according to claim 1, characterized in that The expression of the output of the supernet is: Among them, is the supernetwork of the i th computer node; and are the embedding vector of the i th computer node and the network weight of the supernetwork, respectively; is the output of the supernetwork of the i th computer node, that is, the probability that the supernetwork of the i th computer node selects each layer of the neural network to be trained; The expression for smoothing the difference using the exponential moving average is: Among them, is a hyperparameter; , are respectively the probabilities that each layer of the neural network to be trained generated by the supernet of the i th computer node is selected at the t -1th and t th iterations; is the difference between the output of the supernet of the i th computer node at the t th iteration and the output at the t -1th iteration.

3. The distributed training system of the layer sparsified neural network based on the super network according to claim 2, wherein The expression of the objective function of the supernet is: Among them, is the objective function value of the super network of the i th computer node; is the neural network weight to be trained on the i th computer node; and are respectively the weights of the and th layers in is the loss function; and are respectively the probabilities that the and th layers of the neural network to be trained are selected in is the Hadamard product.

4. The distributed training system of the layer sparsification neural network based on the super network according to claim 2, characterized in that, The expression for updating the gradients of the embedding vector and the supernet network weights during each iteration is: Among them, is 's gradient; is 's gradient; is the neural network weights to be trained on the i th computer node; is the intermediate parameter; and are respectively 's weights of the and th layers; is the loss function; and are respectively 's probabilities of the and th layers of the neural network to be trained being selected; is the Hadamard product.

5. The distributed training system of the layer-sparse neural network based on the super network according to claim 2, characterized in that, The supernet of each computer node includes an input layer, two fully connected layers, and an output layer, and each fully connected layer includes multiple neurons.

6. A distributed training method for layer sparsification neural network based on supernet, characterized in that, It includes steps: S1. Arrange a supernet on each computer node, and each computer node obtains a copy of the neural network to be trained. S2. Based on the local training dataset, train the neural network to be trained and the supernet to obtain the complete gradient of the current training of the neural network to be trained. S3. Perform layer sparsification operations according to the output of the supernet, and select the gradients of multiple layers in the neural network to be trained for transmission. The method for performing layer sparsification operations according to the output of the supernet includes: Calculate the difference between the outputs of the supernet in two adjacent trainings, and use the exponential moving average to smooth the difference to obtain the probability of each layer of the neural network to be trained generated by the supernet being selected. According to the probability of each layer of the neural network to be trained generated by the super network, select the layer with the largest probability k and transmit the gradient of this layer; S4. Receive the gradients uploaded by all computer nodes and aggregate them, and update the global neural network weights. S5. Determine whether the global neural network converges or the number of training times reaches the preset number. If so, enter step S6. Otherwise, distribute the global neural network weights to all computer nodes and return to step S2. S6. Based on the network models trained by multiple computer nodes, determine the trained neural network model. The neural network to be trained is an image classification ResNet or a natural language processing neural network.

7. The distributed training method of the layer-sparse neural network based on the super network according to claim 6, characterized in that The expression of the output of the supernet is: in, is the output of the i-th computer node; For the i A supernet of computer nodes, and Respectively i The embedding vectors of the computer nodes and the network weights of the supernet; For the i The output of the supernet of computer nodes, that is, i The probability of each layer of the neural network being selected by a supernet of computer nodes for training; The expression for smoothing the difference using the exponential moving average is: Among them, is a hyperparameter; , are the probabilities that each layer of the neural network to be trained generated by the supernetwork of the i-th computer node is selected at the (t-1)-th and t-th iterations respectively; is the difference between the output of the supernetwork of the i-th computer node at the t-th iteration and the output at the (t-1)-th iteration.

Citation Information

Patent Citations

  • Hierarchical network structure search method, device and readable storage medium

    CN111860495A

  • Distributed deep learning training method and system based on layer rarefaction

    CN114298277A