Adaptive Distributed Gradient Compression Method Based on Top-k for Supporting Complex Network Conditions

By adopting an adaptive distributed gradient compression method in distributed deep neural networks, the k size in Top-k is dynamically adjusted, and the adaptive compression rate adjustment is combined with training accuracy and communication time, the high efficiency of training time and test accuracy under complex network conditions is solved, and an efficient training process and high-precision test results are achieved.

CN113988266BActive Publication Date: 2025-06-27NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111282366.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-01
Publication Date
2025-06-27
Estimated Expiration
2041-11-01

AI Technical Summary

Technical Problem

The prior art is difficult to shorten the training time of distributed deep neural networks while maintaining high test accuracy under complex network conditions, especially when network communication overhead is high.

Method used

Adaptive distributed gradient compression method based on Top-k is adopted, and the deep neural network model is independently run in distributed nodes, and the adaptive compression rate adjustment is performed in combination with training accuracy and communication time, and the k size in Top-k is dynamically adjusted to adapt to complex network conditions.

Benefits of technology

It effectively reduces the training time of distributed deep neural networks, while maintaining or exceeding the testing accuracy of uncompressed methods, solving the high efficiency of training time and testing accuracy under complex network conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113988266B_ABST
    Figure CN113988266B_ABST
Patent Text Reader

Abstract

The present invention discloses an adaptive distributed gradient compression method based on Top-k for supporting complex network conditions, including each distributed node running a deep neural network learning model to complete the gradient calculation process and save the training accuracy of the current round; using an adaptive gradient compression algorithm pre-deployed on each distributed node to generate compression rate adjustment decisions for different network conditions; adaptively changing the current gradient compression rate in each distributed node according to the generated compression rate adjustment decisions; after the distributed gradient communication process is completed, each distributed node saves the communication time of the current round, and repeats the next round of distributed neural network training. The present invention adaptively adjusts the gradient communication compression rate by combining multi-dimensional evaluation features in distributed deep neural network training, is applicable to complex real-time network conditions, reduces the training time of distributed deep neural networks, and improves the accuracy of test data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of distributed machine learning, and in particular relates to a Top-k based adaptive distributed gradient compression method supporting complex network conditions. Background Art

[0002] With the rapid development of computer hardware (GPU), deep learning has ushered in a wave of revival. It can be widely used in natural language, image recognition, sentiment analysis and other technical processing. However, since ordinary deep neural networks contain millions to tens of millions of parameter settings, it takes a lot of time to train the model. As an alternative, the emergence of distributed deep neural networks can transfer the computing burden of a single GPU to multiple GPU nodes by splitting the training data and deploying it to the remaining working nodes.

[0003] By applying distributed working nodes, the demand for powerful computing power of deep neural networks can be alleviated. However, when distributed neural networks encounter poor network conditions, the network communication time will increase exponentially due to unstable network bandwidth and lags in some nodes during training. In this case, the acceleration benefits brought by distributed working nodes will be severely restricted by the network communication overhead.

[0004] Existing methods to reduce distributed communication overhead include quantization and compression. Quantization methods reduce the bit width of each gradient, thereby reducing the total amount of communication to speed up training. However, they may inevitably lead to a decrease in test accuracy. Although recent work is committed to maintaining the accuracy of test accuracy, they will prolong training time. Compared with quantization, sparsification can flexibly reduce training time by reducing the number of transmitted gradients. Top-k selects the top k% of the absolute size of the gradient for communication aggregation. As a typical representative of sparsification methods, it also has the risk of damaging test accuracy. For the traditional Top-k method using a fixed compression ratio, if this fixed compression ratio is small, that is, the k value is small, the test accuracy may decrease compared with the method where the gradient is not compressed, otherwise the training time may be extended due to the high computational cost of the gradient sorting operation.

[0005] The communication methods of distributed deep neural networks can be mainly divided into two types: All-Reduce and Parameter Server methods. All-Reduce is mainly applied to synchronous communication, and the widely used technology in this is Ring All-Reduce. Through the gradient reduction operation, each distributed node can obtain the global gradient value after the gradient aggregation of all nodes. For the Parameter Server operation, that is, the parameter server can support two communication modes: synchronous and asynchronous, and it can be deployed on the CPU or GPU. After each node sends the calculated gradient to the parameter server, the parameter server will decide whether to set a synchronization barrier according to different types of communication modes. For synchronous communication, the setting of the synchronization barrier can make each worker node wait until the parameter server completes the aggregation operation of all gradients, while asynchronous communication does not require waiting. After the aggregation is completed, the parameter server returns the aggregated global gradient to each worker node, and then the next round of operations is carried out. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide an adaptive distributed gradient compression method based on Top-k that supports complex network conditions in view of the above-mentioned prior art, reduces the training time while maintaining a high test accuracy, and even exceeds the test accuracy of the existing uncompressed methods, and solves the problem that it is difficult to achieve high benefits of test accuracy and training time in the existing complex network conditions for the adaptive distributed gradient compression algorithm.

[0007] To achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:

[0008] An adaptive distributed gradient compression method based on Top-k that supports complex network conditions, characterized by including:

[0009] S1. For complex networks (including static network scenarios and dynamic network scenarios), each distributed node independently runs the held deep neural network learning model, completes the gradient calculation process, and saves the training accuracy of the current round.

[0010] S2. Combining the training accuracy and gradient communication time of several previous rounds, use the adaptive gradient compression algorithm pre-deployed on each distributed node to generate compression rate adjustment decisions for different network conditions.

[0011] S3. For the generated compression rate adjustment decisions, adaptively change the current gradient compression rate in each distributed node, that is, the size of k in Top-k.

[0012] S4. During the distributed communication process, each node completes the aggregated update for the compressed gradient and saves the communication time of the current round. Subsequently, each node independently runs to update the held deep neural network learning model according to the updated gradient, and repeats the operations of S1 to S4. To optimize the above technical solution, the specific measures taken also include:

[0013] In the above step S1, each distributed node independently runs the model and adopts the synchronous communication method.

[0014] The above step S2 includes the following steps:

[0015] S21. For the static network scenario, each node collects the model accuracy and gradient communication time of the previous round segment, where the static network scenario refers to the network bandwidth not changing over time;

[0016] S22. In each distributed node, run the deep decision network model constructed based on the idea of reinforcement learning. This deep decision network model is used to generate the gradient compression adjustment strategy online in real time to adapt to the changing model accuracy;

[0017] Among them, the input of the deep decision network model is the training accuracy and communication time in the previous round. After normalizing the two, a corresponding reward value is generated, and the output is generated through model training, and multiple actions corresponding to the change of the original gradient compression value are output;

[0018] S23. For the dynamic network scenario, each node collects the training accuracy and gradient communication time of the previous several round segments, where the dynamic network scenario refers to the network bandwidth changing over time;

[0019] S24. In each distributed node, run the adaptive compression rate algorithm based on the growth rate of training accuracy and communication time. The growth rate is described by the ratio method. For example, the absolute value of the difference in training accuracy between the two adjacent rounds is divided by the sum of the absolute values of the differences in training accuracy between the adjacent rounds of the previous several rounds to obtain the corresponding ratio to describe the growth rate. By comparing the growth rates of training accuracy and communication time, the compression value is changed with a certain probability.

[0020] The deep decision network model described in the above step S22 is constructed using the DQN method. In the initial stage of training the deep neural network learning model, the deep decision network model has not been trained yet, and it makes random decisions. After the deep decision network model is trained, corresponding decisions are generated according to the input training accuracy and communication time.

[0021] In the above step S22, after normalizing the accuracy and communication time, a reward value is generated and used as part of the reward function. The specific formula is:

[0022] αN(μ - μ min ) + βN(v max - ν);

[0023] N(·) represents the normalization function, μ represents the training accuracy of each round of operation, ν represents the time interval of each round of gradient communication, and μ min and v max represent the acceptable minimum training accuracy and the maximum communication time interval respectively.

[0024] In the above step S22, the normalization function N(·) should be specified for the training accuracy in the static network scenario as:

[0025] N(x) = (Accu - Accu min ) / (Accu max - Accu min );

[0026] Accu is the training accuracy of each round of operation, Accu min is the minimum accuracy that may occur in one round of operation, and Accu max represents the maximum accuracy that may occur in the corresponding one round of operation;

[0027] The normalization function N(·) should be specified for the communication time in the static network scenario as:

[0028] N(x) = (Delay max - Delay) / (Delay max - Delay min );

[0029] Delay max represents the longest training time required to compress the gradient at an arbitrary compression ratio within a limited range and run the complete entire dataset, Delay min represents the shortest training time required to run the complete entire dataset correspondingly, and Delay represents the communication time required for one round of operation.

[0030] In the above step S24, calculate the growth rates of the training accuracy and the communication time within a certain number of rounds of model operation;

[0031] The growth rate of the training accuracy is obtained by dividing the absolute value of the difference in accuracy between the two most recent adjacent rounds by the sum of the absolute values of the differences in accuracy between consecutive adjacent rounds within a specific number of rounds;

[0032] The growth rate of the communication time is obtained by dividing the absolute value of the difference in communication time between the two most recent adjacent rounds by the sum of the absolute values of the differences in communication time between consecutive adjacent rounds within a specific number of rounds;

[0033] Determine the emphasis relationship between accuracy and communication time with respect to the change in compression ratio in the current situation by comparing the growth rates of accuracy and communication time, and thus heuristically and dynamically change the magnitude of the compression ratio with a certain probability.

[0034] In the above-mentioned step S3, the range of adaptively adjusting the gradient compression ratio needs to ensure the reduction of the model training time by compression, and the compression ratio will not be lower than the set extreme minimum value. If the compression ratio after dynamic change exceeds the acceptable range, the gradient compression ratio will be set at the boundary of the acceptable range.

[0035] In the above-mentioned steps S1 and S4, save the corresponding accuracy and communication time to be applicable to different network environments;

[0036] In a static network scenario, after saving the training accuracy and communication time of the previous round and after using them in this round, they can be discarded; while in a dynamic network scenario, set up a queue to continuously save the training accuracy and communication time of several rounds. When the queue is full, dequeue and discard the relevant information that entered the queue earliest.

[0037] The present invention has the following beneficial effects:

[0038] The present invention adaptively adjusts the gradient communication compression ratio by combining multi-dimensional evaluation features in distributed deep neural network training, is applicable to complex real-time network situations, reduces the training time of distributed deep neural networks, and improves the accuracy of test data sets.

[0039] (1) The present invention deeply explores the compression ratio and a series of features of deep neural network models, such as training accuracy, test accuracy, training loss value, and training time. Through the analysis of a large number of experimental results, the present invention first proposes to combine training accuracy and communication time for adaptive compression ratio adjustment.

[0040] (2) The adaptive distributed gradient compression algorithm proposed by the present invention is respectively applied to static networks and dynamic networks, where it is optimized based on DQN in static networks, and optimized based on the growth rates of training accuracy and communication time in dynamic networks. This algorithm explores the best trade-off point between training time and test accuracy while ensuring the convergence of the model.

[0041] (3) The adaptive distributed gradient compression algorithm proposed by the present invention can achieve high efficiency in training time and test accuracy, that is, while maintaining less training time, it also maintains a high model test accuracy, and even in some cases, exceeds the test accuracy in the original uncompressed situation. Description of the Drawings

[0042] Figure 1It is the flowchart of the adaptive distributed gradient compression method based on Top-k that supports complex network conditions of the present invention.

[0043] Figure 2 It is an example of adaptive gradient compression in a static network scenario of the present invention.

[0044] Figure 3 It is an example of adaptive gradient compression in a dynamic network scenario of the present invention. Detailed implementation manners

[0045] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0046] As Figure 1 shown, the adaptive distributed gradient compression method based on Top-k that supports complex network conditions includes:

[0047] S1. Each distributed node runs a deep neural network learning model to complete the gradient calculation process and save the training accuracy of the current round.

[0048] In the embodiment, each distributed node runs the model independently, and aggregates the compressed gradients obtained later in a synchronous communication manner.

[0049] S2. Combining the training accuracy and gradient communication time of several previous rounds, use the adaptive gradient compression algorithm pre-deployed on each distributed node to generate compression rate adjustment decisions for different network conditions, including the following steps:

[0050] S21. For a static network scenario, each node collects the model accuracy and gradient communication time of the previous round, where the static network scenario refers to a network bandwidth that does not change with time;

[0051] S22. Run a deep decision network model constructed based on the idea of reinforcement learning in each distributed node. This deep decision network model is used to generate gradient compression adjustment strategies online in real time to adapt to the changing model accuracy;

[0052] Among them, the input of the deep decision network model is the training accuracy and communication time in the previous round. After normalizing the two, a corresponding reward value is generated, and the output is generated through model training, and multiple actions corresponding to the change of the original gradient compression value are output;

[0053] In the embodiment, the deep decision network model is constructed using the DQN method. In the initial stage of training the deep neural network learning model, the deep decision network model has not been trained yet, and it makes random decisions. After the deep decision network model is trained, corresponding decisions are generated according to the input training accuracy and communication time.

[0054] In the embodiment, after normalizing the accuracy and communication time, a reward value is generated and used as a part of the reward function. The specific formula is as follows:

[0055] αN(μ - μ min ) + βN(v max - ν);

[0056] N(·) represents the normalization function, μ represents the training accuracy of each round of operation, ν represents the time interval of each round of gradient communication, and μ min and v max respectively represent the lowest acceptable training accuracy and the maximum communication time interval.

[0057] In the static network scenario, the normalization function N(·) should be specified for the training accuracy as follows:

[0058] N(x) = (Accu - Accu min ) / (Accu max - Accu min );

[0059] Accu is the training accuracy of each round of operation, Accu min is the minimum accuracy that may occur in a round of operation, and Accu max represents the maximum accuracy that may occur in the corresponding round of operation;

[0060] In the static network scenario, the normalization function N(·) should be specified for the communication time as follows:

[0061] N(x) = (Delay max - Delay) / (Delay max - Delay min );

[0062] Delay max represents the longest training time required to run the complete entire dataset by compressing the gradient at an arbitrary compression ratio within a limited range, Delay min represents the shortest training time required to run the complete entire dataset accordingly, and Delay represents the communication time required to run one round (i.e., a small part of the dataset).

[0063] S23. For the dynamic network scenario, each node collects the training accuracy and gradient communication time of the previous several rounds of data segments, where the dynamic network scenario refers to the network bandwidth changing over time;

[0064] S24. Run the adaptive compression rate algorithm based on the training accuracy and the growth rate of communication time in each distributed node, where the growth rate is described using the proportional method. For example, the corresponding ratio is obtained by dividing the absolute value of the difference in training accuracy between the two most recent adjacent rounds by the sum of the absolute values of the differences in training accuracy between several previous adjacent rounds to describe the growth rate.

[0065] By comparing the growth rates of accuracy and communication time, determine the emphasis relationship between accuracy and communication time with respect to changes in the compression rate in the current situation, and thus heuristically and dynamically change the size of the compression rate with a certain probability (a fixed probability value, and the discriminant basis is detailed in Figure 3 ).

[0066] In the embodiment, calculate the growth rates of training accuracy and communication time within a certain number of model running rounds;

[0067] The growth rate of accuracy is obtained by dividing the absolute value of the difference in accuracy between the two most recent adjacent rounds by the sum of the absolute values of the differences in accuracy within consecutive adjacent rounds within a specific number of rounds;

[0068] The growth rate of communication time is obtained by dividing the absolute value of the difference in communication time between the two most recent adjacent rounds by the sum of the absolute values of the differences in communication time within consecutive adjacent rounds within a specific number of rounds;

[0069] By comparing the growth rates of accuracy and communication time, determine the emphasis relationship between accuracy and communication time with respect to changes in the compression rate in the current situation, and thus heuristically and dynamically change the size of the compression rate with a certain probability.

[0070] S3. For the generated compression rate adjustment decision, adaptively change the current gradient compression rate, that is, the value of k in Top-k, in each distributed node;

[0071] In the embodiment, the range of adaptively adjusting the gradient compression rate should be limited within an acceptable range, which ensures that the compression is effective in reducing the model training time and that the compression rate does not fall below the set extreme minimum value.

[0072] If the compression ratio after the dynamic change exceeds the acceptable range, the gradient compression rate should be set at the boundary of the acceptable range.

[0073] S4. After the distributed gradient communication process is completed, each distributed node saves the communication time of the current round and repeats the next round of distributed neural network training.

[0074] In the embodiment, the corresponding accuracy and communication time saved in steps S1 and S4 are applicable to different network environments;

[0075] In the static network scenario, the training accuracy and communication time of the previous round are saved and can be discarded after being used in this round; while in the dynamic network scenario, a queue needs to be set up, and the training accuracy and communication time of several consecutive rounds need to be continuously saved. When the queue is full, dequeue and discard the relevant information that entered the queue earliest.

[0076] Embodiment

[0077] Use multi-object image classification in the CIFAR-10 and CIFAR-100 datasets. Among them, the CIFAR-10 training dataset includes 10 image classification targets, with 5000 samples in each class, and in the test set, each of the 10 image classification targets contains 1000 samples; the CIFAR-100 training dataset contains 100 image classification targets, with 500 samples in each class, and in the test set, each of the 100 classification targets contains 100 samples.

[0078] To verify the superiority of the complex network condition adaptive distributed gradient compression method supported by the present invention, it is respectively run on the ResNet-18 and VGG-19 deep network models using the CIFAR-10 dataset, and on the ResNet-34 and ResNet-50 deep network models using the CIFAR-100 dataset. Among them, Table 1 is the final test accuracy of the present invention and existing gradient compression algorithms applicable to the static network scenario, Table 2 is the final test accuracy of the present invention and existing gradient compression algorithms applicable to the dynamic network scenario, Table 3 is the running time speedup ratio of the present invention applied to the static network scenario and existing gradient compression algorithms, and Table 4 is the running time speedup ratio of the present invention applied to the dynamic network scenario and existing gradient compression algorithms.

[0079] Table 1 Average test accuracy of each gradient compression algorithm in the last twenty training cycles in the static network scenario

[0080]

[0081] Table 2 Average test accuracy of each gradient compression algorithm in the last twenty training cycles in the dynamic network scenario

[0082]

[0083] Table 3 Speedup ratio of each gradient compression algorithm relative to uncompressed in the static network scenario

[0084]

[0085]

[0086] Table 4 Speedup ratio of each gradient compression algorithm relative to uncompressed in the dynamic network scenario

[0087]

[0088] Among the various gradient compression algorithms presented in Table 1, Table 2, Table 3, and Table 4, Baseline is the algorithm without gradient compression. Top0.001 means that in the Top-k compression algorithm, k is set to 0.001, and Top0.15 means that in the Top-k compression algorithm, k is set to 0.15. DA2, DA4, and DA5 are existing dynamic compression rate adjustment algorithms, and AdaTopK is the adaptive gradient compression algorithm named in the present invention. In the present invention, the parameter α is set to 3 and the parameter β is set to 2.

[0089] Figure 2 This is an example of the adaptive gradient compression algorithm applied to a static network scenario. For simplicity of explanation, this example only sets the current situation at the t-th moment. The neural network decision model in the static network scenario is constructed using the DQN method. In the initial stage of distributed model training, the neural network decision model has not been trained yet, and it makes random decisions. After the neural network decision model is trained, corresponding decisions are generated based on the input training accuracy and communication time.

[0090] After normalization is completed, the calculated corresponding model accuracy reward Acc-Reward and communication time reward Del-Reward are multiplied by the respective influence coefficients α and β and then added together, and used as the final reward value in DQN.

[0091] Figure 3 This is an example of the adaptive gradient compression algorithm applied to a dynamic network scenario. For simplicity of explanation, this example only sets the current situation at the t-th moment. The training accuracy set and communication time set of the previous w window rounds are extracted. The growth rate of the training set accuracy is obtained by dividing the absolute value of the difference in accuracy between the two most recent adjacent rounds by the sum of the absolute values of the differences in accuracy within a specific number of consecutive adjacent rounds, that is, in the figure:

[0092]

[0093] The calculation method of the growth rate of communication time is similar to that of accuracy, and it is generated from the communication time set within a certain number of rounds, that is, in the figure After multiplying the growth rate of the model accuracy and the growth rate of communication time by the respective influence coefficients α and β, by comparing the corresponding growth rate speeds, the emphasis relationship between accuracy and communication time on the compression rate can be determined in the current situation, and thus the size of the compression rate can be heuristically and dynamically changed with a certain probability.

[0094] Combined with Table 1 and Table 2, it can be seen that the proposed AdaTopK algorithm for supporting complex network adaptive gradient compression in the present invention has higher test accuracy in ResNet-18 and ResNet-34. Combining Table 3 and Table 4, it is found that it has a shorter training time, achieving the dual benefits of high accuracy and less training time. In VGG-19 and ResNet-50, the accuracy performance of AdaTopK is not as good as that in the other models, but it still maintains a shorter training time. This is presumably because different deep network models have different preferences for the size of the compression rate, and as a unified adaptive gradient compression algorithm, it is necessary to maintain the high efficiency of the final test accuracy and training time as much as possible in various deep neural network models.

[0095] The novelty of the present invention lies in the Top-k based adaptive gradient compression algorithm for supporting complex networks, which considers various evaluation features in distributed machine learning, uses the training accuracy of the model and the growth rate of communication time as the measurement criteria, and uses different methods to change the gradient compression rate in real time for different network environments, in order to achieve the dual efficiency of the final model test accuracy and training time.

[0096] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. An adaptive distributed gradient compression method based on Top-k and supporting complex network conditions, characterized in that Including: S1. For a complex network, each distributed node independently runs the deep neural network learning model it holds to complete the gradient calculation process and save the training accuracy of the current round. The complex network includes a static network scenario and a dynamic network scenario. S2. Combining the training accuracy and gradient communication time of several previous rounds, use the adaptive gradient compression algorithm pre-deployed on each distributed node to generate compression rate adjustment decisions for different network conditions. Step S2 includes the following steps: S21. For the static network scenario, each node collects the model accuracy and gradient communication time of the previous round. The static network scenario means that the network bandwidth does not change over time. S22. Run the deep decision network model constructed based on the idea of reinforcement learning in each distributed node. This deep decision network model is used to generate gradient compression adjustment strategies online in real time to adapt to the changing model accuracy. The input of the deep decision network model is the training accuracy and communication time in the previous round. After normalizing the two, a corresponding reward value is generated, and the output is generated through model training, and multiple actions corresponding to the change of the original gradient compression value are output. S23. For the dynamic network scenario, each node collects the training accuracy and gradient communication time of several previous rounds. The dynamic network scenario means that the network bandwidth changes over time. S24. Run the adaptive compression rate algorithm based on the growth rate of training accuracy and communication time in each distributed node, and change the compression value with a certain probability by comparing the growth rates of training accuracy and communication time. S3. For the generated compression rate adjustment decision, adaptively change the current gradient compression rate in each distributed node, that is, the size of k in Top-k. S4. In the process of distributed communication, each node completes the aggregation update of the compressed gradient and saves the communication time of the current round. Subsequently, each node independently runs and updates the deep neural network learning model it holds according to the updated gradient. S5. Repeat the operations of S1 to S4.

2. The adaptive distributed gradient compression method based on Top-k for supporting complex network conditions according to claim 1, wherein In step S1, each distributed node independently runs the model and uses the synchronous communication method.

3. The adaptive distributed gradient compression method based on Top-k for supporting complex network conditions according to claim 1, wherein The deep decision network model described in step S22 is constructed using the DQN method. In the initial stage of the training of the deep neural network learning model, the deep decision network model has not been trained yet, and it makes random decisions. After the deep decision network model is trained, corresponding decisions are generated according to the input training accuracy and communication time.

4. The adaptive distributed gradient compression method based on Top-k for supporting complex network conditions according to claim 1, wherein In step S22, the accuracy and communication time are normalized to generate a reward value and used as part of the reward function. The specific formula is: αN(μ - μ min ) + βN(ν max - ν); N(·) represents the normalization function, μ represents the training accuracy of each round of operation, ν represents the time interval of gradient communication per round, and μ min and ν max represent the acceptable minimum training accuracy and the maximum communication time interval respectively, and α and β are the influence coefficients of the accuracy reward and the communication time reward respectively.

5. The adaptive distributed gradient compression method based on Top-k for supporting complex network conditions according to claim 4, wherein In step S22, the normalization function N(·) should be specifically defined for the training accuracy in the static network as: N(x) = (Accu - Accu min ) / (Accu max - Accu min ) ; Accu is the training accuracy for each round of operation, Accu min is the minimum accuracy that may occur in one round of operation, while Accu max represents the maximum accuracy that may occur in the corresponding round of operation; The normalization function N(·) should be specifically defined for the communication time in the static network as: N’(x) = (Delay max - Delay) / (Delay max - Delay min ) Delay max It represents the longest training time required to compress the gradient at an arbitrarily defined compression ratio and complete the entire dataset. Delay min It represents the shortest training time required to complete the entire dataset accordingly. Delay represents the communication time required for one round of operation.

6. The adaptive distributed gradient compression method based on Top-k for supporting complex network conditions according to claim 1, characterized in that, In step S24, the ratio method is used to calculate the growth rates of the training accuracy and communication time of the deep neural network learning model within a certain number of running rounds. The growth rate of training accuracy is obtained by dividing the absolute value of the difference in accuracy between the two most recent adjacent rounds by the sum of the absolute values of the differences in accuracy within consecutive adjacent rounds within a specific number of rounds. The growth rate of communication time is obtained by dividing the absolute value of the difference in communication time between the two most recent adjacent rounds by the sum of the absolute values of the differences in communication time within consecutive adjacent rounds within a specific number of rounds. By comparing the growth rates of accuracy and communication time, the emphasis relationship between accuracy and communication time with respect to changes in the compression ratio in the current situation is determined, and thus the size of the compression ratio is heuristically and dynamically changed with a certain probability.

7. The adaptive distributed gradient compression method based on Top-k for supporting complex network conditions according to claim 1, wherein In step S3, the range of the adaptive adjustment of the gradient compression ratio needs to ensure a reduction in the model training time due to compression, and the compression ratio should not be lower than the set extreme minimum value. If the compression ratio after the dynamic change exceeds the acceptable range, the gradient compression ratio is set at the boundary of the acceptable range.

8. The adaptive distributed gradient compression method based on Top-k for supporting complex network conditions according to claim 1, wherein In step S1 and step S4, the corresponding accuracy and communication time are saved to be applicable to different network environments. In a static network, after saving the training accuracy and communication time of the previous round and after they are used in the current round, they can be discarded. In a dynamic network, a queue is set up to continuously save the training accuracy and communication time for several rounds. When the queue is full, the oldest relevant information is dequeued and discarded.

Citation Information

Patent Citations

  • Federated learning architecture under dynamic bandwidth and unreliable network and compression algorithm of architecture

    CN111447083A

  • Distributed training method for large-scale deep neural network

    CN113515370A