A Method for Training Models with Low Communication Overhead and High Statistical Efficiency in a Distributed Deep Learning System
Through the combination of adaptive adjustment of communication intervals and correction technology, the problems of communication overhead and statistical efficiency in distributed deep learning systems are solved, shorter training time and lower communication overhead are achieved, and training efficiency is improved.
Patent Information
- Application Number
- CN202210023028.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-01-10
AI Technical Summary
In the existing distributed deep learning system, training model methods for synchronous data parallel cannot guarantee low communication overhead and high statistical efficiency at the same time, resulting in too long training time.
The skip communication strategy and correction technology of adaptive communication intervals are adopted, and the communication interval τ is collected through the runtime data collector, and the communication interval τ is adaptively adjusted, and the local model is updated in each iteration, and the global model is updated after several iterations, combining correction technology to reduce the degree of divergence.
It achieves a significant reduction in communication overhead while maintaining high statistical efficiency, shortening training time, and shortening the total training time by 86.76%, 64.83%, 16.37% and 79.26% compared with the existing methods, and reduces communication overhead by 93.3%, 80% and 33.3% respectively.
Smart Images

Figure CN114565007B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of distributed deep learning, and relates to a method for training a model with low communication overhead and high statistical efficiency for synchronous data parallelism in a distributed deep learning system. Background Art
[0002] In the model training for synchronous data parallelism, a distributed deep learning system starts multiple training processes in a cluster, and each training process has a complete model backup. At the same time, the distributed deep learning system divides the entire data set into several data shards and distributes all data shards to each training process. The entire process of model training consists of a series of iterations. In each round of iteration, each training process selects a batch of data from the assigned data shards and calculates parameter updates on the model. In order to synchronize the parameter updates on all data shards, the training processes usually rely on a Parameter Server or an AllReduce communication architecture to aggregate the parameter updates on all training processes. Based on the aggregated parameter updates, each training process updates the parameters of the model. However, the specific process of updating the model parameters is determined by the method of training the model, and different methods of training the model correspond to different update processes. Currently, the methods for training models for synchronous data parallelism in a distributed deep learning system mainly include SSGD, SkipSSGD, LocalSGD, and SMA.
[0003] SSGD is a widely used method for training models in a distributed deep learning system. In each round of iteration, SSGD aggregates the gradients calculated by each training process on a batch of data through network communication and updates the model based on the aggregated gradients. Figure 1 shows three consecutive rounds of iteration in the process of training a model using the SSGD method. In the 10th round of iteration, two training processes calculate the gradients 10 respectively based on the model w and Then, the gradients and are aggregated through network communication, and the average gradient is calculated. Then, SSGD uses the average gradient to update the model from w 10 to w 11 . After completing the update of the model, SSGD enters the next round of iteration. Since each round of iteration involves network communication, the SSGD method usually has a communication bottleneck in a distributed environment.
[0004] Based on SSGD, SkipSSGD adopts a skipping communication strategy to update the model. Instead of aggregating gradients through network communication in each iteration, SkipSSGD conducts communication only once every τ iterations, thus greatly reducing the communication overhead. To retain the obtained gradient information while skipping communication, SkipSSGD maintains a gradient accumulator in each training process to save gradient information. Every τ iterations, SkipSSGD aggregates the gradients in each accumulator through a network communication and updates the model according to the aggregated gradients. Figure 2 shows three consecutive iterations in the process of training a model using the SkipSSGD method with τ = 3. Based on the model w 10 , two training processes calculate the gradients and respectively in the 10th, 11th, and 12th iterations and accumulate these gradients into their respective gradient accumulators. After completing the accumulation of gradients in three consecutive iterations, SkipSSGD aggregates these accumulated gradients through a network communication and updates the model from w 10 to w 11 . Although SkipSSGD reduces the communication overhead, the way of accumulating gradients will result in a large batch size, thereby leading to low statistical efficiency.
[0005] LocalSGD also adopts a skipping communication strategy to reduce the communication overhead. However, different from the way SkipSSGD skips communication by accumulating gradients, LocalSGD maintains multiple local models and a global model. It directly updates the local models with gradients in each iteration and aggregates the parameters of each local model through network communication every τ iterations to update the global model. Figure 3 shows three consecutive iterations in the process of training a model using the LocalSGD method with τ = 3. In the 10th iteration, a training process calculates the gradient and updates the local model to using this gradient. In the 11th and 12th iterations, this training process continues to calculate the gradients and and updates the local model to and Similarly, another training process also updates the local models to and respectively in the 10th, 11th, and 12th iterations. After completing the update of local models in three consecutive iterations, LocalSGD aggregates each local model through a network communication and updates the global model from to Although LocalSGD reduces the communication overhead and maintains a small batch size, the local models diverge more and more as τ increases, leading to low statistical efficiency.
[0006] To reduce the divergence among local models, SMA adopts a correction technique, that is, the global model is used to correct each update of local models. Figure 4 Shows three consecutive iterations during the training of the model by the SMA method. In the 10th iteration, a training process updates the local model from to according to the gradient and correction Another training process updates the local model from to according to the gradient and correction Meanwhile, to ensure that the global model is up-to-date, SMA updates the global model in each iteration. For example, in the 10th iteration, SMA aggregates the corrections calculated in the two training processes (i.e., and ) through network communication to update the global model from to Although SMA maintains a small batch size and reduces the divergence among local models, each of its iterations involves network communication to update the global model, and there are often communication bottlenecks in a distributed environment.
[0007] Generally speaking, the existing methods for training models for synchronous data parallelism all have defects, that is, these methods do not guarantee both low communication overhead and high statistical efficiency at the same time. Summary of the Invention
[0008] To solve the deficiencies of the existing technology, the object of the present invention is to propose a method for training a model with low communication overhead and high statistical efficiency in a distributed deep learning system. This method uses a communication-skipping strategy based on adaptive communication intervals to ensure low communication overhead, and adopts a correction technique to reduce the divergence among local models on the premise of maintaining a small batch size, thereby ensuring high statistical efficiency and ultimately shortening the time for training the model.
[0009] Existing training methods that adopt the skip communication strategy (e.g., SkipSSGD and LocalSGD) require users to manually specify a communication interval τ. However, the communication overhead brought by the same communication interval τ in different running environments (including the computing power of the GPU, the bandwidth of the network, and the size of the model) is often different. That is to say, the communication interval τ suitable for one running environment may not be suitable for another running environment. Therefore, in order to automatically select a suitable communication interval τ in different running environments, the present invention also proposes a method for adaptively adjusting the communication interval during model training. This method is implemented through a runtime data collector and an adaptive communication interval selector. The runtime data collector collects communication and computation time data in the first round of iteration, and the adaptive communication interval selector automatically adjusts the communication interval τ based on the collected communication and computation time data, so that the communication time and computation time in each training epoch are similar, thereby avoiding the communication overhead from becoming the bottleneck in the whole training.
[0010] The specific technical solution for achieving the object of the present invention is as follows:
[0011] A method for training a model with low communication overhead and high statistical efficiency in a distributed deep learning system, the method comprising the following steps:
[0012] Step A: The runtime data collector collects the communication time t cm and computation time t cp data required for the adaptive communication interval during runtime; the adaptive communication interval selector automatically adjusts the communication interval τ through the communication time and computation time data obtained by the above collection;
[0013] Step B: Iteratively train the model, and update the local model using a correction technique in each round of iteration;
[0014] Step C: Update the global model using the skip communication strategy every τ rounds of iteration.
[0015] Wherein, the step that the runtime data collector collects the communication time t cm and computation time t cp data required for the adaptive communication interval during runtime; and the step that the adaptive communication interval selector automatically adjusts the communication interval τ through the communication time and computation time data obtained by the above collection includes:
[0016] Step A1: The system automatically initializes the communication interval τ to 1; sets the communication interval τ = 1, and sets the adaptive communication interval flag bit flag; the adaptive communication interval flag bit is used to indicate whether to enable the adaptive communication interval selector. The adaptive communication interval flag bit flag = true indicates that the adaptive communication interval selector is enabled, that is, the system adaptively selects a suitable communication interval τ; flag = false indicates that the adaptive communication interval selector is disabled, that is, the user needs to specify the communication interval τ; in the present invention, the adaptive communication interval selector is always enabled;
[0017] Step A2: During the operation of the system, the runtime data collector collects the communication time consumption t cm and the calculation time consumption t cp ;
[0018] Step A3: The adaptive communication interval selector adjusts the communication interval τ according to the collected t cm and t cp in the first round of iteration. Specifically, the communication interval τ is adjusted to That is, the communication interval is adjusted to the ceiling of the quotient of t cm and t cp .
[0019] After the communication interval τ is adjusted in the first round of iteration, this communication interval τ is used for all subsequent iterations, that is, in subsequent iterations, the collection of communication and calculation time data and the adjustment of the communication interval τ are no longer performed.
[0020] Among them, the step of iteratively training the model and updating the local model using the correction technique in each round of iteration includes:
[0021] Step B1: Each training process calculates the gradient according to the local model; Among them, i represents the number of the training process, k represents the number of iterations, represents the gradient calculated by the training process numbered i in the kth round of iteration, b (i) represents the batch size on the training process numbered i, represents the local model parameters of the training process numbered i in the kth round of iteration, and L(x, w) represents the loss calculated by the sample x on the model parameters w;
[0022] Step B2: Each training process calculates the correction Among them, i represents the number of the training process, k represents the number of iterations, represents the correction calculated by the training process numbered i in the kth round of iteration, represents the local model parameters of the training process numbered i in the kth round of iteration, represents the global model parameters in the k-th round of iteration, and the correction is the difference between the local model parameters and the global model parameters;
[0023] Step B3: Each training process updates the local model with the calculated gradient and correction. The specific update form is: where i represents the number of the training process, and k represents the number of iterations, and represent the local model parameters of the training process numbered i in the k-th and k + 1-th rounds of iteration, and represent the local model momentum, gradient, and correction of the training process numbered i in the k-th round of iteration respectively. μ, γ, and α represent the momentum coefficient, learning rate, and correction coefficient respectively.
[0024] The local model momentum is initially empty, and the local model momentum is also updated according to the gradient in each round of iteration. The specific update form is: where i represents the number of the training process, and k represents the number of iterations, and represent the local model momentum of the training process numbered i in the k-th and k + 1-th rounds of iteration, represents the gradient of the training process numbered i in the k-th round of iteration, and μ and γ represent the momentum coefficient and learning rate respectively.
[0025] Among them, the step of updating the global model using the skip communication strategy every τ rounds of iteration includes:
[0026] Step C1: Take the remainder of the current iteration number with respect to the adaptively set communication interval τ. If the remainder is 0, it indicates that the global model needs to be updated in the current iteration, and then enter Step C2;
[0027] Step C2: Aggregate the corrections calculated by each training process through a network communication to obtain the aggregated correction where n represents the total number of training processes, and k represents the number of iterations, represents the correction calculated by the training process numbered i in the k-th round of iteration, and use the aggregated correction to update the global model. The specific update form of the global model is: where i represents the number of the training process, and k represents the number of iterations, and represent the global model parameters in the k-th and k + 1-th rounds of iteration respectively, represents the global model momentum in the k-th round of iteration, μ represents the momentum coefficient, and α represents the correction coefficient. The global model momentum is initially empty, and the global model momentum is also updated once every τ rounds of iteration according to the aggregated correction. The specific update form of the global model momentum is: Among them, k represents the number of iterations, and represents the global model momentum in the k-th and (k + 1)-th rounds of iteration, represents the aggregated correction, and μ and α represent the momentum coefficient and the correction coefficient respectively.
[0028] The communication overhead refers to the proportion of communication time in the total time consumption of one Epoch.
[0029] The communication interval refers to the number of iterations between two adjacent iterations when the global model is updated.
[0030] The local model refers to the model updated independently by each training process through locally calculated gradients and corrections.
[0031] The global model refers to the model updated jointly by all training processes through aggregated corrections via network communication, or the global model refers to the model used to correct the local model update in each round of iteration.
[0032] The present invention also proposes a system for implementing the above method, and the system includes: a runtime data collector, an adaptive communication interval selector, and a model update module.
[0033] The runtime data collector collects communication time data t cm and computing time data t cp ;
[0034] The adaptive communication interval selector automatically adjusts the communication interval cm according to the communication time t cp collected by the runtime data collector and the computing time t and uses this communication interval τ for all subsequent iterations;
[0035] The model update module performs iterative updates on the local model and the global model according to the communication interval τ obtained by the adaptive communication interval selection module, that is, in each round of iteration, the local model is updated using the correction technique, and the global model is updated using the skip communication strategy every τ rounds of iteration.
[0036] The beneficial effects of the present invention include:
[0037] In each iteration of the present invention, a correction technique is adopted to update the local models, reducing the divergence degree among the local models, thereby ensuring high statistical efficiency. Meanwhile, during the operation of the system, the time data of communication and computation are collected, and a communication interval τ is adaptively selected based on the collected data. Based on this communication interval τ, the global model is updated through network communication every τ rounds of iteration, thereby reducing the network communication overhead and ultimately shortening the time for training the model. The implementation results show that SkipSMA shortens the total training time by 86.76%, 64.83%, 16.37%, and 79.26% compared with SSGD, SkipSSGD, LocalSGD, and SMA respectively. Because SkipSMA maintains a similar statistical efficiency to SSGD, SkipSSGD, LocalSGD, and SMA, SkipSMA reduces the communication overhead by 93.3%, 80%, and 33.3% compared with SSGD, SkipSSGD, and LocalSGD respectively, and SkipSMA reduces the computation and communication overhead by 79.9% in each epoch compared with SMA. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 FIG. 6 is a schematic diagram of an implementation manner of training a model using the SSGD method in the prior art;
[0039] Figure 2 FIG. 7 is a schematic diagram of an implementation manner of training a model using the SkipSSGD method in the prior art;
[0040] Figure 3 FIG. 8 is a schematic diagram of an implementation manner of training a model using the LocalSGD method in the prior art;
[0041] Figure 4 FIG. 9 is a schematic diagram of an implementation manner of training a model using the SMA method in the prior art;
[0042] Figure 5 FIG. 10 is a schematic diagram of training a model using the SkipSMA method according to an embodiment of the present invention;
[0043] Figure 6 FIG. 11 is a comparison chart of the total training time of SkipSMA and existing training methods in the implementation results of the present invention;
[0044] Figure 7 FIG. 12 is a comparison chart of the communication overhead of SkipSMA and existing training methods in the implementation results of the present invention;
[0045] Figure 8 FIG. 13 is a comparison chart of the statistical efficiency of SkipSMA and existing training methods in the implementation results of the present invention;
[0046] Figure 9This is the flowchart of the present invention. Detailed implementation manners
[0047] In combination with the following specific embodiments and the accompanying drawings, the invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and well-known common sense in the art, and the present invention has no particularly restricted content.
[0048] To reduce the communication overhead of training models in a distributed environment, SkipSMA adopts a skip communication strategy to update the global model, that is, network communication is performed only every several rounds of iteration. Because if the global model is updated through network communication in each round of iteration like SMA, it will lead to high network communication overhead in a distributed environment, resulting in long training time. Therefore, to reduce the network communication overhead, SkipSMA updates the global model only every τ rounds of iteration, thus reducing the frequency of network communication. Among them, the communication interval τ is a key factor affecting the communication overhead. However, the same communication interval τ often causes different communication overheads in different environments (including GPU computing power, network bandwidth, and model size, etc.). Therefore, it is necessary to adopt an adaptive strategy to select the communication interval τ. SkipSMA adaptively adjusts the communication interval τ by collecting the required communication and computing time data during runtime. Specifically, during system initialization, SkipSMA sets the communication interval τ to 1, and then during system runtime, SkipSMA collects the communication time consumption t cm and the computing time consumption t cp . According to the collected t cm and t cp , the adaptive communication interval selector adjusts the communication interval τ to This makes the communication time and computing time in each epoch similar. That is to say, the adaptive communication interval makes the communication overhead no longer become the bottleneck in the entire training process.
[0049] According to the communication interval τ calculated by the adaptive communication interval selector in the first round of iteration, SkipSMA updates the local model using a correction technique in each subsequent round of iteration, and updates the global model every τ rounds of iteration using the skip communication strategy. Specifically, in the iterations where the remainder of the iteration number divided by the communication interval τ is not 0, SkipSMA only updates the local model using the correction technique without involving network communication to update the global model; in the iterations where the remainder of the iteration number divided by the communication interval τ is 0, SkipSMA not only updates the local model but also updates the global model through network communication.
[0050] In each iteration, each training process calculates the gradient and correction based on the local model and the global model, and simultaneously uses the gradient and correction to update the local model. As Figure 5 shown, in the 10th iteration, a training process calculates the gradient based on the local model and combines it with the global model to calculate the correction Then, according to the calculated gradient and correction, the local model is updated from to The correction here represents the penalty when the local model deviates from the global model. Similarly, another training process calculates the gradient based on the local model and combines it with the global model to calculate the correction Then, according to the calculated gradient and correction, the local model is updated from to The same applies to the update of the local model in subsequent iterations. to
[0051] Every τ iterations, the global model is updated once among all training processes through network communication. Figure 5 Taking the communication interval τ = 3 as an example, the process of SkipSMA using the skip communication strategy to update the global model is described. In the 10th and 11th iterations, SkipSMA does not update the global model, that is, SkipSMA skips communication in the 10th and 11th iterations. In the 12th iteration, since the remainder of 12 divided by the communication interval 3 is 0, the global model needs to be updated through network communication in the 12th iteration. Specifically, SkipSMA aggregates the corrections calculated in all training processes through network communication (i.e., and ) and updates the global model from to
[0052] The above is the specific implementation process of the training model method SkipSMA with low communication overhead and high statistical efficiency in a distributed deep learning system. In a distributed deep learning system, this method can be implemented through the relevant code in Method 1. The code of Method 1 is as follows:
[0053]
[0054]
[0055] Embodiment
[0056] Use four machines with CentOS 7 operating system to form a cluster for training models. The machines are connected by a network with a bandwidth of 1000 Mbps. Each machine has two GPUs of the Tesla V100 model. In this cluster, five methods, namely SSGD, SkipSSGD, LocalSGD, SMA, and SkipSMA, are used to train the ResNet50 model, and the total training time, communication overhead, and statistical efficiency of training the ResNet50 model to 68% accuracy are compared among these five methods. It should be noted that SSGD, SkipSSGD, LocalSGD, and SkipSMA all run on the four machines in the cluster, while SMA mainly focuses on the scenario of multiple GPUs on a single machine, so SMA only runs on one machine in the cluster. In addition, the communication interval τ of SkipSMA is determined by the adaptive communication interval selector during operation, while the communication intervals τ of SkipSSGD and LocalSGD are specified as 1, 2, 3, 4, 5, 10, 20, 30, and 40 before operation, and the communication interval τ that results in the shortest training time is selected from them to compare with SkipSMA.
[0057] As Figure 6 shown, compared with SSGD, SkipSMA shortens the total training time by 86.76%. This is mainly because SkipSMA has lower communication overhead than SSGD. Specifically, SSGD involves network communication in each round of iteration, while SkipSMA communicates only once every τ rounds of iteration. As Figure 7 shown, this enables SkipSMA to reduce the communication overhead by 93.3% compared with SSGD.
[0058] As Figure 6 shown, compared with SkipSSGD, SkipSMA shortens the total training time by 64.83%. This is because when the communication interval τ > 3 for SkipSSGD, affected by the batch size, it cannot train the model to 68% accuracy within 50 hours; when the communication interval τ = 3, SkipSSGD obtains the shortest total training time. However, when the communication interval τ = 3 for SkipSSGD, the communication overhead is still relatively large. As Figure 7 shown, SkipSMA with an adaptive communication interval reduces the communication overhead by 80% compared with SkipSSGD with a communication interval τ = 3.
[0059] As Figure 6As shown, compared with LocalSGD, SkipSMA shortens the total training time by 16.37%. LocalSGD has the shortest total training time when the communication interval τ = 10. When the communication interval τ > 10, although LocalSGD can further reduce the communication overhead, it will cause the local models to diverge more, resulting in a decrease in statistical efficiency. However, due to the adoption of the correction technique, SkipSMA effectively alleviates the divergence between local models and thus obtains high statistical efficiency. As Figure 8 shown, SkipSMA with an adaptive communication interval maintains a statistical efficiency similar to that of LocalSGD with a communication interval τ = 10. At the same time, SkipSMA adaptively adjusts the communication interval τ to 15, enabling SkipSMA to have a low communication overhead. As Figure 7 shown, compared with LocalSGD with a communication interval τ = 10, SkipSMA with an adaptive communication interval reduces the communication overhead by 33.3%.
[0060] As Figure 6 shown, compared with SMA, SkipSMA shortens the total training time by 79.26%. Although the communication time in each iteration of SMA is similar to the computing time, that is, communication does not become a bottleneck, SMA only utilizes the computing power of a single node. SkipSMA not only maintains a low communication overhead through the skip communication strategy with an adaptive communication interval but also utilizes the computing power of multiple nodes. As Figure 7 shown, in each epoch, SkipSMA reduces the computing and communication time by 79.9% compared with SMA.
[0061] The protection scope of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept, changes and advantages that can be conceived by those skilled in the art are included in the present invention, and the scope of protection is defined by the appended claims.
Claims
1. A method for training a model with low communication overhead and high statistical efficiency in a distributed deep learning system, characterized in that, the method comprises the following steps: Step A: The runtime data collector collects the communication time t required for the adaptive communication interval and the computing time t during runtime. cm And the computing time t cp data, and the adaptive communication interval selector automatically adjusts the communication interval τ based on the communication time and computing time data obtained through the above collection. The step A further comprises the following steps: Step A1: The system automatically initializes the communication interval τ to 1; sets the communication interval τ = 1, and sets the adaptive communication interval flag bit flag; Step A2: When the system is running, the runtime data collector collects the time consumption t of communication and the time consumption t of computation in the first round of iteration cm during the first round of iteration cp ; Step A3: The adaptive communication interval selector adjusts the communication interval based on t cm and t cp collected in the first round of iteration Step B: Iteratively train the model, and update the local model using a correction technique in each iteration; The step B further comprises the following steps: Step B1: Each training process calculates the gradient according to the local model; Step B2: Each training process calculates the correction, where the correction refers to the difference between the local model and the global model; Step B3: Each training process updates the local model with the calculated gradient and correction; Step C: Update the global model using the skip communication strategy every τ rounds of iteration; The step C further comprises the following steps: Step C1: Take the remainder of the current iteration number with respect to the adaptively set communication interval τ. If the remainder is 0, it indicates that the global model needs to be updated in the current iteration, and then proceed to step C2; Step C2: Aggregate the corrections calculated by each training process through a network communication, and update the global model with the aggregated correction; In step C2, the aggregated correction is represented by the following formula: where n represents the total number of training processes, k represents the number of iterations, represents the correction calculated by the training process numbered i in the k-th iteration; The update form of the global model is as shown in the formula where i represents the number of the training process, k represents the number of iterations, and represent the global model parameters in the k-th and (k + 1)-th iterations respectively, represents the global model momentum in the k-th iteration, μ represents the momentum coefficient, and α represents the correction coefficient; The global model momentum is initially empty and is updated every τ rounds of iteration according to the aggregated correction. The specific update form is as follows: where i represents the number of the training process, k represents the number of iterations, and represent the global model momentum in the k-th and (k + 1)-th rounds of iteration, represents the aggregated correction, and μ and α represent the momentum coefficient and the correction coefficient respectively.
2. The method according to claim 1, characterized in that, In step A1, the adaptive communication interval flag bit is used to indicate whether to enable the adaptive communication interval selector; when the adaptive communication interval flag bit flag = true, the adaptive communication interval selector is enabled, and the system adaptively selects a suitable communication interval τ; when the adaptive communication interval flag bit flag = false, the adaptive communication interval selector is disabled, and the user needs to specify the communication interval τ.
3. The method according to claim 1, characterized in that, After the communication interval τ is adjusted in the first round of iteration, the communication interval τ is used for all subsequent iterations.
4. The method according to claim 1, characterized in that, In step B1, the gradient is calculated by the following formula: where, i represents the number of the training process, k represents the number of iterations, represents the gradient calculated in the k-th iteration of the training process numbered i, b (i) represents the batch size on the training process numbered i, represents the local model parameters in the k-th iteration of the training process numbered i, and L(x, w) represents the loss calculated for the sample x on the model parameters w.
5. The method according to claim 1, characterized in that, In step B2, the correction is calculated by the following formula: where \(i\) represents the number of the training process, \(k\) represents the number of iterations, represents the correction calculated in the \(k\)-th iteration of the training process numbered \(i\), represents the local model parameters of the training process numbered \(i\) in the \(k\)-th iteration, represents the global model parameters in the \(k\)-th iteration.
6. The method according to claim 1, characterized in that, In step B3, the specific update form of the local model is as shown in the following formula: Wherein, i represents the number of the training process, and k represents the number of iterations. and represent the local model parameters of the training process numbered i in the k-th and (k + 1)-th iterations. and represent the local model momentum, gradient, and correction of the training process numbered i in the k-th iteration respectively. μ, γ, and α represent the momentum coefficient, learning rate, and correction coefficient respectively.
7. The method according to claim 6, characterized in that, The local model momentum is initially empty and is updated according to the gradient in each iteration. The specific update form is: where \(i\) represents the number of the training process, and \(k\) represents the number of iterations. and represent the local model momentum of the training process numbered \(i\) in the \(k\)-th and \((k + 1)\)-th iterations. represents the gradient of the training process numbered \(i\) in the \(k\)-th iteration, and \(\mu\) and \(\gamma\) represent the momentum coefficient and the learning rate respectively.
8. A system for implementing the method according to any one of claims 1-7, characterized in that, the system comprises: a runtime data collector, an adaptive communication interval selector, and a model update module.
9. The system according to claim 8, characterized in that, The runtime data collector collects the communication time data t of the system during runtime in the first round of iteration cm and the computing time data t cp ; The adaptive communication interval selector automatically adjusts the communication interval according to the communication time t collected by the runtime data collector cm and the calculation time t cp , and uses this communication interval τ for all subsequent iterations; The model update module iteratively updates the local model and the global model according to the communication interval τ obtained by the adaptive communication interval selection module, that is, updates the local model using a correction technique in each iteration, and updates the global model using the skip communication strategy every τ rounds of iteration.