A cross-domain data parallel training scheduling method for AI

By establishing a hierarchical communication network between cross-domain multi-data centers, using wide area networks and high-speed LANs for parameter synchronization, and using stochastic gradient descent algorithms and parameter server architecture during the training process, the problems of low parameter synchronization efficiency and poor model convergence in cross-domain scenarios are solved, and efficient cross-domain training is achieved.

CN116090528BActive Publication Date: 2025-05-13HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211670584.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-25
Publication Date
2025-05-13
Estimated Expiration
2042-12-25

AI Technical Summary

Technical Problem

The prior art is difficult to achieve efficient neural network model parameter synchronization in cross-domain scenarios, resulting in frequent cross-domain communication and poor model convergence.

Method used

By establishing a hierarchical communication network between cross-domain multi-data centers, using wide area networks and high-speed LANs for parameter synchronization, and using stochastic gradient descent algorithms and parameter server architecture during the training process, the hierarchical synchronization and update of cross-domain neural network model parameters are achieved.

Benefits of technology

The frequency of inter-domain communication is reduced, the convergence of neural network models is improved, and the cross-domain training time and accuracy of AI neural network models are shortened.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116090528B_ABST
    Figure CN116090528B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-domain data parallel training scheduling method for AI. The method comprises the following steps: (1) establishing a wide area network between cross-domain multi-data centers so that the parameter servers of each data center can communicate with the global parameter server under the wide area network; establishing a high-speed local area network inside a single data center so that the working nodes in the data center can communicate under the high-speed local area network; (2) based on the communication network built in step (1), each data center uses a stochastic gradient descent algorithm on a local data set to train a local neural network model, and after reaching the maximum number of iterations, aggregates the neural network model parameters of each data center to obtain the global neural network model parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computers, and in particular, relates to a cross-domain data parallel training scheduling method for AI. Background Art

[0002] In AI neural network model training, valuable datasets are usually owned by independent organizations and located in multiple data centers. Most AI-oriented neural network model training techniques require centralizing data from multiple data centers to improve performance. However, in practice, due to privacy regulations, it is often not feasible to transfer all data from different organizations to a centralized data center. It is very challenging to perform data-parallel training of cross-regionally distributed neural network models between data centers without privacy leakage.

[0003] The existing data parallel training methods for cross-region distributed neural network models mainly face two technical problems. On the one hand, there is a large bandwidth difference between the operator network used for cross-domain communication and the high-speed local area network inside the data center, resulting in the cross-domain synchronization efficiency of the neural network model parameters being much lower than the intra-domain synchronization efficiency. Frequent cross-domain parameter synchronization increases communication time. On the other hand, some existing technologies improve training efficiency through cross-domain asynchronous training, but the asynchronous method leads to poor convergence of AI neural network models. Summary of the invention

[0004] The present invention aims to overcome the deficiencies of the prior art and provide a cross-domain data parallel training scheduling method for AI.

[0005] In order to achieve the above object, the technical solution provided by the present invention is:

[0006] The cross-domain data parallel training scheduling method for AI comprises the following steps:

[0007] (1) Establish a wide area network between multiple data centers across domains, so that the parameter servers in each data center can communicate with the global parameter server under the wide area network; establish a high-speed local area network within a single data center, so that the working nodes in the data center can communicate under the high-speed local area network;

[0008] (2) Based on the communication network built in step (1), each data center uses the stochastic gradient descent algorithm to train the local neural network model on the local data set. After reaching the maximum number of iterations, the neural network model parameters of each data center are aggregated to obtain the global neural network model parameters.

[0009] Preferably, step (2) specifically includes the following steps:

[0010] (1) Set the epoch counter j = 0. The variable j is used to record the number of training rounds, and each round of training contains M iterations. That is, when the number of iterations reaches M, the training enters the next round.

[0011] (2) Set k = 5. The constant k means that each data center performs cross-domain synchronization of neural network model parameters once every k iterations.

[0012] (3) Set the iteration counter i = 0. The variable i is used to record the number of training iterations.

[0013] (4) Within each data center, a stochastic gradient descent algorithm is used on the local data set held by each data center to perform a round of forward and backward propagation calculations, and a parameter server architecture is used within each data center to synchronize the neural network model parameters of this round of iterations under a high-speed local area network;

[0014] (5) Determine whether the remainder of i modulo k is 0. If so, globally synchronize the neural network model parameters obtained in step (4) using a parameter server architecture between data centers and a wide area network, and then proceed to the next step. Otherwise, proceed directly to the next step.

[0015] (6) Determine whether the remainder of i modulo M is 0. If so, set j = j + 1 and proceed to the next step; otherwise, set i = i + 1 and proceed to step (4).

[0016] (7) Determine whether j <= N. If so, proceed to step (3). Otherwise, for the neural network model parameters obtained in step (4), the parameter server architecture is used between the data centers and the wide area network is used to perform global synchronization of the neural network model parameters to obtain the final global neural network model parameters.

[0017] In the present invention, first, a hierarchical communication network is established for cross-domain multi-data centers, a high-speed local area network is established inside the data center, and a wide area network is established between data centers; then, in the cross-domain data parallel training process of the AI ​​neural network model, based on the hierarchical communication network and the stochastic gradient descent algorithm, the neural network model parameters are hierarchically synchronized and updated. Compared with the prior art, the main advantages of the present invention are: a cross-domain data parallel training scheduling method for AI is proposed, which is based on a cross-domain hierarchical communication network and parameter synchronization method, which can reduce the frequency of inter-domain communication and improve the convergence of the neural network model. The present invention can solve the problems of frequent cross-domain communication and difficult model convergence in the existing AI neural network model training technology in cross-domain scenarios, so that the cross-domain training time of the AI ​​neural network model is shortened and the accuracy is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1It is a cross-domain training flowchart for AI.

[0019] Figure 2 It is a cross-domain data parallel training scheduling diagram. DETAILED DESCRIPTION

[0020] See also Figure 1 and Figure 2 , the cross-domain data parallel training scheduling method for AI comprises the following steps:

[0021] (1) Establish a wide area network between multiple data centers across domains, so that the parameter servers in each data center can communicate with the global parameter server under the wide area network; establish a high-speed local area network within a single data center, so that the working nodes in the data center can communicate under the high-speed local area network;

[0022] (2) Based on the communication network built in step (1), each data center uses the stochastic gradient descent algorithm to train the local neural network model on the local data set. After reaching the maximum number of iterations, the neural network model parameters of each data center are aggregated to obtain the global neural network model parameters.

[0023] Wherein, step (2) specifically includes the following steps:

[0024] (1) Set the epoch counter j = 0. The variable j is used to record the number of training rounds, and each round of training contains M iterations. That is, when the number of iterations reaches M, the training enters the next round.

[0025] (2) Set k = 5. The constant k means that each data center performs cross-domain synchronization of neural network model parameters once every k iterations.

[0026] (3) Set the iteration counter i = 0. The variable i is used to record the number of training iterations.

[0027] (4) Within each data center, a stochastic gradient descent algorithm is used on the local data set held by each data center to perform a round of forward and backward propagation calculations, and a parameter server architecture is used within each data center to synchronize the neural network model parameters of this round of iterations under a high-speed local area network;

[0028] (5) Determine whether the remainder of i modulo k is 0. If so, globally synchronize the neural network model parameters obtained in step (4) using a parameter server architecture between data centers and a wide area network, and then proceed to the next step. Otherwise, proceed directly to the next step.

[0029] (6) Determine whether the remainder of i modulo M is 0. If so, set j = j + 1 and proceed to the next step; otherwise, set i = i + 1 and proceed to step (4).

[0030] (7) Determine whether j <= N. If so, proceed to step (3). Otherwise, for the neural network model parameters obtained in step (4), the parameter server architecture is used between the data centers and the wide area network is used to perform global synchronization of the neural network model parameters to obtain the final global neural network model parameters.

Claims

1. A cross-domain data parallel training scheduling method for AI, characterized in that: The method comprises the following steps: (1) Establish a wide area network between multiple data centers across domains, so that the parameter servers in each data center can communicate with the global parameter server under the wide area network; establish a high-speed local area network within a single data center, so that the working nodes in the data center can communicate under the high-speed local area network; (2) Based on the communication network built in step (1), each data center uses the stochastic gradient descent algorithm to train the local neural network model on the local data set. After reaching the maximum number of iterations, the neural network model parameters of each data center are aggregated to obtain the global neural network model parameters; Step (2) specifically includes the following steps: (1) Set the epoch counter j = 0. The variable j is used to record the number of training rounds, and each round of training contains M iterations. That is, when the number of iterations reaches M, the training enters the next round. (2) Set k = 5. The constant k indicates that each data center performs cross-domain synchronization of neural network model parameters once every k iterations. (3) Set the iteration counter i = 0. The variable i is used to record the number of training iterations. (4) Within each data center, a stochastic gradient descent algorithm is used on the local data set held by each data center to perform a round of forward and backward propagation calculations, and a parameter server architecture is used within each data center to synchronize the neural network model parameters of this round of iterations under a high-speed local area network; (5) Determine whether the remainder of i modulo k is 0. If so, globally synchronize the neural network model parameters obtained in step (4) using a parameter server architecture between data centers and a wide area network, and then proceed to the next step; Otherwise, go directly to the next step; (6) Determine whether the remainder of i modulo M is 0. If so, set j = j + 1 and proceed to the next step; Otherwise, set i=i+1 and go to step (4); (7) Determine whether j <= N. If so, proceed to step (3). Otherwise, for the neural network model parameters obtained in step (4), the parameter server architecture is used between the data centers and the wide area network is used to perform global synchronization of the neural network model parameters to obtain the final global neural network model parameters.

Citation Information

Patent Citations

  • Method and system for training neural network model based on federal learning mode

    CN112929223A

  • Network topology construction method and system in hierarchical federated learning scene

    CN114650227A