A cross-domain data parallel training scheduling method for AI
By establishing a hierarchical communication network between cross-domain multi-data centers, using wide area networks and high-speed LANs for parameter synchronization, and using stochastic gradient descent algorithms and parameter server architecture during the training process, the problems of low parameter synchronization efficiency and poor model convergence in cross-domain scenarios are solved, and efficient cross-domain training is achieved.
Patent Information
- Application Number
- CN202211670584.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-12-25
AI Technical Summary
The prior art is difficult to achieve efficient neural network model parameter synchronization in cross-domain scenarios, resulting in frequent cross-domain communication and poor model convergence.
By establishing a hierarchical communication network between cross-domain multi-data centers, using wide area networks and high-speed LANs for parameter synchronization, and using stochastic gradient descent algorithms and parameter server architecture during the training process, the hierarchical synchronization and update of cross-domain neural network model parameters are achieved.
The frequency of inter-domain communication is reduced, the convergence of neural network models is improved, and the cross-domain training time and accuracy of AI neural network models are shortened.
Smart Images

Figure CN116090528B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computers, and in particular, relates to a cross-domain data parallel training scheduling method for AI. Background Art
[0002] In AI neural network model training, valuable datasets are usually owned by independent organizations and located in multiple data centers. Most AI-oriented neural network model training techniques require centralizing data from multiple data centers to improve performance. However, in practice, due to privacy regulations, it is often not feasible to transfer all data from different organizations to a centralized data center. It is very challenging to perform data-parallel training of cross-regionally distributed neural network models between data centers without privacy leakage.
[0003] The existing data parallel training methods for cross-region distributed neural network models mainly face two technical problems. On the one hand, there is a large bandwidth difference between the operator network used for cross-domain communication and the high-speed local area network inside the data center, resulting in the cross-domain synchronization efficiency of the neural network model parameters being much lower than the intra-domain synchronization efficiency. Frequent cross-domain parameter synchronization increases communication time. On the other hand, some existing technologies improve training efficiency through cross-domain asynchronous training, but the asynchronous method leads to poor convergence of AI neural network models. Summary of the invention
[0004] The present invention aims to overcome the deficiencies of the prior art and provide a cross-domain data parallel training scheduling method for AI.
[0005] In order to achieve the above object, the technical solution provided by the present invention is:
[0006] The cross-domain data parallel training scheduling method for AI comprises the following steps:
[0007] (1) Establish a wide area network between multiple data centers across domains, so that the parameter servers in each data center can communicate with the global parameter server under the wide area network; establish a high-speed local area network within a single data center, so that the working nodes in the data center can communicate under the high-speed local area network;
[0008] (2) Based on the communication network built in step (1), each data center uses the stochastic gradient descent algorithm to train the local neural network model on the local data set. After reaching the maximum number of iterations, the neural network model parameters of each data center are aggregated to obtain the global neural network model parameters.
[0009] Preferably, step (2) specifically includes the following steps:
[0010] (1) Set the epoch counter j = 0. The variable j is used to record the number of training rounds, and each round of training contains M iterations. That is, when the number of iterations reaches M, the training enters the next round.
[0011] (2) Set k = 5. The constant k means that each data center performs cross-domain synchronization of neural network model parameters once every k iterations.
[0012] (3) Set the iteration counter i = 0. The variable i is used to record the number of training iterations.
[0013] (4) Within each data center, a stochastic gradient descent algorithm is used on the local data set held by each data center to perform a round of forward and backward propagation calculations, and a parameter server architecture is used within each data center to synchronize the neural network model parameters of this round of iterations under a high-speed local area network;
[0014] (5) Determine whether the remainder of i modulo k is 0. If so, globally synchronize the neural network model parameters obtained in step (4) using a parameter server architecture between data centers and a wide area network, and then proceed to the next step. Otherwise, proceed directly to the next step.
[0015] (6) Determine whether the remainder of i modulo M is 0. If so, set j = j + 1 and proceed to the next step; otherwise, set i = i + 1 and proceed to step (4).
[0016] (7) Determine whether j <= N. If so, proceed to step (3). Otherwise, for the neural network model parameters obtained in step (4), the parameter server architecture is used between the data centers and the wide area network is used to perform global synchronization of the neural network model parameters to obtain the final global neural network model parameters.
[0017] In the present invention, first, a hierarchical communication network is established for cross-domain multi-data centers, a high-speed local area network is established inside the data center, and a wide area network is established between data centers; then, in the cross-domain data parallel training process of the AI neural network model, based on the hierarchical communication network and the stochastic gradient descent algorithm, the neural network model parameters are hierarchically synchronized and updated. Compared with the prior art, the main advantages of the present invention are: a cross-domain data parallel training scheduling method for AI is proposed, which is based on a cross-domain hierarchical communication network and parameter synchronization method, which can reduce the frequency of inter-domain communication and improve the convergence of the neural network model. The present invention can solve the problems of frequent cross-domain communication and difficult model convergence in the existing AI neural network model training technology in cross-domain scenarios, so that the cross-domain training time of the AI neural network model is shortened and the accuracy is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1It is a cross-domain training flowchart for AI.
[0019] Figure 2 It is a cross-domain data parallel training scheduling diagram. DETAILED DESCRIPTION
[0020] See also Figure 1 and Figure 2 , the cross-domain data parallel training scheduling method for AI comprises the following steps:
[0021] (1) Establish a wide area network between multiple data centers across domains, so that the parameter servers in each data center can communicate with the global parameter server under the wide area network; establish a high-speed local area network within a single data center, so that the working nodes in the data center can communicate under the high-speed local area network;
[0022] (2) Based on the communication network built in step (1), each data center uses the stochastic gradient descent algorithm to train the local neural network model on the local data set. After reaching the maximum number of iterations, the neural network model parameters of each data center are aggregated to obtain the global neural network model parameters.
[0023] Wherein, step (2) specifically includes the following steps:
[0024] (1) Set the epoch counter j = 0. The variable j is used to record the number of training rounds, and each round of training contains M iterations. That is, when the number of iterations reaches M, the training enters the next round.
[0025] (2) Set k = 5. The constant k means that each data center performs cross-domain synchronization of neural network model parameters once every k iterations.
[0026] (3) Set the iteration counter i = 0. The variable i is used to record the number of training iterations.
[0027] (4) Within each data center, a stochastic gradient descent algorithm is used on the local data set held by each data center to perform a round of forward and backward propagation calculations, and a parameter server architecture is used within each data center to synchronize the neural network model parameters of this round of iterations under a high-speed local area network;
[0028] (5) Determine whether the remainder of i modulo k is 0. If so, globally synchronize the neural network model parameters obtained in step (4) using a parameter server architecture between data centers and a wide area network, and then proceed to the next step. Otherwise, proceed directly to the next step.
[0029] (6) Determine whether the remainder of i modulo M is 0. If so, set j = j + 1 and proceed to the next step; otherwise, set i = i + 1 and proceed to step (4).
[0030] (7) Determine whether j <= N. If so, proceed to step (3). Otherwise, for the neural network model parameters obtained in step (4), the parameter server architecture is used between the data centers and the wide area network is used to perform global synchronization of the neural network model parameters to obtain the final global neural network model parameters.
Claims
1. A cross-domain data parallel training scheduling method for AI, characterized in that: The method comprises the following steps: (1) Establish a wide area network between multiple data centers across domains, so that the parameter servers in each data center can communicate with the global parameter server under the wide area network; establish a high-speed local area network within a single data center, so that the working nodes in the data center can communicate under the high-speed local area network; (2) Based on the communication network built in step (1), each data center uses the stochastic gradient descent algorithm to train the local neural network model on the local data set. After reaching the maximum number of iterations, the neural network model parameters of each data center are aggregated to obtain the global neural network model parameters; Step (2) specifically includes the following steps: (1) Set the epoch counter j = 0. The variable j is used to record the number of training rounds, and each round of training contains M iterations. That is, when the number of iterations reaches M, the training enters the next round. (2) Set k = 5. The constant k indicates that each data center performs cross-domain synchronization of neural network model parameters once every k iterations. (3) Set the iteration counter i = 0. The variable i is used to record the number of training iterations. (4) Within each data center, a stochastic gradient descent algorithm is used on the local data set held by each data center to perform a round of forward and backward propagation calculations, and a parameter server architecture is used within each data center to synchronize the neural network model parameters of this round of iterations under a high-speed local area network; (5) Determine whether the remainder of i modulo k is 0. If so, globally synchronize the neural network model parameters obtained in step (4) using a parameter server architecture between data centers and a wide area network, and then proceed to the next step; Otherwise, go directly to the next step; (6) Determine whether the remainder of i modulo M is 0. If so, set j = j + 1 and proceed to the next step; Otherwise, set i=i+1 and go to step (4); (7) Determine whether j <= N. If so, proceed to step (3). Otherwise, for the neural network model parameters obtained in step (4), the parameter server architecture is used between the data centers and the wide area network is used to perform global synchronization of the neural network model parameters to obtain the final global neural network model parameters.
Citation Information
Patent Citations
Method and system for training neural network model based on federal learning mode
CN112929223A
Network topology construction method and system in hierarchical federated learning scene
CN114650227A