A Neural Network Adaptive Distributed Parallel Training Method Based on Genetic Algorithm

By transforming the operator placement problem of neural network model into integer linear programming problem, and using the dual population genetic algorithm and cost evaluation model, the training problem of large-scale complex neural networks under the constraints of hardware equipment resources is solved, and efficient model parallel training is achieved.

CN115115052BActive Publication Date: 2025-05-27HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210963656.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-11
Publication Date
2025-05-27
Estimated Expiration
2042-08-11

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively solve the training problem of large-scale complex neural networks when hardware device resources are limited, resulting in the model parallel training strategy relies on expert experience, low cluster utilization rate and large communication overhead.

Method used

The operator placement problem in the neural network model is transformed into integer linear planning problem, and the solution space is simplified by constructing the solution space of the data flow graph and the device topology graph, combining hypothesis constraints and operator grouping constraints; using a two-population genetic algorithm and cost evaluation model, quickly searching and implementing the lowest cost placement strategy.

Benefits of technology

The training speed of neural network models is accelerated, the efficiency of parallel training of models is improved, the communication overhead is reduced, and the utilization rate of clusters is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115052B_ABST
    Figure CN115115052B_ABST
Patent Text Reader

Abstract

The present invention discloses a neural network adaptive distributed parallel training method based on a genetic algorithm. First, this method transforms the operator placement problem in model parallelism into an integer linear programming problem, constructs a solution space, and simplifies the scale of the solution space. Secondly, a cost evaluation model is constructed to evaluate the quality of the solutions to the linear programming problem and guide the iterative update of two sets of placement strategies between two populations. Then, when the termination condition of the iterative update is reached, the two sets of placement strategies are transformed into operator placement strategies for the data flow graph, and the operator placement strategies are used to train the model to select the optimal distributed scheduling strategy. Finally, using the obtained distributed scheduling strategy, the model to be trained is distributedly scheduled, and the data set is input into the neural network for distributed training. The present invention speeds up the algorithm search speed on the premise of ensuring the quality of the solution space, evaluates the execution performance of the operator placement strategy in a fine-grained manner, and effectively avoids falling into the local optimal situation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of distributed machine learning, and mainly relates to a neural network adaptive distributed parallel training method based on a genetic algorithm, providing an optimal model parallel training method for large-scale complex neural networks in artificial intelligence scenarios such as image recognition and natural language processing. Background Art

[0002] In recent years, with the increase in data scale, deep learning models have become increasingly complex, and the parameter scale of deep learning models has grown from the ten-thousand level to the billion level and trillion level. Due to their powerful capabilities of massive data mining and analysis, pre-training, and fast inference, large deep learning models have gradually developed into the next-generation AI technology. However, the rapid growth of the dataset scale and the number of model parameters has restricted the training of AI models by hardware devices. The computing power and storage capacity of a single device are limited and cannot cope with training such large models on large datasets. For example, Bert-Large has 400 million parameters and occupies more than 32GB of memory; GPT-3 has 175 billion parameters and occupies more than 350GB of memory. The ordinary implementation of a 1000-layer residual network requires 48GB of memory, and even after further optimization to reduce memory costs, the DNN model still requires at least 7GB of memory. This makes it extremely difficult to run the entire model on a single device with limited memory and perform data parallel training. At the same time, due to the increase in the scale of deep learning models, the required data and computing volume are also increasing, and the computing power of a single device often cannot meet the requirements. Therefore, using cross-device distributed parallel technology to train large deep learning models has become a key requirement for the development and application of deep learning in the industrial field.

[0003] Current mainstream distributed machine learning systems such as TensorFlow, PyTorch, and MindSpore mainly use the data parallel (Data Parallel) method to accelerate the model training speed. However, data parallel usually can only solve the problem of training with massive data and cannot solve the problem that large models cannot be trained due to hardware device resource limitations. Therefore, in order to meet the training requirements of large-scale network models, it is necessary to partition the neural network and schedule it across devices to achieve model parallel (Model Parallel). In terms of formulating the model parallel strategy, the neural network computation graph is usually horizontally partitioned by layer, vertically partitioned across layers, or randomly partitioned and scheduled to different devices for execution. However, these methods rely heavily on expert experience, the partitioning method is unreasonable, the cluster utilization rate is low, and there is usually a large communication overhead.

[0004] The automatic search method for parallel strategies based on machine learning extracts the features of neural network models and device topologies, and uses machine learning modeling to guide the search for parallel strategies to obtain the optimal distributed parallel strategy. The Google team proposed the automatic strategy search frameworks ColorRL and Hierarchical, which extract neural network model and device features and use reinforcement learning to guide the search for parallel strategies. However, the above methods require frequent sampling, making the strategy search process costly. Gao et al. proposed Spotlight, which first modeled the neural network operator scheduling problem as a Markov decision process (MDP). Subsequently, the proposed POST further improved it by introducing cross-entropy into the sampling process to improve the search efficiency. Ravichandra of MIT proposed Placeto, which introduced the graph embedding method (GraphEmbedding) to enable the neural network to learn the structural features of the model and equip it with the ability to quickly transfer strategies. Hao et al. proposed the AutoSync adaptive framework. Based on the data parallel scenario, it uses machine learning methods to predict the execution time of data parallel strategies, thereby guiding the search for synchronous parallel strategies. Ding, Zhou et al. designed a full-scale search method based on reinforcement learning that can generate the entire graph placement strategy at one time and support the rapid transfer of strategies between models with similar structures. Lan et al. captured the structural information of the neural network by establishing a self-supervised graph neural network and used the Seq2Seq (Sequence-to-Sequence) network to segment and schedule deep learning models. Although the above machine learning-based methods can search for parallel strategies with excellent execution performance, they are usually limited to small and medium-sized neural network models and rely on hardware resources. The search process often takes dozens of hours and is not suitable for scenarios where models need to be quickly deployed or trained.

[0005] The strategy automatic search method based on graph algorithms is another current mainstream method, which has a faster search speed compared to the parallel strategy automatic search method based on machine learning. Jia et al. proposed OptCNN by abstracting the parallel strategy as the slicing of tensors and using the idea of dynamic programming to search for the optimal strategy in the entire search space. However, its coarse-grained model partitioning and rough cost evaluation model limit the performance improvement of the searched strategy. Subsequently, Jia et al. improved and proposed the FlexFlow framework based on OptCNN, introduced the idea of SOAP (Sample, Operator, Attribute, Parameter) to establish a high-dimensional search space, and used the Markov decision process to quickly search for the optimal solution in the search space. However, its implementation is complex and only applicable to a small number of neural network models, and the cost evaluation does not consider the case of parallel execution of operators. In order to search for strategies faster and apply to more models, Yi et al. proposed the FastT algorithm based on DAG scheduling, which places and schedules operators by setting fine-grained operator priorities and using the critical path. However, it does not consider the dynamic memory occupancy during device training and is not applicable to dynamic RNN. Beomyeol et al. proposed the Beachi framework, which searches for strategies based on three graph algorithms: topological sorting, earliest start time, and minimum communication volume. It can search for the model parallel strategy in dozens of seconds at the fastest and supports most models. However, its reliance on specific conditions results in different strategy effects on different neural network models. Although the above methods have a relatively fast search speed, they usually rely on specified search ideas, consider the search direction of strategies from an incomplete dimension, and lack a comprehensive cost evaluation model, making it easy for the methods to search for suboptimal solutions. Summary of the Invention

[0006] In order to solve the above deficiencies and accelerate the training speed of AI models, the present invention designs and implements a neural network adaptive distributed parallel training method based on genetic algorithms.

[0007] A neural network adaptive distributed parallel training method based on genetic algorithms, the steps can be generally summarized as follows:

[0008] Step 1: Transform the operator placement problem in the neural network model into an integer linear programming problem, construct a solution space based on the data flow graph and device topology graph of the neural network model, and simplify the scale of the solution space by setting hypothesis constraints and operator grouping constraints.

[0009] 1-1. Extract the neural network model structure and form a device resource group, and abstract to obtain a computation graph G(O, E) and a device topology graph D. In the computation graph G(O, E), the vertex O represents the set of operators of the neural network model, E represents the set of directed edges between vertices, and the device topology graph D is composed of M computing devices.

[0010] As shown in FIGS. 1-2, given a computation graph G(O, E) and a device topology graph D, on the premise of satisfying the device memory constraint, find a set of placement strategies to minimize the execution cost, thereby further transforming the problem of placing the operators of the neural network model on the device into an integer linear programming problem, and constructing a solution space based on the data flow graph and the device topology graph.

[0011] 1-3, Two assumptions are proposed by analyzing the device characteristics in the homogeneous cluster and the execution characteristics of the neural network model: the computing and communication performance of homogeneous devices are the same; the execution end time of an operator is only related to the operators with direct dependencies.

[0012] Based on the above assumptions, the placement situations of two operators with data dependencies are constrained to two types. One is that the two operators are placed on the same device, and the other is that the two operators are placed on different devices. Among them, when the two operators are placed on the same device, it indicates that the communication cost of the two operators is greater than the computing cost; when the two operators are placed on different devices, it indicates that the computing cost of the two operators is greater than the communication cost.

[0013] 1-4, Subsequently, an operator grouping constraint is further proposed. Among two operators with data dependencies, if one of the operators satisfies the condition that the out-degree or in-degree is 1, then the two operators are constrained to be in the same group, and the operators in the group are regarded as a whole for scheduling, so as to simplify the structure of the computation graph, thereby reducing the scale of the solution space. The above assumption constraints and operator grouping constraints are used to simplify the search space, so as to speed up the search speed of the algorithm.

[0014] Step 2: Construct a cost evaluation model that fuses computing cost, communication cost, and load cost to comprehensively evaluate the quality of the solution of the integer linear programming problem, that is, the execution performance of the operator placement strategy, and guide the double-population genetic algorithm to perform gene mutation and gene exchange iterations between two populations (placement strategies) to update the two sets of placement strategies.

[0015] 2-1, Construct a cost evaluation model that fuses computing cost, communication cost, and load cost.

[0016] First, analyze the key factors affecting the execution performance of the neural network model, and use the actual execution time of the operators in the neural network model on the device to establish the computing cost L i ; use the total tensor size of cross-device communication between operators in the neural network model to establish the communication cost W; use the sum of the ratios of the memory load offset and the average load of each device to establish the load cost f.

[0017] Second, by comprehensively considering the computing cost, communication cost, load cost, and the characteristics of the actual training process of the model, establish a cost evaluation model that fuses computing cost, communication cost, and load cost

[0018]

[0019] Among them, represents the execution cost of the placement strategy; λ represents the weight ratio, simulating the case of parallel computing and communication; Ω is a linear fitting model of tensor size and communication time.

[0020] 2-2. Iteratively update the placement strategy based on the double-population genetic algorithm.

[0021] First, use the initialization method of the placement strategy to generate two sets of placement strategies with initial operator placement information in the solution space. The initialization method of the placement strategy includes random initialization and specified strategy initialization. The random initialization method means randomly allocating the operators in the computational graph G(O, E) to the device topology graph D. The specified strategy initialization method means using the placement strategy searched by the existing parallel strategy search method as the initial state, and arbitrarily selecting one initialization method to initialize the placement strategy.

[0022] Secondly, based on the cost evaluation model, guide the gene mutation and gene exchange methods of the double-population genetic algorithm, and iteratively update the placement information of the operators in the two sets of placement strategies to quickly search for the placement strategy with the minimum execution cost in the solution space.

[0023] Step 3: When the termination condition for the iterative update of the placement strategy is reached, convert the two sets of placement strategies into the operator placement strategies of the data flow graph respectively, and train the neural network model according to the iterative operator placement strategy in the real environment, and select the strategy with the shortest execution time as the optimal distributed scheduling strategy of the neural network model.

[0024] Among them, the iterative termination conditions include that the total number of iterations reaches the specified threshold and the placement strategy does not change within a certain number of iterations. When the iterative process of the double-population genetic algorithm satisfies one of the above two conditions, the iteration is terminated.

[0025] Step 4: Use the distributed scheduling strategy obtained in Step 3 to perform distributed scheduling on the neural network model to be trained, and input the data sets required by models such as image recognition and natural language processing into the neural network model for distributed training.

[0026] The beneficial effects of the present invention are as follows: transforming the operator placement problem in the neural network model into an integer linear programming problem, intuitively describing the problem to be solved, constructing the solution space of this integer linear programming problem based on the data flow graph and the device topology graph, and then introducing hypothesis constraints and operator grouping constraints to simplify the solution space, accelerating the algorithm search speed while ensuring the quality of the solution space; analyzing the key factors affecting the execution performance of the neural network model, constructing a cost evaluation model to finely evaluate the execution performance of the operator placement strategy; the proposed dual-population genetic algorithm uses gene mutation and gene exchange methods to accelerate the iteration speed of the genetic algorithm, effectively avoiding falling into local optimal situations and helping the algorithm approach the global optimal solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a schematic diagram of searching for the optimal parallel strategy based on the genetic algorithm;

[0028] Figure 2 is a schematic diagram of operator pairs;

[0029] Figure 3 is a schematic diagram of gene mutation in the gene mutation method;

[0030] Figure 4 is a schematic diagram of iterative update in the gene mutation and gene exchange methods;

[0031] Figure 5 is a schematic diagram of cross-exchange of high-quality operator pairs in the gene exchange method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The following will further illustrate the present invention in conjunction with the drawings and specific implementation steps:

[0033] A neural network adaptive distributed parallel training method based on the genetic algorithm, the overall flowchart is as Figure 1 shown, including the following specific steps:

[0034] Step 1: Transform the operator placement problem in the neural network model into an integer linear programming problem, construct a solution space based on the data flow graph and the device topology graph of the neural network model, and simplify the scale of the solution space by setting hypothesis constraints and operator grouping constraint conditions:

[0035] 1-1, According to the neural network data flow graph, define the computational graph G(O, E), where O represents the set of operators, and the node o i ∈O represents a single operator (such as matrix multiplication, convolution, etc.), and E represents the set of directed edges between nodes, indicating the data dependency relationship between operators; according to the cluster device topology information, define the device topology graph D, and the device topology graph D is composed of M computing devices, where the node d ∈ D represents a certain computing device (such as CPU or GPU).

[0036] As shown in FIGS. 1-2, given a computation graph G(O, E) and a device topology graph D, on the premise of satisfying the device memory constraint, find a set of placement strategies to minimize the execution cost, further transform the placement problem between the operators of the neural network model and the devices into an integer linear programming problem, and construct the solution space of this integer linear programming problem by using the computation graph G(O, E) abstracted from the data flow graph of the neural network model and the device topology graph D, as shown in Equation (1):

[0037]

[0038] where P represents the placement strategy of the operator, G represents the computation graph G(O, E), D represents the device topology graph, represents the cost overhead of strategy P under the current computation graph G and device topology graph D. In the formula, o i represents the operator in the operator set O, d represents the computing device in the device topology graph D, represents the memory overhead generated by the placement strategy of operator o i , and M(d) represents the maximum memory of the computing device d.

[0039] As shown in FIGS. 1-3, set the hypothesis constraints and operator grouping constraints to simplify the scale of the solution space. Since the performance difference between homogeneous devices is small and the predecessor nodes of the operator directly affect the start time of the operator execution. Therefore, in view of the above characteristics, the following two hypotheses are proposed: the computing and communication performance of homogeneous devices are the same; the execution end time of the operator is only related to the operators with direct dependencies. And based on the above hypotheses, the range of the solution space is constrained, and the hypothesis constraint conditions are shown in Equation (2):

[0040]

[0041] where d i , d j represent the computing devices in the device topology graph D, and represent the performance of homogeneous devices, ≡ represents identically equal to, o i , o j represent the operators in the operator set O, represents the placement strategy of operator o, Succ(o i ) method represents obtaining all the successor nodes of o i operator, σ represents the operator operation end time, represents that the placement of operator o i does not directly affect the execution end time of o j .

[0042] Through the above assumptions, the placement of two operators with data - dependency relationships is constrained to two cases. One is that the two operators are placed on the same device, and the other is that the two operators are placed on different devices. Among them, when the two operators are placed on the same device, it indicates that the communication cost between the two operators is greater than the computational cost; when the two operators are placed on different devices, it indicates that the computational cost between the two operators is greater than the communication cost.

[0043] 1 - 4, in view of the problem of a large number of operators and edges in the neural network, operator grouping constraints are proposed. Based on the operator fusion method, the operators in the computational graph are grouped, and the operators within the group are regarded as a whole for scheduling, so as to simplify the scale of the computational graph and further narrow the scope of the solution space. In order to maintain the characteristic of directed acyclicity of the computational graph and reduce the cost of detecting cycles, the operator grouping constraint adopts a conservative grouping method, that is, if one of the two operators with data - dependency relationships satisfies an out - degree or in - degree of 1, these two operators are constrained to the same group, and the operators within the group are regarded as a whole for scheduling, so as to simplify the structure of the computational graph, thereby reducing the scale of the solution space, and it can be ensured that there is only one data link between these two operators, so as to ensure that the ablation of these two operators will not cause the computational graph to form a cycle. That is, only when operator o i and operator o j meet the following condition of Equation 3, they are constrained to the same group:

[0044]

[0045] where, succ(o i ) represents the set of successor nodes of operator o i , outdegree(o i ) represents the out - degree of operator o i , and indegree(o j ) represents the in - degree of operator o j .

[0046] Step 2: Evaluate the execution performance of the operator placement strategy based on the cost evaluation model that fuses computational cost, communication cost, and load cost, and guide the double - population genetic algorithm to perform gene mutation and gene exchange iterations between two populations (placement strategies) to update the two sets of placement strategies.

[0047] 2 - 1, construct a cost evaluation model that fuses computational cost, communication cost, and load cost

[0048] First, analyze the key factors affecting the execution performance of the neural network model, that is, the computational performance, communication performance, and memory load characteristics of the device. Using the actual execution time of the operators in the neural network model on the device, establish the computational cost L i; Establish the communication cost \(W\) using the total tensor size of cross-device communication between operators in the neural network model; establish the load cost \(f\) using the sum of the ratios of the memory load offsets of each device to the average load. Among them, the computational cost \(L\) i 、The communication cost \(W\) and the established load cost \(f\) are defined as follows:

[0049] For the computational cost, use the end time of the actual execution of the operator on the device minus the start time. Then the computational cost of the operator is shown in Equation (4):

[0050]

[0051] Among them, and are respectively the start execution time and the end execution time of the operator \(o\) i on the device \(d\).

[0052] For the communication cost, use the total tensor size of cross-GPU communication of operators in the neural network. Then the communication cost of the model is shown in Equation (5):

[0053]

[0054] Among them, denote \(S\) i,j as the tensor size transmitted between the operator \(o\) i and the operator \(o\) j , \(\mu\) i,j is whether the operator \(o\) i and the operator \(o\) j need to transmit tensors across devices. When the operator \(o\) i and the operator \(o\) j are placed on the same device, \(\mu\) i,j is 0 indicating no need to transmit across devices, otherwise \(\mu\) i,j is 1 indicating the need to transmit across devices.

[0055] For the load cost, use the sum of the ratios of the memory load offsets and the average load of each device. Then the load balancing cost of the model is shown in Equation (6):

[0056]

[0057] Among them, \(A\) is the average memory size that each device of the current model needs to bear, \(W\) d represents the total tensor size of the operator that needs to communicate across devices on the \(d\)-th device, and \(\Delta\) d represents the parameter size of the operator on the \(d\)-th device.

[0058] Then, according to the computational cost, communication cost, and load cost defined above, fuse and construct a cost evaluation model, and its specific definition is as follows:

[0059] A cost evaluation model that fuses computational cost, communication cost, and load cost characterizes the performance of the model parallel strategy from three dimensions: computational cost, communication cost, and load cost, and constructs a comprehensive cost evaluation model to guide the automatic search and tuning of the operator placement strategy, as shown in Equation (7):

[0060]

[0061] Among them, represents the execution cost of the placement strategy; the weight ratio λ is obtained by empirical tuning; Ω is a linear fitting model of tensor size and communication time; since in the actual environment, part of the computational and communication time often overlaps, the weight ratio λ is used to simulate the situation of computational and communication parallelism.

[0062] 2-2. Iteratively update the placement strategy based on the double-population genetic algorithm

[0063] First, two groups of initial solutions are generated in the solution space using a random or specified strategy, that is, two groups of initial operator placement strategies are generated for the operators in the computational graph. In this method, each group of operator placement strategies is regarded as a population; then, the cost evaluation model that fuses computational cost, communication cost, and load cost is used to guide the gene mutation method and gene exchange method of the double-population genetic algorithm, and iteratively update the placement information of the two groups of placement strategy operators to quickly search for the placement strategy with the minimum execution cost in the solution space.

[0064] Use random initialization or a specified strategy to initialize the initial placement strategy for the operator.

[0065] 1) The random initialization method randomly places the operators in the computational graph into the device cluster to obtain the initial placement strategy. Its implementation method is shown in Algorithm 1, where random_device(D) represents randomly selecting a device d from the available devices. placement(P, o, d) represents placing the operator o on the device d and recording the placement information of the operator in the strategy P.

[0066]

[0067]

[0068] 2) The initialization method based on the specified strategy uses existing distributed parallel strategy search methods such as the m_topo method based on topological sorting placement, the m_etf method based on the earliest start time placement, and the m_sct method based on the minimum communication placement in the Beachi framework for initialization. Its implementation method is shown in Algorithm 2, where, initial_method(G * , D, T) represents using the specified distributed parallel strategy search method T to search the computational graph G *The placement strategy on the device cluster described by the device topology graph D. G′ represents the computation graph with operator placement information obtained after searching by the existing method, and get_device(o) represents obtaining the target placement device of operator o in the computation graph. placement(P, o, d) represents placing operator o on device d and recording the placement information of this operator in policy P.

[0069]

[0070] The gene mutation method consists of three steps: selection of operator pairs, gene mutation, and iterative update.

[0071] The gene mutation method selects the operator to be mutated based on the overall cost of the operator pair and iteratively updates the placement strategy by changing the placement information of the operator, thereby optimizing the strategy.

[0072] 1) Based on the principle of maximum cost first, select the operator pair with the largest overall cost in each group of placement strategies as the gene mutation operator.

[0073] The gene mutation method refers to two operators in the computation graph with data dependency as an operator pair, as Figure 2 shown. The overall cost of the operator pair is obtained by using the computational cost and communication cost of the operators on the gene, and the operator pair to be mutated is selected according to this overall cost. The calculation method of the overall cost of the operator pair is shown in Equation (8):

[0074]

[0075] where, represents the number of times the operator pair R ij has been selected, decay represents the dynamic decay rate, L i and L j respectively represent the computational cost of operator o i and operator o j , S i,j represents the size of the transmission tensor between operator o i and operator o j , Ω represents the linear fitting model of the tensor size and communication time, represents the overall cost of the operator pair R ij .

[0076] Since the operator pair R can be repeatedly selected during the execution of the algorithm, in order to increase the probability of the operator pair with a small overall cost being selected, the dynamic decay rate decay is used to dynamically decay the overall cost of the selected operator pair. The specific implementation of operator pair selection is shown in Algorithm 3:

[0077]

[0078] Among them, get_all_op_pair(G, P) represents obtaining all operator pairs according to the computational graph G and the placement strategy P, and caculate_cost(R) represents calculating the overall cost of the operator pair R; sort_by_cost(op_pair_list) represents sorting all operator pairs in descending order according to the overall cost of the operator pairs; p is the selected threshold.

[0079] 2) Use the gene mutation method to generate a new placement strategy by changing the placement information of the gene mutation operator.

[0080] This method uses the following rules to perform gene mutation on the selected operators:

[0081] (1) If the two operators in the currently selected operator pair are placed on different devices, randomly select an available device and place the two operators on this device.

[0082] (2) If the operators in the currently selected operator pair are placed on the same device, randomly select two available devices and place the two operators on these two different devices.

[0083] Using the above two rules, a new placement strategy can be generated on the basis of the original placement strategy. Rule (1) is as Figure 3 shown. If the operator pair selection method selects the operator pair R to be mutated 24 , then first randomly select an available new device 0, and then mutate R 24 ={0, 1} to R 24 ={2, 2}, which means that the o 2 operator and the o 4 operator are mutated from being originally on device 0 and device 1 respectively to placing the operator o 2 and the operator o 4 on device 2 at the same time, so as to obtain a new placement strategy. Similarly, Rule (2) is as Figure 3 shown. If the operator pair selection method selects the operator pair R 35 , then first randomly select two available new devices, device 2 and device 1, and then mutate R 35 ={1, 1} to R 35 ={2, 0}, which means that the o 3 operator and the o 5 operator are mutated from being originally placed on device 1 at the same time to placing the operator o 3 and the operator o 5 on device 2 and device 0 respectively, so as to obtain a new placement strategy.

[0084] 3) Guided by the cost evaluation model that fuses the computational cost, communication cost, and load cost, iteratively update the placement strategy

[0085] As Figure 4 shown, first, the cost evaluation model that fuses the computing cost, communication cost, and load cost is used to calculate the costs of the original placement strategy and the new placement strategy respectively; then, the costs of the two groups of placement strategies are compared, and the placement strategy with the smaller cost is selected as the starting placement strategy for the next round of iteration; finally, the operator pair R selected in the gene mutation in this round is used as a high-quality operator pair of the current placement strategy, indicating that the current placement information of the upper operator on this operator pair can reduce the strategy cost.

[0086] The overall process of the gene mutation method is as shown in Algorithm 4:

[0087]

[0088]

[0089] Among them, choose_gene represents selecting a gene from the population; genes_info represents the operator placement information on the specified gene; add_good_gene represents taking the gene as a high-quality gene; target_device represents selecting the target placement device; update_population represents updating the gene information of the specified population; cost_model is the comfort level of the specified population calculated by the cost evaluation model that fuses computing, communication, and load

[0090] The gene swapping method includes two steps: the high-quality operator pair cross-swapping method and iterative update.

[0091] 1) Based on the high-quality operator pair cross-swapping method, swap the gene information of the two groups of placement strategies.

[0092] In the high-quality operator pair cross-swapping method, the high-quality operator pairs in the two groups of placement strategies are given to each other, that is, the high-quality operator placement information of each is updated to the same operator of the other, so as to generate two new placement strategies with the operator placement information of the other.

[0093] As Figure 5 shown, the high-quality operator pair group of the first group of placement strategies retains the high-quality operator pairs R 12 ={1,0} and R 35 ={2,0}. The high-quality operator pair group of the second group of placement strategies retains the high-quality operator pairs R 24 ={2,1} and R 13 ={1,0}. The high-quality operator pair cross-swapping method updates the high-quality operator pairs R 12 and R 35 in the first group of placement strategies to the second group of placement strategies, indicating that the operator o in the second group of placement strategies1 and operator o 2 are respectively placed on device 1 and device 0, and operator o 3 and operator o 5 are respectively placed on device 2 and device 0. Similarly, the second set of placement strategies will also update its own high-quality operator pairs R 24 and R 13 into the first set of placement strategies, indicating that operator o 2 and operator o 4 of the first set of placement strategies are respectively placed on device 2 and device 1, and operator o 1 and operator o 3 are respectively placed on device 1 and device 0. Through the above method of cross-swapping high-quality operator pairs, two new placement strategies with each other's operator placement information are generated.

[0094] 2) Guided by the cost evaluation model that fuses computational cost, communication cost, and load cost, iteratively update the placement strategy.

[0095] Similar to the iterative update in the gene mutation method, the survival of the fittest in the gene swapping method also uses the cost evaluation model that fuses computational cost, communication cost, and load cost to judge the quality of the new and old populations generated by genomic cross-swapping and retain the relatively high-quality populations. In addition, in order to maintain the diversity of genes in the sibling populations and prevent the gene information in the sibling populations from converging during the iteration process, the high-quality gene groups of both populations will be cleared after each gene swap.

[0096] The overall process of the gene swapping method is shown in Algorithm 5:

[0097]

[0098] Step 3: When the iteration termination condition is reached, convert the two sets of placement strategies into the operator placement strategies of the data flow graph respectively, and train the model according to the operator placement strategies obtained by iteration in the real environment to select the strategy with the shortest execution time as the optimal distributed scheduling strategy of the model.

[0099] Among them, the iteration termination conditions include that the total number of iterations reaches the specified threshold and the placement strategy does not change within a certain number of iterations. When the iteration process of the double-population genetic algorithm satisfies one of the above two conditions, the iteration is terminated.

[0100] Step 4: Use the obtained distributed scheduling strategy to perform distributed scheduling on the AI model to be trained, and input the data sets required by models such as image recognition and natural language processing into the neural network for distributed training.

[0101] According to the method of constructing the solution space and solving based on the double-population genetic algorithm according to the above steps, the specific process of the deep learning adaptive training method based on the genetic algorithm proposed in this paper is as follows:

[0102] (1) First, construct the solution space of the integer linear programming problem based on the computational graph and the device topology graph, and simplify the scale of the solution space by using the hypothesis constraint and the operator grouping constraint.

[0103] (2) Secondly, initialize two groups of initial placement strategies required for the double-population genetic algorithm based on the methods of random and specified strategies. And based on the cost evaluation model that fuses the computational cost, communication cost, and load cost, guide the gene mutation method and the gene exchange method to iteratively update the two groups of placement strategies.

[0104] (3) Then, when the iteration termination condition is reached, convert the two groups of iteratively obtained placement strategies into the operator placement strategies of the data flow graph respectively, and train the model in the real environment, and select the strategy with the shorter execution time among the two groups of placement strategies as the optimal distributed scheduling strategy of the model.

[0105] (4) Finally, use the obtained distributed scheduling strategy to perform distributed scheduling on the AI model to be trained, and input the data sets required by models such as image recognition and natural language processing into the neural network for distributed training.

[0106] The specific description of this method is shown in Algorithm 6:

[0107] Among them, lines 1-2 extract the computational graph of the model and simplify the computational graph by the ablation method based on the operator grouping constraint.

[0108] Line 3 generates two groups of initial placement strategies P1 and P2 in the solution space by using the placement strategy initialization method described in Algorithm 1 and Algorithm 2.

[0109] Lines 4-11 use the double-population genetic algorithm to iteratively update the two groups of placement strategies, where op_pair_swap represents the iterative method based on gene exchange described in Algorithm 5, op_pair_variation represents the iterative method based on gene mutation described in Algorithm 4, and finish_condition represents the condition for the end of iteration, such as when the two groups of initial placement strategies no longer change in multiple rounds of iteration. Therefore, when the specified number of iterations is reached or the iteration end condition is met, the search process of the double-population genetic algorithm is terminated.

[0110] Line 12 converts the two groups of placement strategies after iteration into the operator placement strategies of the data flow graph.

[0111] Lines 13 - 20 execute the two groups of operator placement strategies in a real execution environment, and use the group of placement strategies with a shorter execution time as the optimal strategy for model parallel operator placement.

[0112] Lines 21 - 22 perform distributed scheduling on the model using the obtained distributed scheduling strategy, input the dataset into the model for distributed training, and return the trained model.

[0113]

[0114] Based on the method in this paper, the training speed of neural network models in the field of artificial intelligence such as image recognition and natural language processing can be accelerated. The comparison experiment results with single - GPU and the Beachi framework are shown in Table 1:

[0115] Table 1 Comparison of execution performance of single - GPU, Beachi, and the dual - population genetic algorithm strategy

[0116]

[0117]

[0118] The experimental results of the present invention and the m_topo, m_etf, and m_sct methods of the single - GPU and the Beachi framework are shown in Table 1. In the table, "random" in GA(random / m_topo / m_etf / m_sct) means using the random initialization method to initialize the two groups of placement strategies, and m_topo, m_etf, and m_sct mean using the specified strategies for initialization, that is, one of the two initial placement strategies required by the present invention is initialized by the m_topo or m_etf or m_sct method, and the other group is initialized by the random initialization method. GA is the experimental result of the present invention. It can be seen from Table 1 that the strategies searched by the present invention are generally better than the three strategy search algorithms of Beachi. Compared with the m_topo algorithm, the maximum improvement can be 8.9%; compared with m_etf, the maximum improvement can be 42.2%; compared with m_sct, the maximum improvement can be 40.4%. The genetic algorithm uses the specified strategy initialization method to initialize the placement strategy, and iteratively updates on this basis. The single - step execution time of the obtained strategy compared with the original strategy can be improved by up to 9%. In addition, Beachi will encounter OOM (out - of - memory) when the batch size is large, making the strategy unable to execute, while the dual - population genetic algorithm can control the load balance of devices in the cluster to prevent the occurrence of OOM.

Claims

1. A neural network adaptive distributed parallel training method based on genetic algorithm, characterized in that it includes the following steps: Step 1: Transform the operator placement problem in the neural network model into an integer linear programming problem; construct a solution space based on the data flow graph and device topology graph of the neural network model, and simplify the scale of the solution space by setting hypothesis constraints and operator grouping constraints: 1-1. Extract the neural network model structure and form a device resource group, and abstract to obtain a computation graph G(O, E) and a device topology graph D. In the computation graph G(O, E), the vertex O represents the operator set of the neural network model, and E represents the set of directed edges between vertices. The device topology graph D is composed of M computing devices; 1-2. Given the computation graph G(O, E) and the device topology graph D, on the premise of satisfying the device memory constraint, find a set of placement strategies to minimize the execution cost, thereby transforming the problem of placing the operators of the neural network model on the devices into an integer linear programming problem, and construct a solution space based on the data flow graph and device topology graph; 1-3. Put forward two hypotheses by analyzing the device characteristics in the homogeneous cluster and the execution characteristics of the neural network model: the computing and communication performance of homogeneous devices are the same; the execution end time of an operator is only related to the operators with direct dependency relationships; Through the above hypotheses, the placement situations of two operators with data dependency relationships are constrained to two types: one is that the two operators are placed on the same device, and the other is that the two operators are placed on different devices; 1-4. Subsequently, put forward an operator grouping constraint. For two operators with data dependency relationships, if one of the operators satisfies the condition that the out-degree or in-degree is 1, then these two operators are constrained to the same group, and the operators in the group are regarded as a whole for scheduling; Step 2: Construct a cost evaluation model that fuses computing cost, communication cost, and load cost, comprehensively evaluate the quality of the solution of the integer linear programming problem, that is, the execution performance of the operator placement strategy, and guide the double-population genetic algorithm to perform gene mutation and gene exchange iteration between two placement strategies to update the two sets of placement strategies: 2-1. Construct a cost evaluation model that fuses computing cost, communication cost, and load cost First, analyze the key factors affecting the execution performance of the neural network model, and establish the computational cost L using the actual execution time of the operators in the neural network model on the device. i ; Establish the communication cost W using the total tensor size of cross-device communication between operators in the neural network model; Establish the load cost f using the sum of the ratios of the memory load offsets and average loads of each device. Secondly, by comprehensively considering the computing cost, communication cost, load cost, and the characteristics of the actual training process of the model, establish a cost evaluation model that fuses computing cost, communication cost, and load cost: π θ = (∑L i + λ × Ω(W)) × (1 + f) Among them, π θ represents the execution cost of the placement strategy, λ represents the weight ratio, and Ω is a linear fitting model of the tensor size and communication time; 2-2. Iteratively update the placement strategy based on the double-population genetic algorithm First, use the initialization method of the placement strategy to generate two sets of placement strategies with initial operator placement information in the solution space; Secondly, based on the cost evaluation model, guide the gene mutation and gene exchange methods of the double-population genetic algorithm to iteratively update the placement information of the operators in the two sets of placement strategies; The gene mutation includes three steps: (1) Based on the principle of maximum cost first, select the operator pair with the largest overall cost in each set of placement strategies as the gene mutation operator; (2) Use the gene mutation method to generate a new placement strategy by changing the placement information of the gene mutation operator; (3) Guided by the evaluation of the cost evaluation model, iteratively update the placement strategy; The gene exchange includes two steps: (1) Exchange the gene information of two groups of placement strategies based on the high-quality operator crossover method; (2) Iteratively update the placement strategy guided by the cost evaluation model evaluation; (3) When the termination condition for the iterative update of the placement strategy is reached, convert the two groups of placement strategies into operator placement strategies for the data flow graph respectively, and train the neural network model according to the iterative operator placement strategy in the real environment, and select the strategy with the shortest execution time as the optimal distributed scheduling strategy for the neural network model; (4) Use the optimal distributed scheduling strategy obtained in step (3) to perform distributed scheduling on the neural network model to be trained, and input the data set required by the neural network model into the neural network model for distributed training.

2. A neural network adaptive distributed parallel training method based on the genetic algorithm according to claim 1, wherein: the computational graph G(O, E) described in step 1-1 is abstracted from the data flow graph of the neural network model.

3. A neural network adaptive distributed parallel training method based on the genetic algorithm according to claim 1, wherein: the key factors described in step 2-1 are the characteristics of the computing performance, communication performance, and memory load of the device.

4. A neural network adaptive distributed parallel training method based on the genetic algorithm according to claim 1, wherein: the initialization methods described in step 2-2 include the random initialization method and the specified strategy initialization method; the random initialization method refers to randomly allocating the operators in the computational graph G(O, E) to the device topology graph D; the specified strategy initialization method refers to using the placement strategy searched by the existing parallel strategy search method as the initial state; Optionally select one of the initialization methods to initialize the placement strategy.

5. A neural network adaptive distributed parallel training method based on the genetic algorithm according to claim 1, wherein: the termination conditions for the iterative update described in step 3 include that the total number of iterations reaches a specified threshold and the placement strategy does not change within a certain number of iterations; When the iterative process of the double-population genetic algorithm satisfies one of the above two conditions, terminate the iteration.

6. A neural network adaptive distributed parallel training method based on the genetic algorithm according to claim 1, wherein: the overall cost is calculated as follows: Among them, represents the number of times the operator pair R ij has been selected, decay represents the dynamic decay rate, L i and L j respectively represent the computational costs of the operator o i and the operator o j respectively, S i,j represents the size of the transmission tensor between the operator o i and the operator o j , represents the overall cost of the operator pair R ij .

7. A neural network adaptive distributed parallel training method based on the genetic algorithm according to claim 1, wherein: the rules for gene mutation are as follows: (1) If the two operators in the currently selected operator pair are placed on different devices, randomly select an available device to place the two operators on this device; (2) If the operators in the currently selected operator pair are placed on the same device, randomly select two available devices to place the two operators on these two different devices.

Citation Information

Patent Citations

  • Adaptive distributed parallel training method for neural network based on reinforcement learning

    CN113128702A

  • Dynamic load balancing method for model parallelism of neural network

    CN114217944A