Distributed training method of deep neural network based on automatic analysis of model structure

The automatic analysis of deep neural network structures optimizes distributed training by classifying operators and identifying bottlenecks, enhancing training efficiency and reducing costs for large-scale models.

GB2638054APending Publication Date: 2025-08-13HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
GB2024016306
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-07
Filing Date
2024-11-05
Publication Date
2025-08-13

AI Technical Summary

Technical Problem

The distributed training of large-scale deep neural networks is challenging due to complex model structures and the difficulty in designing efficient parallel strategies, with existing methods being unsuitable for large models and requiring extensive manual expertise and complex search processes.

Method used

An automatic analysis method for deep neural network structures classifies operators into memory-intensive and compute-intensive types, evaluates their performance, and identifies bottlenecks to optimize distributed training by distributing operators across devices based on computing, storage, and communication costs.

Benefits of technology

This approach improves training efficiency and reduces costs by automatically identifying and addressing bottlenecks, enabling faster training of larger-scale neural networks without manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A distributed training method of a deep neural network model based on an analysis of the structure of the model whereby different types of operators (operations) are being evaluated and classified into memory-intensive operators (i.e. operations where most of the time is spent on memory access such as Concat, Add, ReLU, MaxPooling) and compute-intensive operators (i.e. operations that incur high computational cost such as Conv, FC, MatMul, LSTM) and wherein communication performance or cost (e.g. a bottleneck) of the operations is deduced and wherein the method analyses a data flow structure of the neural network model and searches for a network layer or an operation that consumes most of the time or uses most of the communication traffic i.e. incurs a high cost to the training of the model. The data flow structure may be represented by a computation graph such as a Directed Acyclic Graph (DAG) wherein each node represents a separate neural network operator, such as matrix multiplication (MatMul), convolution Conv operators etc. In this way, an improved computation graph may be obtained so the neural network model may be trained faster and / or more efficiently.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure belongs to the technical field of neural networks, in particular to a distributed training method of a deep neural network based on automatic analysis of a model structure. BACKGROUND

[0002] In recent years, with the development of industrialized information and Internet, deep learning technology has been fully developed. With its powerful feature extraction and classification ability, the deep learning technology has been widely used in the fields such as finance, medicine, recommender systems and molecular dynamics, which changes the lifestyles of people and improves the production efficiency. With the increasing data scale, deep learning models become more and more complex. For example, the ResNet-152 model has approximately 58 million parameters, while the EfficientNet-B7 model has approximately 64 million parameters. Besides, the Generative Pre-trained Transformer (GPT)-3 and Chat Generative Pre-trained Transformer (ChatGPT) network models have about 175 billion parameters, and the memory' size required for the training of the above models needs thousands of GBs at most. Generally, the model accuracy can be improved by training a larger model with more data samples. However, the growth rate of the memory' capacity of computing acceleration devices (such as GPUs) is far less than the growth rate of memory demand for training the neural network. The computing and storage capacity of a single device is limited, so that it is difficult to withstand the training of large-scale data sets and huge models.

[0003] However, the current deep learning seivers generally contain a plurality of deep neural network training cards (GPUs / NPUs). Therefore, how to use the distributed training technology to split the deep learning model into sub-models and schedule them to multiple devices for training to improve the training performance of the model has become a hot spot in the field of deep learning.

[0004] At present, the distributed training of the model mainly depends on the experience of experts, which requires developers to deeply understand the structure features of the neural network model and the training device, and design the distributed training strategy of the model based on these features. However, with the increasing scale of the model and the device cluster, the search space of the strategy increases exponentially, so that it is difficult to design a parallel strategy with better performance in a short time based on expert experience.

[0005] In order to improve the performance of the model parallel strategy and the efficiency of searching strategies, the academic world and the industrial world begin to study the automatic parallel method for models. The automatic parallel method for model analyzes the neural network model automatically by means of a search algorithm by extracting the model structure and the device topology information, so as to simplify the design, process of the model parallel strategy and improve the distributed training performance of the model. At present, the mainstream automatic parallel methods for models are divided into the machine learning-based method and the graph algorithm-based method. The machine leaming-based method needs iterative updating of parameters and a lot amount of computation, which results in a longer search time of the strategy. However, the graph algorithm-based method usually uses the algorithms, such as dynamic programming and shortest path algorithm, for search, which often has problems such as insufficient policy execution performance and unbalanced policy load. At present, the above methods often need to comprehensively take into account the features in many aspects such as the operator structure, the execution performance and the device topology, acquire data of the operators in a real environment, and use a strategy performance simulator to guide the strategy search, which results in a very complex search process, so that the above methods are only suitable for small- and medium-sized models, but not for large models that are very popular in recent years. SUMMARY

[0006] The distributed parallel training strategy of the deep neural network is difficult to be designed and implemented due to complex structure of the deep neural network and difficult analysis. To deal with this problem, the present disclosure designs a distributed training method of a deep neural network based on automatic analysis of a model structure, so as to evaluate the structure of the deep learning model and its computing amount efficiently and accurately and support the design and implementation of the distributed parallel training strategy of the model. First, the present disclosure analyzes model structure operators of a deep neural network, carries out computing and storage evaluating modeling on different types of operators according to the features such as computation, storage and parameters of the operators, classifies the operators in the neural network into memory-intensive operators and compute-intensive operators according to the computing or storage features of the neural network, and deduces the operator performance automatically according to the input and output of the operator through the evaluating model. Second, the present disclosure analyzes features of data flow7 of operators according to the computation graph of the deep neural network, and constructs data flow among operators or a communication evaluating model in combination with the operator parameters and the communication modes, and automatically deduces the operator communication performance through the evaluating model. Finally, the present disclosure, according to the computation graph and the operator features of the model, based on the computing and storage evaluating model, the data flow and the communication evaluating model, analyzes the model structure automatically, identifies parameter-intensive operators and compute-intensive operators in the model structure automatically, and searches for a network layer or an operator that consumes the most time or has the most communication traffic (the network layer or the operator is the computing or communication bottleneck of the overall model structure), so as to carry out distributed training, thereby improving the training efficiency and reducing the training cost.

[0007] The present disclosure provides a distributed training method of a deep neural network based on automatic analysis of a model structure, which includes the following steps.

[0008] Step 1: model structure operators of a deep neural network are analyzed; computing and storage evaluating modeling is carried out on different types of operators according to features such as computation, storage and parameters of the operators; the operators in the neural network are classified into memory-intensive operators and compute-intensive operators according to computing or storage features of the neural network; and an operator computing cost is deduced automatically according to an input and an output of each operator by a computing and storage evaluating model.

[0009] Sub-step 1.1: a computation graph generated by a deep learning framework for the deep neural network is scanned, and a directed acyclic graph G(O,E) record is defined, where nodes represent operators, each node represents a separate neural network operator, such as matrix multiplication MatMu!, convolution Conv operators, etc. Directed edges, representing data dependencies, represent data dependencies among operators. If output of operator B is input of operator A, the directed edge from node A to node B indicates that data flows from A to B. These directed edges can further be used to represent an execution order of operators. And an operator structure and data flow information in the computation graph are recorded in a JSON format.

[0010] Sub-step 1.2: data flow of each operator is formed by its input tensor and its output tensor. The computing and storage evaluating model is constructed by combining the operator and data flow of the operator, the evaluating model is capable of automatically scanning all operators in the computation graph and computing memory7 cost of operators according to type of operators and its data flow (size of the input tensors and data type of the input tensors). The memory cost is mainly related to weight parameters of the operators and intermediate results, so a memory cost of the operators with parameter weights is expressed by following equation:

[0011] Costmemory ~ (1^1 x W2 X...X x sizeof^Wtype) + x Hz X...X x sizeof^Htype)

[0012] where T indicates a number of operators with the parameter weights, MA , W2, ..., Wh indicates size of the parameter weights, sizeof {Wtype} indicates a byte size of the data type Wtype of a parameter, Wtype includes FP16, FP32, etc., K indicates a number of output tensors, H2> ..., Hh indicate a size of h dimensions of the tensor, and sizeof^Htype) indicates the byte size of the data type Htype of acquired data, and Htype includes FP16, FP32, etc.

[0013] Sub-step 1.3: the operator computing cost refers to an overhead generated by tensor computing, which embodies a process of transforming tensors. Therefore, the computing cost Costcompute is computed based on the input tensor and the output tensor of the operator, which is expressed by following equation:

[0014] Costcompute = X 1+e 1+sout

[0015] where S:n indicates sum of the input tensors of the operator and Sout indicates sum of the output tensors of the operator,. When the change between the input tensor and the output tensor is large, it indicates that a computing process of the operator is complex, and the computing cost of the operator is high. R indicates a cost conversion rate, which is obtained from implementation analysis.

[0016] The neural network operators are divided into two categories: memory-intensive operators and compute-intensive operators. The memory-intensive operators refer to operators whose most of time is spent on memory access, such as operators of Concat, Add, ReLU, MaxPooling, etc. The compute-intensive operators refer to operators, with high computing complexity, whose most of time is spent on computing. The compute-intensive operators are often responsible for some complex matrix operations, such as operators of Conv, FC, MatMui, LSTM etc The operators can be successfully classified by computing and comparing Costmemory and Costcorr)-pUf-e of the operators.

[0017] Step 2: features of the data flow of the operators are analyzed according to the computation graph of the deep neural network obtained in Sub-step 1.1, and the data flow among the operators and the communication evaluating model are constructed in combination with operator parameter and communication mode, so as to automatically deduce operator communication performance by the data flow and the communication evaluating model.

[0018] Sub-step 2.1: common inter-cluster communication modes include AllReduce, AllGather, AHToaH, etc. For example, AllReduce communication is to summarize and sum up data in each process before distribution, and each process will communicate K data when there is K data in each process. However, in AllGather communication, each process will collect data from all processes, and each process will communicate K X N data. Therefore, it is necessary to acquire the communication mode of communication operator from the computation graph. In the present disclosure, the communication mode is determined by analyzing data edge in the computation graph. For example, AllReduce usually involves synchronization and reduction operations among a plurality of nodes, while AllGather involves data collection among nodes. An input-output relationship of the operator is determined by scanning the computation graph. The communication mode of operation between the operator and adjacent nodes is analyzed, and a determination result is added to the directed acyclic graph G(O, E).

[0019] Sub-step 2.2: according to the data flow information of the operators, in combination with operator parameters and the communication mode, the data flow' among the operators is analyzed or the communication evaluating modeling analysis is performed on the operators, and an operator communication cost is automatically derived. The communication cost among operators is closely related to a transmission tensor. The communication cost can be obtained by using the output tensor of the operator and its data type, which is expressed by following equation:

[0020] Costcommunicate = x H2 x.,.x Hh) x sizeof (Htype) X T(Ctype)

[0021] where K indicates a number of the output tensors, sizeof (Htype) indicates the byte size of the data type Htype of tensors, the common Htype includes FP16 and FP32, and T(Ctype) indicates a multiple of communication traffic generated when the communication mode is Ctype. For example, each process in AllGather communication needs to collect data from other processes, so T(AllGather) is equal to a number of processes. In the method, the size of the output tensor of the operators is set as the communication cost.

[0022] Step 3: according to the computation graph of the deep neural network obtained in Sub-step 1.1, by utilizing the computing and storage evaluation model along with the data flow and communication evaluation model, the model structure are automatically analyzed, and parameter-intensive operators and compute-intensive operators is automatically identified in the model structure, and a network layer or an operator that consumes the most time or has the most communication traffic is searched for (the network layer or the operator is the computing or communication bottleneck of the overall model structure), so as to carry out distributed training.

[0023] Sub-step 3.1: according to the computing and storage evaluating modeling, the data flow among the operators and modeling analysis of the communication evaluating model, the computing cost Costcompute, the memory cost Costmemory and the communication cost Costc„mmimicate of each operator are obtained, and k operators with the highest cost are searched for automatically from the computation graph according to the above three costs. These operators are often the computing or communication bottleneck in the overall model structure. The operators with a high computing cost are divided and placed on a plurality of devices, and the operators with a high communication cost are gathered and placed on the same device. Alternatively, the computing cost, the memory cost and the communication cost automatically derived from an analysis and computation graph of this present disclosure are input into a model parallel strategy, which can help the strategy to divide the model more effectively and quickly, without the algorithm engineer profiling model training to manually search for the bottleneck of the model structure. An effective model dividing and placing method can make better use of computing resources, storage resources and communication resources, and can train neural network models faster and train larger-scale complex neural network models.

[0024] The present disclosure has the following beneficial effects.

[0025] Model structure operators of a deep neural network are analyzed automatically. Computing and storage evaluating modeling is carried out on different types of operators according to the features such as computation, storage and parameters of the operators. The operators in the neural network are classified into memory-intensive operators and compute-intensive operators according to the computing or storage features of the neural network. The operator performance is deduced automatically according to the input and output of the operator through the evaluating model.

[0026] Features of data flow of operators are analyzed automatically according to the computation graph of the deep neural network. The data flow among operators or the communication evaluating model is constructed in combination with the operator parameters and the communication modes. The operator communication performance is automatically deduced through the evaluating model.

[0027] Through the automatic evaluating modeling analysis of the above two evaluating models, a network layer or an operator that consumes the most time or has the most communication traffic is searched for (the network layer or the operator is the computing or communication bottleneck of the overall model structure), which helps artificial intelligence (AI) engineers to design parallel strategies, thereby facilitating distributed training, and thus improving the training efficiency and reduce the training cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] FIG. 1 is an overall architecture diagram.

[0029] FIG. 2 is a schematic diagram of operator structure information and data flow.

[0030] FIGs. 3A-3C are schematic diagrams of a communication mode. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The present disclosure wall be further explained in conjunction with the attached drawings and specific implementation steps:

[0032] As shown in FIGS. 1 to 3C, a distributed training method of a deep neural network based on automatic analysis of a model structure includes step l-step3.

[0033] In step 1, model structure operator's of a deep neural network are analyzed, computing and storage evaluating modeling is carried out on different types of operators according to the features such as computation, storage and parameters of the operators, and the operators in the neural network are classified into memory-intensive operators and compute-intensive operators according to the computing or storage features of the neural network; and an operator performance is deduced automatically according to the input and the output of the operators by the evaluating model.

[0034] The step 1 includes sub-step l.l, sub-step 1.2 and sub-step 1.3.

[0035] In sub-step 1.1, the computation graph information of the model to be analyzed is obtained, which is represented by the directed acyclic graph G(O, E). As shown in FIG. 2, nodes represent operators, each node represents a separate neural network operator, such as matrix multiplication MatMul, convolution Conv operators, etc. Directed edges, representing data dependencies, represent data dependencies among operators. If the output of operator B is the input of operator A, the directed edge from node A to node B indicates that data flows from A to B. These directed edges can further be used to represent the execution order of operators, and the operator structure and data flow information in the computation graph are recorded in a JSON format. For example, MatMul operator can be converted into [{"name" : "bert / embeddings / MatMul", "inputs": [{"name":"bert / embeddings / one_hot", "shape": [4096, 2], "dtype": "float32"}, {"name":"bert / embeddings / MatMulZReadVariableOp", “shape": [2, 768], "dtype": "float32"}], "outputs": [{"name":"0", "shape": [4096, 768], "dtype": "float32"}]}], where “name” represents the complete operator name of the recorded operator, “inputs” contain all the input operators of the operator, the “shape” information and “type” information of the input operators, and “outputs” contain the “shape” information and “type” information of all of output tensors of the operator.

[0036] In sub-step 1.2, data flow of each tensor is formed by its input tensor and its output tensor. The computing and storage evaluating model constructed by combining the operator and data flow of the operator, is capable of automatically scanning all operators in the computation graph and computing memory cost of operators according to the type of operators and its data flow (the size of the input tensors and the data type of the input tensors). The memory cost is mainly related to the weight parameters of the operators and the intermediate results, so a memory cost of the operators with parameter weights is expressed by the following equation.

[0037] Costmemory = x X...X Wh) x sizeof (Wtype) + x -¾ X...X Hh) x sizeof (Htype)

[0038] where T indicates the number of operators with parameter weights, , VF2 >.... 144 indicates the size of parameter weights, sizeof (Wtype) indicates the byte size of the data type Wtype of a parameter, the common IVtype usually includes FP16 and FP32, K indicates the number of output tensors, H2» ...» Hh indicates the size of h dimensions of the tensor, and sizeof (Htype) indicates the byte size of the data type Htype of acquired data , the common Htype usually includes FP16 and FP32. Htype and Wtype are of the same type. Taking the common Dense fully connected layer operator as an example, the operator contains a weight matrix of hidden„size X ffn_.hidden.size of FP32 type, where the input tensor is a x hidden_size and the output tensor is a X ffri_.hidden_.size. Therefore, the memory' cost is (hidden_size X ffn_.hidden_.size + a x ffn_hidden_size) x 4 w 8 bytes. For operators without weights, such as the MatMul operator, the sum of weights of the equation (i.e., the memory cost of (hidden.size x ffn_hidden_size Tax ffn_hidden_size) x 4 a- 8 bytes) is 0.

[0039] Algorithm 1 Computation of the memory cost of the model operator Input: model computation graph and data flow Output: memory cost of the model operator 1 def get..cost_memory(graphJosn): 2 cost_memory = [] 3 for operator in graph J son: 4 name = operatorf'name"], weights = operator ["weights"] 5 outputs = operator["outputs"],w7eight_cost = 0, output^cost = 0 6 for wei ght i n wei ghts: 1 weight_cost += weightfshape"] * sizeof(weight["dtype"]) 8 for output in outputs: 9 ouput cost+= outputpshape"] * sizeof(output["dtype"]) costjnemonc append( {"name": name, 10 "cosp memory ”: weight j:ost + 11 outpput cost}) 12 return cost_memory

[0040] Sub-step 1.3: the operator computing cost refers to the overhead generated by tensor computing, which embodies the process of transforming tensors. Therefore, the computing cost is computed based on the input tensor and the output tensor of the operator, which is expressed by the following equation: S[n + out Costcompute = 7 x r Pm ^outl 1 g

[0041] where Sin indicates the sum of the input tensors of the operator, and Sout indicates the sum of the output tensors of the operator, in which fl indicates a cost conversion rate, which is obtained from the implementation analysis. For example, the above MatMuI operator has two inputs, the shapes of which are [4096, 2] and [2, 768], respectively, and the dtypes of which are both FP32; so Sin = (4096 x 2 +■ 2 x 768) x 32 h- 8 = 38K bytes. The above MatMuI operator has one output, the shape of which is [4096, 768], so Sma = (4096 x 768) x 32 -:-8 = 12M bytes. The computing cost based on experimental analysis generally accounts for about 20% of the total cost, so the conversion rate R is 0.2. Based on this, the computing cost of the MatMuI operator can be computed.

[0042] Algorithm 2 Computation of the computation cost of the model operator Input: model computation graph and data flow Output: computation cost of the model operator 1 def get„cost„compute(graphjosn): 2 cost_compute = [] 3 for operator in graph J son: 4 name = operator["name"],inputs= operator}"! nputs"] 5 outputs = operator["outputs!!],S_in = 0, S out = 0 6 for input in inputs: 7 SJn += input["shape“] * sizeof(input["dtype"]) 8 for output in outputs: 9 S out += output["shape"] * sizeonfoutputrMtype”]) cost = (S in + S out) / (1 + pow(e, -abs(S in - S out) / (1 + 10 S out))) * R 11 cost_compute.append({"name" :name, "costcompute":cost}) 12 return cost compute

[0043] The neural network operators are divided into two categories: memory-intensive operators and compute-intensive operators. The memory-intensive operators refer to operators whose most of the time is spent on the memory access, such as operators of Concat, Add, ReLU, MaxPooling, etc. The compute-intensive operators refer to operators, with high computing complexity, whose most of time is spent on computing. The compute-intensive operators are often responsible for some complex matrix operations, such as operators of Conv, FC, MatMul, LSTM etc. The operators can be successfully classified by computing and comparing Cc)Stmemory and Costcompute of the operators.

[0044] Step 2: features of data flow of operators are analyzed according to the computation graph of the deep neural network, and data flow among operators or a. communication evaluating model is constructed in combination with the operator parameters and the communication modes, so as to automatically deduce operator communication performance by the communication evaluating model.

[0045] Sub-step 2.1: the common inter-cluster communication modes include AllReduce (all data integration communication), AllGather (all data convergence communication), AllToall (all data to all data communication), etc., as shown in FIGs. 3A-3C. For example, AllReduce communication is to summarize and sum the data in each process before distribution, and each process -will communicate K data when there is K data in each process. However, in AllGather communication, each process will collect data from all processes, and each process will communicate K X N data. Therefore, it is necessary to acquire the communication mode of the communication operator from the computation graph.

[0046] Sub-step 2.2: according to the data flow information of operators, in combination with the operator parameter and the communication mode, the data flow among operators is analyzed or communication evaluating modeling analysis is performed on the operators, and the operator communication performance is automatically derived. The communication cost among operators is closely related to the transmission tensor. The communication cost can be obtained by using the output tensor of the operator and its data type, which is expressed by the following equation:

[0047] Costcommunicate = x W2 X...X x sizeof (Htype) X T(Ctype)

[0048] where K indicates the number of output tensors, sizeof (Htype) indicates the byte size of the data type Htype of tensors, the common Htype includes FP16 and FP32; and T(Ctype) indicates the multiple of communication traffic generated when the communication mode is Ctype. For example, each process in AllGather communication needs to collect data from other processes, so T(AllGather) is equal to the number of processes. In the method, the size of the output tensor of the operators is set as the communication cost. Still taking the above MatMul operator as an example, the operator has an output tensor, K = 1, its shape is equal to [4096, 768], so h = 2, and the dtype is FP32 When tensors parallelism are used, the complete output of the MatMul operator may be obtained through summation of the AllReduce operator, and T(AllReduce) == 1. Therefore, the communication cost of the MatMul operator is (4096 X 768) X 32 X 1 -r- 8 = 12M bytes. For those operators with no communication type being adopted and among which direct transmission is performed, T is equal to I.

[0049] Algorithm 3 Computation of the communication cost of the model operator Input: model computation graph and data flow Output: communication cost, of the model operator 1 def get.cost.communicate(graphjosn): 2 costcommunicate = [] 3 for operator in graph J son: 4 name = operator["name"] 5 outputs == operator["outputs"], output_cost == 0 6 for output in outputs: 7 output_cost += output["shape"] 8 * sizeof(output["dtype"]) * T(output["Ctype”]) 9 cost communicate.append({"name":name, 10 "cost communi cate": output cost}) 11 return cosi^commumcate

[0050] Step 3: according to the computation graph of the model and feature of the operators, based on the computing and storage evaluating model, the data flow and the communication evaluating model, the model structure are automatically analyzed, and parameter-intensive operators and compute-intensive operators in the model structure are automatically identified, and a network layer or an operator that consumes the most time or has the most communication traffic is searched for (the network layer or the operator is the computing or communication bottleneck of the overall model structure), so as to carry out distributed training.

[0051] Sub-step 3.1: according to the computing and storage evaluating modeling, the data flow among operators and the modeling analysis of the communication evaluating model, the computing cost Costcompute, the memory' cost Costmemory and the communication cost C^jst.:omm.imicate of each operator are obtained. List objects returned by the above three computing functions are input into the Topk operator: Topk{ Costcompute J, i=l...n}, Topk{CostmemO!yj, 1=1...11), Topk{Costcomfnun;cateJ, i=l..m), where ii indicates the number of operators, k operators with the highest cost are searched for automatically from the computation graph. These operators are often the computing or communication bottleneck in the overall model structure. It may help AI engineers design parallel strategies upon finding these operators. And computing resources, storage resources and communication resources may be better used when these operators are trained in a distributed manner in parallel reasonably, thereby training neural network models faster and training larger-scale complex neural network models.

[0052] Finally, the present disclosure carries out an experiment on the Bert model. This method automatically reads the O-th layer of the computation graph of the Bert model (due to a similar structure of each layer in the Transformer model). The computing cost, the memory cost and the communication cost are automatically analyzed by the computing and storage evaluating model of the operators and the data flow among operators or the communication evaluating model analysis, k :::: 5 operators with the highest cost are automatically found from the computation graph. The results are shown in Table 1 to Table 3:

[0053] Table 1; Top 5 of computing overhead of the O-th layer of Bert

[0054] Operator name Computing cost bert / encoder / layer_0 / intermediate / dense / Bia sAdd 25165825 bert / encoder / layer__0 / output / dense / MatMul 22593239 bert / encoder / layer_0 / intermediate / dense / add bert / encoder / iayerJ) / intermediate / dense / Po w 20132660 20132660 bert / encoder / layer__0 / intermediate / dense / mul | 15099495 -3 1

[0055] Table 2: Top 5 of memory access overhead of the O-th layer of Bert.

[0056] Operator name bert / encoder / layer__0 / intermedlate / dense / add Memory cost 150994944 bert / encoder / layer_0 / intermediate / dense / Bia sAdd 100663296 bert / encoder / layer__p / intermediate / dense / Tan h 50331648 bert / encoder / layer_0 / intermediate / dense / Po w 37748736 bert / encoder / layer_0 / output / dense / MatMui

[0057] Table 3: Top 5 of communication overhead ol 25165824 f the O-th layer of Bert

[0058] Operator name Communication cost bert / encoder / layer_p / !ntermedlate / dense / add 150994944 bert / encoder / layer_0 / intermediate / dense / Bia sAdd 100663296 bert / encoder / layer___0 / intermediate / dense / Tan h 50331648 bert / encoder / layer_0 / intermediate / dense / Po w bert / encoder / !ayerJ) / output / dense / MatMul 37748736 25165824

[0059] The above operators are the computing or communication bottleneck in the overall model structure. The operators with a high computing cost are divided and placed on a plurality of devices, and the operators with a high communication cost are gathered and placed on the same device. Alternatively, the computing cost, the memory cost and the communication cost automatically derived from the analysis and computation graph of this present disclosure are input into a specific model parallel strategy, which can help the strategy to divide the model more effectively and quickly, without the algorithm engineer profiling model training process to manually search for the bottleneck of the model structure.

Claims

1. A distributed training method of a deep neural network based on automatic analysis of a model structure, comprising:Step 1, analyzing model structure operators of a deep neural network, carrying out computing and storage evaluating modeling on different types of operators, and classifying theoperators in the neural network into memory-intensive operators and compute-intensive operators; deducing an operator computing cost automatically according to an input and an output of each operator by a computing and storage evaluating model;Step 2, analyzing features of data flow of the operators, and constructing data flow among the operators and a communication evaluating model to automatically deduce operator communication performance by the data flow and the communication evaluating model;Step 3: based on the computing and storage evaluating model, the data flow and the communication evaluating model, analyzing model structure automatically, and identifyingautomatically parameter-intensive operators and compute-intensive operators in the model structure, and searching for a network layer or an operator that consumes a most time or has a most communication traffic for distributed training.

2. The method according to claim 1, wherein Step 1 comprises:Step 1.1: scanning a computation graph generated by a deep learning framework for the deep neural network, and defining a directed acyclic graph G(O,E) record, wherein nodesrepresent operators, each node represents a separate neural network operator, and directed edges, representing data dependencies, represent data dependencies among operators and further represent execution order of operators; and recording an operator structure and data flow information in the computation graph in a JSON format,Step 1.2: forming a data flow of each operator by its input tensor and its output tensor,constructing the computing and storage evaluating model by combining the operator and the data flow of the operator, wherein the evaluating model is capable of automatically scanning all operators in the computation graph, and computing memory cost of the operators according totype of the operators and data flows of the operators, which is expressed by following equation:Costmemory~ 5 (Wt x W2 x.. ,xi — 1X sizeof(Wtype)x / 72 x...x / , sizeof(Htype)where T indicates a number of the operators with parameter weights, 1¾ , W2, ..., Wh indicates a size of the parameter weights, sizeof (Wtype) indicates a byte size of data type Wtype of the parameter, K indicates a number of the output tensors, H2, ..., Hh indicatea size of h dimensions of the tensor, and sizeof'CHtype) indicates the byte size of data type Htype of acquired data,Step 1.3: computing a computing cost Costcompute based on the input tensor and the output tensor of the operator by using following equation, wherein the operator computing cost refers to overhead generated by tensor computing:wherein Sin indicates sum of the input tensors of the operator and Sout indicates sum of the output tensors of the operator, and R indicates a cost conversion rate,wherein the neural network operators are divided into memory-intensive operators and compute-intensive operators, and the memory-intensive operators refer to operators whose time is spent on memory access; the compute-intensive operators refer to operators whose time is spent on computing, and the operators are classified by computing and comparing Costmemory and Costcompute of the operators.

3. The method according to claim 2, wherein Step 2 comprises:Step 2.1: determining a communication mode by analyzing data edges in the computation graph, determining an input-output relationship of the operator by scanning the computation graph, analyzing the communication mode of operation between the operator and adjacent nodes, and adding the determination result to the directed acyclic graph G(O, E);Step 2.2: according to the data flow information of the operators, in combination with operator parameters and the communication mode, analyzing the data flow among the operators or performing communication evaluating modeling analysis on the operators, and automatically deriving an operator communication cost by using following equation:KCostcommurdcate = >x H2 x...x Hk) X sizeof(Htype) x T(Ctype) i~lwhere K indicates s number of the output tensors, and T(Ctype) indicates a multiple of communication traffic generated when the communication mode is Ctype.

4. The method according to claim 3, wherein Step 3 comprises:Step 3.1: according to the computing and storage evaluating modeling, the data flow among the operators and the communication evaluating model, obtaining the computing costCostcompute, the memory cost Costmemory and the communication cost Costcommunlcate of each operator, and searching for k operators with a highest cost automatically from the computation graph according to above three costs;Step 3.2: segmenting and placing k operators with the highest computational cost found on different devices, and placing k operators with the highest communication cost found on the same device;alternatively, inputting the computing cost, the memory cost and the communication cost automatically derived from the computation graph into a model parallel strategy to divide the model.