A Model Parallel Method, System, Device and Medium in an Edge Computing Scenario
By adopting flow-through model parallel technology and joint optimization algorithm in edge computing scenarios, the problem of too long training time in deep neural networks is solved, and efficient edge device training is achieved.
Patent Information
- Application Number
- CN202310281020.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-22
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-03-22
AI Technical Summary
In edge computing scenarios, the training of deep neural networks is limited by insufficient memory resources and computing resources, which leads to too long training time and is difficult to achieve online training.
Using the parallel technology of flow model, through the joint optimization of computing time delay and communication time delay, a flow model splitting scheme with training time optimization is generated, and appropriate edge devices are selected to participate in training to shorten the training time.
It realizes efficient training of deep neural networks in edge computing scenarios, shortens training time and saves human and material resources.
Smart Images

Figure CN116302539B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning distributed training, and particularly relates to a model parallel method, system, device and medium in an edge computing scenario. Background Art
[0002] The Internet of Things (IoT) is regarded as the third wave of the world's information industry after computers and the Internet. In 2018, the number of global IoT connections had reached approximately 8 billion, and these IoT devices generate a large amount of data. In the traditional cloud computing architecture, this data needs to be centrally transmitted to the cloud for processing, which will increase the network load, causing transmission congestion and data processing delays. In this case, a new type of computing architecture - edge computing - emerged as the times require. Edge computing refers to providing computing, storage and other services in the vicinity of things or the data source. Compared with the data centralized cloud computing model, edge computing processes data at the edge of the network, which can reduce the network bandwidth load, reduce the request response time and improve the battery life. At the same time, edge computing can ensure the security and privacy of data because it is performed on the side close to the data source. This advantage has led to an increasing demand for using deep learning technology on edge network devices. This is because as the amount of training data increases, the models trained by deep learning will become better, but in many fields, laws do not allow sharing of personal-related data. Using deep learning technology on edge devices can solve this dilemma. Training deep neural networks requires a large amount of computing resources, such as the computing power, memory and video memory of computing devices. The memory of edge devices is not large, which may make it difficult to train deep neural networks; and the computing power of edge devices is insufficient, which will lead to too long training time for deep neural networks. Solving the above difficulties is of great significance for training deep neural network models in edge computing scenarios.
[0003] Existing model parallel methods study how to accelerate the training of large-scale deep neural networks through model parallel technology on computing devices with sufficient memory resources and computing power resources, such as cloud devices; without considering how to achieve online training of deep neural networks on computing devices with insufficient memory resources and computing power resources, such as edge devices.
[0004] Julien Herrmann proposed a graph partitioning method in his paper "Acyclic Partitioning of Large Directed Acyclic Graphs" (IEEE, 2017) to partition a large directed acyclic graph into multiple acyclic partitions, that is, the dependency relationships between all different partitions do not form loops. In this method, a multi-layer method consisting of coarsening, initial partitioning, and refinement phases is adopted for the acyclic partitioning of the directed acyclic graph, and a method for directly partitioning the directed acyclic graph into k partitions is developed. The disadvantage of using this method to generate a model splitting scheme is that when partitioning the model shards, the balance of the computational workload of each shard is not considered, resulting in a longer model iteration time. Summary of the Invention
[0005] In order to overcome the above-mentioned deficiencies of the prior art, the object of the present invention is to propose a model parallel method, system, device, and medium in an edge computing scenario. By using the pipelined model parallel technology, it can jointly optimize the computational delay and communication delay in model parallel training, obtain a pipelined model splitting scheme with optimized training time, and at the same time, appropriate edge devices can be selected to participate in the training, which has the characteristics of shortening the training time and saving manpower and material resources.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A model parallel method in an edge computing scenario, the specific steps are as follows:
[0008] Step S1, calculate the comprehensive score of the computing and storage capabilities of all edge devices that may participate in the training;
[0009] Step S2, estimate the parameters of the neural network model to be trained, and the parameters include the computational workload of each neural network layer in the model and the communication transmission volume between each neural network layer in the model;
[0010] Step S3, obtain all topological sorts of the neural network model to be trained;
[0011] Step S4, for each topological sort of the neural network, use a pipelined model splitting algorithm that jointly optimizes the computational delay and communication delay to find the optimal partitioning solution of the model under this topological sort;
[0012] Step S5, select the solution with the shortest model iteration time among the optimal partitioning solutions of all topological sorts as the final model splitting scheme.
[0013] The specific operation of the said step S1 is as follows:
[0014] Step S1.1, obtain the computing power and memory size of all edge devices that may participate in model parallel training;
[0015] Step S1.2: Based on the computing power and memory size of the edge devices obtained in Step S1.1, conduct a comprehensive evaluation of the computing and storage capabilities of the devices;
[0016] Step S1.3: When selecting edge devices for model parallel training, select training devices to join in descending order of the comprehensive evaluation score.
[0017] The specific operation of conducting a comprehensive evaluation of the computing and storage capabilities of the devices in Step S1.2 is as follows:
[0018] The memory of n edge devices is {r1, r2,..., r n}, and the computing power is {s1, s2,..., s n}; the maximum memory r max of the edge device is:
[0019] r max = max(r i ), i ∈ n;
[0020] The maximum computing power s max of the edge device is:
[0021] s max = max(s i ), i ∈ n;
[0022] Normalize the memory of n edge devices to obtain the normalized device memory {r'1, r'2,..., r' n}:
[0023]
[0024] Normalize the computing power of n edge devices to obtain the normalized device computing power {s'1, s'2,..., s' n}:
[0025]
[0026] The comprehensive evaluation score I i of the computing and storage capabilities of the i-th edge device is:
[0027] I i = 0.5 × r' i + 0.5 × s' i .
[0028] The calculation amount of the neural network layer in Step S2, that is, the calculation amount of the basic layer of the convolutional neural network, is estimated as follows:
[0029] Convolutional layer: The input tensor size of the convolutional layer is n batch×c in ×h in ×w in ,n batch represents the training batch size, c in is the number of channels of the input feature map, h in is the height of the input feature map, w in is the width of the input feature map; the size of the square two-dimensional convolutional kernel is k×k, where k is the height and width of the convolutional kernel; the size of the output tensor of the convolutional layer is n batch ×c out ×h out ×w out ,c out is the number of channels of the output feature map, h out is the height of the output feature map, w out is the width of the output feature map; the computational complexity of the convolutional layer is:
[0030] FLOPs conv =2n batch ×k×k×c in ×c out ×h out ×w out ;
[0031] ReLU layer: The computational complexity of the ReLU layer is:
[0032] FLOPs ReLU =n batch ×c in ×h in ×w in ;
[0033] Pooling layer: The computational complexity of the pooling layer is:
[0034] FLOPs pooling =n batch ×c in ×h in ×w in ;
[0035] Fully connected layer: The computational complexity of the fully connected layer is:
[0036] FLOPs fc =2n batch ×c in ×c out ;
[0037] The specific method for estimating the communication transmission volume between each neural network layer in step S2 is as follows:
[0038] Inter-layer communication transmission volume: The inter-layer communication transmission volume is the size of the output tensor of each neural network layer, which is:
[0039] D = n batch × c out × h out × w out 。
[0040] All topological sorts of the neural network model to be trained are obtained in step S3, and the specific method is as follows:
[0041] Step S3.1: Initially, the status of all nodes is "unvisited".
[0042] Step S3.2: In each round of search, arbitrarily select an "unvisited" node u, start a depth-first search from node u, update the status of node u to "visiting", and for each node v adjacent to node u, judge the status of node v and perform the following operations:
[0043] 1) If the status of node v is "unvisited", continue to search node v;
[0044] 2) If the status of node v is "visiting", a cycle in the directed graph is found, so there is no topological sort;
[0045] 3) If the status of node v is "visited", node v has been searched and added to the output sorted list, and node u has not been searched yet, so the topological order of node u must be before that of node v, and no operation needs to be performed;
[0046] Step S3.3: When the status of all adjacent nodes of node u is "visited", update the status of node u to "visited" and add node u to the output sorted list;
[0047] Step S3.4: After all nodes have been visited, if no cycle in the directed graph is found, there is a topological sort, and the order of all nodes from the top to the bottom of the stack is the topological sort.
[0048] The specific operation of step S4 is as follows:
[0049] Step S4.1: Select the current optimal solution set;
[0050] Step S4.2: Use a local search algorithm to generate the next generation of solution sets from the optimal solution set;
[0051] Step S4.3: If the next generation of solution sets is empty, stop the iteration;
[0052] Step S4.4: Mix the next generation of solution sets with the optimal solution set, and retain a certain proportion of solutions with high fitness as the current optimal solution set;
[0053] Step S4.5: If the number of iterations has not reached the stop condition, repeat Steps S4.1 - S4.5; otherwise, stop the iteration.
[0054] The specific operation of the said Step S5 is as follows:
[0055] Step S5.1: For each optimal partitioning solution under a topological sorting, obtain the model iteration time of the model splitting scheme corresponding to this optimal partitioning solution through code simulation. The specific operations are as follows:
[0056] Step S5.1.1: Abstract the pipeline devices. A single pipeline device class contains attributes such as device computing power, device memory, device communication bandwidth, and the shard number of the bearing model.
[0057] Step S5.1.2: At the iteration initialization, use a shared dictionary as the database of completed tasks to store the information of completed tasks.
[0058] In the iteration: The main thread creates a shared dictionary and sequentially switches to the sub - threads to run the simulation of the pipeline devices; each sub - thread corresponds to a pipeline computing device. The sub - thread reads the database of completed tasks and judges whether the predecessor tasks of the task segment it is about to perform have been completed: If completed, simulate the communication process to obtain the start time and end time of the execution of this task, write the serial number, start time, and end time of this task segment into the database of completed tasks, and then the thread enters the sleep state; if not completed, the thread will directly enter the sleep state; when the main thread observes that the sub - thread is in the sleep state, it switches to the next sub - thread and waits for the next round of task execution.
[0059] Step S5.2: Select the optimal partitioning solution with the shortest iteration time as the final model splitting scheme.
[0060] A model parallel system in an edge computing scenario includes:
[0061] A comprehensive scoring module, used to calculate the comprehensive scores of the computing and storage capabilities of all edge devices that may participate in training.
[0062] A parameter module, used to obtain the parameters of the neural network model to be trained.
[0063] A topological sorting module, used to obtain all topological sortings of the neural network model to be trained.
[0064] An optimal partitioning solution module, used to use a water - flow model splitting algorithm that jointly optimizes the computing delay and communication delay to find the optimal partitioning solution of the model under each topological sorting.
[0065] A model parallel device in an edge computing scenario includes:
[0066] A memory for storing required data;
[0067] A processor for implementing the model parallel method in an edge computing scenario described in steps S1 to S5 when executing the computer program.
[0068] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the model parallel method in an edge computing scenario.
[0069] Compared with the prior art, the present invention has the following advantages:
[0070] 1. In step S1 of the present invention, a comprehensive score is given to the computing and storage of all edge devices that may participate in training. When selecting training devices, those with high scores are selected to participate in training, which can not only ensure that the selected devices have a large amount of storage, increasing the probability of completing model parallel training of a deep neural network with fewer selected devices, but also ensure that the selected devices have a large computing power, enabling them to spend less time during training.
[0071] 2. In step S4 of the present invention, the optimal distribution characteristic of the model shard training duration that minimizes the computing time is added to the optimization objective, which can make the division of the computing amount of each model shard more balanced compared to the acyclic graph partitioning method, reducing the iteration time of the model.
[0072] 3. Step S4 of the present invention provides the optimal distribution characteristic of the model shard training duration that minimizes the computing time in the total time of model parallel training, and takes the difference between the actual computing amount of each model shard and the optimal distribution characteristic as an optimization objective, which can make the splitting scheme change in the direction of reducing the computing time during iteration, optimizing the computing delay in model parallel training.
[0073] In summary, compared with the prior art, the present invention can more likely select devices with large memory to participate in training when selecting edge devices participating in training, reducing the iteration time of the model for pipelined model parallel training, and has the characteristics of shortening the training time and saving manpower and material resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 It is a flowchart for implementing the present invention.
[0075] Figure 2 It is a flowchart of the pipelined model splitting algorithm for joint optimization of computing delay and communication delay in the present invention.
[0076] Figure 3 It is a comparison chart of the model iteration time between the model parallel method provided by the present invention and the model splitting scheme generated by the acyclic graph partitioning method. DETAILED DESCRIPTION OF THE INVENTION
[0077] The present invention will be described in detail below with reference to the accompanying drawings.
[0078] Referring to Figure 1 , a model parallel method in an edge computing scenario is as follows:
[0079] Step S1: Calculate the comprehensive scores of the computing and storage capabilities of all edge devices that may participate in training. The specific operations are as follows:
[0080] Step S1.1: Obtain the computing power and memory size of all edge devices that may participate in model parallel training.
[0081] Step S1.2: According to the computing power and memory size of the edge devices obtained in Step S1.1, perform a comprehensive score for the devices' computing and storage capabilities. The specific operations are as follows:
[0082] The memory of n edge devices is {r1, r2,..., r n}, and the computing power is {s1, s2,..., s n}; the maximum memory r max of the edge devices is:
[0083] r max = max(r i ), i ∈ n;
[0084] The maximum computing power s max of the edge devices is:
[0085] s max = max(s i ), i ∈ n;
[0086] Normalize the memory of n edge devices to obtain the normalized device memory {r'1, r'2,..., r' n}:
[0087]
[0088] Normalize the computing power of n edge devices to obtain the normalized device computing power {s'1, s'2,..., s' n}:
[0089]
[0090] The comprehensive score I i of the computing and storage capabilities of the i-th edge device is:
[0091] I i = 0.5 × r' i + 0.5 × s' i ;
[0092] Step S1.3: When selecting edge devices for model parallel training, select training devices to join in the order from the highest to the lowest comprehensive score.
[0093] Step S2: Estimate the parameters of the neural network model to be trained, where the parameters include the computational volume of each neural network layer in the model and the communication transmission volume between each neural network layer in the model.
[0094] The computational volume of the neural network layer in Step S2, that is, the computational volume of the basic layer of the convolutional neural network, is estimated as follows:
[0095] Convolutional layer: The input tensor size of the convolutional layer is n batch ×c in ×h in ×w in , n batch represents the training batch size, c in is the number of input feature map channels, h in is the height of the input feature map, w in is the width of the input feature map; the size of the square two-dimensional convolutional kernel is k×k, and k is the height and width of the convolutional kernel; the output tensor size of the convolutional layer is n batch ×c out ×h out ×w out , c out is the number of output feature map channels, h out is the height of the output feature map, w out is the width of the output feature map; the computational volume of the convolutional layer is:
[0096] FLOPs conv = 2n batch ×k×k×c in ×c out ×h out ×w out ;
[0097] ReLU layer: The computational volume of the ReLU layer is:
[0098] FLOPs ReLU = n batch ×c in ×h in ×w in ;
[0099] Pooling layer: The computational volume of the pooling layer is:
[0100] FLOPs pooling = n batch ×c in ×h in ×w in ;
[0101] Fully connected layer: The computational complexity of the fully connected layer is:
[0102] FLOPs fc = 2n batch × c in × c out ;
[0103] The specific method for estimating the communication transmission volume between each neural network layer in step S2 is as follows:
[0104] Inter-layer communication transmission volume: The inter-layer communication transmission volume is the size of the output tensor of each neural network layer, which is:
[0105] D = n batch × c out × h out × w out ;
[0106] Step S3, obtain all topological sorts of the neural network model to be trained, specifically:
[0107] In the present invention, the neural network is abstracted as a directed acyclic graph. Therefore, the neural network layer is a node of the directed acyclic graph. The in-degree of this node is the number of neural network layers from all output tensors to this layer, and the out-degree of this node is the number of neural network layers that receive the output tensor of this layer; the basic idea of using the depth-first traversal algorithm to implement topological sorting is: for a specific node, if all adjacent nodes of this node have been searched, then this node will also become a searched node. In topological sorting, this node is in front of all its adjacent nodes; an adjacent node of a node refers to a node that can be reached from this node through a directed edge (in the background of the invention, it refers to the neural network layer that receives the output tensor of this layer); since the order of topological sorting is opposite to the order of search completion, a stack needs to be used to store all the searched nodes. During the depth-first search, the state of each node needs to be maintained. The state of each node may have three situations: unvisited, visiting, and visited; the specific method is as follows:
[0108] Step S3.1, initially, the state of all nodes is "unvisited";
[0109] Step S3.2, in each round of search, arbitrarily select an "unvisited" node u, and start depth-first search from node u. Update the state of node u to "visiting", and for each adjacent node v of node u, judge the state of node v and perform the following operations:
[0110] 1) If the state of node v is "unvisited", then continue to search node v;
[0111] 2) If the status of node v is "being visited", then a cycle is found in the directed graph, so there is no topological sorting;
[0112] 3) If the status of node v is "visited", then node v has been searched and added to the output sorted list, and node u has not been searched yet. Therefore, the topological order of node u must be before that of node v, and no operation needs to be performed;
[0113] Step S3.3: When the status of all adjacent nodes of node u is "visited", update the status of node u to "visited" and add node u to the output sorted list;
[0114] Step S3.4: After all nodes have been visited, if no cycle is found in the directed graph, there is topological sorting, and the order of all nodes from the top to the bottom of the stack is the topological sorting;
[0115] Step S4: For each topological sorting of the neural network, use the pipelining model splitting algorithm that jointly optimizes the computation delay and communication delay to find the optimal partitioning solution of the model under this topological sorting. Specifically:
[0116] The pipelining model splitting algorithm that jointly optimizes the computation delay and communication delay takes into account both the computation time and communication time during training and sets multiple optimization goals;
[0117] First, let the computing power of n computing devices be {s1, s2,..., s n}, and let the computational amounts of n model shards in one training micro-batch be {c1, c2,..., c n}. The training time of the i-th model shard on the i-th computing device is When the time taken for all model shards to train one micro-batch is the same, the maximum training duration of the model shards reaches the shortest; the optimal distribution characteristic of the training duration of the model shard with the shortest computation time is:
[0118]
[0119] Let the computational amounts of n model shards that satisfy the optimal distribution characteristic of the training duration of the model shards be At this time, Let the computational amounts of the actually partitioned n model shards be {c'1, c'2,..., c' n}, and the first optimization goal is set as:
[0120]
[0121] At the same time, in order to reduce the data transmission between model shards, the second optimization goal is set as the sum of the communication amounts between model shards, which can be expressed as:
[0122]
[0123] In pipelined model parallelism, data transfer between model shards is unidirectional. Therefore, it is necessary to ensure that there is no bidirectional connection between any two model shards, that is, no loops are generated:
[0124]
[0125]
[0126] Because the memory of edge devices may not be sufficient to complete the training of the divided model shards. Let the memory of n computing devices be {r1, r2,..., r n}, and the memory sizes required for training n model shards be {R1, R2,..., R n}, and it is necessary to satisfy:
[0127]
[0128] The total computational volume of the entire neural network and the total data transfer volume between layers are:
[0129]
[0130]
[0131] The total computing power S of all devices is:
[0132]
[0133] The pipelined model that jointly optimizes the computational delay and communication delay is split and modeled as the following optimization problem:
[0134]
[0135]
[0136]
[0137]
[0138] The objective function argmin minimizes the computational delay in pipelined model parallel training through the first term and minimizes the communication delay in pipelined model parallel training through the second term; where α and β respectively represent the proportions of the computational delay and communication delay in the model parallel training time; α and β are set as follows, where B represents the average bandwidth of the entire network:
[0139]
[0140]
[0141] After setting the optimization goal, the genetic algorithm is used to solve the problem. The design idea is as follows: Under each topological sorting, in order to find the optimal partitioning solution of the model, according to the optimal distribution characteristics of the training duration of the model shards with the minimum calculation time, the network is split to generate the initial partitioning solution as the first-generation parent generation, and then the local search algorithm is used to obtain the offspring. After mixing the parent generation and the offspring, a certain proportion of individuals with high fitness are retained as the next-generation parent generation for iteration. The local search algorithm will record the number of times each neural network layer moves between different partitioning shards, and based on this number, decide whether to discard this solution and not add it to the new offspring set. The algorithm terminates after reaching the iteration number or no new offspring are generated. At this time, the algorithm result is the optimal solution under this topological sorting. Refer to Figure 2 For step S4, the specific operation is as follows:
[0142] Step S4.1: Select the current optimal solution set;
[0143] Step S4.2: Use the local search algorithm to generate the next-generation solution set from the optimal solution set;
[0144] Step S4.3: If the next-generation solution set is empty, stop the iteration;
[0145] Step S4.4: Mix the next-generation solution set with the optimal solution set, and retain a certain proportion of solutions with high fitness as the current optimal solution set;
[0146] Step S4.5: If the iteration number has not reached the stop condition, repeat steps S4.1 - S4.5, otherwise stop the iteration;
[0147] The design idea of the local search algorithm for generating subsets is as follows: First, find all movable vertices in each shard. In the present invention, the vertices are the neural network layers in the model shards. The movable vertices in model shard P i need to meet at least one of the following two conditions: can move forward, or the vertex has no predecessor nodes, or all predecessor nodes are not in the same shard as this vertex; can move backward, or the vertex has no successor nodes, or all successor nodes are not in the same shard as this node; the algorithm determines whether it can move by analyzing the partitioning shards to which its predecessor nodes and successor nodes belong. If it can move, further determine its moving range between model shards; after finding all movable nodes, the algorithm decides whether to move this node to another partitioning block according to the number of times the movable node has been moved; the more times this point has been moved, the more likely it is to discard this moving operation; this step is to avoid a node being moved too many times and narrow the search space of the algorithm by means of probabilistic discard;
[0148] Step S5. Select the solution with the shortest model iteration time among all the optimal partitioning solutions obtained by topological sorting as the final model splitting scheme. The specific operation is as follows:
[0149] Step S5.1. For each optimal partitioning solution under a certain topological sorting, obtain the model iteration time of the model splitting scheme corresponding to this optimal partitioning solution through code simulation. The specific operation is as follows: In pipelined model parallelism, the operation of a single device is independent. Therefore, a multi-threaded approach can be used to simulate the execution of each pipelined device. During training, after a training micro-batch finishes training on the model shard on the current device, the current device sends the output tensor of the model shard on the device to the device where the next adjacent model shard is located. The start time for the device where the next adjacent model shard is located to start training this training micro-batch depends on the arrival time of the output tensor sent by the current device and the time when the device where the next adjacent model shard is located finishes its previous computational task. Therefore, the execution of a forward process or a backward propagation process depends on two conditions: (1) the device is idle; (2) the calculation and communication processes of the predecessor-dependent process are completed.
[0150] Accordingly, the computing devices can be abstracted. A single device class contains attributes such as device computing power, device memory, device communication bandwidth, and the model shard number it bears. A single thread is used to load and execute the tasks in the iteration. At the iteration initialization, a shared dictionary is used as a database for completed tasks to store the information of the tasks that have been carried out. During the iteration, the main thread creates the shared dictionary and sequentially switches to the child threads to run the simulation of the computing devices. Each thread corresponds to a pipelined computing device, reads the database of completed tasks, and determines whether the predecessor tasks of the task segment it is about to perform have been completed. If the execution conditions are met, simulate the calculation and communication processes to obtain the start time and end time stamp of the execution of this task. Finally, write the serial number, start time, and end time of this task segment into the database of completed tasks, and the thread enters the sleep state. If the execution conditions are not met, the thread will directly go to sleep. When the main thread observes that the child thread is sleeping, it switches to the next child thread and waits for the next round of task execution. Through this code simulation method, the model iteration time corresponding to the model splitting scheme can be obtained.
[0151] Step S5.2. Select the optimal partitioning solution with the shortest iteration time as the final model splitting scheme.
[0152] The following further illustrates the effect of the present invention in combination with simulation experiments:
[0153] 1. Simulation parameter settings:
[0154] In the simulation experiment, the ResNet-50 model is divided into 1 to 8 model shards, and the computing power and memory of 8 edge devices are set as shown in Table 1:
[0155] Table 1 Computing power and memory settings of eight edge devices
[0156]
[0157] The simulation parameters used for training are shown in Table 2:
[0158] Table 2 Training simulation parameter settings
[0159] Parameter Name Parameter Setting Device Bandwidth 10Gbps Iterative Microbatch Number 16 Microbatch Size 32 Maximum Shard Number 8 Split Network ResNet-50 Model
[0160] 2. Simulation content and its result analysis:
[0161] Under the above simulation parameters, the model splitting schemes of the ResNet-50 model divided into 1 to 8 model shards are generated by the method of the present invention and the acyclic graph partitioning method respectively. The model iteration times of the two schemes are as Figure 3 .
[0162] It can be seen from Figure 3 that when the ResNet-50 model is divided into 1 to 8 model shards, the iteration time of the method of the present invention is less than that of the acyclic graph partitioning method.
[0163] The above simulation results show that when setting the optimization goal of the heuristic model splitting algorithm in the present invention, the optimal distribution characteristic of the training duration of the model shard that minimizes the calculation time is taken into account, making the calculation amount division of each model shard more balanced compared with the acyclic graph partitioning method, and reducing the iteration time of the model for pipelined model parallel training.
Claims
1. A model parallel method in an edge computing scenario, characterized in that, The specific steps are as follows: Step S1: Calculate the comprehensive scores of the computing and storage capabilities of all possible edge devices participating in the training; Step S2: Estimate the parameters of the neural network model to be trained, where the parameters include the computational amount of each neural network layer in the model and the communication transmission amount between each neural network layer in the model; Step S3: Obtain all topological sorts of the neural network model to be trained; Step S4: For each topological sort of the neural network, use the pipelining model splitting algorithm jointly optimized by computational latency and communication latency to find the optimal partitioning solution of the model under this topological sort; Step S5: Select the solution that minimizes the model iteration time among the optimal partitioning solutions of all topological sorts as the final model splitting scheme; The specific operation of the said Step S4 is as follows: Step S4.1: Select the current optimal solution set; Step S4.2: Adopt a local search algorithm to generate the next generation of solution sets from the optimal solution set; Step S4.3: If the next generation of solution sets is empty, stop the iteration; Step S4.4: Mix the next generation of solution sets with the optimal solution set, and retain a certain proportion of solutions with high fitness as the current optimal solution set; Step S4.5: If the number of iterations does not reach the stop condition, repeat Steps S4.1 - S4.5, otherwise stop the iteration; The design idea of the local search algorithm in step S4.2 is as follows: First, find all movable vertices in each shard. The vertices are the neural network layers in the model shards and are located in the model shards The movable vertices need to meet at least one of the following two conditions: being movable forward, or the vertex having no predecessor nodes, or all predecessor nodes not being in the same shard as this vertex; being movable backward, or the vertex having no successor nodes, or all successor nodes not being in the same shard as this node. The algorithm determines whether a node can be moved by analyzing the partition shards to which its predecessor nodes and successor nodes belong. If it can be moved, it further determines the range of its movement among the model shards. After finding all movable nodes, the algorithm decides whether to move a node to another partition block according to the number of times the movable node has been moved. The more times this point has been moved, the more likely it is to discard this movement operation.
2. The model parallel method in an edge computing scenario according to claim 1, wherein The specific operation of the said Step S1 is as follows: Step S1.1: Obtain the computing power and memory size of all possible edge devices participating in model parallel training; Step S1.2: Based on the computing power and memory size of the edge devices obtained in Step S1.1, perform a comprehensive score for the devices' computing and storage capabilities; Step S1.3: When selecting edge devices for model parallel training, select training devices to join in descending order of the comprehensive score.
3. The model parallel method in an edge computing scenario according to claim 2, wherein, The specific operation of performing a comprehensive score for the devices' computing and storage capabilities in the said Step S1.2 is as follows: The memory of the edge device is , and the computing power is ; the maximum memory of the edge device is: ; The maximum computing power of the edge device is as follows: ; For normalize the memory of the edge device to obtain the normalized device memory : ; For the computing power of the edge device is normalized to obtain the normalized device computing power : ; The comprehensive score of the computing and storage capabilities of the first edge device is as follows: 。 4. A model parallel method in an edge computing scenario according to claim 1, wherein The computational amount of the neural network layer in the said Step S2, that is, the computational amount of the basic layer of the convolutional neural network, is estimated as follows: Convolutional layer: The input tensor size of the convolutional layer is , represents the training batch size, is the number of input feature map channels, is the input feature map height, is the width of the input feature map; the size of the square two-dimensional convolution kernel is , is the height and width of the convolution kernel; the output tensor size of the convolution layer is , is the number of output feature map channels, is the output feature map height, is the width of the output feature map; the computational complexity of the convolutional layer is: ; ReLU layer: The computational amount of the ReLU layer is: ; Pooling layer: The computational amount of the pooling layer is: ; Fully connected layer: The computational amount of the fully connected layer is: ; The specific method for estimating the communication transmission amount between each neural network layer in the said Step S2 is as follows: Inter-layer communication transmission amount: The inter-layer communication transmission amount is the output tensor size of each neural network layer, which is: 。 5. A model parallel method in an edge computing scenario according to claim 1, wherein, The specific method for obtaining all topological sorts of the neural network model to be trained in the said Step S3 is as follows: Step S3.1: Initially, the status of all nodes is "unvisited"; Step S3.2: In each round of search, arbitrarily select an "unvisited" node u, start a depth-first search from node u, update the status of node u to "visiting", and for each node v adjacent to node u, judge the status of node v and perform the following operations: 1) If the status of node v is "unvisited", then continue to search node v; 2) If the status of node v is "visiting", then a cycle is found in the directed graph, so there is no topological sort; 3) If the status of node v is "visited", then node v has been searched and added to the output sorted list. Node u has not been searched yet. Therefore, the topological order of node u must be before that of node v, and no operation needs to be performed; Step S3.3: When the status of all adjacent nodes of node u is "visited", update the status of node u to "visited" and add node u to the output sorted list; Step S3.4: After all nodes have been visited, if no cycle is found in the directed graph, there is a topological sort, and the order of all nodes from the top to the bottom of the stack is the topological sort.
6. The model parallel method in an edge computing scenario according to claim 1, wherein The specific operation of step S5 is as follows: Step S5.1: For the optimal partitioning solution under each topological sort, obtain the model iteration time of the model partitioning scheme corresponding to this optimal partitioning solution through code simulation. The specific operation is as follows: Step S5.1.1: Abstract the pipeline device. A single pipeline device class contains attributes such as device computing power, device memory, device communication bandwidth, and the shard number of the bearing model. Step S5.1.2: At the iteration initialization, use a shared dictionary as the database of completed tasks to store the information of completed tasks. Step S5.1.3: During the iteration, the main thread creates a shared dictionary and sequentially switches to the child threads to run the simulation of the pipeline device; each child thread corresponds to a pipeline computing device. The child thread reads the database of completed tasks and determines whether the predecessor tasks of the task segment it is about to perform have been completed: if completed, simulate the communication process to obtain the start time and end time of the execution of this task, write the serial number, start time, and end time of this task segment into the database of completed tasks, and then the thread goes to sleep; if not completed, the thread will directly go to sleep; when the main thread observes that the child thread is sleeping, it switches to the next child thread and waits for the next round of task execution; Step S5.2: Select the optimal partitioning solution with the shortest iteration time as the final model partitioning scheme.
7. A model parallel system in an edge computing scenario, which implements a model parallel method in an edge computing scenario as described in any one of claims 1 to 6, characterized in that, Including: A comprehensive scoring module for calculating the comprehensive scores of the computing and storage capabilities of all edge devices that may participate in training; A parameter module for obtaining the parameters of the neural network model to be trained; A topological sorting module for obtaining all topological sorts of the neural network model to be trained; An optimal partitioning solution module for selecting the current optimal solution set, using a local search algorithm to generate the next generation solution set from the optimal solution set, mixing the next generation solution set with the optimal solution set, and retaining a certain proportion of solutions with high fitness as the current optimal solution set.
8. A model parallel device in an edge computing scenario, characterized in that, Including: A memory for storing the required data; A processor for executing a model parallel method in an edge computing scenario as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it can implement a model parallel method in an edge computing scenario as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Deep learning model training acceleration method based on end-to-edge cloud cooperation
CN111242282A
Hybrid pipeline parallel method for accelerating distributed deep neural network training
CN112784968A