A pipeline parallel GPU configuration method and system in an artificial intelligence system
By generating new working partitions based on static and dynamic indicators before the next training of the neural network, and optimizing GPU configuration using meta-network and reinforcement learning, the problem of fixed GPU allocation scheme in shared GPU clusters is solved, and GPU resource utilization and training efficiency of neural networks are improved.
Patent Information
- Application Number
- CN202210797455.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-07-08
AI Technical Summary
In the scenario of shared GPU cluster, in the distributed training system, the GPU allocation scheme is fixed in pipeline parallelism, and it is impossible to follow the actual GPU resource adjustment in time, resulting in low utilization of GPU resource, which affects the performance of pipeline parallelism method and neural network training efficiency.
Provides a GPU configuration method for pipeline parallelism in artificial intelligence systems. By generating new working partitions based on static indicators and dynamic indicators before the next training of the neural network, adding the available bandwidth of each GPU, dynamically reflecting the available resources of the GPU. Use metanet to predict the training speed of each working partition, and use reinforcement learning to determine whether to update the current working partition.
It effectively solves the problem of fixed GPU allocation scheme, realizes more reasonable distributed training, improves GPU resource utilization, and ensures the training efficiency of neural networks.
Smart Images

Figure CN115033388B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep neural network training and is applied to the field of artificial intelligence. In particular, it relates to a pipeline-parallel GPU configuration method and system in an artificial intelligence system, and more particularly, the GPU configuration method is applied to distributed training of neural networks. Background Art
[0002] In the field of artificial intelligence (AI), deep neural network (DNN) training usually uses accelerators for calculations, and the calculation process generally includes forward propagation, directional propagation, and gradient descent. Through the continuous iteration of these three steps, the neural network eventually reaches a convergence state. However, with the continuous increase of data sets and deep learning models, the number of model layers is increasing, which requires a large amount of computing and storage resources to be consumed in the process of training deep neural networks (DNN), that is, the required computing power and video memory are increasing. In order to reduce the time of deep learning training, distributed deep learning that performs cluster parallel computing on multiple machines has gradually become the focus of technological innovation and development.
[0003] Common distributed deep learning technologies are data parallelism, model parallelism, and pipeline parallelism. In data parallelism, the data set is partitioned, that is, each GPU (graphics processing unit) calculates a part of the data set and a copy of the model parameters is maintained on each GPU. In model parallelism, the model is partitioned, that is, each GPU has a different layer of the model, and a copy of the data set is maintained on each GPU. The existing pipeline parallelism makes up for the shortcomings of data parallelism and model parallelism. Pipeline parallelism is based on model parallelism. The model is also partitioned / layered, and a GPU is assigned to each layer to process the data at that level; however, unlike model parallelism, only one GPU is working at the same time. In pipeline parallelism, the GPU of the previous layer can process different data at the same time as the GPU of the next layer, thus reducing the waste of resources. For example, by splitting a data subset into multiple micro-batches (multiple groups of data), during the forward propagation of the model, after a GPU completes the calculation of a group of data, it no longer waits for back propagation, but continues to calculate the forward propagation process of the next group of data. At this time, the GPU of the previous layer and the GPU of the next layer process different data at the same time, thereby greatly reducing the idle time of the GPU and providing parallel efficiency.
[0004] As can be seen from the above, the performance of pipeline parallelism in distributed training technology is highly related to work partitioning. The role of work partitioning is to determine how to allocate the calculation of each layer to the GPU, that is, to configure several GPUs to each layer of the model, and the same GPU can be responsible for multiple layers; however, in the existing pipeline parallel technology, when the training starts, the work partitioning will be performed according to the existing GPUs and bandwidth, and will remain fixed afterwards. However, in actual applications, the DNN training process is very time-consuming, often taking hours or days; at the same time, the GPU cluster in the current DNN training often shares the GPU with other jobs. Therefore, during the DNN training process, the available resources of the GPU are very likely to fluctuate, because during the long model training process, other shared GPU jobs may start, end or pause, which will cause fluctuations in GPU resources, and other non-DNN training jobs will also affect the fluctuation of bandwidth, which greatly affects the running performance of pipeline parallelism, especially when the available resources cannot meet the previously set work partitions, which may cause the previous work partition solution to be outdated and unable to adjust in time with the actual GPU resources, affecting the performance of the pipeline parallel method. Summary of the invention
[0005] The purpose of the present invention is to solve the problem that in the scenario of shared GPU cluster, the GPU allocation scheme in the pipeline parallel is fixed, which makes it impossible to adjust the GPU allocation in time with the actual GPU resources, resulting in low GPU resource utilization, affecting the performance of the pipeline parallel method and the efficiency of neural network training, and then provide a pipeline parallel GPU configuration method and system in an artificial intelligence system, and apply the GPU configuration method to neural network training. The GPU configuration method provided by the technical solution of the present invention obtains many new work partition schemes according to static indicators and dynamic indicators before the next neural network training, and the available bandwidth of each GPU is added to the dynamic indicator, so that the newly generated many work partitions dynamically reflect the available resources of the GPU; further introduces the meta-network to predict the training speed of each work partition scheme to screen the work partition, and introduces reinforcement learning to determine whether to update the current work partition with the screened work partition. In summary, the technical solution of the present invention effectively solves the problem of fixed GPU allocation scheme through the above-mentioned GPU configuration method, can realize more reasonable distributed training, effectively improve GPU resource utilization, and thus ensure the training efficiency of subsequent neural networks.
[0006] On the one hand, the present invention provides a pipeline parallel GPU configuration method in an artificial intelligence system, wherein the GPU configuration method is based on a shared GPU cluster for updating or maintaining the configuration relationship between the current GPU and the network layer before the next training of the neural network, wherein there is a shared GPU in the GPU cluster, and the GPU configuration method comprises the following steps:
[0007] Step 1: Obtain the current static and dynamic indicators in the distributed training system;
[0008] The distributed training system adopts pipeline parallelism, which configures a GPU for each network layer of the neural network. The same GPU is responsible for the training of one or more network layers. The static indicators include: the number of network layers of the neural network, the number of GPUs, and the training characteristics of each network layer; the dynamic indicators include: the available bandwidth of each GPU, the forward propagation time and the backward propagation time of the network layer that each GPU is responsible for;
[0009] Step 2: Generate a number of new work partitions according to the static indicators and dynamic indicators, wherein the work partition represents the configuration relationship between the GPU and the network layer;
[0010] Step 3: Taking the static index, dynamic index, and new work partition as input, using the work partition training speed prediction model to obtain the training speed prediction value corresponding to each new work partition; wherein the input of the work partition training speed prediction model is the static index, dynamic index, and new work partition; the output is the training speed prediction value of the new work partition;
[0011] Step 4: Filter the new work partitions based on the predicted training speed of each new work partition and the number of GPUs whose configuration relationships have changed in each new work partition;
[0012] Step 5: Taking the static indicators, dynamic indicators, the screened new working partition and the current working partition as input, a screening model based on reinforcement learning is used to determine whether to replace the current working partition, that is, to update or maintain the current configuration relationship between the GPU and the network layer.
[0013] In the GPU cluster provided by the technical solution of the present invention, there is a GPU that performs DNN training operations and other operations at the same time, which leads to the possibility that the available resources of the GPU in the DNN training process will fluctuate greatly, and finally affect the utilization rate of the GPU. For this reason, before each training except the first training, if the available resources of the GPU change, the GPU configuration method provided by the technical solution of the present invention will update the configuration relationship between the GPU and the network layer or keep the current configuration relationship between the GPU and the network layer unchanged according to the GPU configuration method of the present invention, that is, before the next training of the neural network, the configuration relationship between the GPU and the network layer is optimized so that it can dynamically match the available resources of the GPU, thereby improving the utilization rate of the GPU. Among them, by constructing a working partition training speed prediction model, the training speed prediction value corresponding to each working partition is obtained, and the training speed prediction value is used as a screening basis to ensure that the training speed of the neural network corresponding to the screened working partition is faster; the number of GPUs with changed configuration relationships is used as another screening basis, which effectively reduces the screening range and reduces the time complexity; finally, reinforcement learning is introduced to determine which of the new working partition and the original working partition is more adapted to the current environment, so as to select the best working partition for the next training of the neural network, and migrate to the optimal partition through the reinforcement learning model.
[0014] For the working partition training speed prediction model, the present invention simultaneously utilizes static indicators, dynamic indicators and new working partitions, and sets them as model inputs. Among them, in order to monitor the changes in the computing power of the GPU, considering that it may be responsible for the training of multiple network layers and responsible for different network layers in different training tasks, the forward propagation and backward propagation time layer by layer are embedded in the feature space as dynamic indicators, so that the trained working partition training speed prediction model is more robust. Secondly, in order to fully understand the computing power of the GPU, the available bandwidth of the GPU is also embedded in the feature space, and the reliability of the evaluation results is effectively improved by combining dynamic and static indicators.
[0015] In summary, the technical solution of the present invention takes into account the dynamic changes of the available resources of the GPU, so that the GPU allocation scheme for each training is no longer fixed, but dynamically matched with the available resources of the current GPU, effectively improving the utilization rate of the GPU and making distributed training more reasonable; furthermore, the meta-network is used to obtain the training speed prediction value, so as to screen out the working partitions with faster preset training speeds and use reinforcement learning to automatically determine whether the working partitions need to be replaced, further ensuring that the working partitions selected for the next training of the neural network are more reasonable and appropriate.
[0016] Further optionally, the work partition training speed prediction model in step 3 is constructed based on a meta-network, and the meta-network includes: each group of static indicators, an embedding layer corresponding to the dynamic indicators, an LSTM network corresponding to the dynamic indicators, and a fully connected layer, wherein the embedding layer corresponding to the dynamic indicators is connected to the LSTM network, and the embedding layer corresponding to the static indicators and the LSTM network are connected to the fully connected layer;
[0017] The dynamic indicator is input into the corresponding embedding layer, and the obtained output result is input into the LSTM network to obtain the sequence characteristics of the dynamic indicator;
[0018] The static indicator is input into the corresponding embedding layer, and the output result, the sequence feature and the new working partition are used as the input of the fully connected layer, and the output of the fully connected layer is the training speed prediction value of the new working partition.
[0019] The embedding layer, LSTM network and fully connected layer in the technical solution of the present invention are all existing network architectures. The present invention does not provide a specific introduction to them. It should be understood that the size of the network is determined according to the size of each set of input features to ensure that the output of the subsequent fully connected layer is a training speed prediction value.
[0020] Further optionally, the work partition training speed prediction model and the screening model are both constructed through offline training;
[0021] Among them, the goal of the reward function in the screening model is to make the training speed corresponding to the working partition selected by the screening model greater than the training speed corresponding to the previous working partition.
[0022] It should be understood that how the reward function in reinforcement learning participates in training and how it is set is already existing technology. The present invention does not optimize it. It only sets its goal to make the training speed corresponding to the working partition selected by the screening model greater than the training speed corresponding to the previous working partition according to the needs of this application when it is applied to this scenario. That is, when the training speed corresponding to the selected working partition is greater than the training speed corresponding to the previous working partition, the output of the screening model is regarded as positive; thereby, the actual meaning of the reward function in reinforcement learning can be adaptively defined according to this goal.
[0023] Further optionally, when screening new work partitions in step 4, the following two rules need to be met at the same time;
[0024] Rule 1: Only two GPUs in the selected new work partition have changed configurations, which are the configuration relationships between the GPU and the network layer.
[0025] Rule 2: The predicted training speed of the selected new working partition is higher than the training speed of the current working partition.
[0026] Further optionally, the training features of each network layer include the output activation size, weight parameters and gradient of the network layer.
[0027] In a second aspect, the present invention provides a neural network distributed training method based on a GPU configuration method, which comprises the following steps:
[0028] Step S1: loading the neural network to be trained and the data set into the distributed training system, and dividing the neural network into multiple network layers;
[0029] Step S2: Initialize the working partition and perform the first training of the neural network;
[0030] Wherein, a GPU is configured for each network layer of the neural network, and the same GPU is responsible for training one or more network layers, that is, the configuration relationship between the GPU and the network layer is determined according to the initialization work partition, the GPU uses the data set to train the neural network, and the distributed training system uses a distributed communication mechanism for communication connection;
[0031] Step S3: Before the next training of the neural network, determine the working partition corresponding to the next training of the neural network according to the method of steps 1 to 5 and then perform training; wherein, if a new working partition is obtained, update the configuration relationship between the GPU and the network layer according to the new working partition;
[0032] Step S4: Determine whether the iterative training termination condition of the neural network is met. If not, return to step S3 to continue training; otherwise, complete the training of the neural network.
[0033] In a third aspect, the present invention provides a GPU allocation device based on the GPU configuration method, comprising:
[0034] The dynamic and static indicator acquisition module is used to obtain the current static and dynamic indicators in the distributed training system;
[0035] The distributed training system adopts pipeline parallelism, which configures a GPU for each network layer of the neural network. The same GPU is responsible for the training of one or more network layers. The static indicators include: the number of network layers of the neural network, the number of GPUs, and the training characteristics of each network layer; the dynamic indicators include: the available bandwidth of each GPU, the forward propagation time and the backward propagation time of the network layer that each GPU is responsible for;
[0036] A configuration module, used to generate a number of new work partitions according to the static indicators and the dynamic indicators; wherein the new work partitions represent the configuration relationship between the GPU and the network layer;
[0037] A training speed prediction value acquisition module is used to take the static index, dynamic index, and new working partition as input, and use the working partition training speed prediction model to obtain a training speed prediction value corresponding to each new working partition;
[0038] A screening module, used for screening new working partitions based on a predicted value of a training speed of each new working partition and the number of GPUs whose configuration relationship has changed in each new working partition;
[0039] The decision module is used to take the static indicators, dynamic indicators, the screened new working partition and the current working partition as inputs, and use the screening model based on reinforcement learning to determine whether to replace the current working partition, that is, to update or maintain the current configuration relationship between the GPU and the network layer.
[0040] In a fourth aspect, the present invention provides a distributed training system based on the GPU configuration method or the training method, wherein the distributed training system comprises at least: a plurality of GPU servers, each of which is provided with a GPU, a CPU, a memory, a network card and a switch;
[0041] Among them, the GPU is used to implement network training; the memory is used to store data; the CPU, network card and switch are used to implement data transmission, and distributed communication is adopted between the GPU servers.
[0042] In a fifth aspect, the present invention provides an electronic device, comprising:
[0043] one or more processors;
[0044] a memory storing one or more computer programs;
[0045] The processor calls the computer program to implement:
[0046] The invention discloses a method for configuring a pipelined and parallel GPU in an artificial intelligence system or a method for distributed training of a neural network based on the GPU configuration method.
[0047] In a sixth aspect, the present invention provides a readable storage medium storing a computer program, wherein the computer program is called by a processor to implement:
[0048] The invention discloses a method for configuring a pipelined and parallel GPU in an artificial intelligence system or a method for distributed training of a neural network based on the GPU configuration method.
[0049] Beneficial Effects
[0050] 1. The technical solution of the present invention provides a pipeline parallel GPU configuration method in an artificial intelligence system. Before each training except the first training, if the available resources of the GPU change, the configuration relationship between the GPU and the network layer will be updated or the current configuration relationship between the GPU and the network layer will be kept unchanged, that is, before the next training of the neural network, the configuration relationship between the GPU and the network layer is optimized so that it can dynamically match the available resources of the GPU and improve the utilization rate of the GPU. In the specific implementation process, the available bandwidth of each GPU, the forward propagation time and the backward propagation time of the network layer where each GPU is responsible are included in the dynamic index, and combined with the static index, several working partitions are first obtained, and then the training speed prediction value of each new working partition is predicted by using the dynamic and static indicators, and it is used as a screening index to screen out new working partitions with faster training speed; finally, reinforcement learning is used to judge whether the screened new working partition is more suitable for the current environment than the current working partition, so as to determine a better working partition for the next training, ensure the utilization rate of the GPU, lay the foundation for accelerating the training of the neural network, and effectively solve the technical challenges brought by the dynamic changes in the available resources of the GPU of the shared GPU cluster in distributed training.
[0051] 2. The GPU configuration method provided by the technical solution of the present invention is applied to neural network training. Except for the first training, the GPU configuration method is used to determine the next GPU allocation plan, so that the entire training process of the neural network is highly dynamically matched with the available resources of the GPU, thereby improving the training speed and efficiency of the neural network. Among them, the core of the present invention is to solve the problem that the GPU allocation in distributed training cannot adapt to the dynamic adjustment process of the available resources of the GPU. This technical idea is applicable to the training of any type of neural network using distributed pipeline parallelism under a shared GPU cluster, and has a wider range of applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a flowchart of a pipeline parallel GPU configuration method in an artificial intelligence system provided by an embodiment of the present invention;
[0053] Figure 2 It is a flowchart of a neural network distributed training method based on a GPU configuration method provided by an embodiment of the present invention;
[0054] Figure 3 It is a schematic diagram of the network structure of a training speed prediction model of a working partition constructed based on a meta-network provided in an embodiment of the present invention;
[0055] Figure 4 Schematic diagram of the network architecture of a screening model based on reinforcement learning provided by an embodiment of the present invention;
[0056] Figure 5Figure 1 is a schematic diagram of the training speed of three deep neural network training schemes in different deep neural network models and two different communication methods as the network bandwidth changes; (a) is a schematic diagram of the training speed of the three training methods in the PS, Tensorflow, ResNet50 scenarios, (b) is a schematic diagram of the training speed of the three training methods in the PS, Tensorflow, VGG16 scenarios, (c) is a schematic diagram of the training speed of the three training methods in the PS, Tensorflow, AlexNet scenarios, (d) is a schematic diagram of the training speed of the three training methods in the PS, MXNet, ResNet50 scenarios, and (e) is a schematic diagram of the training speed of the three training methods in the PS , MXNet, VGG16 scenarios, Figure (f) is a schematic diagram of the training speed of the three training methods in the PS, MXNet, AlexNet scenarios, Figure (g) is a schematic diagram of the training speed of the three training methods in the RingAll-reduce, PyTorch, ResNet50 scenarios, Figure (h) is a schematic diagram of the training speed of the three training methods in the RingAll-reduce, PyTorch, VGG16 scenarios, Figure (i) is a schematic diagram of the training speed of the three training methods in the RingAll-reduce, PyTorch, AlexNet scenarios. DETAILED DESCRIPTION
[0057] In the existing distributed pipeline parallel method, the problem that the available resources of the GPU in the shared GPU cluster will fluctuate and affect the GPU utilization rate is not considered. The GPU configuration corresponding to each network layer of the neural network has been determined before the training starts and remains unchanged during the training process. In order to solve this technical problem, the present invention provides a pipeline parallel GPU configuration method, a neural network distributed training method and system in an artificial intelligence system. Specifically, the available bandwidth of each GPU, the forward propagation and backward propagation time on each GPU at different model layers are monitored as dynamic indicators; then the above-mentioned dynamic indicators and the recorded static indicators are used to give several working partitions adapted to the current environment, and then the training speed of each working partition is predicted through the meta-network, so as to screen out more reasonable new working partitions; finally, reinforcement learning is used to determine whether to switch to the new working partition. The present invention will be further described below in conjunction with the embodiments.
[0058] Example 1
[0059] This embodiment provides a pipeline parallel GPU configuration method in an artificial intelligence system, which is used to update or maintain the current configuration relationship between the GPU and the network layer before the next training of the neural network, and specifically includes the following steps:
[0060] Step 1: Get the current static and dynamic indicators in the distributed training system.
[0061] In this embodiment, the distributed training system adopts pipeline parallelism, which is based on model parallelism. The model is also partitioned, and a batch of data is split into multiple micro-batches. The GPU of the previous layer and the GPU of the next layer can process different micro-batches at the same time. That is, in this embodiment, GPU allocation refers to configuring a GPU for each network layer of the neural network, and the same GPU is responsible for training one or more network layers.
[0062] In this embodiment, the static indicators include the number of network layers of the neural network, the number of GPUs, and the training characteristics of each network layer. The training characteristics of the network layer are the output activation size O i , weight parameter P i and the gradient G i In other feasible embodiments, the training characteristics of the network layer can also be set as the weight parameter P i and the gradient G i , that is, to adjust it adaptively according to the accuracy requirements. Assuming N represents the number of network layers of the neural network, M represents the number of GPUs, and the training features of all network layers are represented by vectors:
[0063] O=[O1,...,O i ,...,O N ]
[0064] P=[P1,...,P i ,...,P N ]
[0065] G=[G1,...,G i ,...G N ]
[0066] Among them, O, P, and G represent the output activation size vector, weight parameter vector, and gradient vector corresponding to all network layers of the neural network, respectively.
[0067] In this embodiment, the dynamic indicators include: the available bandwidth B of each GPU j , the forward propagation time FP of the network layer that each GPU is responsible for i,j Back propagation time BP i,j , j corresponds to the GPU number, i corresponds to the network layer number. The available bandwidth vector B, forward propagation time vector FP, and backward propagation time vector BP of all GPUs are represented by vectors:
[0068] B=[B1,...,B j ,...,B M ]
[0069] FP=[FP1,...,FP j ,...,FP M ],FP ij ∈FP j
[0070] BP=[BP1,...,BP j ,...BP M ],BP ij ∈BP j
[0071] Among them, FP j Represents the forward propagation time vector of the network layer that the jth GPU is responsible for; BP j Represents the backward propagation time vector of the network layer that the j-th GPU is responsible for.
[0072] Step 2: Generate several new work partitions based on the static indicators and dynamic indicators.
[0073] Among them, the means of obtaining new work partitions can refer to the existing technology. For example, this embodiment uses the partitioning algorithm publicly recorded in the paper PipeDream: Generalized Pipeline Parallelism for DNN Training to generate several new work partitions. The paper discloses a partitioning algorithm, that is, using this method and using static indicators and dynamic indicators to generate a series of new work partitions, but this embodiment does not use the dynamic programming screening method disclosed in the paper, but uses the following technology to screen work partitions.
[0074] Step 3: Taking the static index, dynamic index, and new work partition as input, the training speed prediction value corresponding to each new work partition is obtained using the work partition training speed prediction model constructed based on the meta-network.
[0075] The meta-network is an independent network, and its purpose of introduction is to obtain the training speed prediction value of the working partition. Figure 3 As shown, it includes an embedding layer corresponding to each group of features, an LSTM network corresponding to the dynamic features, and a fully connected layer. Among them, the dynamic indicators are input into the embedding layer corresponding to the dynamic features to obtain the output results, and then the respective output results are input into the LSTM network to learn the dynamic environment to obtain the sequence features of the available bandwidth vector B, the forward propagation time vector FP, and the backward propagation time vector BP; the static indicators are input into the embedding layer corresponding to the static features to obtain the output results, and then the output results of the embedding layer corresponding to the static features, the sequence features, and the new working partition are used as the input of the fully connected layer, and the output results are obtained after inputting the fully connected layer. The output result is the training speed prediction value of the new working partition.
[0076] Among them, the embedding layer, LSTM and fully connected layer are all existing network architectures, and the functions realized are also the functions possessed by the network, such as the embedding layer is used to convert the input data into a fixed-size vector, and the fully connected layer is used to connect the features extracted from each component, so as to predict the training speed. The present invention forms a meta-network by combining the embedding layer, LSTM and fully connected layer, and connecting the LSTM network according to the embedding layer corresponding to the dynamic index, and the embedding layer corresponding to the static index, and the LSTM network connected to the fully connected layer, which is finally used to realize the prediction of the training speed of the working partition. It should be understood that the network structure size of the embedding layer, LSTM and fully connected layer is determined based on the input static index, dynamic index and working partition, so as to ensure that the output of the training speed prediction value is obtained after the static index, dynamic index and working partition are combined as the model input.
[0077] It should be understood that based on the meta-network architecture built by the present invention and the determined inputs and outputs; the working partitions with known training speeds and their dynamic and static indicators are collected to construct a sample set, and the working partition training speed prediction model is obtained by offline training of the meta-network. The training process can be understood as obtaining the function f:(FP, BP, B, N, M, P, G, O, S)→V, where V is the predicted value of the working partition training speed, and S represents the new working partition, which is generally represented in the form of an array, such as S[M], S[M]={S[0], S[1]…S[M-1]}, the working partition is described in the form of an array of size M, the array subscript starts from 0, and each element in the array represents the layer assigned to each worker. Since the training process is a conventional technical means in this field, it will not be described in detail.
[0078] Step 4: Filter new work partitions based on their predicted training speed and the number of GPUs whose configuration relationships have changed in each new work partition.
[0079] In this embodiment, when screening new work partitions, the following two rules need to be met at the same time:
[0080] Rule 1: Only two GPUs in the selected new work partitions have changed configurations. The change in the configuration relationship between the GPU and the network layer means that the network layer that the GPU is responsible for has changed.
[0081] Rule 2: The predicted training speed of the selected new working partition is higher than the training speed of the current working partition.
[0082] When the training speed prediction value of each working partition solution is obtained through the meta-network, it is time-consuming to use the enumeration method to find the optimal working partition. Therefore, the present invention preferably first screens out the working partitions in which only two GPUs and the network layer configuration relationship have changed, and then selects the faster working partition according to the training speed.
[0083] It should be understood that in this embodiment, if only one working partition is needed for final screening, the working partition with the largest training speed prediction value is selected based on the training speed.
[0084] Step 5: Taking the static indicators, dynamic indicators, the screened new working partition and the current working partition as input, a screening model based on reinforcement learning is used to determine whether to replace the current working partition, that is, to update or keep the current configuration relationship between the GPU and the network layer unchanged.
[0085] Reinforcement learning is different from deep learning. Deep learning uses existing data to train algorithms to find patterns to solve corresponding problems, and then uses this pattern to predict new data. Reinforcement learning adjusts its actions (outputs) through feedback results (reward functions) to obtain the best results. In this embodiment, the reward function is based on the training speed, specifically comparing the training speed corresponding to the selected working partition with the training speed corresponding to the previous working partition. If the training speed corresponding to the selected working partition is greater than the training speed corresponding to the previous working partition, the output result of the screening model is considered to be correct, that is, in the environment of this training, the output is correct, and the next time this network environment is encountered, the partition solution of this time is preferred; otherwise, it means that the output result of this time is slightly worse, and the next time this network environment is encountered, the priority of this partition solution is reduced. Since the present invention does not optimize the network structure (fully connected neural network) and reward function type of reinforcement learning, it is only to solve the problem of updating the working partition of the present invention, determine the model input and model output related to this application, and obtain the model offline training by constructing a sample set containing model input and output, therefore, the reinforcement learning network architecture and the type of reward function are not constrained or stated.
[0086] It should be understood that according to the above process, a more reasonable GPU configuration can be determined for the next training of the neural network, that is, a working partition result that is more suitable for the current environment can be obtained. The above process is preferably executed after the current training is completed and before the next training of the neural network begins to determine the working partition for the next training.
[0087] Embodiment 2:
[0088] The GPU configuration method provided in Example 1 is applied to the distributed training of a neural network. Therefore, this embodiment provides a neural network distributed training method based on the GPU configuration method, which includes the following steps:
[0089] Step S1: Load the neural network to be trained and the data set into the distributed training system, and layer the neural network.
[0090] Step S2: Initialize the working partition and perform the first training. In this embodiment, the working partition in the existing pipeline parallel solution is selected as the initialization working partition. That is, the configuration relationship between the GPU and the network layer is determined according to the initialization working partition.
[0091] Step S3: Before the next training of the neural network, determine the working partition for the next training of the neural network according to steps 1 to 5 and then perform training; wherein, if a new working partition is obtained, update the configuration relationship between the GPU and the network layer according to the new working partition.
[0092] Since the implementation process of this step can refer to the implementation process of Example 1, no specific description is given for it.
[0093] Step S4: Determine whether the iterative training termination condition of the neural network is met. If not, return to step S3 to continue training; otherwise, complete the training of the neural network.
[0094] The deep neural network model needs to be trained many times to achieve convergence (i.e., the output of the deep neural network model basically meets the actual results). This embodiment proposes a corresponding work partition solution for each training, then compares the training speed with the work partition solution used in the last training, and then allows the system to decide whether to switch to the work partition solution of the current training, and finally starts this training.
[0095] It should be understood that a neural network has many network layers, and distributed training of a neural network means that each GPU is responsible for training some network layers respectively. Therefore, the work partition in the present invention is to determine which GPU is responsible for training which network layer. As for how the GPU uses the data set to realize the training of the neural network, it is an existing technology, and the present invention does not impose specific constraints on this.
[0096] In addition, the neural network distributed training method of the present invention does not constrain the type and application scenario of the neural network model, as long as the neural network distributed training under the shared GPU cluster working condition can be applied to the technical solution of the present invention. The present embodiment is to apply it to image classification, and select three deep neural network models of VGG16, ResNet50 and AlexNet to realize image classification, and then use the training data of image classification, and use the neural network distributed training method provided by the present invention to train the above-mentioned neural network to obtain an image classification model. Wherein, the synthetic data training is set to the format of ImageNet. Other feasible application fields can also be translation (for example, realizing the neural network model in English translation), video subtitles, language recognition, etc.
[0097] Embodiment 3:
[0098] This embodiment provides a distributed training system, which includes at least several GPU servers, each of which is equipped with a GPU, a CPU, a memory, a network card, and a switch. The network layer that each GPU server is responsible for in each training is determined according to the GPU configuration method provided in Example 1; it can also be understood as the neural network distributed training method provided in Example 2, and all GPU servers in the distributed training system are used to implement distributed training of the neural network.
[0099] The GPU is used to implement network training, i.e. data calculation. The memory is used to store data. The CPU, network card and switch are used to implement data transmission. The GPU servers use a distributed communication method, such as PS (Parameter Server) parameter server or RingAll-reduce.
[0100] It should be understood that in some implementations, a GPU server can be selected as a controller. In addition to implementing the network training function, the GPU configuration method described in Example 1 is executed to determine the network layer that each GPU server is responsible for in each training; in other implementations, other external controllers can also be used to determine the network layer that each GPU server is responsible for in each training, and then all GPU servers are used to implement distributed training of neural networks. The specific implementation technology needs to be determined by the selected communication method, and the distributed training system is a prior art, so it will not be described in detail.
[0101] Embodiment 4:
[0102] This embodiment provides a GPU allocation device of the GPU configuration method, which includes: a dynamic and static indicator acquisition module, a configuration module, a training speed prediction value acquisition module, a screening module and a decision module.
[0103] The dynamic and static indicator acquisition module is used to obtain the current static indicators and dynamic indicators in the distributed training system.
[0104] The distributed training system adopts pipeline parallelism, which configures a GPU for each network layer of the neural network. The same GPU is responsible for the training of one or more network layers. The static indicators include: the number of network layers of the neural network, the number of GPUs, and the training characteristics of each network layer, such as output activation size, weight parameters, and gradients; the dynamic indicators include: the available bandwidth of each GPU, the forward propagation time and backward propagation time of the network layer that each GPU is responsible for.
[0105] The configuration module is used to generate a number of new working partitions according to the static indicators and dynamic indicators, wherein the working partition represents the configuration relationship between the GPU and the network layer.
[0106] The training speed prediction value acquisition module is used to take the static index, dynamic index and new working partition as input, and use the working partition training speed prediction model constructed based on the meta-network to obtain the training speed prediction value corresponding to each new working partition.
[0107] A screening module is used to screen new working partitions based on a predicted value of a training speed of each new working partition and the number of GPUs whose configuration relationship has changed in each new working partition.
[0108] The decision module is used to take the static indicators, dynamic indicators, the screened new working partition and the current working partition as inputs, and use the screening model based on reinforcement learning to determine whether to replace the current working partition, that is, to update or maintain the current configuration relationship between the GPU and the network layer.
[0109] Please refer to the above method for the specific implementation process of each module, which will not be repeated here. It should be understood that the division of the above functional modules is only a division of logical functions, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. At the same time, the above integrated units can be implemented in the form of hardware or software functional units.
[0110] Embodiment 5:
[0111] This embodiment provides an electronic device, which includes: one or more processors and a memory storing one or more computer programs; wherein the processor calls the computer program to implement: the steps of a pipeline-parallel GPU configuration method in an artificial intelligence system; or the steps of a neural network distributed training method based on the GPU configuration method.
[0112] In some embodiments, when a processor calls a computer program to implement the steps of a pipeline parallel GPU configuration method in an artificial intelligence system, specifically performing:
[0113] Step 1: Obtain the current static and dynamic indicators in the distributed training system;
[0114] The distributed training system adopts pipeline parallelism, which configures a GPU for each network layer of the neural network. The same GPU is responsible for the training of one or more network layers. The static indicators include: the number of network layers of the neural network, the number of GPUs, and the training characteristics of each network layer; the dynamic indicators include: the available bandwidth of each GPU, the forward propagation time and backward propagation time of the network layer that each GPU is responsible for.
[0115] Step 2: Generate a number of new work partitions based on the static indicators and dynamic indicators; wherein the work partition represents the configuration relationship between the GPU and the network layer.
[0116] Step 3: Taking the static index, dynamic index, and new work partition as input, using the work partition training speed prediction model constructed based on the meta-network to obtain the training speed prediction value corresponding to each new work partition;
[0117] Step 4: Filter the new work partitions based on the predicted training speed of each new work partition and the number of GPUs whose configuration relationships have changed in each new work partition;
[0118] Step 5: Taking the static indicators, dynamic indicators, the screened new working partition and the current working partition as input, a screening model based on reinforcement learning is used to determine whether to replace the current working partition, that is, to update or maintain the current configuration relationship between the GPU and the network layer.
[0119] In some other implementations, when the processor calls a computer program to implement: a step of a neural network distributed training method based on the GPU configuration method, specifically performing:
[0120] Step S1: loading the neural network to be trained and the data set into the distributed training system, and stratifying the neural network;
[0121] Step S2: Initialize the working partition and perform the first training;
[0122] Wherein, a GPU is configured for each network layer of the neural network, and the same GPU is responsible for training one or more network layers, that is, the configuration relationship between the GPU and the network layer is determined according to the initialization work partition, and the distributed training system adopts a distributed communication mechanism for communication connection;
[0123] Step S3: Before the next training of the neural network, determine the working partition for the next training of the neural network according to the method of steps 1 to 5 and then perform training; wherein, if a new working partition is obtained, update the configuration relationship between the GPU and the network layer according to the new working partition;
[0124] Step S4: Determine whether the iterative training termination condition of the neural network is met. If not, return to step S3 to continue training; otherwise, complete the training of the neural network.
[0125] The memory may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0126] If the memory and the processor are implemented independently, the memory, the processor and the communication interface can be connected to each other through a bus and communicate with each other. The bus can be an industrial standard architecture bus, an external device interconnection bus or an extended industrial standard architecture bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0127] Optionally, in a specific implementation, if the memory and the processor are integrated on a chip, the memory and the processor can communicate with each other through an internal interface.
[0128] It should be understood that in the embodiments of the present invention, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the type of device.
[0129] Embodiment 6:
[0130] This embodiment provides a readable storage medium storing a computer program, which is called by a processor to implement: the steps of a pipeline-parallel GPU configuration method in an artificial intelligence system; or the steps of a neural network distributed training method based on the GPU configuration method.
[0131] Among them, in some ways, when the computer program is called by a processor to implement: a pipeline parallel GPU configuration method in an artificial intelligence system, specifically executing:
[0132] Step 1: Get the current static and dynamic indicators in the distributed training system.
[0133] The distributed training system adopts pipeline parallelism, which configures a GPU for each network layer of the neural network. The same GPU is responsible for the training of one or more network layers. The static indicators include the number of network layers of the neural network, the number of GPUs, and the training characteristics of each network layer; the dynamic indicators include: the available bandwidth of each GPU, the forward propagation time and the backward propagation time of the network layer that each GPU is responsible for;
[0134] Step 2: Generate a number of new work partitions based on the static indicators and dynamic indicators; wherein the work partition represents the configuration relationship between the GPU and the network layer.
[0135] Step 3: Taking the static index, dynamic index, and new work partition as input, the training speed prediction value corresponding to each new work partition is obtained using the work partition training speed prediction model constructed based on the meta-network.
[0136] Step 4: Filter new work partitions based on their predicted training speed and the number of GPUs whose configuration relationships have changed in each new work partition.
[0137] Step 5: Taking the static indicators, dynamic indicators, the screened new working partition and the current working partition as input, a screening model based on reinforcement learning is used to determine whether to replace the current working partition, that is, to update or maintain the current configuration relationship between the GPU and the network layer.
[0138] In some other implementations, the computer program is called by a processor to implement: a step of a neural network distributed training method based on the GPU configuration method, specifically performing:
[0139] Step S1: loading the neural network to be trained and the data set into the distributed training system, and stratifying the neural network;
[0140] Step S2: Initialize the working partition and perform the first training;
[0141] Wherein, a GPU is configured for each network layer of the neural network, and the same GPU is responsible for training one or more network layers, that is, the configuration relationship between the GPU and the network layer is determined according to the initialization work partition, and the distributed training system adopts a distributed communication mechanism for communication connection;
[0142] Step S3: Before the next training of the neural network, determine the working partition for the next training of the neural network according to the method of steps 1 to 5 and then perform training; wherein, if a new working partition is obtained, update the configuration relationship between the GPU and the network layer according to the new working partition;
[0143] Step S4: Determine whether the iterative training termination condition of the neural network is met. If not, return to step S3 to continue training; otherwise, complete the training of the neural network.
[0144] For the specific implementation process of each step, please refer to the description of the above method.
[0145] The readable storage medium is a computer-readable storage medium, which may be an internal storage unit of the controller described in any of the foregoing embodiments, such as a hard disk or memory of the controller. The readable storage medium may also be an external storage device of the controller, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the controller. Further, the readable storage medium may also include both an internal storage unit of the controller and an external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0146] Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned readable storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0147] Experimental verification:
[0148] Experimental setup: 10 GPU servers were set up, each of which had an NVIDIA P100 GPU, 40 CPU cores, 128 GB memory, a Mellanox ConnectX5 100 Gbps network card, and a Mellanox SN2100 switch. The Mellanox driver version was 5.1-0.6.6.0. The experimental operating system was Ubuntu 18.04, and the Linux kernel version was 4.15.0-55-generic. The experiment used three deep neural network models, namely VGG16, ResNet50, and AlexNet, and two different parameter synchronization schemes, namely PS (Parameter Synchronization). Server) and RingAll-reduce, using three different machine learning frameworks, namely Tensorflow, MXNet and PyTorch, with the size of training data set to 64 for VGG16, 128 for ResNet and 256 for AlexNet, and link bandwidth of 10Gbps-100Gbps. The present invention performs performance tests with existing deep neural network training schemes (PipeDream, Baseline) under different environments. Among them, Baseline is a common deep neural network training scheme.
[0149] Figure 5 Schematic diagram of the training speed of three deep neural network training schemes in different deep neural network models and two different communication methods as the network bandwidth changes, where: Figure 5 (a)-(c) are schematic diagrams of the training speed of three training methods in the scenarios of PS, Tensorflow and three neural network models (ResNet50, VGG16 and AlexNet) as the network bandwidth changes (10Gbps, 25Gbps, 40Gbps, 100Gbps), wherein Figure (a) is a schematic diagram of the training speed of the three training methods in the scenarios of PS, Tensorflow, ResNet50, Figure (b) is a schematic diagram of the training speed of the three training methods in the scenarios of PS, Tensorflow, VGG16, and Figure (c) is a schematic diagram of the training speed of the three training methods in the scenarios of PS, Tensorflow, AlexNet. The present invention is named AutoPipe. As can be seen from the figure, for ResNet50, AutoPipe's performance is 177% and 89% higher than Baseline and PipeDream, for VGG16, AutoPipe's performance is 113% and 44% higher than Baseline and PipeDream, and for AlexNet, AutoPipe's performance is 143% and 70% higher than Baseline and PipeDream.
[0150] Figure 5 (d)-(f) are schematic diagrams of the training speed of the three training methods in the scenarios of PS, MXNet and three neural network models (ResNet50, VGG16 and AlexNet) with changes in network bandwidth (10Gbps, 25Gbps, 40Gbps, 100Gbps). Among them, (d) is a schematic diagram of the training speed of the three training methods in the scenarios of PS, MXNet, and ResNet50, Figure (e) is a schematic diagram of the training speed of the three training methods in the scenarios of PS, MXNet, and VGG16, and Figure (f) is a schematic diagram of the training speed of the three training methods in the scenarios of PS, MXNet, and AlexNet. It can be seen from the figure that for ResNet50, the performance of AutoPipe is 171% and 82% higher than that of Baseline and PipeDream, for VGG16, the performance of AutoPipe is 104% and 41% higher than that of Baseline and PipeDream, and for AlexNet, the performance of AutoPipe is 124% and 58% higher than that of Baseline and PipeDream.
[0151] Figures (g)-(i) are schematic diagrams of the training speed of the three training methods under Ring All-reduce, PyTorch and three different neural network models (ResNet50, VGG16 and AlexNet) scenarios with changes in network bandwidth (10Gbps, 25Gbps, 40Gbps, 100Gbps), where Figure (g) is a schematic diagram of the training speed of the three training methods under Ring All-reduce, PyTorch, ResNet50 scenarios, Figure (h) is a schematic diagram of the training speed of the three training methods under Ring All-reduce, PyTorch, VGG16 scenarios, and Figure (i) is a schematic diagram of the training speed of the three training methods under Ring All-reduce, PyTorch, VGG16 scenarios. Schematic diagram of the training speed of the three training methods in the All-reduce, PyTorch, and AlexNet scenarios. It can be seen from the figure that for ResNet50, AutoPipe's performance is 148% and 65% higher than Baseline and PipeDream, for VGG16, AutoPipe's performance is 117% and 17% higher than Baseline and PipeDream, and for AlexNet, AutoPipe's performance is 143% and 26% higher than Baseline and PipeDream.
[0152] Therefore, the present invention observes through the above experiments: 1) AutoPipe outperforms PipeDream in all cases, and AutoPipe even obtains more acceleration based on PipeDream in some cases. 2) AutoPipe shows more acceleration in ResNet50. The reason is that ResNet50 contains more layers than the other two models. Therefore, AutoPipe obtains more benefits from more accurate modeling and finer-grained switching.
[0153] It should be emphasized that the examples described in the present invention are illustrative rather than restrictive, and therefore the present invention is not limited to the examples described in the specific embodiments. Any other embodiments derived by those skilled in the art based on the technical solution of the present invention that do not depart from the purpose and scope of the present invention, whether modified or replaced, also fall within the scope of protection of the present invention.
Claims
1. A pipeline parallel GPU configuration method in an artificial intelligence system, characterized by: The GPU configuration method is based on a shared GPU cluster for updating or maintaining the configuration relationship between the current GPU and the network layer before the next training of the neural network, the neural network is divided into multiple network layers, and there is a shared GPU in the GPU cluster. The GPU configuration method includes the following steps: Step 1: Obtain the current static and dynamic indicators in the distributed training system; The distributed training system adopts pipeline parallelism, which configures a GPU for each network layer of the neural network. The same GPU is responsible for the training of one or more network layers. The static indicators include: the number of network layers of the neural network, the number of GPUs, and the training characteristics of each network layer; the dynamic indicators include: the available bandwidth of each GPU, the forward propagation time and the backward propagation time of the network layer that each GPU is responsible for; Step 2: Generate several new work partitions according to the static indicators and dynamic indicators; Among them, the work partition represents the configuration relationship between the GPU and the network layer; Step 3: Taking the static index, dynamic index, and new work partition as input, and using the work partition training speed prediction model to obtain a training speed prediction value corresponding to each new work partition; Step 4: Filter the new work partitions based on the predicted training speed of each new work partition and the number of GPUs whose configuration relationships have changed in each new work partition; Step 5: Taking the static indicators, dynamic indicators, the screened new working partition and the current working partition as input, a screening model based on reinforcement learning is used to determine whether to replace the current working partition, that is, to update or maintain the current configuration relationship between the GPU and the network layer.
2. The GPU configuration method according to claim 1, characterized in that: The work partition training speed prediction model in step 3 is constructed based on a meta-network, and the meta-network includes: each group of static indicators, an embedding layer corresponding to the dynamic indicators, an LSTM network corresponding to the dynamic indicators, and a fully connected layer, wherein the embedding layer corresponding to the dynamic indicators is connected to the LSTM network, and the embedding layer corresponding to the static indicators and the LSTM network are connected to the fully connected layer; The dynamic indicator is input into the corresponding embedding layer, and the obtained output result is input into the LSTM network to obtain the sequence characteristics of the dynamic indicator; The static indicator is input into the corresponding embedding layer, and the output result, the sequence feature and the new working partition are used as the input of the fully connected layer, and the output of the fully connected layer is the training speed prediction value of the new working partition.
3. The GPU configuration method according to claim 2, characterized in that: The work partition training speed prediction model and the screening model are both constructed through offline training; Among them, the goal of the reward function in the screening model is to make the training speed corresponding to the working partition selected by the screening model greater than the training speed corresponding to the previous working partition.
4. The GPU configuration method according to claim 1, wherein: When selecting new work partitions in step 4, the following two rules must be met at the same time: Rule 1: The selected new work partition has only two GPU configurations that have changed, and the configuration is the configuration relationship between the GPU and the network layer; Rule 2: The predicted training speed of the selected new working partition is higher than the training speed of the current working partition.
5. The GPU configuration method according to claim 1, characterized in that: The training features of each network layer include the output activation size, weight parameters and gradient of the network layer.
6. A neural network distributed training method based on the GPU configuration method of claim 1, characterized in that: The following steps are involved: Step S1: loading a neural network to be trained and a data set into a distributed training system, wherein the neural network is divided into multiple network layers; Step S2: Initialize the working partition and perform the first training of the neural network; Wherein, a GPU is configured for each network layer of the neural network, and the same GPU is responsible for training one or more network layers, that is, the configuration relationship between the GPU and the network layer is determined according to the initialization work partition, the GPU uses the data set to train the neural network, and the distributed training system uses a distributed communication mechanism for communication connection; Step S3: Before the next training of the neural network, determine the working partition corresponding to the next training of the neural network according to the method of steps 1 to 5 and then perform training; wherein, if a new working partition is obtained, update the configuration relationship between the GPU and the network layer according to the new working partition; Step S4: Determine whether the iterative training termination condition of the neural network is met. If not, return to step S3 to continue training; otherwise, complete the training of the neural network.
7. A GPU allocation device based on the GPU configuration method according to any one of claims 1 to 5, characterized in that: include: The dynamic and static indicator acquisition module is used to obtain the current static and dynamic indicators in the distributed training system; The distributed training system adopts pipeline parallelism, which configures a GPU for each network layer of the neural network. The same GPU is responsible for the training of one or more network layers. The static indicators include: the number of network layers of the neural network, the number of GPUs, and the training characteristics of each network layer; the dynamic indicators include: the available bandwidth of each GPU, the forward propagation time and the backward propagation time of the network layer that each GPU is responsible for; A configuration module, used to generate a number of new work partitions according to the static indicators and dynamic indicators; Among them, the work partition represents the configuration relationship between the GPU and the network layer; A training speed prediction value acquisition module is used to take the static index, dynamic index, and new working partition as input, and use the working partition training speed prediction model to obtain a training speed prediction value corresponding to each new working partition; A screening module, used for screening new working partitions based on a predicted value of a training speed of each new working partition and the number of GPUs whose configuration relationship has changed in each new working partition; The decision module is used to take the static indicators, dynamic indicators, the screened new working partition and the current working partition as inputs, and use the screening model based on reinforcement learning to determine whether to replace the current working partition, that is, to update or maintain the current configuration relationship between the GPU and the network layer.
8. A distributed training system based on the GPU configuration method of claim 1 or the training method of claim 6, characterized in that: The distributed training system at least includes: a plurality of GPU servers, each of which is provided with a GPU, a CPU, a memory, a network card and a switch; Among them, the GPU is used to implement neural network training; the memory is used to store data; the CPU, network card and switch are used to implement data transmission, and distributed communication is adopted between the GPU servers.
9. An electronic device, characterized in that: include: one or more processors; a memory storing one or more computer programs; The processor calls the computer program to implement: The GPU configuration method of claim 1 or the neural network distributed training method of claim 6.
10. A readable storage medium, characterized in that: A computer program is stored, which is called by a processor to implement: The GPU configuration method of claim 1 or the neural network distributed training method of claim 6.
Citation Information
Patent Citations
Heterogeneous network perception model division and task placement method in pipelined distributed deep learning
CN110533183A
Hybrid pipeline parallel method for accelerating distributed deep neural network training
CN112784968A