Pipeline parallel distributed training method, device and system for deep neural network

By adopting asynchronous parallel collaborative training method in cloud-edge collaborative optical network, the problems of large communication overhead, unbalanced load and unstable asynchronous step-by-step training in pipeline parallel training are solved, and efficient and stable deep neural network training is achieved.

CN119987999AActive Publication Date: 2025-05-13BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202411953201.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

In cloud-edge collaborative optical network, the existing pipeline parallel training methods have problems such as large communication overhead, unbalanced load and unstable asynchronous step-by-step adjustment, which affects training efficiency and model stability.

Method used

By using optical networks to perform asynchronous parallel collaborative training between edge servers and cloud servers, the training tasks of deep neural networks are decomposed, and small batches of training data are transmitted and processed in optical networks, so as to achieve decoupling and collaborative execution of asynchronous parallel training steps.

Benefits of technology

It reduces the communication overhead during the training process, realizes load balancing between devices, improves training efficiency and model stability, and enhances resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987999A_ABST
    Figure CN119987999A_ABST
Patent Text Reader

Abstract

The invention provides an assembly line parallel distributed training method, device and system for a deep neural network, and the method comprises the steps: sequentially transmitting each group of small-batch training data forming the current-batch training data to an edge server through an optical network, enabling the edge server to cooperate with the cloud server to perform asynchronous parallel cooperative training on different sub-task models forming the deep neural network in a communication and training decoupling mode, and outputting gradients corresponding to each group of small-batch training data in sequence by the edge server; and sequentially receiving the gradient of each group of small-batch training data. According to the invention, correct transmission of data in a model training process can be ensured, communication overhead in the training process can be reduced, load balance between devices can be realized, model training efficiency and effectiveness can be improved, and the utilization rate of resources of the devices participating in training can be increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of model training technology, and in particular to a pipelined parallel distributed training method, device and system for deep neural networks. Background Art

[0002] With the rapid development of information technology, such as digital twins, Chat GPT, and large models, the number of applications with extremely high computing and latency requirements is surging. At the same time, the advancement of artificial intelligence, especially the key role of deep neural networks (DNNs) in solving complex problems such as image recognition, object tracking, and language conversion, has led to a growing demand for powerful computing power. However, with the diversification of applications and the surge in data volume, the scale of DNN models is also expanding, which requires more powerful computing resources to support the training of these large models.

[0003] To address this challenge, the European Telecommunications Standards Institute (ETSI) proposed the concept of mobile edge computing (MEC), which allows computing tasks to be transferred to edge servers for timely processing. However, edge servers have limited computing resources, and cloud servers, although they can provide sufficient computing and storage resources, may generate large communication delays due to their long distance from users, which does not meet the requirements of applications with strict low latency requirements. Therefore, in order to meet these requirements, it is necessary to establish a system that has sufficient computing resources and can provide high-bandwidth and low-latency communication links between cloud and edge servers. One possible solution is to use cloud-edge collaborative optical networks to improve the efficiency of communication and computing. Optical networks are known for their large capacity, high reliability, and long-distance transmission capabilities. In cloud-edge collaborative optical networks, in order to reduce the training delay of multi-layer DNN tasks, the tasks can be decomposed, with one part executed on the edge server and the other part processed by the cloud server, which not only reduces the delay of data transmission, but also reduces the computing burden of the central cloud.

[0004] In the field of computer science, the problem of how to train large neural network models on hardware accelerators (such as graphics processors (GPUs)) with multiple interconnect capabilities has attracted widespread attention. Common solutions include parallel training techniques such as data parallelism, model parallelism, and pipeline parallelism, which can improve training efficiency. Distributed parallel training can distribute different stages of model training to different computing nodes to achieve concurrent execution, reduce waiting time, and improve resource utilization. Different parallel training methods have different effects on performance, efficiency, and model accuracy. However, these studies are usually conducted in data center environments and may not fully consider the impact of communication delays. In cloud-edge collaborative optical networks, communication delay becomes an important factor due to the long physical distance between computing nodes. Therefore, studying how to efficiently deploy distributed model training in cloud-edge collaborative optical networks is crucial to accelerate the implementation of AI applications in optical networks.

[0005] However, although existing asynchronous training methods such as pipeline parallelism in cloud-edge collaborative optical networks are faster than synchronous training, they usually lead to a decline in the quality of results because the gradients used in asynchronous training may be calculated based on outdated parameters, which can lead to unstable training. At the same time, the data exchange between different stages requires additional communication overhead, which will also affect the training efficiency. In addition, if the computational complexity of each stage is inconsistent, some devices may be idle while other devices are busy calculating, resulting in load imbalance. Summary of the invention

[0006] In view of this, the embodiments of the present application provide a pipelined parallel distributed training method, device and system for deep neural networks to eliminate or improve one or more defects in the prior art.

[0007] One aspect of the present application provides a pipeline parallel distributed training method for a deep neural network, comprising:

[0008] For the current iteration round of the deep neural network, a preset pipeline parallel distributed training step is executed, wherein the pipeline parallel distributed training step includes: transmitting each group of small batch training data constituting the current batch training data to the edge server via the optical network in sequence, so that the edge server and the cloud server connected to the edge server via the optical network communication perform asynchronous parallel collaborative training on different subtask models constituting the deep neural network based on a preset communication and training decoupling method, and the edge server outputs the gradient corresponding to each group of the small batch training data in sequence;

[0009] The gradients corresponding to the respective groups of the small batch training data returned by the edge server via the optical network are received in sequence.

[0010] In some embodiments of the present application, it also includes:

[0011] If the deep neural network has not converged yet, there is the batch training data that has not been used for model training, or the current iteration round is not the preset last iteration round, then when a gradient corresponding to the small batch training data is received for the first time in the current iteration round, the pipelined parallel distributed training step is performed for the next iteration round of the deep neural network.

[0012] In some embodiments of the present application, the edge server is used to execute the following:

[0013] receiving each group of the small batch training data in sequence, and when receiving a group of the small batch training data for the first time in the current iteration round, starting to execute the preset first asynchronous training step to obtain the intermediate activation values ​​corresponding to each group of the small batch training data respectively; and each time the intermediate activation value corresponding to a group of the small batch training data is obtained, the intermediate activation value corresponding to the group of the small batch training data is transmitted to the cloud server via the optical network;

[0014] Among them, the first asynchronous training step includes: performing forward propagation calculations for the first subtask model corresponding to the deep neural network based on each group of the small batch training data, so that the first subtask model sequentially outputs the intermediate activation values ​​corresponding to each group of the small batch training data.

[0015] In some embodiments of the present application, the cloud server is used to perform the following:

[0016] receiving the intermediate activation values ​​corresponding to each group of the small batch training data in sequence, and when the intermediate activation values ​​corresponding to a group of the small batch training data are received for the first time in the current iteration round, starting to execute the preset second asynchronous training step to obtain the intermediate loss values ​​corresponding to each group of the small batch training data respectively; and each time the intermediate loss values ​​corresponding to a group of the small batch training data are obtained, the intermediate loss values ​​corresponding to the group of the small batch training data are transmitted to the edge server via the optical network;

[0017] Among them, the second asynchronous training step includes: executing preset calculation steps in sequence for the intermediate activation values ​​corresponding to each group of the small batch training data, and the calculation steps include: based on the intermediate activation values ​​corresponding to the current group of the small batch training data, sequentially executing forward propagation calculations and backward propagation calculations for the second subtask model corresponding to the deep neural network to obtain the intermediate loss values ​​corresponding to the group of small batch training data.

[0018] In some embodiments of the present application, the edge server is further configured to execute the following:

[0019] receiving the intermediate loss values ​​corresponding to each group of the small batch training data in sequence, and when the intermediate loss values ​​corresponding to a group of the small batch training data are received for the first time in the current iteration round, starting to execute the preset third asynchronous training step to obtain the gradients corresponding to each group of the small batch training data respectively; and returning the gradients corresponding to the small batch training data via the optical network each time the gradients corresponding to a group of the small batch training data are obtained; wherein, if the current iteration round is not the first iteration round, the edge server alternately executes the third asynchronous training step and the first asynchronous training step;

[0020] Among them, the third asynchronous training step includes: performing back propagation calculations on the first subtask model based on the intermediate loss values ​​corresponding to each group of the small batch training data, so that the first subtask model outputs the gradients corresponding to each group of the small batch training data in turn.

[0021] In some embodiments of the present application, before executing the preset pipeline parallel distributed training step for the current iteration round of the deep neural network, the method further includes:

[0022] Submitting a training task for a deep neural network to an optical network controller, so that the optical network controller obtains a target partition point for the deep neural network according to task requirement data corresponding to the training task and the current resource states corresponding to the edge server and the cloud server, respectively; the optical network controller divides the deep neural network into a first subtask model and a second subtask model according to the target partition point; the optical network controller assigns the first subtask model to the edge server and the second subtask model to the cloud server, and then returns a corresponding model task partition message;

[0023] Receive the model task partition message returned by the optical network controller.

[0024] Another aspect of the present application provides a pipeline parallel distributed training device for a deep neural network, comprising:

[0025] A pipeline parallel distributed training module is used to execute preset pipeline parallel distributed training steps for the current iteration round of the deep neural network, wherein the pipeline parallel distributed training steps include: transmitting each group of small batch training data constituting the current batch training data to the edge server via the optical network in sequence, so that the edge server and the cloud server connected to the edge server via the optical network communication perform asynchronous parallel collaborative training on different subtask models constituting the deep neural network based on a preset communication and training decoupling method, and the edge server outputs the gradient corresponding to each group of the small batch training data in sequence;

[0026] The gradient receiving module is used to sequentially receive the gradients corresponding to each group of the small batch training data returned by the edge server via the optical network.

[0027] A third aspect of the present application provides a pipeline parallel distributed training system for a deep neural network, comprising: a terminal device, an edge server, and a cloud server connected in sequence via optical network communication;

[0028] The terminal device is provided with a pipeline parallel distributed training device for deep neural networks, and the pipeline parallel distributed training device for deep neural networks is used to execute the pipeline parallel distributed training method for deep neural networks.

[0029] The fourth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the pipelined parallel distributed training method for deep neural networks when executing the computer program.

[0030] The fifth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the pipelined parallel distributed training method for a deep neural network.

[0031] The sixth aspect of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the pipelined parallel distributed training method for a deep neural network.

[0032] The pipeline parallel distributed training method for a deep neural network provided in the present application executes a preset pipeline parallel distributed training step for the current iteration round of the deep neural network, wherein the pipeline parallel distributed training step includes: transmitting each group of small batch training data constituting the current batch training data to the edge server via an optical network in sequence, so that the edge server and the cloud server connected to the edge server via optical network communication perform asynchronous parallel collaborative training on different subtask models used to constitute the deep neural network based on a preset communication and training decoupling method, and the edge server sequentially outputs the gradient corresponding to each group of the small batch training data; sequentially receiving the gradient corresponding to each group of the small batch training data returned by the edge server via the optical network, which can ensure the correct transmission of data during the deep neural network training process, reduce the communication overhead during the training process, achieve load balancing between the devices participating in the training, and improve the efficiency and effectiveness of the deep neural network training and the resource utilization of the terminal devices, edge servers and cloud servers participating in the deep neural network training.

[0033] Additional advantages, purposes, and features of the present application will be partially described in the following description, and will become partially apparent to those skilled in the art after studying the following, or may be learned from the practice of the present application. The purposes and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the specification and the drawings.

[0034] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present application are not limited to the above specific description, and the above and other purposes that can be achieved by the present application will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The drawings described herein are used to provide a further understanding of the present application, constitute a part of the present application, and do not constitute a limitation of the present application. The components in the drawings are not drawn to scale, but are only for the purpose of illustrating the principles of the present application. In order to facilitate the illustration and description of some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger relative to other components in the exemplary device actually manufactured according to the present application. In the drawings:

[0036] Figure 1 This is a first flow chart of a pipelined parallel distributed training method for a deep neural network in one embodiment of the present application.

[0037] Figure 2 This is a second flow chart of a pipelined parallel distributed training method for a deep neural network in one embodiment of the present application.

[0038] Figure 3This is a schematic diagram of the deep neural network partitioning and training logic in an example of this application.

[0039] Figure 4 This is a schematic diagram of an example of the timeline of asynchronous parallel training in another example of the present application.

[0040] Figure 5 This is a schematic diagram of the structure of a pipelined parallel distributed training device for a deep neural network in one embodiment of the present application.

[0041] Figure 6 This is a schematic diagram of the structure of a pipelined parallel distributed training system for a deep neural network in one embodiment of the present application.

[0042] Figure 7 A schematic diagram of a process for determining DNN task partitions in an application example of the present application.

[0043] Figure 8 A schematic diagram of the specific process of DNN task training in an application example of the present application. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the implementation modes and the accompanying drawings. Here, the illustrative implementation modes and descriptions of the present application are used to explain the present application, but are not intended to limit the present application.

[0045] It should also be noted here that in order to avoid obscuring the present application due to unnecessary details, only the structures and / or processing steps closely related to the scheme according to the present application are shown in the accompanying drawings, while other details that are not very relevant to the present application are omitted.

[0046] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0047] It should also be noted that, unless otherwise specified, the term “connection” herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.

[0048] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0049] It should be noted that in the existing pipeline parallel training scheme executed in the cloud-edge collaborative optical network, the model is partitioned across multiple GPUs, and each GPU is responsible for only part of the model. Model parallelism can usually achieve faster training time than data parallelism, because not using extremely large mini-batch sizes can improve statistical efficiency. Model parallel DNN training leads to serious underutilization of GPU resources. Therefore, pipeline parallelism is introduced on the basis of model parallelism. The core idea of ​​pipeline parallelism is: on the basis of model parallelism, data parallelism is further introduced, that is, the original data is divided into several batches of batch data, and sent to the GPU for training. The data before division can be called small batch training data, which can be written as mini-batch or minibatch. The training data divided on the mini-batch is called micro-batch. That is to say, the small batch training data appearing in one or more embodiments of the present application refers to the minibatches formed after dividing a batch of data in the training data.

[0050] Specifically, the layers of the model being trained are divided into multiple stages - each stage contains a consecutive set of layers in the model. Each stage is mapped to a separate GPU that performs forward and backward passes for all the layers in that stage. A simple pipeline parallel assignment has the DNN partitioned across four machines. In the simplest case, only one minibatch is active in the system, just like in traditional model parallel training. In the example of the computation timeline for a configuration, there are four machines (machines 1 to 4) and one active minibatch in the configuration. In the forward phase, each stage performs a forward pass of the minibatch for the layers of the stage and sends the results to the next stage. The output stage computes the loss of the minibatch after completing its forward pass. In the backward phase, each stage performs a backward pass and propagates the loss to the previous stage. There is only one active minibatch, and at most one GPU is active at any given point in time. To ensure that no GPU is idle at any point in time, multiple minibatches are injected into the pipeline one after another, thereby enhancing model parallel training through pipelining. The asynchronous communication of forward output activations and backward gradients across stages after completing a minibatch results in a large overlap in communication with the computation of subsequent minibatches, thus achieving better hardware efficiency compared to BSP, where each stage asynchronously sends its output activations to the next stage while starting to process another minibatch.

[0051] It is understandable that the existing pipeline parallel training scheme executed in the cloud-edge collaborative optical network has the following problems:

[0052] (1) Communication overhead: Since data exchange between different stages of the pipeline requires additional communication overhead, which has not been fully considered in the existing technology, it may become a performance bottleneck. In addition, in cloud-edge collaborative optical networks, the distance between the cloud and the edge server is usually far, and the communication time cannot be ignored.

[0053] (2) Load balancing: If the computational complexity of each stage in the pipeline parallel scheme is inconsistent, some devices may be idle while other devices are busy with calculations, resulting in load imbalance.

[0054] (3) Asynchronous pace problem: The parameters and gradients used for updating in existing pipeline parallel solutions do not come from the same iteration. The gradients used for updating may be calculated from the parameters several steps ago. This asynchronous pace is difficult to unify. If the asynchronous pace cannot be precisely controlled, it is difficult to ensure the correct transmission of data.

[0055] Based on this, in order to solve the above problems, the embodiments of the present application respectively provide a pipeline parallel distributed training method for deep neural networks, a pipeline parallel distributed training device for deep neural networks for executing the pipeline parallel distributed training method for deep neural networks, a pipeline parallel distributed training system for deep neural networks, a physical device, a computer-readable storage medium and a computer program product, aiming to accelerate the training of artificial intelligence models by combining computing and communication resources, and reduce the waste of server computing or communication resources caused by the waiting time of service intermediate results. At the same time, the training stability is improved, and high efficiency and low latency are ensured in long-distance transmission and large-scale computing. This solution improves computing efficiency and reduces data transmission delays by executing different stages of model training in parallel on multiple computing nodes. In the cloud-edge collaborative optical network, considering the deterministic delay advantage of optical network transmission, the asynchronous parallel training mode is used to accelerate model training, ensure the stability of asynchronous pace, and reduce the risk of model divergence to meet applications sensitive to low latency.

[0056] The details are described in detail through the following examples.

[0057] Based on this, the embodiment of the present application provides a pipeline parallel distributed training method for a deep neural network that can be implemented by a pipeline parallel distributed training device for a deep neural network, see Figure 1 The pipeline parallel distributed training method for deep neural networks specifically includes the following contents:

[0058] Step 100: For the current iteration round of the deep neural network, execute the preset pipeline parallel distributed training step, wherein the pipeline parallel distributed training step includes: transmitting each group of small batch training data constituting the current batch training data to the edge server via the optical network in sequence, so that the edge server and the cloud server connected to the edge server via the optical network communication can perform asynchronous parallel collaborative training on different subtask models constituting the deep neural network based on the preset communication and training decoupling method, and the edge server can output the gradient corresponding to each group of the small batch training data in sequence.

[0059] In one or more embodiments of the present application, the data used to train the deep neural network can be pre-divided into multiple batches (batch) to form each group of batch training data. Further, each group of batch training data can be divided into multiple small batches (minibatch) to form each small batch training data corresponding to each group of batch training data. It is understandable that the number of data samples contained in the small batch training data is less than the number of data samples in the batch training data. Among them, the type of the data sample can be set according to the application requirements of the deep neural network. For example, if the deep neural network is ultimately used for image classification, the data sample is the image data, which can be specifically set according to the actual application needs, and the present application does not limit this.

[0060] It is understandable that the communication and training decoupling method refers to the edge server and cloud server and other devices involved in the training, which decouple the communication and computing processes and process them asynchronously. For example, while the edge server is receiving the second set of small batch training data of the current batch, it is training the local subtask model based on the first set of small batch training data of the current batch received in advance. Among them, the edge server can adopt one or more.

[0061] It can be understood that the edge server in step 100 and the cloud server connected to the edge server via optical network communication respectively perform asynchronous parallel collaborative training on different subtask models used to constitute the deep neural network based on a preset communication and training decoupling method, and the edge server sequentially outputs the gradients corresponding to each group of the small batch training data. The specific implementation process may include the following contents:

[0062] (1) The edge server receives each group of the small batch training data in sequence, and when receiving a group of the small batch training data for the first time in the current iteration round, starts to execute a preset first asynchronous training step to obtain intermediate activation values ​​corresponding to each group of the small batch training data; and each time an intermediate activation value corresponding to a group of the small batch training data is obtained, the intermediate activation value corresponding to the group of the small batch training data is transmitted to the cloud server via the optical network;

[0063] (2) the cloud server sequentially receives the intermediate activation values ​​corresponding to each group of the small batch training data, and when receiving the intermediate activation values ​​corresponding to a group of the small batch training data for the first time in the current iteration round, starts to execute the preset second asynchronous training step to obtain the intermediate loss values ​​corresponding to each group of the small batch training data; and each time the intermediate loss value corresponding to a group of the small batch training data is obtained, the intermediate loss value corresponding to the group of the small batch training data is transmitted to the edge server via the optical network;

[0064] (3) The edge server receives the intermediate loss values ​​corresponding to each group of the small batch training data in turn, and when receiving the intermediate loss values ​​corresponding to a group of the small batch training data for the first time in the current iteration round, starts to execute the preset third asynchronous training step to obtain the gradients corresponding to each group of the small batch training data respectively; and each time the gradients corresponding to a group of the small batch training data are obtained, the gradients corresponding to the small batch training data are returned to the pipeline parallel distributed training device for the deep neural network via the optical network; wherein, if the current iteration round is not the first iteration round, the edge server alternately executes the third asynchronous training step and the first asynchronous training step.

[0065] Step 200: sequentially receiving the gradients corresponding to the respective groups of the small batch training data returned by the edge server via the optical network.

[0066] From the above description, it can be seen that the pipelined parallel distributed training method for deep neural networks provided in the embodiments of the present application can ensure the correct transmission of data during the deep neural network training process, reduce the communication overhead during the training process, achieve load balancing between the devices participating in the training, and improve the efficiency and effectiveness of deep neural network training as well as the resource utilization of terminal devices, edge servers and cloud servers participating in the deep neural network training.

[0067] In order to further improve the training efficiency of deep neural networks, in a pipeline parallel distributed training method for deep neural networks provided in an embodiment of the present application, see Figure 2 , the pipeline parallel distributed training method for a deep neural network further specifically includes the following contents after step 200:

[0068] Step 300: If the deep neural network has not converged yet, there is the batch training data that has not been used for model training, or the current iteration round is not the preset last iteration round, then when a gradient corresponding to the small batch training data is received for the first time in the current iteration round, the pipeline parallel distributed training step is executed for the next iteration round of the deep neural network.

[0069] It can be understood that in the aforementioned step 100, if the current iteration round is the first iteration round for the deep learning network, then the pipeline parallel distributed training device for the deep neural network will execute step 200 after transmitting each group of small batch training data corresponding to the current batch training data to the edge server via the optical network in sequence, and then wait for several time blocks to start receiving the gradients corresponding to each group of the small batch training data corresponding to the current iteration round returned by the edge server via the optical network, and then execute step 300: if the deep neural network has not converged at present, there is the batch training data that has not been used for model training, or the current iteration round is not the preset last iteration round, then when a gradient corresponding to the small batch training data is received for the first time in the current iteration round, the pipeline parallel distributed training step is executed for the next iteration round of the deep neural network.

[0070] Similarly, in the aforementioned step 100, if the current iteration round is not the first iteration round for the deep learning network, the pipeline parallel distributed training device for the deep neural network will first receive the gradient corresponding to the first group of small batch training data of the previous iteration round, and then immediately start to execute step 100 based on the gradient, and transmit the gradient and the first group of small batch training data corresponding to the current batch training data to the edge server via the optical network, and then the pipeline parallel distributed training device for the deep neural network receives the gradient corresponding to the second group of small batch training data of the previous iteration round, and immediately transmits the second group of small batch training data corresponding to the current batch training data to the edge server via the optical network based on the gradient, and so on. , receiving the gradient of the previous iteration round and sending the small batch training data of the current iteration round are performed alternately in sequence. After sending all the small batch training data corresponding to the current batch training data, immediately execute step 200 to start receiving the gradients corresponding to each group of the small batch training data corresponding to the current iteration round returned by the edge server via the optical network, and then execute step 300: if the deep neural network has not converged at present, there is the batch training data that has not been used for model training, or the current iteration round is not the preset last iteration round, then when a gradient corresponding to the small batch training data is received for the first time in the current iteration round, the pipeline parallel distributed training step is executed for the next iteration round of the deep neural network.

[0071] That is to say, the various steps provided in the embodiments of the present application are not necessarily executed in a time-sequential relationship. Instead, the subsequent step can be executed at the same time as the previous step is started, or the subsequent step can be started during the execution of the previous step. The specific needs are determined according to the conditional limitations recorded in the steps.

[0072] In order to further improve the model training speed and improve the resource utilization of the edge server, in a pipeline parallel distributed training method for a deep neural network provided in an embodiment of the present application, the edge server is used to perform the following:

[0073] Step 01: receiving each group of the small batch training data in sequence, and when receiving a group of the small batch training data for the first time in the current iteration round, starting to execute the preset first asynchronous training step to obtain the intermediate activation values ​​corresponding to each group of the small batch training data respectively; and each time the intermediate activation value corresponding to a group of the small batch training data is obtained, the intermediate activation value corresponding to the group of the small batch training data is transmitted to the cloud server via the optical network;

[0074] Among them, the first asynchronous training step includes: performing forward propagation calculations for the first subtask model corresponding to the deep neural network based on each group of the small batch training data, so that the first subtask model sequentially outputs the intermediate activation values ​​corresponding to each group of the small batch training data.

[0075] In order to further improve the model training speed and improve the resource utilization of the cloud server, in a pipeline parallel distributed training method for a deep neural network provided in an embodiment of the present application, the cloud server is used to perform the following:

[0076] Step 02: receiving the intermediate activation values ​​corresponding to each group of the small batch training data in sequence, and when the intermediate activation values ​​corresponding to a group of the small batch training data are received for the first time in the current iteration round, starting to execute the preset second asynchronous training step to obtain the intermediate loss values ​​corresponding to each group of the small batch training data; and each time the intermediate loss values ​​corresponding to a group of the small batch training data are obtained, the intermediate loss values ​​corresponding to the group of the small batch training data are transmitted to the edge server via the optical network;

[0077] Among them, the second asynchronous training step includes: executing preset calculation steps in sequence for the intermediate activation values ​​corresponding to each group of the small batch training data, and the calculation steps include: based on the intermediate activation values ​​corresponding to the current group of the small batch training data, sequentially executing forward propagation calculations and backward propagation calculations for the second subtask model corresponding to the deep neural network to obtain the intermediate loss values ​​corresponding to the group of small batch training data.

[0078] In order to further improve the model training speed and improve the resource utilization of the edge server, in a pipeline parallel distributed training method for a deep neural network provided in an embodiment of the present application, the edge server is also used to perform the following:

[0079] Step 03: receiving the intermediate loss values ​​corresponding to each group of the small batch training data in sequence, and when receiving the intermediate loss values ​​corresponding to a group of the small batch training data for the first time in the current iteration round, starting to execute the preset third asynchronous training step to obtain the gradients corresponding to each group of the small batch training data respectively; and each time the gradient corresponding to a group of the small batch training data is obtained, the gradient corresponding to the small batch training data is returned through the optical network; wherein, if the current iteration round is not the first iteration round, the edge server alternately executes the third asynchronous training step and the first asynchronous training step;

[0080] Among them, the third asynchronous training step includes: performing back propagation calculations on the first subtask model based on the intermediate loss values ​​corresponding to each group of the small batch training data, so that the first subtask model outputs the gradients corresponding to each group of the small batch training data in turn.

[0081] It can be understood that steps 01 to 03 are executed sequentially to form a complete pipeline parallel distributed training step.

[0082] Specifically, the embodiment of the present application designs a cloud-edge collaborative optical network consisting of a cloud server and multiple edge servers. The training request of the deep neural network can be transmitted through the optical link, forwarded by the router, and processed on multiple servers. Figure 3 , the DNN request model can be divided into two parts, Batch1, Batch2, Batch3 and Batch4 represent different batches of training data. The first part of the deep neural network is deployed on the edge server, and the second part is deployed on the cloud server. Therefore, the training process of the deep neural network can be summarized into four stages:

[0083] 1. The training process of a deep neural network; data transmission from a terminal device to a source edge server.

[0084] 2. Source edge server preprocessing.

[0085] 3. Intermediate data transmission from the source edge server to the target server.

[0086] 4. Target server post-processing.

[0087] Among them, the training process of deep neural network mainly includes three stages: forward propagation, back propagation and parameter update, and the three stages are iteratively carried out.

[0088] In the traditional asynchronous parallel model, each batch updates parameters immediately as needed without waiting for the results of the previous batch. However, the variability of the duration of each batch time block may lead to overly aggressive parameter updates, resulting in an unstable training process. Therefore, maintaining uniform time blocks is critical to ensure data consistency and computation consistency across all servers. Figure 3 In the asynchronous update process of the pipeline shown, each batch of data is trained in an orderly manner, which ensures the stability of the asynchronous steps and reduces the risk of divergence. In the proposed pipeline workflow structure, calculation and communication are performed alternately. Among them, the calculation time includes the forward propagation and backward propagation process of DNN training. Each stage can be described as follows:

[0089] Phase 1: Transmission time between the terminal device and the edge server.

[0090] Phase 2: Computation time at the edge servers, where the DNN layers before the partition point are deployed.

[0091] Phase 3: Transmission time between edge server and cloud server.

[0092] Phase 4: Computation time of the cloud server, deploying the DNN layers after the partition point.

[0093] See also Figure 4 , divides the incoming batch into mini-batches of training data. When the forward propagation of a mini-batch of training data is completed, each stage asynchronously sends the output activations to the next stage while processing another mini-batch of training data. After the forward propagation is completed, the last layer immediately starts backpropagating a mini-batch. After completing the backward propagation, each stage asynchronously sends the gradients to the previous stage while calculating the next mini-batch of training data. In asynchronous parallelism, computation and communication are highly parallel. Asynchronous communication across forward activation and backward gradient stages results in significant overlap between communication and subsequent mini-batch computations. This greatly reduces waiting time. It should be noted that the server's internal storage, which contains large amounts of data such as parameters, will be released when it is no longer needed for future computations. This ensures that the performance of the server is not affected. Figure 4 In the graph, each small square represents a time block.

[0094] In order to further ensure the correct transmission of data during the deep neural network training process and reduce the communication overhead during the training process, in a pipeline parallel distributed training method for a deep neural network provided in an embodiment of the present application, see Figure 2 , the pipeline parallel distributed training method for a deep neural network further specifically includes the following contents before step 100:

[0095] Step 010: Submit a training task for a deep neural network to an optical network controller, so that the optical network controller obtains a target partition point for the deep neural network according to the task requirement data corresponding to the training task and the current resource status of the edge server and the cloud server respectively. The optical network controller divides the deep neural network into a first subtask model and a second subtask model according to the target partition point. The optical network controller assigns the first subtask model to the edge server and the second subtask model to the cloud server, and then returns a corresponding model task partition message.

[0096] The target partition point is also the optimal partition point calculated by the optical network controller.

[0097] Step 020: Receive the model task partition message returned by the optical network controller.

[0098] It can be understood that the optical network controller may specifically be a software defined optical network controller (abbreviated as SDON controller or SDON).

[0099] From the software level, the present application also provides a pipeline parallel distributed training device for a deep neural network for executing all or part of the pipeline parallel distributed training method for a deep neural network, see Figure 5 The pipeline parallel distributed training device for deep neural network specifically includes the following contents:

[0100] The pipeline parallel distributed training module 10 is used to execute the preset pipeline parallel distributed training steps for the current iteration round of the deep neural network, wherein the pipeline parallel distributed training steps include: transmitting each group of small batch training data constituting the current batch training data to the edge server via the optical network in sequence, so that the edge server and the cloud server connected to the edge server via the optical network communication can perform asynchronous parallel collaborative training on different subtask models constituting the deep neural network based on the preset communication and training decoupling method, and the edge server can output the gradient corresponding to each group of the small batch training data in sequence.

[0101] The gradient receiving module 20 is used to sequentially receive the gradients corresponding to each group of the small batch training data returned by the edge server via the optical network.

[0102] The embodiment of the pipelined parallel distributed training device for deep neural networks provided in the present application can be specifically used to execute the processing flow of the embodiment of the pipelined parallel distributed training method for deep neural networks in the above-mentioned embodiment. Its functions are not repeated here, and reference can be made to the detailed description of the above-mentioned embodiment of the pipelined parallel distributed training method for deep neural networks.

[0103] The pipeline parallel distributed training device for deep neural networks can perform pipeline parallel distributed training for deep neural networks in a part that can be completed in the client device. The specific selection can be based on the processing power of the client device and the limitations of the user's usage scenario. This application is not limited to this. If all operations are completed in the client device, the client device may also include a processor for specific processing of pipeline parallel distributed training for deep neural networks.

[0104] The above-mentioned client device (also known as terminal device) may have a communication module (i.e., a communication unit) that can communicate with a remote edge server to achieve data transmission with the edge server. The edge server may include a server on the task scheduling center side, and other implementation scenarios may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The edge server may include a single computer device, or a server cluster consisting of multiple servers, or a server structure of a distributed device.

[0105] The edge server and the client device may communicate with each other using any suitable optical network protocol, including optical network protocols that have not yet been developed on the filing date of this application.

[0106] From the above description, it can be seen that the pipelined parallel distributed training device for deep neural networks provided in the embodiments of the present application can ensure the correct transmission of data during the deep neural network training process, reduce the communication overhead during the training process, achieve load balancing between the devices participating in the training, and improve the efficiency and effectiveness of deep neural network training as well as the resource utilization of terminal devices, edge servers and cloud servers participating in the deep neural network training.

[0107] Based on the aforementioned pipeline parallel distributed training method for deep neural network and / or the pipeline parallel distributed training device for deep neural network, the present application also provides an embodiment of the pipeline parallel distributed training system for deep neural network, see Figure 6 The pipeline parallel distributed training system for deep neural networks specifically includes the following contents:

[0108] Terminal devices, edge servers and cloud servers connected in sequence via optical network communications;

[0109] The terminal device is provided with a pipeline parallel distributed training device for deep neural networks, and the pipeline parallel distributed training device for deep neural networks is used to execute the pipeline parallel distributed training method for deep neural networks provided in the above-mentioned embodiments.

[0110] It is understandable that the pipeline parallel distributed training device for deep neural networks is also communicatively connected to the optical network controller.

[0111] To further illustrate the above embodiments, the present application also provides a specific application example of a pipeline parallel distributed training method for deep neural networks, which belongs to a pipeline parallel distributed training scheme for joint computing and optical network resources. By utilizing the deterministic delay advantage of optical network transmission, the influence of computing and communication delays on task training is fully considered. The determined delay can effectively maintain the pace stability in asynchronous parallel training, thereby improving the convergence speed and reducing the risk of training divergence. Generally, the efficiency of DNN training tasks is determined by the partition point and resource deployment. The computing time of each training task is related to the computing resources and task division method provided on the deployment node, and the time consumed by its communication is related to the bandwidth resources provided and the amount of communication data. The parallel mode also affects the overall efficiency of task training and the network. By decoupling computing and communication, computing tasks and communication tasks can be performed independently, reducing the mutual influence and waiting time between them. This greatly accelerates model training and improves resource utilization.

[0112] In order to achieve cloud-edge collaborative training, the application example of this application outlines the detailed DNN task training process. First, the DNN is stacked by a series of different layers, where the output of one layer is sent to the input of the next layer. For a DNN training request {r|r∈R}, each DNN request r={l1,l2,…,l N}, where N is the total number of layers. For the partition point a of the DNN request r, the front layers {l1,l2,…,l a}Deployed on edge servers s Upper, rear layer {l a+1 ,l a+2 ,…,l N}Deployed on cloud servers c The time complexity of DNN training is mainly concentrated in the forward propagation and backward propagation stages, which is a linear complexity related to the number of network layers and the amount of computation per layer. The parameter update process involves simple parameter update operations, so the update time can be ignored. This application example focuses on the time of DNN forward and backward propagation.

[0113] Forward propagation refers to starting from the input data and calculating the output value of each layer through the neural network until the final result is obtained. For example, on the edge server s s Calculate the first layer l1~l a After the layer, the output value OF la Transfer to cloud servers c , calculate the next layer l a+1 ~l NLayer, get the final model loss. Back propagation refers to calculating the gradient of the loss function for each parameter, returning from the output layer to the input layer layer by layer to adjust and optimize the parameters. For each data sample, the backward phase (i.e., backward propagation using loss to obtain random gradients) starts from the last layer of the DNN. That is, cloud server s c Execute the subsequent l a+1 ~l N In the back propagation phase of the layer, the intermediate data OF la+1 Send to edge servers s Carry out the previous l1~l a The application example of this application is capable of optimizing the computational workload and transmission volume of various tasks by adjusting the sample size of the input data (especially the batch size b).

[0114] Before training a DNN task, it is necessary to determine the optimal partition point of the DNN. Figure 7 As shown in the figure, the terminal first submits the training task to the SDON controller. After receiving the task, SDON analyzes the task requirements and queries the resource status of the current cloud server and edge server. Each computing node of DNN relies heavily on data from the front-end and back-end nodes. However, there is a gap between different time blocks, which means that if the processing time of different stages varies greatly, it will cause idle time and affect the utilization of computing and communication resources. To solve this problem, it is important to balance the time consumed by different stages. In the pipeline workflow model, the training time of DNN depends on the maximum time t of all these stages. block .

[0115]

[0116] According to formula (1), t block , where l i represents the i-th layer of the DNN task, N represents the total number of layers of the DNN task, a represents the model partition point of the DNN task, In the forward propagation stage, l i The output data size of the layer processing 1 sample; Indicates that in the back propagation stage, l i The output data size of a sample processed by the layer; the above are all task-related parameters; in addition, B1 represents the bandwidth from the terminal device to the edge server; B2 represents the bandwidth from the edge server to the cloud server; Indicates processing l on server s i The forward propagation time of 1 sample of the layer; Indicates that l is being i The back propagation time of a layer processing one sample. The above are all network resource status parameters.

[0117] SDON determines the best partition point based on available resources and task requirements. The generated subtask models are then assigned to appropriate cloud and edge servers. Once the model is successfully deployed, training begins. Figure 8 Each minibatch is iteratively trained, alternating between computation and communication. This process continues until the model converges and the training is completed. In theory, the final training time T all It can be calculated according to formula (2), where D is the total training data of the DNN task and b is the data size of the minibatch.

[0118]

[0119] It can be seen that the technical solution provided by the application example of this application makes up for the shortcomings of insufficient computing resources of a single server, and at the same time takes advantage of the large bandwidth of the optical network to achieve cloud-edge collaboration with the optical network to jointly accelerate model training. At the same time, an asynchronous distributed parallel training mode that decouples computing and communication is designed to reduce waiting time. This greatly improves the training efficiency and resource utilization of the model.

[0120] The application example of this application fully utilizes the deterministic delay of optical network transmission, while fully considering the impact of computing and communication on task training. Deterministic delay effectively stabilizes the steps in asynchronous parallelism, thereby accelerating convergence and reducing the risk of training divergence. On this basis, a specific implementation scheme for pipeline parallelism of joint computing and communication is given. Improve the model training speed and greatly increase resource utilization.

[0121] The embodiment of the present application also provides an electronic device, which may include a processor, a memory, a receiver and a transmitter, wherein the processor is used to execute the pipeline parallel distributed training method for a deep neural network mentioned in the above embodiment, wherein the processor and the memory may be connected via a bus or other means, such as by bus connection. The receiver may be connected to the processor and the memory via a wired or wireless manner.

[0122] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0123] As a non-transient computer-readable storage medium, the memory can be used to store non-transient software programs, non-transient computer executable programs and modules, such as the program instructions / modules corresponding to the pipeline parallel distributed training method for deep neural networks in the embodiments of the present application. The processor executes various functional applications and data processing of the processor by running the non-transient software programs, instructions and modules stored in the memory, that is, implementing the pipeline parallel distributed training method for deep neural networks in the above method embodiments.

[0124] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0125] The one or more modules are stored in the memory, and when executed by the processor, perform the pipelined parallel distributed training method for the deep neural network in the embodiment.

[0126] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit, which may include a receiver and a transmitter. The processor, memory, receiver, and transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.

[0127] As an implementation method, the functions of the receiver and the transmitter in the present application can be considered to be implemented through a transceiver circuit or a dedicated chip for transceiver, and the processor can be considered to be implemented through a dedicated processing chip, a processing circuit or a general chip.

[0128] As another implementation method, it is possible to use a general-purpose computer to implement the server provided in the embodiment of the present application, that is, to store the program code for implementing the functions of the processor, receiver, and transmitter in a memory, and the general-purpose processor implements the functions of the processor, receiver, and transmitter by executing the code in the memory.

[0129] The embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned pipeline parallel distributed training method for a deep neural network. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0130] An embodiment of the present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the aforementioned pipeline parallel distributed training method for a deep neural network.

[0131] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0132] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.

[0133] In the present application, features described and / or illustrated for one embodiment may be used in the same manner or in a similar manner in one or more other embodiments, and / or combined with features of other embodiments or replace features of other embodiments.

[0134] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the embodiments of the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A pipeline parallel distributed training method for deep neural networks, characterized in that: include: For the current iteration round of the deep neural network, a preset pipeline parallel distributed training step is executed, wherein the pipeline parallel distributed training step includes: transmitting each group of small batch training data constituting the current batch training data to the edge server via the optical network in sequence, so that the edge server and the cloud server connected to the edge server via the optical network communication perform asynchronous parallel collaborative training on different subtask models constituting the deep neural network based on a preset communication and training decoupling method, and the edge server outputs the gradient corresponding to each group of the small batch training data in sequence; The gradients corresponding to the respective groups of the small batch training data returned by the edge server via the optical network are received in sequence.

2. The pipeline parallel distributed training method for deep neural networks according to claim 1, characterized in that: Also includes: If the deep neural network has not converged yet, there is the batch training data that has not been used for model training, or the current iteration round is not the preset last iteration round, then when a gradient corresponding to the small batch training data is received for the first time in the current iteration round, the pipelined parallel distributed training step is performed for the next iteration round of the deep neural network.

3. The pipeline parallel distributed training method for deep neural network according to claim 1, characterized in that: The edge server is used to execute the following: receiving each group of the small batch training data in sequence, and starting to execute a preset first asynchronous training step when a group of the small batch training data is received for the first time in the current iteration round, so as to obtain the intermediate activation values ​​corresponding to each group of the small batch training data respectively; And each time a group of intermediate activation values ​​corresponding to the small batch training data is obtained, the intermediate activation values ​​corresponding to the group of small batch training data are transmitted to the cloud server via the optical network; Among them, the first asynchronous training step includes: performing forward propagation calculations for the first subtask model corresponding to the deep neural network based on each group of the small batch training data, so that the first subtask model sequentially outputs the intermediate activation values ​​corresponding to each group of the small batch training data.

4. The pipeline parallel distributed training method for deep neural network according to claim 3, characterized in that: The cloud server is used to execute the following: receiving the intermediate activation values ​​corresponding to each group of the small batch training data in sequence, and when the intermediate activation values ​​corresponding to a group of the small batch training data are received for the first time in the current iteration round, starting to execute the preset second asynchronous training step to obtain the intermediate loss values ​​corresponding to each group of the small batch training data respectively; and each time the intermediate loss values ​​corresponding to a group of the small batch training data are obtained, the intermediate loss values ​​corresponding to the group of the small batch training data are transmitted to the edge server via the optical network; Among them, the second asynchronous training step includes: executing preset calculation steps in sequence for the intermediate activation values ​​corresponding to each group of the small batch training data, and the calculation steps include: based on the intermediate activation values ​​corresponding to the current group of the small batch training data, sequentially executing forward propagation calculations and backward propagation calculations for the second subtask model corresponding to the deep neural network to obtain the intermediate loss values ​​corresponding to the group of small batch training data.

5. The pipeline parallel distributed training method for deep neural network according to claim 4, characterized in that: The edge server is also used to execute the following: receiving the intermediate loss values ​​corresponding to each group of the small batch training data in sequence, and when the intermediate loss values ​​corresponding to a group of the small batch training data are received for the first time in the current iteration round, starting to execute the preset third asynchronous training step to obtain the gradients corresponding to each group of the small batch training data respectively; and returning the gradients corresponding to the small batch training data via the optical network each time the gradients corresponding to a group of the small batch training data are obtained; wherein, if the current iteration round is not the first iteration round, the edge server alternately executes the third asynchronous training step and the first asynchronous training step; Among them, the third asynchronous training step includes: performing back propagation calculations on the first subtask model based on the intermediate loss values ​​corresponding to each group of the small batch training data, so that the first subtask model outputs the gradients corresponding to each group of the small batch training data in turn.

6. The pipeline parallel distributed training method for a deep neural network according to any one of claims 1 to 5, characterized in that: Before executing the preset pipeline parallel distributed training steps for the current iteration round of the deep neural network, the method further includes: Submitting a training task for a deep neural network to an optical network controller, so that the optical network controller obtains a target partition point for the deep neural network according to task requirement data corresponding to the training task and the current resource states corresponding to the edge server and the cloud server, respectively; the optical network controller divides the deep neural network into a first subtask model and a second subtask model according to the target partition point; the optical network controller assigns the first subtask model to the edge server and the second subtask model to the cloud server, and then returns a corresponding model task partition message; Receive the model task partition message returned by the optical network controller.

7. A pipeline parallel distributed training device for deep neural networks, characterized in that: include: A pipeline parallel distributed training module is used to execute preset pipeline parallel distributed training steps for the current iteration round of the deep neural network, wherein the pipeline parallel distributed training steps include: transmitting each group of small batch training data constituting the current batch training data to the edge server via the optical network in sequence, so that the edge server and the cloud server connected to the edge server via the optical network communication perform asynchronous parallel collaborative training on different subtask models constituting the deep neural network based on a preset communication and training decoupling method, and the edge server outputs the gradient corresponding to each group of the small batch training data in sequence; The gradient receiving module is used to sequentially receive the gradients corresponding to each group of the small batch training data returned by the edge server via the optical network.

8. A pipelined parallel distributed training system for deep neural networks, characterized in that: include: Terminal devices, edge servers and cloud servers connected in sequence via optical network communications; The terminal device is provided with a pipeline parallel distributed training device for deep neural networks, and the pipeline parallel distributed training device for deep neural networks is used to execute the pipeline parallel distributed training method for deep neural networks described in any one of claims 1 to 6.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the pipelined parallel distributed training method for a deep neural network as described in any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the pipelined parallel distributed training method for a deep neural network as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Assembly line parallel method for accelerating neural network training in heterogeneous GPU cluster

    CN116883229A

  • Unmanned aerial vehicle assisted deep neural network segmentation training method in Internet of Things

    CN116909734A

  • Machine learning method oriented to wireless edge node model segmentation

    CN117035054A

  • Data processing method, device and system, medium and program product

    CN117827418A

  • Training unloading method based on neural network multistage segmentation learning in intelligent Internet of Things

    CN118484301A

Cited By

  • Assembly line parallel training method for multi-modal large model

    CN120179416A

  • A pipeline parallel training method for multimodal large models

    CN120179416B

  • Batch-based parallel splitting federated learning method

    CN120196451A

  • Model asynchronous training method and device of intelligent computing center for providing computing power resources

    CN120508828A

  • Model training method and device, electronic equipment, storage medium and program product

    CN120930831A