Assembly line parallel network, model training method and device and image recognition method and device

By prohibiting the output of routing weights in the control layer of pipeline parallel network, the communication overhead problem caused by routing loss transmission between subnets is solved, and the performance of pipeline parallel model is improved.

CN120031083APending Publication Date: 2025-05-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411834313.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When the hybrid expert model is trained in parallel, the routing loss is passed between each subnet, resulting in an increase in communication overhead and affecting the model performance.

Method used

By prohibiting the output of routing weights in the control layer of pipeline parallel networks, routing losses are avoided to pass between subnets, thereby reducing communication overhead.

Benefits of technology

It effectively reduces the communication overhead of pipeline parallel networks, improves the performance of pipeline parallel models, and realizes a more efficient training and inference process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031083A_ABST
    Figure CN120031083A_ABST
Patent Text Reader

Abstract

The invention provides a pipeline parallel network, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, and can be applied to scenes such as content generation of artificial intelligence. According to the specific implementation scheme, the pipeline parallel network is used for training a pipeline parallel model, the network comprises sub-networks located at different computing devices, and each sub-network comprises a model layer and a control layer; wherein the model layer obtains a routing weight and a hidden state based on the output of the adjacent sub-networks; and the control layer is used for prohibiting the output of the routing weight, so that the communication overhead of the pipeline parallel network is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as computer vision, deep learning, and large models, and can be applied to scenarios such as artificial intelligence content generation, and in particular to a pipeline parallel network, a pipeline parallel model training method and device, an image recognition method and device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Pipeline Parallelism with Mixture-of-Experts (PP-MoE) is a parallel technology for training ultra-large-scale deep learning models. It distributes different parts of the model to different computing devices to reduce the memory consumption of a single computing device, thereby supporting larger-scale model training. Summary of the invention

[0003] The present disclosure provides a pipeline parallel network, a pipeline parallel model training method and device, an image recognition method and device, an electronic device, a computer-readable storage medium, and a computer program product.

[0004] According to a first aspect, a pipeline parallel network is provided for training a pipeline parallel model, wherein the network comprises: sub-networks located on different computing devices, each sub-network comprising: a model layer and a control layer; wherein the model layer obtains routing weights and hidden states based on outputs of adjacent sub-networks; and the control layer is used to prohibit the output of routing weights.

[0005] According to the second aspect, a pipeline parallel model training method is provided, the method comprising: obtaining a training sample set, the training sample set comprising at least one training sample; obtaining a pre-established pipeline parallel network, the pipeline parallel network adopting a pipeline parallel network as described in any implementation method of the first aspect; inputting the training samples selected from the training sample set into the pipeline parallel network to obtain a prediction result output by the pipeline parallel network; based on the prediction result, calculating a network loss value of the pipeline parallel network; based on the network loss value of the pipeline parallel network, training the pipeline parallel network to obtain a pipeline parallel model.

[0006] According to the third aspect, an image recognition method is provided, the method comprising: acquiring an image to be processed; inputting the image to be processed into a pipeline parallel model trained using the method described in any implementation of the second aspect, to obtain a liveness detection result of the image to be processed.

[0007] According to a fourth aspect, a pipeline parallel model training device is provided, which includes: a sample acquisition unit, configured to acquire a training sample set, the training sample set including at least one training sample; a network acquisition unit, configured to acquire a pre-established pipeline parallel network, the pipeline parallel network adopts the pipeline parallel network described in any implementation method of the first aspect; a sample input unit, configured to input the training samples selected from the training sample set into the pipeline parallel network to obtain a prediction result output by the pipeline parallel network; a calculation unit, configured to calculate a network loss value of the pipeline parallel network based on the prediction result; and a network training unit, configured to train the pipeline parallel network based on the network loss value of the pipeline parallel network to obtain a pipeline parallel model.

[0008] According to the fifth aspect, an image recognition device is provided, which includes: an image acquisition unit, configured to acquire an image to be processed; an image input unit, configured to input the image to be processed into a pipeline parallel model trained by the device described in any implementation method of the fourth aspect, to obtain an image recognition result of the image to be processed.

[0009] According to the sixth aspect, an electronic device is provided, which includes: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any implementation of the second aspect or the third aspect.

[0010] According to a seventh aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described in any implementation of the second aspect or the third aspect.

[0011] According to an eighth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the method described in any implementation of the second aspect or the third aspect.

[0012] The pipeline parallel network provided by the embodiment of the present disclosure includes: sub-networks located in different computing devices, each sub-network includes: a model layer and a control layer; wherein the model layer obtains routing weights and hidden states based on the outputs of adjacent sub-networks; and the control layer is used to prohibit the output of routing weights. Thus, by prohibiting the output of routing weights through the control layer, the pipeline parallel network will not transmit routing losses between the sub-networks during training, thereby reducing the communication overhead of the pipeline parallel network and improving the performance of the pipeline parallel model.

[0013] The pipeline parallel model training method and device provided by the embodiment of the present disclosure first obtain a training sample set, the training sample set includes at least one training sample; secondly, obtain a pre-established pipeline parallel network; thirdly, input the training sample selected from the training sample set into the pipeline parallel network to obtain the prediction result output by the pipeline parallel network; thirdly, based on the prediction result, calculate the network loss value of the pipeline parallel network; finally, based on the network loss value of the pipeline parallel network, train the pipeline parallel network to obtain the pipeline parallel model. Therefore, the present disclosure prohibits the output of routing weights through the control layer through the structure of the pipeline parallel network, so that the pipeline parallel network will not transmit the routing loss of the corresponding routing weights between each sub-network during training, thereby reducing the communication overhead of the pipeline parallel network and improving the performance of the pipeline parallel model.

[0014] The image recognition method, image recognition method and device provided by the embodiments of the present disclosure obtain an image to be processed, input the image to be processed into a pipeline parallel model trained by a pipeline parallel model training method, and obtain an image recognition result of the image to be processed. Thus, the image recognition result is generated by using a pipeline parallel model, which improves the reliability and accuracy of the image recognition result.

[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0017] Figure 1 is a schematic diagram of the structure of an embodiment of a pipeline parallel network according to the present disclosure;

[0018] Figure 2 is a flow chart of an embodiment of a pipeline parallel model training method according to the present disclosure;

[0019] Figure 3 is a flow chart of an embodiment of an image recognition method according to the present disclosure;

[0020] Figure 4 is a structural schematic diagram of an embodiment of a pipeline parallel model training device according to the present disclosure;

[0021] Figure 5 is a structural schematic diagram of an embodiment of an image recognition device according to the present disclosure;

[0022] Figure 6It is a block diagram of an electronic device used to implement the pipeline parallel model training method or image recognition method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0024] Pipeline parallelism is a method of distributing different layers of a model to different devices in sequence. By dividing the input data into multiple micro-batches, each device can process the next batch immediately after processing the current batch, thereby improving device utilization. Pipeline parallelism is mainly used in the training of large models. By splitting the model's computing tasks to different devices for execution, it can improve training efficiency and handle larger-scale models.

[0025] Pipeline parallelism can be implemented in various ways, including Gpipe and PipeDream. Gpipe distributes the model's computational tasks to multiple devices, allowing forward propagation and backpropagation to be executed in an overlapping manner, thereby increasing the training speed. PipeDream uses different strategies, such as F-then-B and 1F1B, to meet different training requirements.

[0026] Mixture of Experts (MoE) is an ensemble learning method that combines multiple expert models into an overall model to take advantage of the strengths of each expert model. Each expert model can focus on solving a specific sub-problem, while the overall model can achieve better performance in complex tasks.

[0027] The hybrid expert consists of two key components: the gating network and the expert network. The gating network is responsible for dynamically deciding which expert model should be activated to generate the best prediction based on the characteristics of the input data. The expert network is a group of independent models, each of which is responsible for handling a specific subtask. Through the gating network, the input data will be assigned to the most suitable expert model for processing, and the outputs of different models will be weighted and fused to obtain the final prediction result.

[0028] In the training of a large model that supports mixed experts, each layer of Transformer will generate a routing loss, which is used to train the parameters in the gate of the distribution expert. The routing losses of each layer of Transformer will be superimposed and finally superimposed with the cross entropy loss value. If the mixed expert model is trained using pipeline parallelism, this routing loss will be transmitted between pipeline parallel layers. Because it is ultimately a linear superposition, the communication overhead caused by this transmission is unnecessary.

[0029] In the current implementation, each layer of Transformer has two output values, routing loss and hidden state. The routing loss is accumulated between layers and finally added to the model loss, thus ensuring the contribution of routing loss to the reverse direction.

[0030] After using the pipeline parallel strategy, because each layer of Transformer has two output values, in addition to sending hidden states, routing losses are also sent between different pp-stages. This will introduce pp-stage-1 routing loss transmission, increasing communication overhead.

[0031] In view of the problem that one more routing loss is transmitted between pp-stages in the above pipeline parallel strategy, which increases the communication overhead, the present disclosure first provides a pipeline parallel network for training pipeline parallel models. The control layer in the pipeline parallel network can directly prohibit the output of routing weights to adjacent sub-networks, inhibit the transmission of routing weights of the current sub-network to adjacent sub-networks, and reduce communication overhead. The pipeline parallel network is used to train pipeline parallel models, such as Figure 1 , which is a schematic diagram of the structure of an embodiment of a pipeline parallel network according to the present disclosure, the pipeline parallel network 100 includes: sub-networks 101 located in different computing devices, each sub-network 101 includes: a model layer 1011 and a control layer 1012.

[0032] In this embodiment, the model layer 1011 obtains the routing weight and the hidden state based on the output of the adjacent sub-network; the control layer 1012 is used to prohibit the output of the routing weight.

[0033] In this embodiment, the pipeline parallel network includes multiple sub-networks, each sub-network is located in a different computing device, and the computing device can be a terminal device, such as a mobile phone, a computer, a calculator, etc.

[0034] In this embodiment, each sub-network has the function of a hybrid expert, that is, each sub-network combines multiple expert models to solve a specific sub-problem. Each sub-network can solve a sub-problem, and all sub-networks in the pipeline parallel network can be combined to solve an overall problem.

[0035] In this embodiment, the model layer 1011 in each sub-network is a hybrid expert model, which includes two key components: a gating network and an expert network. The expert network includes multiple expert models, each of which is used to solve a specific problem. The gating network dynamically determines which expert model to use based on the characteristics of the input data, so that the input data is assigned to the most suitable expert model for processing. The routing weight is the information output by the gating network, which is used to characterize the weight size of different sub-networks (such as the above-mentioned expert model). Specifically, in the hybrid expert model, the routing weight is the weight distribution ratio when distributing input data between different expert networks. The hybrid expert model balances the weight distribution of different expert networks through a gating mechanism and a load balancing strategy, thereby improving the training efficiency and overall performance of the model.

[0036] In this embodiment, in the hybrid expert model, hidden states refer to variables or parameters that are in the model but not directly observed. These hidden states are usually determined by the expert network and the gating network to process and distribute input data to different experts.

[0037] In this embodiment, the model layer is used to realize the role of the hybrid expert model, and the control layer is a unit in the pipeline parallel network that prohibits the output of routing weights. The parameters of this unit do not change with the parameters of the pipeline parallel network, and it only has the function of prohibiting routing weights and passing hidden states.

[0038] In this embodiment, the control layer can directly prohibit the routing weight from being output to the adjacent sub-network during the forward propagation of the pipeline parallel network, wherein forward propagation refers to the information transmission process from the input layer to the output layer in a neural network. Specifically, forward propagation is the process of calculating and transmitting data layer by layer through input data and current model parameters to finally obtain the output result of the model.

[0039] The pipeline parallel network provided by the present disclosure includes: sub-networks located in different computing devices, each sub-network includes: a model layer and a control layer; wherein the model layer obtains routing weights and hidden states based on the outputs of adjacent sub-networks; and the control layer is used to prohibit the output of routing weights. Thus, by prohibiting the output of routing weights through the control layer, the pipeline parallel network will not transmit routing losses between the sub-networks during training, thereby reducing the communication overhead of the pipeline parallel network and improving the performance of the pipeline parallel model.

[0040] In some optional implementations of the present disclosure, the model layer includes: a routing layer and a multi-layer perceptron, the routing layer is used to output routing weights; the multi-layer perceptron obtains a hidden state based on the routing weights.

[0041] In this optional implementation, the routing layer is connected to the multi-layer perceptron, and the routing layer outputs the obtained routing weights to the multi-layer perceptron. The routing layer has the same function as the gated network in the hybrid expert model, and outputs the routing weights based on the input data; the multi-layer perceptron has the same function as the expert network in the hybrid expert model, and obtains the hidden state based on the routing weights output by the routing layer.

[0042] The feature recognition subnetwork provided by this optional implementation method includes a routing layer and a multi-layer perceptron. The routing layer is used to output routing weights; the multi-layer perceptron obtains hidden states based on routing weights. This provides a reliable implementation method for the model layer.

[0043] In some optional implementations of the present disclosure, the control layer of each sub-network is also used to generate a tensor of one for the gradient corresponding to the routing weight of the sub-network during backward propagation.

[0044] In this optional implementation, the control layer prohibits the routing weights from being output to adjacent sub-networks, so that during forward propagation, the routing losses corresponding to the routing weights of each sub-network will not be propagated to adjacent sub-networks.

[0045] In this optional implementation, back propagation is an algorithm for training multi-layer feedforward neural networks. It combines the gradient descent method and the chain rule to calculate the partial derivative of the loss function with respect to each parameter layer by layer, thereby updating the weights and biases in the network to minimize the loss function and make the output of the model close to the expected result.

[0046] In this optional implementation, routing loss is the loss caused by the transmission of routing weights to adjacent sub-networks when training the model layer. Specifically, first, each layer of sub-network will generate a routing loss, and the routing loss is superimposed between sub-network layers. Since the pipeline parallel network divides the sub-networks into different computing devices by layer, the process of calculating the routing loss of the Nth sub-network will depend on the result of calculating the routing loss of the N-1th sub-network, and the model parameters and calculation process of the N-1th sub-network occur on the computing device corresponding to the previous pipeline parallel level, so a communication operation will definitely be generated, where the sub-network is a pipeline parallel level.

[0047] In this optional implementation, the control layer is used to reduce the communication operations sent from the N-1th pipeline parallel stage to the Nth pipeline parallel stage (the control layer prohibits the transmission of routing weights between each pipeline parallel stage to achieve the reduction of communication operations). Since the routing loss corresponding to the routing weights of all pipeline parallel stages of the pipeline parallel network is a linear sum value, that is, the routing of each pipeline parallel stage is ultimately linearly superimposed on the loss of the pipeline parallel network, each pipeline parallel stage maintains a local routing loss result, and during back propagation, a tensor with all gradients of 1 is generated for the routing losses of each pipeline parallel stage, so that a fixed value can be generated for the loss corresponding to the routing weights of each pipeline parallel stage, thereby ensuring that the back propagation is carried out correctly.

[0048] Based on this, the present invention proposes a pipeline parallel model training method. Figure 2 A process 200 of an embodiment of a pipeline parallel model training method according to the present disclosure is shown. The pipeline parallel model training method comprises the following steps:

[0049] Step 201: Obtain a training sample set.

[0050] In this embodiment, the execution subject on which the pipeline parallel model training method runs can obtain the training sample set in a variety of ways. For example, the execution subject can obtain the training sample set stored in the database server through a wired connection or a wireless connection. For another example, the user can obtain the training sample set collected by the terminal by communicating with the terminal.

[0051] Here, the training sample set may include at least one training sample, and the content of the training sample is different based on the different multimodal tasks to be implemented by the pipeline parallel network. For example, if the pipeline parallel network is a multimodal network that can implement image recognition, the training sample includes a sample image and an annotated image related to the sample image. The annotated image includes image description text that describes various objects, scenes, and styles in the image, as well as classification information of various objects in the image as living objects or attack objects. When training the pipeline parallel network, the sample image in the training sample can be input into the pipeline parallel network, and the loss calculation is performed using the annotated image and the output result of the pipeline parallel network.

[0052] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of the training sample sets involved are carried out after authorization and comply with relevant laws and regulations.

[0053] Step 202, obtaining a pre-established pipeline parallel network.

[0054] The pipeline parallel network includes: sub-networks located in different computing devices, each sub-network includes: a model layer and a control layer; the model layer obtains routing weights and hidden states based on the outputs of adjacent sub-networks; the control layer is used to prohibit the output of routing weights.

[0055] In step 203, the training samples selected from the training sample set are input into the pipeline parallel network to obtain the prediction results output by the pipeline parallel network.

[0056] In this embodiment, the execution subject can select training samples from the training sample set obtained in step 201, and execute the training steps from step 203 to step 205 to complete an iterative training of the pipeline parallel network. Among them, the selection method and the number of selected training samples from the training sample set are not limited in this application, and the number of iterative training of the pipeline parallel network is not limited. For example, in one iterative training, multiple continuous training samples can be randomly selected, and the network loss value of the pipeline parallel network is calculated through the selected training samples, and the parameters of the pipeline parallel network are adjusted.

[0057] In this embodiment, the prediction result is the actual prediction output of the pipeline parallel network, and the prediction result may be different based on the different tasks of the pipeline parallel network. For example, if the task of the pipeline parallel network is to perform image recognition, the prediction result is the image recognition result after the image is recognized; if the task of the pipeline parallel network is speech recognition, the prediction result is the result obtained after the pipeline parallel network performs speech recognition on the speech. The pipeline parallel network is trained once per iteration and outputs a prediction result.

[0058] In this embodiment, the inhibitory effect of the control layer of each sub-network in the pipeline parallel network on the routing weight does not affect the actual output of the pipeline parallel network; even during the backward propagation of the pipeline parallel network, directly setting the tensor corresponding to the routing weight to 1 will not affect the generation of the prediction results of the pipeline parallel network.

[0059] Step 204, based on the prediction result, calculate the network loss value of the pipeline parallel network.

[0060] In this embodiment, during each iterative training of the pipeline parallel network, a training sample is selected from the training sample set, and the selected training sample is input into the pipeline parallel network. Based on the prediction result output by the pipeline parallel network and the true value in the training sample, the network loss value of the pipeline parallel network is calculated.

[0061] In this embodiment, the loss function of the pipeline parallel network can adopt a cross entropy loss function, which can measure the difference between two different probability distributions in the same random variable, and is represented as the difference between the true probability distribution and the predicted probability distribution in machine learning. The smaller the value of the cross entropy loss function, the better the prediction effect of the pipeline parallel network (i.e., the pipeline parallel model).

[0062] Optionally, the loss function of the pipeline parallel network can adopt a mean square error function, which is the expectation of the square of the difference between the predicted value (estimated value) and the true value of the feature recognition subnetwork and the classification subnetwork. During the iterative training process of the feature recognition subnetwork and the classification subnetwork, the gradient descent algorithm can be used to minimize the loss function of the feature recognition subnetwork and the classification subnetwork, thereby iteratively optimizing the network parameters of the feature recognition subnetwork and the classification subnetwork.

[0063] The original meaning of gradient is a vector, which means that the directional derivative of a loss function at that point reaches the maximum value along that direction, that is, the loss function changes fastest along that direction at that point and the rate of change is the largest. In deep learning, the main task of neural networks is to find the optimal network parameters (weights and biases) during learning. The optimal network parameters are also the parameters when the loss function is minimized.

[0064] Step 205: Based on the network loss value of the pipeline parallel network, train the pipeline parallel network to obtain a pipeline parallel model.

[0065] In this embodiment, the pipeline parallel model is obtained by inputting the selected training samples into the pipeline parallel network for iterative training and adjusting the parameters of the pipeline parallel network. It should be noted that before adjusting the parameters of the pipeline parallel network, the gradient corresponding to the routing weight can be directly set to 1 for the tensor through the model layer of each sub-network in the pipeline parallel network.

[0066] In this embodiment, the network loss value of the pipeline parallel network can be used to detect whether the pipeline parallel network meets the training completion condition. After the pipeline parallel network meets the training completion condition, a trained pipeline parallel model is obtained.

[0067] In this embodiment, the training completion condition includes: the network loss value of the pipeline parallel network is less than a first loss value threshold. The first loss threshold can be determined based on specific training requirements, for example, the first loss threshold is 0.01.

[0068] Optionally, in this embodiment, in response to the pipeline parallel network not meeting the training completion condition, the relevant parameters in the pipeline parallel network are adjusted so that the network loss value of the pipeline parallel network converges, and the above training steps 203-205 are continued to be executed based on the adjusted pipeline parallel network.

[0069] In this embodiment, since the image semantic recognition subnetwork is pre-trained with massive data and migrated to living tasks through fine-tuning, the generalization effect of the model is greatly improved compared with traditional methods. The image semantic recognition subnetwork can be removed during the reasoning stage, so that the obtained pipeline parallel model will not increase too much reasoning time during reasoning.

[0070] In this embodiment, when the pipeline parallel network does not meet the training completion condition, adjusting the relevant parameters of the pipeline parallel network is helpful to help the network loss value of the pipeline parallel network converge.

[0071] In this embodiment, since a control layer is used in the pipeline parallel network, the control layer prohibits the transmission of routing weights to adjacent sub-networks, and the routing loss generated by the routing weights can be removed. Since the routing loss is a linear sum value, each sub-network maintains a local routing loss result, and during back propagation, a tensor with all gradients of 1 is generated for the routing weights, thereby reducing the communication overhead during the training process. It has been measured that the pipeline parallel model can have a performance improvement of about 2%.

[0072] The pipeline parallel model training method provided by the embodiment of the present disclosure first obtains a training sample set, the training sample set includes at least one training sample; secondly, obtains a pre-established pipeline parallel network, and again, inputs the training sample selected from the training sample set into the pipeline parallel network to obtain the prediction result output by the pipeline parallel network; then, based on the prediction result, calculates the network loss value of the pipeline parallel network; finally, based on the network loss value of the pipeline parallel network, trains the pipeline parallel network to obtain the pipeline parallel model. Therefore, the present disclosure prohibits the output of routing weights through the control layer through the structure of the pipeline parallel network, so that the pipeline parallel network will not transmit the routing loss of the corresponding routing weights between each sub-network during training, thereby reducing the communication overhead of the pipeline parallel network and improving the performance of the pipeline parallel model.

[0073] In some optional implementations of the present disclosure, the above-mentioned inputting the training samples selected from the training sample set into the pipeline parallel network to obtain the prediction results output by the pipeline parallel network includes: inputting the training samples selected from the training sample set into the first sub-network to obtain the prediction results output by the last sub-network; wherein, during forward propagation, each sub-network of the pipeline parallel network only outputs hidden states to subsequent adjacent sub-networks.

[0074] In this optional implementation, the first sub-network is a sub-network that receives input data. In order to facilitate the first sub-network to understand the input data, the first sub-network generally has an embedding layer. Based on the first sub-network including the embedding layer, the embedding layer is a technical layer that maps high-dimensional data to a low-dimensional space, which makes the data easier to process and analyze in the low-dimensional space. In fields such as natural language processing and computer vision, embedding layer technology is widely used, such as mapping words, phrases or images to fixed-length vectors that can capture word meanings, contextual information or image features.

[0075] In this optional implementation, in order to facilitate the generation of prediction results, the last sub-network has a conversion layer that converts the hidden state into the prediction result. The conversion layer can be a convolutional network layer, such as a connection layer.

[0076] The method for obtaining the prediction result provided by this optional implementation inputs the training samples selected from the training sample set into the first sub-network to obtain the prediction result output by the last sub-network; wherein, during forward propagation, each sub-network of the pipeline parallel network only outputs the hidden state to the subsequent adjacent sub-network, so that the transmission of the hidden state from the first sub-network to the sub-network in the middle of the last sub-network can be done without the transmission of routing weight parameters, thereby reducing the communication overhead of the pipeline parallel network and improving the performance of the pipeline parallel model.

[0077] In some optional implementations of the present disclosure, the above-mentioned calculation of the network loss value of the pipeline parallel network based on the prediction results includes: obtaining the loss function of each sub-network; calculating the loss value of each sub-network based on the selected training samples and loss function; and obtaining the network loss value of the pipeline parallel network based on the loss values ​​of all sub-networks.

[0078] In this optional implementation, the above-mentioned calculation of the loss value of each sub-network based on the selected training samples and loss function includes: inputting the selected training samples into the first sub-network, obtaining the prediction result output by the last sub-network, and calculating the loss value of each sub-network based on the prediction result and the loss function of each sub-network.

[0079] The above-mentioned obtaining of the network loss value of the pipeline parallel network based on the loss values ​​of all sub-networks may include: assigning weight values ​​to each sub-network respectively, multiplying the loss value of each sub-network with its corresponding weight value and adding them together to obtain the network loss value of the pipeline parallel network.

[0080] The method for calculating the network loss value of a pipelined parallel network provided in this embodiment obtains the loss function of each sub-network; calculates the loss value of each sub-network based on the selected training samples and loss function; obtains the network loss value of the pipelined parallel network based on the loss values ​​of all sub-networks, providing a reliable implementation method for obtaining the network loss value.

[0081] Optionally, the above-mentioned calculation of the network loss value of the pipeline parallel network based on the prediction results includes: obtaining the loss function of the pipeline parallel network; and obtaining the network loss value of the pipeline parallel network based on the prediction results corresponding to the selected training samples and the loss function of the pipeline parallel network.

[0082] In some optional implementations of the present disclosure, the above-mentioned pipeline parallel model training method also includes: when training the pipeline parallel network, no parameter update is performed on the control layer of each sub-network.

[0083] In this optional implementation, the parameters of the control layer of each sub-network have no effect on the training of the pipeline parallel network, and when training the pipeline parallel network, not updating the parameters of the control layer of each sub-network can effectively save the training steps and training time of the pipeline parallel network.

[0084] The pipeline parallel model training method provided in this embodiment does not update the parameters of the control layer of each sub-network when training the pipeline parallel network, which further saves the training time of the pipeline parallel network and reduces the pipeline parallel model training steps.

[0085] In some optional implementations of the present disclosure, when training a pipeline parallel network, the back-propagation gradient corresponding to the loss function controlling each sub-network does not pass through the control layer of each sub-network.

[0086] In this optional implementation, the back-propagation gradient corresponding to the loss function of each sub-network does not pass through the control layer of each sub-network, so that each sub-network only receives the learning information brought by the loss function of the model layer, which can effectively avoid the influence of the control layer on the training of the pipeline parallel model.

[0087] The pipeline parallel model training method provided in this embodiment controls the back-propagation gradient corresponding to the loss function of each sub-network without passing through the control layer of each sub-network when training the pipeline parallel network, thereby avoiding the influence of the control layer on the pipeline parallel model training and improving the training effect of the pipeline parallel model.

[0088] Furthermore, based on the pipeline parallel model training method provided in the above embodiment, the present disclosure also provides an embodiment of an image recognition method. The image recognition method of the present disclosure combines artificial intelligence fields such as computer vision and deep learning.

[0089] See also Figure 3 , shows a process 300 according to an embodiment of the image recognition method of the present disclosure. The image recognition method provided in this embodiment includes the following steps:

[0090] Step 301: Obtain an image to be processed.

[0091] In this embodiment, the image to be processed may include information such as people, objects, and scenery. The image to be processed is processed by a pipeline parallel model to obtain an image recognition result. The execution subject of the image recognition method can obtain the image to be processed in a variety of ways. For example, the execution subject can obtain the image to be processed stored in the database server through a wired connection or a wireless connection. For another example, the execution subject can also receive the image to be processed collected by a terminal or other device in real time.

[0092] In this embodiment, the image recognition result is a result that displays the subject in the image to be processed. Specifically, the image recognition result includes: information about different types of subjects in the image, and the subject information may include the name and type of the subject, wherein the type of the subject can be reflected by the result probability value output by the pipeline parallel model.

[0093] Step 302: input the image to be processed into the pipeline parallel model to obtain the image recognition result of the image to be processed.

[0094] In this embodiment, the execution subject can input the image to be processed obtained from step 301 into the pipeline parallel model to obtain the image recognition result of the image to be processed. The image recognition result is used to indicate whether the image to be processed has a subject and the subject's information.

[0095] In this embodiment, the pipeline parallel model can be adopted as described above. Figure 2 The specific training process can be found in Figure 2 The relevant description of the embodiments will not be repeated here.

[0096] The image recognition method provided by the embodiment of the present disclosure obtains an image to be processed, inputs the image to be processed into a pipeline parallel model generated by a pipeline parallel model training method, and obtains an image recognition result of the image to be processed. Thus, the image recognition result is generated by using a pipeline parallel model, which improves the reliability and accuracy of image recognition.

[0097] Further references Figure 4 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a pipeline parallel model training device. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0098] like Figure 4As shown, the pipeline parallel model training device 400 provided in this embodiment includes: a sample acquisition unit 401, a network acquisition unit 402, a sample input unit 403, a calculation unit 404, and a network training unit 405. Among them, the above-mentioned sample acquisition unit 401 can be configured to obtain a training sample set, and the training sample set includes at least one training sample. The above-mentioned network acquisition unit 402 can be configured to obtain a pre-established pipeline parallel network, and the pipeline parallel network adopts the pipeline parallel network of the above-mentioned embodiment. The above-mentioned sample input unit 403 can be configured to input the training sample selected from the training sample set into the pipeline parallel network to obtain the prediction result output by the pipeline parallel network. The above-mentioned calculation unit 404 can be configured to calculate the network loss value of the pipeline parallel network based on the prediction result. The above-mentioned network training unit 405 can be configured to train the pipeline parallel network based on the network loss value of the pipeline parallel network to obtain a pipeline parallel model.

[0099] In this embodiment, the specific processing and technical effects of the pipeline parallel model training device 400: the sample acquisition unit 401, the network acquisition unit 402, the sample input unit 403, the calculation unit 404, and the network training unit 405 can be referred to respectively. Figure 2 The relevant descriptions of step 201, step 202, step 203, step 204, and step 205 in the corresponding embodiment are not repeated here.

[0100] In some optional implementations of this embodiment, the sample input unit 403 is further configured to: input the training samples selected from the training sample set into the first sub-network to obtain the prediction result output by the last sub-network; wherein, during forward propagation, each sub-network of the pipeline parallel network only outputs the hidden state to the subsequent adjacent sub-network.

[0101] In some optional implementations of this embodiment, the above-mentioned computing unit 404 is further configured to: obtain the loss function of each sub-network; calculate the loss value of each sub-network based on the selected training samples and loss function; and obtain the network loss value of the pipeline parallel network based on the loss values ​​of all sub-networks.

[0102] In some optional implementations of this embodiment, the above-mentioned device 400 also includes: a first control unit (not shown in the figure), and the above-mentioned first control unit is configured not to update the parameters of the control layer of each sub-network when training the pipeline parallel network.

[0103] In some optional implementations of this embodiment, the above-mentioned device 400 also includes: a second control unit (not shown in the figure), and the above-mentioned second control unit is configured to control the back propagation gradient corresponding to the loss function of each sub-network to not pass through the control layer of each sub-network when training the pipeline parallel network.

[0104] The pipeline parallel model training device provided by the embodiment of the present disclosure, first, the sample acquisition unit 401 acquires a training sample set, and the training sample set includes at least one training sample; secondly, the network acquisition unit 402 acquires a pre-established pipeline parallel network; thirdly, the sample input unit 403 inputs the training sample selected from the training sample set into the pipeline parallel network to obtain the prediction result output by the pipeline parallel network; thirdly, the calculation unit 404 calculates the network loss value of the pipeline parallel network based on the prediction result; finally, the network training unit 405 trains the pipeline parallel network based on the network loss value of the pipeline parallel network to obtain the pipeline parallel model. Therefore, the present disclosure prohibits the output of routing weights through the control layer through the structure of the pipeline parallel network, so that the pipeline parallel network will not transmit the routing loss of the corresponding routing weights between each sub-network during training, thereby reducing the communication overhead of the pipeline parallel network and improving the performance of the pipeline parallel model.

[0105] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an image recognition device, which is similar to Figure 5 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0106] like Figure 5 As shown, the image recognition device 500 provided in this embodiment includes: an image acquisition unit 501, and an image input unit 502. The image acquisition unit 501 can be configured to acquire an image to be processed. The image input unit 502 can be configured to input the image to be processed into the image input unit 502 as described above. Figure 4 In the pipeline parallel model generated by the device described in the embodiment, an image recognition result of the image to be processed is obtained.

[0107] In this embodiment, in the image recognition device 500, the specific processing of the image acquisition unit 501 and the image input unit 502 and the technical effects thereof can be referred to in Figure 3 The relevant descriptions of step 301 and step 302 in the corresponding embodiment are not repeated here.

[0108] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0109] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0110] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0111] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0112] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0113] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as a pipeline parallel model training method or an image recognition method. For example, in some embodiments, the pipeline parallel model training method or the image recognition method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the pipeline parallel model training method or the image recognition method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the pipeline parallel model training method or the image recognition method in any other appropriate manner (eg, by means of firmware).

[0114] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0115] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable pipeline parallel model training device, an image recognition device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.

[0116] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0118] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0119] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0120] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0121] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A pipeline parallel network for training a pipeline parallel model, wherein the network includes sub-networks located on different computing devices, each sub-network including a model layer and a control layer; in, The model layer obtains routing weights and hidden states based on the outputs of adjacent sub-networks; the control layer is used to prohibit the output of the routing weights.

2. The network according to claim 1, wherein: The model layer includes a routing layer and a multi-layer perceptron, wherein the routing layer is used to output routing weights; and the multi-layer perceptron is used to obtain a hidden state based on the routing weights.

3. The network according to claim 1, wherein: The control layer of each sub-network is also used to generate a tensor of one for the gradient corresponding to the routing weight of the sub-network during backward propagation.

4. A pipeline parallel model training method, the method comprising: Acquire a training sample set, wherein the training sample set includes at least one training sample; Obtaining a pre-established pipeline parallel network as described in any one of claims 1 to 3; Inputting the training samples selected from the training sample set into the pipeline parallel network to obtain the prediction results output by the pipeline parallel network; Based on the prediction result, calculating the network loss value of the pipeline parallel network; Based on the network loss value of the pipeline parallel network, the pipeline parallel network is trained to obtain a pipeline parallel model.

5. The method according to claim 4, wherein: The step of inputting the training samples selected from the training sample set into the pipeline parallel network to obtain the prediction results output by the pipeline parallel network includes: The training samples selected from the training sample set are input into the first sub-network to obtain the prediction result output by the last sub-network; wherein, during forward propagation, each sub-network of the pipeline parallel network only outputs the hidden state to the subsequent adjacent sub-network.

6. The method according to claim 4, wherein: The calculating the network loss value of the pipeline parallel network based on the prediction result includes: Get the loss function of each sub-network; Based on the selected training samples and the loss function, calculating the loss value of each sub-network; Based on the loss values ​​of all sub-networks, a network loss value of the pipeline parallel network is obtained.

7. The method according to any one of claims 4 to 6, wherein: The method further comprises: When training the pipeline parallel network, no parameter update is performed on the control layer of each sub-network.

8. The method according to any one of claims 4 to 6, further comprising: When training the pipeline parallel network, the back-propagated gradient corresponding to the loss function controlling each sub-network does not pass through the control layer of each sub-network.

9. An image recognition method, the method comprising: Get the image to be processed; The image to be processed is input into a pipeline parallel model trained by the method according to any one of claims 4 to 8 to obtain an image recognition result of the image to be processed.

10. A pipeline parallel model training device, the device comprising: A sample acquisition unit, configured to acquire a training sample set, wherein the training sample set includes at least one training sample; A network acquisition unit, configured to acquire a pre-established pipeline parallel network, wherein the pipeline parallel network adopts the pipeline parallel network according to any one of claims 1 to 3; A sample input unit, configured to input a training sample selected from the training sample set into the pipeline parallel network to obtain a prediction result output by the pipeline parallel network; A calculation unit, configured to calculate a network loss value of the pipeline parallel network based on the prediction result; The network training unit is configured to train the pipeline parallel network based on the network loss value of the pipeline parallel network to obtain a pipeline parallel model.

11. The device according to claim 10, wherein: The sample input unit is further configured to: input the training sample selected from the training sample set into the first sub-network to obtain the prediction result output by the last sub-network; wherein, during forward propagation, each sub-network of the pipeline parallel network only outputs the hidden state to the subsequent adjacent sub-network.

12. The device according to claim 10, wherein: The computing unit is further configured to: obtain the loss function of each sub-network; calculate the loss value of each sub-network based on the selected training samples and the loss function; and obtain the network loss value of the pipeline parallel network based on the loss values ​​of all sub-networks.

13. The device according to any one of claims 10 to 12, further comprising: The first control unit is configured not to update the parameters of the control layer of each sub-network when training the pipeline parallel network.

14. The device according to any one of claims 10 to 12, further comprising: The second control unit is configured to control the back propagation gradient corresponding to the loss function of each sub-network to not pass through the control layer of each sub-network when training the pipeline parallel network.

15. An image recognition device, comprising: An image acquisition unit, configured to acquire an image to be processed; The image input unit is configured to input the image to be processed into a pipeline parallel model trained by the device according to any one of claims 10 to 14 to obtain an image recognition result of the image to be processed.

16. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 4 to 9.

17. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 4 to 9.

18. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 4 to 9.