Model training method and device, computer equipment and storage medium

By communicating and fusion of compressed gradient sets between multiple processors, the problem of inefficiency in training large-scale machine learning models is solved, and training efficiency and model prediction capabilities are improved.

CN120257064APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410004156.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

As the scale of model parameters of machine learning models expands, the difficulty of model training increases, resulting in longer training time and inefficient training.

Method used

By obtaining the original gradient set of the initial model on the target training subset, gradient filtering is performed to obtain the compressed gradient set, and communicate and fuse these compressed gradient sets between multiple processors to update the model parameter set until the training end condition is met.

Benefits of technology

It improves the efficiency and quality of model training, reduces the consumption of communication resources, and enhances the prediction ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257064A_ABST
    Figure CN120257064A_ABST
Patent Text Reader

Abstract

The invention relates to a model training method and device, computer equipment, a storage medium and a computer program product. The embodiment of the invention can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like. The method comprises the steps of obtaining an original gradient set based on a model parameter set corresponding to an initial model and a target training subset; performing gradient screening on the original gradient set to obtain a compression gradient set corresponding to the initial model on the target training subset; sending a corresponding compression gradient set of the initial model on the target training subset to other target processors used for training the initial model, and obtaining corresponding compression gradient sets of the initial model on other training subsets from the target processors; fusing the compression gradient sets to obtain a comprehensive gradient set used for updating a model parameter set corresponding to the initial model; and after the model parameters are updated, carrying out iterative training until a training ending condition is met, and obtaining a target model. By adopting the method, the model training efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a model training method, apparatus, computer device, storage medium, and computer program product. Background Art

[0002] With the development of computer technology, artificial intelligence technology has emerged. In more and more life scenarios, machine learning models can be used to provide model services to meet corresponding business needs.

[0003] In related technologies, a machine learning model is trained with a large number of training samples so that the trained machine learning model can accurately predict and classify new samples. However, as the scale of the model parameters of the machine learning model expands, the difficulty of model training increases, and the time required for training also increases accordingly, resulting in the problem of low model training efficiency. Summary of the Invention

[0004] Based on this, it is necessary to provide a model training method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the model training efficiency for the above technical problems.

[0005] On the one hand, this application provides a model training method, including:

[0006] Obtain a target training subset corresponding to the first processor;

[0007] Based on the model parameter set corresponding to the initial model to be trained and the target training subset, obtain an original gradient set corresponding to the initial model on the target training subset;

[0008] Perform gradient screening on the original gradient set to obtain a compressed gradient set corresponding to the initial model on the target training subset;

[0009] Send the compressed gradient set corresponding to the initial model on the target training subset to other target processors for training the initial model, and obtain the compressed gradient set corresponding to the initial model on other training subsets from the target processors;

[0010] Fuse each compressed gradient set to obtain a comprehensive gradient set; the comprehensive gradient set is used to update the model parameter set corresponding to the initial model;

[0011] After the model parameter set corresponding to the initial model is updated, return to the step of obtaining the target training subset corresponding to the first processor and execute until the training end condition is met to obtain a target model.

[0012] On the other hand, this application also provides a model training apparatus, including:

[0013] A training data acquisition module for acquiring a target training subset corresponding to a first processor;

[0014] A gradient determination module for obtaining an original gradient set corresponding to the initial model on the target training subset based on a model parameter set corresponding to the initial model to be trained and the target training subset;

[0015] A gradient compression module for performing gradient screening on the original gradient set to obtain a compressed gradient set corresponding to the initial model on the target training subset;

[0016] A gradient interaction module for sending the compressed gradient set corresponding to the initial model on the target training subset to other target processors for training the initial model, and obtaining the compressed gradient set corresponding to the initial model on other training subsets from the target processors;

[0017] A gradient fusion module for fusing each compressed gradient set to obtain a comprehensive gradient set; the comprehensive gradient set is used to update the model parameter set corresponding to the initial model;

[0018] A target model determination module for, after the model parameter set corresponding to the initial model is updated, returning to execute the step of acquiring the target training subset corresponding to the first processor until a training end condition is met, to obtain a target model.

[0019] On the one hand, the present application further provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above model training method are implemented.

[0020] On the one hand, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above model training method are implemented.

[0021] On the one hand, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above model training method are implemented.

[0022] The above model training method, device, computer device, storage medium, and computer program product obtain a target training subset corresponding to a first processor, obtain an original gradient set corresponding to the initial model on the target training subset based on the model parameter set corresponding to the initial model to be trained and the target training subset, perform gradient screening on the original gradient set to obtain a compressed gradient set corresponding to the initial model on the target training subset, send the compressed gradient set corresponding to the initial model on the target training subset to other target processors for training the initial model, obtain the compressed gradient sets corresponding to the initial model on other training subsets from the target processors, fuse the compressed gradient sets to obtain a comprehensive gradient set, and the comprehensive gradient set is used to update the model parameter set corresponding to the initial model. After the model parameter set corresponding to the initial model is updated, return to the step of obtaining the target training subset corresponding to the first processor and execute until the training end condition is met to obtain the target model. In this way, each processor for training the initial model performs parallel model training on the initial model based on the model parameter set corresponding to the initial model and its respective training subset, which can improve the model training efficiency. When performing model training, calculate the original gradient set corresponding to the initial model on the training subset, compress the original gradient set to obtain the compressed gradient set corresponding to the initial model on the training subset, and each processor communicates its own compressed gradient set. The data volume of the compressed gradient set is smaller than that of the original gradient set, which can improve the communication efficiency and thus further improve the model training efficiency. The processor fuses its own compressed gradient set with the compressed gradient sets of other processors to obtain a comprehensive gradient set. The comprehensive gradient set integrates the gradient information of the initial model on each training subset, and adjusts the model parameters based on the comprehensive gradient set, enabling the model to learn knowledge from each training subset, which can improve the prediction ability of the model and the quality of model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is an application environment diagram of the model training method in an embodiment;

[0025] Figure 2 It is a flowchart of the model training method in an embodiment;

[0026] Figure 3 It is a schematic diagram of the model training process in an embodiment;

[0027] Figure 4Schematic diagram of communicating gradients in one embodiment;

[0028] Figure 5 Schematic diagram of gradient reduction optimization in one embodiment;

[0029] Figure 6 Schematic diagram of gradient reduction optimization in another embodiment;

[0030] Figure 7 Schematic diagram of model training through multiple GPUs in one embodiment;

[0031] Figure 8 Schematic diagram of weight aggregation between GPUs in one embodiment;

[0032] Figure 9 Schematic diagram of communicating model parameters in one embodiment;

[0033] Figure 10 Schematic diagram of communicating model parameters after zero-padding in one embodiment;

[0034] Figure 11 Schematic flowchart of a model training method in another embodiment;

[0035] Figure 12 Schematic flowchart of communicating model parameters in another embodiment;

[0036] Figure 13 Schematic diagram of optimization steps for model training in a communication scenario in one embodiment;

[0037] Figure 14 Structural block diagram of a model training device in one embodiment;

[0038] Figure 15 Internal structure diagram of a computer device in one embodiment;

[0039] Figure 16 Internal structure diagram of a computer device in another embodiment. Detailed implementation manners

[0040] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0041] Embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.

[0042] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, and mechatronics. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0043] The solution provided by the embodiments of this application relates to technologies such as computer vision technology and machine learning in artificial intelligence, and is specifically described through the following embodiments:

[0044] The model training method provided by the embodiments of this application can be applied to an application environment as Figure 1 shown. The application environment involves multiple computer devices, and the computer devices communicate with each other through a network. It can be understood that the computer device can be a terminal or a server. The terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster or a cloud server composed of multiple servers.

[0045] The computer device can also be called a node. A computer device includes multiple processors, and the multiple processors can include a first processor and a target processor. Further, the multiple processors can also include a second processor.

[0046] Specifically, the first processor obtains the target training subset corresponding to the first processor, and based on the model parameter set corresponding to the initial model to be trained and the target training subset, obtains the original gradient set corresponding to the initial model on the target training subset, performs gradient screening on the original gradient set, and obtains the compressed gradient set corresponding to the initial model on the target training subset. The first processor sends the compressed gradient set corresponding to the initial model on the target training subset to other target processors used to train the initial model, and obtains the compressed gradient set corresponding to the initial model on other training subsets from the target processors. The first processor fuses each compressed gradient set to obtain a comprehensive gradient set, and the comprehensive gradient set is used to update the model parameter set corresponding to the initial model. After the model parameter set corresponding to the initial model is updated, the first processor returns to execute the step of obtaining the target training subset corresponding to the first processor until the training end condition is met, and the target model is obtained.

[0047] In one embodiment, asFigure 2 As shown in Figure 2 , a model training method is provided. Taking the application of this method to a first processor as an example for illustration. Among them:

[0048] Step S202: Obtain a target training subset corresponding to the first processor.

[0049] Among them, the first processor is any one of the processor clusters used to train the initial model. There are multiple processors for training the initial model, and each processor uses a different training subset to perform parallel model training on the initial model. The target training subset is the training subset used by the first processor.

[0050] It can be understood that the training set of the initial model includes multiple training samples. The training set is sampled to obtain multiple training subsets, and each training subset is assigned to each processor.

[0051] Step S204: Based on the model parameter set corresponding to the initial model to be trained and the target training subset, obtain the original gradient set corresponding to the initial model on the target training subset.

[0052] Among them, the initial model refers to the machine learning model to be trained. Through model training on the initial model, a target model can be finally obtained. The target model refers to the machine learning model that has completed training. It can be understood that the method of this application can be applied to the training of any machine learning model to reduce the consumption of communication resources and improve the model training efficiency. Machine learning models include, but are not limited to, neural networks, decision trees, random forest models, etc. The machine learning model can be a model for processing at least one of video, audio, image, text, and other data.

[0053] The model parameter set of the model is an important part of the model. The sample to be processed is input into the model, and after the data processing of the model, the model outputs the prediction label corresponding to the sample to be processed. The model processes the input sample to be processed through the model parameter set to obtain the model output data. During model training, the model parameter set is adjusted to improve the accuracy and precision of the model. Taking the neural network model as an example, the model parameter set of the neural network model includes the weights between neurons and the biases corresponding to the neurons.

[0054] Model parameters are variables that need to be learned during the training process and determine the performance and expressive power of the model. During the training process, the model updates the model parameters through the backpropagation algorithm to minimize the loss function. The gradient is the partial derivative of the loss function with respect to the model parameters, representing the rate of change of the loss function in the parameter space. During the training process, the model parameters are updated by calculating the gradient of the loss function with respect to the model parameters, enabling the model to better fit the training data. The direction of the gradient represents the direction of change of the model parameters, and the magnitude of the gradient represents the speed of change of the model parameters. Based on the gradient, the model parameters are adjusted to enable the model to quickly converge to the optimal solution and minimize the loss function.

[0055] Based on the training subset corresponding to the initial model, model training is performed on the initial model to obtain the original gradient set corresponding to the initial model on the training subset.

[0056] Specifically, the first processor can obtain the target training subset and perform model training on the initial model based on the target training subset to obtain the original gradient set corresponding to the initial model on the target training subset. The target training subset includes multiple training samples, and there are corresponding training labels for the training samples. During model training, the training samples in the training subset are input into the model, and the model outputs the predicted labels corresponding to the training samples. The training labels and predicted labels corresponding to each training sample in the training subset are substituted into the loss function to obtain the model loss. The model loss is backpropagated in the model to calculate the gradient of the loss function with respect to the model parameters, and the gradients corresponding to multiple model parameters in the model parameter set are obtained, forming the original gradient set corresponding to the model on the target training subset.

[0057] Step S206: Perform gradient screening on the original gradient set to obtain the compressed gradient set corresponding to the initial model on the target training subset.

[0058] Among them, gradient screening refers to screening out the gradients that meet the preset conditions from the original gradient set. The preset conditions can be set as needed. For example, gradients with larger absolute values can contribute more to model convergence. The preset condition can be to screen out the gradients whose absolute values are greater than the preset threshold, or the preset condition can be to sort the gradients in descending order of absolute value and screen out the top k gradients.

[0059] It can be understood that gradient screening is performed on the original gradient set to obtain the compressed gradient set. The data volume of the compressed gradient set is smaller than that of the original gradient set.

[0060] Specifically, each processor uses a different training subset to perform parallel model training on the initial model. The training subsets used by each processor are different, so the original gradient sets obtained by each processor are different. The processors need to exchange gradient information with each other in order to integrate the gradient information to adjust the model parameters, so that the model can learn relevant knowledge from multiple training subsets synchronously. The processors need to exchange gradient information with each other, but the original gradient set contains a large amount of data and consumes a lot of communication resources. Therefore, in order to save communication resources and improve communication efficiency, before sending the gradient information to other processors, the first processor performs gradient screening on the original gradient set to obtain the compressed gradient set corresponding to the initial model on the target training subset, and sends the compressed gradient set with a smaller amount of data to other processors.

[0061] Step S208: Send the compressed gradient set corresponding to the initial model on the target training subset to other target processors used to train the initial model, and obtain the compressed gradient set corresponding to the initial model on other training subsets from the target processors.

[0062] The processors used to train the initial model include the first processor and the target processors. The first processor and the target processors can be of the same type. For example, both the first processor and the target processors are GPUs (Graphics Processing Units). It can be understood that the first processor and the target processors can also be of different types. For example, the first processor is a GPU and the target processor is a CPU (Central Processing Unit). There can be at least one target processor.

[0063] Specifically, the processors need to exchange gradient information with each other in order to integrate the gradient information to adjust the model parameters. Therefore, the first processor can send the compressed gradient set corresponding to the initial model on the target training subset to other target processors used to train the initial model. Correspondingly, the first processor can obtain the compressed gradient set corresponding to the initial model on other training subsets sent by the target processors.

[0064] It can be understood that during the model training process, the data processing process of the target processors is similar to that of the first processor. Referring to the method by which the first processor obtains the compressed gradient set corresponding to the initial model on the target training subset, the target processors can obtain the compressed gradient set corresponding to the initial model on other training subsets.

[0065] Step S210: Integrate the compressed gradient sets to obtain a comprehensive gradient set; the comprehensive gradient set is used to update the model parameter set corresponding to the initial model.

[0066] Specifically, the first processor may fuse the compressed gradient set obtained by itself and the compressed gradient set fed back by the target processor to obtain a comprehensive gradient set. Specifically, the gradients corresponding to the same model parameter in each compressed gradient set may be fused to obtain a comprehensive gradient set. Further, it may be that the first processor updates the model parameter set corresponding to the initial model based on the comprehensive gradient set, or it may be that the first processor sends the comprehensive gradient set to the second processor, and the second processor updates the model parameter set corresponding to the initial model based on the comprehensive gradient set.

[0067] In one embodiment, the first processor updates the model parameter set corresponding to the initial model based on the comprehensive gradient set. The first processor obtains the local optimizer parameters matching the first processor, and updates the local model parameters in the model parameter set that match the first processor based on the local optimizer parameters matching the first processor and the comprehensive gradient set. Specifically, the full optimizer parameters for training the model are distributed and stored on the first processor and the target processor, so that the first processor and the target processor respectively store different local optimizer parameters. The local optimizer parameters held by the first processor are the local optimizer parameters matching the first processor. The update of the model parameters requires gradients and optimizer parameters. Since the first processor stores the local optimizer parameters, the first processor can only update the local model parameters in the model parameter set based on the local optimizer parameters and the comprehensive gradient set. The local model parameters are determined by the local optimizer parameters. After the first processor updates the corresponding local model parameters, it can communicate with the target processor about the updated local model parameters, so that both the first processor and the target processor know the updated model parameter set.

[0068] In one embodiment, the first processor updates the model parameter set corresponding to the initial model based on the comprehensive gradient set. The first processor obtains the target gradient set matching the first processor from the comprehensive gradient set and discards the remaining gradients. The first processor obtains the local optimizer parameters matching the first processor, and updates the local model parameters in the model parameter set that match the first processor based on the local optimizer parameters matching the first processor and the target gradient set. Specifically, the comprehensive gradient set for training the model is distributed and stored on the first processor and the target processor. For example, the comprehensive gradient set for training the model is evenly divided, and the first processor and the target processor respectively store a part of the gradients. The part of the gradients responsible for the first processor is the target gradient set matching the first processor.

[0069] In one embodiment, the first processor sends the comprehensive gradient set to the second processor, and the second processor updates the model parameter set corresponding to the initial model based on the comprehensive gradient set. The computing efficiency of the first processor is greater than that of the second processor, and the storage capacity of the second processor is greater than that of the first processor.

[0070] Step S212: After the model parameter set corresponding to the initial model is updated, return to execute the step of obtaining the target training subset corresponding to the first processor until the training end condition is met, and a target model is obtained.

[0071] Among them, the training end condition is the condition for determining the end of model training. The training end condition can be set as needed. For example, the training end condition can be at least one of the model loss being less than a preset loss, the model iteration times being greater than a preset iteration times, etc. The target model refers to the model that has completed training.

[0072] Specifically, after the model parameter set corresponding to the initial model is updated, the first processor can obtain a new target training subset, and perform iterative training on the updated initial model based on the new target training subset until the training end condition is met, obtain the final model parameter set, and obtain the target model based on the final model parameter set.

[0073] In the above model training method, each processor used to train the initial model performs parallel model training on the initial model based on the model parameter set corresponding to the initial model and its respective training subset, which can improve the model training efficiency. When performing model training, calculate the original gradient set corresponding to the initial model on the training subset, compress the original gradient set to obtain the compressed gradient set corresponding to the initial model on the training subset, and each processor communicates its respective compressed gradient set. The data volume of the compressed gradient set is smaller than that of the original gradient set, which can improve the communication efficiency, thereby further improving the model training efficiency. The processor fuses its own compressed gradient set with the compressed gradient sets of other processors to obtain a comprehensive gradient set. The comprehensive gradient set integrates the gradient information of the initial model on each training subset. Adjust the model parameters based on the comprehensive gradient set, so that the model can learn knowledge from each training subset, which can improve the prediction ability of the model and the quality of model training.

[0074] In one embodiment, obtaining the original gradient set corresponding to the initial model on the target training subset based on the model parameter set corresponding to the initial model to be trained and the target training subset includes:

[0075] Perform forward calculation on the model parameter set corresponding to the initial model to be trained based on the training samples in the target training subset to obtain the predicted labels corresponding to the training samples; perform loss calculation based on the training labels and predicted labels corresponding to the training samples to obtain the model loss corresponding to the initial model on the target training subset; perform backward calculation on the model parameter set based on the model loss to obtain the original gradient set corresponding to the initial model on the target training subset.

[0076] Among them, the training subset includes training samples and the training labels corresponding to the training samples. The training labels corresponding to the training samples are the results expected to be output by the model. The training samples are input into the model, and the model outputs the predicted labels corresponding to the training samples. The predicted labels corresponding to the training samples are the results actually output by the model. The training objective of the model is to make the predicted labels corresponding to the training samples close to the training labels corresponding to the training samples, so that the model has the ability to accurately process data and can output the correct labels. It can be understood that the training samples and training labels are determined based on the model function. For example, if the model is an image segmentation model, the training samples are images with known image segmentation results, and the training labels are the correct image segmentation results. If the model is a text classification model, the training samples are texts with known text classification results, and the training labels are the correct text classification results. If the model is a video anomaly detection model, the training samples are videos with known whether there is an anomaly, and the training labels are normal or abnormal.

[0077] The model usually includes multiple network layers. Forward calculation refers to the process of forward propagation of the input data of the model, which is a data processing process that sequentially passes through each network layer from the input layer to the output layer. Forward calculation calculates and stores intermediate variables in order (from the input layer to the output layer). Backward calculation refers to the process of backward propagation of the model loss, which is a data processing process that sequentially passes through each network layer from the output layer to the input layer. Backward calculation sequentially calculates the gradients corresponding to the model parameters in order (from the output layer to the input layer).

[0078] The model loss is used to represent the error between the training labels and the predicted labels corresponding to the training samples, that is, the error between the actual output result and the expected output result of the model. The training labels and the predicted labels corresponding to the training samples can be substituted into the loss function to obtain the model loss.

[0079] Specifically, the first processor inputs the training samples in the target training subset into the model for forward calculation to obtain the predicted labels corresponding to the training samples. That is, the first processor performs forward calculation on the model parameter set corresponding to the initial model to be trained based on the training samples in the target training subset to obtain the predicted labels corresponding to the training samples. The first processor calculates the loss based on the training labels and the predicted labels corresponding to the training samples in the target training subset to obtain the model loss corresponding to the initial model on the target training subset. The first processor performs backward propagation of the model loss in the model to obtain the original gradient set corresponding to the initial model on the target training subset. That is, based on the model loss, backward calculation is performed on the model parameter set to obtain the original gradient set corresponding to the initial model on the target training subset.

[0080] In the above embodiments, based on the training samples in the target training subset, forward calculation is performed on the model parameter set corresponding to the initial model to be trained to obtain the predicted labels corresponding to the training samples. Loss calculation is performed based on the training labels and predicted labels corresponding to the training samples to obtain the model loss corresponding to the initial model on the target training subset. The model loss can reflect the difference between the training labels and the predicted labels. Based on the model loss, backward calculation is performed on the model parameter set to obtain the original gradient set corresponding to the initial model on the target training subset. The gradients in the original gradient set can indicate the adjustment direction and magnitude of the model parameters. The model parameters are adjusted based on the gradients, enabling the model to converge quickly.

[0081] Reference Figure 3 , a complete process of machine learning training mainly includes: data reading, forward calculation, backward calculation, gradient reduction, and weight update. Data reading refers to reading the training data (i.e., the training set, including training samples and the training labels corresponding to the training samples). Forward calculation refers to inputting the training samples into the model to obtain the predicted labels corresponding to the training samples. Backward calculation refers to backpropagating the model loss obtained based on the training labels and predicted labels corresponding to the training samples in the model to obtain the gradients corresponding to the model parameters. Gradient reduction refers to synchronizing the gradients among various processors. Weight update refers to updating the model parameters based on the gradients obtained through gradient reduction.

[0082] In one embodiment, based on the model loss, backward calculation is performed on the model parameter set to obtain the original gradient set corresponding to the initial model on the target training subset, including:

[0083] Backpropagate the model loss in the model parameter set, and sequentially calculate the gradients corresponding to multiple model parameters in the model parameter set; according to the gradient calculation order, form local gradient sets with a reference number of gradients, and sequentially obtain each local gradient set corresponding to the initial model on the target training subset; for each obtained local gradient set, use the local gradient set as the original gradient set corresponding to the initial model on the target training subset, and enter the step of performing gradient screening on the original gradient set to obtain the compressed gradient set corresponding to the initial model on the target training subset.

[0084] Among them, backpropagating the model loss in the model parameter set means that the model loss passes from the output layer of the model through each network layer to the input layer in sequence, and calculates the gradients corresponding to the model parameters along the way.

[0085] Models usually have a relatively large number of model parameters. The first processor can send the gradients obtained by calculation to other processors in a timely manner every time a reference number of gradients is calculated, so as to realize sending gradients while calculating gradients, improve the utilization rate of communication resources, and improve communication efficiency. The reference number can be set as needed. For example, the reference number can be a preset fixed value; the reference number can also be a dynamic value that changes with the communication environment.

[0086] Specifically, the first processor can perform backpropagation of the model loss in the model parameter set, and sequentially calculate the gradients corresponding to multiple model parameters in the model parameter set. The first processor calculates the gradients corresponding to the model parameters in sequence (from the output layer to the input layer). To improve the utilization rate of communication resources, the first processor can, according to the gradient calculation order, every time a reference number of gradients is calculated, form a local gradient set with the reference number of gradients, and use the local gradient set as the original gradient set corresponding to the initial model on the target training subset, and send the local gradient set to other target processors for training the initial model. It can be understood that the first processor can obtain each local gradient set corresponding to the initial model on the target training subset in an orderly manner. Every time the first processor obtains a local gradient set, it uses the local gradient set as the original gradient set corresponding to the initial model on the target training subset, and enters the step of performing gradient screening on the original gradient set to obtain the compressed gradient set corresponding to the initial model on the target training subset, so that the first processor can orderly send each compressed gradient set corresponding to the initial model on the target training subset to the target processor.

[0087] In one embodiment, the first processor can obtain the communication environment characteristics corresponding to the initial model, and determine the reference number according to the communication environment characteristics. The communication environment characteristics include at least one of network bandwidth information, communication link information, and communication topology information. For example, the reference number is positively correlated with the network bandwidth in the network bandwidth information, and the larger the network bandwidth, the larger the reference number. For another example, under the communication environment characteristics, perform communication efficiency tests on multiple candidate numbers, and use the candidate number with the highest communication efficiency as the reference number.

[0088] For example, reference Figure 4 , the first processor can detect the network bandwidth, determine the accumulation amount corresponding to the gradient (i.e., the reference number) based on the detected network bandwidth, accumulate the calculated gradients during the back-calculation process, and when the gradients accumulate to the reference number, transmit multiple gradients at one time to the target processor to improve the bandwidth utilization rate. Further, when the gradients accumulate to the reference number, the first processor can screen the accumulated gradients and transmit the screened gradients to the target processor at one time to reduce the number of communications.

[0089] In the above embodiments, the model loss is backpropagated in the model parameter set, and the gradients corresponding to multiple model parameters in the model parameter set are calculated in sequence. According to the gradient calculation order, a reference number of gradients are grouped to form a local gradient set. Each time a local gradient set is obtained, the local gradient set is used as the original gradient set corresponding to the initial model on the target training subset. The original gradient set is screened for gradients to obtain the compressed gradient set corresponding to the initial model on the target training subset, and the compressed gradient set is sent to the target processor. In this way, accumulating multiple gradients at one time for gradient communication can effectively improve the bandwidth utilization rate and communication efficiency.

[0090] In one embodiment, the model training method further includes:

[0091] When the first processor and the target processor for training the initial model each save the model parameter set, during forward calculation or backward calculation, perform forward calculation or backward calculation on the model parameter set saved by itself;

[0092] When the first processor and the target processor for training the initial model each save the local model parameters in the model parameter set, during forward calculation or backward calculation, obtain the local model parameters saved by the target processor from the target processor, combine the local model parameters saved by itself and the local model parameters obtained from the target processor to form a model parameter set, perform forward calculation or backward calculation on the model parameter set, and after the forward calculation or backward calculation, discard the other model parameters except the local model parameters saved by itself.

[0093] Specifically, the first processor and the target processor for training the initial model can each save all the model parameters. During forward calculation or backward calculation, the processor needs to use the complete model parameter set. Therefore, taking the first processor as an example, when the first processor and the target processor each save the model parameter set, during forward calculation or backward calculation, the first processor performs forward calculation or backward calculation on the model parameter set saved by itself.

[0094] To reduce the storage pressure on the processor used to train the initial model, the model parameter set of the initial model can be split, and a first processor and a target processor used to train the initial model each maintain a part of the model parameters. That is, the first processor and the target processor each save the local model parameters in the model parameter set. However, during forward calculation or backward calculation, the processor needs to obtain the complete model parameter set. Therefore, taking the first processor as an example, when the first processor and the target processor each save the local model parameters in the model parameter set, during forward calculation or backward calculation, the first processor and the target processor communicate, the first processor obtains the local model parameters saved by the target processor, combines the local model parameters saved by itself and the local model parameters obtained from the target processor into a model parameter set, and performs forward calculation or backward calculation on the model parameter set to ensure the accuracy and effectiveness of forward calculation or backward calculation. Further, to reduce the storage pressure on the first processor, after forward calculation or backward calculation, the first processor promptly discards other model parameters except the local model parameters saved by itself.

[0095] In the above embodiment, when each processor used to train the initial model saves the local model parameters in the model parameter set, during forward calculation or backward calculation, the processors communicate with each other to obtain the complete model parameter set, and after forward calculation or backward calculation, each processor promptly discards other model parameters except the local model parameters saved by itself, which can effectively save the storage pressure of the processor.

[0096] In one embodiment, gradient screening is performed on the original gradient set to obtain a compressed gradient set corresponding to the initial model on the target training subset, including:

[0097] Obtain a target number of gradients from the original gradient set in descending order of the absolute value of the gradient, and form a compressed gradient set corresponding to the initial model on the target training subset.

[0098] Among them, the absolute value of the gradient refers to the absolute value of the gradient. The target number can be set as needed. For example, the target number can be a preset fixed value; the target number can also be a dynamic value that changes with the model scale.

[0099] In one embodiment, the target number is positively correlated with the number of model parameters corresponding to the model. The more the number of model parameters of the model, the greater the training difficulty of the model, and involving more gradients in weight update can ensure the training effect of the model.

[0100] Specifically, when performing gradient screening on the original gradient set, the first processor can obtain a target number of gradients from the original gradient set according to the absolute value of the gradient from large to small, and form a compressed gradient set corresponding to the initial model on the target training subset. That is, the compressed gradient set includes the top k gradients with the largest absolute values ​​in the original gradient set, and the target number is represented by k.

[0101] In the above embodiment, gradients with larger absolute values ​​can make more contributions to model convergence. A target number of gradients are obtained from the original gradient set according to the absolute values ​​of the gradients from large to small to form a compressed gradient set. This can not only reduce the amount of communication data, save communication resources, and improve communication efficiency, but also ensure the model training effect.

[0102] In one embodiment, the compressed gradient sets are fused to obtain a comprehensive gradient set, including:

[0103] The gradients corresponding to the same model parameter in each compressed gradient set are accumulated to obtain a cumulative gradient set; and based on the total number of the first processor and the target processor, the accumulated gradient set is averaged to obtain a comprehensive gradient set.

[0104] Specifically, when fusing the compressed gradient sets, the first processor may accumulate the gradients corresponding to the same model parameter in the compressed gradient sets, and form the accumulated gradients into a cumulative gradient set, that is, the gradients corresponding to the same model parameter in the compressed gradient sets are accumulated, and the accumulated gradients form a cumulative gradient set. Further, the first processor averages the gradients in the cumulative gradient set based on the total number of the first processor and the target processor, that is, divides the gradients in the cumulative gradient set by the total number to obtain a comprehensive gradient set.

[0105] In one embodiment, bandwidth utilization is improved by accumulating multiple gradients and sending them at once, and the accumulated gradients are further compressed to reduce communication traffic. Figure 5 , Figure 5The optimization process of gradient planning is shown. After the processor accumulates n gradients, it performs gradient sparsification on the n gradients. Gradient sparsification means selecting the top k gradients with the largest absolute values from the n gradients. After selecting the top k gradients with the largest absolute values, the indices and values of the top k gradients with the largest absolute values are recorded. The indices refer to the non-zero gradient values, and the values refer to the indices of the non-zero gradients. Recording the indices and values can further reduce the communication volume. Each processor performs allgather communication of the indices, values it records and other processors, so that each processor can obtain the global indices and global values. For the global indices and global values, the processor sums and averages the gradients with the same index to obtain the reduced sparse gradient (i.e., the sparse gradient after communication).

[0106] For example, referring to Figure 6 , there are four GPUs used to train the initial model, namely GPU0, GPU1, GPU2, and GPU3. Each GPU performs allgather communication of the indices, values it records and other GPUs. Each GPU sums the gradients with the same index, and then each GPU averages the summed gradients to obtain the reduced sparse gradient.

[0107] In the above embodiment, the gradients corresponding to the same model parameter in each compressed gradient set are accumulated to obtain an accumulated gradient set. Based on the total number of the first processor and the target processor, the accumulated gradient set is averaged to obtain a comprehensive gradient set. The comprehensive gradient set includes the average compressed gradients of the initial model on each training subset. Adjusting the model parameter set of the initial model based on the comprehensive gradient set can enable the model to comprehensively learn the knowledge contained in each training subset and improve the model training efficiency.

[0108] In one embodiment, referring to Figure 7 , when a model is trained in parallel by multiple GPUs, the data that the GPUs need to store includes model parameters, gradients, and optimizer parameters. In the related art, usually each GPU stores the complete model parameters, gradients, and optimizer parameters. However, this greatly occupies the video memory of the GPU, and a GPU with a larger video memory is required to implement model training.

[0109] In order to optimize the video memory occupation of the GPU, the data that the GPU needs to store can be split, which avoids each GPU maintaining a complete copy of the data in parallel, greatly reducing the video memory occupation of the GPU during model training and enabling a larger model to be trained on a single machine.

[0110] In state 1, the optimizer parameters of the model can be sliced, and each GPU maintains a part of the optimizer parameters respectively. In state 2, the optimizer parameters and gradients of the model can be sliced, and each GPU maintains a part of the optimizer parameters and a part of the gradients respectively. In state 3, the optimizer parameters, gradients, and model parameters of the model can be sliced, and each GPU maintains a part of the optimizer parameters, a part of the gradients, and a part of the model parameters respectively. In states 2 and 3, each GPU is responsible for parallelly reducing part of the gradients (i.e., gradient reduction) through reduce, and then updating the weights corresponding to this part of the gradients. After the backpropagation is completed, all GPUs parallelly obtain the full weight information through allgather, so that the communication volume is not increased in order to save video memory. Reduce means that the gradients on different GPUs are calculated to obtain new gradients after communication. Allgather means that the data on different GPUs is transmitted to each GPU after communication.

[0111] Each GPU communicates only with its two adjacent GPUs, sends the data it knows to the next GPU, and receives the data sent by the previous GPU, forming a topological ring among the GPUs.

[0112] In one embodiment, the model training method further includes:

[0113] Obtain a target gradient set matching the first processor from the comprehensive gradient set, and discard the remaining gradients; send the target gradient set matching the first processor to the second processor, so that the second processor updates the local model parameters matching the first processor in the model parameter set based on the target gradient set matching the first processor and the local optimizer parameters; the computing efficiency of the first processor is greater than that of the second processor, and the first processor and the target processor are processors of the same type; after the local model parameters matching the first processor in the model parameter set are updated, obtain the updated local model parameters fed back by the second processor, and send the updated local model parameters to the target processor.

[0114] Among them, in addition to the first processor and the target processor, the model training also requires a second processor. The first processor and the target processor are processors of the same type. For example, both the first processor and the target processor are GPUs. The first processor and the target processor are responsible for data acquisition, forward calculation, backward calculation, and gradient reduction in the model training process. The second processor is responsible for weight update in the model training process. The first processor and the second processor are processors of different types, and the computing efficiency of the first processor is greater than that of the second processor. For example, the first processor is a GPU and the second processor is a CPU, and the computing efficiency of the GPU is greater than that of the CPU, and the storage capacity of the GPU is less than that of the CPU.

[0115] To reduce the storage pressure on the processor used to train the initial model, the gradients of the model can be split, and a first processor and a target processor used to train the initial model each maintain a part of the gradients, that is, the first processor and the target processor each save a part of the gradients in the comprehensive gradient set. The target gradient set matching the first processor includes the gradients that the first processor is responsible for maintaining and saving.

[0116] To reduce the storage pressure on the processor used to train the initial model, the optimizer parameters of the model can be split, and a first processor and a target processor used to train the initial model each maintain a part of the optimizer parameters, that is, the first processor and the target processor each save the local optimizer parameters in the optimizer parameter set. The local optimizer parameters matching the first processor are the optimizer parameters that the first processor is responsible for maintaining and saving.

[0117] It can be understood that the optimizer parameters are the parameters of the optimizer used to train the model. After having the model parameters and the loss function, the learning model parameters are optimized through the optimization function. This optimization function is called an optimizer, and its internal principle is mainly to optimize the model parameters in the model through the method of gradient descent. RMSprop (Root Mean Square Propagation), Adagrad (Adaptive Gradient Algorithm), Adam (Adaptive Moment Estimation), and SGD (Stochastic Gradient Descent) are relatively common optimizers.

[0118] Specifically, to reduce the occupancy of storage space, the first processor can obtain the target gradient set matching the first processor from the comprehensive gradient set, discard the remaining gradients, and send the target gradient set matching the first processor to the second processor. The second processor updates the model parameter set corresponding to the initial model based on the received gradient information.

[0119] The second processor stores the local optimizer parameters matching the first processor. After receiving the target gradient set matching the first processor, the second processor updates the local model parameters in the model parameter set that match the first processor based on the target gradient set and the local optimizer parameters matching the first processor. After the local model parameters in the model parameter set that match the first processor are updated, the second processor returns the updated local model parameters to the first processor. After receiving the updated local model parameters, the first processor can send the updated local model parameters to the target processor so that the first processor and the target processor can synchronize the updated model parameter set in a timely manner.

[0120] It can be understood that there can be multiple second processors. Each second processor can be responsible for updating different local model parameters. The second processor can also use multiple processes for weight update, and each process is responsible for updating different local model parameters.

[0121] A complete process of machine learning training mainly includes: data reading, forward calculation, backward calculation, gradient regularization, and weight update. The first processor has stronger computing power and is responsible for data reading, forward calculation, backward calculation, and gradient regularization. The calculations of forward calculation, backward calculation, and gradient regularization are relatively complex. The second processor has stronger storage capacity and is responsible for weight update. The amount of data involved in weight update is large and the calculation is relatively simple. Combining the first processor and the second processor can reduce the storage pressure of the first processor while ensuring the computing efficiency.

[0122] In the above embodiments, the first processor has stronger computing power and is responsible for data reading, forward calculation, backward calculation, and gradient regularization. The second processor has stronger storage capacity and is responsible for weight update. Combining the first processor and the second processor for model training can reduce the storage occupancy of the first processor while ensuring the computing efficiency.

[0123] In one embodiment, the model training method further includes:

[0124] Obtaining the communication environment characteristics corresponding to the initial model; the communication environment characteristics include at least one of network bandwidth information, communication link information, and communication topology information; determining the target transmission amount for the model parameters based on the communication environment characteristics; and based on the target transmission amount, grouping and sending the model parameters known to the first processor in the model parameter set corresponding to the initial model to the target processor, so as to jointly perform parallel model training on the initial model by the target processor based on their respective corresponding training subsets.

[0125] Among them, the communication environment characteristics are the characteristics reflecting the communication environment of the processors used for training the model. The communication environment characteristics include at least one of network bandwidth information, communication link information, and communication topology information. The network bandwidth information is the information reflecting the network bandwidth. The communication link information refers to the communication method between the nodes where the processors are located. For example, the nodes communicate using the TCP protocol (Transmission Control Protocol), or the nodes communicate using the RDMA technology (Remote Direct Memory Access). The communication topology information refers to the topological relationship between multiple processors inside the node and the topological relationship between the nodes.

[0126] The target transmission amount for the model parameters refers to the amount of data sent at one time when sending the model parameters.

[0127] Specifically, the first processor and the target processor for training the model perform parallel model training on the initial model based on their respective corresponding training subsets. When the first processor and the target processor for training the model each save the local model parameters in the model parameter set, during forward calculation or backward calculation, the first processor needs to send the model parameters it knows to the target processor, and the target processor also needs to send the model parameters it knows to the first processor. The model parameter set of a machine learning model is generally large, so the communication volume is large. To improve communication performance, the model parameters can be sent in batches. The first processor obtains the communication environment characteristics corresponding to the initial model, determines the target transmission volume for the model parameters based on the communication environment characteristics, and based on the target transmission volume, groups and sends the model parameters known to the first processor in the model parameter set corresponding to the initial model to the target processor, so that the target processor can perform forward calculation or backward calculation based on the complete model parameter set.

[0128] In the above embodiments, the processor determines the target transmission volume for the model parameters based on the communication environment characteristics, and groups and sends the known model parameters to other processors based on the target transmission volume, which can make full use of the communication capacity and effectively improve the communication efficiency.

[0129] In one embodiment, determining the target transmission volume for the model parameters based on the communication environment characteristics includes:

[0130] Obtaining a plurality of candidate transmission volumes preset for the communication environment characteristics; based on the candidate transmission volumes, performing communication efficiency tests in the processor cluster for training the initial model to obtain the communication efficiencies corresponding to the respective candidate transmission volumes; and using the candidate transmission volume with the highest communication efficiency as the target transmission volume for the model parameters.

[0131] Among them, the target transmission volume for the model parameters is determined from a plurality of candidate transmission volumes. Different candidate transmission volumes can be preset for different communication environment characteristics. The processor cluster for training the initial model includes the first processor and the target processor.

[0132] Specifically, when determining the target transmission volume for the model parameters based on the communication environment characteristics, the first processor can obtain a plurality of candidate transmission volumes preset for the communication environment characteristics, perform communication efficiency tests in the processor cluster for training the initial model based on the candidate transmission volumes, and obtain the communication efficiencies corresponding to the respective candidate transmission volumes. For example, send data packets in the processor cluster for training the initial model according to the candidate transmission volume, and determine the communication efficiency based on the round-trip delay of the data packets. The first processor uses the candidate transmission volume with the highest communication efficiency as the target transmission volume for the model parameters.

[0133] In the above embodiments, multiple candidate transmission amounts are set in advance according to the characteristics of the communication environment, and communication efficiency tests for the candidate transmission amounts are carried out in the processor cluster used for training the initial model, so as to obtain the communication efficiency corresponding to each candidate transmission amount. The candidate transmission amount with the highest communication efficiency is used as the target transmission amount for the model parameters. Data transmission based on such a target transmission amount can maximize the communication efficiency.

[0134] In one embodiment, based on the target transmission amount, the model parameters known to the first processor in the model parameter set corresponding to the initial model are grouped and sent to the target processor, including:

[0135] Based on the target transmission amount, the model parameters known to the first processor in the model parameter set corresponding to the initial model are grouped to obtain multiple initial combinations of model parameters; the sizes of the multiple initial combinations of model parameters are filled to be integer multiples of the physical storage unit corresponding to the target processor to obtain multiple target combinations of model parameters; and the multiple target combinations of model parameters are sequentially sent to the target processor.

[0136] Among them, the physical storage unit corresponding to the first processor refers to the bit width or word length of the first processor. The bit width of a processor refers to the width of the data bus and is used to represent and process the number of bits of binary data. It determines the number of binary bits that the processor can process simultaneously. The word length of a processor refers to the number of binary bits that the processor can process at one time. Generally, the word length of a processor is determined by the bit width of the processor. The word length of a processor focuses on the size of the data that can be processed at one time, and the bit width of a processor focuses on the width of the register and the data bus. Usually, the word length and bit width of a processor are equal.

[0137] Specifically, the first processor can group the model parameters known to the first processor in the model parameter set corresponding to the initial model based on the target transmission amount to obtain multiple initial combinations of model parameters, and the data volume size of the initial combination of model parameters is the target transmission amount. The first processor fills the size of the initial combination of model parameters to be an integer multiple of the physical storage unit corresponding to the target processor so that the target processor can cache data. The first processor fills the sizes of the multiple initial combinations of model parameters to be integer multiples of the physical storage unit corresponding to the target processor to obtain multiple target combinations of model parameters, and the data volume size of the target combination of model parameters is the target transmission amount. The first processor sequentially sends the multiple target combinations of model parameters to the target processor.

[0138] Taking the first processor and the target processor as GPUs as an example, refer to Figure 8, between GPUs, the local model parameters saved by each are communicated and shared through allgather. In the allgather operation, each of the K processors aggregates the model parameters from other processors to obtain model parameters of K*N. Each processor initially saves N model parameters. Since the optimizer updates the weights and aggregates the weights after each training backward pass, the performance of allgather directly affects the speed of one training in data parallel training.

[0139] As the complexity of machine learning models and the scale of datasets increase, computational efficiency has become an issue that cannot be ignored. The continuous increase in the complexity of machine learning models has led to a large number of communication parameters, even reaching the billion level. If the parameters of the billion level are completed through a single allgather interface call, the performance is very low and the utilization of bandwidth is very low. To improve the bandwidth utilization rate, refer to Figure 9 , detect the network bandwidth of the hardware environment where the GPU is located, automatically determine the optimal packet size between processors based on the detected network bandwidth, split the model parameters to be communicated based on the optimal packet size, and improve the communication efficiency through multiple rounds of sending.

[0140] Furthermore, allgather also has the problem of memory misalignment. When the amount of data sent by each participating processor is not an integer multiple of the physical storage unit of the memory, it will lead to poor performance of allgather.

[0141] To improve the communication quality, encapsulate the allgather interface. In the encapsulated interface, judge the data sending size. Taking 16 bits as the physical storage unit of the memory as an example, if it is found that the data sending size is not an integer multiple of 16, then pad the sending data with zeros (padding) to make the sending data an integer multiple of 16. Refer to Figure 10 , after splitting the model parameters to be communicated based on the optimal packet size, pad the amount of model parameters to be sent in each round with zeros to make it an integer multiple of 16, so that the data sending amount is aligned with the memory of the processor.

[0142] In one embodiment, to reduce the expansion operation, multiple candidate sending amounts preset for the communication environment characteristics can be determined based on the physical storage unit corresponding to the target processor, and each candidate sending amount is a different integer multiple of the physical storage unit.

[0143] In the above embodiment, when grouping and sending model parameters based on the target sending amount, padding the data sending amount in each round to an integer multiple of the physical storage unit corresponding to the target processor can facilitate the target processor to successfully cache the received model parameters and ensure the communication success rate.

[0144] In one embodiment, the first processor and the target processor are processors of the same type. The model training method further includes:

[0145] When the cross-machine bandwidth of the first processor is less than the bandwidth threshold, enter the step of performing gradient screening on the original gradient set to obtain the compressed gradient set corresponding to the initial model on the target training subset;

[0146] When the cross-machine bandwidth of the first processor is greater than or equal to the bandwidth threshold, obtain the original gradient sets corresponding to the initial model on other training subsets from the target processor, fuse the respective original gradient sets to obtain the full gradient set for updating the model parameter set corresponding to the initial model. After updating the model parameter set corresponding to the initial model, return to the step of obtaining the target training subset corresponding to the first processor and execute until the training end condition is met to obtain the target model.

[0147] Wherein, the cross-machine bandwidth of the processor refers to the communication bandwidth between the processor on one device and the processor on another device. The bandwidth threshold can be set according to actual needs. The full gradient set includes the gradients corresponding to each model parameter in the model parameter set.

[0148] Specifically, when the cross-machine bandwidth between the first processor for training the initial model and the target processor is less than the bandwidth threshold, in order to improve communication efficiency, after obtaining the original gradient set, the processor can perform gradient screening on the original gradient set to obtain a compressed gradient set with a smaller data volume, and communicate the compressed gradient set with other processors.

[0149] When the cross-machine bandwidth between the first processor for training the initial model and the target processor is greater than or equal to the bandwidth threshold, in order to improve the accuracy of the model, after obtaining the original gradient set, the processor directly communicates the original gradient set with other processors. After the first processor obtains the original gradient sets sent by all processors, it fuses the respective original gradient sets to obtain the full gradient set. The full gradient set is used to update the model parameter set corresponding to the initial model. After updating the model parameter set corresponding to the initial model, return to the step of obtaining the target training subset corresponding to the first processor to perform model iterative training until the training end condition is met to obtain the target model.

[0150] It can be understood that the process of updating the model parameter set corresponding to the initial model based on the full gradient set can refer to the process of updating the model parameter set corresponding to the initial model based on the comprehensive gradient set.

[0151] In the above embodiments, the first processor and the target processor used for training the initial model are of the same type. When the cross-machine bandwidth of the first processor is less than the bandwidth threshold, the first processor compresses the original gradient set that needs to communicate with the target processor and then sends it, which can improve the communication efficiency, enabling the training of large-scale machine learning models even with a processor cluster having a small cross-machine bandwidth and effectively saving the model training cost. When the cross-machine bandwidth of the first processor is greater than or equal to the bandwidth threshold, the first processor directly communicates the original gradient set with the target processor, so that more model parameters can be updated subsequently, improving the accuracy of the model.

[0152] In one embodiment, the initial model is an initial image processing model, the target training subset and other training subsets are image training subsets, and the target model is a target image processing model.

[0153] Obtaining the original gradient set corresponding to the initial model on the target training subset based on the model parameter set corresponding to the initial model to be trained and the target training subset includes:

[0154] Performing image processing on the image training samples in the target training subset based on the model parameter set corresponding to the initial image processing model to obtain the predicted image labels corresponding to the image training samples; calculating the loss based on the training image labels and the predicted image labels corresponding to the image training samples to obtain the image loss corresponding to the initial image processing model on the target training subset; performing back calculation on the model parameter set based on the image loss to obtain the original gradient set corresponding to the initial image processing model on the target training subset.

[0155] Among them, the initial model to be trained can be an image processing model to be trained. The initial model is an image processing model to be trained, and correspondingly, the target model is a completed image processing model. The input data of the image processing model is an image, and the output data is the predicted image label corresponding to the image. The training subset corresponding to the image processing model includes image training samples and the image training labels corresponding to the image training samples.

[0156] For example, if the image processing model is an image segmentation model, then the image processing is image segmentation, the predicted image label is the predicted image segmentation result, the image training sample is an image with a known image segmentation result, and the training image label is the correct image segmentation result. By way of example, the image segmentation model can specifically be an image foreground segmentation model, an image background segmentation model, a face segmentation model, etc. If the image processing model is an image classification model, then the image processing is image classification, the predicted image label is the predicted image classification result, the image training sample is an image with a known image classification result, and the training image label is the correct image classification result. By way of example, the image classification model can specifically be an animal image classification model, a plant image classification model, etc.

[0157] Specifically, the first processor performs image processing on the image training samples in the target training subset based on the model parameter set corresponding to the initial image processing model, and obtains the predicted image labels corresponding to the image training samples. That is, the first processor inputs the image training samples into the initial image processing model to obtain the predicted image labels corresponding to the image training samples. The first processor calculates the loss based on the training image labels and the predicted image labels corresponding to the image training samples, and obtains the image loss corresponding to the initial image processing model on the target training subset. Specifically, the training image labels and the predicted image labels corresponding to the image training samples can be substituted into the loss function corresponding to the image processing model to obtain the image loss. The first processor performs back-calculation on the model parameter set based on the image loss, and obtains the original gradient set corresponding to the initial image processing model on the target training subset.

[0158] In the above embodiments, the method of the present application can be applied to model training for an image processing model to improve the training efficiency of the image processing model.

[0159] In a specific embodiment, as Figure 11 shown, a model training method is provided. Taking the application of this method to the current first processor as an example for illustration. Each of the first processors for training the initial model stores different local model parameters in the model parameter set corresponding to the initial model. Among them:

[0160] Step S1102: Obtain the target training subset corresponding to the current first processor, obtain the local model parameters stored by other first processors from other first processors, and form the model parameter set corresponding to the initial model to be trained by combining the local model parameters stored by itself and the local model parameters obtained from other first processors.

[0161] Step S1104: Perform forward calculation on the model parameter set corresponding to the initial model based on the training samples in the target training subset, obtain the predicted labels corresponding to the training samples, and calculate the loss based on the training labels and the predicted labels corresponding to the training samples, to obtain the model loss corresponding to the initial model on the target training subset.

[0162] Step S1106: Perform back-calculation on the model parameter set based on the model loss. For every reference number of gradients calculated, form the original gradient set corresponding to the initial model on the target training subset by combining the reference number of gradients. Obtain the target number of gradients from the original gradient set in descending order of the absolute value of the gradients, and form the compressed gradient set corresponding to the initial model on the target training subset; after performing the back-calculation, discard the other model parameters except the local model parameters stored by itself.

[0163] Step S1108: Send the compressed gradient set corresponding to the initial model on the target training subset to other first processors, and obtain the compressed gradient sets corresponding to the initial model on other training subsets from other first processors.

[0164] Step S1110: Accumulate the gradients corresponding to the same model parameter in each compressed gradient set of the initial model on each training subset to obtain an accumulated gradient set, and perform an averaging process on the accumulated gradient set based on the number of first processors corresponding to the initial model to obtain a comprehensive gradient set.

[0165] Step S1112: Obtain a target gradient set that matches the current first processor from the comprehensive gradient set, discard the remaining gradients, and send the target gradient set that matches the current first processor to the second processor, so that the second processor updates the local model parameters that match the current first processor in the model parameter set based on the target gradient set that matches the current first processor and the local optimizer parameters; the computing efficiency of the first processor is greater than that of the second processor.

[0166] Step S1114: Obtain the updated local model parameters fed back by the second processor, send the updated local model parameters to other first processors, and return to execute the step of obtaining the target training subset corresponding to the current first processor until the training end condition is met to obtain the target model.

[0167] In a specific embodiment, as Figure 12 shown, step S1102 includes:

[0168] Step S1202: Obtain the communication environment characteristics corresponding to the initial model, obtain a plurality of candidate transmission amounts preset for the communication environment characteristics, perform a communication efficiency test in the first processor cluster used to train the initial model based on the candidate transmission amounts to obtain the communication efficiency corresponding to each candidate transmission amount, and use the candidate transmission amount with the highest communication efficiency as the target transmission amount for the model parameters.

[0169] Step S1204: Group the model parameters known to the current first processor based on the target transmission amount to obtain a plurality of initial combinations of model parameters, fill the sizes of the plurality of initial combinations of model parameters to an integer multiple of the physical storage unit corresponding to the first processor to obtain a plurality of target combinations of model parameters, and sequentially send the plurality of target combinations of model parameters to other first processors.

[0170] In a specific embodiment, the method of the present application can be applied to a model training scenario based on a low-bandwidth training cluster. Although a high-bandwidth training cluster has high cross-machine communication efficiency, it is costly. For example, a high-bandwidth training cluster is composed of A800 / H800 models, and a low-bandwidth training cluster is composed of V100 / A100 models. Through the method of the present application, the cross-machine communication efficiency can be fully utilized in the low-bandwidth training cluster, and a large-scale machine learning model can be trained through the low-bandwidth training cluster. The low-bandwidth training cluster includes multiple GPUs.

[0171] A complete process of machine learning training mainly includes: data reading, forward calculation, backward calculation, gradient reduction, and weight update. Refer to Figure 13 , in the data parallel training mode, the communication scenarios between processes include two types: gradient reduction and weight aggregation. Weight aggregation means that each GPU collects the weight data of all GPUs. Gradient reduction is performed through the reduce operation, and weight aggregation is performed through the allgather operation. Through the method of the present application, for the model training scenario based on a low-bandwidth training cluster, the gradient reduction process and the weight aggregation process in the large model training process are optimized to improve the cross-machine communication efficiency of large model training. In the gradient reduction process, each time gradient communication occurs between GPUs, the top k gradients with the largest absolute values are selected to participate in the communication to reduce the communication volume and improve the efficiency of gradient reduction. In the weight aggregation process, the optimal packet size is determined by detecting the communication topology / network bandwidth / communication link, and multiple rounds of communication are sent based on the optimal packet size. If necessary, the data packet can be filled with zeros to an integer multiple of the GPU memory bit width during data transmission to improve the efficiency of weight aggregation.

[0172] The allgather operation is also required in the gradient reduction process. Through the allgather operation, each GPU can collect the compressed gradients on other GPUs. Both gradient reduction and weight aggregation require the allgather operation. Therefore, two interfaces are provided for users to integrate and use, which can achieve extreme performance optimization. In some scenarios of large model training, the performance can be improved by 4 times.

[0173] The first interface call is to automatically determine the optimal packet size according to the network bandwidth / communication link / communication topology. Both gradient reduction and weight aggregation need to use the first interface. For gradient reduction, the optimal packet size is the optimal gradient accumulation amount, and for weight aggregation, the optimal packet size is the optimal model parameter transmission amount.

[0174] The second interface call is to perform model parameter segmentation based on the determined optimal model parameter transmission amount, and if necessary, fill the optimal packet size with zeros to an integer multiple of the GPU memory bit width and then perform multiple rounds of communication transmission.

[0175] Referring to Table 1, taking the training scenario of a 10-billion model using 64 cards (i.e., 64 GPUs) with RDMA communication as an example, after optimization using the method of the present application, the performance is improved by about 4 times, and the single iteration time is reduced from 53 s to 13.2 s.

[0176] Table 1

[0177]

[0178] In state 3, a complete training process of the machine learning model is as follows:

[0179] 1. Initialization

[0180] The gradients and model parameters of the model are divided among all GPU processes. After the division, the CPU optimizer performs the initialization work.

[0181] 2. Gradient reduction during the backward calculation process

[0182] Gradients are generated during the backward calculation process of the model. The GPU process can call the reduce interface to perform the gradient reduction operation.

[0183] 3. Unloading the reduced gradients to the CPU

[0184] Each GPU process parallelly unloads its own share of the averaged gradients to the CPU and discards the parts that are not its responsibility.

[0185] 4. The CPU updates the weights

[0186] Since each CPU process has saved a part of the reduced gradients after the backward propagation, each CPU process is responsible for parallelly calling the adam optimizer to update the weights corresponding to this part of the gradients (i.e., the model parameters).

[0187] 5. Moving the updated weights back to the GPU

[0188] After the weights are updated, the CPU moves the part of the model parameters responsible for the GPU back to the GPU. At this time, each GPU process saves a part of the weights.

[0189] 6. Weight aggregation

[0190] All GPU processes obtain the full amount of weight information through the allgather interface.

[0191] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0192] Based on the same inventive concept, an embodiment of the present application further provides a model training device for implementing the above-mentioned model training method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the model training device provided below can refer to the limitations on the model training method in the above text, and will not be elaborated here.

[0193] In one embodiment, as Figure 14 shown, a model training device is provided, including: a training data acquisition module 1402, a gradient determination module 1404, a gradient compression module 1406, a gradient interaction module 1408, a gradient fusion module 1410, and a target model determination module 1412, where:

[0194] The training data acquisition module 1402 is configured to acquire a target training subset corresponding to the first processor.

[0195] The gradient determination module 1404 is configured to obtain an original gradient set corresponding to the initial model on the target training subset based on the model parameter set corresponding to the initial model to be trained and the target training subset.

[0196] The gradient compression module 1406 is configured to perform gradient screening on the original gradient set to obtain a compressed gradient set corresponding to the initial model on the target training subset.

[0197] The gradient interaction module 1408 is configured to send the compressed gradient set corresponding to the initial model on the target training subset to other target processors for training the initial model, and obtain the compressed gradient set corresponding to the initial model on other training subsets from the target processors.

[0198] The gradient fusion module 1410 is configured to fuse each compressed gradient set to obtain a comprehensive gradient set; the comprehensive gradient set is used to update the model parameter set corresponding to the initial model.

[0199] A target model determination module 1412, configured to, after the model parameter set corresponding to the initial model is updated, return to execute the step of obtaining a target training subset corresponding to the first processor until a training end condition is met, to obtain a target model.

[0200] In one embodiment, the gradient determination module 1404 is further configured to:

[0201] Perform forward calculation on the model parameter set corresponding to the initial model to be trained based on the training samples in the target training subset, to obtain predicted labels corresponding to the training samples;

[0202] Perform loss calculation based on the training labels and predicted labels corresponding to the training samples, to obtain a model loss corresponding to the initial model on the target training subset;

[0203] Perform backward calculation on the model parameter set based on the model loss, to obtain an original gradient set corresponding to the initial model on the target training subset.

[0204] In one embodiment, the gradient determination module 1404 is further configured to:

[0205] Perform backward propagation of the model loss in the model parameter set, and sequentially calculate gradients corresponding to multiple model parameters in the model parameter set;

[0206] According to the gradient calculation order, form local gradient sets by using a reference number of gradients, and sequentially obtain respective local gradient sets corresponding to the initial model on the target training subset;

[0207] For each obtained local gradient set, use the local gradient set as the original gradient set corresponding to the initial model on the target training subset, and enter the step of performing gradient screening on the original gradient set to obtain a compressed gradient set corresponding to the initial model on the target training subset.

[0208] In one embodiment, the gradient determination module 1404 is further configured to:

[0209] When the first processor and the target processor used for training the initial model each save a model parameter set, when performing forward calculation or backward calculation, perform forward calculation or backward calculation on the model parameter set saved by itself;

[0210] When the first processor and the target processor used for training the initial model each save local model parameters in the model parameter set, when performing forward calculation or backward calculation, obtain the local model parameters saved by the target processor from the target processor, form a model parameter set with the local model parameters saved by itself and the local model parameters obtained from the target processor, perform forward calculation or backward calculation on the model parameter set, and discard other model parameters except the local model parameters saved by itself after performing forward calculation or backward calculation.

[0211] In one embodiment, the gradient compression module 1406 is further configured to:

[0212] Obtain a target number of gradients from the original gradient set in descending order of gradient absolute value, and form a compressed gradient set corresponding to the initial model on the target training subset.

[0213] In one embodiment, the gradient fusion module 1410 is further configured to:

[0214] Accumulate the gradients corresponding to the same model parameter in each compressed gradient set to obtain an accumulated gradient set;

[0215] Based on the total number of the first processor and the target processor, perform an averaging process on the accumulated gradient set to obtain a comprehensive gradient set.

[0216] In one embodiment, the model training device is further configured to:

[0217] Obtain a target gradient set matching the first processor from the comprehensive gradient set, and discard the remaining gradients;

[0218] Send the target gradient set matching the first processor to the second processor, so that the second processor updates the local model parameters matching the first processor in the model parameter set based on the target gradient set matching the first processor and the local optimizer parameters; the computing efficiency of the first processor is greater than that of the second processor, and the first processor and the target processor are of the same type of processor;

[0219] After the local model parameters matching the first processor in the model parameter set are updated, obtain the updated local model parameters fed back by the second processor, and send the updated local model parameters to the target processor.

[0220] In one embodiment, the model training device is further configured to:

[0221] Obtain the communication environment characteristics corresponding to the initial model; the communication environment characteristics include at least one of network bandwidth information, communication link information, and communication topology information;

[0222] Determine the target transmission amount for the model parameters based on the communication environment characteristics;

[0223] Based on the target transmission amount, group and send the model parameters known to the first processor in the model parameter set corresponding to the initial model to the target processor, so as to jointly perform parallel model training on the initial model by the target processor based on their respective corresponding training subsets.

[0224] In one embodiment, the model training device is further configured to:

[0225] Obtain a plurality of candidate transmission amounts preset for the communication environment characteristics;

[0226] Based on the candidate transmission volume, communication efficiency tests are conducted in the processor cluster used for training the initial model, and the communication efficiency corresponding to each candidate transmission volume is obtained;

[0227] The candidate transmission volume with the highest communication efficiency is used as the target transmission volume for the model parameters.

[0228] In one embodiment, the model training device is further configured to:

[0229] Based on the target transmission volume, group the model parameters known to the first processor in the model parameter set corresponding to the initial model to obtain multiple initial combinations of model parameters;

[0230] Pad the sizes of the multiple initial combinations of model parameters to an integer multiple of the physical storage unit corresponding to the target processor to obtain multiple target combinations of model parameters;

[0231] Send the multiple target combinations of model parameters to the target processor in sequence.

[0232] In one embodiment, the first processor and the target processor are of the same type of processor. The model training device is further configured to:

[0233] When the cross-machine bandwidth of the first processor is less than the bandwidth threshold, enter the step of performing gradient screening on the original gradient set to obtain the compressed gradient set corresponding to the initial model on the target training subset;

[0234] When the cross-machine bandwidth of the first processor is greater than or equal to the bandwidth threshold, obtain the original gradient sets corresponding to the initial model on other training subsets from the target processor, fuse the respective original gradient sets to obtain the full gradient set for updating the model parameter set corresponding to the initial model, and after the model parameter set corresponding to the initial model is updated, return to the step of obtaining the target training subset corresponding to the first processor and execute until the training end condition is met to obtain the target model.

[0235] In one embodiment, the initial model is an initial image processing model, the target training subset and other training subsets are image training subsets, and the target model is a target image processing model. The gradient determination module 1404 is further configured to:

[0236] Based on the model parameter set corresponding to the initial image processing model, perform image processing on the image training samples in the target training subset to obtain the predicted image labels corresponding to the image training samples;

[0237] Calculate the loss based on the training image labels and the predicted image labels corresponding to the image training samples to obtain the image loss corresponding to the initial image processing model on the target training subset;

[0238] Based on the image loss, perform reverse calculation on the model parameter set to obtain the original gradient set corresponding to the initial image processing model on the target training subset.

[0239] The above model training device is used to train each processor of the initial model. Based on the model parameter set corresponding to the initial model and their respective training subsets, parallel model training is performed on the initial model, which can improve the model training efficiency. When performing model training, calculate the original gradient set corresponding to the initial model on the training subset, compress the original gradient set to obtain the compressed gradient set corresponding to the initial model on the training subset, and each processor communicates its respective compressed gradient set. The data volume of the compressed gradient set is smaller than that of the original gradient set, which can improve the communication efficiency and thus further improve the model training efficiency. The processor fuses its own compressed gradient set with the compressed gradient sets of other processors to obtain a comprehensive gradient set. The comprehensive gradient set integrates the gradient information of the initial model on each training subset. Adjust the model parameters based on the comprehensive gradient set, enabling the model to learn knowledge from each training subset, which can improve the prediction ability of the model and the quality of model training.

[0240] Each module in the above model training device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0241] In one embodiment, a computer device is provided. This computer device can be a server, and its internal structure diagram can be as Figure 15 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of this computer device is used to store data such as model parameters and training subsets. The input / output interface of this computer device is used to exchange information between the processor and external devices. The communication interface of this computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a model training method.

[0242] In one embodiment, a computer device is provided. This computer device can be a terminal, and its internal structure diagram can be as Figure 16As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a model training method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0243] Those skilled in the art can understand that Figure 15 , 16 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0244] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0245] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0246] In one embodiment, a computer program product is provided, and the computer program product includes a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0247] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0248] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0249] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0250] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A model training method, characterized in that, Applied to a first processor, the method includes: Obtaining a target training subset corresponding to the first processor; Based on a model parameter set corresponding to an initial model to be trained and the target training subset, obtaining an original gradient set corresponding to the initial model on the target training subset; Performing gradient screening on the original gradient set to obtain a compressed gradient set corresponding to the initial model on the target training subset; Sending the compressed gradient set corresponding to the initial model on the target training subset to other target processors for training the initial model, and obtaining the compressed gradient set corresponding to the initial model on other training subsets from the target processors; Fusing each compressed gradient set to obtain a comprehensive gradient set; the comprehensive gradient set is used to update the model parameter set corresponding to the initial model; After the model parameter set corresponding to the initial model is updated, returning to execute the step of obtaining the target training subset corresponding to the first processor until a training end condition is satisfied to obtain a target model.

2. The method according to claim 1, wherein The obtaining an original gradient set corresponding to the initial model on the target training subset based on a model parameter set corresponding to an initial model to be trained and the target training subset includes: Based on training samples in the target training subset, performing forward calculation on the model parameter set corresponding to the initial model to be trained to obtain predicted labels corresponding to the training samples; Performing loss calculation based on training labels and predicted labels corresponding to the training samples to obtain a model loss corresponding to the initial model on the target training subset; Based on the model loss, performing backward calculation on the model parameter set to obtain an original gradient set corresponding to the initial model on the target training subset.

3. The method according to claim 2, characterized in that, The performing backward calculation on the model parameter set based on the model loss to obtain an original gradient set corresponding to the initial model on the target training subset includes: Performing backward propagation of the model loss in the model parameter set and sequentially calculating gradients corresponding to multiple model parameters in the model parameter set; According to the gradient calculation order, forming a local gradient set with a reference number of gradients, and sequentially obtaining each local gradient set corresponding to the initial model on the target training subset; For each obtained local gradient set, using the local gradient set as the original gradient set corresponding to the initial model on the target training subset, and entering the step of performing gradient screening on the original gradient set to obtain a compressed gradient set corresponding to the initial model on the target training subset for execution.

4. The method according to claim 2, wherein The method further includes: When the first processor and the target processors for training the initial model each save the model parameter set, performing forward calculation or backward calculation on the model parameter set saved by itself during forward calculation or backward calculation; When the first processor and the target processor used for training the initial model each save the local model parameters in the model parameter set, during forward calculation or backward calculation, obtain the local model parameters saved by the target processor from the target processor, combine the local model parameters saved by itself and the local model parameters obtained from the target processor to form the model parameter set, perform forward calculation or backward calculation on the model parameter set, and after the forward calculation or backward calculation, discard other model parameters except the local model parameters saved by itself.

5. The method according to claim 1, wherein The gradient screening of the original gradient set to obtain the compressed gradient set corresponding to the initial model on the target training subset includes: Obtain a target number of gradients from the original gradient set in descending order of the absolute value of the gradients to form the compressed gradient set corresponding to the initial model on the target training subset.

6. The method according to claim 1, wherein The fusion of each compressed gradient set to obtain the comprehensive gradient set includes: Accumulate the gradients corresponding to the same model parameter in each compressed gradient set to obtain an accumulated gradient set; Based on the total number of the first processor and the target processor, perform an averaging process on the accumulated gradient set to obtain the comprehensive gradient set.

7. The method according to claim 1, wherein The method further includes: Obtain a target gradient set matching the first processor from the comprehensive gradient set and discard the remaining gradients; Send the target gradient set matching the first processor to the second processor so that the second processor updates the local model parameters in the model parameter set matching the first processor based on the target gradient set matching the first processor and the local optimizer parameters; the computing efficiency of the first processor is greater than that of the second processor, and the first processor and the target processor are of the same type of processor; After the local model parameters in the model parameter set matching the first processor are updated, obtain the updated local model parameters fed back by the second processor and send the updated local model parameters to the target processor.

8. The method according to claim 1, wherein The method further includes: Obtain the communication environment characteristics corresponding to the initial model; the communication environment characteristics include at least one of network bandwidth information, communication link information, and communication topology information; Determine the target transmission amount for the model parameters based on the communication environment characteristics; Based on the target transmission amount, group and send the model parameters in the model parameter set corresponding to the initial model that are known to the first processor to the target processor to jointly perform parallel model training on the initial model by the target processor based on their respective corresponding training subsets.

9. The method according to claim 8, wherein The determining of the target transmission amount for the model parameters based on the communication environment characteristics includes: Obtain a plurality of candidate transmission amounts preset for the communication environment characteristics; Based on the candidate transmission amounts, perform communication efficiency tests in the processor cluster used for training the initial model to obtain the communication efficiencies corresponding to the respective candidate transmission amounts; Use the candidate transmission amount with the highest communication efficiency as the target transmission amount for the model parameters.

10. The method according to claim 8, wherein Based on the target transmission volume, grouping the model parameters in the model parameter set corresponding to the initial model that are known to the first processor and sending them to the target processor includes: Based on the target transmission volume, grouping the model parameters in the model parameter set corresponding to the initial model that are known to the first processor to obtain a plurality of initial model parameter combinations; Filling the sizes of the plurality of initial model parameter combinations to an integer multiple of the physical storage unit corresponding to the target processor to obtain a plurality of target model parameter combinations; Sequentially sending the plurality of target model parameter combinations to the target processor.

11. The method according to claim 1, wherein The first processor and the target processor are of the same type, and the method further includes: When the cross-machine bandwidth of the first processor is less than the bandwidth threshold, entering the step of performing gradient screening on the original gradient set to obtain the compressed gradient set corresponding to the initial model on the target training subset; When the cross-machine bandwidth of the first processor is greater than or equal to the bandwidth threshold, obtaining the original gradient sets corresponding to the initial model on other training subsets from the target processor, fusing each original gradient set to obtain a full gradient set for updating the model parameter set corresponding to the initial model, and after updating the model parameter set corresponding to the initial model, returning to execute the step of obtaining the target training subset corresponding to the first processor until the training end condition is satisfied to obtain the target model.

12. The method according to any one of claims 1 to 11, characterized in that, The initial model is an initial image processing model, the target training subset and the other training subsets are image training subsets, and the target model is a target image processing model; Based on the model parameter set corresponding to the initial model to be trained and the target training subset, obtaining the original gradient set corresponding to the initial model on the target training subset includes: Based on the model parameter set corresponding to the initial image processing model, performing image processing on the image training samples in the target training subset to obtain the predicted image labels corresponding to the image training samples; Calculating the loss based on the training image labels and the predicted image labels corresponding to the image training samples to obtain the image loss corresponding to the initial image processing model on the target training subset; Based on the image loss, performing reverse calculation on the model parameter set to obtain the original gradient set corresponding to the initial image processing model on the target training subset.

13. A model training device, characterized in that, The device includes: A training data acquisition module, configured to acquire a target training subset corresponding to the first processor; A gradient determination module, configured to obtain the original gradient set corresponding to the initial model on the target training subset based on the model parameter set corresponding to the initial model to be trained and the target training subset; A gradient compression module, configured to perform gradient screening on the original gradient set to obtain the compressed gradient set corresponding to the initial model on the target training subset; A gradient interaction module, configured to send the compressed gradient set corresponding to the initial model on the target training subset to other target processors for training the initial model, and obtain the compressed gradient sets corresponding to the initial model on other training subsets from the target processor; A gradient fusion module, which is used to fuse each compressed gradient set to obtain a comprehensive gradient set; the comprehensive gradient set is used to update the model parameter set corresponding to the initial model; A target model determination module, which is used to, after the model parameter set corresponding to the initial model is updated, return to execute the step of obtaining the target training subset corresponding to the first processor until the training end condition is met, and obtain a target model.

14. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

Citation Information

Cited By

  • Model training method and device, storage medium and computer program product

    CN121352066A