Model Training Method and System, Non-Volatile Storage Medium, and Computer Terminal

By initializing the feature extraction network and fully connected layer in a parallel training framework on multiple GPUs, the problem of large-scale data set consumption is solved, and more efficient training and faster model convergence is achieved.

CN114463158BActive Publication Date: 2025-06-13ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011249605.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-10
Publication Date
2025-06-13
Estimated Expiration
2040-11-10

AI Technical Summary

Technical Problem

When the data set sample size is large, a classification model that supports the same number of classification results needs to be trained, resulting in a large consumption of GPU computing resources.

Method used

Multiple GPUs are used to form a parallel training framework to initialize the feature extraction network of each GPU, and input the initialized features to the corresponding fully connected layer to realize the initialization of the fully connected layer, thereby improving training efficiency and model convergence speed.

Benefits of technology

It effectively reduces the consumption of GPU computing resources during training on large-scale data sets, and improves the training efficiency and convergence speed of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114463158B_ABST
    Figure CN114463158B_ABST
Patent Text Reader

Abstract

The present application discloses a model training method, system, non-volatile storage medium, and computer terminal. Among them, the method includes: each of multiple GPUs has a first feature extraction network and a first fully connected layer, wherein the network structures of the first feature extraction networks in the multiple GPUs are the same; the first GPU among the multiple GPUs is used to initialize the first feature extraction network of the first GPU, and extract first sample features in the target training dataset by using the initialized first feature extraction network; input the first sample features into the first fully connected layer of the first GPU for processing; and determine the prediction error of the first GPU based on the processing result; determine the target prediction error of the target neural network model based on the prediction error of the first GPU and the prediction errors of other received GPUs; and update the parameters of the target neural network model based on the target prediction error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning, and more particularly, to a model training method and system, a non-volatile storage medium, and a computer terminal. Background Art

[0002] The purpose of unsupervised learning or self-supervised learning is to learn a model or features with strong expressive power through unlabeled data. Usually, a pretext task needs to be defined to guide the training of the model. The pretext task includes but is not limited to: change prediction, image completion, spatial or temporal order prediction, clustering, data generation, etc.

[0003] Among them, in the field of unsupervised learning, the method of instance classification regards each data sample in the dataset as a class, and the same training network as supervised classification can be used, which can make full use of all negative examples in the dataset. Therefore, it is a relatively potential solution.

[0004] However, to implement the instance classification method, a classification model with the same size as the dataset sample size needs to be trained. For a dataset with a relatively large amount of data, it requires a large amount of video memory resources of the Graphics Processing Unit (GPU). For example, for a training image dataset, if there are millions or even tens of millions of sample sizes, it is necessary to train the classification model to recognize the same number of sample types, which will seriously consume the computing resources of the GPU.

[0005] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0006] Embodiments of this application provide a model training method and system, a non-volatile storage medium, and a computer terminal, so as to at least solve the technical problem that when the sample size in the dataset is relatively large, the classification model needs to support the same number of classification results through the training process, resulting in a large consumption of the computing resources of the GPU.

[0007] According to one aspect of the embodiments of the present application, a model training system is provided, including: a GPU cluster, where the GPU cluster includes multiple GPUs, and among them: each of the multiple GPUs has a first feature extraction network and a first fully connected layer, where the network structures of the first feature extraction networks in the multiple GPUs are the same, and the first fully connected layer is obtained by splitting the second fully connected layer in the target neural network model; the first GPU among the multiple GPUs is used to initialize the first feature extraction network of the first GPU, and extract the first sample features in the target training dataset by using the initialized first feature extraction network, where the first GPU is any one of the multiple GPUs; input the first sample features into the first fully connected layer of the first GPU for processing; and determine the prediction error of the first GPU based on the processing result; determine the target prediction error of the target neural network model based on the prediction error of the first GPU and the prediction errors of other received GPUs; and update the parameters of the target neural network model based on the target prediction error.

[0008] According to another aspect of the embodiments of the present application, a model training method is further provided. This method is applied to a GPU cluster, where the GPU cluster includes multiple GPUs, and each of the multiple GPUs has a first feature extraction network and a first fully connected layer, where the network structures of the first feature extraction networks in the multiple GPUs are the same, and the first fully connected layer is obtained by splitting the second fully connected layer in the target neural network model; the method includes: the first GPU among the multiple GPUs initializes the first feature extraction network of the first GPU, and extracts the first sample features in the target training dataset by using the initialized first feature extraction network, where the first GPU is any one of the multiple GPUs; the first GPU inputs the first sample features into the first fully connected layer of the first GPU for processing; the first GPU determines the prediction error of the first GPU based on the processing result; determines the target prediction error of the target neural network model based on the prediction error of the first GPU and the prediction errors of other received GPUs; and updates the parameters of the target neural network model based on the target prediction error.

[0009] According to another aspect of the embodiments of the present application, there is also provided a model training method, including: grouping a training data set to obtain a plurality of target training data sets; respectively inputting the plurality of target training data sets into a plurality of electronic devices, wherein each of the plurality of electronic devices includes a fully connected layer and a feature extraction network with the same network structure connected to the fully connected layer, and the fully connected layer is obtained by splitting the fully connected layer in the target neural network model; initializing the feature extraction networks of the plurality of electronic devices, and using the initialized feature extraction networks to extract sample features in the target training data sets; inputting the sample features into the fully connected layers in the plurality of electronic devices for processing to classify the sample features; determining a plurality of prediction errors according to the classification results of the plurality of electronic devices and the sample labels in the target training data sets; and updating the target neural network model based on the plurality of prediction errors.

[0010] According to yet another aspect of the embodiments of the present application, there is also provided a model training method, including: grouping a training data set to obtain a plurality of target training data sets; respectively inputting the plurality of target training data sets into a plurality of electronic devices to train the neural network models in the plurality of electronic devices, wherein each of the neural network models in the plurality of electronic devices includes a fully connected layer and a feature extraction network connected to the fully connected layer, the structures of the feature extraction networks in the plurality of electronic devices are the same, and the fully connected layer is obtained by splitting the fully connected layer in the target neural network model; obtaining a plurality of prediction errors obtained by training the plurality of neural network models; determining the prediction error of the target neural network based on the plurality of prediction errors, and updating the target neural network model based on the prediction error.

[0011] According to yet another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium, wherein the non-volatile storage medium includes a stored program, and when the program runs, it controls the device where the non-volatile storage medium is located to execute the above-mentioned model training method.

[0012] According to still another aspect of the embodiments of the present application, there is also provided a computer terminal, including: a processor; and a memory connected to the processor for providing instructions for the processor to perform the following processing steps: initializing a first feature extraction network of a first GPU, and using the initialized first feature extraction network to extract first sample features in a target training data set, wherein the first GPU is any one of a plurality of GPUs; inputting the first sample features into a first fully connected layer of the first GPU for processing, wherein the first fully connected layer is obtained by splitting a second fully connected layer in the target neural network model; determining a prediction error of the first GPU based on the processing result; determining a target prediction error of the target neural network model based on the prediction error of the first GPU and the prediction errors of other received GPUs; and updating the parameters of the target neural network model based on the target prediction error.

[0013] According to another aspect of the embodiments of the present application, a model training system is further provided, including: a client device and a server, wherein: the client device is configured to provide a human-computer interaction interface to a target object, and call program instructions in the server through the human-computer interaction interface to perform the following steps: group a training data set to obtain a plurality of target training data sets; input the plurality of target training data sets into a plurality of electronic devices respectively to train neural network models in the plurality of electronic devices, wherein the neural network models in the plurality of electronic devices all include a fully connected layer and a feature extraction network connected to the fully connected layer, and the structures of the feature extraction networks in the plurality of electronic devices are the same; obtain a plurality of prediction errors obtained by training the plurality of neural network models; determine a prediction error of a target neural network based on the plurality of prediction errors, and update the target neural network model based on the prediction error.

[0014] In the embodiments of the present application, a parallel training framework is composed of a plurality of GPUs, the feature extraction networks of each GPU are initialized, and the initialized features are input into the corresponding fully connected layer to initialize the fully connected layer, so as to improve the training efficiency while also improving the convergence speed of the model, and further solve the technical problem that when the number of samples in the data set is relatively large, it is necessary to make the classification model support the same number of classification results through the training process, resulting in a large consumption of computing resources of the GPU. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0016] Figure 1 is a schematic structural diagram of a model training system according to an embodiment of the present application;

[0017] Figure 2 is a schematic structural diagram of an optional model training system according to an embodiment of the present application;

[0018] Figure 3 is a schematic flowchart of a model training method according to an embodiment of the present application;

[0019] Figure 4 is a schematic structural diagram of a computer terminal according to an embodiment of the present application;

[0020] Figure 5 is a schematic flowchart of another model training method according to an embodiment of the present application;

[0021] Figure 6 It is a schematic flowchart of another model training method according to an embodiment of the present application;

[0022] Figure 7 It is a schematic structural diagram of a model training system according to an embodiment of the present application. Detailed implementation manners

[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:

[0026] Deep learning: An artificial neural network structure with a high number of layers, which can be used to implement functions such as intelligent image detection and classification.

[0027] Instance classification: A task of classifying each data sample as a class.

[0028] Self-supervised learning: A type of unsupervised learning that does not rely on manual annotation and uses the data itself for model learning.

[0029] Feature extraction network: A network model used to extract sample features from training data, such as a convolutional neural network (CNN), a residual network.

[0030] Fully connected layer: Each node in the fully connected layer is connected to all nodes in the previous layer, and is used to synthesize the features extracted previously. Due to its fully connected nature, the parameters of the fully connected layer are generally the most. In the CNN structure, after multiple convolutional layers and pooling layers, there is one or more fully connected layers connected. Similar to the MLP, each neuron in the fully connected layer is fully connected to all neurons in its previous layer. The fully connected layer can integrate the local information with class discriminability in the convolutional layer or pooling layer. To improve the performance of the CNN network, the activation function of each neuron in the fully connected layer generally uses the ReLU function. The output value of the last fully connected layer is passed to an output, and softmax logistic regression can be used for classification, and this layer can also be called the softmax layer. For a specific classification task, it is very important to select a suitable loss function. There are several commonly used loss functions in CNN, each with different characteristics. Usually, the fully connected layer of CNN is the same as the MLP structure, and the BP algorithm is also mostly used for the training algorithm of CNN

[0031] Embodiment 1

[0032] The purpose of unsupervised learning or self-supervised learning is to learn a model or features with strong expressive ability through unlabeled data. Usually, a pretext task needs to be defined to guide the training of the model. The pretext task includes but is not limited to: change prediction, image completion, spatial or temporal order prediction, clustering, data generation, etc. Among them, the method based on contrastive learning uses a dual-branch network. The input of the network is usually two data augmentations of an image, and the purpose is to make the two data augmentation inputs of the same image closer in the feature space distance, while pulling the feature distances corresponding to the data augmentations of different images farther apart. The contrastive learning method requires a large number of negative samples, needs to be based on a large batch size, or needs a storage queue to save historical feature vectors. However, even so, the diversity of negative samples is still relatively scarce.

[0033] The method of instance classification regards each data sample in the dataset as a class, can use the same training network as supervised classification, does not require a dual-channel network, and can make full use of all negative examples in the dataset. The solution in the embodiments of the present application can be applied to the method of instance classification.

[0034] However, to implement the instance classification method, a classification model of the same size as the number of samples in the dataset needs to be trained. For datasets with a relatively large amount of data, this is often quite difficult and requires extremely large GPU video memory consumption (for example, for the ImageNet dataset with 1.28 million samples, a classification model supporting 1.28 million sample types needs to be trained). To solve this difficulty, one solution is to perform negative sampling, where only a part of other samples are sampled as negative examples during each training process. However, this actually loses the original design concept of the instance classification method. At the same time, there is only one data sample for each class in the instance classification method, and a training cycle can only encounter this sample once, making training difficult. Some algorithms use a more complex data scheduler to alleviate this problem. Thus, no matter which training method is adopted, it will consume a large amount of computing resources and affect the training efficiency. To solve the above technical problems, the embodiments of the present application provide corresponding solutions, which will be described in detail below.

[0035] The GPU is a schematic structural diagram of a model training system according to an embodiment of the present application. As Figure 1 shown, the model training system includes: a GPU cluster 1, and the GPU cluster 1 includes multiple GPUs (GPU1, GPU2, ··· GPUN), where N is a natural number, and:

[0036] Each of the multiple GPUs (GPU1, GPU2, ··· GPUN) has a first feature extraction network 10 and a first fully connected layer 12. Among them, the network structures of the first feature extraction networks 10 in the multiple GPUs are the same, and the first fully connected layer 10 is obtained by splitting the second fully connected layer in the target neural network model;

[0037] The first GPU among the multiple GPUs (for example, Figure 1 any one of GPU1, GPU2, ··· GPUN) is used to initialize the first feature extraction network 10 of the first GPU, and extract the first sample features in the target training dataset using the initialized first feature extraction network. Here, the first GPU is any one of the multiple GPUs; input the first sample features into the first fully connected layer 12 of the first GPU for processing; and determine the prediction error of the first GPU based on the processing result; determine the target prediction error of the target neural network model based on the prediction error of the first GPU and the prediction errors of other received GPUs; and update the parameters of the target neural network model based on the target prediction error.

[0038] As can be seen from the above solution, the same feature extraction network exists in multiple GPUs, and there is a fully connected layer corresponding to the feature extraction network. Therefore, it is possible to split the data in the training dataset and input it in parallel to the corresponding GPUs for processing, reducing the complexity of the training process. At the same time, since the sample features of the training data are extracted using the initialized feature extraction network before the training data is input into the corresponding fully connected layer, rather than initializing the fully connected layer, the convergence speed can be effectively improved, and the convergence accuracy can be increased.

[0039] It should be noted that in the parallel training framework composed of a GPU cluster, this training framework can be applied to the training process of supervised machine learning models. In the embodiments of the present application, the parallel training framework is applied to the self-supervised learning process, and the machine learning model is split into two parts: a feature extraction network and a fully connected layer. Among them, the first feature extraction network of the first GPU (i.e., any one of the GPUs in the GPU cluster) is obtained by copying the first feature extraction network of any one of the multiple GPUs. Since each node in the fully connected layer is connected to all nodes in the previous layer to integrate the features of the previous layer, splitting the fully connected layer is equivalent to grouping the nodes in the fully connected layer of the target neural network model, and each group of nodes corresponds to a sub-fully connected layer (i.e., the first fully connected layer).

[0040] The process of inputting the above first sample features into the first fully connected layer for processing can be manifested as the following implementation process: after the first sample features are input into the fully connected layer, the distributed feature identifiers corresponding to the first sample features are mapped to the sample label space, thereby realizing the function of the classifier. Generally speaking, the fully connected layer and the softmax layer together complete the classification function, where the softmax layer is used to output the final classification result.

[0041] In some embodiments of the present application, during the initialization process of the feature extraction network, after the first GPU (i.e., any one of the multiple GPUs) initializes the feature extraction network corresponding to the first GPU, the initialization parameters used for initialization are sent to other GPUs; other GPUs initialize the feature extraction networks corresponding to other GPUs based on the initialization parameters. By using the above method, the consistency of the initialization process can be ensured. Of course, in some alternative embodiments, the feature extraction networks in the above multiple GPUs can also be initialized simultaneously.

[0042] The above initialization parameters include but are not limited to: the convolutional layer parameters in the CNN, such as the parameters of the residual network itself, and can also include the FC parameters in the multi-layer perceptron (abbreviated as MLP).

[0043] In addition, during the training process of a machine learning model, a large amount of data is often required. At this time, if all the data in the training data set is input into a single GPU, it will inevitably affect the training efficiency. Therefore, to ensure the training efficiency, the training data set can be split. Specifically: The target training data set is obtained by dividing the training data set corresponding to the target neural network model according to the number of samples, where the number of samples in each target training data set is the same. For example, if there are 1 million samples in the training data set and there are 10 GPUs in the model training system (device), then the 1 million samples can be divided into 10 groups, with 100,000 samples in each group, and each group is input into one of the 10 GPUs respectively.

[0044] Since the training process is split in the above embodiments, when determining the prediction error (loss value) of the target neural network model, the loss values obtained from multiple GPUs need to be combined. Specifically, multiple GPUs are used to determine the prediction error of each GPU in the following way: The first GPU determines the loss value of the first GPU and obtains the loss values of other GPUs, where the other GPUs are the GPUs other than the first GPU among the multiple GPUs; the first GPU determines the prediction error based on the loss value of the first GPU and the loss values of other GPUs.

[0045] For example, determine the sum value between the loss value of the first GPU and the loss values of other GPUs, and use this sum value as the above prediction error. When determining the prediction error, each GPU determines the network gradient based on its own loss value, then each GPU adds the gradients of other GPUs to obtain the target gradient, and updates the target neural network model based on this target gradient.

[0046] In some embodiments, the first GPU is further configured to determine the prediction error based on the classification result of the first sample feature and the classification label of the first sample feature according to the processing result. Among them, before determining the prediction error based on the classification result of the first sample feature and the classification label of the first sample feature according to the processing result, calculate the similarity between the first sample feature and other sample features in the target training data set to obtain multiple similarities; sort the multiple similarities in descending order, and determine the first N similarities in the sorting result, and set the classification labels of the sample features corresponding to the first N similarities as the classification label of the first sample feature. Where N is a natural number.

[0047] Taking instance classification as an example, in instance classification, each data sample is regarded as a class. However, in the actual data, there is a high probability that some samples are semantically similar. Considering these samples as negative samples will introduce a lot of noise, which has a negative impact on the training convergence. Therefore, during the training process, by finding the top-K similar classes for each instance class (equivalent to the classes corresponding to the top N similarities mentioned above), and assigning certain positive labels to these similar classes, the label noise can be reduced. Specifically, for the original data x i The label is: Y i =[0,0,…,1,…0], where y i =1 and the other positions are 0. For class i, the top-K negative class set H i ={c 1 ,c 2 ,…,c K} is calculated. Then each bit y i in the smoothed Y j is:

[0048] α is a hyperparameter or a smoothing factor.

[0049] Calculating the top-K can be done once per training epoch. Using the smoothed labels compared to the original labels can significantly improve the final performance. For example, it can effectively suppress noise.

[0050] To better understand the above embodiments, taking instance classification as an example, the following is combined with Figure 2 for illustration.

[0051] Figure 2 is a schematic structural diagram of an optional model training system according to an embodiment of the present application. As Figure 2 shown, Figure 2 provides a hybrid parallel training framework for large-scale classification. This training framework was generally used for supervised model training before. In the embodiments of the present application, the hybrid parallel training framework is used in self-supervised learning. The hybrid parallel splits a classifier based on a deep neural network into two parts: a feature extraction network and an FC layer. In hybrid parallel training, the FC layer is split in a model parallel manner, while the feature extraction network is replicated in a data parallel manner (here, replication means replicating the network, that is, the feature extraction networks on each GPU branch are the same)

[0052] As Figure 2 shown, Figure 2The hybrid parallel training framework in it includes: GPU#1, GPU#2, ······ GPU#T. Taking GPU#1 as an example, the feature extraction part in this branch includes: data loading (sub-batch#1), an encoder (Encoder) for encoding the data, a pooling layer (Pool), and a multi-layer perceptron (MPL). The features extracted by the feature extraction network are input into the fully connected layer. Among them, the fully connected layer includes those for performing forward calculation of the FC layer on the extracted features (W i ) calculation part weights, calculation part logits (partial logits), calculation part Loss (partial loss), and then to backward calculation. After performing label smoothing operation, parameter update is carried out to complete an iteration process.

[0053] In the related technology, when initializing the fully connected layer, the default FC initialization method is to randomly initialize using the Gaussian distribution. This initialization method is particularly difficult to converge in the initial stage of instance classification training (related to the fact that there is only one data sample for each class in instance classification). In Figure 2 the hybrid parallel training architecture shown, a new fully connected layer (fully connected layer, that is Figure 2 the W in i ) initialization method is provided: using the features extracted by the randomly initialized feature extraction network for FC initialization, and it is found that the convergence speed can be effectively improved, and finally there is a relatively high convergence accuracy. Specifically, in the first cycle (epoch) of network training, all the randomly initialized parameters of the feature extraction network are fixed (but the mean and variance of the batch normalization layer, batch norm layer are normally statistically calculated), and at the same time, the corresponding FC parameters are initialized using the features of each data sample extracted. Among them, the FC parameters include: the W matrix, which is of size NxD, N is the number of samples (number of classes), D is the feature dimension. Since it is instance classification and each sample corresponds to one class, after one training cycle, all the W parameters (that is, the W matrix, FC parameters) are exactly initialized once.

[0054] In instance classification, each data sample is regarded as a class, but in the actual data, there is a high probability that some samples are semantically similar. Considering these samples as negative samples will introduce more noise and have a negative impact on training convergence. Therefore, in the solution provided in the embodiments of this application, during the training process, by finding the top-K similar classes for each instance class and assigning certain positive labels to these similar classes, label noise is reduced. Specifically, the original data x i The label is: Y i = [0, 0, …, 1, … 0], where y i= 1, and all other positions are 0. For class i, the top-K negative class set H is calculated. i = {c 1 , c 2 , …, c K}. Then, each bit y i in the smoothed Y j is:

[0055]

[0056] Calculate the top-K once per training epoch. Using the smoothed labels instead of the original labels can significantly improve the final performance.

[0057] It is easy to notice that multiple GPUs can be initialized independently or uniformly. If initialized independently, there is no need for communication between these multiple GPUs. If initialized uniformly, after one GPU finishes initialization, it can broadcast its initialized parameters to other GPUs to complete the unified initialization among all GPUs. For the latter, since multiple GPUs need to cooperate to train the target neural network model using the training dataset, communication is also required among all GPUs. Multiple GPUs can communicate with each other based on an application programming interface function library that supports distributed communication and computing. Such a function library can be, for example, the Message Passing Interface (MPI) library or the Rabit library, etc.

[0058] Embodiment 2

[0059] This embodiment of the present application also provides a model training method, which is applied to a GPU cluster. The GPU cluster includes multiple GPUs, and each GPU in the multiple GPUs has a first feature extraction network and a first fully connected layer. Among them, the network structures of the first feature extraction networks in the multiple GPUs are the same, and the first fully connected layer is obtained by splitting the second fully connected layer in the target neural network model. It should be noted that the optional implementation scheme of the GPU cluster in this embodiment can adopt the implementation scheme of the GPU cluster described in Embodiment 1, but is not limited thereto.

[0060] As Figure 3 shown, the method includes:

[0061] Step S302, the first GPU in the multiple GPUs initializes the first feature extraction network of the first GPU, and uses the initialized first feature extraction network to extract the first sample features in the target training dataset, where the first GPU is any one of the multiple GPUs;

[0062] Step S304, the first GPU inputs the first sample feature into the first fully-connected layer of the first GPU for processing;

[0063] Step S306, the first GPU determines the prediction error of the first GPU based on the processing result; based on the prediction error of the first GPU and the received prediction errors of other GPUs, determine the target prediction error of the target neural network model; and

[0064] Step S308, update the parameters of the target neural network model based on the target prediction error.

[0065] In some embodiments, the process of initializing the feature extraction network may be expressed in the following form, but is not limited thereto: after the first GPU initializes the feature extraction network corresponding to the first GPU, the initialization parameters used for initialization are sent to other GPUs; other GPUs initialize the feature extraction network corresponding to other GPUs based on the initialization parameters. By adopting the above method, the consistency of the initialization process can be ensured. Of course, in some alternative embodiments, the feature extraction networks in the above-mentioned multiple GPUs may also be initialized simultaneously.

[0066] The above-mentioned initialization parameters include but are not limited to: the convolutional layer parameters in the CNN, such as the parameters of the residual network itself, and may also include the FC parameters in the multi-layer perceptron (MLP).

[0067] In some embodiments, the above-mentioned prediction error may be implemented in the following manner: the first GPU determines the loss value of the first GPU and obtains the loss values of other GPUs, where the other GPUs are the GPUs other than the first GPU among the multiple GPUs; the first GPU determines the prediction error based on the loss value of the first GPU and the loss values of other GPUs. Specifically, the first GPU determines the sum value between the loss value of the first GPU and the loss values of other GPUs, and uses this sum value as the prediction error.

[0068] When determining the prediction error, the first GPU determines the prediction error based on the classification result of the first sample feature and the classification label of the first sample feature based on the processing result. Before the first GPU determines the prediction error based on the classification result of the first sample feature and the classification label of the first sample feature based on the processing result, the first GPU calculates the similarity between the first sample feature and other sample features in the target training dataset to obtain multiple similarities; the first GPU sorts the multiple similarities in descending order and determines the first N similarities in the sorting result, and sets the classification labels of the sample features corresponding to the first N similarities as the classification label of the first sample feature.

[0069] Taking instance classification as an example, in instance classification, each data sample is regarded as a class. However, in the actual data, there is a high probability that some samples are semantically similar. Considering these samples as negative samples will introduce a lot of noise, which has a negative impact on the training convergence. Therefore, during the training process, by finding the top-K similar classes for each instance class (equivalent to the classes corresponding to the top N similarities mentioned above), and assigning certain positive labels to these similar classes, the label noise can be reduced.

[0070] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from here.

[0071] It should also be noted that the preferred implementation in this embodiment can refer to the relevant description in Embodiment 1, which will not be elaborated here.

[0072] In the embodiment of the present application, a parallel training framework is composed of multiple GPUs, and the feature extraction networks of each GPU are initialized, and the initialized features are input into the corresponding fully connected layer to initialize the fully connected layer. Thus, while improving the training efficiency, the convergence speed of the model can also be enhanced, and further, the technical problem that when the sample size in the dataset is relatively large, the classification model needs to support the same number of classification results through the training process, resulting in a large consumption of GPU computing resources is solved.

[0073] Embodiment 3

[0074] The embodiment of the present application provides a method embodiment of a model training method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from here.

[0075] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 4 The hardware structure block diagram of a computer terminal (or mobile device) for implementing the model training method is shown. As Figure 4As shown, the computer terminal 40 (or mobile device 40) may include one or more processors 402 (illustrated as 402a, 402b, ……, 402n in the figure) (the processor 402 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 404 for storing data, and a transmission module 406 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 4 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 40 may further include more or fewer components than Figure 4 shown therein, or have a different configuration from Figure 4 that shown.

[0076] It should be noted that the above one or more processors 402 and / or other data processing circuits are generally referred to as "data processing circuits" herein. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 40 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0077] The memory 404 may be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the () method in the embodiments of the present application. The processor 402 executes various functional applications and data processing by running the software programs and modules stored in the memory 404, that is, implements the vulnerability detection method of the above-mentioned application program. The memory 404 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory 404 may further include a memory remotely set relative to the processor 402, and these remote memories may be connected to the computer terminal 40 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0078] The transmission module 406 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 40. In one example, the transmission device 406 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 406 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0079] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 40 (or mobile device).

[0080] Under the above operating environment, this application provides a model training method as Figure 5 shown. As Figure 5 shown, the method includes steps S502 - S508. Among them:

[0081] Step S502: Group the training data set to obtain multiple target training data sets;

[0082] During the training process of a machine learning model, a large amount of data is often required. At this time, if all the data in the training data set is input into a single GPU or electronic device for training, however, this will inevitably affect the training efficiency. Therefore, to ensure the training efficiency, the training data set can be split. When grouping, it can be determined according to the sample size in the training data set and the number of GPUs in the GPU cluster. For example, the sample features in the training data set are evenly distributed according to the number of GPUs, that is, the sample size obtained by each GPU is the same.

[0083] Step S504: Input the multiple target training data sets into multiple electronic devices respectively, where each of the multiple electronic devices includes a fully connected layer and a feature extraction network with the same network structure connected to the fully connected layer, and the fully connected layer is obtained by splitting the fully connected layer in the target neural network model;

[0084] The electronic device includes but is not limited to distributed nodes in a distributed network, or GPUs in the same device or different devices.

[0085] Step S506: Initialize the feature extraction networks of the multiple electronic devices, and use the initialized feature extraction networks to extract the sample features in the target training data sets;

[0086] In some embodiments, after the first GPU (i.e., any one of the multiple GPUs) initializes the feature extraction network corresponding to the first GPU, it sends the initialization parameters used for the initialization to other GPUs; other GPUs initialize the feature extraction networks corresponding to them based on the initialization parameters. By using the above method, the consistency of the initialization process can be ensured. Of course, in some alternative embodiments, the feature extraction networks in the above multiple GPUs can also be initialized simultaneously.

[0087] Step S508: Input the sample features into the fully connected layer in multiple electronic devices for processing to classify the sample features.

[0088] Since each node in the fully connected layer is connected to all nodes in the previous layer to integrate the features of the previous layer, splitting the fully connected layer is equivalent to grouping the nodes in the fully connected layer of the target neural network model, and each group of nodes corresponds to a sub-fully connected layer.

[0089] Step S510: Determine multiple prediction errors based on the classification results of multiple electronic devices and the sample labels in the target training dataset.

[0090] Multiple electronic devices can determine the prediction errors in the following way: For any one electronic device, the electronic device determines the loss value of the learning model (i.e., the branch model of the target neural network model) it runs, and obtains the loss values of other electronic devices, where the other electronic devices are the electronic devices other than the above-mentioned any one electronic device among the multiple electronic devices; the electronic device determines the prediction error based on its own loss value and the loss values of other electronic devices.

[0091] Step S512: Update the target neural network model based on multiple prediction errors.

[0092] It should be noted that the preferred implementation manners in this embodiment can refer to the relevant descriptions in Embodiment 1, and will not be elaborated here.

[0093] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0094] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0095] Embodiment 4

[0096] This embodiment provides a model training method, as Figure 6 shown, the method includes:

[0097] Step S602, group the training data set to obtain multiple target training data sets;

[0098] Step S604, input the multiple target training data sets into multiple electronic devices respectively to train the neural network models in the multiple electronic devices. Among them, the neural network models in the multiple electronic devices all include a fully connected layer and a feature extraction network connected to the fully connected layer. The structures of the feature extraction networks in the multiple electronic devices are the same, and the fully connected layer is obtained by splitting the fully connected layer in the target neural network model;

[0099] Step S606, obtain multiple prediction errors obtained by training the multiple neural network models;

[0100] Step S608, determine the prediction error of the target neural network based on the multiple prediction errors, and update the target neural network model based on the prediction error.

[0101] In some embodiments, the process of initializing the feature extraction network can be manifested in the following form, but not limited to this: after the first GPU initializes the feature extraction network corresponding to the first GPU, it sends the initialization parameters used for initialization to other GPUs; other GPUs initialize the feature extraction networks corresponding to other GPUs based on the initialization parameters. By adopting the above method, the consistency of the initialization process can be ensured. Of course, in some alternative embodiments, the feature extraction networks in the above multiple GPUs can also be initialized simultaneously.

[0102] In the embodiments of the present application, a parallel training framework is composed of multiple GPUs. The feature extraction networks of each GPU are initialized, and the initialized features are input into the corresponding fully connected layers to initialize the fully connected layers. Thus, while improving the training efficiency, the convergence speed of the model can also be enhanced, and furthermore, the technical problem that when the number of samples in the dataset is relatively large, the classification model needs to support the same number of classification results through the training process, resulting in a large consumption of computing resources of the GPUs, is solved.

[0103] It should be noted that the preferred implementation manners in this embodiment can refer to the relevant descriptions in Embodiment 1, which will not be elaborated here.

[0104] Embodiment 5

[0105] This embodiment provides a non-volatile storage medium. The non-volatile storage medium includes a stored program. When the program runs, it controls the device where the non-volatile storage medium is located to execute the model training method described above.

[0106] The non-volatile storage medium is used to store program instructions for implementing the following functions: initializing the first feature extraction network of the first GPU, and using the initialized first feature extraction network to extract the first sample features in the target training dataset, where the first GPU is any one of the multiple GPUs; the first GPU inputs the first sample features into the first fully connected layer of the first GPU for processing; the first GPU determines the prediction error of the first GPU based on the processing result; based on the prediction error of the first GPU and the prediction errors of other received GPUs, determining the target prediction error of the target neural network model; and updating the parameters of the target neural network model based on the target prediction error.

[0107] Optionally, the non-volatile storage medium is used to store program instructions for implementing the following functions: after initializing the feature extraction network corresponding to the first GPU, sending the initialization parameters used for initialization to other GPUs; other GPUs initialize the feature extraction networks corresponding to other GPUs based on the initialization parameters.

[0108] Optionally, in this embodiment, the above storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0109] It should be noted that the preferred implementation manners in this embodiment can refer to the relevant descriptions in Embodiment 1, which will not be elaborated here.

[0110] Embodiment 6

[0111] This embodiment also provides a computer terminal, including: a processor; and a memory connected to the processor for providing instructions for the processor to process the following steps: initializing a first feature extraction network of a first GPU, and extracting first sample features in a target training dataset by using the initialized first feature extraction network, where the first GPU is any one of multiple GPUs; inputting the first sample features into a first fully connected layer of the first GPU for processing, where the first fully connected layer is obtained by splitting a second fully connected layer in a target neural network model; determining a prediction error of the first GPU based on the processing result; determining a target prediction error of the target neural network model based on the prediction error of the first GPU and the prediction errors of other received GPUs; and updating parameters of the target neural network model based on the target prediction error.

[0112] In this embodiment, the above computer terminal may also be replaced with a terminal device such as a mobile terminal.

[0113] Optionally, in this embodiment, the above computer terminal may be located in at least one of multiple network devices in a computer network.

[0114] By adopting the solution provided in the embodiment of the present application, the technical problem that when the number of samples in a dataset is relatively large, it is necessary to make a classification model support the same number of classification results through a training process, resulting in a large consumption of computing resources of the GPU, is solved.

[0115] Embodiment 7

[0116] The embodiment of the present application also provides a model training system, as Figure 7 shown. The system includes: a client device 70 and a server 72. The client device 70 is configured to provide a human-computer interaction interface to a target object, and call program instructions in the server 72 through the human-computer interaction interface, such as Figure 6 the steps shown, but not limited thereto: grouping a training dataset to obtain multiple target training datasets; respectively inputting the multiple target training datasets into multiple electronic devices to train neural network models in the multiple electronic devices, where the neural network models in the multiple electronic devices each include a fully connected layer and a feature extraction network connected to the fully connected layer, and the structures of the feature extraction networks in the multiple electronic devices are the same; obtaining multiple prediction errors obtained by training the multiple neural network models; determining a prediction error of a target neural network based on the multiple prediction errors, and updating the target neural network model based on the prediction error.

[0117] Since a large amount of training data is often required for model training and there are also certain requirements for hardware performance, in order to save costs, the services provided by a Software as a Service (SaaS) platform can be used to implement model training. For example, in some embodiments, the above-mentioned client device 10 includes, but is not limited to, the tenant device of the SaaS platform. Correspondingly, the above-mentioned server includes the server of the SaaS platform. When training the second machine learning model, the tenant device can send a request message to the server. The request message is used to request the training of the second machine learning model. At the same time, the request message can also carry information such as the conditions and resources required for this training. The server provides corresponding services to the tenant device according to the request message to execute the above training process. Among them, the specific training process can refer to the relevant descriptions in Embodiments 1-2 and will not be elaborated here.

[0118] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program. The program can be stored in a computer-readable storage medium. The storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0119] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0120] In the above embodiments of the present application, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0121] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the units or modules can be in an electrical or other form.

[0122] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0123] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0124] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drive, read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk or optical disc and other various media that can store program codes.

[0125] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A model training system, comprising: A GPU cluster, which includes multiple GPUs, where: Each GPU in the multiple GPUs has a first feature extraction network and a first fully connected layer, where the network structures of the first feature extraction networks in the multiple GPUs are the same; The first GPU among the multiple GPUs is used to initialize the first feature extraction network of the first GPU, and extract the first sample features in the target training dataset using the initialized first feature extraction network, where the first GPU is any one of the multiple GPUs; input the first sample features into the first fully connected layer of the first GPU for processing; and determine the prediction error of the first GPU based on the processing result; determine the target prediction error of the target neural network model based on the prediction error of the first GPU and the prediction errors of other received GPUs; and update the parameters of the target neural network model based on the target prediction error; After the first GPU initializes the feature extraction network corresponding to the first GPU, it sends the initialization parameters used for the initialization to other GPUs; the other GPUs initialize the feature extraction networks corresponding to the other GPUs based on the initialization parameters.

2. The system according to claim 1, wherein, The first feature extraction network of the first GPU is obtained by copying the first feature extraction network of any one of the multiple GPUs.

3. The system according to claim 1, wherein, The target training dataset is obtained by dividing the training dataset corresponding to the target neural network model according to the number of samples, where the number of samples in each target training dataset is the same.

4. The system according to claim 1, wherein, The multiple GPUs are further used to determine the prediction errors of each GPU in the following manner: The first GPU determines the loss value of the first GPU and obtains the loss values of other GPUs, where the other GPUs are the GPUs other than the first GPU among the multiple GPUs; The first GPU determines the prediction error based on the loss value of the first GPU and the loss values of other GPUs.

5. The system according to claim 4, wherein, The first GPU is further used to determine the sum value between the loss value of the first GPU and the loss values of other GPUs, and use this sum value as the prediction error.

6. The system according to claim 1, wherein, The first GPU is further used to determine the prediction error based on the classification result of the first sample feature and the classification label of the first sample feature based on the processing result.

7. The system according to claim 6, wherein, The first GPU is further used to, before determining the prediction error based on the classification result of the first sample feature and the classification label of the first sample feature based on the processing result, calculate the similarity between the first sample feature and other sample features in the target training dataset to obtain multiple similarities; Sort the multiple similarities in descending order, and determine the top N similarities in the sorting result. Set the classification labels of the sample features corresponding to the top N similarities as the classification labels of the first sample feature.

8. A model training method, which is applied to a GPU cluster, wherein, the GPU cluster includes multiple GPUs, and each GPU in the multiple GPUs has a first feature extraction network and a first fully connected layer. Among them, the network structures of the first feature extraction networks in the multiple GPUs are the same; the method includes: the first GPU in the multiple GPUs initializes the first feature extraction network of the first GPU, and uses the initialized first feature extraction network to extract the first sample features in the target training dataset, where the first GPU is any one of the multiple GPUs; the first GPU inputs the first sample features into the first fully connected layer of the first GPU for processing; the first GPU determines the prediction error of the first GPU based on the processing result; determines the target prediction error of the target neural network model based on the prediction error of the first GPU and the prediction errors of other received GPUs; and updates the parameters of the target neural network model based on the target prediction error; the method further includes: after the first GPU initializes the feature extraction network corresponding to the first GPU, the first GPU sends the initialization parameters used for the initialization to other GPUs; and the other GPUs initialize the feature extraction networks corresponding to the other GPUs based on the initialization parameters.

9. The method according to claim 8, wherein, the first GPU determining the prediction error of the first GPU based on the processing result includes: the first GPU determines the loss value of the first GPU and obtains the loss values of other GPUs, where the other GPUs are the GPUs other than the first GPU in the multiple GPUs; the first GPU determines the prediction error based on the loss value of the first GPU and the loss values of other GPUs.

10. The method according to claim 9, wherein, the first GPU determining the prediction error based on the loss value of the first GPU and the loss values of other GPUs includes: the first GPU determines the sum value between the loss value of the first GPU and the loss values of other GPUs, and takes this sum value as the prediction error.

11. The method according to claim 8, wherein, the first GPU determining the prediction error of the first GPU based on the processing result includes: the first GPU determines the prediction error based on the classification result of the first sample feature and the classification label of the first sample feature according to the processing result.

12. The method according to claim 11, wherein, before the first GPU determines the prediction error based on the classification result of the first sample feature and the classification label of the first sample feature according to the processing result, the method further includes: The first GPU calculates the similarities between the first sample feature and other sample features in the target training dataset, obtaining multiple similarities; The first GPU sorts the multiple similarities in descending order, determines the top N similarities in the sorting result, and sets the classification labels of the sample features corresponding to the top N similarities as the classification label of the first sample feature.

13. A model training method, including: Grouping the training dataset to obtain multiple target training datasets; Inputting the multiple target training datasets into multiple electronic devices respectively, where each of the multiple electronic devices includes a fully connected layer and a feature extraction network with the same network structure connected to the fully connected layer; Initializing the feature extraction networks of the multiple electronic devices, and using the initialized feature extraction networks to extract the sample features in the target training datasets; Inputting the sample features into the fully connected layers in the multiple electronic devices for processing to classify the sample features; Determining multiple prediction errors based on the classification results of the multiple electronic devices and the sample labels in the target training datasets; Updating the target neural network model based on the multiple prediction errors; The method further includes: after the first electronic device initializes the feature extraction network corresponding to the first electronic device, sending the initialization parameters used for the initialization to other electronic devices; the other electronic devices initialize the feature extraction networks corresponding to the other electronic devices based on the initialization parameters, where the first electronic device is any one of the multiple electronic devices, and the other electronic devices are the electronic devices other than the first electronic device among the multiple electronic devices.

14. A model training method, including: Grouping the training dataset to obtain multiple target training datasets; Inputting the multiple target training datasets into multiple electronic devices respectively to train the neural network models in the multiple electronic devices, where the neural network models in the multiple electronic devices each include a fully connected layer and a feature extraction network connected to the fully connected layer, and the structures of the feature extraction networks in the multiple electronic devices are the same; Obtaining multiple prediction errors obtained by training the multiple neural network models; Determining the prediction error of the target neural network based on the multiple prediction errors, and updating the target neural network model based on the prediction error; The method further includes: after the first electronic device initializes the feature extraction network corresponding to the first electronic device, sending the initialization parameters used for the initialization to other electronic devices; the other electronic devices initialize the feature extraction networks corresponding to the other electronic devices based on the initialization parameters, where the first electronic device is any one of the multiple electronic devices, and the other electronic devices are the electronic devices other than the first electronic device among the multiple electronic devices.

15. A non-volatile storage medium, wherein, The non-volatile storage medium includes a stored program, wherein when the program runs, it controls the device where the non-volatile storage medium is located to execute the model training method described in any one of claims 8 to 14.

16. A computer terminal, wherein, comprises: a processor; and a memory, connected to the processor, for providing instructions for the processor to perform the following processing steps: initializing a first feature extraction network of a first GPU, and using the initialized first feature extraction network to extract first sample features in a target training dataset, wherein the first GPU is any one of a plurality of GPUs; inputting the first sample features into a first fully connected layer of the first GPU for processing; determining a prediction error of the first GPU based on the processing result; determining a target prediction error of a target neural network model based on the prediction error of the first GPU and the prediction errors of other received GPUs; and updating parameters of the target neural network model based on the target prediction error; after the first GPU initializes the feature extraction network corresponding to the first GPU, sending the initialization parameters used for the initialization to other GPUs; and the other GPUs initializing the feature extraction networks corresponding to the other GPUs based on the initialization parameters.

17. A model training system, comprises: a client device and a server, wherein: the client device is configured to provide a human-computer interaction interface to a target object, and call program instructions in the server through the human-computer interaction interface to perform the following steps: grouping a training dataset to obtain a plurality of target training datasets; respectively inputting the plurality of target training datasets into a plurality of electronic devices to train neural network models in the plurality of electronic devices, wherein the neural network models in the plurality of electronic devices each include a fully connected layer and a feature extraction network connected to the fully connected layer, and the structures of the feature extraction networks in the plurality of electronic devices are the same; obtaining a plurality of prediction errors obtained by training the plurality of neural network models; determining a prediction error of a target neural network based on the plurality of prediction errors, and updating the target neural network model based on the prediction error; after a first electronic device initializes the feature extraction network corresponding to the first electronic device, sending the initialization parameters used for the initialization to other electronic devices; and the other electronic devices initializing the feature extraction networks corresponding to the other electronic devices based on the initialization parameters, wherein the first electronic device is any one of the plurality of electronic devices, and the other electronic devices are the electronic devices other than the first electronic device among the plurality of electronic devices.

Citation Information

Patent Citations

  • Convolutional neural network model synchronous training method, cluster and readable storage medium

    CN110705705A

  • Model training method, metal fracture analysis method based on deep learning and application

    CN111209964A

  • Classification model training method, image processing method and device

    CN111242222A