A training method, device, equipment and storage medium for a neural network model
By calculating the L2 norm of the eigenvector of the training sample and configuring larger weights for the difficult samples, the problem of overfitting the deep learning model to simple samples during the training process is solved, and the generalization and robustness of the model is improved.
Patent Information
- Application Number
- CN202110172561.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-02-08
AI Technical Summary
Deep learning models are prone to overfitting to simple samples during training, resulting in poor generalization in practical applications, especially poor performance when dealing with difficult samples.
By calculating the L2 norm of the eigenvector of the training sample, it determines its difficulty level in the training data, and configures a larger cross entropy loss value weight for the difficult sample, so that the model pays more attention to the learning of the difficult sample during the training process.
It improves the generalization of neural network models in practical applications, ensures that the model can better adapt to and predict complex samples, and enhances the robustness and applicability of the model.
Smart Images

Figure CN114943330B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of machine learning technology, and in particular to a training method, apparatus, device and storage medium for a neural network model. Background Art
[0002] At present, technologies based on deep learning have achieved remarkable application results, such as e-commerce recommendation systems and mobile phone face unlocking. Deep learning algorithms are data-driven, that is, they rely on a large amount of training data and often require corresponding label information. For example, to train a face recognition model, you need to first collect a large number of face pictures, then label each face picture with an ID, and then use the labeled face pictures as training data to train the face recognition model. Similar to human learning, different samples may have different learning difficulties for deep models.
[0003] Generally speaking, most of the samples in the training set are simple samples, and a certain proportion of the samples are difficult samples. Taking the training of the classification face recognition model as an example, the model input unit is a single sample. In this case, a simple sample can be a clear frontal face, and a difficult sample can be a side face or a face with occlusion or blur. Taking the training of the metric face recognition model as an example, the model input unit is a combination sample, such as a face pair. In this case, a simple sample can be a negative sample pair with two different IDs and obvious facial differences, or a positive sample pair with two almost identical photos of the same ID. A difficult sample can be a negative sample pair with two different IDs but similar facial appearances, or a positive sample pair with two very different appearances of the same ID.
[0004] During the training process, a situation often occurs where the model performs well on simple samples and the corresponding sample loss is small, but does not learn much on difficult samples and the corresponding sample loss is large. However, due to the large number of simple samples, the overall training loss may decrease. At this time, if the model is considered to have been trained and converged, the generalization of the model will be poor in actual application. Summary of the invention
[0005] The present invention provides a training method, device, equipment and storage medium for a neural network model, so that the model pays more attention to the learning of difficult samples during the training process, thereby improving the generalization of the neural network model in practical applications.
[0006] In a first aspect, an embodiment of the present invention provides a method for training a neural network model, comprising:
[0007] Acquire an i-th batch of data for training a neural network model, wherein the i-th batch of data includes a plurality of training samples, and the training samples are annotated with labels representing categories of the training samples;
[0008] Inputting the plurality of training samples of the i-th batch of data into the neural network model to be trained for processing respectively, and obtaining a first probability vector for predicting that the training samples belong to each category;
[0009] Calculating a cross entropy loss value of the training sample based on the first probability vector and the label of the training sample;
[0010] Calculating the L2 norm of the feature vector of the training sample, where the feature vector is extracted by the neural network model and is used for mapping to obtain the first probability vector;
[0011] Determining a weight of a cross entropy loss value of the training sample based on an L2 norm of the feature vector;
[0012] Calculate the sum of the cross entropy loss values of the training samples in the i-th batch of data multiplied by the weighted loss values to obtain a first total loss value;
[0013] Update parameters of the neural network model based on the first total loss value.
[0014] In a second aspect, an embodiment of the present invention further provides a training device for a neural network model, comprising:
[0015] A data acquisition module, used to acquire an i-th batch of data for training a neural network model, wherein the i-th batch of data includes a plurality of training samples, and the training samples are annotated with labels indicating the categories of the training samples;
[0016] A first probability value acquisition module, used for inputting the plurality of training samples of the i-th batch of data into the neural network model to be trained for processing, and obtaining a first probability vector for predicting that the training samples belong to each category;
[0017] A cross entropy loss value calculation module, used to calculate the cross entropy loss value of the training sample based on the first probability vector and the label of the training sample;
[0018] An L2 norm calculation module, used to calculate the L2 norm of the feature vector of the training sample, where the feature vector is extracted by the neural network model and is used to map to obtain the probability vector;
[0019] A weight determination module, used to determine the weight of the cross entropy loss value of the training sample based on the L2 norm of the feature vector;
[0020] A first total loss value calculation module, used to calculate the sum of the cross entropy loss values of each training sample in the i-th batch of data multiplied by the weighted loss value, to obtain a first total loss value;
[0021] A parameter updating module is used to update the parameters of the neural network model based on the first total loss value.
[0022] In a third aspect, an embodiment of the present invention further provides a computer device, including:
[0023] one or more processors;
[0024] A storage device for storing one or more programs;
[0025] When the one or more programs are executed by the one or more processors, the one or more processors implement the training method of the neural network model provided in the first aspect of the present invention.
[0026] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the training method of the neural network model provided in the first aspect of the present invention.
[0027] The training method of the neural network model provided by the embodiment of the present invention, after obtaining the i-th batch of data for training the neural network model, inputs multiple training samples of the i-th batch of data into the neural network model to be trained for processing, obtains a first probability vector for predicting that the training sample belongs to each category, calculates the cross entropy loss value of the training sample based on the first probability vector and the label of the training sample, calculates the L2 norm of the feature vector of the training sample, determines the weight of the cross entropy loss value of the training sample based on the L2 norm of the feature vector, calculates the sum of the loss values after multiplying the cross entropy loss value of each training sample in the i-th batch of data by the weight, obtains a first total loss value, and updates the parameters of the neural network model based on the first total loss value. The embodiment of the present invention mines difficult samples in the training data based on the L2 norm of the feature vector of the training sample, configures corresponding weights for the cross entropy loss value of the difficult sample, allows the model to pay more attention to the learning of difficult samples during the training process, and thus improves the generalization of the neural network model in the actual application process. In addition, mining difficult samples in the training data based on the L2 norm of the feature vector of the training sample does not require calculating the similarity and sorting operations between the training samples, saves computing resources, and improves training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 A flowchart of a training method for a neural network model provided in Embodiment 1 of the present invention;
[0029] Figure 2A A flowchart of a training method for a neural network model provided in Embodiment 2 of the present invention;
[0030] Figure 2B This is a distribution diagram of the L2 norm of feature vectors based on public data set statistics;
[0031] Figure 3 A classification prediction method provided by Embodiment 3 of the present invention;
[0032] Figure 4 A schematic diagram of the structure of a training device for a neural network model provided in Embodiment 4 of the present invention;
[0033] Figure 5 A schematic diagram of the structure of a classification prediction device provided by an embodiment of the present invention;
[0034] Figure 6 A schematic diagram of the structure of a computer device provided in Embodiment 6 of the present invention. DETAILED DESCRIPTION
[0035] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0036] Embodiment 1
[0037] Figure 1 This is a flowchart of a training method for a neural network model provided in the first embodiment of the present invention. This embodiment can be applied to mining difficult samples in a training set, so that the model can focus on the learning of difficult samples during the training process. The method can be executed by a training device for a neural network model provided in an embodiment of the present invention. The device can be implemented in software and / or hardware, and is usually configured in a computer device, such as Figure 1 As shown, the method specifically comprises the following steps:
[0038] S101. Obtain an i-th batch of data for training a neural network model, where the i-th batch of data includes a plurality of training samples, and the training samples are annotated with labels representing categories of the training samples.
[0039] In the embodiment of the present invention, the neural network model to be trained is used to implement data classification and recognition tasks, and the embodiment of the present invention does not limit the network structure and application scenarios of the neural network model to be trained. Exemplarily, in one embodiment of the present invention, the neural network model is used for face recognition, and the ID of the person to whom the face image belongs is predicted based on the input face image. Taking the classification deep face recognition model as an example, the neural network model can be divided into a backbone network (backbone) and a head network (head). The backbone network processes the input face image and outputs the feature vector of the face image. The head network processes the input feature vector and outputs the probability value used to predict that the face image belongs to each ID.
[0040] The data set used to train the neural network model includes training samples of multiple categories, each category has multiple training samples, a batch of data includes multiple training samples, and each training sample is annotated with a label representing the category of the training sample. For the training of the classification model, the model input unit is a single training sample, and for the training of the metric model, the model input unit is a combined sample, such as a face pair, which is not limited in the embodiment of the present invention.
[0041] Exemplarily, in one embodiment of the present invention, taking the classification deep face recognition model as an example, the data set used to train the neural network model includes face pictures of 10,000 people's IDs, a batch has 64 face pictures, and each face picture is annotated with a label indicating the ID of the person described in the face picture.
[0042] S102: Input multiple training samples of the i-th batch of data into the neural network model to be trained for processing, and obtain a first probability vector for predicting that the training samples belong to each category.
[0043] Multiple training samples of the i-th batch of data are input into the neural network model to be trained for processing. The neural network model can perform operations such as convolution, pooling, and feature fusion on the input training samples, extract feature vectors used to characterize the features of the training samples, and output logits vectors used to predict the categories of the training samples through conversion of the fully connected network layer. The logits vectors are normalized by the softmax function to obtain the first probability vectors used to predict the training samples belonging to each category. Each position in the first probability vector corresponds to each category, and the category corresponding to the position of the maximum element value in the first probability vector is the category to which the training sample belongs predicted by the model.
[0044] Exemplarily, as described above, the backbone network processes the input face image to obtain a feature vector of the face image, and exemplary, the feature vector is 512-dimensional. The head network converts the 512-dimensional feature vector into a 10,000-dimensional logits vector through a 512×10,000 fully connected layer. The softmax function normalizes the logits vector to obtain a 10,000-dimensional first probability vector. The 10,000 positions in the first probability vector correspond to the IDs of 10,000 people, and the ID corresponding to the position of the maximum element value in the first probability vector is used as the predicted ID of the training sample predicted by the model.
[0045] S103: Calculate a cross entropy loss value of the training sample based on the first probability vector and the label of the training sample.
[0046] Specifically, the cross entropy loss function is used to calculate the cross entropy loss value of the training sample. For a binary classification problem, the model finally predicts only two categories, and the probability we predict for each category is p and 1-p. At this time, the cross entropy loss value calculation formula is as follows:
[0047] L=-[y·log(p)+(1-y)·log(1-p)]
[0048] Among them, y is the label of the training sample, the label of the positive sample is 1, and the label of the negative sample is 0. P is the probability value of the training sample predicted by the model to be a positive sample.
[0049] For multi-classification problems, the cross entropy loss value is calculated as follows:
[0050]
[0051] Among them, K is the number of categories, y c is an indicator variable (0 or 1), if the category is the same as the category represented by the label of the training sample, it is 1, otherwise it is 0, p c The probability value of the training sample predicted by the model belongs to category c.
[0052] S104, calculating the L2 norm of the feature vector of the training sample, where the feature vector is extracted by the neural network model and is used for mapping to obtain a first probability vector.
[0053] As mentioned above, the feature vector is extracted by the neural network model and is used to characterize the characteristics of the training sample. It can then be converted through a fully connected network layer to output a logits vector. The logits vector is normalized by the softmax function to obtain the first probability vector used to predict whether the training sample belongs to each category.
[0054] Calculate the L2 norm of the feature vector of the training sample. The L2 norm of the feature vector refers to the sum of the squares of the elements in the feature vector, and then find the square root. The calculation formula is as follows:
[0055]
[0056] Among them, ‖x‖2 is the L2 norm of the feature vector x, x i is the i-th element in the feature vector x.
[0057] S105. Determine the weight of the cross entropy loss value of the training sample based on the L2 norm of the feature vector.
[0058] The weight of the cross entropy loss value of the training sample is determined based on the L2 norm of the feature vector. Specifically, the L2 norm of the feature vector reflects the difficulty of the training sample. Based on this, different weights can be configured for the cross entropy loss values of the training samples corresponding to different L2 norms. For example, the smaller the L2 norm of the feature vector, the more difficult it is for the model to learn the training sample. At this time, a relatively large weight can be configured for the cross entropy loss value of the training sample, so that the model pays more attention to the learning of difficult samples during the training process.
[0059] S106. Calculate the sum of the cross entropy loss values of each training sample in the i-th batch of data multiplied by the weighted loss value to obtain a first total loss value.
[0060] After determining the weight of the cross entropy loss value of the training sample, multiply the weight by the cross entropy loss value of the corresponding training sample to obtain a weighted loss value, and then calculate the sum of the weighted loss values of each training sample in the i-th batch of data to obtain the first total loss value.
[0061] S107. Update the parameters of the neural network model based on the first total loss value.
[0062] After obtaining the first total loss value, the parameters of the neural network model are updated based on the first total loss value. Exemplarily, in some embodiments of the present invention, the first total loss value is compared with the first preset value, and when the first total loss value is greater than the first preset value, the parameters of the neural network model are updated, and the step of obtaining training data is returned to obtain the i+1th batch of data, and the neural network model is trained again, and this cycle is repeated until the first total loss value is less than or equal to the first preset value, and it is determined that the training of the neural network model is completed.
[0063] In the prior art, it is usually necessary to calculate the similarity between the feature vector of each training sample in a batch and the feature vector of the remaining training samples, and to mine the difficult root based on the similarity. For example, if the similarity between the feature vector of a training sample and the feature vectors of all other training samples (or most of the remaining training samples) is less than a set value, then the training sample is determined to be a difficult sample. Since this method needs to calculate the similarity between every two samples, and also sort them to find the maximum or minimum value of the similarity, it takes up additional computing resources and prolongs the training time.
[0064] The training method of the neural network model provided by the embodiment of the present invention, after obtaining the i-th batch of data for training the neural network model, inputs multiple training samples of the i-th batch of data into the neural network model to be trained for processing, obtains a first probability vector for predicting that the training sample belongs to each category, calculates the cross entropy loss value of the training sample based on the first probability vector and the label of the training sample, calculates the L2 norm of the feature vector of the training sample, determines the weight of the cross entropy loss value of the training sample based on the L2 norm of the feature vector, calculates the sum of the loss values after multiplying the cross entropy loss value of each training sample in the i-th batch of data by the weight, obtains a first total loss value, and updates the parameters of the neural network model based on the first total loss value. The embodiment of the present invention mines difficult samples in the training data based on the L2 norm of the feature vector of the training sample, configures corresponding weights for the cross entropy loss value of the difficult sample, allows the model to pay more attention to the learning of difficult samples during the training process, and thus improves the generalization of the neural network model in the actual application process. In addition, mining difficult samples in the training data based on the L2 norm of the feature vector of the training sample does not require calculating the similarity and sorting operations between the training samples, saves computing resources, and improves training efficiency.
[0065] Embodiment 2
[0066] Figure 2A This is a flowchart of a training method for a neural network model provided in Embodiment 2 of the present invention. This embodiment is refined on the basis of Embodiment 1 above, and describes in detail the specific process of determining the weight of the cross entropy loss value of the training sample, such as Figure 2A As shown, the method specifically comprises the following steps:
[0067] S201. Obtain multiple batches of data for training a neural network model.
[0068] In the embodiment of the present invention, in the early stage of model training, only ordinary training is generally performed. Otherwise, if the model training process is focused on difficult samples from the beginning, it may cause the model training process to fail to converge normally. Usually, after ordinary training has been carried out to a certain extent, so that the neural network model meets the preset conditions, the difficult sample mining is started, so that the model that has learned well on simple samples begins to tend to learn difficult samples.
[0069] S202. Perform initial training on the neural network model using multiple batches of data until the neural network model meets preset conditions.
[0070] In an embodiment of the present invention, the neural network model is initially trained using the data of the multiple batches until the neural network model meets the preset conditions. The preset conditions can be set according to the actual situation, such as setting a loss threshold to monitor whether the cross entropy loss value of the model drops below the threshold, or whether it drops below the threshold and undergoes iterative training of a set batch; such as setting an accuracy threshold to monitor whether the accuracy of the model on the validation set exceeds the threshold, or whether the accuracy exceeds the threshold and undergoes iterative training of a set batch; for another example, combining at least two of the above multiple determination methods to monitor whether the model satisfies both conditions.
[0071] For example, in a specific embodiment of the present invention, a loss threshold is set, and the cross entropy loss value of the monitoring model is monitored to see whether it drops below the threshold, thereby determining whether the neural network model meets the preset conditions. Specifically, the preset condition determination process is as follows:
[0072] 1. Multiple training samples of the j-th batch of data are respectively input into the neural network model to be trained for processing to obtain a second probability vector for predicting that the training samples belong to each category.
[0073] As described in the aforementioned embodiment, the neural network model obtains a batch of data (referred to herein as the jth batch of data to distinguish it from the aforementioned embodiment), performs convolution, pooling, feature fusion and other operations on the training samples in the input batch of data, extracts the feature vector used to characterize the features of the training samples, and outputs the logits vector used to predict the category of the training samples through conversion of the fully connected network layer. The logits vector is normalized through the softmax function to obtain a second probability vector for predicting that the training samples belong to each category.
[0074] 2. Calculate the cross entropy loss value of the training sample based on the second probability vector and the label of the training sample.
[0075] The cross entropy loss value of the training sample is calculated based on the second probability vector and the label of the training sample. Specifically, the specific process of calculating the cross entropy loss value of the training sample has been described in detail in the aforementioned embodiment, and the embodiment of the present invention will not be repeated here.
[0076] 3. Calculate the sum of the cross entropy loss values of each training sample in the jth batch of data to obtain the second total loss value.
[0077] Calculate the sum of the cross entropy loss values of each training sample in the jth batch of data to obtain the total loss value of the batch of data, which is called the second total loss value. Specifically, the calculation formula for the second total loss value is as follows:
[0078]
[0079] Among them, m is the number of training samples in a batch, y kc is an indicator variable (0 or 1), if the category is the same as the category represented by the label of the training sample, it is 1, otherwise it is 0, p kc is the probability value predicted by the model that the kth training sample belongs to category c, and K is the number of categories. Exemplarily, in the embodiment of the present invention, m=64, K=10000.
[0080] 4. Determine whether the second total loss is less than or equal to a preset value.
[0081] The second total loss value is compared with a preset second preset value to determine whether the second total loss is less than or equal to the second preset value.
[0082] 5. When the second total loss value is greater than the preset value, update the parameters of the neural network model, and input multiple training samples of the j+1th batch of data into the neural network model to be trained for processing.
[0083] Specifically, when the second total loss value is greater than the second preset value, the parameters of the neural network model are updated, and multiple training samples of the j+1th batch of data are respectively input into the neural network model to be trained for processing.
[0084] 6. When the second total loss value is less than or equal to the preset value, it is determined that the neural network model meets the preset conditions.
[0085] As described above, the above steps 1-5 are executed in a loop until the second total loss value is less than or equal to the preset value, and it is determined that the neural network model meets the preset conditions.
[0086] S203, obtaining the i-th batch of data for training the neural network model, where the i-th batch of data includes a plurality of training samples, and the training samples are annotated with labels representing categories of the training samples.
[0087] After the normal training is carried out to a certain extent, so that the neural network model meets the preset conditions, at this time, the i-th batch of data used to train the neural network model is obtained, and the mining training of difficult samples is started.
[0088] S204: Input multiple training samples of the i-th batch of data into the neural network model to be trained for processing, and obtain a first probability vector for predicting that the training samples belong to each category.
[0089] Multiple training samples of the i-th batch of data are respectively input into the neural network model to be trained for processing, and the feature vectors used to characterize the characteristics of the training samples are extracted. After conversion through the fully connected network layer, the logits vectors used to predict the categories of the training samples are output. The logits vectors are normalized through the softmax function to obtain the first probability vectors used to predict that the training samples belong to each category.
[0090] Exemplarily, in one embodiment of the present invention, a neural network model is used for face recognition, and based on an input face picture, the ID of the person to whom the face picture belongs is predicted. Taking the classification deep face recognition model as an example, the neural network model can be divided into a backbone network (backbone) and a head network (head). The backbone network processes the input face picture and outputs a feature vector of the face picture. The head network processes the input feature vector and outputs a probability value for predicting that the face picture belongs to each ID. The data set used to train the neural network model includes face pictures of 10,000 people's IDs. A batch has 64 face pictures, and each face picture is annotated with a label indicating the ID of the person described in the face picture. Among them, the training sample (face picture) is denoted as x k In the embodiment of the present invention, the backbone network is denoted as f θ (), the head network is where θ and Represent the parameters of the backbone network and the head network respectively, then the training sample x k The eigenvector of θ (x k ).
[0091] In some embodiments of the present invention, in order to speed up the convergence of the model, the feature vector can be standardized to obtain a standardized feature vector, which is denoted as Norm(f θ (x k )). Eigenvector normalization means that the eigenvector f θ (x k ) is divided by the L2 norm of the feature vector. Therefore, the first probability vector can be written as
[0092] S205: Calculate the cross entropy loss value of the training sample based on the first probability vector and the label of the training sample.
[0093] Specifically, the cross entropy loss function is used to calculate the cross entropy loss value of the training sample. Exemplarily, the data set used to train the neural network model includes face pictures of 10,000 people's IDs, that is, the neural network model in the embodiment of the present invention is applied to multi-classification problems. In the multi-classification problem, the calculation process of the cross entropy loss value of the training sample has been described in detail in the aforementioned embodiment, and the embodiment of the present invention will not be repeated here.
[0094] For example, the cross entropy loss function of the training samples in multi-classification is L ce (), then the cross entropy loss function value of the training sample is recorded as
[0095] S206. Calculate the L2 norm of the feature vector of the training sample, where the feature vector is extracted by the neural network model and is used for mapping to obtain a first probability vector.
[0096] Calculate the L2 norm of the feature vector of the training sample. The L2 norm of the feature vector refers to the sum of the squares of the elements in the feature vector, and then find the square root. For example, the feature vector f θ (x k ) is denoted by ‖f θ (x k )‖2.
[0097] S207 , storing the L2 norms corresponding to the multiple training samples of the i-th batch of data into the data queue.
[0098] In an embodiment of the present invention, a data queue Q of size m×p is introduced to store the L2 norm of the feature vector. The distribution of the L2 norm in the data queue Q approximates the distribution of the L2 norm of the feature vector of the training samples of the entire data set. Among them, the data queue Q has m×p storage bits, m is the number of training samples in each batch of data, and p is an integer greater than 1. The value of p should not be too small to avoid a large deviation between the distribution of the L2 norm in the data queue Q and the distribution of the L2 norm of the entire data set. On the other hand, the value of p should not be too large, because the model is constantly optimized and the learned features are gradually getting better. Storing the too early L2 norm may have a negative impact on the subsequent difficult sample mining. For reference, on an MS1M of size 580,000, a queue Q of size 6400 (m=64, p=100) is basically sufficient. Specifically, the process of storing the L2 norm in the data queue is as follows:
[0099] First, determine whether there are vacant storage bits in the data queue. If so, determine the target storage bit to be stored from the vacant storage bits; if not, erase the L2 norms corresponding to multiple training samples of the ipth batch of data to obtain the target storage bit, and then store the L2 norms corresponding to multiple training samples of the i-th batch of data in the target storage bit. That is, the data queue is initially empty and no data is stored. During the first iterative training (including the ordinary training described above), the m L2 norms corresponding to the first batch of data are stored in the data queue Q. In the subsequent iterative training process, when the data queue Q is not full, the m L2 norm values corresponding to the new batch of data are directly stored in the data queue Q. When the data queue Q is full, the first m values in the data queue Q are erased first, and then the m L2 norms corresponding to the new batch of data are stored.
[0100] S208. Determine the weight of the cross entropy loss value of the training sample based on the mean and standard deviation of all L2 norms in the data queue and the L2 norm of the feature vector of the training sample.
[0101] In an embodiment of the present invention, each time the training is iterated, the data in the data queue Q is updated (the m L2 norms corresponding to the batch of data are stored in the data queue Q), and then the weight of the cross entropy loss value of the training sample is determined based on the mean and standard deviation of all L2 norms in the updated data queue Q and the L2 norm of the feature vector of the training sample.
[0102] For the training samples in the current batch, there are multiple methods to determine the weight of the cross entropy loss value of the training samples. Exemplarily, in one embodiment of the present invention, by setting multiple partitions for representing the difficulty of the training samples, the weight is determined according to the partition to which the L2 norm of the training samples belongs. The specific process is as follows:
[0103] 1. Calculate the mean and standard deviation of all L2 norms in the data queue.
[0104] Specifically, the mean μ of all L2 norms in the updated data queue Q and the standard deviation σ of all L2 norms are calculated.
[0105] 2. Based on the mean and standard deviation, multiple partitions are determined to represent the difficulty of the training samples, and each partition has a corresponding weight of the cross entropy loss value.
[0106] Figure 2B This is a distribution diagram of the L2 norm of the feature vector based on public data set statistics, such as Figure 2BAs shown in the figure, taking the face recognition model based on IR-ResNet50 as the backbone network as an example, on 6400 random samples of the public face dataset MS1M, the L2 norm of the statistical feature vector (feature norm in the figure) basically conforms to the Gaussian distribution. The two important parameters of the Gaussian distribution are the mean μ and the standard deviation σ. Therefore, all L2 norms in the data queue Q will also be Gaussian distributed. Based on this, multiple partitions used to represent the difficulty of the training samples can be determined based on the mean μ and standard deviation σ of all L2 norms in the updated data queue Q, and each partition has a corresponding weight of the cross entropy loss value.
[0107] Exemplarily, in a specific embodiment of the present invention, based on the updated L2 norm mean μ and standard deviation σ in the data queue Q, all data in the data queue Q are divided into 8 partitions, namely L4=(-∞,μ-3σ], L3=(μ-3σ,μ-2σ], L2=(μ-2σ,μ-σ], L1=(μ-σ,μ], R1=(μ,μ+σ], R2=(μ+σ,μ+2σ], R3=(μ+2σ,μ+3σ] and R4=(μ+3σ,∞), and the learning difficulty of the training samples corresponding to these partitions shows a downward trend. The difficult samples are mainly concentrated in the two areas L4 and L3 with the smallest L2 norm. In the embodiment of the present invention, the difficult samples are mainly face pictures showing only half of the face, large-angle side face, dark light situation, etc. The samples corresponding to the two areas R3 and R4 are very easy samples. In the embodiment of the present invention, they are mainly complete front faces with uniform lighting. The samples located in the middle partition are ordinary samples.
[0108] After determining the partitions, set the weights of each partition. Specifically, you can introduce prior information based on the actual situation of the data set. For example, you can set the weights of each partition to {1.2, 1.2, 1.1, 1.1, 1.0, 1.0, 1.0, 1.0} in sequence, indicating that you want to increase the model's attention to samples in the first four partitions, with a focus on the first two partitions.
[0109] It should be noted that in the above embodiment, 8 partitions are used as an example to illustrate the embodiment of the present invention. In other embodiments of the present invention, the number of partitions can be set according to the actual situation of the data set, and the embodiment of the present invention is not limited here.
[0110] 3. Determine the partition to which the L2 norm of the feature vector of the training sample belongs.
[0111] Determine the partition into which the L2 norm of the feature vector of the training sample falls, and then determine the weight of the cross entropy loss value of the training sample.
[0112] 4. The weight corresponding to the partition to which the L2 norm of the feature vector of the training sample belongs is used as the weight of the cross entropy loss value of the training sample.
[0113] For example, for the training sample x k , if its eigenvector is denoted by f θ (x k )’s L2 norm ‖f θ (x k )‖2 is located in partition L3, then the training sample x is determined k The weight of the cross entropy loss value is 1.2.
[0114] In another embodiment of the present invention, an adaptive weight function can be set based on prior knowledge, and the weight corresponding to the training sample can be obtained by simply inputting the L2 norm of the feature vector of the training sample into the adaptive weight function. The specific process is as follows:
[0115] 1. Construct a Gaussian distribution function of the L2 norm based on the mean and standard deviation of the L2 norm in the data queue.
[0116] As mentioned above, taking the face recognition model based on IR-ResNet50 as the backbone network as an example, on the random 6400 samples of the public face dataset MS1M, the L2 norm of the statistical feature vector basically conforms to the Gaussian distribution. Therefore, the Gaussian distribution function can be used to approximate the distribution of the L2 norm in the data queue. The Gaussian distribution function is as follows:
[0117]
[0118] Among them, z is the L2 norm of the feature vector of the training sample, μ is the mean of all L2 norms in the data queue, and σ is the standard deviation of all L2 norms in the data queue.
[0119] 2. Construct an adaptive weight function based on the Gaussian distribution function and the standard deviation of the L2 norm in the data queue.
[0120] Exemplarily, in an embodiment of the present invention, the adaptive weight function is as follows:
[0121]
[0122] Among them, z is the L2 norm of the feature vector of the training sample, G(z) is the Gaussian distribution function, σ is the standard deviation of all L2 norms in the data queue, and T is the preset L2 norm threshold. For example, let T = μ, which means that the loss weight of the training samples with L2 norm lower than μ will be adaptively adjusted, without paying attention to the training samples with L2 norm lower than μ, thereby saving computing resources and improving training efficiency.
[0123] 3. Input the L2 norm of the feature vector of the training sample into the adaptive weight function to obtain the weight of the cross entropy loss value of the training sample.
[0124] Specifically, the L2 norm of the feature vector of the training sample is input into the adaptive weight function to obtain the weight of the cross entropy loss value of the training sample.
[0125] In the embodiment of the present invention, during iterative training, the data in the data queue Q is updated each time during iterative training (the m L2 norms corresponding to the batch of data are stored in the data queue Q), and then the weight of the cross entropy loss value of the training sample is determined based on the mean and standard deviation of all L2 norms in the updated data queue Q, and the L2 norm of the feature vector of the training sample. The distribution of the L2 norm in the data queue Q approximates the distribution of the L2 norm of the feature vector of all training samples in the entire data set, which saves computing costs and improves training efficiency. In addition, the data queue Q is dynamically updated to reduce approximation errors and improve the accuracy of model prediction.
[0126] S209, calculating the sum of the cross entropy loss values of each training sample in the i-th batch of data multiplied by the weighted loss value to obtain a first total loss value.
[0127] Specifically, after determining the weight of the cross entropy loss value of the training sample, the weight is multiplied by the cross entropy loss value of the corresponding training sample to obtain a weighted loss value, and then the sum of the weighted loss values of each training sample in the i-th batch of data is calculated to obtain a first total loss value.
[0128] Exemplarily, in a specific embodiment of the present invention, as described above, the cross entropy loss function value of the training sample is recorded as The weighted loss value is Then, the sum of the weighted loss values of each training sample in the i-th batch of data is calculated to obtain the first total loss value L1.
[0129] S210. Update parameters of the neural network model based on the first total loss value.
[0130] After obtaining the first total loss value, the parameters of the neural network model are updated based on the first total loss value. Exemplarily, in some embodiments of the present invention, the first total loss value is compared with the first preset value, and when the first total loss value is greater than the first preset value, the parameters of the neural network model are updated, and the step of obtaining training data is returned to obtain the i+1th batch of data, and the neural network model is trained again, and this cycle is repeated until the first total loss value is less than or equal to the first preset value, and it is determined that the training of the neural network model is completed.
[0131] The training method of the neural network model provided by the embodiment of the present invention mines the difficult samples in the training data based on the L2 norm of the feature vector of the training sample, and configures the corresponding weight for the cross entropy loss value of the difficult sample, so that the model pays more attention to the learning of difficult samples during the training process, thereby improving the generalization of the neural network model in the actual application process. In addition, the difficult samples in the training data are mined based on the L2 norm of the feature vector of the training sample, and there is no need to calculate the similarity and sorting operations between the training samples, which saves computing resources and improves the training efficiency. In addition, a data queue is introduced, and the distribution of the L2 norm in the data queue approximates the distribution of the L2 norm of the feature vectors of all training samples in the entire data set, which saves computing costs and improves training efficiency. In addition, the data queue is dynamically updated during iterative training to reduce the approximation error and improve the accuracy of model prediction.
[0132] In a certain sense, noise samples can be considered as extremely difficult samples, but noise samples should not be learned by the model. On a data set known to have a certain amount of noise, we can still use the aforementioned difficult sample mining method to remove the impact of noise samples on model training, thereby improving the prediction accuracy of the model. Exemplarily, for an embodiment in which weights are determined by setting multiple partitions to represent the difficulty of training samples and the partition to which the L2 norm of the training sample belongs, the weight of the smallest partition can be set to 0.1 or even smaller. For example, the weights of each partition are set to {0.1, 1.2, 1.1, 1.1, 1.0, 1.0, 1.0}, where 0.1 is the weight given to the noise sample, so that the model hardly learns the noise sample. It is not directly set to 0.0 here because the noise sample may play a certain regularization role to prevent the model from overfitting. For an embodiment in which weights are determined by an adaptive weight function, the adaptive weight function can be set to a three-segment formula, as shown below:
[0133]
[0134] Among them, training samples with L2 norm less than or equal to T1 are considered to be noise samples, and the corresponding weight is 0.1, training samples with L2 norm between T1 and T2 are considered to be difficult samples, and samples with L2 norm greater than or equal to T2 are considered to be easy samples.
[0135] Embodiment 3
[0136] Figure 3A classification prediction method is provided in Embodiment 3 of the present invention. This embodiment can be used for classification prediction. The method adopts a neural network model trained by a training method of a neural network model provided in any of the foregoing embodiments of the present invention. The method can be performed by a classification prediction device provided in an embodiment of the present invention. The device can be implemented in software and / or hardware and integrated in a computer device provided in an embodiment of the present invention. Figure 3 As shown, the method specifically comprises the following steps:
[0137] S301: Obtain samples to be classified.
[0138] Exemplarily, in a specific embodiment of the present invention, the sample to be classified is a face image.
[0139] S302: Input the samples to be classified into the trained neural network model for processing to obtain a probability vector for predicting whether the samples to be classified belong to each category.
[0140] Exemplarily, the neural network model can perform convolution, pooling, feature fusion and other operations on the input face image, extract the feature vector used to characterize the features of the face image, and output the logits vector for predicting each ID to which the face image belongs through conversion of the fully connected network layer. The logits vector is normalized through the softmax function to obtain the probability vector for predicting that the face image belongs to each ID.
[0141] S303: Determine the category to which the sample to be classified belongs based on the probability vector.
[0142] Each position in the probability vector corresponds to each category, and the category corresponding to the position with the largest element value in the probability vector is the category predicted by the model to which the sample to be classified belongs. For example, 10,000 positions in the probability vector correspond to the IDs of 10,000 people, and the ID corresponding to the position with the largest element value in the probability vector is used as the predicted ID of the sample to be classified predicted by the model.
[0143] Embodiment 4
[0144] Figure 4 A schematic diagram of a training device for a neural network model provided in Embodiment 4 of the present invention is shown in FIG. Figure 4 As shown, the device comprises:
[0145] A first data acquisition module 401 is used to acquire an i-th batch of data for training a neural network model, wherein the i-th batch of data includes a plurality of training samples, and the training samples are annotated with labels indicating the categories of the training samples;
[0146] A probability value acquisition module 402 is used to input the plurality of training samples of the i-th batch of data into the neural network model to be trained for processing, so as to obtain a first probability vector for predicting that the training samples belong to each category;
[0147] A cross entropy loss value calculation module 403, used to calculate the cross entropy loss value of the training sample based on the first probability vector and the label of the training sample;
[0148] An L2 norm calculation module 404 is used to calculate the L2 norm of the feature vector of the training sample, where the feature vector is extracted by the neural network model and is used to map and obtain the probability vector;
[0149] A weight determination module 405 is used to determine the weight of the cross entropy loss value of the training sample based on the L2 norm of the feature vector;
[0150] A first total loss value calculation module 406 is used to calculate the sum of the cross entropy loss values of the training samples in the i-th batch of data multiplied by the weighted loss values to obtain a first total loss value;
[0151] The parameter updating module 407 is used to update the parameters of the neural network model based on the first total loss value.
[0152] In some embodiments of the present invention, the weight determination module 405 includes:
[0153] An L2 norm storage submodule, used for storing L2 norms corresponding to a plurality of training samples of the i-th batch of data into a data queue;
[0154] A weight determination submodule is used to determine the weight of the cross entropy loss value of the training sample based on the mean and standard deviation of all L2 norms in the data queue and the L2 norm of the feature vector of the training sample.
[0155] In some embodiments of the present invention, the data queue has m×p storage bits, m is the number of training samples in each batch of data, p is an integer greater than 1, and the L2 norm storage submodule includes:
[0156] A judging unit, used for judging whether there is any vacant storage position in the data queue;
[0157] A first target storage bit determination unit, configured to determine a target storage bit to be stored from the vacant storage bits when there are vacant storage bits in the data queue;
[0158] A first target storage position determination unit is used to erase the L2 norms corresponding to multiple training samples of the ipth batch of data to obtain the target storage position when there is no vacant storage position in the data queue;
[0159] The L2 norm storage unit is used to store the L2 norms corresponding to the multiple training samples of the i-th batch of data into the target storage bit.
[0160] In some embodiments of the present invention, the weight determination submodule includes:
[0161] A calculation unit, used to calculate the mean and standard deviation of all L2 norms in the data queue;
[0162] A partition division unit, used to determine a plurality of partitions for representing the difficulty of the training samples based on the mean and the standard deviation, each partition having a corresponding weight of the cross entropy loss value;
[0163] A partition determination unit, used to determine the partition to which the L2 norm of the feature vector of the training sample belongs;
[0164] A weight determination unit is used to use the weight corresponding to the partition to which the L2 norm of the feature vector of the training sample belongs as the weight of the cross entropy loss value of the training sample.
[0165] In some embodiments of the present invention, the weight determination submodule includes:
[0166] A Gaussian distribution function determination unit, configured to construct a Gaussian distribution function of the L2 norm based on a mean and a standard deviation of the L2 norm in the data queue;
[0167] an adaptive weight function construction unit for the Gaussian distribution function and the standard deviation of the L2 norm in the data queue;
[0168] The weight determination unit is used to input the L2 norm of the feature vector of the training sample into the adaptive weight function to obtain the weight of the cross entropy loss value of the training sample.
[0169] In some embodiments of the present invention, the Gaussian distribution function is as follows:
[0170]
[0171] Wherein, z is the L2 norm of the feature vector of the training sample, μ is the mean of the L2 norm in the data queue, and σ is the standard deviation of the L2 norm in the data queue.
[0172] In some embodiments of the present invention, the adaptive weight function is as follows:
[0173]
[0174] Wherein, z is the L2 norm of the feature vector of the training sample, G(z) is the Gaussian distribution function, σ is the standard deviation of the L2 norm in the data queue, and T is the preset L2 norm threshold.
[0175] In some embodiments of the present invention, the training module of the neural network model further includes:
[0176] A second data acquisition module, used to acquire multiple batches of data for training the neural network model before acquiring the i-th batch of data for training the neural network model;
[0177] The initial training module is used to perform initial training on the neural network model using multiple batches of data until the neural network model meets preset conditions.
[0178] In some embodiments of the present invention, the initial training module includes:
[0179] A probability vector acquisition submodule, used to input the plurality of training samples of the j-th batch of data into the neural network model to be trained for processing, and obtain a second probability vector for predicting that the training samples belong to each category;
[0180] A cross entropy loss value calculation submodule, used to calculate the cross entropy loss value of the training sample based on the second probability vector and the label of the training sample;
[0181] A second total loss value calculation submodule, used to calculate the sum of the cross entropy loss values of each training sample in the j-th batch of data to obtain a second total loss value;
[0182] A judging submodule, used for judging whether the second total loss is less than or equal to a preset value;
[0183] A parameter updating submodule, used for updating the parameters of the neural network model when the second total loss value is greater than the preset value, and inputting the plurality of training samples of the j+1th batch of data into the neural network model to be trained for processing;
[0184] The preset condition satisfaction determination submodule is used to determine that the neural network model satisfies the preset condition when the second total loss value is less than or equal to a preset value.
[0185] The training device for the neural network model described above can execute the training method for the neural network model provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0186] Embodiment 5
[0187] Figure 5 A schematic diagram of the structure of a classification prediction device provided by an embodiment of the present invention is shown in FIG. Figure 5 As shown, the device comprises:
[0188] The to-be-classified sample acquisition module 501 is used to acquire the to-be-classified samples.
[0189] The probability vector acquisition module 502 is used to input the samples to be classified into the trained neural network model for processing, and obtain the probability vector used to predict whether the samples to be classified belong to each category.
[0190] The category determination module 503 is used to determine the category to which the sample to be classified belongs based on the probability vector.
[0191] The above-mentioned classification prediction device can execute the classification prediction method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0192] Embodiment 6
[0193] Embodiment 6 of the present invention provides a computer device, Figure 6 A schematic diagram of the structure of a computer device provided in Embodiment 6 of the present invention is shown in FIG. Figure 6 As shown, the computer device includes a processor 601, a memory 602, a communication module 603, an input device 604 and an output device 605; the number of processors 601 in the computer device can be one or more. Figure 6 A processor 601 is taken as an example; the processor 601, memory 602, communication module 603, input device 604 and output device 605 in the computer device can be connected by a bus or other means. Figure 6 The above processor 601, memory 602, communication module 603, input device 604 and output device 605 can be integrated on a control mainboard of a computer device.
[0194] The memory 602, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as the modules corresponding to the training method or classification prediction method of the neural network model in the embodiment of the present invention. The processor 601 executes various functional applications and data processing of the computer device by running the software programs, instructions and modules stored in the memory 602, that is, the training method or classification prediction method of the neural network model provided in the above embodiment is implemented.
[0195] The memory 602 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 602 may further include a memory remotely arranged relative to the processor 601, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0196] The communication module 603 is used to establish a connection with an external device (such as a smart terminal) and realize data interaction with the external device. The input device 604 can be used to receive input digital or character information and generate key signal input related to user settings and function control of the computer device.
[0197] A computer device provided in this embodiment can execute the training method or classification prediction method of the neural network model provided in any of the above embodiments of the present invention, and has specific corresponding functions and beneficial effects.
[0198] Embodiment 7
[0199] Embodiment 7 of the present invention provides a storage medium containing computer executable instructions, on which a computer program is stored. When the program is executed by a processor, the training method or classification prediction method of the neural network model provided in any of the above embodiments of the present invention is implemented.
[0200] The training methods of neural network models include:
[0201] Acquire an i-th batch of data for training a neural network model, wherein the i-th batch of data includes a plurality of training samples, and the training samples are annotated with labels representing categories of the training samples;
[0202] Inputting the plurality of training samples of the i-th batch of data into the neural network model to be trained for processing respectively, and obtaining a first probability vector for predicting that the training samples belong to each category;
[0203] Calculating a cross entropy loss value of the training sample based on the first probability vector and the label of the training sample;
[0204] Calculating the L2 norm of the feature vector of the training sample, where the feature vector is extracted by the neural network model and is used for mapping to obtain the first probability vector;
[0205] Determining a weight of a cross entropy loss value of the training sample based on an L2 norm of the feature vector;
[0206] Calculate the sum of the cross entropy loss values of the training samples in the i-th batch of data multiplied by the weighted loss values to obtain a first total loss value;
[0207] Update parameters of the neural network model based on the first total loss value.
[0208] The classification prediction method adopts a neural network model trained by the training method of the neural network model provided by any of the above embodiments of the present invention, including:
[0209] Obtain samples to be classified;
[0210] The samples to be classified are input into the trained neural network model for processing to obtain the probability vector used to predict whether the samples to be classified belong to each category;
[0211] The category to which the sample to be classified belongs is determined based on the probability vector.
[0212] Of course, the storage medium containing computer executable instructions provided by the embodiment of the present invention is not limited to the method operations described above, but can also execute related operations in the training method or classification prediction method of the neural network model provided by the embodiment of the present invention.
[0213] It should be noted that, for the device, equipment and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0214] Through the above description of the implementation methods, the technicians in the relevant field can clearly understand that the present invention can be implemented by means of software and necessary general hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for a computer device (which can be a robot, a personal computer, a server, or a network device, etc.) to execute the training method or classification prediction method of the neural network model described in any embodiment of the present invention.
[0215] It is worth noting that in the above-mentioned device, the various modules, sub-modules and units included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0216] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution device. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0217] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0218] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A training method for a face recognition neural network model, characterized in that: include: Acquire an i-th batch of facial image data for training a facial recognition neural network model, wherein the i-th batch of facial image data includes a plurality of facial image training samples, wherein the facial image training samples are annotated with labels indicating categories of persons to which the facial image training samples belong; Inputting the plurality of face image training samples of the i-th batch of face image data into the face recognition neural network model to be trained for processing respectively, and obtaining a first probability vector for predicting that the face image training samples belong to the category of each person; Calculate the cross entropy loss value of the face image training sample based on the first probability vector and the label of the face image training sample; Calculating the L2 norm of the feature vector of the face image training sample, where the feature vector is extracted by the face recognition neural network model and is used for mapping to obtain the first probability vector; Determining the weight of the cross entropy loss value of the face image training sample based on the L2 norm of the feature vector; Calculate the sum of the cross entropy loss values of each of the face image training samples in the i-th batch of face image data multiplied by the weighted loss value to obtain a first total loss value; Update the parameters of the face recognition neural network model based on the first total loss value.
2. The training method of the face recognition neural network model according to claim 1, characterized in that: Determining the weight of the cross entropy loss value of the face image training sample based on the L2 norm of the feature vector includes: Storing L2 norms corresponding to multiple face image training samples of the i-th batch of face image data in a data queue; The weight of the cross entropy loss value of the face image training sample is determined based on the mean and standard deviation of all L2 norms in the data queue and the L2 norm of the feature vector of the face image training sample.
3. The training method of the face recognition neural network model according to claim 2, characterized in that: The data queue has m×p storage bits, where m is the number of face image training samples in each batch of face image data, and p is an integer greater than 1, and L2 norms corresponding to multiple face image training samples of the i-th batch of face image data are stored in the data queue, including: Determine whether there is any vacant storage position in the data queue; If yes, determining a target storage location to be stored from the vacant storage locations; If not, then erase the L2 norms corresponding to the multiple face image training samples of the face image data of the ipth batch to obtain the target storage bit; The L2 norms corresponding to the multiple face image training samples of the i-th batch of face image data are stored in the target storage bit.
4. The training method of the face recognition neural network model according to claim 2, characterized in that: Determining the weight of the cross entropy loss value of the face image training sample based on the mean and standard deviation of all L2 norms in the data queue and the L2 norm of the feature vector of the face image training sample includes: Calculate the mean and standard deviation of all L2 norms in the data queue; Determine multiple partitions for representing the difficulty of the face image training samples based on the mean and the standard deviation, each partition having a corresponding weight of the cross entropy loss value; Determine the partition to which the L2 norm of the feature vector of the face image training sample belongs; The weight corresponding to the partition to which the L2 norm of the feature vector of the face image training sample belongs is used as the weight of the cross entropy loss value of the face image training sample.
5. The training method of the face recognition neural network model according to claim 2, characterized in that: Determining the weight of the cross entropy loss value of the face image training sample based on the mean and standard deviation of all L2 norms in the data queue and the L2 norm of the feature vector of the face image training sample includes: Constructing a Gaussian distribution function of the L2 norm based on the mean and standard deviation of the L2 norm in the data queue; Constructing an adaptive weight function based on the Gaussian distribution function and the standard deviation of the L2 norm in the data queue; The L2 norm of the feature vector of the face image training sample is input into the adaptive weight function to obtain the weight of the cross entropy loss value of the face image training sample.
6. The training method of the face recognition neural network model according to claim 5, characterized in that: The Gaussian distribution function is as follows: Wherein, z is the L2 norm of the feature vector of the face image training sample, μ is the mean of the L2 norm in the data queue, and σ is the standard deviation of the L2 norm in the data queue.
7. The training method of the face recognition neural network model according to claim 5, characterized in that: The adaptive weight function is as follows: Wherein, z is the L2 norm of the feature vector of the face image training sample, G(z) is the Gaussian distribution function, σ is the standard deviation of the L2 norm in the data queue, and T is the preset L2 norm threshold.
8. The method for training a face recognition neural network model according to any one of claims 1 to 7, characterized in that: Before obtaining the i-th batch of face image data for training the face recognition neural network model, it also includes: Obtain multiple batches of face image data for training a face recognition neural network model; The face recognition neural network model is initially trained using multiple batches of face image data until the face recognition neural network model meets preset conditions.
9. The training method of the face recognition neural network model according to claim 8, characterized in that: Initially training the face recognition neural network model using multiple batches of face image data until the face recognition neural network model meets preset conditions, including: Inputting the plurality of face image training samples of the jth batch of face image data into the face recognition neural network model to be trained for processing respectively, and obtaining a second probability vector for predicting that the face image training samples belong to the category of each person; Calculate the cross entropy loss value of the face image training sample based on the second probability vector and the label of the face image training sample; Calculate the sum of the cross entropy loss values of each face image training sample in the j-th batch of face image data to obtain a second total loss value; Determining whether the second total loss is less than or equal to a preset value; When the second total loss value is greater than the preset value, updating the parameters of the face recognition neural network model, and inputting the plurality of face image training samples of the j+1th batch of face image data into the face recognition neural network model to be trained for processing; When the second total loss value is less than or equal to a preset value, it is determined that the face recognition neural network model meets a preset condition.
10. A training device for a face recognition neural network model, characterized in that: include: A data acquisition module, used to acquire an i-th batch of face image data for training a face recognition neural network model, wherein the i-th batch of face image data includes a plurality of face image training samples, wherein the face image training samples are annotated with labels indicating categories of persons to which the face image training samples belong; A first probability value acquisition module, used for inputting the plurality of face image training samples of the i-th batch of face image data into the face recognition neural network model to be trained for processing, and obtaining a first probability vector for predicting that the face image training samples belong to the category of each person; A cross entropy loss value calculation module, used to calculate the cross entropy loss value of the face image training sample based on the first probability vector and the label of the face image training sample; An L2 norm calculation module, used to calculate the L2 norm of the feature vector of the face image training sample, where the feature vector is extracted by the face recognition neural network model and is used to map to obtain the probability vector; A weight determination module, used to determine the weight of the cross entropy loss value of the face image training sample based on the L2 norm of the feature vector; A first total loss value calculation module is used to calculate the sum of the cross entropy loss values of each of the face image training samples in the i-th batch of face image data multiplied by the weighted loss value to obtain a first total loss value; A parameter updating module is used to update the parameters of the face recognition neural network model based on the first total loss value.
11. A computer device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the training method of the face recognition neural network model as described in any one of claims 1-9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the training method of the face recognition neural network model as described in any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Classifier, neural network model training method, data processing device, and medium
CN109902722A
Data processing method, device and equipment and readable storage medium
CN110020022A