Face recognition model training method, device and equipment based on difficult case mining

By setting sample difficulty parameters and iterative training methods based on these parameters, the problem of difficult samples being ignored in the existing technology is solved, and the training effect of face recognition models and the accuracy of difficult samples mining are improved.

CN120014677APending Publication Date: 2025-05-16JINAN BOGUAN INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311528869.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, margin-based loss function ignores difficult samples when training face recognition models, resulting in poor model training effect, and existing difficult sample mining methods are inefficient and subjective, and cannot guarantee the accuracy of mining results.

Method used

By setting sample difficulty parameters, determining the sample difficulty influence factor based on sample information, and iterative training of the training samples based on these parameters, the accuracy, robustness and objectivity of sample difficulty are improved.

Benefits of technology

The effect of face recognition model training is improved, and difficult samples are mined online through the adaptive difficulty indicator loss, which enhances the model's utilization of sample information of different difficulty levels and improves the final performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014677A_ABST
    Figure CN120014677A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a face recognition model training method, device and equipment based on difficult case mining. The method comprises the following steps: acquiring sample information of a to-be-trained sample; determining a sample difficulty influence factor according to the sample information, and determining a sample difficulty parameter of the to-be-trained sample according to the sample difficulty influence factor; and performing iterative training on the to-be-trained sample based on the sample difficulty parameter to obtain a training result of the face recognition model. According to the technical scheme, a universal sample difficulty parameter is set based on the sample difficulty influence factor, and compared with other simple definition of mining loss on a difficult sample, the provided sample difficulty parameter can adaptively indicate the sample difficulty and the change of the sample difficulty along with the model training process in a numerical form; the accuracy, robustness and objectivity of sample difficulty indication are improved. In addition, due to the universality of the sample difficulty parameter, the method can also be used for offline difficult sample mining, data set denoising and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of model training technology, and in particular to a face recognition model training method, device and equipment based on hard example mining. Background Art

[0002] With the development of biometric technology, face recognition has been widely used in payment, security, entertainment and other fields. Among them, face recognition is based on facial feature information for identity recognition, and has the characteristics of non-compulsory, non-contact and concurrency.

[0003] At present, the face recognition model is often trained by using the margin-based loss function, and face recognition is performed through the face recognition model. However, the margin-based loss function often ignores the important role of difficult samples in training, resulting in poor model training results.

[0004] In the prior art, difficult samples are usually mined manually, which is inefficient and highly subjective, and cannot guarantee the accuracy of the mining results, thus failing to describe the difficulty of the samples well. Therefore, how to effectively mine difficult samples to fully utilize sample information of different difficulty levels for model training, thereby improving the model training effect, is one of the issues worthy of attention in current model training. Summary of the invention

[0005] The present invention provides a face recognition model training method, device and equipment based on difficult example mining. By setting sample difficulty parameters, the sample difficulty can be quantitatively described, thereby improving the accuracy, robustness and objectivity of the sample difficulty indication.

[0006] According to one aspect of the present invention, a face recognition model training method based on hard example mining is provided, the method comprising:

[0007] Obtain sample information of samples to be trained;

[0008] Determining a sample difficulty influencing factor according to the sample information, and determining a sample difficulty parameter of the sample to be trained according to the sample difficulty influencing factor;

[0009] The samples to be trained are iteratively trained based on the sample difficulty parameter to obtain a training result of the face recognition model.

[0010] According to another aspect of the present invention, a face recognition model training device based on hard example mining is provided, comprising:

[0011] A sample information acquisition module is used to obtain sample information of samples to be trained;

[0012] A sample difficulty parameter determination module, used to determine a sample difficulty influencing factor according to the sample information, and determine a sample difficulty parameter of the sample to be trained according to the sample difficulty influencing factor;

[0013] The face recognition model training module is used to iteratively train the samples to be trained based on the sample difficulty parameter to obtain the training result of the face recognition model.

[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the face recognition model training method based on hard example mining described in any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the face recognition model training method based on hard example mining described in any embodiment of the present invention when executed.

[0019] The technical solution of the embodiment of the present invention obtains sample information of samples to be trained; determines a sample difficulty influencing factor based on the sample information, and determines a sample difficulty parameter of the samples to be trained based on the sample difficulty influencing factor; iteratively trains the samples to be trained based on the sample difficulty parameter to obtain the training result of the face recognition model. This technical solution sets a universal sample difficulty parameter based on the sample difficulty influencing factor. Compared with the simple definition of difficult samples in other mining losses, the proposed sample difficulty parameter can adaptively indicate the sample difficulty and its changes with the model training process in numerical form, thereby improving the accuracy, robustness and objectivity of the sample difficulty indication. In addition, due to the universality of the sample difficulty parameter, it can also be used for offline difficult sample mining and data set denoising.

[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 is a flowchart of a face recognition model training method based on hard example mining provided according to Embodiment 1 of the present invention;

[0023] Figure 2 is a schematic diagram of changes in training phase parameters on a face data set during model training provided by Embodiment 1 of the present invention;

[0024] Figure 3 is a schematic diagram of a face recognition model training process based on hard example mining provided according to Embodiment 1 of the present invention;

[0025] Figure 4 is a flowchart of a face recognition model training method based on hard example mining provided according to Embodiment 2 of the present invention;

[0026] Figure 5 is a structural schematic diagram of a face recognition model training device based on hard example mining provided according to Embodiment 3 of the present invention;

[0027] Figure 6 It is a structural schematic diagram of an electronic device for implementing a face recognition model training method based on hard example mining according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first", "second", "target", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] Embodiment 1

[0031] Figure 1 The flowchart of a face recognition model training method based on hard example mining provided in the first embodiment of the present invention is applicable to the case where the face recognition model is trained through hard example mining. The method can be executed by a face recognition model training device based on hard example mining. The face recognition model training device based on hard example mining can be implemented in the form of hardware and / or software. The face recognition model training device based on hard example mining can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes:

[0032] S110, obtaining sample information of samples to be trained.

[0033] The samples to be trained as input of the face recognition model may be face images used for training the face recognition model or face videos composed of face images. Exemplarily, the sample information may include the loss value of the samples to be trained, the cosine similarity of the samples to be trained, and the number of samples of the category to which the samples to be trained belong.

[0034] S120, determining a sample difficulty influencing factor according to the sample information, and determining a sample difficulty parameter of the sample to be trained according to the sample difficulty influencing factor.

[0035] Among them, the sample difficulty factor can reflect the difficulty level of the training sample more objectively and robustly in numerical form. It has sufficient robustness to alleviate the impact of noise data and can intuitively display the difficulty changes of the training sample during the model training process, thereby more finely distinguishing the training samples of different difficulties. In addition, the sample difficulty factor can also be used as a general indicator of difficult examples, which can be used for offline difficult sample mining and data set denoising.

[0036] In this embodiment, optionally, the sample difficulty influencing factor includes a sample loss influencing factor, a sample misclassification influencing factor and a sample imbalance influencing factor, wherein the sample loss influencing factor is used to indicate the contribution of sample loss to reducing sample difficulty during network training; the sample misclassification influencing factor is used to indicate the contribution of the number of times a sample is misclassified during training to reducing sample difficulty during network training; the sample imbalance influencing factor is used to indicate the contribution of the distribution balance of the category to which the sample belongs to reducing sample difficulty during network training.

[0037] Among them, sample loss is an important feature of sample difficulty, and the training loss of samples of different difficulty levels shows great differences. For example, the loss of a simple sample with a frontal face and clear face can converge quickly, while the loss of a semi-difficult sample with a side face but relatively clear face decreases relatively slowly. For difficult samples such as severe blur, the loss decreases slowly or cannot decrease during training. The sample loss impact factor can be used to evaluate the contribution of sample loss to the loss of the entire model. Among them, for difficult samples, the contribution of sample loss is relatively high. Therefore, the sample loss impact factor can be used to indicate that the smaller the sample loss, the greater the contribution to reducing the difficulty of the sample.

[0038] In addition to sample loss, whether the sample is misclassified during the training process (i.e., sample misclassification) is another important feature of difficult samples. There are obvious differences in the correct and incorrect classification of samples of different difficulty levels during the training process. For example, samples with clear frontal faces can be correctly classified in most training stages except for the first few rounds of misclassification during the training process. For semi-difficult samples such as half-profile faces, it is often necessary to wait until the middle and late stages of training to be correctly classified by the network, and it is easy to be misclassified again. For difficult samples such as complete profile faces, it is difficult for the network to correctly classify them during the entire training process. The sample misclassification impact factor can be used to evaluate the misclassification of samples as the iterations proceed. Among them, simple samples will be correctly classified by the model quickly as the number of iterations increases, while difficult samples are more difficult to be correctly classified during the entire training process. Simple samples are easily correctly classified by the network, and the more times they are correctly classified, the greater their contribution to reducing the difficulty of the sample. For samples that are more difficult for the network to learn, it is not easy to be correctly classified during the training process. Therefore, the sample misclassification impact factor can be specifically used to indicate that the more times a sample is misclassified during the training process, the smaller its contribution to reducing the difficulty of the sample.

[0039] In addition to the above two influencing factors, difficult samples have another important feature that is easily overlooked, namely, the balance of sample categories. There is a serious sample imbalance problem in the existing face training data set, and most categories have a small number of samples. For categories with sparse samples, the space of existing samples may only be a small part of the true distribution of the category, and the samples in this category are relatively more difficult to learn. The sample imbalance influencing factor can be used to evaluate the balance of the sample, that is, the proportion of the number of samples in the class to which the sample belongs in the overall training set. Among them, samples of categories with a higher proportion are easier to train, while samples of categories with a smaller proportion are more difficult to train. Therefore, the sample imbalance influencing factor can be used to specifically indicate that the higher the proportion of samples in the category to which the sample belongs, the greater the contribution to reducing the difficulty of the sample.

[0040] In this embodiment, optionally, determining the sample loss impact factor according to the sample information includes: determining the sample loss impact factor according to the loss value of the sample to be trained in the sample information based on the following formula: Among them, H1(x i ) represents the training sample x i The sample loss influence factor, α1 represents the first preset hyperparameter, α1>0, Represents the training sample x i The loss value in the current iteration, L max and L min Indicates the maximum loss value and the minimum loss value of the sample in the current iteration.

[0041] Among them, S(x i )express The normalized value of . Since the smaller the sample loss, the greater the contribution to reducing the sample difficulty, H1(x i ) is the sample loss is a monotonically decreasing function. In order to reduce the impact of different training stages on H1(x i ) (the sample loss is large in the early stage of training, while the sample loss is small in the later stage of training), it is necessary to adjust the sample loss Normalize. Then we need a suitable monotonically decreasing function to convert S(x i ) is mapped to [0, 1]. It should be noted that, in addition to selecting H1(x i ) function, any other monotonically decreasing function that can map the normalized sample loss value to [0, 1] is acceptable.

[0042] In this embodiment, optionally, determining the sample misclassification influence factor according to the sample information includes: determining the sample misclassification influence factor according to the cosine similarity of the samples to be trained in the sample information based on the following formula: Among them, H2(x i) represents the training sample x i The sample misclassification influence factor is is the sample x to be trained i The weights in round a of training, and y i is the category label of the sample to be trained, cosθ j Represents the training sample x i The corresponding cosine similarity of the j-th category, j = 1, 2, 3, ..., K, K represents the number of sample categories.

[0043] Among them, cosθ i By taking the training sample x i The feature vector of is normalized and then inner-producted with the weight of the j-th class to be used to characterize the training sample x i The similarity between the features of and the jth category. It should be noted that in the process of model training, samples are usually trained according to the batch size, that is, all samples are trained in batches, and each batch of samples is different. After completing the training of all samples through multiple iterations, a round of training ends. In other words, a round of training includes multiple iterations of training.

[0044] In this embodiment, optionally, determining the sample imbalance impact factor according to the sample information includes: determining the sample imbalance impact factor according to the number of samples of the category to which the to-be-trained sample belongs in the sample information based on the following formula: Among them, H3(x i ) represents the training sample x i The sample imbalance impact factor is Represents the training sample x i The number of samples belonging to the category, Represents the training sample x i The distribution estimation result of the number of samples belonging to the category.

[0045] in, Determined based on the kernel probability density function. Specifically, first estimate the kernel probability density of the sample number distribution of each category, then calculate the corresponding cumulative distribution function based on the estimated kernel probability density, and then obtain the distribution estimation result.

[0046] In this embodiment, optionally, determining the sample difficulty parameter of the sample to be trained according to the sample difficulty influencing factor includes: determining the sample difficulty parameter of the sample to be trained based on the following sample difficulty indication function: Among them, H(x i ) represents the training sample xi The sample difficulty parameter, H1(x i ) represents the training sample x i The sample loss influence factor, H2(x i ) represents the training sample x i The sample misclassification influence factor, H3(x i ) represents the training sample x i , λ1 represents the weight coefficient of the sample loss factor, and λ2 represents the weight coefficient of the sample misclassification factor.

[0047] Among them, λ1 and λ2 can be preset according to actual application requirements, and satisfy λ1+λ2=1. i ) can be used to characterize the difficulty of a sample. The larger the value, the greater the difficulty of the sample.

[0048] S130, iteratively training the training samples based on the sample difficulty parameter to obtain a training result of the face recognition model.

[0049] In this embodiment, after determining the sample difficulty parameter, the training sample can be iteratively trained based on the sample difficulty parameter, thereby obtaining the training result of the face recognition model. Optionally, iteratively training the training sample based on the sample difficulty parameter includes: determining an adaptive difficulty indicator loss according to the sample difficulty parameter based on the following formula, and iteratively training the training sample according to the adaptive difficulty indicator loss; Among them, L represents the training sample x i The adaptive difficulty indicator loss of , N represents the number of samples to be trained in the current batch, s represents the second preset hyperparameter, N(t (k) ,H(x i ), cosθ j )=I(t (k) ,H(x i ), cosθ j )·cosθ j , m represents the constraint margin parameter, cosθ j Represents the training sample x i The cosine similarity of the corresponding j-th category, j = 1, 2, 3, ..., K, K represents the number of sample categories, y i is the category label of the sample to be trained, t (k) Represents the training phase parameters corresponding to the k-th iteration training.

[0050] It should be noted that in order to improve the final performance of the model, it is necessary to make full use of the difficult sample information. However, if the difficult samples are always emphasized throughout the training phase, it may cause convergence problems, so it is necessary to consider the training phase information. Exemplarily, in this embodiment, the sliding average of the sine and cosine similarity means is used to estimate the model training phase. Optionally, the training phase parameters are determined by the following formula: (k) =αr (k) +(1-α)t (k-1) ; where α represents the momentum parameter, t (0) =0.

[0051] Figure 2 A schematic diagram of the change of training phase parameters on a face dataset during model training provided in the first embodiment of the present invention. CASIA-WebFace and MS1MV2 represent two different face datasets. Figure 2 It can be seen that t (k) The value of can accurately represent the training stage of the model. By introducing the training stage parameter t (k) , the constraints on difficult samples can be adaptively adjusted according to different training stages, making the training process more flexible.

[0052] In this embodiment, the sample difficulty parameter H(x i ) and training phase parameter t (k) Integrate into the adaptive difficulty indicator loss. Where I(t (k) ,H(x i ), cosθ j ) is the negative cosine similarity modulation coefficient that combines the training phase information and the sample difficulty information. It should be noted that the sine-cosine similarity can use any margin-based loss function, such as the sine-cosine similarity function in ArcFace. Specifically, the modulation coefficient I(t (k) ,H(x i ), cosθ j ) depends on t (k) 、H(x i ) and cosθ j In the early training stage, learning from simple samples is conducive to model convergence, and in this training stage, t (k) Close to 0 and t (k) +cosθ j is less than 1, so I(t (k) ,H(x i ), cosθ j ) = t (k) +cosθ j, so that the weight of difficult samples is reduced, so that simple samples can be relatively more concerned. As training progresses, the loss function should gradually emphasize difficult samples in order to utilize sample difficulty information and improve the final performance of the model. In the later stage of training, t (k) Close to 1 and t (k) +cosθ j Greater than 1. Therefore, in this training phase, I(t (k) ,H(x i ), cosθ j ) = t (k) +(t (k) +H(x i ))cosθ j , where t (k) +H(x i )>1 is the global difficulty factor that includes information from the training phase. Negative cosine similarity cosθ of difficult samples j By t (k) +H(x i ) is weighted to further enhance the discrimination of the weights assigned to difficult samples. (k) +H(x i ))cosθ j It can be regarded as a new sample difficulty index that combines global and local difficulty indexes, which can further increase the weight assigned to more difficult samples, while less difficult samples will be assigned smaller weights. (k) +H(x i ))cosθ j , the contribution of difficult samples of different difficulty levels to the total loss can be further expanded, so that the face recognition model can make more full use of the information of difficult samples.

[0053] In order to solve the problem that the existing mining loss cannot effectively distinguish samples of different difficulty, this scheme proposes to mine difficult samples online through adaptive difficulty indicator loss, which integrates global and local sample difficulty information and can impose more unique constraints on each sample according to the training difficulty of the sample at the current stage, so as to mine the sample difficulty information online. Among them, the modulation coefficient of the negative cosine similarity of difficult samples combines global and local sample difficulty indicators to further distinguish the constraints imposed on different samples and make full use of sample information of different difficulty. At the same time, in order to avoid the influence of difficult samples on network convergence in the early stage of training, the training stage parameters are introduced into the loss to impose training stage-related constraints on the samples, so as to adaptively adjust the training strategy to accelerate model convergence. Therefore, the proposed adaptive difficulty indicator loss can achieve a good balance between convergence speed and the final performance of the model.

[0054] Figure 3A schematic diagram of a face recognition model training process based on hard example mining provided in the first embodiment of the present invention. Figure 3 As shown, first input relevant parameters into the network, including x i (sample to be trained), y i (x i The category label), W (the last fully connected layer parameter), cosθ j (x i The corresponding cosine similarity of the jth class), θ (network parameters) and λ (learning rate), and the number of iterations k and its corresponding training phase parameter t (k) Initialize to k = 0 and t (0) = 0. Then determine whether the network converges. If not, calculate the sample difficulty parameter H(x i ) and training phase parameter t (k) , then calculate and N(t (k) ,H(x i ), cosθ j ), and then calculate the adaptive difficulty indicator loss L, and then calculate the gradient and update the parameters through SGD (stochastic gradient descent). The specific update is: Then, iterative training is performed again based on the updated parameters until the network converges, and the parameters W and θ obtained in the last iteration are output. If the network converges, the parameters W and θ are directly output.

[0055] The technical solution of the embodiment of the present invention obtains sample information of samples to be trained; determines a sample difficulty influencing factor based on the sample information, and determines a sample difficulty parameter of the samples to be trained based on the sample difficulty influencing factor; iteratively trains the samples to be trained based on the sample difficulty parameter to obtain the training result of the face recognition model. This technical solution sets a universal sample difficulty parameter based on the sample difficulty influencing factor. Compared with the simple definition of difficult samples in other mining losses, the proposed sample difficulty parameter can adaptively indicate the sample difficulty and its changes with the model training process in numerical form, thereby improving the accuracy, robustness and objectivity of the sample difficulty indication. In addition, due to the universality of the sample difficulty parameter, it can also be used for offline difficult sample mining and data set denoising.

[0056] Embodiment 2

[0057] Figure 4 A flowchart of a face recognition model training method based on hard example mining is provided in the second embodiment of the present invention. This embodiment is optimized based on the above embodiment.

[0058] like Figure 4 As shown, the method of this embodiment specifically includes the following steps:

[0059] S210: Obtain sample information of samples to be trained.

[0060] S220, determining a sample loss influence factor according to the loss value of the sample to be trained in the sample information.

[0061] Among them, the sample loss impact factor can be determined based on the following formula:

[0062]

[0063] Among them, H1(x i ) represents the training sample x i The sample loss influence factor, α1 represents the first preset hyperparameter, α1>0, Represents the training sample x i The loss value in the current iteration, L max and L min Indicates the maximum loss value and the minimum loss value of the sample in the current iteration.

[0064] S230, determining a sample misclassification influencing factor according to the cosine similarity of the samples to be trained in the sample information.

[0065] Among them, the sample misclassification influencing factor can be determined based on the following formula:

[0066]

[0067] Among them, H2(x i ) represents the training sample x i The sample misclassification influence factor is is the sample x to be trained i The weights in round a of training, and y i is the category label of the sample to be trained, cosθ j Represents the training sample x i The corresponding cosine similarity of the j-th category, j = 1, 2, 3, ..., K, K represents the number of sample categories.

[0068] S240: Determine a sample imbalance impact factor according to the number of samples of the category to which the to-be-trained samples belong in the sample information.

[0069] Among them, the sample imbalance impact factor can be determined based on the following formula:

[0070]

[0071] Among them, H3(xi ) represents the training sample x i The sample imbalance impact factor is Represents the training sample x i The number of samples belonging to the category, Represents the training sample x i The distribution estimation result of the number of samples belonging to the category.

[0072] S250, determining a sample difficulty parameter of the sample to be trained according to the sample loss influence factor, the sample misclassification influence factor and the sample imbalance influence factor.

[0073] Among them, the sample difficulty parameter can be determined based on the following formula:

[0074]

[0075] Among them, H(x i ) represents the training sample x i The sample difficulty parameter, H1(x i ) represents the training sample x i The sample loss influence factor, H2(x i ) represents the training sample x i The sample misclassification influence factor, H3(x i ) represents the training sample x i , λ1 represents the weight coefficient of the sample loss factor, and λ2 represents the weight coefficient of the sample misclassification factor.

[0076] S260, determining an adaptive difficulty indicator loss according to the sample difficulty parameter.

[0077] Among them, the adaptive difficulty indicator loss can be determined based on the following formula:

[0078]

[0079] Among them, L represents the training sample x i The adaptive difficulty indicator loss of , N represents the number of samples to be trained in the current batch, s represents the second preset hyperparameter, N(t (k) ,H(x i ), cosθ j )=I(t (k) ,H(x i ), cosθ j )·cosθ j , m represents the constraint margin parameter, cosθ j Represents the training sample x iThe cosine similarity of the corresponding j-th category, j = 1, 2, 3, ..., K, K represents the number of sample categories, y i is the category label of the sample to be trained, t (k) Represents the training phase parameters corresponding to the k-th iteration training.

[0080] Among them, the parameters of the training phase can be determined by the following formula:

[0081] t (k) =αr (k) +(1-α)t (k-1) ;

[0082] Where α represents the momentum parameter, t (0) =0.

[0083] S270, iteratively training the training samples according to the adaptive difficulty indicator loss to obtain the training result of the face recognition model.

[0084] The technical solution of the embodiment of the present invention aims at the problem that the existing difficult example measurement cannot robustly and finely mine the sample difficulty information, and proposes a universal sample difficulty parameter for sample difficulty mining. Compared with the simple definition of difficult samples by other mining losses, the proposed sample difficulty parameter can adaptively indicate the sample difficulty and its changes with the model training process in numerical form, thereby improving the accuracy, robustness and objectivity of the sample difficulty indication. In addition, due to the universality of the sample difficulty parameter, it can also be used for offline difficult sample mining and data set denoising. Aiming at the problem that the existing mining loss cannot effectively distinguish samples of different difficulties, a face recognition model training method based on adaptive difficulty indicator loss is proposed, which integrates global and local sample difficulty information, can impose more unique constraints on each sample according to the training difficulty of the sample at the current stage, mine the sample difficulty information online, and introduce training stage information in the adaptive difficulty indicator loss. The training stage parameters are used to impose training stage-related constraints on the sample to improve the convergence speed, which can achieve a better balance between the convergence speed and the final performance of the model.

[0085] Embodiment 3

[0086] Figure 5 This is a schematic diagram of the structure of a face recognition model training device based on hard example mining provided in the third embodiment of the present invention. The device can execute the face recognition model training method based on hard example mining provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method. Figure 5 As shown, the device comprises:

[0087] The sample information acquisition module 310 is used to acquire the sample information of the sample to be trained;

[0088] A sample difficulty parameter determination module 320, configured to determine a sample difficulty influencing factor according to the sample information, and determine a sample difficulty parameter of the sample to be trained according to the sample difficulty influencing factor;

[0089] The face recognition model training module 330 is used to iteratively train the samples to be trained based on the sample difficulty parameter to obtain the training result of the face recognition model.

[0090] Optionally, the sample difficulty influencing factor includes a sample loss influencing factor, a sample misclassification influencing factor and a sample imbalance influencing factor, wherein the sample loss influencing factor is used to indicate the contribution of sample loss to reducing sample difficulty during network training; the sample misclassification influencing factor is used to indicate the contribution of the number of times a sample is misclassified during training to reducing sample difficulty during network training; the sample imbalance influencing factor is used to indicate the contribution of the distribution balance of the category to which the sample belongs to reducing sample difficulty during network training;

[0091] Accordingly, the sample difficulty parameter determination module 320 is used to:

[0092] The sample difficulty parameter of the sample to be trained is determined based on the following sample difficulty indication function:

[0093]

[0094] Among them, H(x i ) represents the training sample x i The sample difficulty parameter, H1(x i ) represents the training sample x i The sample loss influence factor, H2(x i ) represents the training sample x i The sample misclassification influence factor, H3(x i ) represents the training sample x i , λ1 represents the weight coefficient of the sample loss factor, and λ2 represents the weight coefficient of the sample misclassification factor.

[0095] Optionally, the sample difficulty parameter determination module 320 is further configured to:

[0096] Based on the following formula, the sample loss impact factor is determined according to the loss value of the sample to be trained in the sample information:

[0097]

[0098] Among them, H1(xi ) represents the training sample x i The sample loss influence factor, α1 represents the first preset hyperparameter, α1>0, Represents the training sample x i The loss value in the current iteration, L max and L min Indicates the maximum loss value and the minimum loss value of the sample in the current iteration.

[0099] Optionally, the sample difficulty parameter determination module 320 is further configured to:

[0100] Based on the following formula, the sample misclassification influence factor is determined according to the cosine similarity of the samples to be trained in the sample information:

[0101]

[0102] Among them, H2(x i ) represents the training sample x i The sample misclassification influence factor is is the sample x to be trained i The weights in round a of training, and y i is the category label of the sample to be trained, cosθ j Represents the training sample x i The corresponding cosine similarity of the j-th category, j = 1, 2, 3, ..., K, K represents the number of sample categories.

[0103] Optionally, the sample difficulty parameter determination module 320 is further configured to:

[0104] Based on the following formula, the sample imbalance impact factor is determined according to the number of samples in the category to which the to-be-trained samples in the sample information belong:

[0105]

[0106] Among them, H3(x i ) represents the training sample x i The sample imbalance impact factor is Represents the training sample x i The number of samples belonging to the category, Represents the training sample x i The distribution estimation result of the number of samples belonging to the category.

[0107] Optionally, the face recognition model training module 330 is used to:

[0108] Based on the following formula, an adaptive difficulty indicator loss is determined according to the sample difficulty parameter, and the to-be-trained sample is iteratively trained according to the adaptive difficulty indicator loss;

[0109]

[0110] Among them, L represents the training sample x i The adaptive difficulty indicator loss of , N represents the number of samples to be trained in the current batch, s represents the second preset hyperparameter, N(t (k) ,H(x i ), cosθ j )=I(t (k) ,H(x i ), cosθ j )·cosθ j , m represents the constraint margin parameter, cosθ j Represents the training sample x i The corresponding cosine similarity of the jth class, j = 1, 2, 3, ..., K, K represents the number of sample categories, yi is the category label of the sample to be trained, t (k) Represents the training phase parameters corresponding to the k-th iteration training.

[0111] Optionally, the training phase parameters are determined by the following formula:

[0112] t (k) =αr (k) +(1-α)t (k-1) ;

[0113] Where α represents the momentum parameter, t (0) =0.

[0114] A face recognition model training device based on hard example mining provided by an embodiment of the present invention can execute a face recognition model training method based on hard example mining provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.

[0115] Embodiment 4

[0116] Figure 6A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0117] like Figure 6 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0118] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0119] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The processor 11 executes the various methods and processes described above, such as a face recognition model training method based on hard example mining.

[0120] In some embodiments, the face recognition model training method based on hard example mining can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the face recognition model training method based on hard example mining described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the face recognition model training method based on hard example mining in any other appropriate manner (for example, by means of firmware).

[0121] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0122] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0123] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0124] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0125] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0126] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0127] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0128] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A face recognition model training method based on hard example mining, characterized in that: The method comprises: Obtain sample information of samples to be trained; Determining a sample difficulty influencing factor according to the sample information, and determining a sample difficulty parameter of the sample to be trained according to the sample difficulty influencing factor; The samples to be trained are iteratively trained based on the sample difficulty parameter to obtain a training result of the face recognition model.

2. The method according to claim 1, characterized in that The sample difficulty influencing factors include sample loss influencing factors, sample misclassification influencing factors and sample imbalance influencing factors, wherein the sample loss influencing factors are used to indicate the contribution of sample loss to reducing sample difficulty during network training; the sample misclassification influencing factors are used to indicate the contribution of the number of times a sample is misclassified during training to reducing sample difficulty during network training; the sample imbalance influencing factors are used to indicate the contribution of the distribution balance of the sample category to reducing sample difficulty during network training; Correspondingly, determining the sample difficulty parameter of the sample to be trained according to the sample difficulty influencing factor includes: The sample difficulty parameter of the sample to be trained is determined based on the following sample difficulty indication function: Among them, H(x i ) represents the training sample x i The sample difficulty parameter, H1(x i ) represents the training sample x i The sample loss influence factor, H2(x i ) represents the training sample x i The sample misclassification influence factor, H3(x i ) represents the training sample x i , λ1 represents the weight coefficient of the sample loss factor, and λ2 represents the weight coefficient of the sample misclassification factor.

3. The method according to claim 2, characterized in that Determining a sample loss influencing factor according to the sample information includes: Based on the following formula, the sample loss impact factor is determined according to the loss value of the sample to be trained in the sample information: Among them, H1(x i ) represents the training sample x i The sample loss influence factor, α1 represents the first preset hyperparameter, α1>0, Represents the training sample x i The loss value in the current iteration, L max and L min Indicates the maximum loss value and the minimum loss value of the sample in the current iteration.

4. The method according to claim 2, characterized in that: Determining a sample misclassification influencing factor according to the sample information includes: Based on the following formula, the sample misclassification influence factor is determined according to the cosine similarity of the samples to be trained in the sample information: Among them, H2(x i ) represents the training sample x i The sample misclassification influence factor is is the sample x to be trained i The weights in round a of training, and y i is the category label of the sample to be trained, cosθ j Represents the training sample x i The corresponding cosine similarity of the j-th category, j = 1, 2, 3, ..., K, K represents the number of sample categories.

5. The method according to claim 2, characterized in that: Determining a sample imbalance influencing factor according to the sample information includes: Based on the following formula, the sample imbalance impact factor is determined according to the number of samples in the category to which the to-be-trained samples in the sample information belong: Among them, H3(x i ) represents the training sample x i The sample imbalance impact factor is Represents the training sample x i The number of samples belonging to the category, Represents the training sample x i The distribution estimation result of the number of samples belonging to the category.

6. The method according to claim 2, characterized in that Iteratively training the sample to be trained based on the sample difficulty parameter includes: Based on the following formula, an adaptive difficulty indicator loss is determined according to the sample difficulty parameter, and the to-be-trained sample is iteratively trained according to the adaptive difficulty indicator loss; Among them, L represents the training sample x i The adaptive difficulty indicator loss of , N represents the number of samples to be trained in the current batch, s represents the second preset hyperparameter, N(t (k) ,H(x i ),cosθ j )=I(t (k) ,H(x i ),cosθ j )·cosθ j , m represents the constraint margin parameter, cosθ j Represents the training sample x i The corresponding cosine similarity of the jth class, j = 1, 2, 3, ..., K, K represents the number of sample categories, y i is the category label of the sample to be trained, t (k) Represents the training phase parameters corresponding to the k-th iteration training.

7. The method according to claim 6, characterized in that The training phase parameters are determined by the following formula: t (k) =αr (k) +(1-α)t (k-1) ; Where α represents the momentum parameter, t (0) =0.

8. A face recognition model training device based on hard example mining, characterized in that: The device comprises: A sample information acquisition module is used to obtain sample information of samples to be trained; A sample difficulty parameter determination module, used to determine a sample difficulty influencing factor according to the sample information, and determine a sample difficulty parameter of the sample to be trained according to the sample difficulty influencing factor; The face recognition model training module is used to iteratively train the samples to be trained based on the sample difficulty parameter to obtain the training result of the face recognition model.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the face recognition model training method based on hard example mining as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the face recognition model training method based on hard example mining according to any one of claims 1 to 7 when executed.

Citation Information

Cited By

  • Text data processing method and device, storage medium and electronic equipment

    CN120821828A