A face recognition method for machine learning of non - identically - distributed data

Through the combination of convolutional neural network and attention mechanism, multi-stage training is carried out for non-distributed data, which solves the robustness of the face recognition model under light and occlusion, and achieves efficient recognition of makeup and occlusion.

CN119693989BActive Publication Date: 2025-07-11SOUTH CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510178689.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-07-11
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

Existing facial recognition technologies have challenges in model generalization capabilities and privacy protection under non-distributed data. Especially in the field of facial recognition, data distribution varies greatly, labeled data is scarce, and computational costs are high, and the model is not robust to lighting and occlusion.

Method used

The convolutional neural network model is adopted, attention mechanism and feature extraction layer are added, incremental learning is performed by adding perturbation data, and face data after makeup is trained for multi-stage training. The attention mechanism and loss function are used to optimize model parameters to enhance the robustness of lighting and occlusion.

Benefits of technology

It improves the robustness of the face recognition model for lighting and occlusion, adapts to scenes and characters changes, maintains high accuracy and generalization capabilities, and is suitable for non-distributed data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693989B_ABST
    Figure CN119693989B_ABST
Patent Text Reader

Abstract

The present invention provides a face recognition method for machine learning of non - identically - distributed data, including performing first - stage training on a face recognition model using a training set to obtain a first classification set. Perturbations are added to the training data in the training set to obtain supplementary data. Based on the attention mechanism, face features in the supplementary data are extracted, and the face recognition model is subjected to second - stage training to obtain a second classification set. The post - makeup face dataset is input into the face recognition model, and with the similarity between the output third classification set and the second classification set as a constraint term, the face recognition model is subjected to third - stage training so that the face recognition model can pay attention to more subtle change regions, obtaining a trained model, and the trained model is used to identify the user's identity. The trained model after the above three - stage training can adapt to changes in scenarios and people, will not affect the face recognition accuracy due to the introduction of new data, and the trained model is robust to data perturbations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of face recognition technology, and in particular to a face recognition method for machine learning of non-i.i.d. (independent and identically distributed) data. Background Art

[0002] In the field of machine learning of non-i.i.d. data, especially in face recognition technology, the existing technologies mainly include federated learning and federated clustering. Federated learning, as a distributed machine learning method, aims to solve the privacy and scalability problems of centralized methods, allowing multiple clients to collaboratively train a model while protecting data privacy. However, federated learning faces challenges in dealing with non-independent and identically distributed data. Especially in the field of face recognition, the data distributions of different users vary greatly, resulting in limited model generalization ability and personalization effect. Federated clustering creates a dedicated global model for groups of similar users, thus achieving better personalization effects. However, federated clustering relies on the availability of labeled data for each client, which is often difficult to achieve in practical applications, especially in face recognition scenarios where the cost of obtaining labeled data is high. In addition, the existing technologies have problems in aspects such as privacy protection, scarcity of labeled data, and model efficiency. The challenge of privacy protection lies in how to achieve effective data collaborative training while protecting user privacy. The scarcity of labeled data limits the training effect and accuracy of the model, especially in face recognition tasks that require a large amount of high-quality labeled data. Finally, the existing semi-supervised federated learning methods may require a large number of active learning queries, increasing the computational cost and communication cost of model training.

[0003] In security surveillance, face recognition systems are used to identify and track people entering the surveillance area. Since surveillance systems usually face multiple challenges, such as changes in lighting, angle, facial expression, and the population over different time periods, the system needs to be highly robust and can continuously learn new data to adapt to parameter changes. In addition, the surveillance system may receive a large amount of data, and the newly added data may disrupt the existing features, that is, the new features are quite different from the existing features, resulting in the face recognition model being unable to be compatible with the new dataset and the existing dataset. Moreover, when the dataset has interferences, such as the dataset includes faces with makeup, and faces under low light and occlusion conditions, the face recognition model will be perturbed by the input, making the face recognition model not very robust.

[0004] Therefore, there is a need for a face recognition method that can continuously update the model to adapt to changes in scenarios and people, and is applicable to face data with interferences, such as face data under low light and with facial occlusion, that is, robust to data perturbations. Summary of the Invention

[0005] To overcome the problems existing in the related art, the object of the present invention is to provide a face recognition method for machine learning of non - identically - distributed data, which can continuously update the model to adapt to changes in scenarios and people, and is applicable to face data with interference, such as low - light and face - occluded face data, that is, it is robust to data perturbations.

[0006] A face recognition method for machine learning of non - identically - distributed data, comprising:

[0007] Obtain a training set, and perform the first - stage training on a face recognition model using the training set to obtain a first classification set;

[0008] Add perturbations to the training data in the training set to obtain supplementary data; extract face features in the supplementary data based on an attention mechanism;

[0009] Perform the second - stage training on the face recognition model using the face features in the supplementary data to obtain a second classification set;

[0010] Input the post - makeup face dataset into the face recognition model, and perform the third - stage training on the face recognition model with the similarity between the output third classification set and the second classification set as a constraint term;

[0011] If the face recognition performance of the face recognition model meets the expectation, stop the third - stage training to obtain a trained model;

[0012] Use the trained model to identify the user's identity.

[0013] In a preferred technical solution of the present invention, before extracting the face features in the supplementary data based on the attention mechanism, it further includes:

[0014] Use a convolutional neural network model as the face recognition model;

[0015] Add an attention module between the feature extraction module and the output module of the face recognition model.

[0016] In a preferred technical solution of the present invention, adding an attention module between the feature extraction module and the output module of the face recognition model includes:

[0017] Add N feature extraction layers between the feature extraction module and the output module of the convolutional neural network model;

[0018] Or add a Transformer model between the feature extraction module and the output module of the convolutional neural network model.

[0019] In a preferred technical solution of the present invention, adding perturbations to the training data in the training set to obtain supplementary data includes:

[0020] Converting the training data in the training set into a plurality of training vectors; wherein, the training data is in matrix form;

[0021] Randomly generating a small vector;

[0022] Adding the small vector to each of the training vectors to obtain a plurality of perturbed vectors;

[0023] Converting the perturbed vectors into matrix form to obtain supplementary data.

[0024] In a preferred technical solution of the present invention, using the face features in the supplementary data to perform a second-stage training on the face recognition model to obtain a second classification set includes:

[0025] Inputting the face features in the supplementary data into the face recognition model to output recognition labels;

[0026] Comparing the recognition labels with the true labels and calculating a second loss function;

[0027] According to the second loss function, updating the model parameters of the face recognition model by means of backpropagation;

[0028] Forming all the recognition labels into a second classification set.

[0029] In a preferred technical solution of the present invention, performing a third-stage training on the face recognition model includes:

[0030] Using a third loss function to measure the difference between the second classification set and the third classification set;

[0031] Adjusting the model parameters of the face recognition model according to the third loss function.

[0032] In a preferred technical solution of the present invention, if the face recognition performance of the face recognition model meets the expectation, stopping the third-stage training to obtain a trained model includes:

[0033] If the accuracy rate of the face recognition model is greater than or equal to an accuracy rate threshold and the recall rate is less than a recall rate threshold, then detecting whether the face recognition model is robust to occlusion and illumination;

[0034] If the face recognition model is robust to occlusion and illumination, then stopping the third-stage training to obtain a trained model.

[0035] In a preferred technical solution of the present invention, calculating the second loss function includes:

[0036] Calculate the second loss function according to the following formula:

[0037] ;

[0038] Wherein, t is the true label, x is the training data, and x' is the supplementary data; f(x) is the first predicted label of the face recognition model for predicting the training data, and f(x') is the second predicted label of the face recognition model for predicting the supplementary data; ||f(x)-t||2 is the standard classification loss, ||f(x')-f(x)||2 is the consistency loss, λ(t) is the first dynamic weight function, and L2(x, x', t) is the second loss function.

[0039] In a preferred technical solution of the present invention, the measuring the difference between the second classification set and the third classification set by using the third loss function includes:

[0040] ;

[0041] Wherein, β(t) is the second dynamic weight function, L3(x', x'', t) is the third loss function, f(x') is the second predicted label of the face recognition model for predicting the supplementary data, f(x'') is the third predicted label of the face recognition model for predicting the post-makeup face dataset, t is the true label, ||||2 represents the two-norm operation, and the true label t includes the actual labels of the training data and the post-makeup face dataset.

[0042] In a preferred technical solution of the present invention, the first-stage training of the face recognition model by using the training set to obtain the first classification set includes:

[0043] Input the training set into the face recognition model to extract the global features in the training set;

[0044] Output the first classification set according to the global features;

[0045] Calculate the cross-entropy loss function between the first classification set and the true label, and update the model parameters of the face recognition model according to the value of the cross-entropy loss function.

[0046] The beneficial effects of the present invention are:

[0047] The face recognition method for machine learning of non - identically distributed data provided by the present invention includes obtaining a training set, performing the first - stage training on a face recognition model using the training set to obtain a first classification set. During the first - stage training, a face recognition model based on a convolutional neural network is used to extract face features, so that the face recognition model can recognize conventional face features such as eyes, nose, and mouth, etc. Perturbations are added to the training data in the training set to obtain supplementary data; face features in the supplementary data are extracted based on an attention mechanism. The face features in the supplementary data are used to perform the second - stage training on the face recognition model to obtain a second classification set. On the basis of the first - stage training, with the perturbed supplementary data, an incremental learning method is used to enhance the robustness of the face recognition model to illumination and occlusion. The post - makeup face dataset is input into the face recognition model, and the similarity between the output third classification set and the second classification set is used as a constraint term to perform the third - stage training on the face recognition model. If the face recognition performance of the face recognition model meets the expectations, the third - stage training is stopped to obtain a trained model, and the trained model is used to recognize the user identity. The third - stage training is to add post - makeup face features on the basis of the existing face recognition model, so that the face recognition model can pay attention to more subtle change areas, such as makeup marks on the eyes and mouth, thereby increasing the generalization ability of the model. The trained model after the above three - stage training can adapt to changes in scenarios and people, will not affect the face recognition accuracy due to the introduction of new data, and the trained model is applicable to face data with interference such as low - illumination and face - occluded face data, that is, it is robust to data perturbations. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flowchart of the face recognition method for machine learning of non - identically distributed data of the present invention;

[0049] Figure 2 is a flowchart of the first - stage training of the face recognition model using the training set of the present invention;

[0050] Figure 3 is a flowchart of adding perturbations to the training data in the training set of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention will be more thorough and complete, and the scope of the present invention can be fully conveyed to those skilled in the art.

[0052] Example 1

[0053] AsFigure 1 As shown in the figure, this embodiment provides a face recognition method for machine learning of non-i.i.d. data, including:

[0054] S1: Obtain a training set, and use the training set to perform the first-stage training on the face recognition model to obtain a first classification set.

[0055] S2: Add perturbations to the training data in the training set to obtain supplementary data; extract the face features in the supplementary data based on the attention mechanism.

[0056] S3: Use the face features in the supplementary data to perform the second-stage training on the face recognition model to obtain a second classification set.

[0057] S4: Input the post-makeup face dataset into the face recognition model, and use the similarity between the output third classification set and the second classification set as a constraint term to perform the third-stage training on the face recognition model.

[0058] S5: If the face recognition performance of the face recognition model meets the expectations, stop the third-stage training to obtain a trained model.

[0059] S6: Use the trained model to identify the user's identity.

[0060] The training set uses a conventional face dataset, that is, a face dataset with normal lighting and no occlusion. Add light adjustment and facial occlusion to the training data in the training set, and the training data can also be rotated and cropped, so as to enhance the robustness of the face recognition model to light and occlusion in the second-stage training. The post-makeup face dataset includes the face features after makeup. In this stage, the self-attention mechanism can be used to make the face recognition model pay attention to more subtle change regions, such as the makeup marks on the eyes and mouth.

[0061] As Figure 2 shown, the step of using the training set to perform the first-stage training on the face recognition model to obtain a first classification set includes:

[0062] S12: Input the training set into the face recognition model to extract the global features in the training set.

[0063] S13: Output a first classification set according to the global features.

[0064] S14: Calculate the cross-entropy loss function between the first classification set and the true label, and update the model parameters of the face recognition model according to the value of the cross-entropy loss function.

[0065] Before step S12, it further includes step S11: obtaining a training set. The training set is input into a face recognition model, which includes a convolutional layer, a fully connected layer, and an output layer. The convolutional layer is used to extract face features from the training set, and the fully connected layer is used to integrate the face features extracted by different convolutional layers. The face features extracted in the first-stage training include eyes, nose, mouth, etc. In the first-stage training, the cross-entropy loss function, i.e., the first loss function, is used to measure the difference between the recognition labels output by the output layer and the true labels, and the face recognition model is trained according to the difference between the labels.

[0066] As Figure 3 shown, adding perturbations to the training data in the training set to obtain supplementary data includes:

[0067] S21: converting the training data in the training set into multiple training vectors; wherein, the training data is in matrix form.

[0068] S22: randomly generating a small vector.

[0069] S23: adding the small vector to each of the training vectors to obtain multiple perturbed vectors.

[0070] S24: converting the perturbed vectors into matrix form to obtain supplementary data.

[0071] In the second-stage training, when given slightly different inputs, the face recognition model is required to produce consistent predictions, thereby reducing the sensitivity of the face recognition model to small perturbations in the input. During the second-stage training process, adversarial perturbations are added to prevent the face recognition model from overfitting when facing changing data distributions.

[0072] Adding a small vector to each training vector, and the small vector can represent illumination, occlusion, rotation, translation, and cropping. By using a random generation method, multiple perturbed vectors can correspond to different scenarios. Adding perturbations to the training data and using the supplementary data to perform the second-stage training on the face recognition model to avoid the model relying too much on specific data patterns. The perturbations are mainly used to simulate image noise, illumination changes, and facial occlusions in the real environment, making the learning process of the face recognition model diverse.

[0073] The similarity between the perturbation vector and the corresponding training vector is greater than 0.6 to ensure the consistency of the identity of the person in the training vector, that is, the person corresponding to the training vector does not change. On the premise of fixing the size of the occluded area, the light distribution and the position of the occluded area can be changed simultaneously; or on the premise of fixing the position of the occluded area, the light distribution and the size of the occluded area can be changed simultaneously; or on the premise of fixing the position and size of the occluded area, the light distribution and the shape of the occluded area can be changed simultaneously, for example, changing the occluded area from a square to a rectangle and then from a rectangle to a circle.

[0074] The light distribution is described by the following formula:

[0075] K(z)=J(z)t(z)+B(z)(1 - t(z));

[0076] Where K(z) is the light intensity of the z-th pixel of the current face image in the training data. The abscissa of the z-th pixel is m1, and the ordinate is n1. t(z) is the transmittance corresponding to the z-th pixel in the transmittance spectrum, J(z) is the distant-view light at the z-th pixel, and B(z) is the near-view light at the z-th pixel. The near-view light refers to the light within a spherical region with a radius of R1 centered on the target face. The distant-view light refers to the light outside the spherical region with a radius of R1 centered on the target face.

[0077] Since the light source of the near-view light is relatively close to the target face, the near-view light mainly generates reflections in the area near the target face. The present invention assumes that the near-view light only undergoes reflection and the distant-view light only undergoes transmission. Therefore, the product of t(z) and J(z) is used to describe the influence of the distant-view light on the target face, that is, the light intensity of the distant light source irradiating the target face. 1 - t(z) represents the reflectivity corresponding to the z-th pixel, and the product of B(z) and (1 - t(z)) is used to describe the influence of the near-view light on the target face, that is, the light intensity of the near light source irradiating the target face. Combining the light intensity after the near-view light source is reflected and the light intensity after the distant-view light source is transmitted, the light intensity of the z-th pixel of the current face image in the training data can be obtained.

[0078] After the third-stage training, the face recognition model of the present invention is applicable to conventional face data without perturbations and makeup, face data with perturbations, and face data after makeup, making the face recognition model of the present invention have strong robustness to non-homogeneous distribution data.

[0079] The face recognition method for machine learning of non - identically - distributed data provided in this embodiment includes obtaining a training set, performing the first - stage training on the face recognition model using the training set to obtain a first classification set. During the first - stage training, a face recognition model based on a convolutional neural network is used to extract face features so that the face recognition model can recognize conventional face features such as eyes, nose, and mouth, etc. Perturbations are added to the training data in the training set to obtain supplementary data; face features in the supplementary data are extracted based on an attention mechanism. The face features in the supplementary data are used to perform the second - stage training on the face recognition model to obtain a second classification set. On the basis of the first - stage training, with the perturbed supplementary data, an incremental learning method is used to enhance the robustness of the face recognition model to illumination and occlusion. The post - makeup face dataset is input into the face recognition model, and taking the similarity between the output third classification set and the second classification set as a constraint term, the third - stage training is performed on the face recognition model. If the face recognition performance of the face recognition model meets the expectations, the third - stage training is stopped to obtain a trained model, and the trained model is used to recognize the user identity. The third - stage training is based on the existing face recognition model and adds post - makeup face features, enabling the face recognition model to pay attention to more subtle change regions, such as makeup marks around the eyes and mouth, thereby increasing the generalization ability of the model. The trained model after the above three - stage training can adapt to changes in scenarios and people, will not affect the face recognition accuracy due to the introduction of new data, and the trained model is applicable to face data with interference such as low - illumination and face - occluded face data, that is, it is robust to data perturbations.

[0080] Embodiment 2

[0081] This embodiment provides a face recognition method for machine learning of non - identically - distributed data. This embodiment only describes the differences from Embodiment 1. Before extracting the face features in the supplementary data based on the attention mechanism, it further includes:

[0082] S11’: Use a convolutional neural network model as the face recognition model.

[0083] S12’: Add an attention module between the feature extraction module and the output module of the face recognition model.

[0084] Adding the attention module between the feature extraction module and the output module of the face recognition model includes:

[0085] S121’: Add N feature extraction layers between the feature extraction module and the output module of the convolutional neural network model.

[0086] S122’: Or add a Transformer model between the feature extraction module and the output module of the convolutional neural network model.

[0087] The attention mechanism is added to the neural network model, enabling the face recognition model to focus on more subtle change regions, such as makeup marks on the eyes and mouth. When calculating attention using the traditional self-attention mechanism, only the global relationships in the input are considered. However, in face recognition, local changes may occur in the facial features after makeup. Therefore, the multi-head self-attention mechanism is introduced to enable the face recognition model to learn the relationships between local features and global features from multiple perspectives.

[0088] Furthermore, through the multi-head attention mechanism, the face recognition model can learn the relationships between local regions such as the eyes, mouth, and nose, while maintaining attention to the global facial features.

[0089] A Transformer model is added between the feature extraction module and the output module to gradually fuse local features such as the eyes and mouth, and global features such as the overall facial contour. By dynamically adjusting the attention weights, the face recognition model can better handle face recognition scenarios under makeup and occlusion. Combining the multi-head attention mechanism and local feature fusion can enhance the face recognition model's ability to focus on complex facial features such as makeup faces, making the face recognition model more robust and accurate during the face recognition process.

[0090] If the face recognition performance of the face recognition model meets the expectations, the third-stage training is stopped to obtain the trained model, including:

[0091] S51: If the accuracy rate of the face recognition model is greater than or equal to the accuracy rate threshold and the recall rate is less than the recall rate threshold, then it is detected whether the face recognition model is robust to occlusion and illumination.

[0092] S52: If the face recognition model is robust to occlusion and illumination, the third-stage training is stopped to obtain the trained model.

[0093] The accuracy rate of face recognition refers to the proportion of samples that are correctly predicted as positive samples among all samples predicted as positive samples. The calculation method of the accuracy rate is to divide the number of positive samples predicted as positive by the total number of samples predicted as positive samples. Among them, the total number of samples predicted as positive samples is the sum of the number of positive samples correctly predicted as positive and the number of negative samples mispredicted as positive. The recall rate of face recognition refers to the proportion of positive samples that are correctly identified among all positive samples. The higher the recall rate, the more positive samples the face recognition model can identify. The calculation method of the recall rate is to divide the number of correctly identified faces by the total number of positive samples. The total number of positive samples is the sum of the number of positive samples predicted as positive and the number of positive samples mispredicted as negative. Positive samples are images containing the target object, and negative samples are images not containing the target object.

[0094] Before extracting the face features in the supplementary data based on the attention mechanism in this embodiment, it further includes using a convolutional neural network model as a face recognition model and adding an attention module between the feature extraction module and the output module of the face recognition model. The attention module is an N-layer feature extraction layer or a Transformer model. By dynamically adjusting the attention weights, the face recognition model can better handle the face recognition scenarios under makeup and occlusion. Combining the multi-head attention mechanism and local feature fusion can improve the attention ability of the face recognition model to complex face features such as faces with makeup, making the face recognition model more robust and accurate in the face recognition process.

[0095] Embodiment 3

[0096] This embodiment provides a face recognition method for machine learning for non-i.i.d. data. This embodiment only describes the differences from Embodiment 1. Using the face features in the supplementary data to perform the second-stage training on the face recognition model to obtain a second classification set, including:

[0097] S31: Input the face features in the supplementary data into the face recognition model and output an identification label.

[0098] S32: Compare the identification label with the true label and calculate the second loss function.

[0099] S33: According to the second loss function, use the backpropagation method to update the model parameters of the face recognition model.

[0100] S34: Combine all the identification labels to form a second classification set.

[0101] The supplementary data is the data obtained by adding perturbations to the training data in the training set. The forms of perturbations include increasing the light intensity, decreasing the light intensity, increasing facial occlusion, rotating the image, translating the image, and cropping the image. In an ideal situation, when the input image is perturbed, the output classification label is still the same as the true label. Therefore, the present invention uses the second loss function to train the face recognition model. The second loss function is composed of a standard classification loss and a consistency loss, which can improve the robustness of the face recognition model in face recognition when the face data is perturbed.

[0102] The calculation of the second loss function includes:

[0103] Calculate the second loss function according to the following formula:

[0104] ;

[0105] Among them, t is the true label, x is the training data, and x' is the supplementary data; f(x) is the first predicted label of the face recognition model for predicting the training data, and f(x') is the second predicted label of the face recognition model for predicting the supplementary data; ||f(x) - t||2 is the standard classification loss, ||f(x') - f(x)||2 is the consistency loss, λ(t) is the first dynamic weight function, and L2(x, x', t) is the second loss function.

[0106] The standard classification loss is used to measure the difference between the first predicted label of the face recognition model for predicting the training data and the true label. The consistency loss is used to measure the difference between the second predicted label of the face recognition model for predicting the supplementary data and the first predicted label of the face recognition model for predicting the training data. By combining the standard classification loss and the consistency loss, it is possible to reduce the impact of perturbations on face recognition while enabling the face recognition model to accurately identify the user's identity. The value of the first dynamic weight function is variable and is used to adjust the proportion of the standard classification loss and the consistency loss.

[0107] After the face recognition model undergoes the first stage of training, face data with illumination changes and occlusions is gradually introduced. To avoid catastrophic forgetting, the present invention adopts an incremental learning method, enabling the face recognition model to maintain its memory of regular face data while learning new data.

[0108] During each round of training in the second stage of training, the difference between the training data and the supplementary data is calculated to ensure that the feature learning of the new data does not disrupt the feature learning of the old data. During the second stage of training, when there is less supplementary data with perturbations, the value of the first dynamic weight function gradually increases to enhance the adaptability of the face recognition model to perturbations such as illumination changes and occlusions.

[0109] The third stage of training the face recognition model includes:

[0110] Using a third loss function to measure the difference between the second classification set and the third classification set;

[0111] Adjusting the model parameters of the face recognition model according to the third loss function.

[0112] The using a third loss function to measure the difference between the second classification set and the third classification set includes:

[0113] ;

[0114] Among them, β(t) is the second dynamic weight function, L3(x’, x’’, t) is the third loss function, f(x’) is the second prediction label of the face recognition model for predicting supplementary data, f(x’’) is the third prediction label of the face recognition model for predicting the post-makeup face dataset, t is the true label, and ||||2 represents the two-norm operation.

[0115] The key to dynamic consistency regularization is to automatically adjust the weight of the consistency loss according to the stability requirements of the model at different stages. During the third-stage training, as the face recognition model gradually stabilizes, the value of the second dynamic weight function is continuously reduced, thereby reducing the proportion of the consistency loss and focusing on optimizing the classification ability.

[0116] The third loss function not only focuses on the consistency of the perturbed data but also considers the distribution differences between the data. Especially in the face recognition scenario of non-i.i.d. data, it dynamically adjusts the intensity of the consistency loss to prevent the supplementary data with perturbations from affecting the original ability of the face recognition model, and can effectively alleviate the shift of the data distribution on the premise of introducing the post-makeup face dataset.

[0117] After the second-stage training, the face recognition model of the present invention already has good robustness to face data with illumination changes and occlusions. In the third-stage training, the post-makeup face dataset is added to enhance the adaptability of the face recognition model to post-makeup face images, enabling the face recognition model to handle complex face changes.

[0118] In this embodiment, as the dataset changes, the proportion between the standard classification loss and the consistency loss needs to be adjusted. During the second-stage training, since supplementary data with perturbations such as illumination changes and occlusions are introduced, the weight of the consistency loss, that is, the first dynamic weight function, is gradually increased, thereby enhancing the adaptability of the face recognition model to the perturbed data. After introducing the post-makeup face dataset, it may lead to insufficient attention of the face recognition model to non-critical regions. During the third-stage training, since the post-makeup face data has an increasing impact on the face recognition model, the weight of the consistency loss, that is, the second dynamic weight function, is continuously reduced. At this time, the face recognition model focuses on learning the post-makeup face data.

[0119] Embodiment 4

[0120] In an embodiment of the present application, a computer device is further provided. The computer device may be a server. Among them, the computer device includes a processor, a memory, a network interface, and a database connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection.

[0121] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the face recognition method for machine learning of non-i.i.d. data described in any one of Embodiments 1-3. It can be understood that the computer-readable storage medium in this embodiment may be a volatile readable storage medium or a non-volatile readable storage medium.

[0122] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, device, article or method including that element.

[0123] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present application.

Claims

1. A face recognition method for machine learning of non - identically - distributed data, characterized in that, Including: Obtain a training set, and perform first-stage training on the face recognition model using the training set to obtain a first classification set; Add perturbations to the training data in the training set to obtain supplementary data; extract face features from the supplementary data based on the attention mechanism; Perform second-stage training on the face recognition model using the face features in the supplementary data to obtain a second classification set; Input the post-makeup face dataset into the face recognition model, and use the similarity between the output third classification set and the second classification set as a constraint term to perform third-stage training on the face recognition model; If the face recognition performance of the face recognition model meets the expectations, stop the third-stage training to obtain a trained model; Use the trained model to identify the user's identity; The performing second-stage training on the face recognition model using the face features in the supplementary data to obtain a second classification set includes: Input the face features in the supplementary data into the face recognition model to output recognition labels; Compare the recognition labels with the true labels and calculate the second loss function; Update the model parameters of the face recognition model in a backpropagation manner according to the second loss function; Form all the recognition labels into a second classification set; The performing third-stage training on the face recognition model includes: Use a third loss function to measure the difference between the second classification set and the third classification set; Adjust the model parameters of the face recognition model according to the third loss function; during the third-stage training, continuously reduce the weight of the consistency loss; The calculating the second loss function includes: The second loss function consists of a standard classification loss and a consistency loss. The standard classification loss is used to measure the difference between the first predicted label of the face recognition model for predicting the training data and the true label; the consistency loss is used to measure the difference between the second predicted label of the face recognition model for predicting the supplementary data and the first predicted label of the face recognition model for predicting the training data; During the second-stage training, gradually increase the weight of the consistency loss to enhance the adaptability of the face recognition model to the perturbed data.

2. The face recognition method for machine learning of non - identically - distributed data according to claim 1, wherein, Before extracting the face features from the supplementary data based on the attention mechanism, it further includes: Use a convolutional neural network model as the face recognition model; Add an attention module between the feature extraction module and the output module of the face recognition model.

3. The face recognition method for machine learning of non - identically - distributed data according to claim 2, wherein The adding an attention module between the feature extraction module and the output module of the face recognition model includes: Add N layers of feature extraction layers between the feature extraction module and the output module of the convolutional neural network model; Or add a Transformer model between the feature extraction module and the output module of the convolutional neural network model.

4. The face recognition method for machine learning of non - identically - distributed data according to claim 1, characterized in that, The adding perturbations to the training data in the training set to obtain supplementary data includes: Convert the training data in the training set into multiple training vectors; where the training data is in matrix form; Randomly generate tiny vectors; Add the tiny vectors to each of the training vectors to obtain multiple perturbed vectors; the similarity between the perturbed vectors and the corresponding training vectors is greater than 0.6; Convert the perturbation vector into a matrix form to obtain supplementary data; the tiny vector represents illumination, occlusion, rotation, translation, and cropping.

5. The face recognition method for machine learning of non - identically - distributed data according to claim 1, wherein If the face recognition performance of the face recognition model meets the expectations, stop the third-stage training to obtain the trained model, including: If the accuracy of the face recognition model is greater than or equal to the accuracy threshold and the recall rate is less than the recall rate threshold, detect whether the face recognition model is robust to occlusion and illumination; If the face recognition model is robust to occlusion and illumination, stop the third-stage training to obtain the trained model.

6. The face recognition method for machine learning of non - identically - distributed data according to claim 1, characterized in that, Perform the first-stage training on the face recognition model using the training set to obtain the first classification set, including: Input the training set into the face recognition model to extract the global features in the training set; Output the first classification set according to the global features; Calculate the cross-entropy loss function between the first classification set and the true labels, and update the model parameters of the face recognition model according to the value of the cross-entropy loss function.

Citation Information

Patent Citations

  • Model training method, face recognition method, device, equipment, medium and product

    CN114140862A

  • Fraud detection method and system for face makeup

    CN117935380A