Monitoring face training optimization method based on fully supervised and semi-supervised joint training

By adopting a fully supervised and semi-supervised joint training method in facial recognition technology, combined with ResNet50 network and dual model technology, the accuracy and generalization performance of face recognition are optimized, and the problem of insufficient facial recognition capabilities in complex scenarios is solved.

CN120014404APending Publication Date: 2025-05-16LINKER
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411945890.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing face recognition technology is difficult to effectively recognize faces in complex scenarios, and due to the high cost of data labeling, it is difficult to obtain high-quality labeling data.

Method used

The monitoring face recognition optimization method based on full-supervised and semi-supervised joint training is adopted, and feature extraction and training is performed through ResNet50 network structure and dual models (model p and model g), combined with ArcFace loss function and OHEM difficult sample mining technology, the discriminant ability and generalization performance of the model are optimized.

Benefits of technology

It significantly improves the accuracy and generalization performance of face recognition, reduces dependence on a large amount of labeled data, and improves the robustness and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005213631500000051
    Figure BDA0005213631500000051
  • Figure BDA0005213631500000052
    Figure BDA0005213631500000052
  • Figure FDA0005213631490000021
    Figure FDA0005213631490000021
Patent Text Reader

Abstract

The invention discloses a face monitoring training optimization method based on fully-supervised and semi-supervised joint training. The method comprises the following steps: step 1, preprocessing all input image data; 2, initializing two models p and g; 3, generating a feature list of which the length is equal to the total number of the face IDs, wherein the feature list is used for storing the feature representation of each ID; step 4, sampling each training batch according to the number of all face IDs generated in the step 3; 5, carrying out feature similarity calculation; step 6, optimizing the model trained in the step 4 by using an ArcFace loss function; 7, dynamic updating of the list is achieved; and 8, calculating semi-supervised loss and full supervised loss. The monitored face training optimization method based on full-supervised and semi-supervised joint training aims at improving the performance and stability of a monitored face recognition model through key point alignment correction, dynamic feature list updating, cosine similarity optimization and difficult sample mining.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an industrial inspection industry and the field of face recognition, and more specifically to a monitoring face training optimization method based on full-supervision and semi-supervision joint training. Background Art

[0002] With the continuous development of face recognition technology in surveillance scenarios, it is particularly important to optimize the model's ability to recognize faces in complex scenarios. For example, in the prior art, there is a patent number 202011530619.5, titled "A student attention recognition method for online education scenarios optimized for CPU operations." It records a student attention recognition method optimized for CPU operations. When used, only a single student face image is required to complete attention recognition, which effectively reduces the algorithm performance overhead and the attention recognition calculation process is completed directly on the student's computer with a low CPU occupancy rate. This method mainly optimizes the face recognition algorithm by using the MTCNN face recognition model to perform face detection on each image in the face image dataset to obtain the face image I and key point detection to obtain the face key point L. However, due to the high cost of data annotation and the difficulty in obtaining high-quality annotated data in some scenarios, the existing face recognition technology is still insufficient in its ability to recognize faces in complex scenes. Summary of the invention

[0003] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a surveillance face recognition optimization method based on full supervision and semi-supervised joint training, which can effectively realize the recognition of faces in complex scenes.

[0004] To achieve the above object, the present invention provides the following technical solution: a monitoring face training optimization method based on full supervision and semi-supervised joint training, characterized in that it includes the following steps:

[0005] Step 1: Preprocess all input image data, align and correct facial key points, rotate the image to the frontal orientation to reduce the impact of posture differences, and use the ResNet50 network structure as the basic model;

[0006] Step 2: Initialize two models p and g. The parameters of model p will be updated through normal back propagation, while the parameters of model g will be updated through the moving average of model p, so that it retains the relatively stable feature extraction ability in the previous round of training.

[0007] Step 3: Generate a feature list with a length equal to the total number of face IDs to store the feature representation of each ID;

[0008] Step 4: Sample each training batch according to the number of all face IDs generated in step 3, randomly select a specific batchsize IDs from them, randomly select two images for each ID, and input them into two models p and g for training after data augmentation processing;

[0009] Step 5: Calculate feature similarity and destroy it immediately after calculation;

[0010] Step 6: Use the ArcFace loss function to optimize the model trained in step 4 to improve the model's discrimination ability;

[0011] Step 7: During the training process, randomly select the feature vectors of models p and g and add them to the feature list to achieve dynamic update of the list;

[0012] Step 8: Calculate the semi-supervised loss and the fully supervised loss, and use the calculated semi-supervised loss and the fully supervised loss to perform gradient updates to optimize the model parameters.

[0013] As a further improvement of the present invention, the feature list generated in step 3 may change randomly and dynamically in each epoch to avoid overfitting and sample bias.

[0014] As a further improvement of the present invention, the specific method of calculating the similarity in step 5 is: the temporary feature lists l' and l" respectively store the features generated in the current batch, and then calculate the similarity by the following formula:

[0015] sim p1,l′ =cos(p1,l ′

[0016] sim p2,l″ =cos(p2,l″

[0017] Among them, the temporary feature list is a copy list.

[0018] As a further improvement of the present invention, the specific method of using the ArcFace loss function in step 6 to optimize the trained model in step 4 is: using the ArcFace loss function to adjust the cosine similarity so that the inter-class distance is increased and the intra-class distance is reduced. The specific formula is as follows:

[0019] ArcFace=cos(θ+m)

[0020] Among them, θ is the original cosine similarity, and m is the preset angle margin.

[0021] As a further improvement of the present invention, the semi-supervised loss in step eight is calculated by the following method: based on cosine similarity, the similarity calculation result and the true label are used to calculate the cross entropy loss, and the average value is taken. The formula is as follows:

[0022]

[0023] As a further improvement of the present invention, the fully supervised loss is calculated by: calculating the loss by comparing the classification head output of models p and g with the true label, and strengthening the training of difficult samples by the difficult sample mining (OHEM) method, the formula is as follows:

[0024]

[0025] Then calculate the total loss function:

[0026] L total =L semi +L sup .

[0027] Beneficial effects of the present invention:

[0028] 1. Enhanced inter-class distinction and intra-class aggregation: Optimizing cosine similarity through ArcFace loss significantly improves inter-class distinction and intra-class consistency, and has higher recognition accuracy than traditional methods.

[0029] 2. Use unlabeled data to improve generalization capabilities: Semi-supervised learning reduces the reliance on large amounts of labeled data and improves the generalization performance of the model in monitoring scenarios where labeling is difficult.

[0030] 3. Moving average stable feature representation: A dual model is introduced, in which the parameters of one model are updated through moving average to ensure stable feature representation. Compared with traditional single model training, it has stronger robustness.

[0031] 4. Dynamic feature list enhances diversity: Random feature updates prevent feature concentration, improve model adaptability and reduce overfitting.

[0032] 5. OHEM difficult sample mining: Focusing on the training of difficult samples, enhancing the model's recognition ability in complex scenarios, it has more advantages than traditional methods. Reduce data scale and annotation costs: Semi-supervised methods reduce annotation requirements, reduce training costs, and improve the application value of monitoring scenarios. DETAILED DESCRIPTION

[0033] The present invention will be further described in detail with reference to the given embodiments below.

[0034] A surveillance face recognition optimization method based on full-supervision and semi-supervision joint training in this embodiment includes the following steps:

[0035] Data preprocessing: First, all the input images are aligned and corrected for facial key points, and the images are rotated to the front to reduce the impact of posture differences. The aligned images are used as model input, and the ResNet50 network structure is used as the basic model.

[0036] Model initialization: Initialize two models p and g, where the parameters of model p will be updated through normal back propagation, while the parameters of model g will be updated through the moving average of model p, so that it retains the relatively stable feature extraction capability in the previous round of training.

[0037] Feature list construction: Generate a feature list with a length equal to the total number of face IDs to store the feature representation of each ID. The list will change randomly and dynamically in each epoch to avoid overfitting and sample bias.

[0038] Batch Training: Each training batch is sampled according to the number of all face IDs, from which a specific batchsize IDs are randomly selected. Two images are randomly selected for each ID and input into two models p and g after data augmentation:

[0039] For example, two images image1 and image2 under a certain ID are input into models p and g respectively to generate 512-dimensional feature representations p1, p2, g1, g2.

[0040] In addition, the model also generates specific category outputs through the classification head for full supervision loss calculation

[0041] Feature similarity calculation: Temporary feature lists l' and l" store the features generated in the current batch for similarity calculation:

[0042] sim p1,l′ =cos(p1,l ′

[0043] sim p2,l″ =cos(p2,l″

[0044] Among them, the temporary feature list is a copy list and is destroyed immediately after the similarity is calculated.

[0045] ArcFace loss optimization: ArcFace loss function is used to adjust the cosine similarity, so that the distance between classes is increased and the distance within classes is reduced, thereby improving the discrimination ability of the model. The cosine similarity is optimized and adjusted by ArcFace, and the formula is as follows:

[0046] ArcFace=cos(θ+m)

[0047] Among them, θ is the original cosine similarity, and m is the preset angle margin.

[0048] Dynamic feature update: During the training process, the feature vectors of models p and g are randomly selected and added to the feature list to achieve dynamic update of the list.

[0049] The loss function consists of two parts:

[0050] Semi-supervised loss: Based on cosine similarity, the similarity calculation result and the true label are used to calculate the cross entropy loss and the average is taken. The formula is as follows:

[0051]

[0052] Fully supervised loss: The classification head output of models p and g is compared with the true label to calculate the loss, and the difficult sample mining (OHEM) method is used to strengthen the training of difficult samples. The formula is as follows:

[0053]

[0054] Total loss function:

[0055] L total =L semi +L sup

[0056] Gradient update: Use the above total loss to perform gradient update to optimize the model parameters.

[0057] In summary, the optimization method of this embodiment is significantly better than the traditional training method in monitoring face recognition tasks and has higher accuracy, generalization and practicality.

[0058] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A surveillance face training optimization method based on full-supervision and semi-supervision joint training, characterized in that: The steps include: Step 1: Preprocess all input image data, align and correct facial key points, rotate the image to the frontal orientation to reduce the impact of posture differences, and use the ResNet50 network structure as the basic model; Step 2: Initialize two models p and g. The parameters of model p will be updated through normal back propagation, while the parameters of model g will be updated through the moving average of model p, so that it retains the relatively stable feature extraction ability in the previous round of training. Step 3: Generate a feature list with a length equal to the total number of face IDs to store the feature representation of each ID; Step 4: Sample each training batch according to the number of all face IDs generated in step 3, randomly select a specific batch size of IDs from them, randomly select two images for each ID, and input them into two models p and g for training after data augmentation processing; Step 5: Calculate feature similarity and destroy it immediately after calculation; Step 6: Use the ArcFace loss function to optimize the model trained in step 4 to improve the model's discrimination ability; Step 7: During the training process, randomly select the feature vectors of models p and g and add them to the feature list to achieve dynamic update of the list; Step 8: Calculate the semi-supervised loss and the fully supervised loss, and use the calculated semi-supervised loss and the fully supervised loss to perform gradient updates to optimize the model parameters.

2. The monitoring face training optimization method based on full-supervision and semi-supervision joint training according to claim 1 is characterized in that: The feature list generated in step 3 will change randomly and dynamically in each epoch to avoid overfitting and sample bias.

3. The monitoring face training optimization method based on full-supervision and semi-supervision joint training according to claim 2 is characterized in that: The specific method of calculating the similarity in step 5 is: the temporary feature lists l' and l" respectively store the features generated in the current batch, and then calculate the similarity by the following formula: Sat. p1,l′ =basket ( p1,l ′) sim p2,l″ =cos ( p2,l″ ) Among them, the temporary feature list is a copy list.

4. The monitoring face training optimization method based on full-supervision and semi-supervision joint training according to claim 3 is characterized in that: The specific method of using the ArcFace loss function in step 6 to optimize the trained model in step 4 is: using the ArcFace loss function to adjust the cosine similarity so that the distance between classes is increased and the distance within the class is reduced. The specific formula is as follows: ArcFace=cos(θ+m) Among them, θ is the original cosine similarity, and m is the preset angle margin.

5. The monitoring face training optimization method based on full-supervision and semi-supervision joint training according to claim 4 is characterized in that: The semi-supervised loss in step eight is calculated by the following method: based on cosine similarity, the cross entropy loss is calculated by comparing the similarity calculation result with the true label and taking the average value. The formula is as follows:

6. The monitoring face training optimization method based on full-supervision and semi-supervision joint training according to claim 5 is characterized in that: The fully supervised loss is calculated by comparing the classification head outputs of models p and g with the true labels, and strengthening the training of difficult samples through the difficult sample mining (OHEM) method. The formula is as follows: Then calculate the total loss function: L total =L semi +L sup .

Citation Information

Patent Citations

  • A method for identifying student attention in online education scenarios optimized for CPU computing

    CN112597888B