High-precision face recognition method and system based on transfer learning
Through a multi-stage transfer strategy based on transfer learning and deep data enhancement technology, combined with attention mechanism and adaptive model optimization, data labeling problems in face recognition technology and low cross-scene recognition accuracy are solved, and high-precision and adaptive face recognition effect are achieved.
Patent Information
- Application Number
- CN202510644854.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing facial recognition technology has challenges in data labeling problems and low cross-scene recognition accuracy.
Using a transfer learning-based method, a multi-stage transfer strategy and deep data enhancement technology is combined with attention mechanism and adaptive model optimization to achieve adaptability and high-precision recognition of different scenario features.
It reduces the need for large-scale annotation data, improves the accuracy and consistency of annotation, and significantly improves the recognition accuracy of the model in different scenarios such as complex lighting, diverse poses and low resolution.
Smart Images

Figure CN120183019A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face recognition, and particularly relates to a high-precision face recognition method and system based on transfer learning. Background Art
[0002] As a key technology in the field of biometrics, face recognition technology has been widely used in many scenarios such as security monitoring, access control systems, financial payments, and unlocking of intelligent devices. This technology identifies individual identities by analyzing and comparing the feature information of faces, and has the advantages of convenience, high efficiency, non-contact, etc.
[0003] However, the current face recognition technology still faces many challenges in practical applications. On the one hand, the problem of data annotation severely restricts the improvement of face recognition accuracy. When constructing a face recognition model, a large amount of face data with accurate labels is required for training. However, manual annotation of data not only consumes a large amount of manpower, material resources, and time, but is also easily affected by the subjective factors of the annotators, resulting in errors in the annotation results. Especially in the process of annotating large-scale data sets, it is difficult to ensure the consistency and accuracy of the annotation, which will cause the model to learn incorrect information during training, thereby affecting the recognition accuracy.
[0004] On the other hand, there are significant differences in face recognition in different scenarios. For example, in the security monitoring scenario, the lighting conditions are complex and changeable, and there may be situations such as strong light direct shooting and shadow occlusion; in the access control system scenario, the face collection angles are diverse, and there may be non-frontal postures such as side faces, upward faces, and downward angles; in some low-resolution monitoring environments, the quality of face images is poor, and a large amount of detail information is lost. Most of the existing face recognition models may perform well in a single scenario, but when applied across scenarios, due to the lack of adaptability to the characteristics of different scenarios, the recognition accuracy will drop significantly. Summary of the Invention
[0005] The purpose of the present invention is to provide a high-precision face recognition method and system based on transfer learning, so as to solve the problems of data annotation difficulties and low cross-scenario recognition accuracy existing in face recognition in the prior art.
[0006] The purpose of the present invention can be achieved through the following technical solutions: A high-precision face recognition method based on transfer learning, comprising the following steps: S1: Select multiple general deep learning models and perform pre-training through a general face data set; screen the pre-trained deep learning models according to a preset plurality of evaluation indicators, and obtain an excellent model as the basic model; adopt a multi-stage transfer strategy to sequentially perform shallow feature transfer, middle-layer feature fine-tuning, and personalized training of high-level features; S2: Augment the data volume by adopting deep data augmentation techniques, and introduce generative adversarial networks for data augmentation; establish an annotation optimization system and introduce active learning algorithms; S3: Extract multi-dimensional scene features from the collected face images, use the attention mechanism to perform weighted fusion on the scene features of different dimensions, and design an adaptive model structure to dynamically adjust the parameters and structure of the model according to different scene features; S4: Use the face data fused with scene features and augmented to jointly train multiple models with different architectures, adopt knowledge distillation technology to let the teacher model guide the student model to learn; introduce a multi-objective optimization algorithm to optimize multiple performance indicators of the model at the same time; during the training process, adopt an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to the training effect of the model, and introduce regularization technology; regularly evaluate the model, and iteratively optimize the parameters, structure and training strategy of the model according to the evaluation results.
[0007] As a further solution of the present invention: in the S1, the evaluation indicators include accuracy rate, recall rate and F1 value.
[0008] As a further solution of the present invention: in the S2, it specifically includes the following steps: Let the original data be X = {x1, x2,..., x n}, and the data after traditional data augmentation operation is X aug1 ; The data generated by the generative adversarial network is X gan , then the augmented data set is X aug = X aug1 ∪ X gan ; Let the number of crowdsourcing annotators be m, and the annotation result of each annotator for the sample x i be l i,j , i = 1, 2,..., n; j = 1, 2,..., m; The final annotation result is , where I is the indicator function; The uncertainty of the sample is measured by entropy, and the formula is , and select the sample with a larger entropy value for annotation.
[0009] As a further solution of the present invention: in the S3, it specifically includes: Let the scene feature vector be , and the attention weight vector be ; The fused feature is ; Among them, the attention weight , where W and b are learnable parameter matrices and bias vectors respectively; is a learnable vector, and σ is an activation function; Let the scene feature be , the parameter set of the model is θ, and the adaptively adjusted parameter θ′ = θ + Δθ( ).
[0010] As a further solution of the present invention: in the said S4, the loss function of knowledge distillation is: L KD = (1 - λ)L CE (P S (y|x), y) + λT 2 KL(P T (y|x) / T, P S (y|x) / T); The multi-objective optimization problem is expressed as ; The adaptive learning rate adjustment rule is: If L val (t) > L val (t - 1), then η(t) = αη(t - 1), 0 < α < 1, Otherwise η(t) = η(t - 1); The loss function after adding regularization is: L reg (θ) = L(θ) + λ1 + λ2 .
[0011] As a further solution of the present invention: the shallow feature transfer is to directly transfer the parameters of the shallow convolutional layer of the pre-trained model to the new model; the middle feature fine-tuning is to perform small-batch and low learning rate fine-tuning on the middle convolutional layer on the new dataset; the personalized training of the high-level features is to perform large-scale training on the high-level convolutional layer and the fully connected layer according to the specific requirements of the target scene.
[0012] As a further solution of the present invention: in the multi-dimensional scene feature extraction, it includes illumination dimension features, pose dimension features, and image quality dimension features. The illumination dimension features include the brightness, contrast, color histogram, and illumination direction of the image; the pose dimension features include the pitch angle, yaw angle, and roll angle of the face; the image quality dimension features include the resolution, sharpness, and noise level of the image.
[0013] A high-precision face recognition system based on transfer learning, comprising: Model Migration and Initialization Module: Select multiple general deep learning models and pre-train them with a general face dataset; Screen the pre-trained deep learning models according to multiple preset evaluation metrics, and obtain the models with excellent performance as the basic models; Adopt a multi-stage migration strategy, and perform shallow feature migration, middle-layer feature fine-tuning, and personalized training of high-level features in sequence; Data Augmentation and Annotation Module: Use deep data augmentation technology to expand the data volume, and introduce generative adversarial networks for data augmentation; Establish an annotation optimization system and introduce active learning algorithms; Feature Fusion and Model Training Module: Extract multi-dimensional scene features from the collected face images, use the attention mechanism to weight and fuse the scene features of different dimensions, and design an adaptive model structure to dynamically adjust the parameters and structure of the model according to different scene features; Model Evaluation and Optimization Module: Use the face data that combines scene features and enhanced to jointly train multiple models with different architectures, adopt knowledge distillation technology, and let the teacher model guide the student model to learn; Introduce multi-objective optimization algorithms to optimize multiple performance indicators of the model at the same time; During the training process, adopt an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to the training effect of the model, and introduce regularization technology; Regularly evaluate the model, and iteratively optimize the parameters, structure, and training strategy of the model according to the evaluation results.
[0014] Advantages of the present invention: By using multi-stage transfer learning with pre-trained models, the need for large-scale labeled data is reduced, and the data annotation cost is lowered. At the same time, by using deep data augmentation technology and an optimized annotation system, the accuracy and consistency of annotation are improved, enabling the model to learn more accurate face features and enhancing the recognition accuracy; Fusing multi-dimensional scene features and adopting an adaptive model optimization strategy enables the model to better adapt to face changes in different scenarios, significantly improving the recognition accuracy of the model in different scenarios such as complex lighting, diverse poses, and low resolution, and broadening the application scope of face recognition technology; Adopting joint training, knowledge distillation, and multi-objective optimization algorithms speeds up the convergence rate of the model, improves the generalization ability and performance of the model. Through iterative optimization, the overall performance of the model is continuously improved, making it more reliable in practical applications. Description of the Drawings
[0015] The following further describes the present invention with reference to the drawings.
[0016] Figure 1 It is a schematic flowchart of a high-precision face recognition method based on transfer learning of the present invention. Specific Embodiments
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0018] Please refer to Figure 1 as shown in the figure, the present invention is a high-precision face recognition method based on transfer learning, including the following steps: Pre-trained model selection and multi-stage transfer: Select deep learning models pre-trained on multiple large-scale general face datasets, such as models with different architectures like ResNet and DenseNet. Comprehensively evaluate these pre-trained models and select the model that performs excellently on multiple evaluation metrics (such as accuracy, recall, F1 value) as the base model. The calculation formulas of the evaluation metrics are as follows: Accuracy: Accuracy = (TP + TN) / (FP + FN + TP + TN); Among them, TP (True Positive) represents the number of true positive examples, that is, the number of samples that are actually positive examples and are predicted as positive examples by the model; TN (True Negative) represents the number of true negative examples, that is, the number of samples that are actually negative examples and are predicted as negative examples by the model; FP (False Positive) represents the number of false positive examples, that is, the number of samples that are actually negative examples but are predicted as positive examples by the model; FN (False Negative) represents the number of false negative examples, that is, the number of samples that are actually positive examples but are predicted as negative examples by the model.
[0019] Recall: Recall = TP / (TP + FN); F1 value: F1 = 2 × Precision × Recall / (Precision + Recall); Among them, Precision = TP / (TP + FP) is the precision rate.
[0020] Adopt a multi-stage transfer strategy. First, perform shallow feature transfer, directly transfer the parameters of the shallow convolutional layer of the pre-trained model to the new model. Then, perform mid-level feature fine-tuning, and perform small-batch and low learning rate fine-tuning on the mid-level convolutional layer on the new dataset. Finally, perform personalized training on the high-level features, and perform large-scale training on the high-level convolutional layer and the fully connected layer according to the specific requirements of the target scenario.
[0021] During the multi-stage migration process, in the first step of shallow feature migration, the parameters of the first few convolutional layers of ResNet101 are directly copied to the new model. These shallow convolutional layers mainly learn general features such as edges and textures of face images. Then, middle-layer feature fine-tuning is carried out. On a small amount of labeled face data collected in the mall, the middle-layer convolutional layers are trained in small batches with a low learning rate of 0.0001, enabling the model to gradually adapt to some features of faces in the mall scene. Finally, personalized training of high-level features is conducted. In view of the characteristics of the mall scene, such as different lighting conditions and personnel postures, large-scale training is carried out on the high-level convolutional layers and fully connected layers to learn unique features related to the mall scene.
[0022] Deep Enhancement and Annotation Optimization System for Small-Sample Data: To address the problem of insufficient labeled data in the target scene, deep data augmentation techniques are used to expand the data volume. In addition to traditional operations such as rotation, translation, scaling, and adding noise, generative adversarial networks (GANs) are introduced for data augmentation. Let the original data be X = {x1, x2,..., x n}, and the data after traditional data augmentation operations (such as rotation angle θ, translation amount (t x , t y ), scaling ratio s, etc.) is represented as X aug1 . The data generated by GAN is represented as X gan , then the augmented data set X aug is: X aug = X aug1 ∪ X gan ; An annotation optimization system is established, adopting a combination of crowdsourcing annotation and expert review. Let the number of crowdsourcing annotators be m, and the annotation result of each annotator for the sample x i be l i,j , i = 1, 2,..., n; j = 1, 2,..., m; the final annotation result is determined through a voting mechanism: ; where I is the indicator function, l represents the class label. When l i,j = l, I(l i,j = l) = 1, otherwise I(l i,j = l) = 0. An active learning algorithm is introduced to select the most valuable samples for annotation according to the uncertainty of the model. arg max is the abbreviation of "argument of the maximum", and its function is to find the value of the independent variable that makes a certain function reach the maximum value. If there is a function f(x), argmax x f(x) represents the value of x that can make f(x) reach the maximum value.
[0023] Let the predicted probability distribution of the model for the sample \(x\) be \(P(y|x)\), then the uncertainty of the sample can be measured by entropy: ; Select samples with larger entropy values for annotation.
[0024] The data collection and preprocessing module collects face image data through high-definition cameras and infrared cameras at different positions in the mall. The collected data undergoes preprocessing operations such as grayscaling, normalization, histogram equalization, and denoising to improve the image quality.
[0025] In the data augmentation and annotation module, for the small amount of labeled face data collected, first perform traditional data augmentation operations, such as rotating the face image by ±20°, horizontally or vertically translating by ±8 pixels, scaling by ±15%, and adding an appropriate amount of salt-and-pepper noise. Then, use the GAN data augmentation sub-module to train a generative adversarial network. The generator network continuously learns to generate realistic face images in the mall scene, and the discriminator network determines whether the input image is a real image or a generated image. After multiple iterative trainings, a large number of samples similar to the actual face data distribution in the mall are generated.
[0026] Multi-dimensional scene feature fusion and model adaptive optimization: To improve the recognition accuracy of the model in different scenarios, multi-dimensional scene features are extracted from the collected face images. From the illumination dimension, features such as the brightness \(L\), contrast \(C\), color histogram \(H_{color}\), and illumination direction \(d\) of the image are extracted; from the pose dimension, pose information such as the pitch angle \(\alpha\), yaw angle \(\beta\), and roll angle \(\gamma\) of the face are extracted; from the image quality dimension, features such as the resolution \(R\), sharpness \(Q\), and noise level \(N\) of the image are extracted.
[0027] The attention mechanism is used to perform weighted fusion on scene features in different dimensions. Let the scene feature vector be , and the attention weight vector be , then the fused feature is: ; where the attention weight \(w\) i is calculated in the following way: ; Here, \(W\) and \(b\) are learnable parameter matrices and bias vectors, is a learnable vector, \(\sigma\) is an activation function (such as ReLU), \(\exp()\) is the exponential operation, which is the mathematical representation of the natural exponential function (the exponential function with the natural constant \(e\) as the base), and \(T\) represents the matrix transpose operation.
[0028] At the same time, design an adaptive model structure to dynamically adjust the parameters and structure of the model according to different scene features. Let the scene feature be , if the parameter set of the model is θ, then the adaptively adjusted parameter θ′ can be expressed as: θ′ = θ + Δθ(s); where Δθ(s) is a function related to the scene feature s, used to adjust the model parameters according to the scene feature.
[0029] Model joint training and iterative optimization mechanism: Input the fused scene features and enhanced face data into ResNet101 and a small MobileNet model for joint training. During the joint training process, use the knowledge distillation training sub-module, with ResNet101 as the teacher model and MobileNet as the student model. According to the loss function L of knowledge distillation KD for training to adjust the parameters of the student model.
[0030] Specifically, let a larger teacher model T guide a smaller student model S to learn. Suppose the output probability distribution of the teacher model is P T (y|x), and the output probability distribution of the student model is P S (y|x). The loss function L of knowledge distillation KD is: L KD = (1 - λ)L CE (P S (y|x), y) + λT 2 KL(P T (y|x) / T, P S (y|x) / T); where L CE is the cross-entropy loss function, KL is the KL divergence, T is the temperature parameter, and λ is the weight coefficient.
[0031] Introduce a multi-objective optimization algorithm to optimize multiple performance metrics of the model simultaneously, such as accuracy, recall, recognition speed, etc. Suppose the objective function is F(θ), which contains multiple sub-objective functions F i (θ) (i = 1, 2,..., n), then the multi-objective optimization problem can be expressed as: ; where w i is the weight coefficient of each sub-objective function.
[0032] During the training process, adopt an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to the training effect of the model. Suppose the initial learning rate is η0, the current training round is t, and the loss function value on the validation set is L val (t), then the learning rate η(t) can be adjusted according to the following rules: Lval (t) > L val If (t - 1), then η(t) = αη(t - 1), 0 < α < 1, otherwise η(t) = η(t - 1); Meanwhile, introduce regularization techniques such as Dropout, L1, and L2 regularization; Train according to the loss function after adding regularization to prevent the model from overfitting.
[0033] Let the original loss function be L(θ), then the loss function L reg (θ) is: L reg (θ) = L(θ) + λ1 + λ2 ; Wherein, is the L1 norm, is the L2 norm, and λ1 and λ2 are regularization coefficients.
[0034] Regularly evaluate the model, use the validation set and test set data, and calculate indicators such as the recognition accuracy rate, recall rate, F1 value, misrecognition rate, and rejection rate of the model. According to the evaluation results, iteratively optimize the parameters, structure, and training strategy of the model to continuously improve the performance of the model.
[0035] The model evaluation and optimization module regularly evaluates the trained model. On the validation set and test set, calculate indicators such as the recognition accuracy rate, recall rate, F1 value, misrecognition rate, and rejection rate of the model. For example, in the initial training stage, the accuracy rate of the model on the test set is 80%, and the recall rate is 75%. According to the evaluation results, it is found that the accuracy rate of the model is low when recognizing large-angle side faces. To address this issue, adjust the structure of the model, add a layer for extracting side face features, and further enhance the training data to increase the proportion of side face samples. After multiple iterations of optimization, evaluate again on the test set, and the recognition accuracy rate is increased to 93%, and the recall rate is increased to 90%, significantly improving the performance of the model in the mall scenario.
[0036] The above has described a detailed embodiment of the present invention, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the present invention application shall still fall within the scope covered by the patent of the present invention.
Claims
1. A high-precision face recognition method based on transfer learning, characterized in that: The following steps are involved: S1: Select multiple common deep learning models and pre-train them using common face datasets; screen the pre-trained deep learning models according to multiple preset evaluation indicators to obtain models with excellent performance as basic models; adopt a multi-stage migration strategy to sequentially perform shallow feature migration, mid-level feature fine-tuning, and personalized training of high-level features; S2: Use deep data enhancement technology to expand the data volume and introduce generative adversarial networks for data enhancement; establish a labeling optimization system and introduce active learning algorithms; S3: Extract multi-dimensional scene features from the collected face images, use the attention mechanism to weightedly fuse scene features of different dimensions, and design an adaptive model structure to dynamically adjust the model parameters and structure according to different scene features; S4: Use the fusion of scene features and enhanced face data to jointly train multiple models of different architectures, and use knowledge distillation technology to let the teacher model guide the student model to learn; introduce a multi-objective optimization algorithm to optimize multiple performance indicators of the model at the same time; during the training process, use an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to the training effect of the model, and introduce regularization technology; regularly evaluate the model, and iteratively optimize the model's parameters, structure, and training strategy based on the evaluation results.
2. According to claim 1, a high-precision face recognition method based on transfer learning is characterized in that: In S1, the evaluation indicators include accuracy, recall and F1 value.
3. According to the high-precision face recognition method based on transfer learning in claim 1, it is characterized in that: In said S2, the following steps are specifically included: Assume the original data is X={x1,x2,...,x n }, the data after traditional data enhancement operation is X aug1 ; The data generated by the generative adversarial network is X gan , then the enhanced data set is X aug =X aug1 ∪X gan ; Assume that the number of crowdsourced annotation personnel is m, and each annotation personnel has i The labeling result is l i,j , i=1,2,...,n; j=1,2,...,m; The final annotation result is , where I is the indicator function and l represents the category label; The uncertainty of the sample is measured by entropy, and the formula is , select samples with larger entropy values for labeling.
4. According to claim 1, a high-precision face recognition method based on transfer learning is characterized in that: In said S3, it specifically includes: Let the scene feature vector be , the attention weight vector is ; The fused features are ; The attention weight , where W and b are the learnable parameter matrix and bias vector respectively; is a learnable vector, σ is an activation function; Assume the scene feature is , the parameter set of the model is θ, and the adaptively adjusted parameter θ′=θ+Δθ( ).
5. The high-precision face recognition method based on transfer learning according to claim 1, characterized in that: In S4, the loss function of knowledge distillation is: L KD =(1−λ)L CE (P S (y∣x),y)+λT 2 KL(P T (y∣x) / T,P S (y∣x) / T); The multi-objective optimization problem is expressed as ; The adaptive learning rate adjustment rule is: if L val (t)>L val (t−1), then η(t)=αη(t−1), 0<α<1, Otherwise η(t)=η(t−1); The loss function after adding regularization is: L reg (θ)=L(θ)+λ1 +λ2 。 6. The high-precision face recognition method based on transfer learning according to claim 1, characterized in that: The shallow feature migration is to directly migrate the shallow convolution layer parameters of the pre-trained model to the new model; the middle-layer feature fine-tuning is to fine-tune the middle-layer convolution layer in small batches and at a low learning rate on the new data set; Personalized training of high-level features is to perform large-scale training on high-level convolutional layers and fully connected layers according to the specific needs of the target scene.
7. The high-precision face recognition method based on transfer learning according to claim 1, characterized in that: The multi-dimensional scene feature extraction includes illumination dimension features, posture dimension features and image quality dimension features. The illumination dimension features include image brightness, contrast, color histogram and illumination direction; the posture dimension features include the pitch angle, yaw angle and roll angle of the face; the image quality dimension features include image resolution, clarity and noise level.
8. A high-precision face recognition system based on transfer learning, characterized in that: include: Model migration and initialization module: select multiple common deep learning models and pre-train them using common face datasets; The pre-trained deep learning models are screened according to multiple preset evaluation indicators to obtain models with excellent performance as the basic models; a multi-stage migration strategy is adopted to sequentially perform shallow feature migration, mid-level feature fine-tuning, and personalized training of high-level features; Data enhancement and annotation module: Deep data enhancement technology is used to expand the data volume, and generative adversarial networks are introduced for data enhancement; Establish a labeling optimization system and introduce active learning algorithms; Feature fusion and model training module: extract multi-dimensional scene features from the collected face images, use the attention mechanism to weightedly fuse scene features of different dimensions, and design an adaptive model structure to dynamically adjust the model parameters and structure according to different scene features; Model evaluation and optimization module: Use the fusion of scene features and enhanced face data to jointly train multiple models of different architectures, adopt knowledge distillation technology to let the teacher model guide the student model to learn; introduce multi-objective optimization algorithm to optimize multiple performance indicators of the model at the same time; during the training process, adopt an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to the training effect of the model, and introduce regularization technology; evaluate the model regularly, and iteratively optimize the model parameters, structure and training strategy according to the evaluation results.
Citation Information
Patent Citations
A face image age recognition method based on transfer learning
CN109815864A
Lightweight real-time human body posture recognition method based on MobileViT
CN118609163A
Face recognition method and system in complex environment based on transfer learning
CN118609190A
Personal identification-oriented face quality perception method and system
WO2022073453A1
Cited By
Multi-scene pavement PCI (Peripheral Component Interconnect) prediction method based on cross-regional transfer learning
CN120912573A
Bruxism recognition system based on multi-angle face image
CN121506438A
Face recognition method and system based on cloud machine cooperation
CN121527601A