Face attribute recognition method based on transfer learning

By applying transfer learning technology in the field of face attribute recognition, using pre-trained models to extract features and adjust the model structure, the problems of high training costs and insufficient generalization capabilities in traditional methods are solved, and efficient and accurate face attribute recognition is achieved.

CN120126199APending Publication Date: 2025-06-10SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510267650.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Traditional face attribute recognition technology faces high training costs and limitations of model generalization capabilities, and it is difficult to adapt to new scenarios or unprecedented changes in face attributes.

Method used

The face attribute recognition method based on transfer learning is adopted, and the feature extraction is performed using a pre-trained deep neural network model, and the last layer of the basic model is structurally adjusted to adapt to specific face attribute recognition tasks.

Benefits of technology

It significantly reduces the dependence on a large amount of labeled data, improves model training efficiency and recognition accuracy, enhances generalization performance, and is suitable for data scarce environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126199A_ABST
    Figure CN120126199A_ABST
Patent Text Reader

Abstract

The invention discloses a face attribute recognition method based on transfer learning, and belongs to the technical field of artificial intelligence, and the method comprises the following steps: obtaining a data set containing a face image; performing feature extraction on the data set by using a pre-trained deep neural network model; performing structure adjustment on the last layer of the basic model to adapt to a specific face attribute recognition task; and performing attribute recognition on the new face image by using the adjusted model. According to the method, the dependence on a large amount of labeled data is effectively reduced, the model training efficiency and the recognition precision are improved, and the wide application potential in a data scarcity environment is shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and specifically to a face attribute recognition method based on transfer learning. Background Art

[0002] With the development of artificial intelligence technology, face recognition technology has been widely applied in many fields such as security monitoring and personalized services. Traditional face attribute extraction models often require a large amount of labeled data for training, which is not only time-consuming and laborious, but also difficult to obtain sufficient samples in specific application scenarios.

[0003] In the prior art, face attribute recognition technology faces two major challenges: firstly, the high training cost. Since building an efficient face attribute recognition model usually requires a large and diverse labeled data set, which is not only difficult to collect, but also has a very high labeling cost; secondly, the limitation of the model generalization ability. Traditional methods often train from scratch for specific tasks, resulting in the model being difficult to adapt to new scenarios or unseen face attribute changes. Summary of the Invention

[0004] The technical task of the present invention is to provide a face attribute recognition method based on transfer learning in view of the above deficiencies, which can reduce data requirements, improve training efficiency, enhance generalization performance, and has wide applicability.

[0005] The technical solution adopted by the present invention to solve its technical problems is: A face attribute recognition method based on transfer learning, the implementation of this method includes the following steps: Obtain a data set containing face images; Use a pre-trained deep neural network model to extract features from the data set; Adjust the structure of the last layer of the base model to adapt to a specific face attribute recognition task; Use the adjusted model to perform attribute recognition on new face images.

[0006] As an effective method for utilizing pre-trained models, transfer learning can transfer knowledge from one domain to another related but data-scarce domain, thereby reducing the dependence on large-scale labeled data. This method overcomes the problems of large data requirements and high training costs in traditional methods. By selecting a recognition model pre-trained on a large-scale face dataset as a starting point and combining transfer learning strategies, the model is fine-tuned and optimized for face attribute recognition tasks such as age, gender, expression, and facial features (such as single / double eyelids, facial feature shapes). It includes: selecting a basic face recognition model with strong generalization ability; defining a set of target face attributes and corresponding labeled data; implementing feature-level transfer and model structure adjustment; and training a new model on the adjusted model. Thus, it effectively reduces the dependence on a large amount of labeled data, improves the model training efficiency and recognition accuracy, and demonstrates broad application potential in data-scarce environments.

[0007] Furthermore, the face attributes include age, gender, face shape, eyelid features, expression status, glasses wearing situation, etc.

[0008] Furthermore, the pre-trained deep neural network model is pre-trained on a large-scale image dataset, and the model at least includes a convolutional layer, a fully connected layer, and an output layer.

[0009] Furthermore, the structure of the last layer of the basic model is adjusted. The adjustment steps include fine-tuning the last layer or multiple layers of the pre-trained model to optimize the classification performance for target face attributes.

[0010] Furthermore, it also includes using data augmentation techniques to increase the diversity of the training dataset to improve the generalization ability of the model.

[0011] Furthermore, the specific implementation of this method includes: 1) Select a basic model. Select a face recognition model pre-trained on a large-scale face dataset as the basic model. This model should have good generalization ability and be able to capture the basic features of faces; 2) Feature-level transfer: Add a fully connected layer after the output layer of the basic model as a new classifier layer; 3) Define face attribute labels: According to the application scenario requirements, clarify the face attributes to be extracted, such as age groups, genders, face shapes, eyelid features, etc., and prepare the corresponding attribute label datasets, even if the scale is small; 4) Data preparation and preprocessing: Dataset construction: Collect a dataset containing face attribute labels, such as age, gender, expression, etc.; the dataset should contain sufficient diversity to cover different face attributes; Data preprocessing: Crop and scale the original images to fit the input size of the base model; Apply data augmentation techniques such as random rotation, flipping, brightness adjustment, etc. to enhance the generalization ability of the model; Perform normalization to convert pixel values into the range [0,1]. 5) The training process includes: transfer learning strategy, loss function design, and regularization techniques; Monitor the performance on the validation set and stop training when the performance no longer improves. Split the overall task into classification tasks for multiple attributes, and train each attribute classifier separately. 6) Perform optimization, including: Dynamic learning rate adjustment: Automatically adjust the learning rate according to the loss changes during training to accelerate the convergence process. Multi-GPU parallel training: Utilize multiple GPUs for parallel training to shorten the overall training time. 7) Model evaluation and deployment, including: Performance evaluation: Use metrics including accuracy, recall, and F1-score to evaluate the model performance; Compare the impacts of different data augmentation techniques and regularization techniques on the model performance. Model deployment: Deploy the trained model to a server or mobile device for real-time face attribute recognition; Depending on the deployment platform, quantization or other optimization measures may be required for the model. When the trained model is used for inference tasks, since the pre-trained model weights are frozen during the training of different classifiers, different classifiers can share the model output.

[0012] After inputting a face image into the pre-trained model, obtain the feature vector (EMBEDDING) of the image, and then send this vector through a fully connected manner to different trained attribute classifiers to further obtain different attribute feature classifications.

[0013] Furthermore, use Google FaceNet as the base model, which is a face embedding model based on a deep convolutional neural network (CNN). The FaceNet model can map the input face image into a fixed-length vector representation. The FaceNet model is trained using the triplet loss function and can learn the similarities and differences between faces. The FaceNet model includes: a batch input layer, a deep convolutional network, and an L2-norm normalization layer (outputting the face embedding vector) connected behind the deep convolutional network; During the training process, connect the triplet loss to minimize the model loss. Training a model from scratch requires a large amount of labeled data, computing power, and time. However, the weights of the pre-trained model provided have already achieved good accuracy. The pre-trained model weights can be directly used for the next step of transfer learning; Add a fully connected layer after the output layer of the FaceNet model as the new classifier layer; determine the number of output nodes of the classifier layer according to the number of face attribute categories to be recognized; use the Softmax activation function to obtain the probability distribution of each attribute category.

[0014] The FaceNet network pre-trained model has completed a large amount of training on general face recognition tasks and has good recognition accuracy. The vectors in its output EMBEDDING layer contain rich face features. Therefore, different classifiers can be trained for different features to complete the classification tasks of specific features. The classifier consists of a fully connected layer of classification nodes and a Softmax layer, and the classifier is optimized by the cross-entropy loss method.

[0015] Furthermore, the training process Transfer learning strategy: Keep the parameters of the pre-trained FaceNet model frozen and only train the newly added classifier layer; use a small learning rate to fine-tune the classifier layer to avoid destroying the feature learning in the pre-trained model; Loss function design: Use the cross-entropy loss function (Cross-Entropy Loss) as the loss function for classification tasks; if multiple attributes need to be recognized simultaneously, a multi-task learning framework can be used, and each attribute corresponds to a loss function; Regularization techniques: To prevent overfitting, techniques such as Dropout or L2 regularization can be applied; use the early stopping method (Early Stopping), monitor the performance on the validation set, and stop training when the performance no longer improves.

[0016] The present invention also claims to protect a face attribute recognition device based on transfer learning, including: at least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program to implement the above method.

[0017] The present invention also claims to protect a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the above method is implemented.

[0018] Compared with the prior art, a face attribute recognition method based on transfer learning of the present invention has the following beneficial effects: 1. Reduce data requirements: Significantly reduce the dependence on a large amount of labeled data, suitable for data-scarce scenarios.

[0019] 2. Improve efficiency and accuracy: Quickly build a high-performance face attribute extraction model with limited resources and improve the accuracy of attribute recognition.

[0020] 3. Wide applicability: This method has high flexibility and can adapt to different face attribute recognition tasks, promoting the application of the technology in multiple fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic diagram of the principle of a face attribute recognition method based on transfer learning provided by an embodiment of the present invention; Figure 2 is a schematic diagram of the FaceNet model structure provided by an embodiment of the present invention; Figure 3 is a schematic diagram of the network structure of the age classifier provided by an embodiment of the present invention; Figure 4 is a schematic diagram of the simplified structure of the classifier for each attribute provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0023] An embodiment of the present invention provides a face attribute recognition method based on transfer learning. The implementation of this method includes the following steps: Obtain a data set containing face images; Use a pre-trained deep neural network model to extract features from the data set; Adjust the structure of the last layer of the base model to adapt to a specific face attribute recognition task; Use the adjusted model to perform attribute recognition on new face images.

[0024] Among them, the face attributes include age, gender, face shape, eyelid features, expression state, glasses wearing situation, etc.

[0025] The pre-trained deep neural network model is pre-trained on a large-scale image data set, and the model at least includes a convolutional layer, a fully connected layer, and an output layer.

[0026] The adjustment of the structure of the last layer of the base model includes fine-tuning the last layer or multiple layers of the pre-trained model to optimize the classification performance for the target face attributes.

[0027] This method also includes using data augmentation techniques to increase the diversity of the training data set to improve the generalization ability of the model.

[0028] As an effective method for utilizing pre-trained models, transfer learning can transfer knowledge from one domain to another related but data-scarce domain, thereby reducing the dependence on large-scale labeled data. This method overcomes the problems of large data requirements and high training costs in traditional methods. By selecting an identification model pre-trained on a large-scale face dataset as the starting point and combining transfer learning strategies, the model is fine-tuned and optimized for face attribute recognition tasks such as age, gender, expression, and facial features (such as single / double eyelids, facial feature shapes). It includes: selecting a basic face recognition model with strong generalization ability; defining the target face attribute set and corresponding label data; implementing feature-level transfer and model structure adjustment; and training a new model on the adjusted model. Thus, it effectively reduces the dependence on a large amount of labeled data, improves the model training efficiency and recognition accuracy, and demonstrates broad application potential in data-scarce environments.

[0029] Combined with the attached Figures 1-4 As described, the specific implementation of this method includes: 1. Select the basic model: First, select a face recognition model pre-trained on a large-scale face dataset as the basic model. This model should have good generalization ability and be able to capture the basic features of faces.

[0030] Use Google FaceNet as the basic model. It is a face embedding model based on a deep convolutional neural network (CNN) that can map the input face image to a fixed-length vector representation.

[0031] The FaceNet model is trained using the triplet loss function and can learn the similarities and differences between faces.

[0032] The structure of the FaceNet model is as Figure 2 shown. This model consists of a batch input layer, a deep convolutional network followed by an L2 norm normalization layer (outputting the face embedding vector). During the training process, triplet loss is connected at the back to minimize the model loss. In the figure, EMBEDDING represents the fused face feature vector.

[0033] Training the model from scratch requires a large amount of labeled data, computing power, and time. And the pre-trained model weights it provides have reached a good accuracy. The pre-trained model weights can be directly used for the next step of transfer learning.

[0034] 2. Feature-level transfer: Add a fully connected layer after the output layer of FaceNet as the new classifier layer.

[0035] Determine the number of output nodes of the classifier layer according to the number of categories of face attributes to be recognized.

[0036] Use the Softmax activation function to obtain the probability distribution of each attribute category.

[0037] The pre-trained model of the FaceNet network has completed a large amount of training on the general face recognition task and has good recognition accuracy. The vector of its output EMBEDDING layer contains rich face features. Therefore, different classifiers can be trained for different features to complete the classification tasks of specific features. The classifier consists of a fully connected layer of classification nodes and a Softmax layer, and the classifier is optimized by the cross-entropy loss method.

[0038] Taking the age classifier as an example, the network structure is as Figure 3 shown.

[0039] Construct classifiers for other attributes in the same way. The simplified schematic diagram is as Figure 4 shown.

[0040] 3. Define face attribute labels: According to the requirements of the application scenario, clarify the face attributes to be extracted, such as age group, gender, face shape, eyelid features, etc., and prepare the corresponding attribute label dataset, even if the scale is small.

[0041] Examples of face attributes are as follows: The age group is divided into juvenile, young, middle-aged, and elderly; The gender is divided into male and female; The face shape is divided into round face and long face; The eyelids are divided into single eyelids and double eyelids.

[0042] 4. Data preparation and preprocessing: Dataset construction: Collect a dataset containing face attribute labels, such as age, gender, expression, etc.; the dataset should contain sufficient diversity to cover different face attributes.

[0043] Data preprocessing: Crop and scale the original images to adapt to the input size of the base model. Apply data augmentation techniques, such as random rotation, flipping, brightness adjustment, etc., to increase the generalization ability of the model. Perform normalization processing to convert the pixel values into the range of [0,1].

[0044] 5. Training process: Transfer learning strategy: Keep the parameters of the pre-trained FaceNet model frozen and only train the newly added classifier layer.

[0045] Fine-tune the classifier layer with a smaller learning rate to avoid destroying the feature learning in the pre-trained model.

[0046] Loss function design: Use the Cross-Entropy Loss function as the loss function for the classification task.

[0047] If multiple attributes need to be recognized simultaneously, a multi-task learning framework can be used, with each attribute corresponding to a loss function.

[0048] Regularization techniques: To prevent overfitting, techniques such as Dropout or L2 regularization can be applied.

[0049] Use Early Stopping, monitor the performance on the validation set, and stop training when the performance no longer improves.

[0050] Split the overall task into classification tasks for multiple attributes. Each attribute classifier is trained separately.

[0051] 6. Training techniques: Dynamic learning rate adjustment: Automatically adjust the learning rate according to the loss change during training to accelerate the convergence process.

[0052] Multi-GPU parallel training: Utilize multiple GPUs for parallel training to shorten the overall training time.

[0053] 7. Model evaluation and deployment: Performance evaluation: Use metrics such as accuracy, recall, and F1-score to evaluate the model performance.

[0054] Compare the effects of different data augmentation techniques and regularization techniques on the model performance.

[0055] Model deployment: Deploy the trained model to a server or a mobile device for real-time face attribute recognition. Depending on the deployment platform, quantization or other optimization measures may be required; When the trained model is used for inference tasks, since the weights of the FaceNet pre-trained model are frozen during the training of different classifiers, different classifiers can share the output of the FaceNet model.

[0056] After inputting a face image into the pre-trained FaceNet model, obtain the feature vector (EMBEDDING) of the image, and then send this vector through a fully connected manner to different trained attribute classifiers to further obtain different attribute feature classifications.

[0057] This method significantly reduces the new model's dependence on large datasets by reusing the knowledge of pre-trained models, while accelerating the training speed and improving the model's generalization ability in different face attribute recognition tasks. Specifically, this method achieves the following goals: Reduce data requirements: Utilize transfer learning techniques to enable the model to be effectively fine-tuned on small-scale or domain-specific data, thereby reducing the need for a large number of labeled samples.

[0058] Improve training efficiency: By adjusting and optimizing based on pre-trained models, shorten the training time and accelerate the model convergence process.

[0059] Enhance generalization performance: Ensure that the model adjusted by transfer learning can accurately recognize diverse face attributes and maintain a high recognition rate even when facing new samples or changing conditions.

[0060] Wide applicability: The proposed method should be easily extensible to various face attribute recognition application scenarios, including but not limited to security monitoring, human-computer interaction, personalized recommendation systems, etc., to improve the practical value and market adaptability of the technology.

[0061] In summary, this method constructs an efficient, accurate, and adaptable face attribute extraction model through transfer learning strategies.

[0062] An embodiment of the present invention also provides a face attribute recognition device based on transfer learning, including: at least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program to implement the face attribute recognition method based on transfer learning described in the above embodiments.

[0063] An embodiment of the present invention also provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the face attribute recognition method based on transfer learning described in the above embodiments is implemented. Specifically, a system or device equipped with a storage medium can be provided, on which software program codes for implementing the functions of any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device is made to read and execute the program codes stored in the storage medium.

[0064] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments, so the program code and the storage medium storing the program code constitute a part of the present invention.

[0065] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.

[0066] In addition, it should be clear that not only can the functions of any one of the above embodiments be realized by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer to perform part or all of the actual operations based on the instructions of the program code.

[0067] Furthermore, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is caused to perform part and all of the actual operations, thereby realizing the functions of any one of the above embodiments.

[0068] The present invention has been described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that the code review means in different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.

Claims

1. A face attribute recognition method based on transfer learning, characterized in that: The implementation of this method includes the following steps: Get a dataset containing face images; Using a pre-trained deep neural network model to extract features from the data set; The last layer of the base model is structurally adjusted to suit the specific face attribute recognition task; Use the adjusted model to perform attribute recognition on new face images.

2. The method for face attribute recognition based on transfer learning according to claim 1, characterized in that: The facial attributes include age, gender, face shape, eyelid features, facial expression, and glasses wearing condition.

3. The method for face attribute recognition based on transfer learning according to claim 1, characterized in that: The pre-trained deep neural network model is pre-trained on a large-scale image dataset, and the model at least includes a convolutional layer, a fully connected layer and an output layer.

4. The method for face attribute recognition based on transfer learning according to claim 1, characterized in that: The structural adjustment is performed on the last layer of the basic model, and the adjustment step includes fine-tuning the last layer or multiple layers of the pre-trained model to optimize the classification performance of the target facial attributes.

5. The method for face attribute recognition based on transfer learning according to claim 1, characterized in that: It also includes the use of data augmentation techniques to increase the diversity of the training data set to improve the generalization ability of the model.

6. The method for face attribute recognition based on transfer learning according to claim 1, characterized in that: The specific implementation of this method includes: 1) Select a basic model. A face recognition model pre-trained on a large-scale face dataset is selected as the basic model. The model should have good generalization ability and be able to capture the basic features of the face. 2) Feature layer migration: Add a fully connected layer after the output layer of the base model as a new classifier layer; 3) Define facial attribute labels: According to the application scenario requirements, clarify the facial attributes to be extracted and prepare the corresponding attribute label dataset; 4) Data preparation and preprocessing: Dataset construction: Collect a dataset containing facial attribute labels; Data preprocessing: crop and scale the original image to fit the input size of the basic model; apply data augmentation techniques to increase the generalization ability of the model; perform normalization to convert pixel values ​​to the [0,1] range; 5) The training process includes: transfer learning strategy, loss function design, and regularization techniques; monitoring the performance on the validation set and stopping training when the performance no longer improves; The overall task is split into classification tasks for multiple attributes, and each attribute classifier is trained separately; 6) Optimize, including: Dynamic learning rate adjustment: Automatically adjust the learning rate according to the change in loss during training to accelerate the convergence process; Multi-GPU parallel training: Use multiple GPUs for parallel training to shorten the overall training time; 7) Model evaluation and deployment, including: Performance evaluation: Use indicators including accuracy, recall, and F1 score to evaluate model performance, and compare the impact of different data augmentation and regularization techniques on model performance; Model deployment: Deploy the trained model to a server or mobile device for real-time facial attribute recognition; After the face image is input into the pre-trained model, the feature vector of the image is obtained, and then the vector is sent to the trained different attribute classifiers through a full connection method to further obtain different attribute feature classifications.

7. The method for face attribute recognition based on transfer learning according to claim 6, characterized in that: Google FaceNet is used as the basic model. The FaceNet model can map the input face image into a vector representation of a fixed length. The FaceNet model is trained using a triplet loss function, which can learn the similarities and differences between faces; The FaceNet model includes: a batch input layer, a deep convolutional network, and an L2 norm normalization layer connected to the deep convolutional network; during the training process, a triplet loss is used to minimize the model loss; Add a fully connected layer after the output layer of the FaceNet model as a new classifier layer; determine the number of output nodes of the classifier layer according to the number of face attribute categories that need to be recognized; use the Softmax activation function to obtain the probability distribution of each attribute category.

8. The method for face attribute recognition based on transfer learning according to claim 7, characterized in that: The training process, Transfer learning strategy: keep the pre-trained FaceNet model parameters frozen and only train the newly added classifier layer; Fine-tune the classifier layer using a small learning rate to avoid destroying the features learned in the pre-trained model; Loss function design: Use the cross entropy loss function as the loss function for the classification task; if there are multiple attributes that need to be recognized at the same time, use a multi-task learning framework with one loss function for each attribute; Regularization technique: Apply Dropout or L2 regularization technique; Use early stopping to monitor the performance on the validation set and stop training when the performance stops improving.

9. A face attribute recognition device based on transfer learning, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 8.

10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.