A method, system, electronic device and storage medium for estimating face age
By introducing attribute guidance module and composite loss function in the face age estimation technology, the problems of insufficient robustness and failure to effectively utilize attribute information in the prior art are solved, and more efficient and accurate face age estimation is achieved.
Patent Information
- Application Number
- CN202111538375.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-15
AI Technical Summary
The existing face age estimation technology has problems with insufficient robustness, difficulty in expressing the nonlinear and continuity characteristics of age-related growth, and failure to effectively utilize age-related label information such as gender and race.
The face age estimation method based on attribute guidance is adopted, and the feature extraction module and attribute guidance module are constructed, and the composite loss function of the error compression sorting loss based on the sorting label and the attribute guidance classification loss is realized to achieve accurate estimation of face age.
It improves the robustness and accuracy of face age estimation, can more effectively utilize age-related attribute information, and improves the robustness of the network and the accuracy of age estimation.
Smart Images

Figure CN114399808B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pattern recognition and digital image processing, and in particular relates to a method, system, electronic equipment and storage medium for estimating the age of a human face. Background Art
[0002] At present, face age estimation is a biometric recognition technology that infers a person's age information based on his or her facial features. With the continuous development of big data related technologies, face age estimation has been widely used in auxiliary identity authentication, human-computer interaction, demography and other fields. Among these tasks, age estimation based on face images has gradually become an important and challenging topic.
[0003] From the current research, according to the different ways of extracting features, face age estimation can be divided into two categories: methods based on traditional machine learning and methods based on deep learning. Since the manual feature extraction method of traditional machine learning methods is limited by conditions such as a single facial posture and accurate positioning of facial feature points, its robustness is far less than that of deep learning methods supported by big data. In the early days, deep learning methods regarded age estimation as a type of pattern recognition problem. Since the growth of age values is a process of change of ordered numbers, and each age value can be regarded as a class separately, age estimation can be considered as a regression problem or a classification problem. At the same time, based on the expression of the dynamic nature of age growth, recent studies have introduced ranking models into the problem of age estimation and achieved good results.
[0004] Both estimation models have their limitations. For the two major characteristics of age growth: nonlinearity and continuity, classification and regression models can only tend to express one of them. The sorting model has improved this problem to a certain extent, but there is still room for development in the application of the sorting model. On the other hand, age information is highly correlated with label information such as gender and race, which has not been considered and studied in the field of age estimation.
[0005] Through the above analysis, the problems and defects of the prior art are as follows:
[0006] (1) Since the manual feature extraction method of traditional machine learning methods is limited by conditions such as a single facial posture and accurate positioning of facial feature points, its robustness is far lower than that of deep learning methods supported by big data.
[0007] (2) In the existing age estimation models, classification and regression models can only tend to express one of the two major characteristics of age growth: nonlinearity and continuity, and there is still room for development in the application of sorting models.
[0008] (3) In existing age estimation models, age information is highly correlated with label information such as gender and race, which has not been considered and studied in the field of age estimation.
[0009] The difficulty of solving the above problems and defects is:
[0010] (1) There is a certain conflict in the expression of age characteristics (nonlinearity and continuity), which often cannot be well expressed at the same time. A loss function that satisfies both as much as possible is needed;
[0011] (2) How to use information highly correlated with age, such as gender and race, in neural networks is a difficult and innovative point.
[0012] The significance of solving the above problems and defects is:
[0013] Using a loss function that is close to the characteristics of age changes to calculate age values is more conducive to accurate age estimation; using attribute-constrained information for the calculation of age values can make the age learning process more deterministic, which is conducive to the accurate expression of the results and improving the robustness of the network. Summary of the invention
[0014] In view of the problems existing in the prior art, the present invention provides a method, system, electronic device and storage medium for estimating the age of a face, and more particularly, relates to a method, system, electronic device and storage medium for estimating the age of a face based on attribute guidance.
[0015] The present invention is implemented as follows: a method for estimating the age of a human face, the method comprising the following steps:
[0016] Step 1: Obtain a face age image set and perform preprocessing to obtain a preprocessed face age image set. The purpose of preprocessing is to make the face images have similar image sizes and eliminate background interference as much as possible;
[0017] Step 2: Construct a face age estimation model. Specifically, it includes expressing the characteristics of age by constructing an error compression sorting loss based on sorting labels, constructing a feature extraction module for convolution operations, and constructing an attribute guidance module and attribute guidance classification loss to establish the connection between the fully connected layer and the corresponding attributes.
[0018] Step 3: training the face age estimation model according to the face preprocessing image set to obtain a trained face age estimation model;
[0019] Step 4: testing the trained face age estimation model according to the test data set to obtain a face age estimation result to achieve face age estimation.
[0020] Further, the face age image set is preprocessed in step 1 to obtain the preprocessed face age image set, which includes:
[0021] Performing face detection, cropping and scaling on the face age image set to obtain the preprocessed face age image set;
[0022] The preprocessed face age image set is randomly divided into a training set, a validation set and a test set according to a certain ratio.
[0023] Furthermore, the face age estimation model constructed in step 2 includes a feature extraction module and an attribute guidance module;
[0024] Among them, the feature extraction module includes a basic convolution unit that is repeatedly connected in sequence for different times, namely a multi-scale attention mechanism residual convolution unit; wherein the basic convolution unit includes a multi-scale convolution mechanism and a channel attention mechanism that are connected in sequence; the multi-scale convolution mechanism includes convolution layers with different convolution kernel sizes and different numbers of output channels; the channel attention mechanism includes a global pooling layer, a fully connected layer and an activation layer.
[0025] The attribute guidance module includes a first attribute fully connected layer, a second attribute fully connected layer and a global fully connected layer; wherein the first attribute fully connected layer is connected to a corresponding number of output neurons, the first attribute fully connected layer is connected to the second attribute fully connected layer, and the second attribute fully connected layer is spliced with the global fully connected layer;
[0026] The input of the feature extraction module is a face image, and the output of the feature extraction module is the input of the attribute guidance module; the output of the attribute guidance module is a single age value result.
[0027] Further, the step 3 of training the face age estimation model according to the face preprocessing image set to obtain a trained face age estimation model includes:
[0028] A composite loss function including an error compression sorting loss based on sorting labels and an attribute-guided classification loss is constructed; the face age estimation model is trained according to the preprocessed face age image set and using the composite loss function to obtain a trained face age estimation model.
[0029] The construction includes a composite loss function of error compression ranking loss based on ranking labels and attribute guided classification loss, including:
[0030] L total =L ecr +L attr ;
[0031] Among them, the error compression sorting loss based on sorting labels is:
[0032]
[0033] In the formula, x i is the i-th input sample image, h(x i ) is the single-value age value output by the network model, b k is the starting endpoint of the kth age interval, K is the total number of age categories, N is the total number of samples, and y i is the true ranking label, σ(·) is the S-type activation function;
[0034] The attribute-guided classification loss is used to establish the connection between the attribute and the true label, and is calculated as:
[0035]
[0036] In the formula, α, β and γ are weight coefficients, a(x i ) is the age category for calculation, a i is the image age group label, g(x i ) is the sex classification for calculation, g i is the image gender label, e(x i ) is the racial category for calculation, e i To label the image ethnicity.
[0037] Further, the step 4 of testing the trained face age estimation model according to the test data set to obtain the face age estimation result to achieve face age estimation includes:
[0038] The trained face age estimation model is used to estimate the face age of the test set images, and the mean absolute error (MAE) value between the result and the true label is calculated to achieve face age estimation, and the MAE value is used as an evaluation indicator of the quality of the model.
[0039] The MAE value is calculated as follows:
[0040]
[0041] Another object of the present invention is to provide a face age estimation system using the face age estimation method, the face age estimation system comprising:
[0042] An image set acquisition and preprocessing module is used to acquire and preprocess a face age image set to obtain a preprocessed face age image set;
[0043] Age estimation model building module, used to build a face age estimation model;
[0044] An age estimation model training module is used to train a face age estimation model according to a face preprocessing image set to obtain a trained face age estimation model;
[0045] The face age estimation module is used to test the trained face age estimation model according to the test data set to obtain the face age estimation result to realize face age estimation.
[0046] Another object of the present invention is to provide an electronic device for estimating face age using the method for estimating face age, wherein the electronic device for estimating face age includes an image acquisition device, a display, a graphics processor, a communication interface, a memory, a central processing unit and a communication bus.
[0047] Wherein, the image acquisition device, the display, the graphics processor, the communication interface, the memory and the central processing unit communicate with each other through the communication bus;
[0048] The image acquisition device is used to acquire image data;
[0049] The display is used to display image recognition data;
[0050] The graphics processor is used to calculate image data;
[0051] The memory is used to store computer programs;
[0052] The central processing unit is used to implement the face age estimation method when executing the computer program stored in the memory.
[0053] Another object of the present invention is to provide an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and when the computer program is executed by the processor, the processor performs the following steps:
[0054] A face age image set is obtained and preprocessed to obtain a preprocessed face age image set; a face age estimation model is constructed; the face age estimation model is trained according to the face preprocessed image set to obtain a trained face age estimation model; the trained face age estimation model is tested according to a test data set to obtain a face age estimation result to achieve face age estimation.
[0055] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the following steps:
[0056] A face age image set is obtained and preprocessed to obtain a preprocessed face age image set; a face age estimation model is constructed; the face age estimation model is trained according to the face preprocessed image set to obtain a trained face age estimation model; the trained face age estimation model is tested according to a test data set to obtain a face age estimation result to achieve face age estimation.
[0057] Another object of the present invention is to provide an information data processing terminal, which is used to implement the face age estimation system.
[0058] Combining all the above technical solutions, the advantages and positive effects of the present invention are as follows: the face age estimation method provided by the present invention achieves the goal of robust and efficient face age estimation with high estimation performance indicators by introducing a high-performance multi-scale attention mechanism residual convolution unit, an attribute guidance module, and a composite loss function including error compression sorting loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0060] Figure 1 It is a flow chart of a method for estimating face age provided by an embodiment of the present invention.
[0061] Figure 2 It is a principle diagram of a method for estimating face age provided by an embodiment of the present invention.
[0062] Figure 3 is a structural block diagram of a face age estimation system provided by an embodiment of the present invention;
[0063] In the figure: 1. Image set acquisition and preprocessing module; 2. Age estimation model construction module; 3. Age estimation model training module; 4. Face age estimation module.
[0064] Figure 4 It is a schematic diagram of the structure of the overall network of the face age estimation method provided by an embodiment of the present invention.
[0065] Figure 5 It is a schematic diagram of the basic convolution unit structure of the feature extraction module in the face age estimation method provided by an embodiment of the present invention.
[0066] Figure 6 It is a schematic diagram of the design of age group points corresponding to the error compression sorting loss based on the sorting label provided in an embodiment of the present invention.
[0067] Figure 7 It is a schematic diagram of the structure of an attribute guidance module in the face age estimation method provided by an embodiment of the present invention.
[0068] Figure 8 It is a schematic diagram of the structure of an electronic device for estimating face age provided by an embodiment of the present invention.
[0069] Fig. 9 It is a schematic diagram of the structure of a computer-readable storage medium for estimating face age provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0070] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0071] In view of the problems existing in the prior art, the present invention provides a method, system, electronic device and storage medium for estimating the age of a face. The present invention is described in detail below in conjunction with the accompanying drawings.
[0072] like Figure 1 As shown, the face age estimation method provided by the embodiment of the present invention includes the following steps:
[0073] S101, acquiring a face age image set and performing preprocessing to obtain a preprocessed face age image set;
[0074] S102, constructing a face age estimation model;
[0075] S103, training a face age estimation model according to the face preprocessing image set to obtain a trained face age estimation model;
[0076] S104, testing the trained face age estimation model according to the test data set to obtain a face age estimation result to achieve face age estimation.
[0077] The principle diagram of the face age estimation method provided by the embodiment of the present invention is as follows: Figure 2 shown.
[0078] like Figure 3 As shown, the face age estimation system provided by the embodiment of the present invention includes:
[0079] The image set acquisition and preprocessing module 1 is used to acquire and preprocess the face age image set to obtain the preprocessed face age image set;
[0080] Age estimation model building module 2, used to build a face age estimation model;
[0081] The age estimation model training module 3 is used to train the face age estimation model according to the face preprocessing image set to obtain a trained face age estimation model;
[0082] The face age estimation module 4 is used to test the trained face age estimation model according to the test data set to obtain the face age estimation result to achieve face age estimation.
[0083] The technical solution of the present invention is further described below in conjunction with specific embodiments.
[0084] Example 1
[0085] At present, a series of research has been developed in the field of face age estimation. Since the existing popular estimation models tend to express only one of the age characteristics, that is, the regression algorithm tends to express continuous information, and the classification algorithm tends to express nonlinear information, these algorithms have certain limitations because the two characteristics cannot be well expressed at the same time.
[0086] Based on the above problems, please see Figure 2 , Figure 2 : is a flow chart of a method for estimating the age of a human face based on attribute guidance provided by an embodiment of the present invention. This embodiment provides a method for estimating the age of a human face based on attribute guidance, and the method comprises the following steps:
[0087] Step 1: Obtain a face age image set and preprocess it to obtain a preprocessed face age image set.
[0088] Specifically, this embodiment uses two general face age image sets as face age image sets, namely Morph and UTKFace face age image sets. Among them, the Morph dataset is one of the most popular age estimation datasets, including 55,134 images of 13,000 people, including the age and gender information of the characters; the UTKFace dataset contains more than 20,000 face images from 0 to 116, including the age, gender, and race information of the task. In this embodiment, only the age images from 1 to 100 are retained for experiments.
[0089] In order to achieve better face age estimation, this embodiment first performs face detection, cropping and scaling on the face age image set before age estimation, and randomly divides the images into a training set, a validation set and a test set in a ratio of 8:1:1.
[0090] Among them, for face detection and cropping, this embodiment adopts the detection function and cropping function in the open source library dlib in the field of image processing, only retains images with 1 detected face, and crops the detected faces.
[0091] The cropped images are scaled. In this embodiment, the images are uniformly scaled to an image size of 256×256.
[0092] In this embodiment, each face image in the face age image set is processed as above, thereby obtaining a training set, a verification set and a test set of the preprocessed face age image set.
[0093] Step 2: Construct a feature extraction module, an attribute guidance module, etc. to obtain a face age estimation model.
[0094] Specifically, since the current classification algorithm and sorting algorithm for face age estimation each have corresponding defects, and the attribute information strongly related to age is not effectively utilized, this embodiment proposes a face age estimation model based on attribute guidance. Figure 4 , Figure 4 This is the overall model architecture of the face age estimation method based on attribute guidance provided by an embodiment of the present invention. It can be seen that the constructed face age estimation model includes a feature extraction module and an attribute guidance module connected in sequence.
[0095] The feature extraction module is composed of basic convolution units (i.e., multi-scale attention mechanism residual convolution units) repeatedly connected in the order of [6, 8, 12, 6]. The specific structure of the basic convolution unit is as follows: Figure 5 As shown. The basic convolution unit includes a multi-scale convolution mechanism and a channel attention mechanism; the multi-scale convolution mechanism includes a 1×1 convolution with an output of 1 / 4 of the number of channels, a 3×3 convolution with an output of 1 / 2 of the number of channels, and a 5×5 convolution with an output of 1 / 4 of the number of channels; the channel attention mechanism includes a global pooling layer, a fully connected layer, and an activation layer. The global pooling layer obtains a fully connected layer of the same dimension through a one-dimensional convolution, which is activated by the activation layer as the weight of each channel of the original feature map, and multiplied element by element to obtain a new feature map.
[0096] The attribute guidance module is as follows Figure 6 As shown in the figure, Gender_FC, Age_Group_FC and Ethnicity_FC are the first attribute fully connected layer, and Attribute_FC is the second attribute fully connected layer. The first attribute fully connected layer connects the corresponding output neurons to calculate the attribute loss. At the same time, the first attribute fully connected layer is cascaded and then convolved with a one-dimensional convolution to obtain the second attribute fully connected layer. The second attribute fully connected layer is concatenated with the global fully connected layer to calculate the final age value, and the loss is calculated using the error compression sorting loss based on the sorting label.
[0097] Step 3: Construct a composite loss function of error compression ranking loss based on ranking labels and attribute-guided classification loss;
[0098] Specifically, this embodiment randomly crops the preprocessed face age image set obtained in step 2 into a size of 224×224 and inputs it into the face age estimation model constructed in step 3 for training. Specifically, the training process divides the preprocessed face age image set into three parts: training set, validation set, and test set, and the training set does not overlap with the validation set and the test set. Then, the training set and the validation set are used for training, and a composite loss function including the error compression sorting loss based on the sorting label and the attribute-guided classification loss is constructed at the same time. The specific design is as follows:
[0099] L total =L ecr +L attr (1)
[0100] Among them, the error compression sorting loss based on sorting labels is:
[0101]
[0102] In the above formula, x i is the i-th input sample image, h(x i ) is the single-value age value output by the network model, b k is the starting endpoint of the kth age interval, K is the total number of age categories, N is the total number of samples, and y i is the true ranking label, σ(·) is the S-type activation function. L ecr is the output value h(x i ) minus b k The cross entropy between the binary vector obtained by activating the S-type function and the true label is obtained. After back propagation and chain derivation, the final age regression value will converge to the range of ±0.5 of the true value. The advantage of this is that the use of sorting labels can utilize the continuity information of age, and the use of the corresponding segment points between 1 / 0 can utilize the nonlinear information of age. In addition, the label vector dimension of the sorting algorithm is generally K-1, while the label vector dimension proposed in the present invention is K, so that the information of the starting age value can be utilized.
[0103] The attribute-guided classification loss is used to establish the connection between the attribute and the true label, and is calculated as follows:
[0104]
[0105] Among them, α, β and γ are weight coefficients, a(x i ) is the age category for calculation, a i is the image age group label, g(x i ) is the sex classification for calculation, g i is the image gender label, e(x i ) is the racial category for calculation, e iTo label the image ethnicity.
[0106] Since the different face age datasets used in this embodiment contain different specific attributes, the corresponding attribute guidance loss weights are also different. Specifically, the Morph dataset contains the age and gender labels of the face, so α and β are set to 1 and γ is set to 0 in its application network; while the UTKFace dataset contains the age, gender and race labels of the face, so α, β and γ are all set to 1 in its application network.
[0107] Step 4: training the face age estimation model using the face preprocessed image set to obtain a trained face age estimation model;
[0108] Specifically, this embodiment uses a composite loss function including attribute-guided classification loss and error compression sorting loss based on sorting labels to train the above-mentioned face age estimation model. The training process simultaneously uses the training set and validation set in the preprocessed face age image set. The preferred optimizer in the training is Adam, the learning rate is fixed at 0.0005, and the batch size is 64. The face age estimation model that is finally successfully trained is obtained by continuously saving the model that minimizes the validation set loss.
[0109] Step 5: Use the test data set to test the face age estimation model trained in step 4 to obtain the face age estimation result to achieve face age estimation.
[0110] Specifically, in order to compare this example with other mainstream face age estimation methods, this embodiment performs age estimation on the training set of the trained face age estimation model obtained in step 4 to obtain the estimated age value result, and uses the MAE value between the estimated age value and the true value as the evaluation index to compare the results. The MAE calculation method is:
[0111]
[0112] In summary, the present invention introduces the attribute guidance concept for the problem of face age estimation, and gradually improves the accuracy of the final age estimation by designing an error compression sorting loss based on sorting labels and a feature extraction module that adds multi-scale feature extraction and channel attention. Specifically: first, before performing face age estimation, the face image is detected, cropped and scaled; then the feature extraction module proposed in the present invention is used to extract features; then the idea of attribute guidance is used to establish a connection between a part of the feature fully connected layer and the attribute information, and finally the convolutional feature fully connected layer is spliced with the global fully connected layer, and the age value of the image is calculated using the regression task loss based on sorting labels proposed in the present invention to achieve face age estimation.
[0113] The parameter design of each layer of the feature extraction module in the face age estimation network model in the verification process of this embodiment is specifically shown in Table 1. In the channel attention mechanism of the basic convolution unit (i.e., the multi-scale attention mechanism residual convolution unit), in order to effectively utilize the relationship between adjacent channels, in the basic convolution units of layers 3 to 16, the one-dimensional convolution kernel size is 1×3, and in layers 17 to 34, it is 1×5.
[0114] Table 1. Parameter design of feature extraction module in the face age estimation network model of the present invention
[0115]
[0116] Since the face age dataset used in the verification process of this embodiment contains different attributes, the attribute guidance module in the face age estimation network model uses attribute fully connected layers of different dimensions. Specifically, in the application network of the Morph dataset, the dimensions of Gender_FC, Age_Group_FC, Ethnicity_FC, Attribute_FC and Global_FC are 512, 512, 0, 1024 and 1024 respectively; while in the application network of the UTKFace dataset, the corresponding fully connected layer dimensions are 256, 256, 256, 768 and 1024 respectively.
[0117] The experiment designed in this embodiment is compared and demonstrated from the following three aspects:
[0118] (1) In order to illustrate the effectiveness of the error compression sorting loss based on sorting labels in the face age estimation network model of the present invention, a ResNet34 is constructed as a feature extraction network, and a comparison network that outputs a single-value age value using the loss function is compared with the current mainstream age estimation method. The results are shown in Table 2.
[0119] Table 2. Comparison of mean absolute error (MAE) of the age estimation method in this embodiment with other public methods after using error compression sorting loss term
[0120]
[0121] The comparison results in Table 2 show that when the same feature extraction network output age value is maintained, the composite loss function including the error compression and sorting loss term of the present invention is used, and the MAE values of the estimated results and the true labels on the test set reach the best and second best on the two data sets, specifically 3.61 on UTKFace and 2.42 on Morph. Therefore, this experiment confirms the effectiveness of the error compression and sorting loss term proposed by the present invention and the composite loss function including the loss term in improving the estimation of face age.
[0122] (2) In order to illustrate the effectiveness of the feature extraction module in the face age estimation network model of the present invention, the above-mentioned calculation method using the error compression sorting loss based on the sorting label as the loss function is retained, and a 34-layer repetitive connection network model is constructed using the basic convolutional unit proposed in the present invention to compare the results of ResNet34. In addition, the experiment designed in this embodiment provides the results of using classic features as the extraction network model. Please see Table 3 for details.
[0123] Table 3. Comparison of mean absolute error (MAE) between the feature extraction module used in the age estimation method of this embodiment and other mainstream feature extraction modules
[0124]
[0125] The results in Table 3 show that the method proposed in the present invention, i.e., using the aforementioned composite loss function including the error compression ranking loss and adopting our feature extraction module based on the multi-scale attention mechanism residual unit, has improved test results on two public datasets Morph and UTKFace compared to the method in Table 2 that only uses the error compression ranking loss but does not add our feature extraction module; at the same time, the results also show that the method that retains the error compression ranking loss of the present invention as the loss function and adopts the feature extraction module of the present invention exceeds the results of the feature extraction modules (such as VGG19, DenseNet, ResNet50, ResNet34, etc.) adopted by multiple existing mainstream age estimation network models, and the MAE errors are all smaller. Therefore, this experiment confirms the effectiveness of the feature extraction module based on the multi-scale attention mechanism residual unit proposed in the present invention in improving the age estimation of human faces.
[0126] (3) In order to illustrate the effectiveness of the attribute guidance module in the face age estimation network model proposed in the present invention, an attribute guidance module is added to the above structure for experiments. The results are shown in Table 4.
[0127] Table 4. Comparison of mean absolute error (MAE) before and after the age estimation method in this embodiment with and without the attribute guidance module
[0128]
[0129] The results in Table 4 show that the attribute guidance module proposed in the present invention improves the MAE results of the test sets of the two datasets and reaches the best, specifically reaching 2.36 on Morph and 3.51 on UTKFace. Therefore, this experiment confirms the effectiveness of the attribute guidance module proposed in the present invention in improving face age estimation.
[0130] Judging from the above three face age estimation experimental results, the error compression sorting loss based on sorting labels, the feature extraction module based on multi-scale attention mechanism residual convolution unit, and the attribute guidance module proposed in the face age estimation method of the present invention are all conducive to improving the model performance, making the final age estimation result significantly better than multiple existing mainstream face age estimation network models.
[0131] It can be seen that this embodiment aims at the limitation that the classification bias tends to express the nonlinearity of age and the regression bias tends to express the continuity of age in the traditional face age estimation method, and proposes an error compression sorting loss based on sorting labels, which effectively utilizes the continuity information and nonlinear information of age; this embodiment aims at the problem of feature extraction network performance, designs a residual basic convolution unit with a multi-scale convolution mechanism and a channel attention mechanism, and builds a network with the construction structure of ResNet34 to obtain a feature extraction module; this embodiment aims at the problem that the information strongly related to age is not effectively utilized in the face age estimation problem, and designs an attribute guidance module, which uses the information of the fully connected layer related to the attributes to make the final age result have the expression of relevant labels, so as to improve the performance of the final face age estimation.
[0132] This embodiment proposes a complete set of face age estimation technology based on attribute guidance, which can solve many defects of traditional face age estimation technology, such as the inability to effectively utilize the continuity and nonlinearity of age at the same time, the failure to utilize attribute information strongly related to age, and the low performance of feature extraction network; this embodiment provides new theory and new algorithm support for the practical application of age estimation, making face age estimation technology more practical, reliable and popular; this embodiment can be widely used in auxiliary identity recognition, demography and other application scenarios in outdoor, indoor, network and other environments.
[0133] Embodiment 2
[0134] Based on the above embodiment 1, see Figure 8 , Figure 8 1 is a schematic diagram of a structure of an electronic device for estimating face age based on attribute guidance provided by an embodiment of the present invention. This embodiment provides an electronic device for estimating face age based on attribute guidance, the electronic device comprises an image acquisition device, a display, a graphics processor, a communication interface, a memory, a central processing unit and a communication bus, wherein the image acquisition device, the display, the graphics processor, the communication interface, the memory and the central processing unit communicate with each other through the communication bus;
[0135] The image acquisition device is used to collect face image data;
[0136] The display is used to display facial image recognition data;
[0137] The graphics processor is used to calculate the face image data;
[0138] Memory, used to store computer programs;
[0139] The central processing unit is used to execute the computer program stored in the memory. When the computer program is executed by the processor, the following steps are implemented:
[0140] Step 1: Control the image acquisition device to acquire face images, obtain a face age image set, and perform preprocessing to obtain a preprocessed face age image set.
[0141] Specifically, in step 1 of this embodiment, the face age image set is preprocessed to obtain the preprocessed face age image set, including:
[0142] The face age image set is subjected to face detection, cropping and scaling to obtain the preprocessed face age image set.
[0143] The preprocessed face age image set is randomly divided into a training set, a validation set and a test set in a ratio of 8:1:1.
[0144] Step 2: Build a face age estimation network model, train it, and save the training model parameters.
[0145] Specifically, step 2 of this embodiment includes:
[0146] (1) Constructing a face age estimation model, including a feature extraction module and an attribute guidance module connected in sequence, wherein the feature extraction part includes a basic convolution unit repeatedly connected in sequence by the number of [6, 8, 12, 6], each basic convolution unit includes a multi-scale convolution mechanism and a channel attention mechanism, the multi-scale convolution mechanism includes a 1×1 convolution with an output of 1 / 4 of the number of channels, a 3×3 convolution with an output of 1 / 2 of the number of channels, and a 5×5 convolution with an output of 1 / 4 of the number of channels, and the channel attention mechanism includes a global pooling layer, a fully connected layer, and an activation layer; the attribute guidance module includes a first attribute fully connected layer, a second attribute fully connected layer, and a global fully connected layer, wherein the first attribute fully connected layer is connected to a corresponding number of output neurons, the first attribute fully connected layer is connected to the second attribute fully connected layer, and the second attribute fully connected layer is spliced with the global fully connected layer.
[0147] (2) Construct a composite loss function based on the error compression ranking loss and attribute classification loss of the ranking label;
[0148] (3) Using the face age image set training set and a composite loss function including the error compression ranking loss based on the ranking label and the attribute classification loss to train the face age estimation network model, a trained face age estimation model is obtained, and the model parameters with the smallest MAE with the true label on the validation set are retained;
[0149] Step 3: Use the trained network model to estimate face age.
[0150] Specifically, step 3 of this embodiment inputs the face age preprocessing test image into the trained face age estimation network model to perform face age estimation, and obtains the corresponding age value and the MAE between the real label.
[0151] This embodiment provides an electronic device for estimating face age based on attribute guidance, which can execute the above face age estimation embodiment. Its implementation principle and technical effect are similar and will not be repeated here.
[0152] Embodiment 3
[0153] Based on the above Example 2, see Fig. 9 , Fig. 9 1 is a schematic diagram of a computer-readable storage medium provided by an embodiment of the present invention. This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0154] Step 1: Obtain a face age image set and perform preprocessing to obtain a preprocessed face age image set.
[0155] Specifically, in step 1 of this embodiment, the face age image set is preprocessed to obtain the preprocessed face age image set, including:
[0156] The face age image set is subjected to face detection, cropping and scaling to obtain the preprocessed face age image set.
[0157] The preprocessed face age image set is randomly divided into a training set, a validation set and a test set in a ratio of 8:1:1.
[0158] Step 2: Build a face age estimation network model, train it, and save the training model parameters.
[0159] Specifically, step 2 of this embodiment includes:
[0160] A face age estimation model is constructed, including a feature extraction module and an attribute guidance module connected in sequence. The feature extraction part includes basic convolution units that are repeatedly connected in sequence by [6, 8, 12, 6] times. Each basic convolution unit includes a multi-scale convolution mechanism and a channel attention mechanism. The multi-scale convolution mechanism includes a 1×1 convolution with an output of 1 / 4 of the number of channels, a 3×3 convolution with an output of 1 / 2 of the number of channels, and a 5×5 convolution with an output of 1 / 4 of the number of channels. The channel attention mechanism includes a global pooling layer, a fully connected layer, and an activation layer; the attribute guidance module includes a first attribute fully connected layer, a second attribute fully connected layer, and a global fully connected layer. The first attribute fully connected layer is connected to a corresponding number of output neurons, the first attribute fully connected layer is connected to the second attribute fully connected layer, and the second attribute fully connected layer is spliced with the global fully connected layer.
[0161] Construct a composite loss function based on the error compression ranking loss and attribute classification loss of the ranking label;
[0162] Use the training set of the face age preprocessing training image set and use the composite loss function of the error compression sorting loss based on the sorting label and the attribute classification loss to train the face age estimation network model, and obtain the trained face age estimation model, and retain the model parameters with the smallest MAE with the true label on the validation set;
[0163] Step 3: Use the trained network model to estimate face age.
[0164] Specifically, step 3 of this embodiment inputs the face age preprocessing test image into the trained face age estimation network model to perform face age estimation, and obtains the corresponding age value and the MAE between the real label.
[0165] This embodiment provides a computer-readable storage medium that can execute the above-mentioned face age estimation method embodiment and the above-mentioned face age estimation electronic device embodiment. The implementation principles and technical effects are similar and will not be repeated here.
[0166] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When the use is implemented in whole or in part in the form of a computer program product, the computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL) or wireless (e.g., infrared, wireless, microwave, etc.) mode) to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.
[0167] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for estimating face age, characterized in that: The method for estimating face age comprises the following steps: Step 1, obtaining a face age image set and preprocessing it to obtain a preprocessed face age image set; Step 2: construct a face age estimation model; Step 3: training the face age estimation model according to the face preprocessing image set to obtain a trained face age estimation model; Step 4: testing the trained face age estimation model according to the test data set to obtain a face age estimation result to achieve face age estimation; The step 1 of preprocessing the face age image set to obtain the preprocessed face age image set includes: Performing face detection, cropping and scaling on the face age image set to obtain the preprocessed face age image set; The preprocessed face age image set is randomly divided into a training set, a validation set and a test set according to a certain ratio; The face age estimation model constructed in step 2 includes a feature extraction module and an attribute guidance module; The feature extraction module includes a basic convolution unit that is repeatedly connected in sequence at different times, namely a multi-scale attention mechanism residual convolution unit; the basic convolution unit includes a multi-scale convolution mechanism and a channel attention mechanism that are connected in sequence; the multi-scale convolution mechanism includes convolution layers with different convolution kernel sizes and different numbers of output channels; the channel attention mechanism includes a global pooling layer, a fully connected layer and an activation layer; The attribute guidance module includes a first attribute fully connected layer, a second attribute fully connected layer and a global fully connected layer; wherein the first attribute fully connected layer is connected to a corresponding number of output neurons, the first attribute fully connected layer is connected to the second attribute fully connected layer, and the second attribute fully connected layer is spliced with the global fully connected layer; The input of the feature extraction module is a face image, and the output of the feature extraction module is the input of the attribute guidance module; the output of the attribute guidance module is a single age value result.
2. The method for estimating face age as claimed in claim 1, characterized in that: The step 3 of training the face age estimation model according to the face preprocessing image set to obtain a trained face age estimation model comprises: Constructing a composite loss function including an error compression ranking loss based on a ranking label and an attribute-guided classification loss; training the face age estimation model according to the preprocessed face age image set and using the composite loss function to obtain a trained face age estimation model; The composite loss function including the error compression ranking loss based on the ranking label and the attribute-guided classification loss is constructed as follows: L total =L ecr +L attr ; Among them, the error compression sorting loss based on sorting labels is: In the formula, x i is the i-th input sample image, h(x i ) is the single-value age value output by the network model, b k is the starting endpoint of the kth age interval, K is the total number of age categories, N is the total number of samples, and y i is the true ranking label, σ(·) is the S-type activation function; The attribute-guided classification loss is used to establish the connection between the attribute and the true label, and is calculated as: In the formula, α, β and γ are weight coefficients, a(x i ) is the age category for calculation, a i is the image age group label, g(x i ) is the sex classification for calculation, g i is the image gender label, e(x i ) is the racial category for calculation, e i To label the image ethnicity.
3. The method for estimating face age as claimed in claim 1, characterized in that: The step 4 of testing the trained face age estimation model according to the test data set to obtain the face age estimation result to achieve face age estimation includes: The trained face age estimation model is used to estimate the face age of the test set images, and the mean absolute error (MAE) between the result and the true label is calculated to achieve face age estimation, and the MAE value is used as an evaluation indicator of the quality of the model.
4. A face age estimation system implementing the face age estimation method according to any one of claims 1 to 3, characterized in that: The face age estimation system comprises: An image set acquisition and preprocessing module is used to acquire and preprocess a face age image set to obtain a preprocessed face age image set; Age estimation model building module, used to build a face age estimation model; An age estimation model training module is used to train a face age estimation model according to a face preprocessing image set to obtain a trained face age estimation model; The face age estimation module is used to test the trained face age estimation model according to the test data set to obtain the face age estimation result to realize face age estimation.
5. An electronic device for estimating face age using the face age estimation method according to any one of claims 1 to 3, characterized in that: The electronic device for estimating face age includes an image acquisition device, a display, a graphics processor, a communication interface, a memory, a central processing unit and a communication bus; Wherein, the image acquisition device, the display, the graphics processor, the communication interface, the memory and the central processing unit communicate with each other through the communication bus; The image acquisition device is used to acquire image data; The display is used to display image recognition data; The graphics processor is used to calculate image data; The memory is used to store computer programs; The central processing unit is used to implement the face age estimation method when executing the computer program stored in the memory.
6. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps according to any one of claims 1 to 3.
7. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the steps according to any one of claims 1 to 3.
8. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the face age estimation system as described in claim 4.
Citation Information
Patent Citations
Age estimation method and device
CN107545245A