Face age recognition method and system based on uncertain inhibition network model

By adopting an uncertain suppression network model in face age recognition, combining lightweight Resnet network and L-Net branch network, the image uncertainty evaluation module of batch attention mechanism is used to solve the problems of high model complexity and large differences between individuals, and high-precision and low-complexity face age recognition are achieved.

CN114267060BActive Publication Date: 2025-05-16HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111376110.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-19
Publication Date
2025-05-16
Estimated Expiration
2041-11-19

AI Technical Summary

Technical Problem

The prior art has the problem of high model weight and computational complexity in facial age recognition, and at the same time, there are large differences between individuals, which leads to difficult learning process and low accuracy.

Method used

The face age recognition method based on the uncertainty suppression network model is adopted, image feature extraction and facial contour difference culling is performed through lightweight Resnet network and L-Net branch network, and the image uncertainty evaluation module of the batch attention mechanism is used to reduce interference from uncertain data.

Benefits of technology

The weight and calculation complexity of the recognition model are effectively reduced, while improving the recognition accuracy of face age and reducing the impact of inter-individual differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114267060B_ABST
    Figure CN114267060B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for face age recognition based on an uncertainty suppression network model, including: a training set acquisition step, preprocessing a face image containing an age label to obtain a training image with a label distribution; a model training step, inputting the training image into the uncertainty suppression network model for iterative training, until the prediction accuracy of the uncertainty suppression network model on the validation set no longer increases during the iteration process, and obtaining a weight file of the trained uncertainty suppression network model; an age prediction step, using the trained uncertainty suppression network model weight file to identify the face image to be identified, and obtaining a face age prediction result. The model uses a lightweight Resnet network as the backbone network to extract image features, uses an L‑Net branch network to eliminate the influence of facial contours on age prediction, and simultaneously calculates the uncertainty of the image itself to reduce the interference of uncertainty data on age prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of face attribute recognition, and in particular relates to a face age recognition method and system based on an uncertain inhibition network model. Background Art

[0002] Traditional methods divide the face age estimation task into two stages: feature extraction and age estimation. The feature extraction stage uses manually designed features, which have poor stability and low robustness, and are extremely sensitive to changes in lighting, posture, and expression in natural scenes; the classification and regression models in the age estimation stage have poor feature resolution. In recent years, the rise of deep neural networks has pushed the field of computer vision to a new level of development. Related studies have shown that convolutional neural networks have advantages in constructing high-level semantic features of images that manual features do not have; in addition, this feature extraction capability can adapt to different visual scenes and has strong generalization capabilities, and the face age estimation task is no exception.

[0003] Deep learning models are not conducive to deployment in practical applications due to their huge model weights and computational complexity. Lightweight networks such as MobileNet perform poorly in the field of face attribute recognition, especially in age estimation tasks. How to effectively reduce model weights and computational complexity while ensuring model accuracy has become a hot topic in age estimation tasks. In addition, due to the particularity of facial age labels and the great individual differences, the learning process is difficult and the accuracy is low in practical applications. How to avoid the influence of individual differences and construct a unique age label distribution for each individual is also a difficulty. Summary of the invention

[0004] In view of the above problems, the present invention provides a face age recognition method, system and storage medium based on an uncertain suppression network model, which can effectively reduce the recognition model weight and calculation complexity while improving the face age accuracy.

[0005] A first aspect of the present invention provides a method for face age recognition based on an uncertain inhibition network model, comprising the following steps:

[0006] The training set acquisition step preprocesses the face images with age labels to obtain training images with label distribution;

[0007] Model training step: input the training image into the uncertainty suppression network model for iterative training until the prediction accuracy of the uncertainty suppression network model on the validation set no longer increases during the iteration process, and obtain the weight file of the trained uncertainty suppression network model;

[0008] In the age prediction step, the trained uncertainty suppression network model weight file is used to identify the face image to be identified and obtain the face age prediction result;

[0009] The uncertainty suppression network model includes a lightweight Resnet network for image feature extraction and an L-Net branch network for eliminating facial contour differences based on facial key point information. The model training steps include:

[0010] The smooth first-order regularized loss function value L is calculated using the global pooling feature vector obtained by the lightweight Resnet network and the facial contour feature vector obtained by the L-Net branch network. sa ;

[0011] The KLD loss function value L is calculated using the predicted probability density distribution and label distribution obtained by the uncertainty suppression network model. r ;

[0012] The prediction results obtained by using the uncertainty suppression network model are compared with the real age to calculate the smoothed first-order regularized loss function value L a ;

[0013] Using two smooth first-order regularized loss function values ​​L s , L a and a KLD loss function value L r Perform backpropagation to update the weight file of the uncertainty suppression network model.

[0014] According to some embodiments of the present invention, the uncertainty suppression network model also includes an image uncertainty assessment module based on a batch attention mechanism, the image uncertainty assessment module includes a fully connected layer, a batch pooling layer and a normalization layer for performing global feature conversion on each training image, the image uncertainty assessment module adopts a Query-Key matching mechanism to calculate the uncertainty score of the training image sample, the Query is obtained by summing and averaging the global pooling features of all training images in the batch samples, each sample corresponds to a Key value vector, which is obtained through a fully connected layer with input and output dimensions of 512. The image uncertainty assessment module is mainly used to reduce the interference of uncertain images to the network during the training process and accelerate the convergence speed of the network.

[0015] According to some embodiments of the present invention, the preprocessing of the face image containing the age label specifically comprises the following steps:

[0016] Label encoding, converting the age label into a label distribution. Specifically, it presupposes that the label distribution of face age conforms to the normal distribution, sets the mean of the normal distribution as the true age label, and the variance as the prior. The age label is converted from a specific value to its corresponding label distribution. Its beneficial effect is to smooth the change process of face age and effectively learn the correlation between age labels.

[0017] Data enhancement: Use key point detection tools to obtain key point information and pupil coordinates of the face, and perform face alignment based on the pupil coordinates. The beneficial effect is that data can be expanded through data enhancement to improve the robustness of the model.

[0018] According to some embodiments of the present invention, before the model training step, the weight file pre-trained on the MS-Celeb-1M data is first used and loaded into the lightweight Resnet network, and the Kaiming Normal is used to initialize the weights of the L-Net branch network and the image uncertainty assessment module. The beneficial effect is that it can help the model converge quickly and alleviate the problem of insufficient training data.

[0019] According to some embodiments of the present invention, the lightweight Resnet network is based on the Resnet18 network and is improved, including:

[0020] Resize the input image to 112*112;

[0021] Reduce the number of repetitions of all residual blocks in the Resnet18 network to 1;

[0022] Reduce the number of channels of all convolutional layers in the Resnet18 network to 1 / 2 of the original number.

[0023] According to some embodiments of the present invention, the L-Net branch network is composed of a fully connected network and an orthogonal separation plane. The L-Net branch network uses the key point coordinates of the human face as input features. The orthogonal separation plane is composed of a feature vector extracted by the facial key point coordinates and a feature vector extracted by a lightweight Resnet network, which is a dot product operation between the two feature vectors. The beneficial effect of using the L-Net branch network is that it can construct the facial contour of the human face through key point auxiliary information, and then learn the differences in facial contours between individuals, thereby achieving the effect of eliminating differentiated information.

[0024] According to some embodiments of the present invention, the KLD loss function value calculation steps are as follows:

[0025] The first iteration calculates the KLD loss function value between the predicted probability density distribution of the uncertain inhibition network model and the label distribution;

[0026] Record the prediction results of the uncertainty inhibition network model during each iteration;

[0027] Starting from the second iteration, the first KLD loss function value between the predicted probability density distribution of the uncertain inhibition network model and the label distribution is calculated, and the second KLD loss function value between the predicted probability density distribution of the uncertain inhibition network model and the predicted probability density distribution of the uncertain inhibition network model in the previous iteration is calculated, and the first KLD loss function value and the second KLD loss function value are balanced through hyperparameters. Based on the KLD loss function, in the training process, not only the label distribution after label encoding is used as the learning target, but also the knowledge learned by the model itself is taken into account. Its beneficial effect is that it can correct the difference between the variance set by the prior and the true variance of the sample, while reducing the interference of the age label error on the learning direction of the model.

[0028] According to some embodiments of the present invention, the age prediction step is specifically:

[0029] The uncertainty suppression network model is set to prediction mode, in which the L-Net branch network and image uncertainty assessment module are not involved;

[0030] Load the weight file obtained when the uncertainty suppression network model training is completed;

[0031] Input the face image to be recognized into the loaded uncertainty suppression network model to obtain the age prediction probability density distribution;

[0032] The face age prediction result output by the uncertainty suppression network model is obtained by calculating the expected value of the age prediction probability density distribution.

[0033] A second aspect of the present invention provides a face age recognition system based on an uncertain inhibition network model, comprising:

[0034] The training set acquisition module is used to pre-process the face images with age labels to obtain training images with label distribution;

[0035] A model training module is used to input the training images into the uncertainty suppression network model for iterative training until the prediction accuracy of the uncertainty suppression network model on the validation set no longer increases during the iteration process, thereby obtaining a weight file of the trained uncertainty suppression network model;

[0036] The age prediction module is used to use the trained uncertainty suppression network model weight file to identify the face image to be identified and obtain the face age prediction result;

[0037] The uncertainty suppression network model includes a lightweight Resnet network for image feature extraction, an L-Net branch network for eliminating facial contour differences based on facial key point information, and an image uncertainty assessment module based on a batch attention mechanism. The model training module specifically includes:

[0038] The smooth first-order regularized loss function value L is calculated using the global pooling feature vector obtained by the lightweight Resnet network and the facial contour feature vector obtained by the L-Net branch network. sa ;

[0039] The KLD loss function value L is calculated using the predicted probability density distribution and label distribution obtained by the uncertainty suppression network model. r ;

[0040] The prediction results obtained by using the uncertainty suppression network model are compared with the real age to calculate the smoothed first-order regularized loss function value L a ;

[0041] Using two smooth first-order regularized loss function values ​​L s , L a and a KLD loss function value L r Perform backpropagation to update the weight file of the uncertainty suppression network model.

[0042] According to a third aspect of the present invention, a computer-readable storage medium is provided, on which instructions are stored, characterized in that when the instructions are executed by a processor, the processor executes the face age recognition method based on the uncertain inhibition network model as described above.

[0043] The present invention provides a face age recognition method, system and storage medium based on an uncertainty suppression network model. First, face image sample data with age labels are used to train the uncertainty suppression network model, and then the trained network model is used to estimate the age of the face image to be identified. The uncertainty suppression network model uses label distribution as the basic mode of age estimation, and converts the age estimation into a prediction of its probability density distribution; the uncertainty suppression network model uses a lightweight ResnNet network as the backbone network to extract image features, and then uses the L-Net branch network to eliminate the influence of facial contour on age estimation. At the same time, the image uncertainty assessment module based on the batch attention mechanism calculates the uncertainty of the image itself to reduce the interference of uncertainty data on age estimation. The present invention proposes an uncertainty assessment algorithm for images based on a batch attention mechanism, and at the same time, in order to be able to learn the corresponding probability density distribution for each image, an improved KLD loss function is proposed for supervised learning. The beneficial effects finally achieved are:

[0044] 1. Use the weight file trained in MS-Celeb-1M, load it into the lightweight Resnet backbone network, and use KaimingNormal to initialize the L-Net and image uncertainty assessment modules. This can help the model converge quickly and alleviate the problem of insufficient training data.

[0045] 2. Convert the age label from a specific value to its corresponding label distribution, which has the beneficial effect of smoothing the change process of face age and effectively learning the correlation between age labels;

[0046] 3. Use key point detection tools to obtain key point information and pupil coordinates of the face, and perform face alignment based on the pupil coordinates. This has the beneficial effect of being able to expand data through data enhancement and improve the robustness of the model;

[0047] 4. The beneficial effect of using the L-Net branch network is that it can construct the facial contour of the face through the auxiliary information of key points, and then learn the differences in facial contours between individuals, achieving the effect of eliminating differentiated information;

[0048] 5. Based on the KLD loss function, in the training process, not only the label distribution after label encoding is used as the learning target, but also the knowledge learned by the model itself is taken into account. Its beneficial effect is that it can correct the difference between the variance set by the prior and the true variance of the sample, and at the same time reduce the interference of age label errors on the learning direction of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 1. It is a flow chart of a method for face age recognition based on an uncertain inhibition network model in an embodiment of the present invention;

[0050] Figure 2 Schematic diagram of the L-Net branch network structure for eliminating facial contour differences based on facial key point information in an embodiment of the present invention;

[0051] Figure 3 This is a flow chart of uncertainty suppression network model training in an embodiment of the present invention;

[0052] Figure 4 This is a flowchart of age prediction in an embodiment of the present invention;

[0053] Figure 5 Schematic diagram of the structure of a face age recognition system based on an uncertain inhibition network model in an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to further explain the technical solution of the present invention in detail, this embodiment is implemented on the premise of the technical solution of the present invention, and provides a detailed implementation method and specific steps.

[0055] like Figure 1 As shown, a face age recognition method based on an uncertain inhibition network model is proposed, including the following steps:

[0056] S01, a training set acquisition step, preprocessing the face images with age labels to obtain training images with label distribution;

[0057] In the specific implementation process, the pre-processed face image with key points and age annotations is obtained, and the process includes:

[0058] Label encoding, converting the age label into a label distribution. Specifically, it presupposes that the label distribution of the face age conforms to the normal distribution, sets the mean of the normal distribution to the true age label, and the variance is a priori. In the specific implementation, the initial value of the variance is 3, and the age label is converted from a specific value to its corresponding label distribution. The uncertainty suppression network model will correct the gap between the variance set by the prior and the true value during the training process. Its beneficial effect is to smooth the change process of the face age and effectively learn the correlation between age labels.

[0059] Data enhancement: Use the key point detection tool to obtain the 68 key points of the face and the pupil coordinates. By calculating the angle between the straight line where the pupil coordinates are located and the horizontal line, the face is aligned. The beneficial effect is that data can be expanded through data enhancement to improve the robustness of the model. norm Use the following formula for regularization, where P represents the key point coordinates, p c represents the center point between pupils, d represents the distance between pupils,

[0060]

[0061] In addition, the preprocessing also includes face image color, brightness adjustment, image normalization, random center cropping and proportional scaling.

[0062] S02, a model training step, inputting the training image into the uncertainty suppression network model for iterative training, until the prediction accuracy of the uncertainty suppression network model on the validation set no longer increases during the iteration process, and obtaining a weight file of the trained uncertainty suppression network model;

[0063] In the specific implementation process, before the model training step, the weight file pre-trained on the MS-Celeb-1M data is first used to load into the lightweight Resnet network, and Kaiming Normal is used to initialize the weights of the L-Net branch network and the image uncertainty assessment module. The beneficial effect is that it can help the model converge quickly and alleviate the problem of insufficient training data. At the same time, the model contains facial features of faces in various scenarios, which can improve the robustness of the model. MS-Celeb-1M is a face dataset publicly available by Microsoft, and Kaiming Normal is a model weight initialization method.

[0064] Specifically, the uncertainty suppression network model structure includes three parts: a lightweight Resnet network for image feature extraction, an L-Net branch network for eliminating facial contour differences based on facial key point information, and an image uncertainty assessment module based on a batch attention mechanism. The lightweight Resnet network is used as the backbone network for image feature extraction. The lightweight Resnet network is based on the Resnet18 network and is improved, including adjusting the input image size to 112*112; reducing the number of repetitions of all residual blocks in the Resnet18 network to 1; and reducing the number of channels of all convolutional layers in the Resnet18 network to 1 / 2 of the original. The lightweight Resnet network gives the model the ability to learn identity mappings through skip-layer connections between convolutional layers, greatly improving the feature extraction capability of the convolutional network, and making the number of weights of the final model only 1 / 10 of that of the Resnet18 network. The L-Net branch network that eliminates facial contour differences based on facial key point information consists of a fully connected network and orthogonal separation planes, such as Figure 2As shown in the figure, the fully connected layer is divided into four layers with dimensions of 136, 512, 256, and 512 respectively. 136 is a one-dimensional vector of 68 facial key points, and the middle layer is 256 dimensions, which can effectively remove noise information. The L-Net branch network uses the coordinates of 68 key points on the face as input features and normalizes the key point information according to the pupil distance. The orthogonal separation plane is composed of the feature vector extracted by the facial key point coordinates and the feature vector extracted by the lightweight Resnet network. It is a dot product operation between the two feature vectors. By reducing the correlation between the two feature vectors, the facial contour information is removed from the global features. The beneficial effect of using the L-Net branch network is that it can construct the facial contour of the face through the auxiliary information of the key points, and then learn the differences in facial contours between individuals, so as to achieve the effect of eliminating differential information. The image uncertainty assessment module includes a fully connected layer, a batch pooling layer, and a normalization layer for global feature conversion of each training image in the Mini-Batch. The image uncertainty assessment module is similar to the Attention mechanism. It uses the Query-Key matching mechanism to calculate the uncertainty score of the training image sample. The uncertainty of each sample in the batch data is calculated by the dot product operation of the Query and the Key. The Query is obtained by summing and averaging the global pooling features of all training images in the batch sample. Each sample will correspond to a Key value vector, which is obtained by a fully connected layer with an input and output dimension of 512. Specifically, a fully connected layer converts the global pooling features of each image in the batch to obtain the Key value, where the input and output dimensions of the fully connected layer are both 512. The image uncertainty assessment module is mainly used to reduce the interference of uncertain images to the network during the training process and accelerate the convergence speed of the network. The image uncertainty assessment module uses the Query and Key matching mechanism to filter out samples with low image quality and high uncertainty from a batch of samples, reducing their guiding role in the weight update process. The uncertainty score of the image sample is obtained by the dot product operation of the query and the key through the sigmoid activation function. It is used to measure the degree of deviation of the input sample from the overall distribution of the batch samples. The closer the score is to 1, the more consistent the sample is with the overall distribution, while the score close to 0 indicates that the sample does not conform to the overall distribution and has a large uncertainty.

[0065] The specific steps of uncertainty suppression network model training are as follows: Figure 3As shown in the figure, the specific process of obtaining the prediction result of the training image after the network forward propagation is as follows: This process is the network forward propagation calculation stage. According to the characteristics of the network structure, the image first passes through the feature extraction lightweight Resnet network, and then is sent to the L-Net branch network and the image uncertainty evaluation module respectively, and performs orthogonal calculations with the facial contour features learned in the L-Net branch network; in the image uncertainty evaluation module, the batch attention mechanism is used to calculate the deviation degree of the sample from the overall distribution: Specifically, assuming X = {x 1 ,x 2 ,x 3 ,...,x b} is the input training image data of the model in each iteration; F = {f 1 ,f 2 ,f 3 ,...,f b} represents the facial image feature vector obtained by the feature extraction lightweight Resnet network; Represents the mean of all feature vectors; then for each input image data x i , and its uncertainty is calculated by the following formula: i =sigmoid(f a ·W T f i ), where W is the feature vector f i Convert to the weight matrix of the Key value vector, f a Represents the Query vector. In addition, the global pooling features of the image in the lightweight Resnet network will be converted into the number of categories corresponding to the age label through the fully connected layer, and then converted into the predicted probability density distribution through the softmax function; in order to calculate the predicted age of the model, it is necessary to calculate the expectation of the predicted probability density distribution.

[0066] like Figure 3 The specific steps of training the uncertainty suppression network model are shown in the figure. The specific process of calculating the loss value of the prediction result according to the real label is as follows: each image corresponds to a real age and a prior distribution. The real age is the encoding of the age label. It is assumed that the distribution of face age labels conforms to the normal distribution. The age label of each image is taken as the mean of the normal distribution, and the variance is set by prior. In the specific implementation, 3 is used as the variance of the normal distribution to calculate the label distribution. Specifically, in the age range of i∈[1,100], each age corresponds to a probability value ρ i , the specific encoding formula is as follows:

[0067]

[0068] Where y is the real age corresponding to the training image, and σ is the artificially set normal distribution variance.

[0069] The specific process of calculating the loss function is:

[0070] The global pooling feature vector obtained by the lightweight Resnet network and the facial contour feature vector obtained by the L-Net branch network are used to calculate the smooth first-order regularization loss function value; the output feature of the last fully connected layer of the L-Net branch network and the global pooling feature in the lightweight Resnet network are used to construct an orthogonal loss, specifically calculating the dot product of the two 521 features and using a smooth first-order regularization for loss calculation. This orthogonal loss is used to remove facial contour information from the global pooling feature of the face image and constrain the correlation between the two feature vectors. Specifically, assuming f global represents the global pooled feature vector of the face image obtained by the lightweight Resnet network, f shape Represents the facial contour feature vector output by the last fully connected layer of the L-Net branch network, s = f global ·f shape For the click operation between two vectors, the loss function is as follows:

[0071]

[0072] The KLD loss function value is calculated using the predicted probability density distribution and label distribution of the uncertainty inhibition network model; the predicted probability density distribution and the label distribution obtained by age label encoding are calculated using the KLD loss function. At the same time, in the tth iteration of the model, the predicted output at time t-1 is used as the learning target and the KLD loss is also used to smooth the learning process of the sample. Specifically, the KLD loss function value between the predicted probability density distribution and the label distribution of the uncertainty inhibition network model is calculated in the first iteration; the prediction results of the uncertainty inhibition network model in each iteration are recorded; starting from the second iteration, the first KLD loss function value between the predicted probability density distribution of the uncertainty inhibition network model and the label distribution is calculated, and the second KLD loss function value between the predicted probability density distribution of the uncertainty inhibition network model and the predicted probability density distribution of the uncertainty inhibition network model in the previous iteration is calculated, and the first KLD loss function value and the second KLD loss function value are balanced through hyperparameters. Based on the KLD loss function, in the training process, not only the label distribution after label encoding is used as the learning target, but also the knowledge learned by the model itself is taken into account. The beneficial effect is that it can correct the difference between the variance set by the prior and the true variance of the sample, while reducing the interference of age label errors on the learning direction of the model. The definition of the KLD loss function is as follows:

[0073]

[0074] Where η is a hyperparameter used to balance the loss function, and is preferably set to 0.1. The label distribution obtained by encoding the age label is still the dominant factor to guide the learning process of the network. t-1 Represents the predicted probability density distribution of the sample at time t-1, p t represents the predicted probability density distribution of the sample at time t, p represents the label distribution of the sample after label encoding, k is the specific age label, and p k Represents the probability that label k can be used as a sample to estimate age.

[0075] During the forward propagation process, the model's predicted age y' is calculated by the expectation of its predicted probability density distribution, as follows:

[0076]

[0077] Where K represents the number of age label categories.

[0078] Finally, a smooth first-order regularized loss function is used to calculate the difference between the model's predicted age y' and the actual age y. The formula is as follows:

[0079]

[0080] The total loss value is composed of the above two smooth first-order regularized loss function values ​​L s , L a and a KLD loss function value L r Taking into account the size of the loss function and the balance of task importance, the global pooling feature vector obtained by the lightweight Resnet network and the facial contour feature vector obtained by the L-Net branch network calculate the smooth first-order regularized loss function value L s The weight is set to 0.5, and the other two L a , L r Set to 1. Multiply the obtained sum loss value by the uncertainty score α calculated by the image uncertainty assessment module. i , and get the final loss function value.

[0081]

[0082] Where N is the number of training image samples.

[0083] like Figure 3The specific steps of uncertainty suppression network model training are shown, in which the back propagation algorithm is used to update the weights. The weight file of the uncertainty suppression network model is updated by back propagating two smoothed first-order regularized loss function values ​​and a KLD loss function value. The specific process is: the total loss function value calculated above is passed layer by layer according to the chain rule until the input layer, and the weight parameters of the model are updated during the transmission process. The global optimal solution is approached through multiple iterations, and each iteration involves dividing all training samples into multiple batches for forward propagation and gradient update.

[0084] like Figure 3 The specific steps of uncertainty suppression network model training are shown in the figure, in which the accuracy calculation of the validation set is performed after each iteration, that is, after completing one forward propagation and gradient update of all training samples, the validation set data is also divided into multiple batches, and forward propagation is performed without gradient update. The prediction results of the model on the validation set are calculated, and the mean absolute error and cumulative accuracy are used as the measurement indicators on the validation set. The formula is defined as follows:

[0085]

[0086] Where y is the true age label, N is the number of training samples, {E≤3} represents the number of training samples with age prediction error within 3 years, MAE is the mean absolute error, CA E is the cumulative accuracy.

[0087] When the model's mean absolute error MAE and cumulative accuracy CA on the validation set E When there is no more improvement, the training ends and the final model weight file is obtained. At this point, the training process is complete.

[0088] S03, age prediction step, using the trained uncertainty suppression network model weight file to identify the face image to be identified, and obtain the face age prediction result;

[0089] Specifically, the age prediction steps are as follows: Figure 4 As shown, specifically:

[0090] The uncertainty suppression network model is set to prediction mode, in which the L-Net branch network and image uncertainty assessment module are not involved;

[0091] Load the weight file obtained when the uncertainty suppression network model training is completed;

[0092] The face image to be recognized is input into the loaded uncertainty suppression network model, the image is forward propagated and the age prediction probability density distribution is obtained. Here the image size is scaled to 112*112. In addition, no other transformation operations are required. Age key points are not required as auxiliary information in the prediction process.

[0093] The face age prediction result output by the uncertainty suppression network model is obtained by calculating the expected value of the age prediction probability density distribution.

[0094] Below, refer to Figure 5 To describe the embodiment of the present disclosure Figure 1 The system corresponding to the method shown is a face age recognition system based on an uncertainty suppression network model, wherein the system 100 includes: a training set acquisition module 101, which is used to preprocess face images containing age labels to obtain training images with label distribution; a model training module 102, which is used to input the training images into the uncertainty suppression network model for iterative training until the prediction accuracy of the uncertainty suppression network model on the validation set no longer increases during the iteration process, and obtain the weight file of the trained uncertainty suppression network model; an age prediction module 103, which is used to use the trained uncertainty suppression network model weight file to identify the face image to be identified and obtain the face age prediction result; wherein the uncertainty suppression network model includes a lightweight The model training module 102 comprises a lightweight Resnet network, an L-Net branch network for eliminating facial contour differences based on facial key point information, and an image uncertainty assessment module based on a batch attention mechanism. The model training module 102 specifically includes: calculating a smoothed first-order regularized loss function value using the global pooling feature vector obtained by the lightweight Resnet network and the facial contour feature vector obtained by the L-Net branch network; calculating a KLD loss function value using the predicted probability density distribution and label distribution obtained by the uncertainty suppression network model; calculating a smoothed first-order regularized loss function value using the predicted result obtained by the uncertainty suppression network model and the real age; and updating the weight file of the uncertainty suppression network model by back-propagating two smoothed first-order regularized loss function values ​​and one KLD loss function value. In addition to the above modules, the system 100 may also include other components. However, since these components are irrelevant to the contents of the embodiments of the present disclosure, their illustrations and descriptions are omitted here.

[0095] The embodiment of the present invention may also be implemented as a computer-readable storage medium. The computer-readable storage medium according to the embodiment stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the face age recognition method based on the uncertain inhibition network model according to the embodiment of the present invention described with reference to the above figures may be executed.

[0096] In summary, the present invention provides a method, system and storage medium for face age recognition based on an uncertainty suppression network model. First, face image sample data with age labels are used to train the uncertainty suppression network model, and then the trained network model is used to estimate the age of the face image to be identified. The uncertainty suppression network model uses label distribution as the basic mode of age estimation, and converts the age estimation into a prediction of its probability density distribution; the uncertainty suppression network model uses a lightweight ResnNet network as the backbone network to extract image features, and then uses the L-Net branch network to eliminate the influence of facial contour on age estimation. At the same time, the uncertainty of the image itself is calculated based on the image uncertainty evaluation module to reduce the interference of uncertainty data on age estimation. The present invention proposes an uncertainty evaluation algorithm for images based on a batch attention mechanism, and in order to be able to learn the corresponding probability density distribution for each image, an improved KLD loss function is proposed for supervised learning. The beneficial effects finally achieved are: using the weight file trained in MS-Celeb-1M, loading it into the lightweight Resnet backbone network, and using KaimingNormal to initialize the L-Net and image uncertainty assessment modules, which can help the model converge quickly and alleviate the problem of insufficient training data; converting the age label from a specific value to its corresponding label distribution, which has the beneficial effect of smoothing the change process of face age and effectively learning the correlation between age labels; using the key point detection tool to obtain the key point information and pupil coordinates of the face, and aligning the face according to the pupil coordinates, which has the beneficial effect of being able to perform data augmentation and improve the robustness of the model; the beneficial effect of using the L-Net branch network is that it can construct the facial contour of the face through the auxiliary information of the key points, and then learn the differences in facial contours between individuals, achieving the effect of eliminating differentiated information; based on the KLD loss function, in the training process, not only the label distribution after label encoding is used as the learning target, but also the knowledge learned by the model itself is taken into account. The beneficial effect is that it can correct the difference between the variance set by the prior and the true variance of the sample, while reducing the interference of age label errors on the learning direction of the model.

[0097] In this document, the terms "comprises," "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, such that a step or method that includes a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such step or method.

[0098] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A face age recognition method based on an uncertain inhibition network model, characterized in that: The following steps are involved: The training set acquisition step preprocesses the face images with age labels to obtain training images with label distribution; Model training step: input the training image into the uncertainty suppression network model for iterative training until the prediction accuracy of the uncertainty suppression network model on the validation set no longer increases during the iteration process, and obtain the weight file of the trained uncertainty suppression network model; In the age prediction step, the trained uncertainty suppression network model weight file is used to identify the face image to be identified and obtain the face age prediction result; The uncertainty suppression network model includes a lightweight Resnet network for image feature extraction and an L-Net branch network for eliminating facial contour differences based on facial key point information. The model training steps include: The smooth first-order regularized loss function value L is calculated using the global pooling feature vector obtained by the lightweight Resnet network and the facial contour feature vector obtained by the L-Net branch network. sa ; The KLD loss function value L is calculated using the predicted probability density distribution and label distribution obtained by the uncertainty suppression network model. r ; The prediction results obtained by using the uncertainty suppression network model are compared with the real age to calculate the smoothed first-order regularized loss function value L a ; Using two smooth first-order regularized loss function values ​​L s , L a and a KLD loss function value L r Perform backpropagation to update the weight file of the uncertainty suppression network model.

2. According to claim 1, a method for face age recognition based on an uncertain inhibition network model is characterized in that: The uncertainty suppression network model also includes an image uncertainty assessment module based on a batch attention mechanism, which includes a fully connected layer, a batch pooling layer and a normalization layer for performing global feature conversion on each training image. The image uncertainty assessment module uses a Query-Key matching mechanism to calculate the uncertainty score of the training image sample. The Query is obtained by summing and averaging the global pooling features of all training images in the batch sample. Each sample corresponds to a Key value vector, which is obtained through a fully connected layer with input and output dimensions of 512.

3. According to claim 1, a method for face age recognition based on an uncertain inhibition network model is characterized in that: The face image containing the age label is preprocessed, and the specific steps include: Label encoding, converting the age label into a label distribution. Specifically, it presupposes that the label distribution of the face age conforms to the normal distribution, sets the mean of the normal distribution as the true age label, and the variance as the prior, converting the age label from a specific value into its corresponding label distribution; Data enhancement: Use key point detection tools to obtain key point information and pupil coordinates of the face, and perform face alignment based on the pupil coordinates.

4. According to claim 2, a method for face age recognition based on an uncertain inhibition network model is characterized in that: Before the model training step, the weight file pre-trained on the MS-Celeb-1M data is loaded into the lightweight Resnet network, and Kaiming Normal is used to initialize the weights of the L-Net branch network and the image uncertainty assessment module.

5. The method for face age recognition based on the uncertain inhibition network model according to claim 1 is characterized in that: The lightweight Resnet network is based on the Resnet18 network and is improved, including: Resize the input image to 112*112; Reduce the number of repetitions of all residual blocks in the Resnet18 network to 1; Reduce the number of channels of all convolutional layers in the Resnet18 network to 1 / 2 of the original number.

6. The method for face age recognition based on the uncertain inhibition network model according to claim 1 is characterized in that: The L-Net branch network is composed of a fully connected network and an orthogonal separation plane. The L-Net branch network uses the key point coordinates of the face as input features. The orthogonal separation plane is composed of a feature vector extracted from the facial key point coordinates and a feature vector extracted from a lightweight Resnet network, and is a dot product operation between the two feature vectors.

7. The method for face age recognition based on the uncertain inhibition network model according to claim 1 is characterized in that: The KLD loss function value calculation steps are as follows: The first iteration calculates the KLD loss function value between the predicted probability density distribution of the uncertain inhibition network model and the label distribution; Record the prediction results of the uncertainty inhibition network model during each iteration; Starting from the second iteration, the first KLD loss function value between the predicted probability density distribution of the uncertain inhibition network model and the label distribution is calculated, and the second KLD loss function value between the predicted probability density distribution of the uncertain inhibition network model and the predicted probability density distribution of the uncertain inhibition network model in the previous iteration is calculated, and the first KLD loss function value and the second KLD loss function value are balanced through hyperparameters.

8. The method for face age recognition based on the uncertain inhibition network model according to claim 2 is characterized in that: The age prediction step is specifically as follows: The uncertainty suppression network model is set to prediction mode, in which the L-Net branch network and image uncertainty assessment module are not involved; Load the weight file obtained when the uncertainty suppression network model training is completed; Input the face image to be recognized into the loaded uncertainty suppression network model to obtain the age prediction probability density distribution; The face age prediction result output by the uncertainty suppression network model is obtained by calculating the expected value of the age prediction probability density distribution.

9. A face age recognition system based on an uncertain inhibition network model, characterized in that: include: The training set acquisition module is used to pre-process the face images with age labels to obtain training images with label distribution; A model training module is used to input the training images into the uncertainty suppression network model for iterative training until the prediction accuracy of the uncertainty suppression network model on the validation set no longer increases during the iteration process, thereby obtaining a weight file of the trained uncertainty suppression network model; The age prediction module is used to use the trained uncertainty suppression network model weight file to identify the face image to be identified and obtain the face age prediction result; The uncertainty suppression network model includes a lightweight Resnet network for image feature extraction and an L-Net branch network for eliminating facial contour differences based on facial key point information. The model training module specifically includes: The smooth first-order regularized loss function value L is calculated using the global pooling feature vector obtained by the lightweight Resnet network and the facial contour feature vector obtained by the L-Net branch network. sa ; The KLD loss function value L is calculated using the predicted probability density distribution and label distribution obtained by the uncertainty suppression network model. r ; The prediction results obtained by using the uncertainty suppression network model are compared with the real age to calculate the smoothed first-order regularized loss function value L a ; Using two smooth first-order regularized loss function values ​​L s , L a and a KLD loss function value L r Perform backpropagation to update the weight file of the uncertainty suppression network model.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Face recognition method and device, computer equipment and storage medium

    CN109522872A

  • Real-time license plate recognition method based on deep learning in complex scene

    CN110619327A