Face Aging System and Establishment Method Based on Pixel-Level Alignment Generative Adversarial Network

Through a face aging system based on pixel-level alignment generation adversarial network, combined with orderly regression and ResNet50, the problems of poor face aging image quality and incomplete feature retention in the prior art are solved, and high-quality face aging image generation and identity information retention are achieved.

CN115035007BActive Publication Date: 2025-05-27SHANGHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210478631.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-05
Publication Date
2025-05-27
Estimated Expiration
2042-05-05

AI Technical Summary

Technical Problem

The existing face aging methods produce poor images of the images, incomplete preservation of facial features, and high-quality face images of the age group cannot be synthesized.

Method used

A face aging system based on pixel-level alignment generation adversarial network is adopted to extract face features through a deeper network structure, combining an age estimation model based on orderly regression and ResNet50 as feature extractors to ensure that the generated aged face is within the target age range, and reduce background artifacts through pixel-level constraints to improve image quality.

Benefits of technology

It realizes the generation of high-quality face aging images, ensures the retention of facial identity information, and meets technical needs in various application fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035007B_ABST
    Figure CN115035007B_ABST
Patent Text Reader

Abstract

The present invention relates to a face aging system based on pixel-level alignment generative adversarial network and an establishment method thereof, which sequentially comprises a task module, a face image preprocessing module, a face feature extraction module, an age estimation module and a face synthesis module; an age estimation method based on ordered regression is adopted to ensure that the synthesized face is within a target age range, a residual network structure is introduced into a convolutional neural network to extract face features to ensure the identity consistency of the face, and pixel-level loss is also added to reduce the error caused by background noise and improve the quality of the synthesized image. Compared with the existing face aging methods at this stage, the pixel-level alignment generative adversarial network model proposed by the present invention has obvious advantages in practical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image recognition technology, and in particular to a face aging system and a building method based on a pixel-level alignment generative adversarial network. Background Art

[0002] In recent years, the synthesis of faces of different ages using a single picture has been widely applied in many fields, such as criminal case investigation, movie special effects production, face-swapping software, etc. The specific implementation is that given a source image and specifying the age of the generated image, a new image with the identity information of the source image and the target age information can be generated. With the emergence of artificial intelligence technology, deep learning methods have many excellent technologies in computational vision, especially in face processing. However, in existing methods, there are many deficiencies. Some methods generate pictures that only contain the target age information, while the identity information is not maintained; some methods maintain the identity information and can generate faces that meet the target age, but the picture quality is not high, and there are even noises and artifacts. There are also some methods that cannot synthesize face images of multiple age groups.

[0003] Recently, the emergence of generative adversarial networks has provided a new direction for the research of face image generation. The main reason is that generative adversarial networks can synthesize different face images according to different feature vectors. As long as the embedding vector information that meets the target age and face identity characteristics can be extracted, generative adversarial networks can be used to synthesize face images that meet the requirements. However, on the one hand, due to the problem of uneven distribution of face data in different age stages, many methods cannot synthesize high-quality face images well. On the other hand, most of the research on faces is at the image level and does not pay much attention to the processing of the picture background, which will cause more background artifacts and lower image quality in the generated face images. Therefore, existing methods cannot construct a method for synthesizing high-quality face aging images and cannot meet the technical requirements of various application fields. Summary of the Invention

[0004] Aiming at the problems of poor generation quality and incomplete retention of face features in the results obtained by existing face aging methods, a face aging method based on a pixel-level alignment generative adversarial network is proposed. A deeper network structure is used to extract face features for retaining face identity information. An age estimation model based on ordinal regression is adopted to ensure that the generated aged faces are within the required age range. Finally, pixel-level constraints are used to reduce the influence of background artifacts and improve the quality of the aged pictures generated by the generator.

[0005] The technical solution of the present invention is: a face aging system based on a pixel-level alignment generative adversarial network, which sequentially includes a task module, a face image preprocessing module, a face feature extraction module, an age estimation module, and a face synthesis module;

[0006] Task module: used for face image acquisition and task issuing, collecting face images using a camera or a handheld photography device and inputting the face aging task;

[0007] Face image preprocessing module: used for detecting and cropping the face area of the input face image, and outputting the face itself;

[0008] Face feature extraction module: for the preprocessed face image, using a deeper trained feature extraction network ResNet50 to extract face features, and obtaining a feature map of the face image to represent the face;

[0009] Age estimation module: for the preprocessed face image, adopting an age estimation framework based on ordinal regression, ranking consistency regression, to distribute the face aging results within the expected age range;

[0010] Face synthesis module: sending the preprocessed face image, age estimation and feature extraction structures into the face synthesis module to synthesize a face image that meets the target requirements.

[0011] Preferably, the task module receives two tasks: the first is to collect faces in real time through a shooting and video recording device online and upload them to the system to synthesize faces of different ages; the second is to upload the face images collected by the device offline to the system in the form of file upload.

[0012] A method for establishing a face aging system based on a pixel-level alignment generative adversarial network, comprising the following steps:

[0013] 1) Collect face images of each age group and add age labels;

[0014] 2) Face image preprocessing: using a face detection tool to detect the face images collected in step 1), obtaining the position coordinates of the face in the image and the face key points, and performing subsequent cropping on the framed face box to remove redundant background information as training samples;

[0015] 3) Face feature extraction training: given the training samples, introducing a residual network structure into the convolutional neural network to extract face features;

[0016] 4) Age estimation training: given the training samples, using ordinal regression to find the mapping relationship from the image to age ranking;

[0017] 5) Face synthesis training: given the training samples, sending them into the generator of the generative adversarial network, and the generator generates an image with the aim of reducing the difference in each pixel value between the source image and the generated image, and the generated image is sent to the discriminator for scoring to achieve face synthesis training.

[0018] Furthermore, in step 2), the face image preprocessing adopts the MTCNN algorithm to simultaneously complete face detection and face alignment.

[0019] Furthermore, for the age estimation training method in step 4): Assume that there is a corresponding age label for each face image. The age label is extended to a K-dimensional feature space and represented as an order relationship, monotonically increasing from 1 to K. That is, all dimensions less than this age label are assigned 1, and the rest are 0. Given the training samples, K-1 binary classifiers are trained. The result predicted by each classifier is the relationship between the predicted value and its belonging rank. To achieve rank monotonicity, the K-1 binary classifiers share weight parameters but have different biases.

[0020] Furthermore, for the face synthesis training in step 5), the face synthesis is trained through a conditional generative adversarial network combined with age estimation, identity feature preservation, and background denoising.

[0021] Given the input image x s , the age label y s , and the target age y t , the ultimate goal is to synthesize a face image x t with the age feature of y t ;

[0022] First, a conditional generative adversarial network is used to generate a target age picture x s of the input image x t . The input image x s , the age label y s , and the target age label y t , that is:

[0023] x t = G(x s , y s , y t )

[0024] G represents the generator used to generate face images of a specific age range;

[0025] For the age estimation task, the face x t generated by the generator is used as the input to a CNN network AE for age estimation. Then there is:

[0026] y t ' = AE(x t )

[0027] Estimation is carried out through Euclidean distance calculation. The estimation formula:

[0028]

[0029] For identity feature preservation, that is, ensuring that the generated face is recognized as the same person as the source face, ResNet50 is used as the backbone network of the feature extractor to extract x s and x t The feature maps of are respectively ft s and ft t , in addition, in order to make the age estimation task better adapt to different faces, two fully connected layers and a softmax layer are connected after the ResNet50 backbone network as an age classifier, and its ResNet50 structure shares parameters with the identity feature preservation task:

[0030] L identity = MSE(ft s , ft t );

[0031] For the background denoising task, aiming to shift the focus of training the generator to a finer-grained pixel level, the goal is to narrow the gap between the pixel values of the source image and the generated image, then there is:

[0032]

[0033] where represents the error between the estimated age and the true age y s , and w, h, and c respectively correspond to the width, height, and number of channels of the image;

[0034] Next, the obtained target face x t is fed into a discriminator D, and the discriminator will score the input image, and the more realistic the image, the higher the score.

[0035] Furthermore, when training the conditional generative adversarial network, an age classifier with ResNet50 as the backbone network is used to identify which age group C t the face comes from. If the generated face is indeed in C t , the classifier will give it a smaller penalty. On the contrary, if it is not in C t , the classifier will give a larger penalty. The age classification loss L age of the overall age estimation task is defined as follows:

[0036]

[0037] where CE(.) is the cross-entropy loss function; is the predicted age group; C t is the target age group; The face aging training objective function using the conditional generative adversarial network is as follows:

[0038] L = λg L G + λ i L identity + λ p L pixel + λ α L age

[0039] where λ g , λ i , λ p , λ α are hyperparameters that respectively control the losses of each part, and the hyperparameters are used to balance the losses between face accuracy and identity feature preservation.

[0040] The beneficial effects of the present invention are as follows: The face aging system and its establishment method based on pixel-level alignment generative adversarial network of the present invention adopt an age estimation method based on ordinal regression to ensure that the synthesized face is within the target age range, and use ResNet50 as the face feature extractor to ensure the identity consistency of the face. In addition, pixel-level loss is added to reduce the error caused by background noise and improve the quality of the synthesized image. Compared with the existing face aging methods at the present stage, the pixel-level alignment generative adversarial network model proposed by the present invention has obvious advantages in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is the overall implementation flowchart of the system of the present invention;

[0042] Figure 2 is the structural diagram of the age estimation network based on ordinal regression in the system construction method of the present invention;

[0043] Figure 3 is the structural diagram of the network based on pixel-level alignment generative adversarial network in the system construction method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0045] As Figure 1 shown in the implementation flowchart of the face aging system based on pixel-level alignment generative adversarial network, the specific implementation steps are as follows:

[0046] a) Task module: used for face image acquisition and task issuance, and a camera or a handheld photography device is used to collect face images and input face aging tasks.

[0047] The system developed by the present invention receives two types of input data: online and offline. Online data is collected in real time through devices such as mobile phones, and the uploaded human faces are synthesized by the system into human faces of different ages. Offline data refers to the face images collected by devices and uploaded to the system in the form of file uploads. Age labels need to be input synchronously.

[0048] b) Face image preprocessing module: Detect and crop the face region of the input face image to ensure that the subsequent content only focuses on the preprocessed face.

[0049] c) Face feature extraction module: Use the deeper trained feature extraction network ResNet50 to extract face features, and obtain the feature map of the preprocessed face image to represent the face.

[0050] d) Age estimation module: For the preprocessed face image, adopt the age estimation framework of Consistent Rank Logits (CORAL) based on ordinal regression to distribute the face aging results within the expected age range.

[0051] e) Face synthesis module: Send the preprocessed face image, age estimation, and feature extraction structure into the face synthesis module to synthesize the face image that meets the target requirements.

[0052] The construction method of the face aging system based on the pixel-level alignment generative adversarial network is as follows:

[0053] 1) Collect face data:

[0054] The system developed by the present invention receives two types of input data: online and offline. Online data is collected in real time through devices such as mobile phones, and the uploaded human faces are synthesized by the system into human faces of different ages. Offline data refers to the face images collected by devices and uploaded to the system in the form of file uploads. Age labels need to be input synchronously.

[0055] 2) Face image preprocessing: Face detection and cropping.

[0056] There are many face images obtained in real scenarios that may contain not only one face but also a lot of additional information, which is redundant noise information for our next research work. Therefore, before using it as the input of the model, it is necessary to complete the preprocessing of the face image, that is, face detection and cropping to a suitable position, removing the redundant background information and only focusing on the face itself. The purpose of face detection is to circle the position corresponding to the face in a picture and output the coordinates of the face in the image, providing cropping coordinates for the subsequent cropping work. We use MTCNN, which is a deep learning-based face detection and face alignment method. It can complete the tasks of face detection and face alignment simultaneously, and do the data preprocessing work for the training of subsequent face aging, age estimation and other models. MTCNN can output the position coordinates of the face and the face key points in a picture. When there are multiple faces in a picture, it can also frame multiple face boxes for subsequent cropping. Just according to the position coordinates bbox estimated by MTCNN, the Python library PIL can be used to crop the picture according to the coordinates.

[0057] 3) Face feature extraction training:

[0058] When studying some face-related tasks, the first step is to extract face features. The work of extracting image features is to extract some characteristics that can correctly represent the image to ensure the successful completion of many downstream computer vision tasks. Deep learning can simulate the cognitive mechanism of the human brain by constructing complex network structures, and this principle can be well used to construct image features at different levels. When the computer learns these image features, it regards these images as matrices composed of pixels. Therefore, analyzing this image is to analyze the matrix composed of these numbers. These analyses use convolutional neural networks (CNNs).

[0059] As the depth of the model increases, the phenomenon of gradient disappearance will occur. Therefore, when performing deeper face feature extraction tasks, the ResNet model based on the residual network is selected. The innovation of this model is to introduce the residual network structure into the convolutional neural network. It can construct the original model structure as a shallow network plus its own mapping, and cleverly connect the trained shallow structure with its own mapping through residuals. In this way, it can be ensured that the next layer of the network must contain more image information than the previous layer. Different numbers of residual network layers (ResNet-34, ResNet-50) are used for feature extraction work in Figure 2 and Figure 3

[0060] 4) Age estimation training:

[0061] Ordered regression is a model for solving problems where there is a certain order relationship between categories, such as age. In addition to considering classification loss, it also takes into account the order relationship between different categories. The loss of a result closer to the true label order is less than that of a result farther from the true label. The age estimation framework CORAL based on ordered regression is adopted to ensure that the distribution of the face aging results is within the expected age range.

[0062] Assume that there is a corresponding age label for each face image. The age label is extended to a K-dimensional feature space and represented as an order relationship, monotonically increasing from 1 to K. That is, all dimensions less than this age label are assigned 1, and the rest are 0. Given the training samples, use ordered regression to find the mapping relationship from the image to the age ranking. Therefore, K - 1 binary classifiers need to be trained. As Figure 2 shown, the result predicted by each classifier is the relationship between the predicted value and its corresponding ranking. To achieve rank monotonicity, these K - 1 binary classifiers share weight parameters but have different biases.

[0063] 5) Face synthesis training:

[0064] First, divide the ages into 7 non-overlapping age groups, corresponding to 0 - 10, 10 - 20, 20 - 30, 30 - 40, 40 - 50, 50 - 60, and over 60 years old respectively. Given a face image, use a 7-dimensional feature encoding to represent the age group where the input image belongs. The height and width of each dimension feature map are the same as those of the input image. 7 represents the number of age groups. Only the channel corresponding to the age group to which the input picture belongs is all 1, and the others are all 0. The main work of this step is to synthesize a face image belonging to the target age group, and i) it looks close to a real face; ii) it has the same features as the input face; iii) its age estimation category is within the target age group.

[0065] To obtain face images of different age groups, the present invention divides the face synthesis training task into three parts: age estimation, identity feature preservation, and background denoising. The overall process is as Figure 3 shown. Given the input image x s , the age label y s , and the target age y t . Our ultimate goal is to synthesize a face image x t with the age characteristics of y t . First, use the Conditional Generative Adversarial Networks (CGAN) to generate the corresponding target-aged picture x s of the input image x t . Here, not only the input image x s and y sAs the input of the conditional generative adversarial network, it is also necessary to know the age label y corresponding to the output image t , that is:

[0066] x t = G(x s , y s , y t )

[0067] G represents the generator used to generate face images of a specific age group. First, for the age estimation task, the face x generated by the generator t is used as the input of a CNN network AE for age estimation, then there is:

[0068] y t ' = AE(x t )

[0069] Estimation is carried out through Euclidean distance calculation, and the estimation formula is:

[0070]

[0071] To achieve identity feature preservation, that is, to ensure that the generated face and the source face are recognized as the same person. The present invention uses ResNet50 as the backbone network of the feature extractor, and the feature maps of x s and x t are respectively ft s and ft t . In addition, in order to make the age estimation task better adapt to different faces, two fully connected layers and one softmax layer are connected after the ResNet50 backbone network as an age classifier, and its ResNet50 structure shares parameters with the identity feature preservation task.

[0072] L identity = MSE(ft s , ft t )

[0073] When the training task focuses too much on the face itself, the background noise of the generated image will gradually increase, and the authenticity of the synthesized face will gradually decrease. Therefore, the present invention proposes a background denoising task, aiming to shift the focus of training the generator to a finer-grained pixel level. The goal is to narrow the gap between the pixel values of the source image and the generated image, then there is:

[0074]

[0075] where represents the estimated age and the real age y sThe error between them, where w, h, and c correspond to the width, height, and number of channels of the image respectively. Next, the obtained target face x t is fed into a discriminator D, which scores the input image, and the more realistic the image, the higher the score. The purpose of the present invention is to improve the quality of the synthesized face and ensure that the identity features of the synthesized face are consistent with the input face. Therefore, pixel-level alignment can effectively reduce background artifacts and improve the quality of the synthesized face.

[0076] Specifically, the main innovation of the present invention is to add pixel-level face alignment loss to the generator loss part and use the Euclidean distance loss between the generated image predicted by the CORAL framework and the target age label.

[0077] The main generation structure of the present invention is to use CGAN to generate new pictures with the characteristics of pictures of another age group from pictures of one age group. When training CGAN, the synthesized face with the target age group C t will not be judged as a fake sample by the discriminator D. The more realistic the face, the higher the discriminator score. Therefore, the present invention adopts Conditional Least Squares Generative Adversarial Network (Conditional LSGAN), which is a special CGAN. It will try to push the generated face and the real face to a position close to the decision boundary, making them indistinguishable, so as to ensure the generation of high-quality pictures close to reality and more stable training.

[0078] To further ensure that the generated humans conform to the age group C t , the present invention uses an age classifier with ResNet50 as the backbone network to identify which age group the face belongs to. If the generated face is indeed in C t , the classifier will give it a smaller penalty. On the contrary, if it is not in C t , the classifier will give a larger penalty. In many classification problems, fusing information from different views can improve the prediction performance, especially when the information from different perspectives is complementary. Then, for the overall age estimation task, the age classification loss L age is defined as follows:

[0079]

[0080] where CE(.) is the cross-entropy loss function; is the predicted age group; C t is the target age group.

[0081] Generally speaking, the objective function in the face aging training stage using CGAN is as follows:

[0082] L = λ g L G + λ iL identity +λ p L pixel +λ α L age

[0083] where λ g 、λ i 、λ p 、λ α are hyperparameters that respectively control the losses of each part. Using these hyperparameters, the loss balance between face accuracy and identity feature preservation can be weighed.

[0084] Finally, the result is returned to the uploader and displayed on a displayable device. If it is real-time face aging, the target age can also be selected in real-time in the system, and the system will return the synthesis result in real-time.

[0085] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent shall be subject to the appended claims.

Claims

1. A face aging system based on a pixel-level alignment generative adversarial network, characterized in that, it sequentially includes a task module, a face image preprocessing module, a face feature extraction module, an age estimation module, and a face synthesis module; Task module: used for face image acquisition and task issuance, using a camera or a handheld photography device to collect face images and input face aging tasks; Face image preprocessing module: used to detect and crop the face region of the input face image, and output the face itself; Face feature extraction module: for the preprocessed face image, use a deeper trained feature extraction network ResNet50 to extract face features, and obtain a feature map of the face image to represent the face; Age estimation module: for the preprocessed face image, adopt an age estimation framework based on ordinal regression, ranking consistency regression, to distribute the face aging results within the expected age range; Face synthesis module: the preprocessed face image, age estimation, and feature extraction structure are sent to the face synthesis module to synthesize a face image that meets the target requirements; In the face synthesis module, face synthesis training is performed by combining age estimation, identity feature preservation, and background denoising through a conditional generative adversarial network for face synthesis training; Given the input image x s , the age label y s , and the target age y t , the ultimate goal is to synthesize a face image x with the age feature of y t ; t ; First, a conditional generative adversarial network is used to generate the target age image x s of the input image x t , the input image x s , age label y s , target age label y t , that is: x t = G(x s , y s , y t ) G represents a generator for generating face images of a specific age group; For the age estimation task, the face x generated by the generator t is used as the input of a CNN network AE for age estimation, and we have: y t ' = AE(x t ) Estimation is performed through Euclidean distance calculation, and the estimation formula is: For identity feature preservation, that is, ensuring that the generated face is recognized as the same person as the source face, ResNet50 is used as the backbone network of the feature extractor to extract the features of x s and x t as the feature maps ft s and ft t respectively. In addition, in order to make the age estimation task better adapt to different faces, two fully connected layers and one softmax layer are connected after the ResNet50 backbone network as an age classifier, and its ResNet50 structure shares parameters with the identity feature preservation task: L identity = MSE(ft s , ft t ); For the background denoising task, it aims to shift the focus of training the generator to a finer-grained pixel level. The goal is to narrow the gap between the pixel values of the source image and the generated image. Then there is: where represents the estimated age and the true age y s is the error between them, where w, h, and c correspond to the width, height, and number of channels of the image, respectively; Next, the obtained target face x t is fed into a discriminator D, which scores the input image, and the more realistic the image, the higher the score.

2. The face aging system based on a pixel-level alignment generative adversarial network according to claim 1, characterized in that, the task module receives two types of tasks: the first is to collect faces in real time through shooting and video recording devices and upload them to the system to synthesize faces of different age groups; the second is to upload the face images collected by the device to the system in the form of file uploads when not online.

3. A method for establishing a face aging system based on a pixel-level alignment generative adversarial network, characterized in that, it includes the following steps: 1) Collect face images of each age group and add age labels; 2) Face image preprocessing: Use a face detection tool to detect the face images collected in step 1), obtain the position coordinates of the face in the image and the face key points, and perform subsequent cropping on the framed face box to remove redundant background information as training samples; 3) Face feature extraction training: Given training samples, introduce a residual network structure into the convolutional neural network to extract face features; 4) Age estimation training: Given training samples, use ordinal regression to find the mapping relationship from the image to age ranking; 5) Face synthesis training: Given training samples, send them into the generator of the generative adversarial network. The generator generates images with the goal of narrowing the gap between the pixel values of the source image and the generated image. The generated images are sent to the discriminator for scoring to achieve face synthesis training. Specifically as follows: Face synthesis training is performed by combining age estimation, identity feature preservation, and background denoising through a conditional generative adversarial network; Given the input image x s , the age label y s , and the target age y t , the ultimate goal is to synthesize a face image x with the age feature of y t ; t ; First, a conditional generative adversarial network is used to generate the target age picture x of the input image x s The input image x t , the input image x s , age label y s , target age label y t , that is: x t = G(x s , y s , y t ) G represents a generator for generating face images of a specific age group; For the age estimation task, the face x generated by the generator t is used as the input of a CNN network AE for age estimation, then we have: y t ' = AE(x t ) Estimation is performed using the Euclidean distance calculation, and the estimation formula is: For identity feature preservation, that is, ensuring that the generated face is recognized as the same person as the source face, ResNet50 is used as the backbone network of the feature extractor to extract x s and x t The feature maps of are respectively ft s and ft t . In addition, in order to make the age estimation task better adapt to different faces, two fully connected layers and one softmax layer are connected after the ResNet50 backbone network as an age classifier, and its ResNet50 structure shares parameters with the identity feature preservation task: L identity = MSE(ft s , ft t ); For the background denoising task, the aim is to shift the focus of the training generator to a finer-grained pixel level. The goal is to narrow the gap between the pixel values of the source image and the generated image. Thus, we have: where represents the estimated age and the true age y s is the error between them, where w, h, and c correspond to the width, height, and number of channels of the image, respectively; Next, the obtained target face x t is fed into a discriminator D, which scores the input image, and the more realistic the image, the higher the score.

4. The method for establishing a face aging system based on a pixel-level alignment generative adversarial network according to claim 3, characterized in that in step 2), the MTCNN algorithm is used for face image preprocessing to simultaneously complete face detection and face alignment.

5. The method for establishing a face aging system based on a pixel-level alignment generative adversarial network according to claim 3, characterized in that the age estimation training method in step 4): Assume that each face image has a corresponding age label. The age label is extended to a K-dimensional feature space and represented as an order relationship, monotonically increasing from 1 to K. That is, all dimensions less than this age label are assigned 1, and the rest are 0. Given the training samples, K - 1 binary classifiers are trained. The result predicted by each classifier is the relationship between the predicted value and its belonging rank. To achieve rank monotonicity, the K - 1 binary classifiers share weight parameters but have different biases.

6. The method for establishing a face aging system based on a pixel-level alignment generative adversarial network according to claim 3, characterized in that When training the conditional generative adversarial network, an age classifier with ResNet50 as the backbone network is used to identify which age group C the face belongs to. t , if the generated face is indeed in C t , the classifier will give it a smaller penalty. On the contrary, if it is not in C t , the classifier will give a larger penalty. The age classification loss L of the overall age estimation task age is defined as follows: where CE(.) is the cross-entropy loss function; is the predicted age group; C t is the target age group; the face aging training objective function using conditional generative adversarial network is as follows: L = λ g L G + λ i L identity + λ p L pixel + λ α L age where λ g , λ i , λ p , λ α are hyperparameters that control the losses of each part respectively, and the hyperparameters are used to balance the losses between face accuracy and identity feature preservation.

Citation Information

Patent Citations

  • Face aging method based on a conditional generative adversarial network

    CN109523463A

  • Age estimation

    GB201818948D0