A Face Image Editing Method Based on a Generative Adversarial Network with Graph Convolutional Network
By introducing a graph convolutional network as a discriminator in the generative adversarial network, the problem of poor attribute coupling and generation quality in multi-attribute face image editing is solved, and more efficient multi-attribute editing and better generation effects are achieved.
Patent Information
- Application Number
- CN202211181562.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-09-27
AI Technical Summary
The traditional face image editing method based on generative adversarial networks is difficult to handle simultaneous editing of multiple attributes, and there is a coupling relationship between different attributes, resulting in the non-target attributes being modified when editing target attributes, and the generation quality is poor.
The graph convolution network is used as a discriminator, combined with the supporting loss function and data, and trained the generator to improve the multi-attribute face editing performance through adversarial training, and use the graph convolution network to process the association between different attributes.
It improves the performance of multi-attribute face editing, improves the quality of generated images, and can effectively decouple editing between different attributes, with good generalization and generation authenticity.
Smart Images

Figure CN115565223B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face image editing, and in particular to a face image editing method based on a generative adversarial network of graph convolutional networks. Background Art
[0002] Face image editing refers to editing a face image according to target attributes. Attributes refer to the styles and shapes of various parts of the face in the face image, and attribute editing refers to modifying face attributes. For example, thickening the eyebrows in a face image. Face image editing has broad application scenarios in fields such as video beauty and photo retouching.
[0003] Traditional face image editing methods based on generative adversarial networks are difficult to handle simultaneous editing of multiple attributes. There is a coupling relationship between different attributes, resulting in the modification of non-target attributes when editing target attributes, and the generated quality is poor. In the present invention, a new generative adversarial network is designed. In this generative adversarial network, the discriminator network part uses graph convolutional networks to help process the associations between different attributes, thereby improving the performance of the model for multi-attribute face editing. Summary of the Invention
[0004] The purpose of the present invention is to alleviate the problems of multi-attribute coupling and poor generation quality existing in existing face image editing models, and provide a face image editing method based on a generative adversarial network of graph convolutional networks. It uses a discriminator with graph convolutional networks, combines a supporting loss function and supporting data to train a generator for face image editing together, and finally can be applied to face image editing tasks.
[0005] To achieve the above purpose, the technical solution provided by the present invention is: A face image editing method based on a generative adversarial network of graph convolutional networks, including the following steps:
[0006] S1. Prepare a face attribute data set. The face attribute data set refers to a set of data containing face images cropped in proportion and corresponding attribute annotations. Denote the image as X and the corresponding attribute annotation as Y;
[0007] S2. Initialize the generative adversarial network model. The model contains three neural networks: a generator G, a discriminator D, and a classifier C. First is the generator G, whose input is the face image x to be processed src and the target attribute y trg , and the output is the processed image x with the target attribute genSecondly, there is a discriminator D, whose input is a real face image or an image processed by the generator G, and the output is the discrimination result of each attribute of the input image. The discriminator D contains a graph convolutional network for modeling the associations between image attributes. Finally, there is a classifier C, whose input is a face image, and the output is the prediction score of each attribute of the input image.
[0008] S3. Use X and Y as training data to perform adversarial training on the generator G, the discriminator D, and the classifier C. Finally, obtain a well-trained generator G, discard the discriminator D and the classifier C, and finally use only the generator G to complete the task of face image editing.
[0009] Furthermore, in step S1, the face images in the face attribute dataset should be processed first. They need to be cropped to a size of 256×256 and ensure that the face is centered and aligned. To prevent model overfitting due to insufficient sample quantity, it is also necessary to perform random horizontal flipping data augmentation on the images.
[0010] Furthermore, in step S2, the generator G needs to convert the input image x src into an edited image. For this purpose, the input image needs to be first processed by an encoder to be converted into the latent space to obtain the corresponding latent code, and then a decoder converts the latent code obtained by the former back into the image space. At the same time, the generator G also contains a mapping network for converting the input target attribute vector y trg into a style code, which is provided to the encoder together with the latent code. Finally, the encoder outputs the edited image x gen ; Since the operations performed by the encoder and the decoder are inverse to each other, their structures are symmetric; for the encoder, first use a convolutional network to convert the input RGB image into a feature map, and then further convert it into a latent code through multiple residual block networks; for the decoder, first use multiple residual block networks that can accept style information to process the latent code, and then use a convolutional network to convert the feature map into the RGB image space; the mapping network consists of multiple fully connected layers, and it processes the input target attribute vector y trgProcess it to convert it into the style code required by the aforementioned residual block network, thereby merging the target attribute information with the processed image information, and finally outputting the processed image by the decoder. Then, input the generated image into the discriminator D and the classifier C; the discriminator D is responsible for discriminating the authenticity of the input image to ensure the authenticity of the model-edited image; the backbone of the discriminator D is similar to the encoder part of the generator. First, the input image is processed by multiple convolutional layers to convert it from the RGB image space to the feature map, and then a flattening operation of the feature map is performed to obtain a feature vector. Then, through a mapping module, the feature map is adaptively processed for each attribute to obtain a set of attribute-related feature vectors. Then, together with the attribute-related matrix calculated from the dataset, they are sent into the subsequent graph convolutional network. Finally, the output layer of the graph convolutional network outputs a vector with the same size as the number of attributes, where is the discrimination score of the discriminator D for each attribute. The classifier C is responsible for scoring each attribute of the input image to ensure that the model-edited image meets the editing target; the backbone of the classifier C is also similar to the encoder part of the generator. First, it is a convolutional network block stacked with multiple convolutional layers, and then a fully connected layer is connected to output the scores of the classifier C for each attribute.
[0011] Further, in step S3, the generator G processes the input image and generates the processed image, which performs adversarial learning with the discriminator D. The purpose of the discriminator D is to discriminate the authenticity of the input image, while the generator G tries to deceive the discriminator D so that it is difficult to distinguish whether the input image is a real image or a generated image. Therefore, the generator G needs to generate as realistic an image as possible; this is achieved through the adversarial training of the discriminator D and the generator G, and the adversarial loss function L of the discriminator D and the generator G is defined. adv :
[0012]
[0013] In the formula, a represents the total number of attributes, y (i) represents the value of the i-th attribute, D (i) (·) represents the discrimination result of the discriminator for the i-th attribute of the input image, x src represents the input image, y trg represents the target attribute;
[0014] Another key point of attribute editing is that the irrelevant part should be as consistent with the original image as possible. Therefore, a loss function L is defined to promote the generated image to be as consistent as possible with the original image in the attributes that should not be changed. psv :
[0015]
[0016] In the formula, ysrc represents the i-th attribute of the original input image, represents the i-th target attribute, represents the feature representation of the discriminator D for the input image on attribute i; in addition, the generated image should also satisfy the constraint of the target attribute y trg and the generated image x is constrained by the classifier C gen to define a classification loss L cls :
[0017]
[0018] In the formula, C (i) (·) represents the classification result of the classifier for the input image on attribute i;
[0019] Finally, the overall training objective is:
[0020]
[0021] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0022] 1. The present invention combines multiple neural networks and jointly trains between the neural networks, ultimately improving the effect of face image editing, and having good generalization, and can be applied to a variety of face datasets and real face data.
[0023] 2. The present invention proposes a neural network for image conversion, which can be well adapted to face image editing. By using the reconstruction loss, it is ensured that the parts that should not be modified in the generated image are consistent with the original image. At the same time, the adversarial loss is used to make the generated image sufficiently real, and the additional classification loss makes the generated image of the model contain the target attributes of face image editing, so it has good applicability.
[0024] 3. The present invention combines the graph convolutional network with the discriminator in the generative adversarial network, enabling the discriminator to more effectively utilize the correlation information between attributes and decouple the editing between different attributes well, effectively improving the performance of the model in the case of multi-attribute editing, and providing a good solution for the task of using the generative adversarial network for face image editing. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flow block diagram of the method of the present invention.
[0026] Figure 2 is a structural diagram of the generative adversarial network model in the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0027] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.
[0028] As Figure 1 shown, this embodiment provides a face image editing method based on a generative adversarial network of graph convolutional networks. Taking the implementation of the generative adversarial network model on the CelebA-HQ dataset as an example, the method includes the following steps:
[0029] S1. The CelebA-HQ face attribute dataset contains 30,000 face images, including 28,000 training set images and 2,000 test set images. Each image has 40 binary attribute annotations, and the image size is 1024×1024. In our experiment, we first resize the images to 256×256, then perform random horizontal flipping for image augmentation, and finally normalize the images.
[0030] S2. As Figure 2 shown, our generative adversarial network model contains three neural networks: First is the generator G, whose input is the face image x src to be processed and the target attribute y trg , and the output is the processed image x gen with the target attribute. The generator first needs to convert the input image into the latent space through an encoder to obtain the corresponding latent code, and then a decoder converts the latent code obtained previously back into the image space. At the same time, the generator also contains a mapping network for converting the input target attribute vector y trg into a style code, which is provided to the encoder together with the latent code. Finally, the encoder outputs the edited image x gen . Since the operations performed by the encoder and the decoder are inverse to each other, their structures are symmetric. For the encoder, first use a convolutional network to convert the input RGB image into a feature map, and then further convert it into a latent code through multiple residual block networks; for the decoder, first use multiple residual block networks that can accept style information to process the latent code, and then use a convolutional network to convert the feature map into the RGB image space. The mapping network consists of multiple fully connected layers, which processes the input target attribute vector y trgProcess it and convert it into the style code required by the aforementioned residual block network, thereby merging the target attribute information with the processed image information, and finally outputting the processed image by the decoder. Then, the generated image is input into the discriminator D and the classifier C. Secondly, there is the discriminator D. Its input is a real face image or the image processed by the generator, and its output is the discrimination result of each attribute of the input image. The discriminator D is responsible for discriminating the authenticity of the input image to ensure the authenticity of the model-edited image. The backbone part of our discriminator is similar to the encoder part of the generator. First, a multi-layer convolutional layer processes the input image, converting it from the RGB image space to a feature map. Then, through an operation of flattening the feature map, a feature vector is obtained. Then, through a mapping module, the feature map is adaptively processed for each attribute, and a group of attribute-related feature vectors are obtained. Then, together with the attribute-related matrix calculated from the dataset, they are sent into the subsequent graph convolutional network. Finally, the output layer of the graph convolutional network inputs a vector with the same size as the number of attributes, where is the discrimination score of the discriminator for each attribute. Finally, there is the classifier C. Its input is a face image, and its output is the prediction score of each attribute of the input image. The classifier C is responsible for scoring each attribute of the input image to ensure that the model-edited image meets the editing goal. The backbone part of our classifier is also similar to the encoder part of the generator. First, it is a convolutional network block stacked with multi-layer convolutional layers, and then a fully connected layer is connected to output the scores of the classifier for each attribute. After defining the model architecture, we need to initialize the model parameters. Here, we adopt the He parameter initialization strategy.
[0031] S3. After the model initialization is completed, we need to train it. Among them, the generator and the discriminator are trained adversarially. The purpose of the discriminator is to discriminate the authenticity of the input image, while the generator tries to deceive the discriminator so that it is difficult to distinguish whether the input image is a real image or a generated image. For this reason, the generator needs to generate as real an image as possible. We achieve this through the adversarial training of the discriminator D and the generator G. Specifically, we define the adversarial loss function L of the discriminator D and the generator G adv :
[0032]
[0033] In the formula, a represents the total number of attributes, y (i) represents the value of the i-th attribute, D (i) (·) represents the discrimination result of the discriminator for the i-th attribute of the input image, x src represents the input image, y trg represents the target attribute;
[0034] Another key point of attribute editing is that the irrelevant parts should be as consistent with the original image as possible. Therefore, a loss function \(L\) is defined to promote the generated image to be as consistent as possible with the original image in terms of the attributes that should not be changed. psv :
[0035]
[0036] In the formula, represents the \(i\)-th attribute of the original input image, represents the \(i\)-th target attribute, represents the feature representation of the discriminator \(D\) for the input image on attribute \(i\); in addition, the generated image should also satisfy the constraint of the target attribute \(y\) trg By using the classifier \(C\) to constrain the generated image \(x\) gen a classification loss \(L\) is defined cls :
[0037]
[0038] In the formula, \(C\) (i) (·) represents the classification result of the classifier for the input image on attribute \(i\);
[0039] Finally, the overall training objective is:
[0040]
[0041] Finally, after the training is completed, the method is evaluated on the test set of the CelebA-HQ dataset. The evaluation criteria are the Fréchet inception distance (FID) and the Peak signal-to-noise ratio (PSNR). The lower the FID value, the higher the generation quality and the better the effect of the model, while the higher the PSNR value, the higher the reconstruction quality and the better the effect of the model. After evaluation, the effect of the present invention is significantly higher than that of the baseline method and is worthy of promotion.
[0042] The above-described embodiments are only the preferred embodiments of the present invention, but do not limit the application scope of the method of the present invention. Therefore, all changes made according to the shape and principle of the present invention should be covered within the protection scope of the present invention.
Claims
1. A face image editing method based on a generative adversarial network of graph convolutional networks, characterized in that, The steps include the following: S1. Prepare a face attribute dataset. The face attribute dataset refers to a set of data that includes face images cropped proportionally and corresponding attribute annotations. Denote the image as X and the corresponding attribute annotation as Y; S2. Initialize the generative adversarial network model, which contains three neural networks: generator G, discriminator D, and classifier C. First is the generator G, whose input is the face image x to be processed src and the target attribute y trg , and the output is the processed image x with the target attribute gen ; Second is the discriminator D, whose input is a real face image or the image processed by the generator G, and the output is the discrimination result of each attribute of the input image. The discriminator D contains a graph convolutional network for modeling the associations between image attributes; Finally is the classifier C, whose input is a face image, and the output is the prediction score of each attribute of the input image; The generator G needs to convert the input image x src into the edited image. To this end, the input image needs to be first processed by an encoder to be converted into the latent space to obtain the corresponding latent code, and then a decoder converts the latent code obtained by the former back into the image space. At the same time, the generator G also includes a mapping network for converting the input target attribute vector y trg into the style code, which is provided to the encoder together with the latent code. Finally, the encoder outputs the edited image x gen ; Since the operations performed by the encoder and the decoder are inverse to each other, their structures are symmetric; for the encoder, first use a convolutional network to convert the input RGB image into a feature map, and then further convert it into a latent code through multiple residual block networks; for the decoder, first use multiple residual block networks that can accept style information to process the latent code, and then use a convolutional network to convert the feature map into the RGB image space; the mapping network consists of multiple fully connected layers, which processes the input target attribute vector y trg to convert it into the style code required by the aforementioned residual block network, thereby combining the target attribute information with the processed image information, and finally outputting the processed image by the decoder. Then, the generated image is input into the discriminator D and the classifier C; S3. Use X and Y as training data to conduct adversarial training on the generator G, discriminator D, and classifier C. Finally, obtain a well-trained generator G, discard the discriminator D and classifier C, and finally use only the generator G to complete the task of face image editing.
2. The face image editing method of a generative adversarial network based on a graph convolutional network according to claim 1, characterized in that: In step S1, the face images in the face attribute dataset should first be processed. They need to be cropped to a size of 256×256 and ensure that the face is centered and aligned. To prevent model overfitting due to insufficient sample quantity, data augmentation by randomly horizontally flipping the images is also required.
3. The face image editing method of a generative adversarial network based on a graph convolutional network according to claim 1, characterized in that: In step S2, the discriminator D is responsible for discriminating the authenticity of the input image to ensure the authenticity of the model-edited image. The main part of the discriminator D first processes the input image through multiple convolutional layers, converting it from the RGB image space to a feature map. Then, through an operation of flattening the feature map to obtain a feature vector. After that, through a mapping module, the feature map is adaptively processed for each attribute to obtain a set of attribute-related feature vectors. Then, together with the attribute-related matrix calculated from the dataset, they are fed into the subsequent graph convolutional network. Finally, the output layer of the graph convolutional network outputs a vector with the same size as the number of attributes, where is the discrimination score of the discriminator D for each attribute. The classifier C is responsible for scoring each attribute of the input image to ensure that the model-edited image meets the editing goal. The main part of the classifier C is first a convolutional network block stacked with multiple convolutional layers, and then followed by a fully connected layer to output the scores of the classifier C for each attribute.
4. A face image editing method based on a generative adversarial network of graph convolutional networks according to claim 1, characterized in that: In step S3, the generator G processes the input image and generates a processed image, which undergoes adversarial learning with the discriminator D. The purpose of the discriminator D is to distinguish the authenticity of the input image, while the generator G attempts to deceive the discriminator D so that it is difficult to distinguish whether the input image is a real image or a generated image. Therefore, the generator G needs to generate as realistic an image as possible; this is achieved through the adversarial training of the discriminator D and the generator G, and an adversarial loss function L for the discriminator D and the generator G is defined. adv : where a represents the total number of attributes, and y (i) represents the value of the i-th attribute, D (i) (·) represents the discrimination result of the discriminator for the i-th attribute of the input image, x src represents the input image, y trg represents the target attribute; Another key point in attribute editing is that the irrelevant parts should be as consistent with the original image as possible. Therefore, a loss function L is defined to promote the generated image to be as consistent as possible with the original image in terms of attributes that should not be changed. psv : Wherein, represents the i-th attribute of the original input image, represents the i-th target attribute, represents the feature representation of the input image by the discriminator D on the attribute i; in addition, the generated image should also satisfy the target attribute y trg constraint, and the classifier C is used to constrain the generated image x gen to define a classification loss L cls : Where C (i) (·) represents the classification result of the classifier on the input image with respect to attribute i; Finally, the overall training objective is:
Citation Information
Patent Citations
Method and device for generating description information of multimedia data, equipment and medium
CN111723937A
Depth map convolution model defense method based on generative adversarial network
CN112287997A