A Face Mask Editing Method Based on Generative Adversarial Networks

Through the training of the generation of adversarial network models and the optimization of loss function, the editing problem of wearing mask images is solved, the realistic mask editing effect is achieved, and the application ability of face recognition is improved.

CN115439311BActive Publication Date: 2025-07-29NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211023935.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-24
Publication Date
2025-07-29
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

In the prior art, when processing face images of masks, it is difficult to effectively edit, remove or add masks, resulting in limited facial recognition applications, and the existing methods are complex or have poor results.

Method used

The generative adversarial network model is adopted, and the generative adversarial network model is trained. The editing function of masks is achieved by using no translation path, self-translation path, loop translation path and triple consistent translation path, combining adversarial loss, reconstruction loss, style loss and triple consistent loss.

Benefits of technology

Effective editing of mask-wearing images is realized, and realistic mask-free images or mask-wearing images can be generated, which enhances the application effect of face recognition, and improves the practicality of the algorithm and the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FHA0000011680470000023
    Figure FHA0000011680470000023
  • Figure FHA0000011680470000024
    Figure FHA0000011680470000024
  • Figure FHA0000011680470000026
    Figure FHA0000011680470000026
Patent Text Reader

Abstract

The present invention discloses a face mask editing method based on a generative adversarial network, which includes the following steps: randomly mark one-tenth of the face images in the face dataset with masks and add annotations on whether a mask is worn to obtain a dataset; build a generative adversarial network model, whose input is the face images and corresponding annotations in the dataset, and the target image is the appearance of the input face image with a mask on or the mask removed. After training the generative adversarial network model, the functions of wearing a mask and removing a mask are realized; train the generative adversarial network model, which has four training paths: non-translation path, self-translation path, cyclic translation path, and triple-consistent translation path. A total of four loss functions, namely adversarial loss, reconstruction loss, style loss, and triple-consistent loss, are calculated, and the network parameters are updated by backpropagation to achieve the optimal editing effect. The present invention can restore different faces from a single picture according to different reference pictures and has a good editing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a face mask editing method implemented by applying machine learning. Background Art

[0002] During the COVID-19 pandemic, for safety reasons, people often wear masks when traveling. However, today when face recognition is widely used, the wearing of masks obscures many facial features, bringing many difficulties to the application of this technology. Therefore, designing a face mask editing method that can obtain the possible image of a person without a mask when an image of a person wearing a mask is input, and can obtain the image of a person wearing a mask when an image of a person without a mask is input, can greatly alleviate the inconvenience brought by wearing masks to real life.

[0003] Currently, there is little research on removing masks from faces, but it can be regarded as a face completion problem; on this basis, to achieve the effect of adding a mask, whether wearing a mask can be regarded as a facial attribute, and this problem is thus transformed into a facial attribute editing problem. Essentially, the facial attribute editing problem is an image generation problem. However, different from the style conversion of images, since facial attribute editing only needs to modify some image features while keeping other details of the image unchanged, the facial attribute editing problem has a higher difficulty. When the technology was not yet mature, the key principle of facial editing attribute technology was to find key points on the face, which might distort the shape of the face and result in inconsistent left and right faces. With the continuous improvement and development of the generative model and the realization of the mapping function of images between two domains by the generative adversarial network, remarkable research progress has been made in the facial attribute editing problem. Summary of the Invention

[0004] Existing facial attribute editing methods often require very complex networks or have poor editing attributes, and there is no special optimization for putting on and taking off masks, resulting in poor editing effects. To solve the above problems, the purpose of the present invention is to propose a face mask editing method based on the generative adversarial network.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] A face mask editing method based on the generative adversarial network, comprising the following steps:

[0007] Step 1, randomly put masks on one-tenth of the face images in the face dataset and add labels indicating whether a mask is worn to obtain a dataset;

[0008] Step 2: Build a generative adversarial network model. The input is the face images and corresponding annotations in the dataset obtained in Step 1, and the target images are the face images with masks on or masks removed. After training the generative adversarial network model, the functions of wearing masks and removing masks are realized;

[0009] Step 3: Train the generative adversarial network model built in Step 2. There are four training paths: the no-translation path, the self-translation path, the cyclic translation path, and the triple-consistent translation path. A total of four loss functions, namely the adversarial loss, the reconstruction loss, the style loss, and the triple-consistent loss, are calculated, and the network parameters are updated by backpropagation to achieve the optimal editing effect.

[0010] In Step 1, the face images in the obtained dataset are randomly cropped and randomly flipped horizontally.

[0011] In Step 2, the overall structure of the generative adversarial network model is as follows:

[0012] Input → Encoder E → Editor T → Generator G → Output;

[0013] Among them, the input is the image x to be edited. The encoder E encodes the direct features of the input face image as e = E(x). The editor T edits the encoding according to the style code s as e' = T(e, s). The generator G decodes the encoding e' as x' = G(e') to obtain the final output image x'. All model parameters are jointly trained with the discriminator;

[0014] The style code refers to the code that matches the style of the edited image, which is obtained by passing random noise through a mapper or by passing a reference image through an extractor;

[0015] The mapper module consists of an MLP, i.e., a multi-layer perceptron. The input is random noise, the label is injected before the first layer, and the attribute is injected in the middle layer for indexing. The output is the style code;

[0016] The extractor module consists of two convolutional layers, five downsampling blocks, and a pooling layer. Among them, the downsampling block inherits the ResBlock, i.e., the pre-activated residual unit. The input is the reference image, the label is injected before the last layer for indexing, and the output is the style code;

[0017] The encoder module consists of one convolutional layer and two downsampling blocks. The input is the image, and the output is the encoding;

[0018] The generator module consists of two upsampling blocks and one convolutional layer. The input is the encoding, and the output is the image;

[0019] The shared module in the encoder module and the generator module uses IN, i.e., the instance normalization layer;

[0020] The converter module consists of two convolutional layers and eight adaptive instance normalization instant blocks that reduce the channel dimension. The attribute-related style code is injected into all adaptive instance normalization layers, and scaling and shift vectors are provided through a linear layer. The label is used to establish an index before the first layer. The input is the encoding, and the output is the encoding modified according to the style code.

[0021] The discriminator module uses the same architecture as the extractor module, but the attribute is injected for indexing before the first layer, and the label-unrelated condition is injected into the last layer. The input is the image, and the output is the true or false of the input image.

[0022] For all remaining units, the LReLU function is used as the activation function.

[0023] The feature map is resampled using average pooling and nearest neighbor upsampling.

[0024] In step 3, the four training paths are as follows:

[0025] Non-translation path: Encode the original image x i,j using the encoder E, where i represents the label and j represents the style, and then decode it using the generator G to obtain the non-translated reconstructed image of the original image, i.e., x'. i,j = G(E(x i,j ));

[0026] Self-translation path: Use the extractor module F to extract the style code s i,j of label i from the original image x i,j , i.e., s i,j = F i (x i,j ), then edit the style code and the original image encoding by inputting them into the converter T, and finally decode the encoding using the generator G to obtain the self-translated reconstructed image of the original image, i.e., x'' i,j = G(T(E(x i,j ), s i,j ));

[0027] Cyclic translation path: First, input the random noise z into the mapper module M to generate a random style code related to the target label i. Secondly, inject the style code and the original image into the converter T and decode it through the generator G to obtain the transformed image Finally, encode the transformed image and inject it together with the style code obtained from the original image through the extractor into the converter and the generator to obtain the cyclic translation reconstructed image, i.e.,

[0028] Triple-consistent translation path: First, use the mapper module to generate a style code related to the target label and another random style code s' i,j′ = M i,j′ (z), and then inject the two style codes into the original image injection converter and decode through the generator to obtain the converted image x' i,j′ = G(T(E(x i,j ), s i,j′ )) and Then inject x' i,j′ and into the converter and decode through the generator to obtain the converted image

[0029] In step 3, the four loss functions are as follows:

[0030] Adversarial loss: Encourage the network to perform realistic operations on the generated and extracted styles. Input the real image x i,j , the intermediate generated image obtained by cyclic translation and the output image obtained by cyclic translation into the discriminator D to obtain the data D(x) of the authenticity of each image. If the image is a generated image, it is processed with 1 - D(x), then take the logarithm of the data and calculate the maximum likelihood estimate, and add the coefficients to obtain the adversarial loss, that is

[0031] Reconstruction loss: All the final outputs of the non - translation path, self - translation path, and cyclic translation path are reconstructed images of the original image. Therefore, the reconstruction target should be equal to the original image. Calculate the distance between the output images of the above translation paths and the original image, and further calculate the maximum likelihood estimate, and add them to obtain the reconstruction loss, that is

[0032] Style loss: The style code extracted from the translated image should be equal to the generated style code. Therefore, calculate the distance between the two style codes and calculate the maximum likelihood estimate to obtain the style loss, that is

[0033] Triple consistency loss: Experiments show that when the network calculates the reconstruction loss, the generator will inject hidden conditions into the intermediate image, causing the image after re - editing to restore to the original image, thus losing the meaning of calculating the reconstruction loss. Therefore, the triple consistency loss is adopted. The images obtained after one - time editing and two - time editing of the style of the same attribute should be equal. Calculate the distance between the above two images and calculate the maximum likelihood estimate to obtain the triple consistency loss, that is

[0034] The overall objective function is: where λ is a hyperparameter and L is each loss function Represents maximizing the discrimination effect of the discriminator D. Represents minimizing the loss functions of E, G, T, F, and M on the premise of maximizing D.

[0035] Beneficial effects: In the present invention, in the aspect of processing the data set, the processing of adding masks to the real face data set is adopted, enhancing the practicability of the algorithm; in terms of the network structure, the present invention uses a generative adversarial network, enhancing the robustness of the model to the removal and wearing of masks; in terms of the network objective function, different from ordinary adversarial networks, the present invention adds a triple consistency loss function, making the effect of the generator better. Through experiments, it is found that the model provided by this method can basically realize the removal and wearing of masks for any frontal face or semi-profile face, and the editing results are diverse, and the finally obtained face images are also realistic. Specific implementation manner

[0036] The following further explains the present invention with reference to the accompanying drawings.

[0037] A face de-occlusion method based on a conditional generative adversarial network of the present invention includes the following steps:

[0038] Step 1: Prepare the data set. Obtain CelebAhq as the data set for training the model this time. This data set has a total of 30,000 face images annotated with 40 face attributes; perform alignment, cropping, classification, etc. on it, search for possible mask styles on the Internet, randomly add masks to one-tenth of the face images, constituting a data set composed of original face images and face images with masks added, and add annotations on whether the masks are worn to all images; the first 3,000 images in the data set are used as test examples, and the latter 27,000 images are used as training examples; randomly shear and randomly flip the images in the training set to alleviate the overfitting situation of the model.

[0039] Step 2: Build a generative adversarial network model, whose input is the face images and corresponding annotations in the data set obtained in Step 1, and the target images are the appearances of the input images with masks worn or removed. After training the network, the functions of wearing masks and removing masks are realized.

[0040] In the network, for the labeled tags, two concepts of attribute and style are introduced: the attribute represents the manifestation of the tag, and the style represents the specific style of the attribute. For example, for the tag "mask", whether the mask is worn is the attribute, and for the attribute of wearing a mask, the specific mask style is the style.

[0041] The overall structure of the generative adversarial network is as follows:

[0042] Input → Encoder E → Editor T → Generator G → Output;

[0043] Among them, the input is the image x to be edited. The encoder E encodes the direct features of the input image as e = E(x). The editor T edits the encoding according to the style code s to get e' = T(e, s). The generator G decodes the encoding e' to get x' = G(e'), obtaining the final output image x'; all model parameters are jointly trained with the discriminator;

[0044] The style code refers to the code that matches the style of the edited image, which can be obtained from random noise through a mapper or from a reference picture through an extractor;

[0045] The mapper module consists of an MLP, i.e., a multi-layer perceptron. Labels are injected before the first layer, and attributes are injected in the middle layer for indexing. The structure of the mapper is as follows:

[0046] input->I->LinearReLU_1->LinearReLU_2->LinearReLU_3->j->LinearReLU_4->LinearReLU_5->Linear>output

[0047] Among them, LinearReLU_i represents the i-th convolutional layer of the mapper, i = 1, 2,..., 5; i represents the label; j represents the attribute; input is the random noise; output is the style code.

[0048] The extractor module consists of two convolutional layers, five downsampling blocks, and one pooling layer. Labels are injected before the last layer for indexing. The structure of the extractor is as follows:

[0049] input->conv_1->DownResBlock_1->DownResBlock_2->DownResBlock_3->DownResBlock_4->DownResBlock_5->GlobalAvgPool->i->conv_2->output

[0050] Among them, conv_i represents the i-th convolutional layer of the extractor, i = 1, 2; DownResBlock_i represents the i-th downsampling block of the extractor, i = 1, 2,..., 5; the downsampling block inherits ResBlock, i.e., the pre-activated residual unit; GlobalAvgPool represents the pooling layer; i represents the label; input is the reference image; output is the style code.

[0051] The encoder module consists of one convolutional layer and two downsampling blocks, and the structure is as follows:

[0052] input->conv->DownResBlockIN_1->DownResBlockIN_2->output

[0053] The generator module consists of two upsampling blocks and a convolutional layer, and its structure is as follows:

[0054] input->UpResBlockIN_1->UpResBlockIN_2->conv->output

[0055] The converter module consists of a convolutional layer, eight adaptive instance normalization instant blocks that reduce the channel dimension, and a convolutional layer. The style code related to the attribute is injected into all adaptive instance normalization layers, and the scaling and shift vectors are provided through a linear layer. The label is used to establish an index before the first layer. The input is the encoding, and the output is the encoding modified according to the style code. The structure is as follows:

[0056] input->conv_1->ResBlockAdaIN_1->ResBlockAdaIN_2->ResBlockAdaIN_3->ResBlockAdaIN_4->ResBlockAdaIN_5->ResBlockAdaIN_6->ResBlockAdaIN_7->ResBlockAdaIN_8->conv_2->Attention->output

[0057] The discriminator module consists of two convolutional layers, five downsampling blocks, and a pooling layer. The attribute is injected before the first layer, and the label-independent condition is injected at the last layer. The input is the image, and the output is the true or false of the input image. The structure is as follows:

[0058] input->i->conv_1->DownResBlock_1->DownResBlock_2->DownResBlock_3->DownResBlock_4->DownResBlock_5->GlobalAvgPool->conv_2->output

[0059] For different attributes, there is a widespread imbalance in implicit conditions in the dataset. In the CelebAhq dataset, 83.3% of the images with the label "glasses" attribute being "yes" are male and 65.7% are elderly, while the percentages in the images with the label "glasses" attribute being "no" drop to 36.0% and 20.0% respectively. The discriminator will force the translation operation of these implicit conditions. Therefore, label-independent conditions (such as the labels "male" and "young") are injected into the discriminator to solve this problem.

[0060] Step 3: Train the generative adversarial network built in Step 2.

[0061] The network has four training paths, described as follows:

[0062] Non-translation path: Encode the original image x i,j using the encoder E, where i represents the label and j represents the style, and then decode it using the generator G to obtain the non-translation reconstructed image of the original image, i.e., x′ i,j = G(E(x i,j ));

[0063] Self-translation path: Use the extractor module F to extract the style code s i,j of label i from the original image x i,j , i.e., s i,j = F i (x i,j ), then edit the style code and the original image encoding by inputting them into the converter T, and finally decode the encoding using the generator G to obtain the self-translation reconstructed image of the original image, i.e., x” i,j = G(T(E(x i,j ), s i,j ));

[0064] Cyclic translation path: First, input the random noise z into the mapper module M to generate a random style code related to the target label i Secondly, inject the style code and the original image into the converter T and decode them through the generator G to obtain the converted image

[0065] Finally, encode the converted image and inject it, together with the style code obtained from the original image through the extractor, into the converter and the generator to obtain the cyclic translation reconstructed image, i.e., i,j′ = M i,j′ (z), secondly, inject the two style codes and the original image into the converter and decode them through the generator to obtain the converted image x′ i,j′ = G(T(E(x i,j ), s i,j′ )) and Then inject x′ i,j′ and into the converter and decode them through the generator to obtain the converted image

[0066] The four loss functions are as follows:

[0067] Adversarial loss: Encourages the network to perform realistic operations on the generated and extracted styles. Given a real image x i,j , the intermediate generated image obtained through cycle translation and the output image obtained through cycle translation are input into the discriminator D to obtain the data D(x) representing the authenticity of each image. If the image is a generated image, it is processed with 1 - D(x). Then, the logarithm of the data is taken and the maximum likelihood estimate is calculated. The adversarial loss is obtained by adding the coefficients, i.e.,

[0068] Reconstruction loss: All final outputs of the non-translation path, self-translation path, and cycle translation path are reconstructed images of the original image. Therefore, the reconstruction target should be equal to the original image. The distance between the output images of the above translation paths and the original image is calculated, and the maximum likelihood estimate is further calculated and added to obtain the reconstruction loss, i.e.,

[0069] Style loss: The style codes extracted from the translated images should be equal to the generated style codes. Therefore, the distance between the two style codes is calculated, and the maximum likelihood estimate is calculated to obtain the style loss, i.e.,

[0070] Triple consistency loss: Experiments show that when the network calculates the reconstruction loss, the generator injects hidden conditions into the intermediate image, causing the image after re-editing to return to the original image, thus losing the meaning of calculating the reconstruction loss. Therefore, the triple consistency loss is adopted. The images obtained after the first and second edits of the style of the same attribute should be equal. The distance between the above two images is calculated and the maximum likelihood estimate is calculated to obtain the triple consistency loss, i.e.,

[0071] The overall objective function is: where λ is a hyperparameter and L is each loss function, represents maximizing the discrimination effect of the discriminator D, represents minimizing the loss functions of E, G, T, F, and M on the premise of maximizing D.

[0072] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A face mask editing method based on a generative adversarial network, characterized in that: It includes the following steps: Step 1: Randomly put masks on one-tenth of the face images in the face dataset, and add labels indicating whether a mask is worn or not to obtain a dataset; Step 2: Build a generative adversarial network model. Its input is the face images and corresponding labels in the dataset obtained in Step 1, and the target images are the appearances of the input face images with masks on or masks removed. After training the generative adversarial network model, the functions of putting on masks and removing masks are realized; Among them, the overall structure of the generative adversarial network model is as follows: Input → Encoder E → Editor T → Generator G → Output; Among them, the input is the image x to be edited. The encoder E encodes the direct features of the input face image as e = E(x). The editor T edits the encoding according to the style code s as e' = T(e, s). The generator G decodes the encoding e' as x' = G(e') to obtain the final output image x'; all model parameters are jointly trained with the discriminator; The style code refers to the code that matches the style of the edited image, which is obtained by a mapper from random noise or by an extractor from a reference picture; The mapper module consists of an MLP, that is, a multi-layer perceptron. The input is random noise, the label is injected before the first layer, and the attribute is injected in the middle layer for indexing. The output is the style code; The extractor module consists of two convolutional layers, five downsampling blocks, and one pooling layer. Among them, the downsampling block inherits ResBlock, that is, the pre-activated residual unit. The input is the reference image, the label is injected for indexing before the last layer, and the output is the style code; The encoder module consists of one convolutional layer and two downsampling blocks. The input is the image, and the output is the encoding; The generator module consists of two upsampling blocks and one convolutional layer. The input is the encoding, and the output is the image; The shared modules in the encoder module and the generator module use IN, that is, the instance normalization layer; The converter module consists of two convolutional layers and eight adaptive instance normalization instant blocks that reduce the channel dimension. The style code related to the attribute is injected into all adaptive instance normalization layers, and the scaling and shift vectors are provided through a linear layer. The label is used to establish an index before the first layer. The input is the encoding, and the output is the encoding modified according to the style code; The discriminator module uses the same architecture as the extractor module, but the attribute is injected for indexing before the first layer, and the label-independent condition is injected in the last layer. The input is the image, and the output is the authenticity of the input image; For all residual units, the LReLU function is used as the activation function; Average pooling and nearest neighbor upsampling are used to resample the feature maps; Step 3: Train the generative adversarial network model built in Step 2. There are four training paths: the no-translation path, the self-translation path, the cyclic translation path, and the triple-consistent translation path. A total of four loss functions, namely the adversarial loss, the reconstruction loss, the style loss, and the triple-consistent loss, are calculated, and the network parameters are updated by backpropagation to achieve the optimal editing effect.

2. The method for editing a face mask based on a generative adversarial network according to claim 1, wherein: In Step 1, the face images in the obtained dataset are randomly cropped and randomly flipped horizontally; 3. The face mask editing method based on generative adversarial network according to claim 1, characterized in that: In Step 3, the four training paths are as follows: Translation-free path: the original image x i,j is encoded by the encoder E, where i represents the label and j represents the style, and then decoded by the generator G to obtain the translation-free reconstructed image of the original image, i.e., x′ i,j = G(E(x i,j )); Self - translation path: Use the extractor module F to process the original image x i,j to extract the style code s of label i, that is i,j s i,j = F i (x i,j ). Then, input this style code and the original image encoding into the converter T for editing. Finally, use the generator G to decode the encoding to obtain the self - translation reconstructed image of the original image, that is x'' i,j = G(T(E(x i,j ), s i,j )); Cyclic translation path: First, input the random noise z into the mapper module M to generate a random style related to the target label i Code Secondly, inject the style code and the original image into the converter T and decode through the generator G to obtain the transformed image Finally, encode the transformed image and inject it together with the style code obtained from the original image by the extractor into the converter and the generator to obtain the cyclic translation reconstructed image, that is Triple-consistent translation path: First, use the mapper module to generate a style code related to the target label and another random style code s' i,j′ = M i,j′ (z). Secondly, inject the two style codes and the original image into the converter respectively and decode through the generator to obtain the transformed image x' i,j′ = G(T(E(x i,j ), s i,j′ )) and Then inject x' i,j′ and into the converter and decode through the generator to obtain the transformed image 4. The method for editing a face mask based on a generative adversarial network according to claim 1 or 3, characterized in that: In step 3, the four loss functions are as follows: Adversarial loss: Encourages the network to operate realistically on the generated and extracted styles, transforming the real image x i,j , the intermediate generated image obtained by cyclic translation And the output image obtained by loop translation Input the discriminator D and obtain the data D(x) of the authenticity of each image. If the image is a generated image, it is processed with 1-D(x). Then, the logarithm of the data is taken and the maximum likelihood estimate is calculated. The adversarial loss is obtained by adding the coefficients, that is, Reconstruction Loss: The final outputs of all the non-translation path, self-translation path, and cyclic translation path are the reconstructed images of the original image. Therefore, the reconstruction target should be equal to the original image; calculate the distance between the output image of the above translation paths and the original image, and further calculate the maximum likelihood estimation, and add them up to obtain the reconstruction loss, that is Style loss: The style code extracted from the translated image should be equal to the generated style code. Therefore, calculate the distance between the two style codes and calculate the maximum likelihood estimate to obtain the style loss, that is Triple consistency loss: Experiments show that when the network calculates the reconstruction loss, the generator will inject hidden conditions into the intermediate image, causing the image after re - editing to return to the original image, thus losing the meaning of calculating the reconstruction loss. Therefore, the triple consistency loss is adopted. The images obtained after the first and second edits of the style of the same attribute of the image should be equal; calculate the distance between the above two images and calculate the maximum likelihood estimate to obtain the triple consistency loss, that is The overall objective function is as follows: where λ is a hyperparameter and L is each loss function, which represents maximizing the discrimination effect of the discriminator D, and which represents minimizing the loss functions of E, G, T, F, and M on the premise of maximizing D.