A facial image restoration method based on StyleGAN
By segmenting the real face image into face area and background area, using the encoder to extract hidden code vectors and feature maps, and input them into the StyleGAN generator network for mixing processing, the problem of large differences between the face image repair results in the prior art is solved, and the repair effect with similar structure is achieved and the image is given a real skin texture and gloss.
Patent Information
- Application Number
- CN202210736142.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-06-27
AI Technical Summary
The results after reconstruction in the prior art are quite different from those of the original image, and the structural similarity may not be well guaranteed during the repair process, and it is not possible to effectively impart texture and luster to the real skin.
By segmenting the real face image into face area and background area, the encoder extracts hidden code vectors and feature maps, and inputs them into the StyleGAN generator network, and performs mixed processing to achieve the repair of face images.
It has achieved a significant improvement in the facial image repair ability, ensuring the similarity of structure during the repair process, and giving the image a true skin texture and shiny.
Smart Images

Figure CN115049556B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision, and in particular to a facial image restoration method based on StyleGAN. Background Art
[0002] In recent years, the quality of images generated by Generative Adversarial Networks has been significantly improved, especially for face images. Existing technologies can randomly generate high-quality face images through neural networks. Among them, the most advanced generative adversarial network StyleGAN achieves the most advanced visual quality on high-resolution images. In addition, StyleGAN has a latent space W that can be disentangled by attributes. By randomly sampling in the W space, face images are randomly generated. By embedding the real image into the W space, that is, obtaining the hidden code vector of the real image, and then inputting it into the generator network of StyleGAN, the reconstruction result can be obtained. Existing research has found that embedding the real image into the extended W+ space can obtain a more refined reconstructed image. There are two main methods for embedding the real image into the W+ space. One method is to obtain the best reconstructed image by continuously optimizing the hidden code vector; the other method is to obtain the hidden code vector by a single forward propagation through the encoder method, thereby obtaining the reconstruction result. Since the generator model of StyleGAN contains rich face image information, the face prior information in the generator can be used to complete image restoration. At the same time, StyleGAN uses hidden code vectors to control the generated content. By inputting the hidden code vectors into different layers of the StyleGAN generator network, it can control the generation results of different scales.
[0003] Current facial image restoration technology usually uses a preset algorithm. The reconstructed result is quite different from the original image. During the restoration process, it may not be able to ensure structural similarity and cannot give the texture and luster of real skin, resulting in an unsatisfactory overall effect and inconvenience to the restoration work. Traditional restoration methods rely on the boundary information and texture features of the image to be restored. These methods are generally based on mathematical principles, have poor information generation capabilities, and poor robustness and universality. In summary, there is still much room for improvement in facial image restoration methods. Summary of the invention
[0004] The embodiment of the present application provides a facial image restoration method based on StyleGAN, which solves the technical problem in the prior art that the reconstructed result is quite different from the original image and the structural similarity may not be well guaranteed during the restoration process. The facial image restoration capability is greatly improved and the structural similarity is well guaranteed during the restoration process.
[0005] The embodiment of the present application provides a facial image restoration method based on StyleGAN, comprising the following steps: dividing a real facial image into a facial region and a background region, and using the regions as a training set; performing data enhancement on the data set by horizontal flipping, and setting the original image as a label; training an encoder by using the training set and the label to obtain an encoder network; using the encoder network to respectively extract a latent vector of the real facial image, a latent vector of the facial region of the image to be restored, and a latent feature map of the background region of the image to be restored; mixing the latent vector of the real facial image with the latent vector of the facial region of the image to be restored to obtain a latent vector of the mixed face, and mixing the latent vector of the mixed face with the latent feature map of the background region of the image to be restored. Figure 1 The same is input into the StyleGAN generator network to obtain the repaired face image.
[0006] Furthermore, the encoder is trained using the training set and the label, including the following steps: encoding the image, dividing the face area and the background area into two parts for encoding, wherein, for the face area, the encoder structure combining ResNet50 and SE attention module is used to encode the input face area image to obtain the latent vector of the face part. For the background area, the convolutional neural network is used to extract the background features to obtain the latent feature map of the background part; reconstructing the image, inputting the latent vectors of the face part and the background part into the StyleGAN generator to obtain the reconstructed image; encoder optimization, calculating the L2 distance between pixels, the perceptual similarity score, and the L2 distance of the face identity feature based on the label image and the reconstructed image, and optimizing the encoder network to obtain a trained encoder network.
[0007] Furthermore, the encoder structure combining ResNet50 and SE attention module is used to extract the latent vector of the face area image.
[0008] Furthermore, the dimension of the face latent code vector is 18*512, and the dimension of the background latent code feature map is 512*64*64.
[0009] Furthermore, three loss functions are used to optimize the encoder; among them, the first loss function is to calculate the L2 distance between the image label and the generated image based on the pixel value; the second loss function is to use the VGG16 neural network to extract the deep feature information of the image label and the generated image respectively, and calculate the L2 distance between the deep feature information of the two; the third loss function is to use the face recognition neural network to extract the facial feature information between the image label and the generated image respectively, and calculate the L2 distance for the facial features of the two.
[0010] Furthermore, the latent vector of the real face image and the latent vector of the face region of the image to be restored are mixed in a ratio of 8:10.
[0011] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0012] 1. Due to the use of the encoder method, the reconstruction of the damaged image can be completed in one forward propagation, which is fast; at the same time, because the repair method utilizes the rich face prior knowledge in StyleGAN, the repair details of the facial features are more accurate and realistic.
[0013] 2. Since the damaged facial image can be accurately restored through a pre-trained model, it can give the image real skin texture and luster. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flowchart of a facial image restoration method based on StyleGAN in an embodiment of the present application;
[0015] Figure 2 This is a flowchart of encoder training in an embodiment of the present application;
[0016] Figure 3 Schematic diagram of the structure of the face image restoration method in the embodiment of the present application. DETAILED DESCRIPTION
[0017] The embodiment of the present application discloses a facial image restoration method based on StyleGAN, which solves the technical problem in the prior art that the reconstructed result is quite different from the original image and the structural similarity may not be well guaranteed during the restoration process.
[0018] In response to the above technical problems, the overall idea of the technical solution provided by the present application is as follows: a real face image is divided into a face area and a background area, and the areas are used as a training set; the data set is enhanced by horizontal flipping, and the original image is set as a label; the encoder is trained by using the training set and the label to obtain an encoder network; the encoder network is used to extract the latent vector of the real face image, the latent vector of the face area of the image to be repaired, and the latent feature map of the background area of the image to be repaired; the latent vector of the real face image is mixed with the latent vector of the face area of the image to be repaired to obtain the latent vector of the mixed face, and the latent vector of the mixed face is mixed with the latent feature map of the background area of the image to be repaired. Figure 1 The same is input into the StyleGAN generator network to obtain the repaired face image.
[0019] In order to make the above basic method of the embodiment of the present application more obvious and easy to understand, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0020] Figure 1 This is a facial image restoration method based on StyleGAN in an embodiment of the present application, which is described in detail through specific steps below.
[0021] S1, divides the real face image into the face area and the background area, and uses it as the training set.
[0022] In a specific implementation, a real face image can be segmented into a face area and a background area through a semantic segmentation network.
[0023] In a specific implementation, the photos in the dataset of real face images are personal selfies in the real world, and the dataset of real face images is obtained after collection.
[0024] In a specific implementation, in the face area image, we use RGB (0, 0, 0) to fill the missing background part, and in the background area image, we use RGB (0, 0, 0) to fill the missing face part.
[0025] In the specific implementation, a large number of real face images are used to train StyleGAN to train a StyleGAN generator model that can stably generate diverse face images.
[0026] S2, uses horizontal flipping to perform data augmentation on the dataset and sets the original image as the label.
[0027] In a specific implementation, the unsegmented original image can be used as a label image.
[0028] S3: Train the encoder using the training set and the label to obtain an encoder network.
[0029] In the specific implementation, refer to Figure 2 As shown, training can be performed by:
[0030] S31, encoding the image, dividing the face area and the background area into two parts for encoding, wherein, for the face area, using the encoder structure combining ResNet50 and SE attention module, the input face area image is encoded to obtain a latent vector of the face part. For the background area, using a convolutional neural network to extract features from the background to obtain a latent feature map of the background part.
[0031] In the specific implementation, for the encoder network processing the face area, a structure combining ResNet50 and SE attention module can be used. There are 23 convolution blocks in total. Each convolution block contains a BatchNormal layer, a two-dimensional convolution layer, a LeakyReLU activation function and a SE attention module. The input will be connected to the output of the SE module after maximum pooling. This jump connection structure improves the information flow and effectively avoids the gradient vanishing problem caused by the network being too deep.
[0032] And the feature map f1 output by the 6th convolution block, the feature map f2 output by the 20th convolution block, and the feature map f3 output by the 23rd convolution block can be taken out, added and connected by upsampling, and converted into feature maps c1, c2 and c3, where c1 = f3, c2 = upsample(c1) + f2, c3 = upsample(c2) + f1. Shallow features contain more detailed information, while deep features pay more attention to the overall situation and do not pay attention to image details. Using the network structure of the feature pyramid to fuse deep and shallow features can maintain the global features and semantic information of the image while paying attention to the detailed information.
[0033] In the specific implementation, for the network module for constructing feature map conversion latent vector, the module is composed of two-dimensional convolution, LeakyReLU activation function, and fully connected layer, which processes c1, c2, and c3 respectively, converts feature map c1 into a 3*512-dimensional latent vector, converts feature map c2 into a 4*512-dimensional latent vector, and converts feature map c3 into a 11*512-dimensional latent vector. The obtained latent vectors are concatenated to obtain the final 18*512-dimensional latent vector.
[0034] In the specific implementation, for the encoder network that processes the background area image, this method uses the same convolution block as the face encoder. Since this method processes the background into a latent feature map, only 6 layers of convolution blocks are used to process the background image. Each convolution block contains a BatchNormal layer, a two-dimensional convolution layer, a ReLU activation function, and an SE attention module, and the input is connected to the output of the SE module after maximum pooling, and the background area image is processed into a 512*64*64-dimensional latent feature map through the background encoder network.
[0035] S32, reconstructing the image, inputting the latent vectors of the face part and the background part into the StyleGAN generator to obtain a reconstructed image.
[0036] In the specific implementation, the output of the encoder network is connected to the input of the StyleGAN network. The output of the face image encoder is connected to the input of the StyleGAN generator, and the output of the background image encoder is fused with the feature map of the middle layer of the StyleGAN generator. The 18*512 dimensional latent code vector output by the face image encoder is input into different layers of the StyleGAN generator to control the face generation effect of different scales. The 512*64*64 output of the background image encoder is weightedly fused with the feature map of the middle layer of the StyleGAN generator, and accurate reconstruction of the background is achieved by suppressing and enhancing certain areas of the feature map of the middle layer of the generator.
[0037] In the specific implementation, when training the encoder, the weights of the StyleGAN generator network are fixed, and the loss is calculated using the image generated by the StyleGAN generator and the pre-set label image to optimize the encoder. Because the weights of the StyleGAN generator network are fixed, when the image generated by StyleGAN is similar to the label image, it proves that the latent vector and latent feature map generated by the encoder can accurately express the face image and background image.
[0038] In order to measure the similarity between the generated image and the label image and use the similarity to calculate the loss, the encoder is optimized. The total loss function is L, which consists of three loss functions. The first loss function is to calculate the mean square error L between the image label and the generated image based on the pixel value. mse The second loss function uses the VGG16 neural network to extract the deep feature information of the image label and the generated image, and calculates the mean square error L between the deep feature information of the two. lpips The third loss function uses the face recognition neural network to extract the facial feature information between the image label and the generated image, and calculates the mean square error L for the facial features of the two. id .
[0039] L mse =‖IG(E(I))‖2
[0040] L lpips =‖LPIPS(I)-LPIPS(G(E(I)))‖2
[0041] L id =‖ID(I)-ID(G(E(I)))‖2
[0042] Where I is the input image, E is the trained encoder network, and G is the trained StyleGAN generator network. LPIPS is a pre-trained VGG16 network used to extract deep features of images and calculate the perceptual similarity of two images. ID is a pre-trained face recognition network used to extract identity features of faces in images.
[0043] The total loss function is L total L total =λ mse L mse +λ lpips L lpips +λ id L id
[0044] Among them, L mse is the mean square error between the pixel values of the two images, λ mse =1.0 is the weight coefficient of the loss. lpips is the mean square error of the deep features of the two images, λ lpips =0.8 is the weight coefficient of the loss. id is the mean square error of the facial features of the two images, λ id =0.5 is the weight coefficient of the loss.
[0045] S33, encoder optimization, calculates the L2 distance between pixels, the perceptual similarity score, and the L2 distance of the face identity feature according to the label image and the reconstructed image, and optimizes the encoder network to obtain a trained encoder network.
[0046] In the specific implementation, the batch size can be set to 8, the number of iterations to 300,000 times, and the learning rate to 1e-4. According to the batch size of 8, 8 samples are taken from the real face image each time, and the face image and background image of these 8 samples are obtained by using the semantic segmentation algorithm. They are input into the face encoder network and the background encoder network respectively to obtain the corresponding latent vector and latent feature map, which are then input into the StyleGAN generator to obtain the generated image, complete the forward propagation, and then calculate the loss through the carefully set loss function and weights, and back propagate to optimize the face encoder and background encoder networks.
[0047] S4, using the encoder network to respectively extract the latent vector of the real face image, the latent vector of the face area of the image to be restored, and the latent feature map of the background area of the image to be restored.
[0048] In the specific implementation, please refer to Figure 3As shown, the face recognition library Dlib can be used to locate the facial key points of the image to be repaired, crop the face image to be repaired, and then use the semantic segmentation algorithm to divide the face image to be repaired into a face area image and a background area image.
[0049] S5, mixing the hidden code vector of the real face image with the hidden code vector of the face area of the image to be restored to obtain a hidden code vector of the mixed face, and comparing the hidden code vector of the mixed face with the hidden code feature of the background area of the image to be restored. Figure 1 The same is input into the StyleGAN generator network to obtain the repaired face image.
[0050] In the specific implementation, the latent vector of the human image to be repaired and the latent vector of the real face image are mixed to obtain a mixed latent vector, and the mixing ratio is 8:10. The first 8*512 dimensions of the latent vector of the face to be repaired are used, and the last 10*512 dimensions of the latent vector of the real face are used, and they are spliced into a new 18*512-dimensional latent vector. Since StyleGAN generates different face images by controlling the latent vector, different dimensions in the latent vector control the generation of images of different scales. The mixing ratio is set to 8:10. While fully considering the facial prior information contained in the StyleGAN generator network, the rough facial features style, appearance and other information in the face image to be repaired are retained.
[0051] And, the mixed hidden code vector is combined with the background hidden code feature of the image to be repaired Figure 1 The face image and the background image are processed separately, and the background information is saved separately using the latent feature map, which helps to reconstruct diverse background information.
[0052] In summary, due to the adoption of a facial image restoration method based on StyleGAN, the facial identity information of the image to be restored is maintained while the facial features, skin, texture, and gloss are restored. First, by training the StyleGAN generator, rich facial prior knowledge is obtained. Secondly, the encoder network sets pixel-level loss, overall perceptual similarity loss, and facial attribute similarity loss on the image, so that the encoder can accurately express facial information and background information through hidden vectors and feature maps. Under the dual control of hidden vectors and hidden feature maps, the reconstructed image has both the facial features and appearance information of the facial image to be restored, and can increase the skin gloss and texture. While maintaining the identity information of the face to be restored, the facial prior knowledge in the StyleGAN generator is used to greatly supplement the detailed information of the facial image to be restored, ensuring the accuracy and robustness of the restoration.
[0053] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0054] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A system that specifies the functions of a box or multiple boxes.
[0055] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction system, which is implemented in the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0056] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps of the functions specified in a box or multiple boxes. Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the attached claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0057] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A facial image restoration method based on StyleGAN, characterized in that: The following steps are involved: Segment the real face image into the face area and the background area and use them as the training set; The dataset is augmented with horizontal flipping, with the original image set as the label; Using the training set and the labels to train an encoder to obtain an encoder network; The encoder network is used to respectively extract a latent vector of a real face image, a latent vector of a face region of the image to be restored, and a latent feature map of a background region of the image to be restored; Mixing the latent vector of the real face image with the latent vector of the face area of the image to be repaired to obtain a latent vector of the mixed face, and inputting the latent vector of the mixed face together with the latent feature map of the background area of the image to be repaired into the StyleGAN generator network to obtain a repaired face image; Training the encoder using the training set and the label comprises the following steps: Encoding the image, dividing the face area and the background area into two parts for encoding, wherein, for the face area, using the encoder structure combining ResNet50 with the SE attention module, the input face area image is encoded to obtain a latent vector of the face part, and for the background area, using a convolutional neural network to extract features from the background to obtain a latent feature map of the background part; Reconstructing the image, inputting the latent vectors of the face part and the background part into the StyleGAN generator to obtain a reconstructed image; Encoder optimization: Calculate the L2 distance between pixels, the perceptual similarity score, and the L2 distance of face identity features based on the labeled image and the reconstructed image, and optimize the encoder network to obtain a trained encoder network.
2. The method for facial image restoration based on StyleGAN according to claim 1, characterized in that: The encoder structure combining ResNet50 and SE attention module is used to extract the latent vector of the face area image.
3. The method for facial image restoration based on StyleGAN as claimed in claim 1, characterized in that: The dimension of the face latent vector is 18*512, and the dimension of the background latent feature map is 512*64*64.
4. The method for facial image restoration based on StyleGAN as claimed in claim 1, characterized in that: The encoder is optimized using three loss functions. The first loss function calculates the L2 distance between the image label and the generated image based on the pixel value. The second loss function uses the VGG16 neural network to extract the deep feature information of the image label and the generated image respectively, and calculates the L2 distance between the deep feature information of the two. The third loss function uses the face recognition neural network to extract the facial feature information between the image label and the generated image respectively, and calculates the L2 distance based on the facial features of the two.
5. The method for facial image restoration based on StyleGAN as claimed in claim 1, characterized in that: The latent vector of the real face image and the latent vector of the face area of the image to be restored are mixed in a ratio of 8:10.
Citation Information
Patent Citations
Face image generation method and device
CN108197525A
Face image restoration method introducing attention mechanism
CN111612718A