Face makeup transfer system and method based on image conversion and stylegan hybrid framework
Patent Information
- Application Number
- CN202310542002.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-15
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-05-15
AI Technical Summary
然而,这种“先反演,后编辑”的策略存在一个关键问题,即根据信息有损压缩理论,将图像压缩到低比特风格编码不可避免地会丢失一些信息
[0016]对于妆容色彩迁移,本发明考虑了图像反演中丢失的妆容信息,并借助于生成对抗网络强大的生成能力,合成了图像真实且妆容自然的妆容迁移结果;
Smart Images

Figure CN116681580B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of face editing technology, specifically to a face makeup transfer system and method based on a hybrid framework of image conversion and StyleGAN. Background Technology
[0002] With its potential for widespread applications in e-commerce and video chat, makeup transfer technology has attracted great attention from scholars both at home and abroad and has become a hot topic in computer vision. Given a pair of source images and reference images, the goals of makeup transfer are mainly twofold: (1) to render the makeup style from the reference image into the source image. Makeup styles include low-frequency colors (e.g., lipstick, eyeshadow, foundation) and high-frequency textures (e.g., skin texture, eyelashes). (2) to preserve the content of the source image, including identity, hairstyle, background, etc.
[0003] Many deep learning methods based on image transformation frameworks have been widely applied to makeup transfer tasks. For example... Figure 2 As shown, these methods first employ a fully convolutional encoder to capture content feature maps from the source image. Then, they are fused with makeup styles captured from a reference image to generate the makeup style transfer result. The feature maps contain rich spatial details, and the direct extraction of makeup styles from the original image makes these methods effective in content preservation and low-frequency color transfer. Besides low-frequency colors, high-frequency textures (such as eyelashes and skin texture) are also an important component of makeup styles. However, image transformation-based methods lack the ability to render high-frequency textures. Figure 3 In contrast, EleGANt is unable to synthesize eyelash details and transfer skin textures. Furthermore, its de novo training and recurring consistency paradigm make it difficult to apply to high resolutions, limiting its use on today's high-definition devices.
[0004] Recently, generating high-resolution images with realistic textures and details using the StyleGAN (Analyzing and Improving the Image Quality of StyleGAN) framework has become a common technique in facial image editing. Figure 2 In traditional makeup transfer methods, the StyleGAN framework primarily inverts the real image into the latent space and inputs the modified style encoding into a pre-trained StyleGAN to obtain the desired result. However, this "invert first, then edit" strategy has a key problem: according to lossy compression theory, compressing the image to low-bit style encoding inevitably results in the loss of some information. In other words, low-bit style encoding cannot guarantee high-fidelity image inversion. This problem leads to content distortion and color deviation in the makeup transfer results. Figure 3In this study, methods based on the StyleGAN framework suffer from issues such as facial structure deformation, background variation, and color differences. Summary of the Invention
[0005] The purpose of this invention is to provide a face makeup transfer system and method based on a hybrid framework of image transformation and StyleGAN. The hybrid method of this invention sacrifices some content details, but in exchange for better image quality, more accurate makeup transfer performance, and higher resolution.
[0006] To achieve this objective, the face makeup transfer system based on the hybrid framework of image conversion and StyleGAN designed in this invention is characterized by comprising an image inversion module, a feature map acquisition module, a first-stage training module, a preliminary transfer result acquisition module, a second-stage training module, a color correction module, a third-stage training module, and a makeup transfer module.
[0007] The image inversion module is used to calculate the inversion style codes of the source image and the reference image respectively using the inversion encoder, and to generate the inversion images of the source image and the reference image using the StyleGAN2 generator. The inversion images are compressed to obtain the residual map between the source image and the inversion compressed image of the source image and the residual map between the reference image and the inversion compressed image of the reference image.
[0008] The feature map acquisition module is used to input the residual map of the source image and the inverted compressed image of the source image into the multi-scale residual encoder to obtain multi-resolution feature maps, and inject the multi-resolution feature maps into the feature map of the StyleGAN2 coarse layer based on the fusion network to obtain the hybrid feature map.
[0009] The first-stage training module is used to train a multi-scale residual encoder and a fusion network on the CelebA-HQ dataset using a loss function;
[0010] The preliminary transfer result acquisition module is used to invert the style code using the detail layer of the source image and the reference image, obtain the change in style code through the first mapping network, and input the change in the mixed feature map and style code into the StyleGAN2 generator to generate the preliminary transfer result.
[0011] The second-stage training module is used to input the initial transfer results into the overall loss function to obtain the overall loss information, and then optimize the first mapping network using backpropagation based on the overall loss information.
[0012] The color correction module is used to distill the lost makeup information from the residual map of the reference image and the inverted compressed image of the reference image using a residual encoder, and to embed the lost makeup information into the inverted image of the source image using a second mapping network, so as to obtain the source image fine-layer inversion style code embedded with the reference makeup style. The fine-layer style code and the mixed feature map are input into the StyleGAN2 generator to generate the final transfer result.
[0013] The third-stage training module is used to input the final transfer result into the overall loss function to obtain the overall loss information, and then use backpropagation to optimize the residual encoder and the second mapping network using the overall loss information.
[0014] The makeup transfer module is used to construct a hybrid framework CPTR (Content-Preserving and Texture-Rendering Framework) based on image transformation and StyleGAN, which consists of the image inversion module, feature map acquisition module, preliminary transfer result acquisition module and color correction module. Then, the source image and reference image are input into this hybrid framework to obtain the makeup transfer result.
[0015] The beneficial effects of this invention are:
[0016] For makeup color transfer, this invention takes into account the makeup information lost in image inversion and uses the powerful generative capabilities of generative adversarial networks to synthesize makeup transfer results that are realistic in image and natural in makeup.
[0017] For makeup texture transfer, this invention benefits from StyleGAN’s powerful property control capabilities, which can generate eyelashes and change skin texture, making it the only method capable of rendering makeup texture.
[0018] This invention is based on the pre-trained StyleGAN architecture. Benefiting from StyleGAN's powerful image generation capabilities, this invention is also the only method that supports high-resolution (1024×1024) makeup transfer.
[0019] Qualitative Comparison: A qualitative comparison was conducted between this invention and five state-of-the-art makeup transfer methods: LADN (Local Adversarial Disentangling Network for Facial Makeup and De-Makeup), PSGAN (Pose and Expression Robust Spatial-Aware GAN for Customizable Makeup Transfer), SCGAN (Spatially-invariant Style-codes Controlled Makeup Transfer), SSAT (A Symmetric Semantic-Aware Transformer Network for Makeup Transfer and Removal), and EleGANt (Exquisite and Locally Editable GAN for Makeup Transfer). Color transfer comparisons are also included. Figure 8 As shown. LADN's results exhibit some artifacts and blurring. PSGAN renders eyeshadow lighter than the reference color. SCGAN fails to transfer eyeshadow to the appropriate locations, resulting in dark patches. Recent methods such as SSAT and EleGANt produce more visually acceptable results. However, SSAT fails to transfer the correct foundation color, while EleGANt produces uneven facial makeup with unnatural lighting. In contrast, the method of this invention generates the most realistic images and renders accurate and natural makeup. For texture transfer, comparisons are made as follows: Figure 9 As shown. The proposed CPTR (Content-Preserving and Texture-Rendering Framework) hybrid framework is the only method with the ability to transfer makeup textures, generating eyelashes and altering skin textures. Regarding content preservation, while the method of this invention recovers most of the source image content information, there is still room for improvement in details such as hair strands, earrings, and eye protection under extreme gaze directions. Arrows indicate details missing during the inversion.
[0020] Quantitative Comparison: A quantitative comparison of the proposed CPTR method with other methods is presented here. Images generated by all methods are aligned to the same resolution (256×256).
[0021] Regarding image quality, the quality and realism of the generated images were evaluated using FID (GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium). This invention calculated the FID score between the generated image and the reference image (lower scores are better). The FID scores for each method are reported in Table 2. The method of this invention achieved the lowest FID scores on both datasets compared to other methods, meaning that the results produced by the method of this invention are closer to real images.
[0022] Regarding makeup diversity, the average LPIPS (The Unreasonable Effectiveness of Deep Features as a Perceptual Metric) distance (higher scores are better) between makeup transfer output pairs with the same source image but random reference images is used to measure makeup diversity. Table 2 shows that, based on the LPIPS distance, the makeup diversity of the method in this invention is significantly higher than other methods.
[0023] For content protection, the high-level feature maps (relu5_2) of the VGG (Very Deep Convolutional Networks for Large-Scale Image Recognition) model are used to calculate the perceptual distance between the generated image and the corresponding source image (lower scores are better). PSGAN, which directly uses perceptual loss to optimize the network, obtained the lowest score. Consistent with visual observation, the hybrid framework of this invention is slightly inferior to the image transformation framework in terms of content protection.
[0024] Overall, the hybrid method of this invention sacrifices some content detail, but in exchange for better image quality, more accurate makeup transfer performance, and higher resolution. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the structure of the present invention;
[0026] Figure 2 The diagrams show a comparison of the overall workflow of different frameworks in this invention. (a) The image transformation framework uses a fully convolutional encoder to capture content feature maps and fuses them with makeup styles. This method is effective for content preservation and color transfer. (b) The StyleGAN-based framework inverts real faces into the latent space and applies the modified style encoding to a pre-trained StyleGAN. It can render textures and supports high resolution. (c) This invention proposes a hybrid framework, CPTR, to inherit the advantages of both frameworks.
[0027] Figure 3 The images show a comparison of makeup transfer effects using different frameworks in this invention. The image conversion framework cannot accurately transfer eyelash details and skin texture. The method based on the StyleGAN framework has problems such as facial structure deformation, background changes, and color differences. The proposed hybrid framework avoids these shortcomings.
[0028] Figure 4 This diagram illustrates content distortion and color degradation. Content distortion: Problems such as facial structure deformation, background and hairstyle alterations exist during source image inversion. Color degradation: Problems such as disappearing eyeshadow and lipstick color deviation appear during reference image inversion.
[0029] Figure 5 This is a network detail diagram of the hybrid framework CPTR of this invention. In CPTR, residual learning is used to recover information lost during the inversion process. Based on the decoupling property of StyleGAN, the Multi-Scale Residual Inversion (MSRI) module modifies the coarse layer feature map of StyleGAN to preserve content, while the Style Modulation (SM) module and the Color Refinement (CR) module adjust the style encoding of the detail layer to achieve the transfer of makeup effects.
[0030] Figure 6 This is a detailed diagram of the fusion network of the present invention. The fusion network adaptively controls and propagates residual information through a gating mechanism.
[0031] Figure 7 This is a detailed map of the mapping network of the present invention. AdaIN is used to conditionally embed a reference makeup look into the source image.
[0032] Figure 8 This is a comparison chart with existing methods in color transfer. Compared with other methods, the images generated by this invention are more realistic and the makeup is more natural.
[0033] Figure 9 This is a comparison chart with existing methods in texture transfer. The present invention can generate eyelashes and change skin texture, and is the only method capable of rendering makeup textures.
[0034] Figure 10 This is an ablation study diagram of the proposed modules. The invention proposes a multi-scale residual inversion (MSRI) and color correction (CR) modules to address content distortion and color degradation issues in inversion. This invention has conducted ablation studies on these two modules to verify their effectiveness.
[0035] Figure 11This is a makeup intensity control chart. The present invention transfers makeup style by predicting the change in style coding. Therefore, the makeup intensity can be controlled by interpolating the change.
[0036] Figure 12 This is a robustness test result diagram for complex scenarios. This invention tested the robustness of CPTR under complex conditions such as pose, lighting, and occlusion.
[0037] Figure 13 This is a flowchart of the present invention. Detailed Implementation
[0038] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0039] like Figure 1 The system shown is a face makeup transfer system based on a hybrid framework of image conversion and StyleGAN. It includes an image inversion module, a feature map acquisition module, a first-stage training module, a preliminary transfer result acquisition module, a second-stage training module, a color correction module, a third-stage training module, and a makeup transfer module.
[0040] The image inversion module is used to calculate the inversion style codes of the source image and the reference image respectively using the inversion encoder, and to generate inversion images of the source image and the reference image using the StyleGAN2 generator. The inversion images are compressed to obtain the residual map between the source image and the inversion compressed image of the source image and the residual map between the reference image and the inversion compressed image of the reference image. The residual map obtained by this module provides the necessary data for subsequent missing information completion and makeup transfer operations.
[0041] The feature map acquisition module (Multi-Scale Residual Inversion (MSRI) module) is used to input the residual map of the source image and the inverted compressed image of the source image into the multi-scale residual encoder to obtain multi-resolution feature maps. Based on the fusion network, the multi-resolution feature maps are injected into the feature map of the StyleGAN2 coarse layer to obtain a hybrid feature map. Inspired by the fact that feature maps in the image transformation framework contain more local spatial information, this design extracts feature maps of residual information instead of style encoding, thereby better protecting the content information of the source image.
[0042] The first-stage training module is used to train a multi-scale residual encoder and a fusion network on the CelebA-HQ dataset using a loss function. After this training, the proposed framework can achieve high-fidelity image inversion, thereby better preserving the content attributes of the source image.
[0043] The preliminary transfer result acquisition module is used to invert style coding through the detail layers of the source and reference images, obtain the change in style coding through the first mapping network, and input the changes in the mixed feature map and style coding into the StyleGAN2 generator to generate preliminary transfer results. The decoupling injection strategy for different layers enables this invention to inherit the powerful attribute control capability of StyleGAN, protecting content attributes without reducing the makeup editing capability of the framework.
[0044] The second-stage training module is used to input the preliminary transfer results into the overall loss function to obtain the overall loss information. The first mapping network is then optimized using backpropagation based on the overall loss information. After this training, the inverted makeup attributes are injected into the source image inversion results, thus enabling the transfer of the inverted makeup style and obtaining the preliminary makeup transfer results.
[0045] The color correction module uses a residual encoder to distill the lost makeup information from the residual map of the reference image and the inverted compressed image of the reference image. It then uses a second mapping network to embed the lost makeup information into the inverted image of the source image, resulting in a fine-layer inversion style code of the source image embedded with the reference makeup style. This fine-layer style code and the mixed feature map are input into the StyleGAN2 generator to generate the final transfer result. The color correction module then injects the lost makeup information back into the source image inversion style code, thereby recalling the lost makeup information and correcting the color information.
[0046] The third-stage training module is used to input the final transfer result into the overall loss function to obtain the overall loss information. The overall loss information is then used to optimize the residual encoder and the second mapping network through backpropagation. After this training, the residual encoder learns the information of the lost makeup and injects it into the source image inversion result through the second mapping network to generate the final accurate makeup result.
[0047] The makeup transfer module, named CPTR, integrates the image inversion module, feature map acquisition module, preliminary transfer result acquisition module, and color correction module into a hybrid framework based on image transformation and StyleGAN. The source and reference images are then input into this hybrid framework to obtain the makeup transfer result. This framework benefits from a decoupled injection strategy that modifies feature maps at the coarse layer and edits style codes at the detail layer, thus inheriting the content preservation advantages of the image transformation framework and the texture editing and high-resolution synthesis advantages of the StyleGAN framework.
[0048] In the above technical solution, the specific method by which the image inversion module uses the e4e (Designing an encoder for StyleGAN image manipulation) inversion encoder E0 to calculate the inversion style codes of the source image x and the reference image y respectively is as follows: This invention is based on the inversion results of existing methods, and on this basis, residual information is injected to obtain high-fidelity inversion results. Such residual learning reduces the difficulty of network training.
[0049]
[0050]
[0051] Among them, w x For the inversion style encoding of the source image x, w y For the inversion style code of the reference image y, E0(x) represents the style code output from the inversion encoder E0 when the source image x is input into it. Each represents 18 cascaded components forming w x Individual style coding, Each represents 18 cascaded components forming w y Individual style coding;
[0052] Based on the inversion style encoding of the source and reference images, the inversion image is generated using the StyleGAN2 generator G0 pre-trained on FFHQ data. This allows the framework of this invention to leverage the powerful image generation and attribute editing capabilities of StyleGAN2. The specific pre-training process of StyleGAN2 is described in the paper "Analyzing and Improving the ImageQuality of StyleGAN".
[0053] x0=G0(w x ), y0=G0(w y ), where x0 is the first inversion result of the source image, y0 is the first inversion result of the reference image, and G0(W x ) represents the inversion style encoding w of the source image x. x The input is fed into the StyleGAN2 generator G0, G0(w y ) indicates that the inversion style encoding w of the reference image y will be performed. y Input into the StyleGAN2 generator G0;
[0054] Bilinear downsampling is used to compress the inverted image, which is then scaled to a resolution of 256*256. Residual images are obtained between the source image and the inverted compressed image, and between the reference image and the inverted compressed image.res =x(i)-x 01 (i), y res =y(i)-y 01 (I), where x res The residual image represents the difference between the source image and the inverted compressed image. res This represents the residual image between the reference image and the inverted compressed image of the reference image, where x(i) represents the pixel value of each pixel in the source image. 01 (i) represents the pixel value of each pixel in the inverted compressed image of the source image, and y(i) represents the pixel value of each pixel in the reference image. 01 (i) represents the pixel values of each pixel in the inverted compressed image of the reference image. This operation is used to obtain the residual map, providing the necessary data for subsequent missing information completion and makeup transfer operations.
[0055] In the above technical solution, the feature map acquisition module is used to obtain the residual map x between the source image and the inverted compressed image of the source image. res Input to multi-scale residual encoder E c Obtaining multi-resolution feature maps C=E c (x res Furthermore, a multi-resolution feature map C is injected into the feature map of the StyleGAN2 coarse layer through a fusion network based on a gating mechanism, resulting in a hybrid feature map. The multi-resolution feature map C has the same resolution as the feature map of the StyleGAN2 coarse layer. Inspired by the fact that feature maps in image transformation frameworks contain more local spatial information, this module extracts feature maps of residual information instead of style encoding, thereby better preserving the content information of the source image.
[0056]
[0057] Among them, C i Represents the multi-resolution feature map of the i-th layer, with gate variable f. i g (C i ) controls the original variable (the i-th layer multi-resolution feature map C) i The retention of ) and the modification of variable f i c (C i ) is used to correct errors in the original variables, f i g (·) and f i c (·) represents a 1×1 convolutional layer, F i This represents the feature map of the i-th layer of StyleGAN, such as... Figure 6 As shown.
[0058] In the above technical solution, the first-stage training module is used to train a multi-scale residual encoder and a fusion network on the CelebA-HQ dataset using the loss function for image inversion and reconstruction. The loss function for image inversion and reconstruction can be found in reference 1: Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen Or. 2021. Designing an encoder for StyleGAN image manipulation. In ACM Transactions on Graphics, Vol. 40. 1–14.
[0059] The optimizer's hyperparameter settings are consistent with those in Reference 1 above. During training, the parameters of the inversion encoder E0 and generator G0 are fixed and not updated. After this training, the proposed framework can achieve high-fidelity image inversion, thereby better preserving the content attributes of the source image.
[0060] In the above technical solution, the preliminary migration result acquisition module is used to invert the style encoding of the detail layer of the source image. Detail layer inversion style coding of reference image The input is fed into the first mapping network M0 to calculate the style coding increment of the inverted makeup representation, using the following formula: Inverted style coding of the detail layer of the source image. Detail layer inversion style coding of reference image By extracting the inversion style code w from the source image x x Inversion style encoding w of source image x x get.
[0061]
[0062] Where M0 is the mapping network, These represent the inversion style encodings of the detail layers of the source and reference images, respectively. and This indicates that a single style code is selected for the source and reference images with codes from 8 to 18 (0-7 for coarse layers, 8-18 for fine layers). This indicates that the style encoding of the source image x belongs to the fine layer.
[0063] The mapping network M0 consists of a fully connected layer FC, an adaptive layer AdaIN (Arbitrary Style Transfer in Real-Time with Adaptive Instance Normalization), and an activation layer LeakyReLU stacked together. The adaptive layer AdaIN propagates the reference makeup by modulating the mean and variance of the style encoding, thereby embedding the inverted reference makeup as a condition into the source image.
[0064] Mixed feature maps Changes in style coding The input is fed into the StyleGAN2 generator G0 for progressive image generation, yielding preliminary transfer results.
[0065] In the above technical solution, the overall loss function is:
[0066] L all =λ protect L pritect +λ trans L trans
[0067] Among them, L all Let L be the overall loss function. protect To protect the loss function, L trans Let λ be the migration loss function. protect and λ trans Hyperparameter weights are used to balance two different training objectives: content protection and makeup transfer; the overall loss function L is utilized. all The first mapping network M0 is optimized. During the optimization process, the inversion encoder E0, the generator G0, and the multi-scale residual encoder E0, which was trained in the first stage, are also optimized. c The fusion network has fixed parameters and does not perform gradient updates. After this training, the inverted makeup attributes are injected into the source image inversion results, thus enabling the transfer of the inverted makeup style and obtaining preliminary makeup transfer results.
[0068] In the above technical solution, the protection loss function L protect as follows:
[0069] L protect =λ id L id +λ bg L bg
[0070] Among them, L id Let L be the identity loss function. bg Let λ be the background loss function. id To balance the hyperparameter weights for identity protection, λbg To balance the hyperparameter weights for the protection background, for the identity loss function L id By calculating the migration result x trans The cosine distance between the source image x and the source image x is obtained as follows:
[0071] L id =1-cos(R(x),R(x) trans ))
[0072] Where R represents the pre-trained ArcFace (ArcFace: Additive Angular MarginLoss for Deep Face Recognition) network used for face recognition, and L represents the background loss. bg The calculation is as follows:
[0073] L bg =||(xx) trans )·P bg (x)||2
[0074] Where ||·||2 represents the L2 norm constraint, · represents element-wise multiplication, and P represents the pre-trained face segmentation network BiSeNet (BiSeNet: Bilateral Segmentation Network for Real-Time Semantic Segmentation). bg (x) represents the mask of the background region of the source image x based on BiSeNet prediction. In makeup transfer, the clothing and hair regions are also considered as background.
[0075] In makeup transfer, the two main components that need to be protected are the background and facial identity. Therefore, the protection loss function described above was designed for training the framework of this invention.
[0076] Due to the unique nature of makeup transfer tasks, it is difficult to prepare ground truth values for a pair of source and reference images. Existing methods commonly employ pseudo-paired data synthesis based on histogram matching or geometric distortion for network training. Unlike previous methods that require synthesizing pseudo-paired data, this invention does not. Instead, it introduces a contextual loss (The Contextual Loss for Image Transformation with Non-aligned Data) to guide the network in transferring makeup styles under unsupervised conditions.
[0077] L trans =L CX (y·P face (y),x trans ·P face(x trans ))
[0078] Among them, L trans L represents the migration loss function. CX P represents contextual loss. face (y) represents the face mask of the predicted reference image y.
[0079] This invention avoids the tedious synthesis of pseudo-paired data. Furthermore, during the optimization process, the parameters of the coarse layer controlling image shape are frozen, and only the vector code in the detail layer is optimized, thereby avoiding the identity drift problem caused by upper contextual loss. Note that in this optimization, the transfer loss function L... trans The result of the first inversion, y0, is used to replace y to guide the makeup transfer.
[0080] In the second stage of training, the proposed overall loss function L is used. all Optimize the mapping network M0. During optimization, the inversion encoder E0, generator G0, and the multi-scale residual encoder E0 (trained in the first stage) are optimized. c The fusion network has fixed parameters and does not undergo gradient updates. Where λ... id =0.3, x bg =10, x protect =1 and λ trans =2. For contextual loss, a ReLU2_2 layer from VGG (Very Deep Convolutional Networks for Large-Scale Image Recognition) is used because low-level features capture richer style information, which is beneficial for conveying makeup style. The training iterations are 500,000, using the Adam optimizer, with β1 and β2 set to 0.95 and 0.999, respectively.
[0081] In the above technical solution, the color correction module uses a trainable residual encoder pSp (Encoding in Style: A StyleGAN Encoder for Image-to-Image Translation) to distill the lost makeup information from the residual image of the reference image and the inverted compressed image of the reference image. Specifically, it is expressed as follows:
[0082]
[0083] Among them, E res pSp represents the trainable residual encoder. This represents the lost makeup information obtained from distillation, i.e., the reference residual image y. res The fine-grained part of the inversion style coding;
[0084] The lost makeup information is embedded into the inverted image of the source image using the second mapping network M1, resulting in a fine layer of the source image x inverted style encoding that has been embedded with the reference makeup style, specifically represented as:
[0085]
[0086] Where M1 is the second mapping network, The style encoding output by the first mapping network M0, This indicates the loss of makeup obtained through distillation. This indicates that a fine layer of source image x inversion style encoding has been embedded with reference makeup style, and mapping networks M0, M1 and have the same structure;
[0087] A fine layer containing a source image with embedded reference makeup style and inverse style encoding. With mixed feature maps The input is fed into the StyleGAN2 generator G0 to generate the final transfer result. The color correction module re-injects the lost makeup information into the style encoding of the source image, thereby recalling the lost makeup information and correcting the color information.
[0088] The third training module is used to input the final transfer result into the overall loss function to obtain the overall loss information. This overall loss information is then used to optimize the residual encoder and the second mapping network via backpropagation. The hyperparameter settings are the same as in the second training module. After this training, the residual encoder learns the information about the lost makeup and injects it into the source image inversion result through the second mapping network, generating the final accurate makeup result.
[0089] In the above technical solution, the makeup transfer module integrates the image inversion module, feature map acquisition module, preliminary transfer result acquisition module, and color correction module into a hybrid framework based on image transformation and StyleGAN. The specific process of inputting the source image and reference image into this hybrid framework to obtain the makeup transfer result is as follows:
[0090] Given a source image x and a reference image y, first obtain their respective inversion style codes w. x =E0(x), w y =E0(x), and the inversion result x0=G0(w x ), x0 = G o (w x and residual plot x res =x - x0, y res= y - y0, where E0 is the inversion encoder, G0 is the pre-trained StyleGAN2 generator, and then, for the residual image x of the source image and the inversion compressed image of the source image... res It is input into the multi-scale residual encoder E. c To obtain multi-resolution feature maps C=E c (x res Then, it is injected into the corresponding feature map of the StyleGAN2 coarse layer with the same resolution through a base fusion network, as described in the following formula:
[0091]
[0092] Among them, C i Represents the multi-resolution feature map of the i-th layer, with gate variable f. i g (C i ) controls the retention of the original variable and modifies the variable f i c (C i ) is used to correct errors in the original variables, f i g (·) and f i c (·) represents a 1×1 convolutional layer, F i This represents the feature map of the i-th layer of StyleGAN;
[0093] Invert style encoding of the detail layer of the source image Detail layer inversion style coding of reference image The input is fed into the first mapping network M0 to calculate the style encoding increment of the inverted makeup representation, as shown in the following formula:
[0094]
[0095] M0 is a mapping network. These represent the inversion style encodings of the detail layers of the source and reference images, respectively. This indicates that the style encoding of the source image x belongs to the fine layer.
[0096] For the residual map y of the reference image res Using encoder E res To distill the lost makeup information from the residual image;
[0097]
[0098] Among them, E res This represents the trained encoder pSp. The lost makeup representation obtained by distillation is then used to embed the lost makeup representation into the inversion result of the source image using the second mapping network M1.
[0099]
[0100] Where M1 is the second mapping network, The style encoding output by the first mapping network M0, This indicates the loss of makeup obtained through distillation. This indicates that a fine layer containing the source image with the reference makeup style and the inversion style code has been embedded. Finally, this fine layer containing the source image with the reference makeup style and the inversion style code is... With mixed feature maps The input is fed into the StyleGAN2 generator G0 to generate the final transfer result. This framework benefits from a decoupled injection strategy that modifies feature maps at the coarse layer and edits style codes at the detail layer, thus inheriting the content protection advantages of image transformation frameworks and the texture editing and high-resolution synthesis advantages of the StyleGAN framework.
[0101] This invention designs a hybrid framework, CPTR, which inherits the advantages of content protection and color accuracy from image transformation frameworks, and the advantages of texture transfer and high-resolution support based on the StyleGAN framework. The properties of different frameworks are shown in Table 1.
[0102] Table 1 Attribute Tables for Different Frames
[0103]
[0104] Table 2 shows the quantitative results of FID (GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium.), LPIPS (The Unreasonable Effectiveness of Deep Features as a Perceptual Metric), and perceptual distance on the Makeup Transfer (MT) and High-Resolution Makeup Transfer (HRMT) datasets.
[0105]
[0106] This invention addresses the issues of facial structure deformation and background changes in the transfer results of the StyleGAN framework by proposing a multi-scale residual inversion MSRI module, which utilizes multi-scale spatial variables and vector encoding to jointly invert the image and better preserve the image content.
[0107] This invention addresses the color difference problem in the transfer results of the StyleGAN framework by proposing a color correction (CR) module, which uses residual learning to repair degraded makeup transfer.
[0108] Based on the semantic decoupling property of StyleGAN, this invention designs a layered decoupling injection strategy, which restores the source content in the coarse layer and injects the reference makeup style in the detail layer.
[0109] In CPTR, this invention transfers makeup style by predicting the increment of vector variables. Therefore, the intensity of makeup can be controlled by interpolating the increment. Figure 11 As shown.
[0110] exist Figure 12 In this invention, the robustness of CPTR under other complex conditions was also tested. Vector encoding discards spatial information, therefore CPTR is robust to spatial misalignment caused by pose. Regarding illumination, CPTR transfers illumination from the reference image to the source image, thus producing a natural output. For facial occlusion, CPTR can effectively fill in the covered makeup information and generate realistic results.
[0111] A method for facial makeup transfer based on a hybrid framework of image transformation and StyleGAN is characterized by comprising the following steps;
[0112] Step 1: The image inversion module uses the inversion encoder to calculate the inversion style code of the source image and the reference image respectively, and uses the StyleGAN2 generator to generate the inversion images of the source image and the reference image. The inversion images are compressed to obtain the residual map between the source image and the inversion compressed image of the source image and the residual map between the reference image and the inversion compressed image of the reference image.
[0113] Step 2: The feature map acquisition module inputs the residual map of the source image and the inverted compressed image of the source image into the multi-scale residual encoder to obtain multi-resolution feature maps, and injects the multi-resolution feature maps into the feature maps of the StyleGAN2 coarse layer based on the fusion network to obtain the hybrid feature map.
[0114] Step 3: The first-stage training module uses a loss function to train a multi-scale residual encoder and a fusion network on the CelebA-HQ dataset;
[0115] Step 4: The preliminary transfer result acquisition module uses the detail layer of the source image and the reference image to invert the style code and obtain the change in style code through the first mapping network. The changes in the mixed feature map and style code are then input into the StyleGAN2 generator to generate the preliminary transfer result.
[0116] Step 5: The second-stage training module optimizes the first mapping network using the overall loss function;
[0117] Step 6: The color correction module distills the lost makeup information from the residual image of the reference image and the inverted compressed image of the reference image, and uses the second mapping network to embed the lost makeup information into the inverted image of the source image, thus obtaining a fine layer of the source image inversion style encoding that has been embedded with the reference makeup style. This fine layer and the mixed feature map are input into the StyleGAN2 generator to generate the final transfer result.
[0118] Step 7: The third-stage training module inputs the final transfer result into the overall loss function to obtain the overall loss information, and uses backpropagation to optimize the residual encoder and the second mapping network using the overall loss information;
[0119] Step 8: The makeup transfer module combines the image inversion module, feature map acquisition module, preliminary transfer result acquisition module, and color correction module into a hybrid framework based on image transformation and StyleGAN. Then, the source image and reference image are input into this hybrid framework to obtain the makeup transfer result.
[0120] In CPTR, this invention proposes Multi-Scale Residual Inversion (MSRI) and Color Refinement (CR) modules to address content distortion and color degradation issues during inversion. We conducted ablation studies on these two modules to verify their effectiveness.
[0121] The MSRI module integrates multi-scale spatial variables into a vector-controlled StyleGAN for high-fidelity inversion of source images. For example... Figure 9 As shown, compared to the basic encoder e4e (Designing an encoder for StyleGAN image manipulation), MSRI has higher accuracy in reconstructing images, especially in terms of facial structure.
[0122] CPTR designed a Style Modulation (SM) module to inject the reference-inverted makeup into the output, and added a Color Refinement (CR) module to correct color deviations. For example... Figure 9As shown, using SM, the makeup style of the reference inversion image is effectively propagated to the source image. However, due to color degradation during the reference inversion process, the transferred result (MSRI+SM) differs slightly from the original makeup style. After equipping the CR module, the missing color information was successfully repaired, and the makeup similarity was further enhanced.
[0123] The contents not described in detail in this specification are existing technologies known to those skilled in the art.
Claims
1. A facial makeup transfer system based on a hybrid framework of image transformation and StyleGAN, characterized in that, It includes an image inversion module, a feature map acquisition module, a first-stage training module, a preliminary transfer result acquisition module, a second-stage training module, a color correction module, a third-stage training module, and a makeup transfer module; The image inversion module is used to calculate the inversion style codes of the source image and the reference image respectively using the inversion encoder, and to generate the inversion images of the source image and the reference image using the StyleGAN2 generator. The inversion images are compressed to obtain the residual map between the source image and the inversion compressed image of the source image and the residual map between the reference image and the inversion compressed image of the reference image. The feature map acquisition module is used to input the residual map of the source image and the inverted compressed image of the source image into the multi-scale residual encoder to obtain multi-resolution feature maps, and inject the multi-resolution feature maps into the feature map of the StyleGAN2 coarse layer based on the fusion network to obtain the hybrid feature map. The first-stage training module is used to train a multi-scale residual encoder and a fusion network on the CelebA-HQ dataset using a loss function; The preliminary transfer result acquisition module is used to invert the style code using the detail layer of the source image and the reference image, obtain the change in style code through the first mapping network, and input the change in the mixed feature map and style code into the StyleGAN2 generator to generate the preliminary transfer result. The second-stage training module is used to input the initial transfer results into the overall loss function to obtain the overall loss information, and then optimize the first mapping network using backpropagation based on the overall loss information. The color correction module is used to distill the lost makeup information from the residual map of the reference image and the inverted compressed image of the reference image using a residual encoder, and to embed the lost makeup information into the inverted image of the source image using a second mapping network, so as to obtain the source image fine-layer inversion style code embedded with the reference makeup style. The fine-layer style code and the mixed feature map are input into the StyleGAN2 generator to generate the final transfer result. The third-stage training module is used to input the final transfer result into the overall loss function to obtain the overall loss information, and then use backpropagation to optimize the residual encoder and the second mapping network using the overall loss information. The makeup transfer module is used to construct a hybrid framework based on image transformation and StyleGAN by combining the image inversion module, feature map acquisition module, preliminary transfer result acquisition module and color correction module. Then, the source image and reference image are input into this hybrid framework to obtain the makeup transfer result.
2. The face makeup transfer system based on a hybrid framework of image transformation and StyleGAN as described in claim 1, characterized in that: The specific method by which the image inversion module calculates the inversion style codes of the source image x and the reference image y using the inversion encoder E0 is as follows: In x =E0(x), In y =E0(x), Among them, w x For the inversion style encoding of the source image x, w y For the inversion style code of the reference image y, E0(x) represents the style code output from the inversion encoder E0 when the source image x is input into it. Each represents 18 cascaded components forming w x Individual style coding, Each represents 18 cascaded components forming w y Individual style coding; Inversion images are generated using the StyleGAN2 generator G0, which is pre-trained on FFHQ data, based on the inversion style encoding of the source and reference images. x0=G0(w x ), y0=G0(w y ), where x0 is the first inversion result of the source image, y0 is the first inversion result of the reference image, and G0(w x ) represents the inversion style encoding w of the source image x. x The input is fed into the StyleGAN2 generator G0, G0(w y ) indicates that the inversion style encoding w of the reference image y will be performed. y Input into the StyleGAN2 generator G0; Bilinear downsampling is used to compress the inverted image, and residual maps of the source image and the inverted compressed image, and residual maps of the reference image and the inverted compressed image are obtained. res =x(i)-x 01 (i), y res =y(i)-y 01 (i), where x res The residual image represents the difference between the source image and the inverted compressed image. res This represents the residual image between the reference image and the inverted compressed image of the reference image, where x(i) represents the pixel value of each pixel in the source image. 01 (i) represents the pixel value of each pixel in the inverted compressed image of the source image, and y(i) represents the pixel value of each pixel in the reference image. 01 (i) represents the pixel value of each pixel in the inverted compressed image of the reference image.
3. The face makeup transfer system based on a hybrid framework of image transformation and StyleGAN as described in claim 2, characterized in that: The feature map acquisition module is used to obtain the residual map x between the source image and the inverted compressed image. res Input to multi-scale residual encoder E c Obtaining multi-resolution feature maps C=E c (x res Furthermore, a multi-resolution feature map C is injected into the feature map of the StyleGAN2 coarse layer through a fusion network based on a gating mechanism, resulting in a hybrid feature map. The multi-resolution feature map C has the same resolution as the feature map of the StyleGAN2 coarse layer; Among them, C i Represents the multi-resolution feature map of the i-th layer, with gate variable f. i g (C i ) controls the retention of the original variable and modifies the variable f i c (C i ) is used to correct errors in the original variables, f i g (·) and f i c (·) represents a 1×1 convolutional layer, F i This represents the feature map of the i-th layer of StyleGAN.
4. The face makeup transfer system based on a hybrid framework of image transformation and StyleGAN as described in claim 3, characterized in that: The first-stage training module is used to train a multi-scale residual encoder and a fusion network on the CelebA-HQ dataset using a loss function that enables image inversion and reconstruction.
5. The face makeup transfer system based on a hybrid framework of image transformation and StyleGAN according to claim 4, characterized in that: The preliminary transfer result acquisition module is used to invert style encoding of the detail layer of the source image. Detail layer inversion style coding of reference image The input is fed into the first mapping network M0 to calculate the style encoding increment of the inverted makeup representation, as shown in the following formula: Where M0 is the mapping network, These represent the inversion style encodings of the detail layers of the source and reference images, respectively. This indicates that the style encoding of the source image x belongs to the fine layer. The mapping network M0 consists of a fully connected layer FC, an adaptive layer AdaIN, and an activation layer stacked together. The adaptive layer AdaIN propagates the reference makeup by modulating the mean and variance of the style encoding, thereby conditionally embedding the inverted reference makeup into the source image. Mixed feature maps Changes in style coding The input is fed into the StyleGAN2 generator G0 for progressive image generation, yielding preliminary transfer results.
6. The face makeup transfer system based on a hybrid framework of image transformation and StyleGAN according to claim 5, characterized in that: The overall loss function is: L all =λ protect L protect +λ trans L trans Among them, L all Let L be the overall loss function. protect To protect the loss function, L trans Let λ be the migration loss function. protect and λ trans Hyperparameter weights are used to balance two different training objectives: content protection and makeup transfer, utilizing the overall loss function L. all The first mapping network M0 is optimized. During the optimization process, the inversion encoder E0, the generator G0, and the multi-scale residual encoder E0, which was trained in the first stage, are also optimized. c The fusion network has fixed parameters and does not perform gradient updates.
7. The face makeup transfer system based on a hybrid framework of image transformation and StyleGAN as described in claim 5, characterized in that: Protective loss function L protect as follows: L protect =λ id L id +λ bg L bg Among them, L id Let L be the identity loss function. bg Let λ be the background loss function. id To balance the hyperparameter weights for identity protection, λ bg To balance the hyperparameter weights for the protection background, for the identity loss function L id By calculating the migration result x trans The cosine distance between the source image x and the source image x is obtained as follows: L id =1-cos(R(x),R(x trans )) Where R represents the pre-trained ArcFace network used for face recognition, and L represents the background loss. bg The calculation is as follows: L bg =||(x-x trans )·P bg (x)||2 Where ||·||2 represents the L2 norm constraint, · represents element-wise multiplication, and P represents the pre-trained face segmentation network BiSeNet. bg (x) represents the mask of the background region of the source image x based on BiSeNet prediction; L trans =L CX (y·P face (y),x trans ·P face (x trans )) Among them, L trans L represents the migration loss function. CZ P represents contextual loss. face (y) represents the face mask of the predicted reference image y.
8. The face makeup transfer system based on the image transformation and StyleGAN hybrid framework according to claim 7, characterized in that: The color correction module distills the lost makeup information from the residual image of the reference image and the inverted compressed image of the reference image, specifically as follows: Among them, E res This represents the trainable encoder pSp. This indicates the lost makeup information obtained through distillation; The lost makeup information is embedded into the inverted image of the source image using the second mapping network M1, resulting in a fine layer of the source image x inverted style encoding that has been embedded with the reference makeup style, specifically represented as: Where M1 is the second mapping network, The style encoding output by the first mapping network M0, This indicates the loss of makeup obtained through distillation. This indicates that a fine layer has been embedded with the source image of the reference makeup style and the inversion style code; The source image, which already contains the reference makeup style, is then processed by a fine layer of inversion style encoding. With mixed feature maps The input is fed into the StyleGAN2 generator G0 to generate the final transfer result.
9. The face makeup transfer system based on a hybrid framework of image transformation and StyleGAN as described in claim 8, characterized in that: The makeup transfer module integrates the image inversion module, feature map acquisition module, preliminary transfer result acquisition module, and color correction module into a hybrid framework based on image transformation and StyleGAN. The specific process of inputting the source image and reference image into this hybrid framework to obtain the makeup transfer result is as follows: Given a source image x and a reference image y, first obtain their respective inversion style codes w. x =E0(x), w y =E0(x), and the inversion result x0=G0(w x ), x0=G0(w x and residual plot x res =x - x0, y res = y - y0, where E0 is the inversion encoder, G0 is the pre-trained StyleGAN2 generator, and then, for the residual image x of the source image and the inversion compressed image of the source image... res It is input into the multi-scale residual encoder E. c To obtain multi-resolution feature maps C=E c (x res Then, it is injected into the corresponding feature map of the StyleGAN2 coarse layer with the same resolution through a base fusion network, as described in the following formula: Among them, C i Represents the multi-resolution feature map of the i-th layer, with gate variable f. i g (C i ) controls the retention of the original variable and modifies the variable f i c (C i This is used to correct errors in the original variables. i g (·) and f i c (·) represents a 1×1 convolutional layer, F i This represents the feature map of the i-th layer of StyleGAN; Invert style encoding of the detail layer of the source image Detail layer inversion style coding of reference image The input is fed into the first mapping network M0 to calculate the style encoding increment of the inverted makeup representation, as shown in the following formula: M0 is a mapping network. These represent the inversion style encodings of the detail layers of the source and reference images, respectively. This indicates that the style encoding of the source image x belongs to the fine layer. For the residual map y of the reference image res Using encoder E res To distill the lost makeup information from the residual image; Among them, E res This represents the trained encoder pSp. The lost makeup representation obtained by distillation is then used to embed the lost makeup representation into the inversion result of the source image using the second mapping network M1. Where M1 is the second mapping network, The style encoding output by the first mapping network M0, This indicates the loss of makeup obtained through distillation. This indicates that a fine layer containing the source image with the reference makeup style and the inversion style code has been embedded. Finally, this fine layer containing the source image with the reference makeup style and the inversion style code is... With mixed feature maps The input is fed into the StyleGAN2 generator G0 to generate the final transfer result.
10. A method for facial makeup transfer based on a hybrid framework of image transformation and StyleGAN, characterized in that, It includes the following steps; Step 1: The image inversion module uses the inversion encoder to calculate the inversion style code of the source image and the reference image respectively, and uses the StyleGAN2 generator to generate the inversion images of the source image and the reference image. The inversion images are compressed to obtain the residual map between the source image and the inversion compressed image of the source image and the residual map between the reference image and the inversion compressed image of the reference image. Step 2: The feature map acquisition module inputs the residual map of the source image and the inverted compressed image of the source image into the multi-scale residual encoder to obtain multi-resolution feature maps, and injects the multi-resolution feature maps into the feature maps of the StyleGAN2 coarse layer based on the fusion network to obtain the hybrid feature map. Step 3: The first-stage training module uses a loss function to train a multi-scale residual encoder and a fusion network on the CelebA-HQ dataset; Step 4: The preliminary transfer result acquisition module uses the detail layer of the source image and the reference image to invert the style code and obtain the change in style code through the first mapping network. The changes in the mixed feature map and style code are then input into the StyleGAN2 generator to generate the preliminary transfer result. Step 5: The second-stage training module optimizes the first mapping network using the overall loss function; Step 6: The color correction module distills the lost makeup information from the residual image of the reference image and the inverted compressed image of the reference image, and uses the second mapping network to embed the lost makeup information into the inverted image of the source image, thus obtaining a fine layer of the source image inversion style encoding that has been embedded with the reference makeup style. This fine layer and the mixed feature map are input into the StyleGAN2 generator to generate the final transfer result. Step 7: The third-stage training module inputs the final transfer result into the overall loss function to obtain the overall loss information, and uses backpropagation to optimize the residual encoder and the second mapping network using the overall loss information; Step 8: The makeup transfer module combines the image inversion module, feature map acquisition module, preliminary transfer result acquisition module, and color correction module into a hybrid framework based on image transformation and StyleGAN. Then, the source image and reference image are input into this hybrid framework to obtain the makeup transfer result.