Face exchange method and system based on StyleGAN2 architecture

By converting feature maps to S space and performing feature fusion in the StyleGAN2 face switching method, the problem of insufficient decoupling capability in the prior art is solved, and a more natural and high-fidelity face switching effect is achieved.

CN120219149AActive Publication Date: 2025-06-27ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510688626.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-27
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The existing face exchange method based on StyleGAN2 is insufficient in decoupling capabilities in W/W+ space, resulting in the generated face exchange results that cannot effectively maintain the expression, posture and other attribute characteristics of the target face.

Method used

By converting the feature map of W+ space to S space, the source identity features and target attribute features are separated using the higher level decoupling characteristics of S space, and fuse the identity feature extraction network and the attribute feature extraction network for fusion. Enter the StyleGAN2 generator to generate a face exchange image, and smoothly fuse it with the target image through soft mask fusion technology.

Benefits of technology

It significantly improves the naturalness and detail retention of generated images, avoids artifacts caused by identity and attribute conflicts, and ensures high fidelity and efficiency of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219149A_ABST
    Figure CN120219149A_ABST
Patent Text Reader

Abstract

The invention discloses a human face exchange method and system based on a StyleGAN2 architecture, and belongs to the technical field of human face exchange, and the method comprises the steps: achieving the decoupling of a source identity and a target attribute through feature space conversion, respectively extracting the identity and attribute features through an independent network, and fusing the features through a dynamic balance mechanism to eliminate artifacts; an unmodified StyleGAN2 generator is adopted to process fusion features, seamless fusion of a human face and a background is realized in combination with a soft mask technology, and a generation process is simplified; based on a multi-level detection mechanism, the face alignment precision is improved, and the adaptability to complex scenes is enhanced. According to the method, the problems of identity attribute conflict, artifact generation and insufficient scene robustness are systematically solved, and efficient and high-fidelity face exchange is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of face swapping, and particularly relates to a face swapping method and system based on the StyleGAN2 architecture. Background Art

[0002] The currently more common face swapping methods are based on generative adversarial networks (GANs). Such methods usually use a generator network to implement the conversion between the face feature maps and face images in the source image and the target image, and use a discriminator network to determine whether the input image is real or generated and give corresponding probability values. Then, by optimizing the adversarial loss function between the generator network and the discriminator network, a high-quality and realistic face swapping effect is achieved. For example, in the design solution with the patent application number CN202410010020.0, the ID information of the face and several feature maps are input into the MobileSwap model to obtain the generated face. The generated face image is combined with the transparency map to obtain the final face swapping image. Such methods can achieve a good face swapping effect, but there are also some disadvantages. For example, they rely more on cumbersome network and loss designs, still struggle with the information balance between the source face and the target face, and often produce visible artifacts.

[0003] In addition, face swapping methods based on the StyleGAN2 latent space have been proposed. Such methods often first encode the image into the editable latent space of StyleGAN2 through GAN inversion and perform operations in the latent space to achieve the purpose of face swapping. For example, in the design solution StyleSwap of ECCV 2022, a very small modification is made to the StyleGAN2 generator so that it can successfully process the required information from the source and the target, and then perform face swapping. Such methods rely on the inherent prior knowledge of the pre-trained GAN model and rely on the powerful generation ability of StyleGAN2 to generate high-resolution and high-fidelity face swapping results. However, such methods are limited by the decoupling ability of the latent encoding in the W / W+ space, and the generated face swapping results often cannot well maintain the attribute features such as the expression and pose of the target face. Therefore, a new face swapping method is urgently needed. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a face swapping method and system based on the StyleGAN2 architecture to solve the problems existing in the above prior art.

[0005] In a first aspect, to achieve the above object, the present invention provides a face swapping method based on the StyleGAN2 architecture, including the following steps:

[0006] Performing face detection and alignment on the source image and the target image through a face detection algorithm;

[0007] Extract the feature maps of the face information of the source image and the target image in the W+ space, and transform the feature maps to the S space through an affine transformation;

[0008] Extract the identity features of the source image and the attribute features of the target image in the S space respectively, and fuse the identity features and the attribute features;

[0009] Input the fused features into the StyleGAN2 generator to generate a face-swapped image;

[0010] Smoothly fuse the generated face-swapped image with the target image.

[0011] Optionally, the process of face detection and alignment includes:

[0012] Construct an image pyramid for the source image and the target image through the MTCNN algorithm, and use the P-Net to generate candidate face windows and calibrate the bounding boxes;

[0013] Further refine the candidate windows through the R-Net and reject the false candidates;

[0014] Output the final bounding box and the facial landmark positions through the O-Net.

[0015] Optionally, the process of transforming the feature maps to the S space includes:

[0016] Extract the feature maps of the W+ space through a pre-trained pSp encoder;

[0017] Use an affine transformation to transform the feature maps of the W+ space to the S space to decouple the identity features and the attribute features.

[0018] Optionally, the process of extracting the identity features and the attribute features includes:

[0019] Extract the identity information from the source feature maps in the S space through an identity feature extraction network;

[0020] Extract the attribute information from the target feature maps in the S space through an attribute feature extraction network;

[0021] The identity feature extraction network and the attribute feature extraction network share the backbone structure but have independent parameters.

[0022] Optionally, the process of the fusion includes:

[0023] Dynamically balance the contributions of the identity features and the attribute features in each layer in the S space through a weight matrix;

[0024] Weight the identity feature and the attribute feature according to the weight matrix to generate the fused S-space feature map.

[0025] Optionally, the process of smooth fusion includes:

[0026] Perform element-wise weighted fusion of the generated face-swapped image and the target image through a soft mask;

[0027] The soft mask is extended by channels to match the RGB channels of the target image.

[0028] In a second aspect, the present invention also provides a face-swapping system based on the StyleGAN2 architecture for implementing a face-swapping method based on the StyleGAN2 architecture. The system includes:

[0029] A face processing module for performing face detection and alignment on the source image and the target image;

[0030] A feature extraction module for extracting the feature map in the W+ space from the source image and the target image and converting the feature map to the S space;

[0031] A feature fusion module for extracting the identity feature of the source image and the attribute feature of the target image in the S space and fusing the two;

[0032] An image generation module for generating a face-swapped image based on the StyleGAN2 generator according to the fused features;

[0033] An image fusion module for smoothly fusing the generated face-swapped image with the target image.

[0034] Optionally, the image fusion module includes:

[0035] A soft mask generation unit for generating a soft mask for fusion;

[0036] A weighted fusion unit for performing element-wise weighted fusion of the generated face-swapped image and the target image through a soft mask;

[0037] A channel matching unit for extending the channels of the soft mask to match the RGB channels of the target image to achieve smooth fusion.

[0038] In a third aspect, the present invention also provides a computer terminal device, including:

[0039] One or more processors;

[0040] A memory coupled to the processor for storing one or more programs;

[0041] When the one or more programs are executed by the one or more processors, the one or more processors implement a face swapping method based on the StyleGAN2 architecture.

[0042] In a fourth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, a face swapping method based on the StyleGAN2 architecture is implemented.

[0043] Compared with the prior art, the present invention has the following advantages and technical effects:

[0044] The present invention provides a face swapping method and system based on the StyleGAN2 architecture. By converting the feature mapping in the W+ space to the S space through affine transformation and utilizing its hierarchical decoupling characteristics to accurately separate the source identity features and target attribute features, the present invention solves the problem of attribute loss caused by insufficient decoupling ability in the W / W+ space in traditional methods, and significantly improves the naturalness of the generated images; by separately extracting identity information and attribute information from the S space through an identity feature extraction network and an attribute feature extraction network with independent parameters, and dynamically balancing the contributions of features in each layer by combining a weight matrix, the present invention effectively avoids artifacts such as facial blurring or edge breaks caused by identity-attribute conflicts, and ensures high-fidelity details of the generated images; by directly using the unmodified StyleGAN2 generator to process the fused S space features, the present invention avoids the limitations of existing solutions that require adjusting the generator structure. At the same time, a soft mask fusion technology is adopted to achieve seamless connection between the generated face and the target background, reducing the dependence on complex post-processing algorithms, significantly reducing the implementation complexity and improving the processing efficiency; a multi-level detection mechanism based on MTCNN is used to achieve high-precision face alignment, combined with the adaptability of the S space to complex scenarios such as occlusion and extreme lighting, so that the method can still output stable and high-quality face swapping results under non-ideal conditions, expanding the coverage of practical application scenarios. Through the systematic improvement of the above technical means, the present invention simplifies the process and reduces the structural dependence while overcoming the problems of difficult balance between identity and attribute, artifact generation, and insufficient robustness in complex scenarios, providing a solution with both high efficiency and reliability for high-precision face swapping. Description of the Drawings

[0045] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0046] Figure 1 It is a flowchart of the face swapping method based on the StyleGAN2 architecture in the specific implementation manner of Embodiment 1 of the present invention;

[0047] Figure 2Schematic diagram of the feature mapping extraction structure according to an embodiment of the present invention;

[0048] Figure 3 Schematic diagram of the conversion from the W+ space to the S space according to an embodiment of the present invention;

[0049] Figure 4 Schematic diagram of the identity feature extraction network according to an embodiment of the present invention;

[0050] Figure 5 Schematic diagram of the attribute feature extraction network according to an embodiment of the present invention;

[0051] Figure 6 Image of the face after swapping generated by the StyleGAN2 generator according to an embodiment of the present invention;

[0052] Figure 7 Schematic diagram of the process according to an embodiment of the present invention. In the figure, is the identity loss. ArcFace is a pre-trained state-of-the-art face recognition model, which is used as an identity encoder to extract identity features. is the feature matching loss, represents the output of the m-th layer of the discriminator, is the layer where the feature matching loss starts to be calculated, is the total number of layers. Detailed implementation manners

[0053] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0054] It should be noted that the steps shown in the flowchart of the accompanying drawings may be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0055] First, the technical terms involved in the following embodiments will be described below.

[0056] The present invention uses the powerful generation ability of StyleGAN2 for face swapping to ensure the high resolution and high fidelity of the generated images.

[0057] The present invention adopts the decoupling characteristics at a higher level in the S space to more effectively separate the source identity from the target attributes and enhance the flexibility of the latent encoding. The design mechanism effectively balances the transfer of the source identity and the preservation of the target face features (such as expressions and postures), avoiding information loss and artifact generation.

[0058] The present invention reduces the dependence on complex network architectures and loss functions, simplifies the implementation process, and improves the application convenience. The method exhibits better robustness under various conditions and can handle different types of source and target images.

[0059] The key points of the present invention include:

[0060] 1. A technical method for decoupling w latent encoding using the S space to ensure the effective separation of source identity and target attributes.

[0061] 2. Reducing the dependence on complex network architectures and loss functions, simplifying the implementation process, and improving the application convenience.

[0062] 3. Steps and methods for extracting and fusing features through an identity feature extraction network and an attribute feature extraction network.

[0063] Embodiment 1

[0064] As Figure 7 shown, in this embodiment, a face swapping method based on the StyleGAN2 architecture is provided, including:

[0065] Performing face detection and alignment on the source image and the target image through a face detection algorithm;

[0066] Extracting the feature maps of the face information of the source image and the target image in the W+ space, and converting the feature maps to the S space through an affine transformation;

[0067] Extracting the identity features of the source image and the attribute features of the target image in the S space respectively, and fusing the identity features and the attribute features;

[0068] Inputting the fused features into the StyleGAN2 generator to generate a face swapping image;

[0069] Smoothing and fusing the generated face swapping image with the target image.

[0070] Specifically, based on the goal of retaining the source image identity information and the target image attribute information, the process of the method is as Figure 1 shown, mainly having four steps:

[0071] Obtaining the source image and the target image to be swapped, and performing face detection and alignment on the source image and the target image through the face detection algorithm MTCNN.

[0072] Extracting the feature maps of the source image face information and the target image face information in the W+ space through a pre-trained pSp encoder, and then converting the feature maps in the W+ space to the S space through an affine transformation A to achieve better information decoupling.

[0073] Extract the identity information of the source image in the S space through the identity feature extraction network, extract the attribute information of the target image in the S space through the attribute information extraction network, and then fuse the two.

[0074] Send the fusion information in the S space obtained in the previous step into the StyleGAN2 generator to generate the image after face swapping.

[0075] As an implementation manner in this embodiment, the process of face detection and alignment includes:

[0076] Construct an image pyramid for the source image and the target image through the MTCNN algorithm, and use the P-Net to generate candidate face windows and calibrate the bounding boxes;

[0077] Further refine the candidate windows through the R-Net and reject the false candidates;

[0078] Output the final bounding box and the facial landmark positions through the O-Net.

[0079] Specifically, the process of extracting the source image and the target image

[0080] When performing face swapping, it is necessary to obtain the source image and the target image first. However, the source image and the target image often contain not only face information, but also background information and other information. To achieve a better face swapping effect, the corresponding face information should be extracted from the source image and the target image first to reduce the interference of non-face regions during the face swapping process. Therefore, the present invention performs face detection and alignment on the source image and the target image through the face detection algorithm MTCNN.

[0081] The steps are as follows:

[0082] Step 1-1: As Figure 2 shown, given the input image, initially adjust it to different scales to construct an image pyramid. Then use a fully convolutional network, called the proposal network (P-Net), to obtain candidate face windows and their bounding box regression vectors. Then calibrate the candidate data according to the estimated bounding box regression vectors. After that, use non-maximum suppression (NMS) to merge the highly overlapping candidate data.

[0083] Step 1-2: All candidates are fed into another CNN, called the refinement network (R-Net). Then refine these candidates through the refinement network (R-Net), which further rejects a large number of false candidates, performs calibration using bounding box regression, and performs NMS.

[0084] Step 1-3: Identify the face region through more supervision, and the output network (O-Net) generates the final bounding box and the facial landmark positions.

[0085] As an implementation manner in this embodiment, the process of converting the feature map to the S space includes:

[0086] Extract the feature map in the W+ space through the pre-trained pSp encoder;

[0087] Use an affine transformation to convert the feature map in the W+ space to the S space to decouple the identity feature and the attribute feature.

[0088] Specifically, the source image and the target image The process of decoupling the feature information includes:

[0089] Because face swapping requires the source image to provide the identity feature and the target image to provide the attribute feature, after face detection and alignment of the source image and the target image, it is necessary to extract the identity feature of the source image and the attribute feature of the target image. Before feature extraction, it is necessary to decouple the feature information of the source image and the target image. The steps are as follows:

[0090] Step 2-1: Extract the face information of the source image and the face information of the target image The feature map in the W+ space , , first use a standard feature pyramid on the ResNet backbone to extract the feature map. For each of the 18 target styles, train a small mapping network to extract the learned style from the corresponding feature map, where the styles (0-2) are generated from the small feature map, (3-6) are from the medium feature map, and (7-18) are from the largest feature map. The mapping network map2style is a small fully convolutional network that gradually reduces the spatial size using a set of 2-stride convolutions and LeakyReLU activations. The structural framework is as Figure 2 shown.

[0091] The "style" mentioned above generally refers to the feature vectors injected into different layers (different resolutions) of the generator. These style vectors will respectively affect different scales of image generation, and the overall shape to local texture is controlled by them. For example, the early style layers (injected into the low-resolution layers) often determine the "coarse" features such as the overall contour, large color blocks, and composition of the image; while the later style layers (injected into the high-resolution layers) more determine the "fine" features such as local texture, details, and colors.

[0092] Step 2-2: The decoupling ability of the latent encoding in the W+ space is limited, and the generated face swapping results often cannot well preserve the attribute features such as the expression and pose of the target face. Therefore, as Figure 3 shown, the present invention maps the feature mapping in the W+ space to the S space through the affine transformation A matching with , to obtain , , so as to achieve better information decoupling, and thus better extract the identity features of the source image and the attribute features of the target image.

[0093] As an implementation manner in this embodiment, the process of extracting the identity features and the attribute features includes:

[0094] Extracting identity information from the source feature mapping in the S space through an identity feature extraction network;

[0095] Extracting attribute information from the target feature mapping in the S space through an attribute feature extraction network;

[0096] The identity feature extraction network and the attribute feature extraction network share the backbone structure but have independent parameters.

[0097] Specifically, after information decoupling, the identity features of the source image in the S space are extracted through the identity feature extraction network, and the attribute features of the target image in the S space are extracted through the attribute feature extraction network, and then the two are fused.

[0098] Step 3-1: Input the feature mapping of the source image face information in the S space into the identity feature extraction network to obtain the identity feature mapping of the source image face information in the S space , as Figure 4 shown.

[0099] Step 3-2: Input the feature mapping of the target image face information in the S space into the attribute feature extraction network to obtain the attribute feature mapping of the target image face information in the S space , as Figure 5 shown. The identity feature extraction network and the attribute feature extraction network share a set of network structures with different parameters.

[0100] Step 3-3: Fuse the identity feature mapping and the attribute feature mapping to obtain the fusion information , as shown in Equation (1). Where M is the weight parameter matrix and 1 is the all-1 matrix. Because the identity feature mapping and the attribute feature mapping Both are 18 layers, so M is used to balance the contributions of the identity information of the source images and the attribute information of the target images in each layer of the S space.

[0101]

[0102] As an implementation manner in this embodiment, the fusion process includes:

[0103] Dynamically balance the contributions of the identity features and the attribute features in each layer of the S space through a weight matrix;

[0104] Perform weighted summation on the identity features and the attribute features according to the weight matrix to generate a fused S space feature map.

[0105] Specifically, send the fusion information into the StyleGAN2 generator to generate an image after face swapping. Since the image after face swapping only includes face information, it needs to be smoothly fused with the target image smoothly.

[0106] Step 4-1: Send the fusion information into each layer of the StyleGAN2 generator to generate an image after face swapping. Since the fusion information is already in the S space, the StyleGAN2 generator without the affine transformation A is used here, as Figure 6 shown.

[0107] As an implementation manner in this embodiment, the smooth fusion process includes:

[0108] Perform element-wise weighted fusion of the generated face-swapped image and the target image through a soft mask;

[0109] The soft mask is extended by channels to match the RGB channels of the target image.

[0110] Specifically, Step 4-2: Since the image after face swapping only includes face information, it needs to be smoothly fused with the original target image to obtain the final image . The present invention uses a soft mask Mask for fusion, as shown in Equation (2). Where Mask is a soft mask with each value between 0 and 1, where * represents element-wise multiplication with broadcasting, and 1 is a tensor of all 1s. The mask repeats itself three times by channels to match the RGB channels.

[0111]

[0112] Based on this, a face swapping method based on the StyleGAN2 architecture provided by an embodiment of the present invention utilizes the higher-level decoupling characteristics of the S space to more effectively separate the source identity and the target attributes, enhancing the flexible operation ability of the latent encoding. It achieves a more natural face swapping effect while maintaining the characteristics of the source face and the target face.

[0113] The present invention effectively balances the relationship between source identity transfer and target attribute preservation, avoiding the common information loss and artifact problems in traditional methods. Through the improved encoding decoupling method, the loss of latent structural information is reduced, and the quality and detail retention of the generated images are improved.

[0114] The present invention shows better robustness under various conditions and can handle different types of source and target images. Compared with the dependence on complex networks and loss functions in traditional methods, the new architecture can reduce the design complexity.

[0115] Embodiment 2

[0116] In this embodiment, a computer terminal device is provided, including:

[0117] One or more processors;

[0118] A memory coupled to the processor for storing one or more programs;

[0119] When the one or more programs are executed by the one or more processors, the one or more processors implement the methods in the above embodiments.

[0120] In this embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the methods in the above embodiments are implemented.

[0121] In this embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the methods in the above embodiments.

[0122] The above program can run in a processor or can also be stored in a memory (or referred to as a computer-readable medium). A computer-readable medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0123] These computer programs can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate computer-implemented processing. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps for the functions specified in one process Figure 1 one process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one box or multiple boxes. Corresponding different steps can be implemented by different modules.

[0124] This embodiment provides such a device or system. The system is called a face swapping system based on the StyleGAN2 architecture and includes:

[0125] A face processing module for performing face detection and alignment on a source image and a target image;

[0126] A feature extraction module for extracting feature maps in the W+ space from the source image and the target image and converting the feature maps to the S space;

[0127] A feature fusion module for extracting the identity feature of the source image and the attribute feature of the target image in the S space and fusing the two;

[0128] An image generation module for generating a face swapping image based on the StyleGAN2 generator according to the fused features;

[0129] An image fusion module for smoothly fusing the generated face swapping image with the target image.

[0130] As an implementation manner in this embodiment, the face processing module includes:

[0131] An image pyramid construction unit for constructing image pyramids for the source image and the target image through the MTCNN algorithm;

[0132] A candidate window generation unit for generating candidate face windows by using P-Net and calibrating the bounding boxes;

[0133] A candidate window refinement unit for further refining the candidate windows through R-Net and excluding false candidates;

[0134] A bounding box detection unit for outputting the final bounding box and the facial landmark positions through O-Net.

[0135] As an implementation manner in this embodiment, the feature extraction module includes:

[0136] A W+ spatial feature extraction unit for respectively extracting the feature maps in the W+ space from the source image and the target image by using a pre-trained pSp encoder;

[0137] An affine transformation unit for converting the feature maps in the W+ space to the S space by using affine transformation to decouple the identity features and the attribute features.

[0138] As an implementation manner in this embodiment, the feature fusion module includes:

[0139] An identity feature extraction unit for extracting identity information from the source image feature maps in the S space;

[0140] An attribute feature extraction unit for extracting attribute information from the target image feature maps in the S space;

[0141] A dynamic weighted fusion unit for performing weighted summation on the identity features and the attribute features of each layer in the S space according to a preset weight matrix to generate a fused feature map in the S space.

[0142] As an implementation manner in this embodiment, the image generation module includes:

[0143] A feature input unit for receiving the fused feature map in the S space;

[0144] A StyleGAN2 generation unit for generating a face-swapped image based on the StyleGAN2 generator.

[0145] As an implementation manner in this embodiment, the image fusion module includes:

[0146] A soft mask generation unit for generating a soft mask for fusion;

[0147] A weighted fusion unit for performing element-wise weighted fusion on the generated face-swapped image and the target image through the soft mask;

[0148] A channel matching unit for expanding the channels of the soft mask to match the RGB channels of the target image to achieve smooth fusion.

[0149] The system or device is used to implement the functions of the method in the above embodiments. Each module in the system or device corresponds to each step in the method, and those that have been described in the method will not be repeated here.

[0150] Through the above implementation manner, the problem of face swapping based on the StyleGAN2 architecture in the related art is solved. The present invention uses the powerful generation ability of StyleGAN2 for face swapping to ensure the high resolution and high fidelity of the generated images.

[0151] The present invention adopts the decoupling characteristics at a higher level in the S space to more effectively separate the source identity from the target attributes and enhance the flexibility of the latent encoding. The design mechanism effectively balances the transfer of the source identity and the preservation of the target face features (such as expressions and postures), avoiding information loss and artifact generation.

[0152] The present invention reduces the dependence on complex network architectures and loss functions, simplifies the implementation process, and improves the application convenience. The method shows better robustness under various conditions and can process different types of source and target images.

[0153] The above is only a preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A face swapping method based on the StyleGAN2 architecture, characterized in that, The steps include: Performing face detection and alignment on the source image and the target image through a face detection algorithm; Extracting the feature maps of the face information of the source image and the target image in the W+ space, and transforming the feature maps to the S space through an affine transformation; Extracting the identity feature of the source image and the attribute feature of the target image in the S space respectively, and fusing the identity feature and the attribute feature; Inputting the fused feature into the StyleGAN2 generator to generate a face-swapped image; Smoothly fusing the generated face-swapped image with the target image.

2. The method according to claim 1, characterized in that, The process of the face detection and alignment includes: Constructing an image pyramid for the source image and the target image through the MTCNN algorithm, and using the P-Net to generate candidate face windows and calibrate the bounding boxes; Further refining the candidate windows and rejecting the false candidates through the R-Net; Outputting the final bounding box and the facial landmark positions through the O-Net.

3. The method according to claim 1, characterized in that, The process of transforming the feature maps to the S space includes: Extracting the feature maps in the W+ space through a pre-trained pSp encoder; Using an affine transformation to transform the feature maps in the W+ space to the S space to decouple the identity feature and the attribute feature.

4. The method according to claim 1, wherein The process of extracting the identity feature and the attribute feature includes: Extracting the identity information from the source feature maps in the S space through an identity feature extraction network; Extracting the attribute information from the target feature maps in the S space through an attribute feature extraction network; The identity feature extraction network and the attribute feature extraction network share the backbone structure but have independent parameters.

5. The method according to claim 1, wherein The process of the fusion includes: Dynamically balancing the contributions of the identity feature and the attribute feature in each layer in the S space through a weight matrix; Performing weighted summation on the identity feature and the attribute feature according to the weight matrix to generate the fused feature maps in the S space.

6. The method according to claim 1, wherein The process of the smooth fusion includes: Performing element-wise weighted fusion on the generated face-swapped image and the target image through a soft mask; The soft mask is extended by channels to match the RGB channels of the target image.

7. A face swapping system based on the StyleGAN2 architecture, characterized in that, The system includes: A face processing module for performing face detection and alignment on the source image and the target image; A feature extraction module for extracting the feature maps in the W+ space from the source image and the target image, and transforming the feature maps to the S space; A feature fusion module for extracting the identity feature of the source image and the attribute feature of the target image in the S space, and fusing the two; An image generation module for generating a face-swapped image based on the StyleGAN2 generator according to the fused feature; An image fusion module for smoothly fusing the generated face-swapped image with the target image.

8. The system according to claim 7, wherein The image fusion module includes: A soft mask generation unit for generating a soft mask for fusion; A weighted fusion unit for performing element-wise weighted fusion on the generated face-swapped image and the target image through a soft mask; A channel matching unit for extending the channels of the soft mask to match the RGB channels of the target image to achieve smooth fusion.

9. A computer terminal device, characterized in that, It includes: One or more processors; A memory coupled to the processor for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the face swapping method based on the StyleGAN2 architecture as described in any one of claims 1-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the face swapping method based on the StyleGAN2 architecture as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Panoramic stitching seam smoothing method and panoramic stitching seam smoothing device

    CN105678719A

  • Face changing method and device, electronic equipment and storage medium

    CN112734634A

  • Face image fusion method and device, storage medium and electronic equipment

    CN113361387A

  • Face exchange method based on identity information response

    CN115527258A

  • Face exchange method and device and terminal equipment

    CN118609179A