Image Processing Method, Apparatus, Computer Device, Storage Medium and Program Product

By acquiring the identity characteristics of the source image and the initial attribute characteristics of the target image, and using the control mask technology of the generator and convolutional layer, the problem of large images in the existing face swap technology is solved, and the face swap effect with high definition and high accuracy is achieved.

CN114972016BActive Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210626467.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-07-08
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

There is a big difference between the images obtained by the existing face-changing technology and the ideal face-changing image, resulting in poor face-changing effect.

Method used

By acquiring the identity characteristics of the source image and the initial attribute characteristics of the target image, the generator in the trained face swap module is used to combine the convolution layer and the control mask to accurately locate and filter the features outside the identity characteristics of the target face in the target image to generate a high-definition face swap image.

Benefits of technology

It achieves a high-definition face change effect, improves the clarity and accuracy of the face in the face change image, and ensures that the attributes and detailed characteristics of the target face are effectively preserved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972016B_ABST
    Figure CN114972016B_ABST
Patent Text Reader

Abstract

The present application provides an image processing method, apparatus, computer device, storage medium, and program product, which relate to technical fields such as artificial intelligence, machine learning, and intelligent transportation. By inputting the identity features of a source image into a generator in a trained face swapping module, and respectively inputting the initial attribute features of at least one scale of a target image into the convolutional layers corresponding to the respective scales in the generator, a target face-swapped image is obtained; and in each convolutional layer of the generator, a second feature map can be generated based on the identity features and the first feature map output by the previous convolutional layer; and based on the second feature map and the initial attribute features, a control mask corresponding to the target image at the corresponding scale is determined to accurately locate the pixel points of the features other than the identity features of the target face; by screening out the target attribute features based on the control mask, generating a third feature map based on the target attribute features and the second feature map, and outputting it to the next convolutional layer, the accuracy of face swapping is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of artificial intelligence, machine learning, intelligent transportation, etc. This application relates to an image processing method, apparatus, computer device, storage medium, and program product. Background Art

[0002] Face swapping is an important technology in the field of computer vision. Face swapping is widely used in scenarios such as content production, film and television portrait production, entertainment video production, virtual avatars, or privacy protection. Face swapping refers to replacing the face of an object in an image with another face.

[0003] In related technologies, neural network models are usually used to implement face swapping. For example, an image is input into a neural network model for face swapping, and the neural network model outputs an image obtained by swapping the face of the input image. However, there is a large difference between the image obtained by the existing face swapping technology and the ideal face-swapped image, resulting in a poor face swapping effect. Summary of the Invention

[0004] This application provides an image processing method, apparatus, computer device, storage medium, and program product, which can solve the problem of poor face swapping effect in related technologies. The technical solutions are as follows:

[0005] On the one hand, an image processing method is provided. The method includes:

[0006] In response to a received face swapping request, obtain the identity feature of the source image and at least one scale of initial attribute features of the target image;

[0007] The face swapping request is used to request to replace the target face in the target image with the source face in the source image. The identity feature represents the object to which the source face belongs, and the initial attribute feature represents the three-dimensional attributes of the target face;

[0008] Input the identity feature into the generator in the trained face swapping module, and input the at least one scale of initial attribute features into the convolutional layer corresponding to the scale in the generator respectively, and output a target face-swapped image. The face in the target face-swapped image fuses the identity feature of the source face and the target attribute feature of the target face;

[0009] Wherein, through each convolutional layer of the generator, the following steps are performed on the input identity feature and the initial attribute feature corresponding to the scale:

[0010] Obtain the first feature map output by the previous convolutional layer of the current convolutional layer;

[0011] Generate a second feature map based on the identity feature and the first feature map, and determine a control mask of the target image at a corresponding scale based on the second feature map and the initial attribute feature, where the control mask represents pixel points carrying features other than the identity feature of the target face;

[0012] Screen the initial attribute feature based on the control mask to obtain a target attribute feature;

[0013] Generate a third feature map based on the target attribute feature and the second feature map, and output the third feature map to the next convolutional layer of the current convolutional layer as the first feature map of the next convolutional layer.

[0014] On the other hand, an image processing apparatus is provided, and the apparatus includes:

[0015] A feature acquisition module, configured to acquire an identity feature of a source image and at least one scale of initial attribute features of a target image in response to a received face swapping request;

[0016] The face swapping request is used to request to swap the target face in the target image with the source face in the source image, the identity feature represents an object to which the source face belongs, and the initial attribute feature represents three-dimensional attributes of the target face;

[0017] A face swapping module, configured to input the identity feature into a generator in a trained face swapping module, and input the at least one scale of initial attribute features into convolutional layers corresponding to the scales in the generator respectively, and output a target face swapped image, where the face in the target face swapped image fuses the identity feature of the source face and the target attribute feature of the target face;

[0018] Wherein, when the face swapping module processes the input identity feature and the initial attribute feature corresponding to the scale through each convolutional layer of the generator, it includes:

[0019] An acquisition unit, configured to acquire a first feature map output by the previous convolutional layer of the current convolutional layer;

[0020] A generation unit, configured to generate a second feature map based on the identity feature and the first feature map;

[0021] A control mask determination unit, configured to determine a control mask of the target image at a corresponding scale based on the second feature map and the initial attribute feature, where the control mask represents pixel points carrying features other than the identity feature of the target face;

[0022] An attribute screening unit, configured to screen the initial attribute feature based on the control mask to obtain a target attribute feature;

[0023] The generating unit is further configured to generate a third feature map based on the target attribute feature and the second feature map, and output the third feature map to the next convolutional layer of the current convolutional layer as the first feature map of the next convolutional layer.

[0024] In a possible implementation, the generating unit is further configured to perform an affine transformation on the identity feature to obtain a first control vector; based on the first control vector, map the first convolutional kernel of the current convolutional layer to a second convolutional kernel, and perform a convolution operation on the first feature map based on the second convolutional kernel to generate a second feature map.

[0025] In a possible implementation, the control mask determining unit is configured to perform feature splicing on the second feature map and the initial attribute feature to obtain a spliced feature map; map the spliced feature map to the control mask based on a pre-configured mapping convolutional kernel and an activation function.

[0026] In a possible implementation, when the device trains a face-swapping model, it further includes:

[0027] A sample acquisition module, configured to acquire a sample data set, where the sample data set includes at least one pair of sample images, and each pair of sample images includes a sample source image and a sample target image;

[0028] A sample feature acquisition module, configured to input the sample data set into an initial model to obtain the sample identity feature of the sample source image in each pair of sample images, and at least one scale of sample initial attribute features of the sample target image;

[0029] A generating module, configured to determine at least one scale of sample masks based on the sample identity feature of the sample source image in each pair of sample images and at least one scale of sample initial attribute features of the sample target image through an initial generator of the initial model, and generate a sample generated image corresponding to each pair of sample images based on the sample identity feature, at least one scale of sample masks, and sample initial attribute features;

[0030] A discrimination module, configured to input the sample source image and the sample generated image in each pair of sample images into an initial discriminator of the initial model to obtain discrimination results of the initial discriminator on the sample source image and the sample generated image respectively;

[0031] A loss determination module, configured to, for each pair of sample images, determine a first loss value based on at least one scale of sample masks of the sample target image in each pair of sample images, and determine a second loss value based on the discrimination results of the initial discriminator on the sample source image and the sample generated image respectively;

[0032] The loss determination module is further configured to obtain the total training loss based on the first loss value and the second loss value corresponding to the at least one pair of sample images;

[0033] The training module is configured to train the initial model based on the total training loss until the training stops when the target condition is met, and obtain the face swapping model.

[0034] In a possible implementation, the at least one pair of sample images includes at least one pair of first sample images and at least one pair of second sample images, the first sample images include a sample source image and a sample target image belonging to the same object, and the second sample images include a sample source image and a sample target image belonging to different objects;

[0035] The loss determination module is further configured to:

[0036] Based on the sample generated images and the sample target images included in the at least one pair of first sample images, obtain the third loss value corresponding to the at least one pair of first sample images;

[0037] Based on the third loss value corresponding to the at least one pair of first sample images, and the first loss value and the second loss value corresponding to the at least one pair of sample images, obtain the total training loss.

[0038] In a possible implementation, the initial discriminator includes at least one initial convolutional layer; the loss determination module is further configured to:

[0039] For each pair of sample images, determine the first similarity between the non-face regions of the first discriminative feature map and the non-face regions of the second discriminative feature map, where the first discriminative feature map is the feature map corresponding to the sample target image output by the first part of the at least one initial convolutional layer, and the second discriminative feature map is the feature map corresponding to the sample generated image output by the first part of the at least one initial convolutional layer;

[0040] Determine the second similarity between the third discriminative feature map and the fourth discriminative feature map, where the third discriminative feature map is the feature map corresponding to the sample target image output by the second part of the at least one initial convolutional layer, and the fourth discriminative feature map is the feature map corresponding to the sample generated image output by the second part of the at least one initial convolutional layer;

[0041] Based on the first similarity and the second similarity corresponding to each pair of sample images, determine the fourth loss value corresponding to the at least one pair of sample images;

[0042] Based on the first loss value, the second loss value, and the fourth loss value corresponding to the at least one pair of sample images, obtain the total training loss.

[0043] In a possible implementation, the loss determination module is further configured to:

[0044] For each pair of sample images, respectively extract the first identity feature of the sample source image, the second identity feature of the sample target image, and the third identity feature of the sample generated image;

[0045] Based on the first identity feature and the third identity feature, determine the first identity similarity between the sample source image and the sample generated image;

[0046] Based on the second identity feature and the third identity feature, determine the first identity distance between the sample generated image and the sample target image;

[0047] Based on the first identity feature and the second identity feature, determine the second identity distance between the sample source image and the sample target image;

[0048] Based on the first identity distance and the second identity distance, determine the distance difference;

[0049] Based on the first identity similarity and the distance difference corresponding to each pair of sample images, determine the fifth loss value corresponding to at least one pair of sample images;

[0050] Based on the first loss value, the second loss value, and the fifth loss value corresponding to the at least one pair of sample images, obtain the total training loss.

[0051] On the other hand, a computer device is provided, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the above-mentioned image processing method.

[0052] On the other hand, a computer-readable storage medium is provided, on which a computer program is stored, and the computer program implements the above-mentioned image processing method when executed by a processor.

[0053] On the other hand, a computer program product is provided, including a computer program, and the computer program implements the above-mentioned image processing method when executed by a processor.

[0054] The beneficial effects brought by the technical solution provided by the embodiments of the present application are:

[0055] The image processing method according to the embodiment of the present application includes obtaining the identity feature of the source image and at least one scale of initial attribute features of the target image; inputting the identity feature into the generator in the trained face swapping module, and respectively inputting the at least one scale of initial attribute features into the convolutional layers corresponding to the scales in the generator to obtain the target face-swapped image; in each convolutional layer of the generator, a second feature map can be generated based on the identity feature and the first feature map output by the previous convolutional layer; and based on the second feature map and the initial attribute features, a control mask corresponding to the target image at the corresponding scale is determined to accurately locate the pixel points of the features other than the identity feature of the target face carried in the target image; by screening out the target attribute features in the initial attribute features based on the control mask, generating a third feature map based on the target attribute features and the second feature map, and outputting it to the next convolutional layer, after being processed layer by layer through at least one convolutional layer, the attributes and detailed features of the target face are effectively retained in the final target face-swapped image, greatly improving the clarity of the face in the face-swapped image, achieving high-definition face swapping, and improving the accuracy of face swapping. Description of the Drawings

[0056] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0057] Figure 1 Schematic diagram of the implementation environment of an image processing method provided by an embodiment of the present application;

[0058] Figure 2 Schematic diagram of the process of an image processing method provided by an embodiment of the present application;

[0059] Figure 3 Schematic diagram of the structure of a face swapping model provided by an embodiment of the present application;

[0060] Figure 4 Schematic diagram of the structure of a block in a generator provided by an embodiment of the present application;

[0061] Figure 5 Schematic diagram of the process of a face swapping model training method provided by an embodiment of the present application;

[0062] Figure 6 Schematic diagram of a control mask of at least one scale provided by an embodiment of the present application;

[0063] Figure 7 Schematic diagram of a comparison of face swapping results provided by an embodiment of the present application;

[0064] Figure 8 Schematic diagram of the structure of an image processing device provided by an embodiment of the present application;

[0065] Figure 9 Schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0066] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the implementation manners described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0067] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include plural forms. The terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, but do not exclude the implementation of other features, information, data, steps, operations, etc. supported by the art of the present technology.

[0068] It can be understood that in the specific implementation manners of the present application, any data related to an object, such as the source image, target image, source face, target face, and at least one pair of samples in the sample dataset used for model training, and any data related to an object, such as the image to be face-swapped, face features of the target face, and attribute parameters used for face swapping using the face-swapping model, are all obtained after the consent or permission of the relevant object; when the following embodiments of the present application are applied to specific products or technologies, the consent or permission of the object needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. In addition, the face-swapping process performed on the face image of any object using the image processing method of the present application is a face-swapping process that is executed after the face-swapping service or face-swapping request triggered by the relevant object and after the consent or permission of the relevant object.

[0069] The image processing method provided by the embodiments of the present application involves the following artificial intelligence, computer vision and other technologies. Exemplarily, for example, using technologies such as cloud computing and big data processing in artificial intelligence technology to implement processes such as face-swapping model training and extraction of multi-scale attribute features in an image. For example, using computer vision technology to perform face recognition on an image to obtain the identity features corresponding to the face in the image.

[0070] It should be understood that artificial intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0071] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0072] It should be understood that computer vision technology (CV) is a science that studies how to enable machines to "see". Further, it refers to using cameras and computers to replace human eyes to identify and measure targets, etc., for machine vision, and further performing graphic processing to make the computer process images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0073] Figure 1 Schematic diagram of the implementation environment for an image processing method provided for this application. As Figure 1 shown, the implementation environment includes: server 11 and terminal 12.

[0074] The server 11 is configured with a trained face-swapping model, and the server 11 can provide a face-swapping function to the terminal 12 based on the face-swapping model. The face-swapping function can be used to generate a face-swapped image based on a source image and a target image. The generated face-swapped image has the identity features of the source face in the source image and the attribute features of the target face in the template image. The identity features represent the object to which the source face belongs, and the initial attribute features represent the three-dimensional attributes of the target face.

[0075] In a possible scenario, the terminal 12 is installed with an application program, and the application program can be pre-configured with a face-swapping function. The server 11 can be the background server of the application program. The terminal 12 and the server 11 can perform data interaction based on the application program to implement the face-swapping process. Exemplarily, the terminal 12 can send a face-swapping request to the server 11, and the face-swapping request is used to request to swap the target face in the target image with the source face in the source image. The server 11 can execute the image processing method of the present application to generate a target face-swapped image based on the face-swapping request and return the target face-swapped image to the terminal 12. For example, the application program is any application that supports the face-swapping function. For example, the application program includes but is not limited to: video editing applications, image processing tools, video applications, live broadcast applications, social applications, content interaction platforms, game applications, and so on.

[0076] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The above network can include but is not limited to: wired networks, wireless networks. Among them, the wired network includes: local area networks, metropolitan area networks, and wide area networks. The wireless network includes: Bluetooth, Wi-Fi, and other networks that implement wireless communication. The terminal can be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a laptop computer, a digital broadcast receiver, a MID (Mobile Internet Devices), a PDA (Personal Digital Assistant), a desktop computer, a vehicle-mounted terminal (such as a vehicle-mounted navigation terminal, a vehicle-mounted computer, etc.), a smart home appliance, an aircraft, a smart speaker, a smart watch, etc. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, but it is not limited to this. Specifically, it can also be determined based on the actual application scenario requirements and is not limited here.

[0077] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.

[0078] First, the technical terms related to this application will be introduced as follows:

[0079] Face swapping: It is to replace the face in an image with another face. Exemplarily, given a source image X s and a target image X t , use the image processing method of this application to generate a face-swapped image Y s,t . Among them, the face-swapped image Y s,t has the identity feature of the face in the source image X s ; at the same time, it retains the attribute features irrelevant to identity in the target image X t .

[0080] Face-swapping model: It is used to replace the target face in the target image with the source face in the source image.

[0081] Source image: An image that provides identity features, and the face in the generated face-swapped image has the identity features of the face in this source image;

[0082] Target image: An image that provides attribute features, and the face in the generated face-swapped image has the attribute features of the face in this target image. For example, if the source image is an image of object A and the target image is an image of object B, and the face of object B in the target image is replaced with the face of object A to obtain a face-swapped image; then the identity of the face in the face-swapped image is the face of object A, the face in the face-swapped image is the same as the face of object A in terms of identity features such as eye shape, interpupillary distance, and nose size, and the face in the face-swapped image has the attribute features of the face of object B such as expression, hair, lighting, wrinkles, pose, and facial occlusion.

[0083] Figure 2 FIG. [Specific figure number] is a schematic flowchart of an image processing method provided by an embodiment of this application. The execution subject of this method can be a computer device. As Figure 2 shown, this method includes the following steps.

[0084] Step 201, the computer device, in response to the received face-swapping request, obtains the identity features of the source image and at least one scale of the initial attribute features of the target image.

[0085] The face swapping request is used to request swapping the target face in the target image with the source face in the source image. The computer device can utilize a trained face swapping model to obtain a face swapped image, thereby providing a face swapping function. Among them, the identity feature characterizes the object to which the source face belongs; exemplarily, the identity feature can be a feature that identifies the identity of the object, and the identity feature can include at least one of the target face facial feature or the target face contour feature of the object; the target face facial feature refers to the feature corresponding to the facial features, and the target face contour feature refers to the feature corresponding to the contour of the target face. For example, the identity feature can include but is not limited to eye shape, interpupillary distance, nose size, eyebrow shape, facial contour, etc. The initial attribute feature characterizes the three-dimensional attributes of the target face. Exemplarily, the initial attribute feature can characterize attributes such as the posture and spatial environment of the target face in three-dimensional space; for example, the initial attribute feature can include but is not limited to background, lighting, wrinkles, posture, expression, hair, facial occlusion, etc.

[0086] In a possible implementation manner, the face swapping model can include an identity recognition network. The computer device can input the source image into the face swapping model, and perform face recognition on the source image through the identity recognition network in the face swapping model to obtain the identity feature of the source image. Exemplarily, the identity recognition network is used to recognize the identity to which the face in the input image belongs. For example, the identity recognition network can be the Fixed FR Net (Fixed Face Recognition Network, fixed face recognition network) in the face swapping model. For example, when the source image is a human face image, the identity recognition network can be a trained face recognition model, and the face recognition model is used to recognize the object to which the human face in the source image belongs. The identity feature can be a 512-dimensional feature vector output by the face recognition model. The 512-dimensional feature vector can characterize features such as eye shape, interpupillary distance, nose size, eyebrow shape, facial contour, etc.

[0087] In a possible implementation, the face-swapping model further includes an attribute feature map extraction network, which can be a U-shaped deep network including an encoder and a decoder. The steps for the computer device to obtain the initial attribute features of at least one scale of the target image include: the computer device can perform layer-by-layer downsampling on the target image through at least one encoding network layer of the encoder to obtain encoded features; and perform layer-by-layer upsampling on the encoded features through at least one decoding network layer of the decoder, and use the decoded features of different scales output by at least one decoding network layer as the initial attribute features of at least one scale. Exemplarily, each encoding network layer is used to perform an encoding operation on the target image to obtain encoded features, and each decoding network layer is used to perform a decoding operation on the encoded features to obtain initial attribute features. When running, the decoder can perform reverse operations according to the operating principle of the encoder. For example, the encoder can perform downsampling on the target image, and the decoder can perform upsampling on the downsampled encoded features. For example, if the encoder is an Autoencoder (AE), then the decoder can be the decoder corresponding to the autoencoder.

[0088] In a possible example, each encoding network layer is used to perform downsampling on the encoded features output by the previous encoding network layer to obtain encoded features of at least one scale; each decoding network layer is used to perform upsampling on the decoded features output by the previous decoding network layer to obtain the initial attribute features of at least one scale; wherein, each decoding network layer can also perform upsampling on the initial attribute features output by the previous decoding network layer in combination with the encoded features of the encoding network layer of the corresponding scale. As Figure 3 shown, Figure 3 a U-shaped deep network is used in to extract features from the target image X t For example, the target image is input into the encoder, and through multiple encoding network layers of the encoder, the target image X is output tThe resolutions of the feature maps of the encoding features are 1024×1024, 512×512, 256×256, 128×128, and 64×64 in sequence. Then, the 64×64 feature map is input into the first decoding network layer of the decoder for upsampling to obtain a decoded feature map of size 128×128. The decoded feature map of size 128×128 is concatenated with the encoding feature map of 128×128, and then the concatenated feature map is upsampled to obtain a 256×256 decoded feature map, and so on. The feature map of at least one resolution decoded by the network structure of the sampling U-shaped deep network is used as the initial attribute feature of at least one scale. Among the initial attribute features of at least one scale, the initial attribute feature of each scale prominently represents the attribute feature of the target image at the corresponding scale. The features highlighted by the initial attribute features of different scales can be different. The initial attribute feature of a relatively small scale prominently represents the global position, pose, etc. of the target face in the target image, and the initial attribute feature of a relatively large scale prominently represents the local details of the target face in the target image, so that the initial attribute features of at least one scale encompass features of multiple levels. For example, the initial attribute features of at least one scale can be multiple feature maps with increasing resolutions from small to large. The feature map with resolution R1 can represent the face position of the target face in the target image, the feature map with resolution R2 can represent the pose expression of the target face in the target image, and the feature map with resolution R3 can represent the facial details of the face position of the target face in the target image; where the resolution R1 is less than R2, and R2 is less than R3.

[0089] Step 202, the computer device inputs the identity feature into the generator in the trained face swapping module, and inputs the initial attribute features of at least one scale into the convolutional layers corresponding to the scales in the generator, and outputs a target face swapped image.

[0090] The identity feature of the source face and the target attribute feature of the target face are fused in the face of the target face swapped image. Exemplarily, the computer device can input the identity feature into each convolutional layer of the generator. The computer device inputs the initial attribute features of at least one scale into the convolutional layer that matches the scale of the initial attribute feature in the generator. Among them, the scales of the feature maps output by each convolutional layer of the generator are different, and the convolutional layer that matches the scale of the initial attribute feature refers to the scale of the feature map to be output by the convolutional layer being the same as the scale of the initial attribute feature. For example, if a certain convolutional layer in the generator is used to process a 64×64 size feature map from the previous convolutional layer and output a 128×128 size feature map, then the initial attribute feature of size 128×128 can be input into this convolutional layer.

[0091] In a possible implementation, in the generator, the computer device may determine a control mask for at least one scale of the target image based on the identity feature and the initial attribute features of at least one scale, and obtain the target face-swapped image based on the identity feature, the control mask of at least one scale, and the initial attribute features. Exemplarily, the control mask represents the pixel points carrying features other than the identity feature of the target face. Then, the computer device may determine the target attribute features of at least one scale based on the control mask of at least one scale and the initial attribute features, and generate the target face-swapped image based on the identity feature and the target attribute features of at least one scale.

[0092] The computer device may obtain the target face-swapped image through successive processing by the convolutional layers of the generator. In a possible example, the computer device performs the following steps S1 to S4 on the input identity feature and the initial attribute features of the corresponding scale through the convolutional layers of the generator:

[0093] Step S1: The computer device obtains the first feature map output by the previous convolutional layer of the current convolutional layer.

[0094] In the generator, each convolutional layer may process the feature map output by the previous convolutional layer and output it to the next convolutional layer. Among them, for the first convolutional layer, the computer device may input the initial feature map into the first convolutional layer. For example, the initial feature map may be a 4×4×512 all-zero feature vector. For the last convolutional layer, the computer device may generate the final target face-swapped image based on the feature map output by the last convolutional layer.

[0095] Step S2: The computer device generates a second feature map based on the identity feature and the first feature map, and determines the control mask of the target image at the corresponding scale based on the second feature map and the initial attribute features.

[0096] The control mask represents the pixel points carrying features other than the identity feature of the target face.

[0097] In a possible implementation, the computer device adjusts the weights of the convolutional kernels of the current convolutional layer based on the identity feature, and obtains the second feature map based on the first feature map and the adjusted convolutional kernels. Exemplarily, the steps for the computer device to generate the second feature map may include: the computer device performs an affine transformation on the identity feature to obtain a first control vector; the computer device maps the first convolutional kernel of the current convolutional layer to a second convolutional kernel based on the first control vector, and performs a convolution operation on the first feature map based on the second convolutional kernel to generate the second feature map. Exemplarily, the identity feature may be represented in the form of an identity feature vector, and the affine transformation refers to an operation of linearly transforming and translating the identity feature vector to obtain the first control vector. The affine transformation operation includes, but is not limited to, translation, scaling, rotation, and flipping transformations. Each convolutional layer of the generator includes a trained affine parameter matrix, and the computer device may perform transformations such as translation, scaling, rotation, and flipping on the identity feature vector based on the affine parameter matrix. Exemplarily, the computer device may perform modulation operation (Mod) and demodulation operation (Demod) on the first convolutional layer of the current convolutional layer through the first control vector to obtain the second convolutional kernel. Among them, the modulation operation may be a scaling process of the weights of the convolutional kernels of the current convolutional layer, and the demodulation operation may be a normalization process of the scaled weights of the convolutional kernels. For example, the computer device may perform a scaling process on the convolutional kernel weights through the scaling ratio corresponding to the first feature map input to the current convolutional layer and the first control vector.

[0098] In a possible implementation, the computer device obtains a control mask corresponding to the corresponding scale based on the second feature map and the initial attribute feature of the corresponding scale input to the current convolutional layer. This process may include: the computer device splices the second feature map and the initial attribute feature to obtain a spliced feature map; the computer device maps the spliced feature map to the control mask based on a pre-configured mapping convolutional kernel and activation function. Exemplarily, the control mask is a binary map. In this binary map, the pixel points carrying features other than the identity feature of the target face, such as pixel points in the hair area and background area, take a value of 1, and the pixel points carrying the identity feature take a value of 0. Exemplarily, the mapping convolutional kernel may be a convolutional kernel with a size of 1×1, and the activation function may be a Sigmoid function. For example, the second feature map and the initial attribute feature may be represented in the form of feature vectors. The computer device may merge the feature vector corresponding to the second feature map and the feature vector corresponding to the initial attribute feature to obtain the spliced vector, and perform a convolution operation and an activation operation on the spliced vector to obtain the control mask.

[0099] Exemplarily, the generator may include multiple blocks, each block including multiple layers. The computer device inputs the identity feature and the initial attribute feature at each scale into the block corresponding to that scale. In this block, the input identity feature and initial attribute feature can be processed layer by layer through at least one layer. Exemplarily, Figure 4 shows the network structure of the i-th block (i-th GAN block, the i-th adversarial network block) in the generator, where N represents the Attribute Injection module, and the internal structure of this Attribute Injection module is enlarged and shown in the right dashed box. As Figure 4 shown, the i-th block includes two layers. Here, the first layer is taken as an example for illustration. Figure 4 On the left side in, w represents the identity feature f of the source image id , A represents the Affine Transform operation. After performing the affine transform operation on the identity feature vector, a first control vector is obtained. Figure 4 In, Mod and Demod represent the modulation and demodulation operations on the convolutional kernel Conv 3×3. After the computer device performs the upsampling operation on the first feature map input to the current layer of the current block, it uses the convolutional kernel Conv 3×3 after the Mod and Demod operations to perform a convolution operation on the upsampled first feature map to obtain a second feature map. Then, the computer device concatenates the second feature map and the initial attribute feature f i att input to the current block, and uses the convolutional kernel Conv 1×1 and the Sigmoid function to map the concatenated feature vector obtained by concatenation to the control mask M corresponding to the current layer i,j att .

[0100] Step S3: The computer device filters the initial attribute feature based on this control mask to obtain the target attribute feature.

[0101] The computer device can perform a dot product of the feature vector corresponding to this control mask and the feature vector corresponding to the initial attribute feature to filter out the target attribute feature in the initial attribute feature.

[0102] As Figure 4 shown, the computer device can perform a dot product of the control mask M i,j att and the initial attribute feature f i att , and add the feature vector obtained by the dot product to the feature vector corresponding to the second feature map to obtain this target attribute feature.

[0103] Step S4: The computer device generates a third feature map based on the target attribute feature and the second feature map, and outputs the third feature map to the next convolutional layer of the current convolutional layer as the first feature map of the next convolutional layer.

[0104] The computer device can add the feature vector corresponding to the second feature map to the feature vector corresponding to the target attribute feature to obtain the third feature map.

[0105] It should be noted that for each convolutional layer included in the generator, the computer device can repeatedly execute the above steps S1 to S4 until the above steps S1 to S4 are repeatedly executed for the last convolutional layer of the generator, obtaining the third feature map output by the last convolutional layer, and generating a target face-swapped image based on the third feature map output by the last convolutional layer.

[0106] As Figure 4 shown, if the i-th block includes two layers, the third feature map can be input into the second layer of the i-th block, and the operations of the first layer can be repeated, and the feature map obtained by the second layer can be output to the next block, and so on until the last block. As Figure 3 shown, the Figure 3 where N represents the Attribute Injection module, and the dashed box represents the Generator of the StyleGAN2 model. For the N blocks included in the generator, the identity feature f s of the source image X id is respectively input, and the corresponding initial attribute features f1 att 、f2 att 、……、f i att 、……、f N-1 att 、f N att are respectively input into N blocks, and the processes of the above steps S1 to S4 are executed in each block until the features output by the last block are obtained. Based on the feature map output by the last block, the final target face-swapped image Y s,t is generated, thus completing the face swap.

[0107] Figure 5 is a schematic flowchart of a method for training a face-swapping model provided by an embodiment of the present application. The execution subject of this method can be a computer device. As Figure 5 shown, this method includes:

[0108] Step 501: The computer device obtains a sample data set.

[0109] The sample data set includes at least a pair of sample images, and each pair of sample images includes a sample source image and a sample target image. In a possible implementation, the at least a pair of sample images includes at least a pair of first sample images and at least a pair of second sample images. The first sample images include a sample source image and a sample target image belonging to the same object, and the second sample images include a sample source image and a sample target image belonging to different objects. For example, the multiple pairs of sample images include a source image X s and a target image X t that form a pair of sample images, and also include a source image X s and a target image X t of object B that form a pair of sample images. Each pair of sample images is correspondingly labeled with a ground-truth label, and the ground-truth label represents whether the pair of sample images belongs to the same object.

[0110] Step 502: The computer device inputs the sample data set into the initial model, and obtains the sample identity features of the sample source image in each pair of sample images, and the sample initial attribute features of at least one scale of the sample target image.

[0111] The initial model may include an initial identity recognition network and an attribute feature map extraction network. The computer device can respectively extract the sample identity features of the sample source image and the sample initial attribute features of at least one scale of the sample target image through the initial identity recognition network and the attribute feature map extraction network. It should be noted that the implementation of obtaining the sample identity features and the sample initial attribute features in step 502 is a process similar to the way of obtaining the identity features and the initial attribute features in step 201 above, and will not be elaborated here one by one.

[0112] Step 503: The computer device determines the sample masks of at least one scale based on the sample identity features of the sample source image in each pair of sample images and the sample initial attribute features of at least one scale of the sample target image through the initial generator of the initial model, and generates the sample generated images corresponding to each pair of sample images based on the sample identity features, the sample masks of at least one scale, and the sample initial attribute features.

[0113] The initial generator includes multiple initial convolutional layers. For each pair of sample images, the computer device can input the sample identity features into each initial convolutional layer, input the sample initial attribute features of at least one scale into the initial convolutional layer that matches the scale of the sample initial attribute features, and obtain the sample generated images after being processed layer by layer by each initial convolutional layer.

[0114] Exemplarily, the computer device can perform the following steps on the input sample identity feature and the sample initial attribute feature of the corresponding scale through each initial convolutional layer of the initial generator: The computer device obtains the first sample feature map output by the previous initial convolutional layer of the current initial convolutional layer; based on the sample identity feature and the first sample feature map, generates a second sample feature map, and based on the second sample feature map and the sample initial attribute feature, determines the sample mask of the sample target image at the corresponding scale; the computer device filters the sample initial attribute feature based on the sample mask to obtain the sample target attribute feature; the computer device generates a third sample feature map based on the sample target attribute feature and the second sample feature map, and outputs the third sample feature map to the next initial convolutional layer of the current initial convolutional layer as the first sample feature map of the next initial convolutional layer. Repeat this cycle until the above steps are repeatedly executed on the last initial convolutional layer of the initial generator to obtain the third initial feature map output by the last initial convolutional layer, and obtain the sample generated image based on the third initial feature map output by the last convolutional layer.

[0115] It should be noted that in the model training stage, the steps performed through each initial convolutional layer are the same as the steps performed by each convolutional layer in the generator of the trained face-swapping model (i.e., the above steps S1-S4), which will not be elaborated here one by one.

[0116] Step 504: The computer device inputs the sample source image and the sample generated image in each pair of sample images into the initial discriminator of the initial model to obtain the discrimination results of the initial discriminator on the sample source image and the sample generated image respectively.

[0117] The initial model may further include an initial discriminator. For each pair of sample images, the computer device inputs the sample source image and the sample generated image into the discriminator, and the discriminator outputs the first discrimination result on the sample source image and the second discrimination result on the sample generated image. Among them, the first discrimination result can represent the probability that the sample source image is a real image; the second discrimination result can represent the probability that the sample generated image is a real image.

[0118] In a possible implementation manner, the initial discriminator includes at least one initial convolutional layer; each initial convolutional layer is configured to process the discriminative feature map output by the previous initial convolutional layer of the initial discriminator and output it to the next initial convolutional layer of the initial discriminator. Each initial convolutional layer can output a discriminative feature map for feature extraction of the sample source image and a discriminative feature map for feature extraction of the sample generated image. Until the last initial convolutional layer of the initial discriminator, a first discrimination result is obtained based on the discriminative feature map of the sample source image output by the last initial convolutional layer; and a second discrimination result is obtained based on the discriminative feature map of the sample generated image output by the last initial convolutional layer.

[0119] Step 505: For each pair of sample images, the computer device determines a first loss value based on at least one scale of the sample mask of the sample target image in each pair of sample images, and determines a second loss value based on the discrimination results of the initial discriminator for the sample source image and the sample generated image respectively.

[0120] The computer device can accumulate at least one scale of the sample mask, and use the accumulated value corresponding to at least one scale of the sample mask as the first loss value. For example, the sample mask can be a binary map, and the computer device can accumulate the values of each pixel point in the binary map to obtain the first sum value corresponding to each sample mask, and accumulate the first sum values corresponding to at least one scale of the sample mask to obtain the first loss value.

[0121] Exemplarily, taking the initial generator including at least one initial block, and each initial block including at least one layer as an example, for each pair of sample images, the computer device can determine the first loss value based on at least one scale of the sample mask of the sample target image in each pair of sample images through the following formula (1):

[0122] Formula (1): L mask = ∑ i,j |M i,j |1;

[0123] Where, L mask represents the first loss value, i represents the i-th block of the initial generator, and j represents the j-th layer of the i-th block. M i,j represents the sample mask of the j-th layer of the i-th block. The computer device can accumulate the sample masks of at least one layer of at least one block through the above formula (1), and in the training stage, minimize the first loss value L mask, to train the generator so that the obtained control mask can effectively represent the pixel points of the key attribute features other than the identity features, and then the key attribute features in the initial attribute features can be screened out through the control mask, the redundant features in the initial attribute features can be filtered out, and the key and necessary features in the initial attribute features can be retained, so as to avoid redundant attributes and finally improve the accuracy of the generated face-swapped image.

[0124] It should be noted that the refinement degrees of the pixel points representing the features other than the identity features of the target face carried by the binary images of different scales are different. Figure 6 Shows the sample masks of different scales corresponding to each of the three target images. Each row of sample masks is the sample masks of each scale corresponding to one of the target images. As Figure 6 shown, for any target image, the resolutions of the sample masks from left to right increase in turn. Taking the change of the sample masks of each scale in the first row as an example, from 4×4, 8×8, 16×16, 32×32, the position of the face in the target image is gradually and clearly located. Among them, the pixel values of the pixel points corresponding to the face area are 0, and the pixel values of the pixel points corresponding to the background area outside the face area are 0; from 64×64, 128×128, 16×16, 256×256, 512×512, 1024×1024, the facial pose and expression of the face in the target image are gradually and clearly shown, and then the facial details of the face in the target image are gradually reflected.

[0125] Exemplarily, the computer device can determine the second loss value based on the discrimination results of the sample source image and the sample generated image by the initial discriminator through the following formula two:

[0126] Formula two: L GAN = min G max D E[log(D(X s ))]+E[log(1 - D(Y s,t ))];

[0127] Among them, L GAN represents the second loss value, D(X s ) represents the first discrimination result of the initial discriminator on the sample source image. This first discrimination result can be the probability that the sample source image X s is a real image; D(Y s,t ) represents the second discrimination result of the initial discriminator on the sample generated image Y s,t . This second discrimination result can be the probability that the sample generated image is a real image; E[log(D(X s ))] refers to the expectation of log(D(X s )) and can represent the loss value of the initial discriminator; E[log(1 - D(Ys,t ))] refers to the expectation of log(1 - D(Y s,t )) and can represent the loss value of the initial generator. It represents that the initial generator expects to minimize the loss function value, and represents that the initial discriminator maximizes the loss function value; it should be noted that this initial model includes an initial generator and an initial discriminator and can be an initial adversarial network. The initial adversarial network learns by having the initial generator and the initial discriminator play against each other to obtain the desired machine learning model, which is a method of unsupervised learning. The training objective of the initial generator is to obtain the desired output based on the input; the training objective of the initial discriminator is to distinguish the images generated by the initial generator from the real images as much as possible. The input of the initial discriminator includes the sample source image and the sample generated image generated by the initial generator. The two network models learn against each other and continuously adjust the parameters. The ultimate goal is for the initializer to deceive the initial discriminator as much as possible so that the initial discriminator cannot determine whether the image generated by the initial generator is real.

[0128] Step 506, the computer device obtains the total training loss based on the first loss value and the second loss value corresponding to the at least one pair of sample images.

[0129] In a possible implementation manner, the computer device may use the sum value of the first loss value and the second loss value as the total training loss.

[0130] In a possible implementation manner, the at least one pair of sample images includes at least one pair of first sample images and at least one pair of second sample images. Then the computer device may also train based on the sample images of the same object. Before the computer device determines the total training loss, the computer device obtains the third loss value corresponding to the at least one pair of first sample images based on the sample generated image and the sample target image included in the at least one pair of first sample images. Then the steps for the computer device to determine the total training loss may include: the computer device obtains the total training loss based on the third loss value corresponding to the at least one pair of first sample images, and the first loss value and the second loss value corresponding to the at least one pair of sample images.

[0131] Exemplarily, the computer device may obtain the third loss value corresponding to the at least one pair of first sample images through the following formula three based on the sample generated image and the sample target image included in the at least one pair of first sample images:

[0132] Formula three: L rec =|Y s,t -X t |1;

[0133] where, L rec represents the third loss value, Ys,t Represents a sample-generated image corresponding to a pair of first sample images, X t Represents the sample target image in the pair of first sample images. It should be noted that when the sample source image and the sample target image belong to the same object, by constraining the face-swapping result to be the same as the sample target image, when the trained face-swapping model performs face swapping on images of the same object, the generated face-swapped images are close to the target image, thereby improving the accuracy of model training.

[0134] In a possible implementation, the initial discriminator includes at least one initial convolutional layer; the computer device can calculate the loss based on the output results of each initial convolutional layer of the initial discriminator. Before determining the total training loss, for each pair of sample images, the computer device determines the first similarity between the non-face regions of the first discriminative feature map and the non-face region of the second discriminative feature map. The first discriminative feature map is the feature map corresponding to the sample target image output by the first part of the at least one initial convolutional layer, and the second discriminative feature map is the feature map corresponding to the sample-generated image output by the first part of the initial convolutional layer; the computer device determines the second similarity between the third discriminative feature map and the fourth discriminative feature map. The third discriminative feature map is the feature map corresponding to the sample target image output by the second part of the at least one initial convolutional layer, and the fourth discriminative feature map is the feature map corresponding to the sample-generated image output by the second part of the initial convolutional layer; the computer device determines the fourth loss value corresponding to at least one pair of sample images based on the first similarity and the second similarity corresponding to each pair of sample images; then the step of determining the total training loss may include: the computer device obtains the total training loss based on the first loss value, the second loss value, and the fourth loss value corresponding to the at least one pair of sample images.

[0135] Exemplarily, the computer device can determine the first similarity through a trained segmentation model. For example, the computer device can obtain the segmentation mask corresponding to the first discriminative feature map or the second discriminative feature map through the segmentation model, and determine the first similarity between the non-face regions of the first discriminative feature map and the non-face region of the second discriminative feature map based on the segmentation mask. Among them, the segmentation mask can be a binary map corresponding to the first discriminative feature map or the second discriminative feature map. In the binary map, the pixel values corresponding to the non-face regions are 1, and the pixel values corresponding to the regions outside the non-face regions are 0, so as to effectively extract the background regions outside the face.

[0136] Exemplarily, the computer device can determine the fourth loss value corresponding to the at least one pair of sample images through the following formula four:

[0137] Formula four:

[0138]

[0139] Among them, L FM represents the fourth loss value, and M bg represents the segmentation mask. The initial discriminator includes M initial convolutional layers. Among them, the first to the m-th initial convolutional layers are the first part of the initial convolutional layers, and the m-th to the M-th initial convolutional layers are the second part of the initial convolutional layers. D i (X t ) represents the feature map corresponding to the sample target image output by the i-th initial convolutional layer in the first part of the initial convolutional layers; D i (Y s,t ) represents the feature map corresponding to the sample generated image output by the i-th initial convolutional layer in the first part of the initial convolutional layers; D j (X t ) represents the feature map corresponding to the sample target image output by the j-th initial convolutional layer in the second part of the initial convolutional layers; D j (Y s,t ) represents the feature map corresponding to the sample generated image output by the j-th initial convolutional layer in the second part of the initial convolutional layers. It should be noted that the value of m is a positive integer not less than 0 and not greater than M, and the value of m can be configured based on needs, and this application does not make any limitations in this regard.

[0140] In a possible implementation manner, the computer device can also obtain the similarity between the identity features of each image respectively and perform loss calculation. Exemplarily, before determining the total training loss, for each pair of sample images, the computer device can extract the first identity feature of the sample source image, the second identity feature of the sample target image, and the third identity feature of the sample generated image respectively; based on the first identity feature and the third identity feature, determine the first identity similarity between the sample source image and the sample generated image; and, the computer device determines the first identity distance between the sample generated image and the sample target image based on the second identity feature and the third identity feature, and determines the second identity distance between the sample source image and the sample target image based on the first identity feature and the second identity feature; the computer device can determine the distance difference based on the first identity distance and the second identity distance; the computer device determines the fifth loss value corresponding to at least one pair of sample images based on the first identity similarity and the distance difference corresponding to each pair of sample images. Then the steps for the computer device to determine the total training loss may include: the computer device obtains the total training loss based on the first loss value, the second loss value, and the fifth loss value corresponding to the at least one pair of sample images.

[0141] Exemplarily, the computer device can determine the fifth loss value of the at least one pair of sample images through the following formula five:

[0142] Formula five:

[0143] L ICL = 1 - cos(z id (Y s,t ), z id (X s )) + (cos(z id (Y s,t ), z id (X t )) - cos(z id (X s ), z id (X t ))) 2 ;

[0144] Wherein, L ICL represents the fifth loss value, z id (X s ) represents the first identity feature of the source sample image, z id (X t ) represents the second identity feature of the target sample image, z id (Y s,t ) represents the third identity feature of the generated sample image; 1 - cos(z id (Y s,t ), z id (X s )) represents the first identity similarity between the source sample image and the generated sample image; cos(z id (Y s,t ), z id (X t )) represents the first identity distance between the generated sample image and the target sample image; cos(z id (X s ), z id (X t )) represents the second identity distance between the source sample image and the target sample image; (cos(z id (Y s,t ), z id (X t )) - cos(z id (X s ), z id (X t ))) 2 represents the distance difference.

[0145] It should be noted that the distance difference is determined by the first identity distance and the second identity distance. Since the distance between the sample source image and the sample target image is measured by the second identity distance, by minimizing the distance difference, the first identity distance, that is, the distance between the sample generated image and the sample target image, has a certain distance, and this distance is equivalent to the distance between the sample source image and the sample target image. Through the first identity similarity, it is ensured that the generated image has the identity characteristics of the target image, thereby improving the accuracy of model training and the accuracy of face swapping.

[0146] Taking the total training loss including the above five loss values as an example, the computer device can determine the total training loss through the following formula six;

[0147] Formula six: L total =L GAN +L mask +L FM +10*L rec +5*L ICL ;

[0148] Among them, L total represents the total training loss, L GAN represents the second loss value, L mask represents the first loss value, L rec represents the third loss value, L FM represents the fourth loss value, L ICL represents the fifth loss value.

[0149] Step 507, the computer device trains the initial model based on the total training loss until the training stops when the target conditions are met, and the face swapping model is obtained.

[0150] It should be noted that the computer device can perform iterative training on the initial model based on the above steps 501 to 507, and obtain the total training loss corresponding to each iterative training. The parameters of the initial model are adjusted based on the total training loss of each iterative training. For example, the parameters included in the encoder, decoder, initial generator, initial discriminator, etc. in the initial model are optimized. Until the total training loss meets the target conditions, the computer device stops training and uses the initial model optimized for the last time as the face swapping model. For example, the computer device can use the Adam algorithm optimizer, with a learning rate of 0.0001, to perform iterative training on the initial model until the target conditions are reached, and it is considered that the training converges and stops training; for example, the target conditions can be that the numerical value of the total loss is within the target numerical range, for example, the total loss is less than 0.5; or the time consumed by multiple iterative trainings exceeds the maximum duration, etc.

[0151] Figure 3The following is a schematic framework diagram of a face swap model provided by an embodiment of the present application. As Figure 3 shown, the computer device can use the face image of object A as the source image X s , and use the face image of object B as the target image X t . The computer device obtains the identity feature f of the source image through the Fixed FR Net (fixed face recognition network) id , and the computer device inputs the identity feature f id into N blocks included in the generator respectively; the computer device obtains at least one scale of initial attribute features f1 att , f2 att , ……, f i att , ……, f N-1 att , f N att of the target image through the encoder and decoder of the U-shaped depth network structure, and inputs them into the blocks corresponding to the respective scales. The computer device executes the processes of the above steps S1 to S4 for each block until the feature map output by the last block is obtained. The computer device generates the final target face swap image Y s,t based on the feature map output by the last block, thereby completing the face swap.

[0152] It should be noted that through the image processing method of the present application, high-definition face swapping can be realized. For example, a face swap image with a high resolution such as 1024 2 can be generated. At the same time, the generated high-resolution face swap image takes into account a high image quality, the identity consistency with the source face in the source image, and effectively retains the key attributes of the target face in the target image with high precision. Different from the method A in the related technology, which can only generate a face swap image with a lower resolution such as 256 2 , through the image processing method of the present application, by processing the initial attribute features and identity features of at least one scale in the successive convolutions of the generator, and using the control masks of at least one scale to screen the initial attribute features, redundant information such as the target face identity features is effectively filtered out in the obtained target attribute features, and the key attribute features of the target face are effectively retained; and the initial attribute features of at least one scale highlight the features corresponding to different scales. Through the control masks of larger scales corresponding to the initial attribute features of larger scales, high-clarity screening of the key attributes can be realized, so as to retain facial detail features such as hair, wrinkles, and facial occlusions of the target face with high precision, greatly improving the accuracy and clarity of the generated face swap image and the authenticity of the face swap image.

[0153] Moreover, the image processing method of the present application can directly generate the entire face-swapped image after face swapping. This entire face-swapped image includes both the face after face swapping and the background area, without the need for further fusion or enhancement processing in related technologies; greatly improving the processing efficiency of the face-swapping process.

[0154] Moreover, the face-swapping model training method of the present application can perform end-to-end training on the entire generation framework used to generate sample-generated images in the initial model during model training, avoiding the situation of error accumulation caused by multi-stage training, enabling the face-swapping model trained by the present application to generate face-swapped images more stably, and improving the stability and reliability of the face-swapping process.

[0155] Moreover, the image processing method of the present application can generate face-swapped images with higher resolution, and precisely retain details such as the texture, skin brightness, and hair strands of the target face in the target image, improving the accuracy, clarity, and authenticity of face swapping, and being applicable to scenarios with higher requirements for face-swapping quality such as games and movies. And for the virtual image maintenance scenario, the image processing method of the present application can achieve face swapping of the face of any object to the face of any object. For a specific virtual image, the face of the specific virtual image is swapped into the face image of any object, facilitating the maintenance of the virtual image and improving the convenience of virtual image maintenance.

[0156] The following shows a comparison of the face-swapping results using the image processing method of the present application and the face-swapping results of related technologies. It can be seen from the comparison that the high-definition face-swapping results generated by the image processing method of the present application show obvious superiority over related technologies in both qualitative and quantitative comparisons.

[0157] As Figure 7 shown, Figure 7 shows a comparison of the high-definition face-swapping results of some methods in related technologies (hereinafter referred to as Method A) and the solution proposed by the present application. It can be seen from the comparison that Method A has obvious problems of inconsistent skin brightness and cannot retain the hair occlusion on the face; the results generated by the solution proposed by the present application retain the skin brightness, expression, skin texture, occlusion and other attribute features of the target human face, and have better image quality and greater authenticity.

[0158] Table 1 below shows the quantitative comparison of the high-definition face swapping results of Method A in the related art and the proposed solution of this application. The experimental data in Table 1 compares the identity similarity (ID Retrieval) between the face in the generated face-swapped image and the face in the source image, the pose difference (Pose Error) between the face in the face-swapped image and the face in the target image, and the picture quality difference (FID) between the face in the face-swapped image and the real face image. From the experimental data in Table 1, it can be concluded that the identity similarity of the high-definition face swapping result of the proposed solution in this application is significantly higher than that of Method A in the related art; the pose difference of the high-definition face swapping result of the proposed solution in this application is lower than that of Method A in the related art, and the pose difference of this application's solution is even lower; the picture quality difference of the high-definition face swapping result of the proposed solution in this application is significantly lower than that of Method A in the related art, and the picture quality difference between the face-swapped image obtained by this application's solution and the real image is smaller. Therefore, the solution proposed in this application takes into account image quality, identity consistency with the source human face, and retention of the attributes of the target human face, and has significant superiority over Method A in the related art.

[0159] Table 1

[0160] ID Retrieval↑ Pose Error↓ FID↓ Method A in the related art 90.83 2.64 16.64 The solution proposed in this application 96.34 2.52 2.04

[0161] The image processing method of the embodiment of this application obtains the identity feature of the source image and the initial attribute features of at least one scale of the target image; inputs the identity feature into the generator in the trained face swapping module, and inputs the initial attribute features of the at least one scale into the convolutional layer corresponding to the scale in the generator respectively to obtain the target face-swapped image; in each convolutional layer of the generator, a second feature map can be generated based on the identity feature and the first feature map output by the previous convolutional layer; and based on the second feature map and the initial attribute features, the control mask of the target image at the corresponding scale is determined to accurately locate the pixel points of the features other than the identity feature of the target face carried in the target image; by screening out the target attribute features in the initial attribute features based on the control mask, generating a third feature map based on the target attribute features and the second feature map, and outputting it to the next convolutional layer, after being processed layer by layer through at least one convolutional layer, the attributes and detail features of the target face are effectively retained in the final target face-swapped image, greatly improving the clarity of the face in the face-swapped image, realizing high-definition face swapping, and improving the accuracy of face swapping.

[0162] Figure 8 It is a schematic structural diagram of an image processing device provided by an embodiment of this application. As Figure 7 shown, the device includes

[0163] A feature acquisition module 801, configured to obtain the identity feature of the source image and the initial attribute features of at least one scale of the target image in response to a received face swapping request;

[0164] The face swapping request is used to request swapping the target face in the target image with the source face in the source image. The identity feature characterizes the object to which the source face belongs, and the initial attribute feature characterizes the three-dimensional attributes of the target face;

[0165] The face swapping module 802 is configured to input the identity feature into the generator in the trained face swapping module, and respectively input the initial attribute features of at least one scale into the convolutional layers of corresponding scales in the generator, and output a target face-swapped image, in which the face in the target face-swapped image fuses the identity feature of the source face and the target attribute feature of the target face;

[0166] Wherein, when the face swapping module 802 processes the input identity feature and the initial attribute features of corresponding scales through each convolutional layer of the generator, it includes:

[0167] An obtaining unit, configured to obtain a first feature map output by the previous convolutional layer of the current convolutional layer;

[0168] A generating unit, configured to generate a second feature map based on the identity feature and the first feature map;

[0169] A control mask determining unit, configured to determine a control mask of the target image at the corresponding scale based on the second feature map and the initial attribute feature, where the control mask characterizes the pixel points carrying features other than the identity feature of the target face;

[0170] An attribute screening unit, configured to screen the initial attribute feature based on the control mask to obtain a target attribute feature;

[0171] The generating unit is further configured to generate a third feature map based on the target attribute feature and the second feature map, and output the third feature map to the next convolutional layer of the current convolutional layer as the first feature map of the next convolutional layer.

[0172] In a possible implementation, the generating unit is further configured to perform an affine transformation on the identity feature to obtain a first control vector; based on the first control vector, map the first convolutional kernel of the current convolutional layer to a second convolutional kernel, and perform a convolutional operation on the first feature map based on the second convolutional kernel to generate a second feature map.

[0173] In a possible implementation, the control mask determining unit is configured to perform feature splicing on the second feature map and the initial attribute feature to obtain a spliced feature map; map the spliced feature map to the control mask based on a pre-configured mapping convolutional kernel and activation function.

[0174] In a possible implementation, when the device trains the face swapping model, it further includes:

[0175] A sample acquisition module, configured to acquire a sample data set, where the sample data set includes at least one pair of sample images, and each pair of sample images includes a sample source image and a sample target image;

[0176] A sample feature acquisition module, configured to input the sample data set into an initial model, and acquire the sample identity feature of the sample source image in each pair of sample images, and at least one scale of sample initial attribute features of the sample target image;

[0177] A generation module, configured to, through an initial generator of the initial model, determine at least one scale of sample masks based on the sample identity feature of the sample source image in each pair of sample images and at least one scale of sample initial attribute features of the sample target image, and generate a sample generated image corresponding to each pair of sample images based on the sample identity feature, at least one scale of sample masks, and sample initial attribute features;

[0178] A discrimination module, configured to input the sample source image and the sample generated image in each pair of sample images into an initial discriminator of the initial model, and obtain discrimination results of the initial discriminator for the sample source image and the sample generated image respectively;

[0179] A loss determination module, configured to, for each pair of sample images, determine a first loss value based on at least one scale of sample masks of the sample target image in each pair of sample images, and determine a second loss value based on the discrimination results of the initial discriminator for the sample source image and the sample generated image respectively;

[0180] The loss determination module is further configured to obtain a total training loss based on the first loss value and the second loss value corresponding to the at least one pair of sample images;

[0181] A training module, configured to train the initial model based on the total training loss, and stop training until a target condition is met, to obtain the face swapping model.

[0182] In a possible implementation, the at least one pair of sample images includes at least one pair of first sample images and at least one pair of second sample images, the first sample images include a sample source image and a sample target image belonging to the same object, and the second sample images include a sample source image and a sample target image belonging to different objects;

[0183] The loss determination module is further configured to:

[0184] Acquire a third loss value corresponding to the at least one pair of first sample images based on the sample generated images and the sample target images included in the at least one pair of first sample images;

[0185] Based on the third loss value corresponding to the at least one pair of first sample images, as well as the first loss value and the second loss value corresponding to the at least one pair of sample images, obtain the total training loss.

[0186] In a possible implementation, the initial discriminator includes at least one initial convolutional layer; the loss determination module is further configured to:

[0187] For each pair of sample images, determine a first similarity between the non-face regions of the first discriminative feature map and the non-face regions of the second discriminative feature map, where the first discriminative feature map is the feature map corresponding to the sample target image output by the first part of the at least one initial convolutional layer, and the second discriminative feature map is the feature map corresponding to the sample generated image output by the first part of the at least one initial convolutional layer;

[0188] Determine a second similarity between the third discriminative feature map and the fourth discriminative feature map, where the third discriminative feature map is the feature map corresponding to the sample target image output by the second part of the at least one initial convolutional layer, and the fourth discriminative feature map is the feature map corresponding to the sample generated image output by the second part of the at least one initial convolutional layer;

[0189] Based on the first similarity and the second similarity corresponding to each pair of sample images, determine a fourth loss value corresponding to the at least one pair of sample images;

[0190] Based on the first loss value, the second loss value, and the fourth loss value corresponding to the at least one pair of sample images, obtain the total training loss.

[0191] In a possible implementation, the loss determination module is further configured to:

[0192] For each pair of sample images, respectively extract a first identity feature of the sample source image, a second identity feature of the sample target image, and a third identity feature of the sample generated image;

[0193] Based on the first identity feature and the third identity feature, determine a first identity similarity between the sample source image and the sample generated image;

[0194] Based on the second identity feature and the third identity feature, determine a first identity distance between the sample generated image and the sample target image;

[0195] Based on the first identity feature and the second identity feature, determine a second identity distance between the sample source image and the sample target image;

[0196] Based on the first identity distance and the second identity distance, determine a distance difference;

[0197] Based on the first identity similarity and the distance difference corresponding to each pair of sample images, determine a fifth loss value corresponding to the at least one pair of sample images;

[0198] Based on the first loss value, the second loss value, and the fifth loss value corresponding to the at least one pair of sample images, the total training loss is obtained.

[0199] The image processing device according to the embodiment of the present application obtains the identity feature of the source image and the initial attribute features of at least one scale of the target image; inputs the identity feature into the generator in the trained face swapping module, and inputs the initial attribute features of at least one scale into the convolutional layers corresponding to the scales in the generator respectively to obtain the target face-swapped image; in each convolutional layer of the generator, a second feature map can be generated based on the identity feature and the first feature map output by the previous convolutional layer; and based on the second feature map and the initial attribute features, the control mask of the target image at the corresponding scale is determined to accurately locate the pixel points of the features other than the identity feature of the target face carried in the target image; by screening out the target attribute features in the initial attribute features based on the control mask, a third feature map is generated based on the target attribute features and the second feature map and output to the next convolutional layer, and after being processed layer by layer through at least one convolutional layer, the attributes and detailed features of the target face are effectively retained in the final target face-swapped image, greatly improving the clarity of the face in the face-swapped image, realizing high-definition face swapping, and improving the accuracy of face swapping.

[0200] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed function description of each module of the device, reference can be specifically made to the description in the corresponding method shown above, and details are not described herein again.

[0201] Figure 9 It is a schematic structural diagram of a computer device provided in the embodiment of the present application. As Figure 9 shown, the computer device includes: a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the image processing method. Compared with the related art, it can achieve:

[0202] The image processing method according to the embodiment of the present application includes: obtaining the identity feature of the source image and at least one scale of initial attribute features of the target image; inputting the identity feature into the generator in the trained face swapping module, and respectively inputting the at least one scale of initial attribute features into the convolutional layers corresponding to the scales in the generator to obtain the target face-swapped image; in each convolutional layer of the generator, a second feature map can be generated based on the identity feature and the first feature map output by the previous convolutional layer; and based on the second feature map and the initial attribute features, the control mask of the target image at the corresponding scale is determined to accurately locate the pixel points of the features other than the identity feature of the target face carried in the target image; by screening out the target attribute features in the initial attribute features based on the control mask, generating a third feature map based on the target attribute features and the second feature map, and outputting it to the next convolutional layer, and through the processing of at least one convolutional layer, the attributes and detail features of the target face are effectively retained in the final target face-swapped image, greatly improving the clarity of the face in the face-swapped image, realizing high-definition face swapping, and improving the accuracy of face swapping.

[0203] In an alternative embodiment, a computer device is provided, as Figure 9 shown Figure 9 The computer device 900 shown in the figure includes: a processor 901 and a memory 903. Among them, the processor 901 and the memory 903 are connected, such as through a bus 902. Optionally, the computer device 900 may further include a transceiver 904, and the transceiver 904 may be used for data interaction between the computer device and other computer devices, such as data sending and / or data receiving, etc. It should be noted that in actual applications, the transceiver 904 is not limited to one, and the structure of the computer device 900 does not constitute a limitation to the embodiment of the present application.

[0204] The processor 901 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure of the present application. The processor 901 may also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0205] The bus 902 may include a path for transmitting information among the above components. The bus 902 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 902 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 it is only represented by a thick line in Figure 9 , but it does not mean that there is only one bus or one type of bus.

[0206] The memory 903 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media \ other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited here.

[0207] The memory 903 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 901 to execute. The processor 901 is used to execute the computer program stored in the memory 903 to implement the steps shown in the foregoing method embodiments.

[0208] Among them, the electronic device includes but is not limited to: servers, terminals, or cloud computing center devices, etc.

[0209] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0210] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0211] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. The terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, but do not exclude the implementation of other features, information, data, steps, operations, etc. supported by the technical field of the present application.

[0212] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than the illustrated or described order.

[0213] It should be understood that although the flowchart of the embodiments of the present application indicates various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless clearly stated herein, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0214] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, adopting other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.

Claims

1. An image processing method, characterized in that, The method includes: In response to a received face swapping request, obtaining the identity feature of the source image and the initial attribute features of at least one scale of the target image; The face swapping request is used to request to swap the target face in the target image with the source face in the source image. The identity feature represents the object to which the source face belongs, and the initial attribute feature represents the three-dimensional attributes of the target face; Inputting the identity feature into the generator in the trained face swapping module, and respectively inputting the initial attribute features of at least one scale into the convolutional layers corresponding to the scales in the generator, and outputting a target face swapped image, in which the face in the target face swapped image fuses the identity feature of the source face and the target attribute feature of the target face; Wherein, through each convolutional layer of the generator, the following steps are performed on the input identity feature and the initial attribute features of the corresponding scale: Obtaining a first feature map output by the previous convolutional layer of the current convolutional layer; Based on the identity feature and the first feature map, generating a second feature map, and based on the second feature map and the initial attribute feature, determining a control mask of the target image at the corresponding scale, where the control mask represents pixel points carrying features other than the identity feature of the target face; Based on the control mask, screening the initial attribute features to obtain target attribute features; Based on the target attribute feature and the second feature map, generating a third feature map, and outputting the third feature map to the next convolutional layer of the current convolutional layer as the first feature map of the next convolutional layer.

2. The method according to claim 1, wherein The generating a second feature map based on the identity feature and the first feature map includes: Performing an affine transformation on the identity feature to obtain a first control vector; Based on the first control vector, mapping the first convolutional kernel of the current convolutional layer to a second convolutional kernel, and performing a convolutional operation on the first feature map based on the second convolutional kernel to generate a second feature map.

3. The method according to claim 1, wherein The determining a control mask of the target image at the corresponding scale based on the second feature map and the initial attribute feature includes: Performing feature splicing on the second feature map and the initial attribute feature to obtain a spliced feature map; Based on a pre-configured mapping convolutional kernel and activation function, mapping the spliced feature map to the control mask.

4. The method according to claim 1, wherein The training method of the face swapping model includes: Obtaining a sample data set, where the sample data set includes at least one pair of sample images, and each pair of sample images includes a sample source image and a sample target image; Inputting the sample data set into the initial model, and obtaining the sample identity feature of the sample source image in each pair of sample images and the sample initial attribute features of at least one scale of the sample target image; Through the initial generator of the initial model, based on the sample identity feature of the sample source image in each pair of sample images and the sample initial attribute features of at least one scale of the sample target image, determining sample masks of at least one scale, and based on the sample identity feature, the sample masks of at least one scale and the sample initial attribute features, generating a sample generated image corresponding to each pair of sample images; Input the sample source image and the sample generated image in each pair of sample images into the initial discriminator of the initial model, and obtain the discrimination results of the initial discriminator for the sample source image and the sample generated image respectively; For each pair of sample images, determine a first loss value based on the sample masks of at least one scale of the sample target image in the pair of sample images, and determine a second loss value based on the discrimination results of the initial discriminator for the sample source image and the sample generated image respectively; Obtain the total training loss based on the first loss value and the second loss value corresponding to the at least one pair of sample images; Train the initial model based on the total training loss until the training stops when the target conditions are met, and obtain the face swapping model.

5. The method according to claim 4, wherein The at least one pair of sample images includes at least one pair of first sample images and at least one pair of second sample images. The first sample images include a sample source image and a sample target image belonging to the same object, and the second sample images include a sample source image and a sample target image belonging to different objects; The obtaining the total training loss based on the first loss value and the second loss value corresponding to the at least one pair of sample images includes: Obtain a third loss value corresponding to the at least one pair of first sample images based on the sample generated images and the sample target images included in the at least one pair of first sample images; Obtain the total training loss based on the third loss value corresponding to the at least one pair of first sample images, and the first loss value and the second loss value corresponding to the at least one pair of sample images.

6. The method according to claim 4, characterized in that, The initial discriminator includes at least one initial convolutional layer; The obtaining the total training loss based on the first loss value and the second loss value corresponding to the at least one pair of sample images includes: For each pair of sample images, determine a first similarity between the non-face regions of the first discriminative feature map and the non-face regions of the second discriminative feature map. The first discriminative feature map is the feature map corresponding to the sample target image output by the first part of the at least one initial convolutional layer, and the second discriminative feature map is the feature map corresponding to the sample generated image output by the first part of the initial convolutional layer; Determine a second similarity between the third discriminative feature map and the fourth discriminative feature map. The third discriminative feature map is the feature map corresponding to the sample target image output by the second part of the at least one initial convolutional layer, and the fourth discriminative feature map is the feature map corresponding to the sample generated image output by the second part of the initial convolutional layer; Determine a fourth loss value corresponding to the at least one pair of sample images based on the first similarity and the second similarity corresponding to each pair of sample images; Obtain the total training loss based on the first loss value, the second loss value, and the fourth loss value corresponding to the at least one pair of sample images.

7. The method according to claim 4, characterized in that The obtaining the total training loss based on the first loss value and the second loss value corresponding to the at least one pair of sample images includes: For each pair of sample images, extract the first identity feature of the sample source image, the second identity feature of the sample target image, and the third identity feature of the sample generated image respectively; Determine a first identity similarity between the sample source image and the sample generated image based on the first identity feature and the third identity feature; Determine a first identity distance between the sample generated image and the sample target image based on the second identity feature and the third identity feature; Determine a second identity distance between the sample source image and the sample target image based on the first identity feature and the second identity feature; Determine a distance difference based on the first identity distance and the second identity distance; Determine a fifth loss value corresponding to at least one pair of sample images based on the first identity similarity and the distance difference corresponding to each pair of sample images; Obtain the total training loss based on the first loss value, the second loss value, and the fifth loss value corresponding to the at least one pair of sample images; 8. An image processing apparatus, characterized in that, The device includes: A feature acquisition module, configured to acquire an identity feature of a source image and initial attribute features of at least one scale of a target image in response to a received face swapping request; The face swapping request is used to request swapping the target face in the target image with the source face in the source image, the identity feature represents the object to which the source face belongs, and the initial attribute feature represents the three-dimensional attributes of the target face; A face swapping module, configured to input the identity feature into a generator in a trained face swapping module, and respectively input the initial attribute features of the at least one scale into convolutional layers corresponding to the scales in the generator, and output a target face swapped image, where the face in the target face swapped image fuses the identity feature of the source face and the target attribute feature of the target face; Wherein, when the face swapping module processes the input identity feature and the initial attribute features of the corresponding scale through each convolutional layer of the generator, it includes: An acquisition unit, configured to acquire a first feature map output by the previous convolutional layer of the current convolutional layer; A generation unit, configured to generate a second feature map based on the identity feature and the first feature map; A control mask determination unit, configured to determine a control mask of the target image at the corresponding scale based on the second feature map and the initial attribute feature, where the control mask represents pixel points carrying features other than the identity feature of the target face; An attribute screening unit, configured to screen the initial attribute features based on the control mask to obtain target attribute features; The generation unit is further configured to generate a third feature map based on the target attribute feature and the second feature map, and output the third feature map to the next convolutional layer of the current convolutional layer as the first feature map of the next convolutional layer.

9. The device according to claim 8, characterized in that, The generation unit is further configured to perform an affine transformation on the identity feature to obtain a first control vector; based on the first control vector, map the first convolutional kernel of the current convolutional layer to a second convolutional kernel, and perform a convolution operation on the first feature map based on the second convolutional kernel to generate a second feature map.

10. The device according to claim 8, wherein The control mask determination unit is configured to perform feature splicing on the second feature map and the initial attribute feature to obtain a spliced feature map; map the spliced feature map to the control mask based on a pre-configured mapping convolutional kernel and activation function.

11. The device according to claim 8, wherein When training the face-swapping model, the device further includes: A sample acquisition module, configured to acquire a sample data set, where the sample data set includes at least one pair of sample images, and each pair of sample images includes a sample source image and a sample target image; A sample feature acquisition module, configured to input the sample data set into an initial model to acquire the sample identity features of the sample source image in each pair of sample images, and at least one scale of sample initial attribute features of the sample target image; A generation module, configured to, through an initial generator of the initial model, determine at least one scale of sample masks based on the sample identity features of the sample source image in each pair of sample images and at least one scale of sample initial attribute features of the sample target image, and generate a sample generated image corresponding to each pair of sample images based on the sample identity features, at least one scale of sample masks, and sample initial attribute features; A discrimination module, configured to input the sample source image and the sample generated image in each pair of sample images into an initial discriminator of the initial model to obtain the discrimination results of the initial discriminator for the sample source image and the sample generated image respectively; A loss determination module, configured to, for each pair of sample images, determine a first loss value based on at least one scale of sample masks of the sample target image in each pair of sample images, and determine a second loss value based on the discrimination results of the initial discriminator for the sample source image and the sample generated image respectively; The loss determination module is further configured to obtain a total training loss based on the first loss value and the second loss value corresponding to the at least one pair of sample images; A training module, configured to train the initial model based on the total training loss until the training stops when the target conditions are met, and obtain the face-swapping model.

12. The device according to claim 11, characterized in that, The at least one pair of sample images includes at least one pair of first sample images and at least one pair of second sample images. The first sample images include a sample source image and a sample target image belonging to the same object, and the second sample images include a sample source image and a sample target image belonging to different objects; The loss determination module is further configured to: Acquire a third loss value corresponding to the at least one pair of first sample images based on the sample generated images and the sample target images included in the at least one pair of first sample images; Obtain the total training loss based on the third loss value corresponding to the at least one pair of first sample images, and the first loss value and the second loss value corresponding to the at least one pair of sample images.

13. The device according to claim 11, wherein The initial discriminator includes at least one initial convolutional layer; the loss determination module is further configured to: For each pair of sample images, determine a first similarity between the non-face regions of a first discrimination feature map and the non-face regions of a second discrimination feature map. The first discrimination feature map is a feature map corresponding to the sample target image output by a first part of the at least one initial convolutional layer, and the second discrimination feature map is a feature map corresponding to the sample generated image output by the first part of the initial convolutional layer; Determine a second similarity between a third discriminative feature map and a fourth discriminative feature map, where the third discriminative feature map is a feature map corresponding to a sample target image output by a second part of the at least one initial convolutional layer, and the fourth discriminative feature map is a feature map corresponding to a sample generated image output by the second part of the initial convolutional layer; Based on the first similarity and the second similarity corresponding to each pair of sample images, determine a fourth loss value corresponding to at least one pair of sample images; Based on the first loss value, the second loss value, and the fourth loss value corresponding to the at least one pair of sample images, obtain the total training loss.

14. The device according to claim 11, wherein The loss determination module is further configured to: For each pair of sample images, respectively extract a first identity feature of the sample source image, a second identity feature of the sample target image, and a third identity feature of the sample generated image; Based on the first identity feature and the third identity feature, determine a first identity similarity between the sample source image and the sample generated image; Based on the second identity feature and the third identity feature, determine a first identity distance between the sample generated image and the sample target image; Based on the first identity feature and the second identity feature, determine a second identity distance between the sample source image and the sample target image; Based on the first identity distance and the second identity distance, determine a distance difference; Based on the first identity similarity and the distance difference corresponding to each pair of sample images, determine a fifth loss value corresponding to at least one pair of sample images; Based on the first loss value, the second loss value, and the fifth loss value corresponding to the at least one pair of sample images, obtain the total training loss.

15. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the image processing method according to any one of claims 1 to 7.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the image processing method according to any one of claims 1 to 7.

17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the image processing method according to any one of claims 1 to 7.