Image processing method, device, computer equipment, storage medium and program product
By obtaining the attribute parameters of the image to be changed and the attribute parameters of the target face, combined with the adaptive instance regularization method, more accurate fusion encoding features are generated, which solves the problem of unnatural face deformation in the face change image and improves the accuracy of face change.
Patent Information
- Application Number
- CN202210334052.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2042-03-30
AI Technical Summary
In the prior art, when the posture difference is large during the face change process, the face deformation in the face change image is unnatural, the similarity is low, and the accuracy is low.
By obtaining the attribute parameters of the image to be changed, combining the pre-stored target face attribute parameters, the target comprehensive features are determined, and migrating to the image encoding features through the adaptive instance regularization method, fusion encoding features are generated, and finally decoding is obtained for the fused face image.
It improves the sensory similarity between the fused face and the target face in the face change image, and enhances the accuracy of the face change.
Smart Images

Figure CN114972010B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to technical fields such as artificial intelligence and computer vision, and relates to an image processing method, apparatus, computer equipment, storage medium, and program product. Background Art
[0002] Face swapping is a key technology in the field of computer vision and is widely used in content production, film and television portrait creation, and entertainment video production. Face swapping involves taking two images, A and B, and transferring the facial features from image A to image B to create a swapped image.
[0003] In related technologies, image processing methods may include: usually achieving face-changing based on shape fitting; for example, based on the detected facial key points in image A and the facial key points in image B, the shape change relationship between the two images regarding facial features, contours, etc. can be calculated, and the faces in image A and image B are fused according to the shape transformation relationship to obtain a face-changing image.
[0004] The aforementioned shape fitting process achieves face swapping by deforming and fusing the face. However, when there are significant differences in posture between Image A and Image B, simple shape fitting cannot meet the requirements of face swapping. This can easily lead to unnatural facial deformation in the swapped image, resulting in a low similarity between the swapped image and the face in Image A, and thus low accuracy in face swapping. Summary of the Invention
[0005] This application provides an image processing method, apparatus, computer device, storage medium, and program product that can solve the problems of low facial similarity in face-swapped images and low accuracy of face-swapped methods in related technologies. The technical solution is as follows:
[0006] In one aspect, an image processing method is provided, the method comprising:
[0007] In response to a received face-swapping request, obtaining attribute parameters of the image to be face-swapping, wherein the face-swapping request is used to request that a face in the image to be face-swapping be replaced with a target face, and the attribute parameters of the image to be face-swapping indicate three-dimensional attributes of the face in the image to be face-swapping;
[0008] Determining target attribute parameters based on the attribute parameters of the image to be face-swapped and pre-stored attribute parameters of the target face;
[0009] Determining target comprehensive features based on the target attribute parameters and pre-stored facial features of the target face;
[0010] Encoding the image to be face-swapped to obtain image coding features of the image to be face-swapped;
[0011] Migrating the target comprehensive features to the image coding features of the face-swapped image through an adaptive instance regularization method to obtain a fused coding feature;
[0012] The fusion coding feature is decoded to obtain a target face-swapped image including a fused face, where the fused face is a fusion of the face in the image to be face-swapped and the target face.
[0013] In another aspect, an image processing apparatus is provided, the apparatus comprising:
[0014] an attribute parameter acquisition module, configured to acquire attribute parameters of the image to be face-swapped in response to a received face-swap request, wherein the face-swap request is used to request that the face in the image to be face-swapped be replaced with a target face, and the attribute parameters of the image to be face-swapped indicate three-dimensional attributes of the face in the image to be face-swapped;
[0015] a target attribute parameter determination module, configured to determine target attribute parameters based on the attribute parameters of the image to be face-swapped and pre-stored attribute parameters of the target face;
[0016] a comprehensive feature determination module, configured to determine a target comprehensive feature based on the target attribute parameters and pre-stored facial features of the target face;
[0017] An encoding module, configured to encode the image to be face-swapped to obtain image encoding features of the image to be face-swapped;
[0018] A migration module, configured to migrate the target comprehensive features to the image coding features of the face-swapped image through an adaptive instance regularization method to obtain a fused coding feature;
[0019] A decoding module is used to decode the fusion coding features to obtain a target face-changing image including a fused face, where the fused face is a fusion of the face in the image to be face-changed and the target face.
[0020] In one possible implementation, the attribute parameter of the image includes at least one of a shape coefficient, an expression coefficient, an angle coefficient, a texture coefficient, or an illumination coefficient;
[0021] The target attribute parameter determination module is used to determine the shape coefficient of the target face and the pre-configured parameters of the image to be face-swapped as the target attribute parameters, where the pre-configured parameters include at least one of an expression coefficient, an angle coefficient, a texture coefficient or an illumination coefficient.
[0022] In one possible implementation, the migration module is configured to:
[0023] Obtaining a mean value and a standard deviation of the image coding feature in at least one feature channel, and obtaining a mean value and a standard deviation of the target comprehensive feature in at least one feature channel;
[0024] The mean and standard deviation of the image coding feature in each feature channel are aligned with the mean and standard deviation of the target comprehensive feature in the corresponding feature channel to obtain the fused coding feature.
[0025] In one possible implementation, the target face-swap image is obtained by a pre-trained face-swap model; the face-swap model is used to swap the target face into any facial image based on pre-stored attribute data and facial features of the target face;
[0026] The device further includes a model training module, which includes:
[0027] an acquiring unit, configured to acquire facial features and attribute parameters of a first sample image, and acquire attribute parameters of a second sample image, wherein the first sample image includes the target face and the second sample image includes the face to be replaced;
[0028] a sample attribute parameter determining unit, configured to determine sample attribute parameters based on the attribute parameters of the first sample image and the attribute parameters of the second sample image, wherein the sample attribute parameters are used to indicate desired attributes of the face in the sample face-swapped image to be generated;
[0029] a sample comprehensive feature acquisition unit, configured to determine a sample comprehensive feature based on the sample attribute parameters and the facial features of the first sample image;
[0030] an encoding unit, configured to input the second sample image into an encoder of an initial model for encoding to obtain a sample encoding feature;
[0031] a migration unit, configured to migrate the sample comprehensive features to the sample encoding features of the second sample image through an adaptive instance regularization method to obtain a sample fusion feature;
[0032] A decoding unit, configured to input the sample fusion features into a decoder of the initial network for decoding to obtain a sample face-swapped image;
[0033] A training unit is used to determine the total loss of the initial model based on the differences between the sample face-swapped image and the sample attribute parameters, the facial features of the first sample image, and the second sample image, and to train the initial model based on the total loss until the training is stopped when the target conditions are met, thereby obtaining the face-swapped model.
[0034] In one possible implementation, the training unit is specifically configured to:
[0035] Obtaining a first similarity between the attribute parameters of the sample face-swapped image and the attribute parameters of the sample;
[0036] Obtaining a second similarity between facial features of the sample face-swapped image and facial features of the first sample image;
[0037] obtaining, by the discriminator of the initial network, a third similarity between the second sample image and the sample face-swapped image;
[0038] The total loss is determined based on the first similarity, the second similarity, and the third similarity.
[0039] In one possible implementation, the training unit is further configured to:
[0040] Inputting the second sample image as a real image into the discriminator, and inputting the sample face-swapped image into the discriminator;
[0041] Obtaining, by the discriminator, a scaled image of the second sample image at at least one scale and a scaled image of the sample face-swapped image at the corresponding at least one scale;
[0042] Obtaining a discrimination probability corresponding to at least one scaled image of the second sample image, and obtaining a discrimination probability corresponding to at least one scaled image of the sample face-swapped image, wherein the discrimination probability of an image is used to indicate a probability that the image is a real image;
[0043] The third similarity is determined based on at least one discrimination probability corresponding to the second sample image and at least one discrimination probability corresponding to the sample face-swapped image.
[0044] In one possible implementation, the acquiring unit is specifically configured to:
[0045] Acquire at least two posture images as the first sample images, wherein the at least two posture images include at least two facial postures of the target face;
[0046] Based on the at least two posture images, acquiring facial features and attribute parameters corresponding to the at least two facial postures;
[0047] using an average of the facial features corresponding to the at least two facial postures as the facial features of the first sample image, and using an average of the attribute parameters corresponding to the at least two facial postures as the attribute parameters of the first sample image;
[0048] Correspondingly, the device further includes a storage unit, which is used to store the facial features and attribute parameters of the first sample image.
[0049] In one possible implementation, the acquiring unit is further configured to:
[0050] Performing facial recognition on at least two image frames included in the video of the target object to obtain at least two image frames including the target face, where the target face is the face of the target object;
[0051] Perform face cropping on the at least two image frames to obtain the at least two posture images, and use the at least two posture images as the first sample images.
[0052] On the other hand, a computer device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above-mentioned image processing method.
[0053] On the other hand, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the image processing method described above is implemented.
[0054] On the other hand, a computer program product is provided, comprising a computer program, wherein the computer program implements the above-mentioned image processing method when executed by a processor.
[0055] The beneficial effects of the technical solution provided by the embodiments of the present application are:
[0056] The image processing method provided by the present application obtains attribute parameters of the image to be replaced, which are used to indicate the three-dimensional attributes of the face in the image, and determines target attribute parameters based on the attribute parameters of the image to be replaced and the attribute parameters of the pre-stored target face, thereby locating the three-dimensional attribute features of the face in the image to be generated; and, based on the target attribute parameters and the facial features of the pre-stored target face, obtains a target comprehensive feature that can comprehensively characterize the image to be replaced and the target face; and encodes the image to be replaced to obtain the image coding feature of the image to be replaced, thereby obtaining the detailed features of the image to be replaced at the pixel level through the image coding feature; and further The features are migrated to the image coding features of the image to be face-changed through an adaptive instance regularization method to obtain a fused coding feature. The present application combines the coding features refined to the pixel level with the global comprehensive features, and aligns the feature distribution of the image coding features with the target comprehensive features, thereby improving the accuracy of the generated fused coding features; by decoding the fused coding features, a target face-changed image including the fused face is obtained, so that the decoded image can be refined to each pixel point to show the target comprehensive features, so that the sensory perception of the fused face in the decoded image is closer to the target face, thereby improving the sensory similarity between the fused face and the target face, thereby improving the accuracy of face-changing. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0058] Figure 1 A schematic diagram of an implementation environment for an image processing method provided in an embodiment of the present application;
[0059] Figure 2 A flowchart of a face-swapping model training method provided in an embodiment of the present application;
[0060] Figure 3 A schematic diagram of the training process framework of a face-swapping model provided in an embodiment of the present application;
[0061] Figure 4 A signaling interaction diagram of an image processing method provided in an embodiment of the present application;
[0062] Figure 5 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application;
[0063] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0064] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0065] It is understandable that in the specific implementation of the present application, the facial images involved, such as the first sample image, the second sample image, the posture image, the video of the target object, and any other object-related data used in the face-changing model training, as well as the image to be changed, the facial features of the target face, the attribute parameters, and any other object-related data used when changing faces using the face-changing model, are all obtained after the consent or permission of the relevant object; when the following embodiments of the present application are applied to specific products or technologies, it is necessary to obtain the permission or consent of the object, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. In addition, the face-changing process of any object's facial image using the image processing method of the present application is based on the face-changing service or face-changing request triggered by the relevant object, and is performed after the permission or consent of the relevant object.
[0066] The following is an introduction to the technical terms involved in this application:
[0067] Face swapping: swapping the target face in one face image to the face in another image.
[0068] Face-changing model: used to replace the target face into any facial image based on the pre-stored attribute data and facial features of the target face; the image processing method provided in this application can use the face-changing model to replace the face in the image to be replaced with an exclusive target face.
[0069] Image to be replaced: an image of the face to be replaced, for example, the target face can be replaced with the face in the image to be replaced; it should be noted that the image to be replaced is replaced by the image processing method of the present application to obtain a target face-changing image, and the fused face included in the target face-changing image is a fusion of the face in the image to be replaced and the target face, and the sensory similarity between the fused face and the target face is higher. The fused face also integrates the expression, angle and other postures of the face in the image to be replaced, thereby making the target face image more vivid and realistic.
[0070] Attribute parameters: The attribute parameters of an image are used to indicate the three-dimensional attributes of the face in the image, and can represent attributes such as the posture and spatial environment of the face in three-dimensional space.
[0071] Facial features: Characterize the two-dimensional features of the face in the image, such as the distance between the eyes and the size of the nose; facial features can represent the identity of the object with these facial features.
[0072] Target face: A unique face used to replace the face in the image; this application provides a face-swapping service using the target face as the unique face, that is, the unique target face can be swapped to any facial image; for example, target face A can be swapped to the face in image B, and A can also be swapped to the face in image C.
[0073] First sample image: The first sample image includes the target face and is the image used in the face-changing model training.
[0074] Second sample image: This second sample image contains the face to be replaced and is used when training the face-swapping model. During training, the target face in the first sample image is used as the exclusive face. The target face in the first sample image is then swapped into the second sample image. This process is used to train the face-swapping model.
[0075] Figure 1 This is a schematic diagram of an implementation environment of an image processing method provided by this application. Figure 1 As shown, the implementation environment includes: a server 11 and a terminal 12.
[0076] The server 11 is configured with a pre-trained face-swapping model and can provide a face-swapping function to the terminal 12 based on this face-swapping model. This face-swapping service involves swapping the face in the target face image with the target face, so that the fused face in the generated target face image can be a fusion of the original face in the image and the target face. In one possible scenario, the terminal 12 can send a face-swapping request to the server 11. This face-swapping request can include the target face-swapping image. Based on the face-swapping request, the server 11 can execute the image processing method of the present application to generate the target face-swapping image and return the target face-swapping image to the terminal 12. In one example scenario, the server 11 can be the backend server of an application. The terminal 12 has an application installed, and the terminal 12 and the server 11 can exchange data based on this application to implement the face-swapping process. The application can be configured with the face-swapping function. The application can be any application that supports face-swapping, including, but not limited to, video editing applications, image processing tools, video applications, live streaming applications, social applications, content interaction platforms, gaming applications, and the like.
[0077] The server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The aforementioned networks may include, but are not limited to, wired networks and wireless networks. Wired networks include local area networks, metropolitan area networks, and wide area networks, and wireless networks include Bluetooth, Wi-Fi, and other networks that enable wireless communication. Terminals may include smartphones (such as Android phones and iOS phones), tablets, laptops, digital broadcast receivers, MIDs (Mobile Internet Devices), PDAs (Personal Digital Assistants), desktop computers, in-vehicle terminals (such as in-vehicle navigation terminals, in-vehicle computers, etc.), smart appliances, aircraft, smart speakers, smart watches, etc. The terminals and servers may be connected directly or indirectly via wired or wireless communications, but are not limited thereto. The specific configuration may also be determined based on the actual application scenario requirements and is not limited here.
[0078] The image processing methods provided in the embodiments of this application involve the following exemplary technologies, such as artificial intelligence and computer vision. For example, they utilize cloud computing and big data processing within artificial intelligence to implement processes such as extracting attribute parameters from a first sample image and training a face-swapping model. For example, computer vision technology is used to perform facial recognition on image frames in a video to crop a first sample image that includes the target face.
[0079] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also encompasses the study of the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0080] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and smart transportation.
[0081] Computer vision (CV) is the science of making machines "see." Specifically, it refers to machine vision, where cameras and computers replace the human eye in identifying and measuring objects, performing further image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, and smart transportation. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0082] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0083] Figure 2 This is a flow chart of a face-changing module training method provided in an embodiment of the present application. The execution subject of this method can be a computer device. Figure 2 As shown, the method includes the following steps.
[0084] Step 201: The computer device obtains facial features and attribute parameters of a first sample image, and obtains attribute parameters of a second sample image.
[0085] The first sample image includes the target face, and the second sample image includes the face to be replaced. The computer device can collect data including any face as the second sample image; and collect images including the target face at various poses and angles as the first sample image. The computer device can obtain attribute parameters of the first sample image and attribute parameters of the second sample image using a facial parameter estimation model. The computer device can obtain facial features of the first sample image using a facial recognition model.
[0086] The facial parameter estimation model is used to estimate the three-dimensional attribute parameters of the face based on the input two-dimensional facial image. The facial parameter estimation model can be a model with a CNN network structure. For example, the facial parameter estimation model can be a 3DMM (3D Morphable models, three-dimensional deformable face model). The present application can regress the three-dimensional attribute parameters of the input two-dimensional facial image through the ResNet (Residual Network, residual network) part in the 3DMM. Of course, the facial parameter estimation model can also be any other model with the function of extracting the three-dimensional attribute parameters of the face in the two-dimensional image. Here, only the 3DMM model is used as an example.
[0087] This attribute parameter indicates the three-dimensional attributes of the face in the image, representing the face's posture and spatial environment in three-dimensional space. The attribute parameters include, but are not limited to, shape coefficient (id_coeff), expression coefficient (expression_coeff), texture coefficient (texture_coeff), angle coefficient (angles_coeff), and illumination coefficient (gamma_coeff). The shape coefficient represents the shape of the face and the shape of the facial features; the angle coefficient represents the face's pitch angle, left and right yaw angle, and other angles; the texture coefficient represents the condition of the face's skin and hair; and the illumination coefficient represents the lighting conditions of the face's surroundings in the image.
[0088] In this application, the computer device may extract one or more of the shape coefficient, expression coefficient, texture coefficient, angle coefficient, and illumination coefficient as attribute parameters of each sample image, or may extract all of the coefficients as attribute parameters of the corresponding sample image. Accordingly, the attribute parameters of the first sample image and the second sample image may be obtained in the following three ways.
[0089] Method 1: The computer device extracts the shape coefficient of the target face in the first sample image as the attribute parameter of the first sample image, and the computer device extracts the expression coefficient and the angle coefficient in the second sample image as the attribute parameter of the second sample image.
[0090] In method 1, the attribute parameters of the first sample image include the shape coefficient of the target face in the first sample image. The attribute parameters of the second sample image include the expression coefficient and angle coefficient of the face in the second sample image. By obtaining the shape coefficient of the first sample image and the expression coefficient and angle coefficient of the second sample image, the shape features of the target face and the expression, angle, and other features of the face to be replaced can be subsequently fused. This ensures that the face in the fused sample face-swap image has the facial features of the target face, as well as the expression, angle, and other features of the face to be replaced, thereby improving the similarity in facial features between the fused face and the target face.
[0091] Method 2: For the second sample image, the computer device may obtain pre-configured parameters of the second sample image as attribute parameters of the second sample image. For the first sample image, the computer device may extract shape coefficients of the target face in the first sample image as attribute parameters of the first sample image.
[0092] In Method 2, the computer device configures which attributes of the second sample image may be included based on need. That is, the attribute parameters of the second sample image may include pre-configured parameters. In one example, the pre-configured parameters may include at least one of an expression coefficient, a texture coefficient, an angle coefficient, and an illumination coefficient. The pre-configured parameters are pre-configured based on need. For example, by configuring the pre-configured parameters to include the illumination coefficient and the expression coefficient, the resulting fused face can have the characteristics of the surrounding environment, such as illumination and expression, of the face to be replaced. Of course, the pre-configured parameters may also be configured to include texture coefficients, angle coefficients, etc., which will not be detailed here.
[0093] Mode 3: The computer device may also extract multiple parameters of the first sample image and the second sample image as corresponding attribute parameters, and further extract required parameters from the multiple parameters in subsequent steps.
[0094] In one example, the attribute parameters of the first sample image may include five parameters: the shape coefficient, expression coefficient, texture coefficient, angle coefficient, and illumination coefficient of the target face in the first sample image. For example, if the attribute parameters can be represented by a vector, then when the attribute parameters of the first sample image include the above five parameters, the attribute parameters of the first sample image can be represented as a 257-dimensional feature vector. The attribute parameters of the second sample image may also include the five parameters: the shape coefficient, expression coefficient, texture coefficient, angle coefficient, and illumination coefficient of the second sample image. Accordingly, the attribute parameters of the second sample image can also be represented as a 257-dimensional feature vector.
[0095] In one possible implementation, the computer device may obtain pose images of the target face at multiple pose angles and extract facial features and attribute parameters of a first sample image based on the multiple pose images. Exemplarily, the process of the computer device obtaining facial features and attribute parameters of the first sample image may include: the computer device obtaining at least two pose images as the first sample images, the at least two pose images including at least two facial poses of the target face; the computer device obtaining facial features and attribute parameters corresponding to the at least two facial poses based on the at least two pose images; the computer device using the average of the facial features corresponding to the at least two facial poses as the facial features of the first sample image, and the average of the attribute parameters corresponding to the at least two facial poses as the attribute parameters of the first sample image. In another example, the computer device may input the at least two pose images into a facial parameter estimation model, extract attribute parameters of each pose image using the facial parameter model, calculate the average of the attribute parameters of the at least two pose images, and use the average of the attribute parameters of the at least two pose images as the attribute parameters of the first sample image. In one example, the computer device may input the at least two pose images into a facial recognition model, extract facial features of each pose image in a two-dimensional plane through the facial recognition model, calculate the average of the facial features of the at least two pose images, and use the average of the facial features of the at least two pose images as the facial features of the first sample image. For example, the facial features of the first sample image may be a 512-dimensional feature vector. The facial features represent the identity of the target object, and the target face is the face of the target object.
[0096] In one possible example, the computer device may extract multiple pose images including the target face from a video. Exemplarily, the step of the computer device obtaining the at least two pose images as first sample images includes: performing facial recognition on at least two image frames included in a video of the target subject to obtain at least two image frames including the target face, wherein the target subject's face is the target face; performing facial cropping on the at least two image frames to obtain the at least two pose images, and using the at least two pose images as the first sample images. Facial poses may include, but are not limited to, any attribute such as facial expression, angle, shape of facial features, movement, glasses, and facial makeup. The computer device may distinguish poses based on any of these attributes. For example, a smiling face and an angry face may be considered as two faces with different poses; a face with glasses and a face without glasses may also be considered as two faces with different poses; and a face with eyes closed and tilted 45° upward or tilted 30° downward or tilted 30° downward or tilted 30° downward may also be considered as two faces with different poses. In another possible example, the computer device may also obtain multiple independent still images of the target face and extract the multiple pose images from the multiple independent still images. The computer device may also perform face cropping on the multiple still images to obtain the at least two pose images, and use the at least two pose images as the first sample images.
[0097] In a possible technical implementation, the computer device may perform face cropping on the image frame to obtain the posture image through the process of the following steps a to c.
[0098] Step a: The computer device performs face detection on the image frame to obtain a face coordinate frame of the image frame.
[0099] The facial coordinate frame encircles the facial region where the target face is located in the image frame.
[0100] Step b: The computer device performs facial registration on the image frame according to the facial coordinate frame of the image frame to obtain target facial key points in the image frame.
[0101] The target facial key points may include facial feature key points, facial contour key points, and of course hair key points of the target face in the image frame.
[0102] It should be noted that the computer device can perform face configuration on the image frame through a face configuration algorithm. The input information of the face configuration algorithm is the facial image and the facial coordinate frame, and the output information is a facial key point coordinate sequence including the target facial key points. The number of key points included in the facial key point coordinate sequence can be pre-configured based on needs. For example, the number of key points included in the facial key point coordinate sequence can be a fixed value such as 5 points, 68 points, or 90 points.
[0103] Step c: The computer device performs facial cropping on the image frame based on the target facial key points to obtain the posture image.
[0104] In one example, the computer device may also use the same process for acquiring the second sample image as for acquiring the first sample image. For example, the computer device may acquire an object image including an arbitrary object, perform facial cropping on the object image, obtain an image including the face of the object, and use the cropped image as the second sample image. The facial cropping method is similar to steps a to c of performing facial cropping on an image frame to obtain a posture image, and will not be further described here. Furthermore, the computer device may input the second sample image into a facial parameter estimation model, and extract attribute parameters of the second sample image through the facial parameter estimation model.
[0105] In one possible implementation, the computer device may store the facial features and attribute parameters of the first sample image; the process may include: the computer device storing the facial features and attribute parameters of the first sample image to a target address. By fixedly storing the facial features and attribute parameters of the target face, it is convenient to directly extract data from the target address during subsequent use; for example, when using a trained face-swapping model to provide an exclusive face-swapping service, the fixed storage method allows the computer device to directly extract the stored facial features and attribute parameters of the target face, thereby implementing an exclusive face-swapping process in which the exclusive target face is replaced with any facial image; for another example, during the iterative training phase, the facial features and attribute parameters of the target face can be directly extracted from the target address for training.
[0106] Step 202: The computer device determines a sample attribute parameter based on the attribute parameter of the first sample image and the attribute parameter of the second sample image.
[0107] The sample attribute parameter is used to indicate the expected attributes of the face in the sample face-swapped image to be generated.
[0108] Corresponding to method 1 in step 201 , in this step, the computer device may determine the shape coefficient of the first sample image and the expression coefficient and angle coefficient of the second sample image as the sample attribute parameters.
[0109] Corresponding to Method 2 and Method 3 in step 201, in this step, the computer device may select various attribute parameters of the first sample image and the second sample image as sample attribute parameters based on need. Step 202 may include: the computer device determining the shape coefficient of the first sample image and the preconfigured parameters of the second sample image as the target attribute parameters, where the preconfigured parameters of the second sample image include at least one of an expression coefficient, an angle coefficient, a texture coefficient, or an illumination coefficient. Corresponding to Method 2 in step 201, the preconfigured parameters may be the preconfigured parameters obtained in Method 2 in 201. In this step, the computer device may directly obtain the preconfigured parameters of the second sample image. Corresponding to Method 3 in step 201, the preconfigured parameters may also be preconfigured parameters extracted from the attribute parameters including the five coefficients. In this step, the computer device may extract the preconfigured parameters corresponding to the preconfigured parameter identifier from the second sample image according to the preconfigured parameter identifier. For example, the preconfigured parameter identifier may include a parameter identifier for at least one of the expression coefficient, the angle coefficient, the texture coefficient, or the illumination coefficient. For example, the preconfigured parameters may include expression coefficients and angles, that is, the face in the sample face-swap image to be generated is expected to have the shape of the target face, facial features, etc., as well as the expression and angle of the face in the second sample image; the computer device may then determine the shape coefficient of the target face, and the expression coefficient and angle of the second sample image as the target attribute parameters. For another example, the preconfigured parameters may also include texture and illumination, that is, the face in the sample face-swap image is expected to have the shape of the target face, as well as the texture and illumination of the face in the second sample image; the computer device may then determine the shape coefficient of the target face, and the texture coefficient and illumination coefficient of the second sample image as the sample attribute parameters.
[0110] Step 203: The computer device determines the comprehensive features of the sample based on the sample attribute parameters and the facial features of the first sample image.
[0111] The computer device may concatenate the sample attribute parameters and the facial features of the first sample image, and use the concatenated features as the sample comprehensive features. The sample comprehensive features may represent the comprehensive features of the face in the desired sample facial features. For example, the sample attribute parameters and the facial features may be represented in the form of feature vectors. The computer device may combine the first feature vector corresponding to the sample attribute parameters and the second feature vector corresponding to the facial features to obtain a third feature vector corresponding to the sample comprehensive features.
[0112] Step 204: The computer device inputs the second sample image into the encoder of the initial model for encoding to obtain sample coding features.
[0113] The computer device inputs the second sample image into the encoder of the initial model, and the encoder encodes the second sample image to obtain a coding vector corresponding to the second sample image, and uses the coding vector as the sample coding feature. It should be noted that the sample coding feature obtained by encoding the second sample image accurately refines the pixel-level information of each pixel included in the second sample image.
[0114] Step 205: The computer device transfers the sample comprehensive features to the sample coding features of the second sample image through an adaptive instance regularization method to obtain a sample fusion feature.
[0115] The computer device may implement step 205 to fuse the sample comprehensive feature and the sample coding feature. In this step, the computer device may utilize the adaptive instance regularization method to align the feature distribution of the sample comprehensive feature with the feature distribution of the second sample image to obtain the sample fusion feature. In one possible implementation, the feature distribution may include a mean and a standard deviation. Accordingly, step 205 may include: the computer device obtaining the mean and standard deviation of the sample coding feature in at least one feature channel, and obtaining the mean and standard deviation of the sample comprehensive feature in at least one feature channel; the computer device aligning the mean and standard deviation of the sample coding feature in each feature channel with the mean and standard deviation of the sample comprehensive feature in the corresponding feature channel to obtain the sample fusion feature. Exemplarily, the computer device may normalize each feature channel of the sample coding feature and align the mean and standard deviation of the normalized sample coding feature with the mean and standard deviation of the sample comprehensive feature to generate the sample fusion feature.
[0116] In one possible example, the computer device may calculate the sample fusion feature based on the sample coding feature and the sample comprehensive feature using the following formula 1.
[0117] Formula 1:
[0118] Here, x represents the sample encoding feature, y represents the sample comprehensive feature, σ(x) and μ(x) represent the mean and standard deviation of the sample encoding feature, respectively, and σ(y) and μ(y) represent the mean and standard deviation of the sample comprehensive feature, respectively. AdaIN(x,y) represents the sample fusion feature generated based on the AdaIN (Adaptive Instance Normalization) algorithm.
[0119] Step 206: The computer device inputs the sample fusion feature into the decoder of the initial network for decoding to obtain a sample face-swapped image.
[0120] The computer device passes the sample fusion feature through a decoder to restore the image corresponding to the sample fusion feature, and the computer device uses the image output by the decoder as the sample face-changing image. The decoder can restore the image corresponding to the injected feature based on the injected feature. The computer device decodes the sample fusion image through the decoder to obtain a sample face-changing image; for example, the encoder can perform a convolution operation on the input image, so the decoder can perform a reverse operation according to the operating principle of the encoder during operation, that is, a deconvolution operation, to restore the image corresponding to the sample fusion feature. For example, the encoder can be an autoencoder (AE), and the decoder can be a decoder corresponding to the autoencoder.
[0121] It should be noted that, through the above-mentioned step 205, the feature migration is performed using the adaptive instance regularization method, which can support the migration of the sample comprehensive features to the coding features of any image, thereby realizing the mixing of the sample comprehensive features and the sample coding features; and the sample coding features can represent the features of each pixel in the second sample image, and the sample comprehensive features integrate the features of the first sample image and the second sample image from a global perspective. Therefore, through the adaptive instance regularization method, the mixing between the coding features refined to the pixel level and the global comprehensive features is achieved, and the feature distribution of the sample coding features is aligned with the sample comprehensive features, thereby improving the accuracy of the generated sample fusion features; through step 206, the sample fusion features are used to decode the image, so that the decoded image can be refined to each pixel to show the sample comprehensive features, thereby improving the sensory similarity between the face in the decoded image and the target face, and improving the accuracy of face swapping.
[0122] Step 207: The computer device determines the total loss of the initial model based on the differences between the sample face-swapped image and the sample attribute parameters, the facial features of the first sample image, and the second sample image, and trains the initial model based on the total loss until the training is stopped when the target conditions are met, thereby obtaining the face-swapped model.
[0123] The computer device can determine multiple similarities between the sample face-swapped image and the sample attribute parameters, the facial features of the first sample image, and the second sample image, respectively, and obtain the total loss based on the multiple similarities. In one possible implementation, the initial network can include a discriminator, and the computer device can use the discriminator to determine the authenticity of the sample face-swapped image. Exemplarily, the process of the computer determining the total loss can include the following steps: the computer device obtains a first similarity between the attribute parameters of the sample face-swapped image and the sample attribute parameters; the computer device obtains a second similarity between the facial features of the sample face-swapped image and the facial features of the first sample image; the computer device obtains a third similarity between the second sample image and the sample face-swapped image through the discriminator of the initial network; and the computer device determines the total loss based on the first similarity, the second similarity, and the third similarity.
[0124] In one example, the computer device may extract the attribute parameters of the sample face-swapped image based on a first similarity between the attribute parameters of the sample face-swapped image and the sample attribute parameters using the following formula 2.
[0125] Formula 2: 3dfeatureloss=abs(gt3dfeature–result3dfeature);
[0126] Wherein, 3D feature loss represents the first similarity. The smaller the value of the first similarity, the closer the attribute parameters of the sample face-swapped image are to the sample attribute parameters. Result 3D feature represents the attribute parameters of the sample face-swapped image, gt 3D feature represents the sample attribute parameters, and abs represents the absolute value of (gt 3D feature – result 3D feature). In a possible example, the sample attribute parameters can be the shape coefficients of the target face and the expression coefficients and angles of the second sample image. Accordingly, gt 3D feature can be expressed as the following formula 3:
[0127] Formula 3: gt 3d feature=source 3d feature id+target 3d featureexpression+target 3d feature angles;
[0128] Among them, source 3D feature id represents the shape image number of the first sample image, target 3D feature expression represents the expression coefficient of the second sample image, and target 3D feature angles represents the angle of the second sample image.
[0129] In one example, the computer device may extract facial features of the sample face-swapped image based on a second similarity between the facial features of the sample face-swapped image and the facial features of the first sample image using the following formula 4.
[0130] Formula 4: id loss=1-cosine similarity(result id feature,Mean SourceID);
[0131] Wherein, id loss represents the second similarity. The smaller the value of the second similarity, the closer the facial features of the sample face-swapped image are to the facial features of the first sample image. result id feature represents the facial features of the sample face-swapped image, and mean source ID represents the facial features of the first sample image. cosine similarity (result id feature, mean source ID) represents the cosine similarity between result id feature and mean source ID. The method for determining the preliminary similarity can be as follows:
[0132] Formula 5:
[0133] Wherein, A and B can represent the feature vector corresponding to the facial features of the sample face-swapped image and the feature vector corresponding to the facial features of the first sample image respectively; θ represents the angle between the two feature vectors A and B; A i Represents the component of the i-th feature channel in the facial features of the sample face-swapped image; B i It represents the component of the i-th feature channel in the facial features of the first sample image. similarity and cos(θ) represent cosine similarity.
[0134] In one example, the computer device may input the second sample image as a real image into the discriminator and input the sample face-swapped image into the discriminator. The computer device, through the discriminator, obtains a scaled image of the second sample image at at least one scale and a scaled image of the sample face-swapped image at at least one corresponding scale. The computer device obtains a discrimination probability corresponding to at least one scaled image of the second sample image and a discrimination probability corresponding to at least one scaled image of the sample face-swapped image, where the discrimination probability of the image indicates the probability that the image is a real image. The computer device determines the third similarity based on the at least one discrimination probability corresponding to the second sample image and the at least one discrimination probability corresponding to the sample face-swapped image. For example, the initial network may include a generator and a discriminator. The computer device obtains a discrimination loss value corresponding to the discriminator and a generation loss value corresponding to the generator, and determines the third similarity based on the generation loss and the discrimination loss values. The generator is configured to generate the sample face-swapped image based on the second sample image and the first sample image. For example, the generator may include the encoder and decoder used in steps 204 to 206 above. In one example, the third similarity may include a generation loss value and a discrimination loss value; wherein, the computer device may use the discrimination probability of the sample face-swapped image to represent the generation loss value. For example, the computer device calculates the generation loss value based on the discrimination probability of the sample face-swapped image through the following formula six.
[0135] Formula 6: Gloss = log(1 – D(result));
[0136] Among them, D(result) represents the discrimination probability of the sample face-swapped image, which refers to the probability that the sample face-swapped image belongs to the real image, and G loss represents the generation loss value.
[0137] In one example, the discriminator may be a multi-scale discriminator. The computer device may use the discriminator to scale the sample face-swapped image to obtain scaled images at multiple scales, for example, obtaining scaled images of the sample face-swapped image at a first scale, a second scale, and a third scale. Similarly, the computer device may use the discriminator to obtain scaled images of the second sample image at the first scale, a second scale, and a third scale. The first scale, the second scale, and the third scale may be set as needed. For example, the first scale may be the original scale of the sample face-swapped image or the second sample image, the second scale may be 1 / 2 of the original scale, and the third scale may be 1 / 4 of the original scale. The computer device may use the multi-scale discriminator to obtain the discrimination probabilities corresponding to the scaled images at each scale and calculate the discrimination loss value based on the discrimination probabilities of the scaled images at the multiple scales. For example, based on at least one discrimination probability corresponding to the second sample image and at least one discrimination probability corresponding to the sample face-swapped image, the computer device obtains the discrimination loss value using the following formula 7:
[0138] Formula 7:
[0139] D loss=1 / 3*{–logD(template img)–log(1–D(result))–logD(template img1 / 2)–log(1–D(result1 / 2))–logD(template img1 / 4)–log(1–D(result1 / 4))};
[0140] Wherein, D(template img), D(template img1 / 2), and D(template img1 / 4) represent the discrimination probabilities of the second sample image at the original scale, the second sample image at the scaled image of 1 / 2, and the second sample image at the scaled image of 1 / 4, respectively; D(result), D(result1 / 2), and D(result1 / 4) represent the discrimination probabilities of the sample face-swapped image at the original scale, the sample face-swapped image at the scaled image of 1 / 2, and the sample face-swapped image at the scaled image of 1 / 4, respectively. In this application, the second sample image frame can be used as the real image.
[0141] The computer device may determine a third similarity based on the discriminant loss and the generation loss, for example, third similarity = G loss + D loss. For the discriminator, when a balance is reached between the generation loss and the discriminant loss, the discriminator may be considered to have reached a training stop condition and no further training is required.
[0142] Furthermore, the computer device may determine the total loss based on the first similarity, the second similarity, and the third similarity using the following formula 8:
[0143] Formula 8: loss=id loss+3d feature loss+D loss+G loss;
[0144] Among them, loss represents the total loss, 3d feature loss represents the first similarity, id loss represents the second similarity, and (D loss + G loss) represents the third similarity.
[0145] The computer device may iteratively train the initial model based on steps 201 to 206 above, and obtain the total loss corresponding to each iterative training. The parameters of the initial model may be adjusted based on the total loss of each iterative training. For example, the parameters of the encoder, decoder, and discriminator in the initial model may be optimized multiple times until the total loss meets a target condition. The computer device then stops training and uses the initial model obtained from the last optimization as the face-swapping model. The target condition may be that the total loss is within a target range, for example, the total loss is less than 0.5; or that the time consumed by multiple iterative trainings exceeds a maximum duration.
[0146] Figure 3 A schematic diagram of a framework of a dedicated face-changing model training process provided in an embodiment of the present application is shown as follows: Figure 3As shown, the computer device can use the face of subject A as the exclusive target face, obtain facial images of subject A's face in multiple poses as the first sample image, extract attribute parameters of the first sample image using a 3D facial parameter estimation model, extract facial features of the first sample image using a facial recognition model, and extract attribute parameters of the second sample image using the 3D facial parameter estimation model. The computer device integrates the facial features and shape coefficients of the first sample image and pre-configured parameters (such as expression coefficients and angle coefficients) of the second sample image into sample attribute parameters. The computer device can input the second sample image into an initial face-swapping network, which may include an encoder and a decoder. The computer device can encode the second sample image using the encoder to obtain encoded features of the second sample image, for example, encoding the second sample image into a corresponding feature vector. The computer device obtains sample fusion features based on the sample attribute parameters and the encoded features of the second sample image, and injects the sample fusion features into the decoder of the initial face-swapping network. The decoder can restore the image corresponding to the injected features based on the injected features. The computer device decodes the sample fusion image through the decoder to obtain a sample face-swapped image; for example, the encoder performs a deconvolution operation according to the operating principle of the encoder to restore the image corresponding to the sample fusion feature.
[0147] Furthermore, the computer device obtains a third similarity through a multi-scale discriminator, and obtains a first similarity and a second similarity based on the facial features and attribute parameters of the extracted sample face-swapped image, and calculates a total loss based on the first similarity, the second similarity, and the third similarity to optimize the model parameters according to the total loss; the computer device iteratively trains the above process until the target conditions are met, and stops training, thereby obtaining a face-swapped model that can replace the face in any image with a dedicated target face.
[0148] Figure 4 This is a signaling interaction diagram of an image processing method provided by this application. Figure 4 As shown, the image processing method can be implemented by the interaction between the server and the terminal. The interaction process of the image processing method is as follows:
[0149] Step 401: The terminal displays an application page of a target application, where the application page includes a target trigger control, and the target trigger control is used to trigger a face-swapping request for the image to be face-swapped.
[0150] The target application may provide a face-swapping function, which may replace the face in the image to be swapped with a specific target face. In one example, the target application's application page may include a target trigger control. Based on the user's triggering operation on the target trigger control, the terminal may send a face-swapping request to the server. For example, the target application may be an image processing application, a live broadcast application, a photo-taking tool, a video editing application, etc. The server may be the backend server of the target application, or it may be any computer device that provides the face-swapping function, such as a cloud computing center device equipped with a face-swapping model.
[0151] Step 402: The terminal obtains the image to be face-swapped in response to the trigger operation for the target trigger control in the received application page, and sends a face-swapped request to the server based on the image to be face-swapped.
[0152] Scenario example one, the target application can provide a face-swapping function for a single image. For example, the target application can be an image processing application, a live broadcast application, a social application, etc. The image to be replaced can be a selected image obtained by the terminal from the local storage space, or it can also be an image obtained by the terminal by shooting the object in real time. Scenario example two, the target application can provide a face-swapping function for each image frame included in a video. For example, the target application can be a video editing application, a live broadcast application, etc. The server can replace the entire image frame including the face of object A in the video with the target face. Then the image to be replaced can include each image frame in the video, or the terminal can perform initial face detection on each image frame in the video, and use each image frame including the face of object A in the video as the image to be replaced.
[0153] Step 403: The server receives the face-changing request sent by the terminal, and obtains attribute parameters of the image to be face-changed in response to the received face-changing request.
[0154] The face-swapping request is used to request that the face in the image to be swapped be replaced with the target face. The attribute parameters of the image to be swapped are used to indicate the three-dimensional attributes of the face in the image to be swapped. The server may obtain the attribute parameters of the image to be swapped using a 3D facial parameter estimation model. Exemplarily, the attribute parameters of the image include at least one of a shape coefficient, an expression coefficient, an angle coefficient, a texture coefficient, or an illumination coefficient.
[0155] Step 404: The server determines target attribute parameters based on the attribute parameters of the image to be replaced and the pre-stored attribute parameters of the target face.
[0156] In one possible implementation, the server may determine the target attribute parameter as the target attribute parameter, the shape coefficient of the target face and preconfigured parameters of the image to be face-swapped, where the preconfigured parameters include at least one of an expression coefficient, an angle coefficient, a texture coefficient, or an illumination coefficient. For example, the preconfigured parameters may include an expression coefficient and an angle coefficient. Alternatively, the preconfigured parameters may also include a texture coefficient, an illumination coefficient, and the like.
[0157] Step 405: The server determines the target comprehensive features based on the target attribute parameters and pre-stored facial features of the target face.
[0158] The server may combine the target attribute parameters and the facial features of the target face to obtain the target comprehensive features.
[0159] It should be noted that the server can be configured with a pre-trained face-changing model, and the server can perform the process from step 404 to step 405 above through the face-changing model. The face-changing model is trained based on the above steps 201 to step 207. When the server trains the face-changing model, it can store the facial features and attribute parameters of the target face in a fixed manner, for example, it can store them at the target address. When executing steps 404 and 405, the server can extract the attribute parameters of the target face from the target address and execute step 404, and the server can extract the facial features of the target face from the target address and execute step 405. In addition, the server can perform the process from step 406 to step 408 below through the face-changing model.
[0160] Step 406: The server encodes the image to be face-swapped to obtain image coding features of the image to be face-swapped.
[0161] Step 407: The server transfers the target comprehensive features to the image coding features of the image to be face-swapped through an adaptive instance regularization method to obtain a fused coding feature.
[0162] In one possible implementation, the computer device may align the mean and standard deviation of the image coding feature with the target comprehensive feature. Exemplarily, step 407 may include: the server obtains the mean and standard deviation of the image coding feature in at least one feature channel, and obtains the mean and standard deviation of the target comprehensive feature in at least one feature channel; the server aligns the mean and standard deviation of the image coding feature in each feature channel with the mean and standard deviation of the target comprehensive feature in the corresponding feature channel to obtain the fused coding feature. For example, the server may also use Formula 1 in step 205 above to calculate the fused coding feature.
[0163] Step 408: The server decodes the fused coding feature to obtain a target face-swapped image including a fused face, where the fused face is a fusion of the face in the image to be face-swapped and the target face.
[0164] It should be noted that the server executes steps 403 to 408 to obtain the target face-swapped image in the same manner as the computer device executes steps 201 to 206 to generate a sample face-swapped image, and will not be described in detail here.
[0165] Step 409: The server returns the target face-swapped image to the terminal.
[0166] When the image to be face-swapped is a single image, the server may return the target face-swapped image corresponding to the single image to be face-swapped to the terminal. When the image to be face-swapped is a plurality of image frames included in a video, the server may generate a target face-swapped image corresponding to each image frame to be face-swapped in the video through steps 403 to 408 above. The server may then return the face-swapped video corresponding to the video to the terminal, the face-swapped video including the target face-swapped image corresponding to each image frame.
[0167] Step 410: The terminal receives the target face-swapped image returned by the server and displays the target face-swapped image.
[0168] The terminal may display the target face-swapped image in the application page. Alternatively, the terminal may also play each target face-swapped image in the face-swapped video in the application page.
[0169] The image processing method provided by the present application obtains attribute parameters of the image to be replaced, which are used to indicate the three-dimensional attributes of the face in the image, and determines target attribute parameters based on the attribute parameters of the image to be replaced and the attribute parameters of the pre-stored target face, thereby locating the three-dimensional attribute features of the face in the image to be generated; and, based on the target attribute parameters and the facial features of the pre-stored target face, obtains a target comprehensive feature that can comprehensively characterize the image to be replaced and the target face; and encodes the image to be replaced to obtain the image coding feature of the image to be replaced, thereby obtaining the pixel-level refined features of the image to be replaced through the image coding feature; and further migrates the target comprehensive feature to the image coding feature of the image to be replaced through an adaptive instance regularization method to obtain a fused coding feature. The present application mixes coding features refined to the pixel level and global comprehensive features, and aligns the feature distribution of image coding features with the target comprehensive features, thereby improving the accuracy of the generated fused coding features; by decoding the fused coding features, a target face-changing image including the fused face is obtained, so that the decoded image can be refined to each pixel point to show the target comprehensive features, making the sensory perception of the fused face in the decoded image closer to the target face, improving the sensory similarity between the fused face and the target face, thereby improving the accuracy of face-changing.
[0170] Figure 5 This is a structural diagram of an image processing device provided in an embodiment of the present application. Figure 5 As shown, the device includes:
[0171] Attribute parameter acquisition module 501, configured to acquire attribute parameters of the image to be face-swapped in response to a received face-swap request, wherein the face-swap request is used to request that the face in the image to be face-swapped be replaced with the target face, and the attribute parameters of the image to be face-swapped are used to indicate three-dimensional attributes of the face in the image to be face-swapped;
[0172] The target attribute parameter determination module 502 is used to determine the target attribute parameters based on the attribute parameters of the image to be replaced and the pre-stored attribute parameters of the target face;
[0173] A comprehensive feature determination module 503 is configured to determine a target comprehensive feature based on the target attribute parameters and pre-stored facial features of the target face;
[0174] The encoding module 504 is used to encode the image to be face-swapped to obtain image encoding features of the image to be face-swapped;
[0175] A migration module 505 is configured to migrate the target comprehensive feature to the image coding feature of the face-swapped image through an adaptive instance regularization method to obtain a fused coding feature;
[0176] The decoding module 506 is used to decode the fused coding feature to obtain a target face-swapped image including a fused face, where the fused face is a fusion of the face in the image to be face-swapped and the target face.
[0177] In one possible implementation, the attribute parameter of the image includes at least one of a shape coefficient, an expression coefficient, an angle coefficient, a texture coefficient, or an illumination coefficient;
[0178] The target attribute parameter determination module is used to determine the shape coefficient of the target face and the pre-configured parameters of the image to be face-swapped as the target attribute parameters, and the pre-configured parameters include at least one of an expression coefficient, an angle coefficient, a texture coefficient or an illumination coefficient.
[0179] In one possible implementation, the migration module is configured to:
[0180] Obtaining a mean value and a standard deviation of the image coding feature in at least one feature channel, and obtaining a mean value and a standard deviation of the target comprehensive feature in at least one feature channel;
[0181] The mean and standard deviation of the image coding feature in each feature channel are aligned with the mean and standard deviation of the target comprehensive feature in the corresponding feature channel to obtain the fused coding feature.
[0182] In one possible implementation, the target face-swap image is obtained by a pre-trained face-swap model; the face-swap model is used to swap the target face into any facial image based on pre-stored attribute data and facial features of the target face;
[0183] The device also includes a model training module, which includes:
[0184] an acquiring unit, configured to acquire facial features and attribute parameters of a first sample image, and acquire attribute parameters of a second sample image, wherein the first sample image includes the target face and the second sample image includes the face to be replaced;
[0185] a sample attribute parameter determining unit, configured to determine a sample attribute parameter based on the attribute parameter of the first sample image and the attribute parameter of the second sample image, the sample attribute parameter being used to indicate desired attributes of the face in the sample face-swapped image to be generated;
[0186] a sample comprehensive feature acquisition unit, configured to determine the sample comprehensive feature based on the sample attribute parameter and the facial feature of the first sample image;
[0187] an encoding unit, configured to input the second sample image into an encoder of an initial model for encoding to obtain a sample encoding feature;
[0188] a migration unit, configured to migrate the sample comprehensive feature to the sample encoding feature of the second sample image through an adaptive instance regularization method to obtain a sample fusion feature;
[0189] A decoding unit, configured to input the sample fusion features into the decoder of the initial network for decoding to obtain a sample face-swapped image;
[0190] The training unit is used to determine the total loss of the initial model based on the differences between the sample face-swapped image and the sample attribute parameters, the facial features of the first sample image, and the second sample image, and to train the initial model based on the total loss until the training is stopped when the target conditions are met, thereby obtaining the face-swapped model.
[0191] In one possible implementation, the training unit is specifically configured to:
[0192] Obtaining a first similarity between the attribute parameters of the sample face-swapped image and the attribute parameters of the sample;
[0193] Obtaining a second similarity between facial features of the sample face-swapped image and facial features of the first sample image;
[0194] Obtaining, by the discriminator of the initial network, a third similarity between the second sample image and the sample face-swapped image;
[0195] The total loss is determined based on the first similarity, the second similarity, and the third similarity.
[0196] In one possible implementation, the training unit is further configured to:
[0197] Inputting the second sample image as a real image into the discriminator, and inputting the sample face-swapped image into the discriminator;
[0198] Obtaining, by the discriminator, a scaled image of the second sample image at at least one scale and a scaled image of the sample face-swapped image at the corresponding at least one scale;
[0199] Obtaining a discrimination probability corresponding to at least one scaled image of the second sample image, and obtaining a discrimination probability corresponding to at least one scaled image of the sample face-swapped image, wherein the discrimination probability of an image is used to indicate a probability that the image is a real image;
[0200] The third similarity is determined based on at least one discrimination probability corresponding to the second sample image and at least one discrimination probability corresponding to the sample face-swapped image.
[0201] In one possible implementation, the acquiring unit is specifically configured to:
[0202] Acquire at least two posture images as the first sample images, where the at least two posture images include at least two facial postures of the target face;
[0203] Based on the at least two posture images, acquiring facial features and attribute parameters corresponding to the at least two facial postures;
[0204] using an average of the facial features corresponding to the at least two facial postures as the facial features of the first sample image, and using an average of the attribute parameters corresponding to the at least two facial postures as the attribute parameters of the first sample image;
[0205] Correspondingly, the device further includes a storage unit, which is used to store the facial features and attribute parameters of the first sample image.
[0206] In one possible implementation, the acquiring unit is further configured to:
[0207] Performing facial recognition on at least two image frames included in the video of the target object to obtain at least two image frames including the target face, where the target face is the face of the target object;
[0208] Perform face cropping on the at least two image frames to obtain the at least two posture images, and use the at least two posture images as the first sample images.
[0209] The image processing device provided by the present application obtains attribute parameters of the image to be replaced, which are used to indicate the three-dimensional attributes of the face in the image, and determines target attribute parameters based on the attribute parameters of the image to be replaced and the attribute parameters of the pre-stored target face, thereby locating the three-dimensional attribute features of the face in the image to be generated; and, based on the target attribute parameters and the facial features of the pre-stored target face, obtains a target comprehensive feature that can comprehensively characterize the image to be replaced and the target face; and encodes the image to be replaced to obtain the image coding feature of the image to be replaced, thereby obtaining the refined features of the image to be replaced at the pixel level through the image coding feature; and further migrates the target comprehensive feature to the image coding feature of the image to be replaced through an adaptive instance regularization method to obtain a fused coding feature. The present application mixes coding features refined to the pixel level and global comprehensive features, and aligns the feature distribution of image coding features with the target comprehensive features, thereby improving the accuracy of the generated fused coding features; by decoding the fused coding features, a target face-changing image including the fused face is obtained, so that the decoded image can be refined to each pixel point to show the target comprehensive features, making the sensory perception of the fused face in the decoded image closer to the target face, improving the sensory similarity between the fused face and the target face, thereby improving the accuracy of face-changing.
[0210] The device of the embodiment of the present application can execute the image processing method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the image processing device of each embodiment of the present application correspond to the steps in the image processing method of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description of the corresponding image processing method shown in the previous text, and will not be repeated here.
[0211] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 6 As shown, the computer device includes: a memory, a processor, and a computer program stored in the memory. The processor executes the above computer program to implement the steps of the image processing method. Compared with the related art, it can achieve:
[0212] The image processing device provided by the present application obtains attribute parameters of the image to be replaced, which are used to indicate the three-dimensional attributes of the face in the image, and determines target attribute parameters based on the attribute parameters of the image to be replaced and the attribute parameters of the pre-stored target face, thereby locating the three-dimensional attribute features of the face in the image to be generated; and, based on the target attribute parameters and the facial features of the pre-stored target face, obtains a target comprehensive feature that can comprehensively characterize the image to be replaced and the target face; and encodes the image to be replaced to obtain the image coding feature of the image to be replaced, thereby obtaining the refined features of the image to be replaced at the pixel level through the image coding feature; and further migrates the target comprehensive feature to the image coding feature of the image to be replaced through an adaptive instance regularization method to obtain a fused coding feature. The present application mixes coding features refined to the pixel level and global comprehensive features, and aligns the feature distribution of image coding features with the target comprehensive features, thereby improving the accuracy of the generated fused coding features; by decoding the fused coding features, a target face-changing image including the fused face is obtained, so that the decoded image can be refined to each pixel point to show the target comprehensive features, making the sensory perception of the fused face in the decoded image closer to the target face, improving the sensory similarity between the fused face and the target face, thereby improving the accuracy of face-changing.
[0213] In an alternative embodiment, a computer device is provided, such as Figure 6 As shown, Figure 6The computer device 600 shown includes a processor 601 and a memory 603. The processor 601 and the memory 603 are connected, for example, via a bus 602. Optionally, the computer device 600 may also include a transceiver 604, which can be used for data exchange between the computer device and other computer devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 604 is not limited to one, and the structure of the computer device 600 does not constitute a limitation on the embodiments of this application.
[0214] The processor 601 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor 601 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0215] The bus 602 may include a path for transmitting information between the above components. The bus 602 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 602 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0216] The memory 603 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media\other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation here.
[0217] The memory 603 is used to store the computer program for executing the embodiment of the present application, and the execution is controlled by the processor 601. The processor 601 is used to execute the computer program stored in the memory 603 to implement the steps shown in the above method embodiment.
[0218] Among them, electronic equipment includes but is not limited to: servers or cloud computing center equipment, etc.
[0219] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0220] An embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiment when executed by a processor.
[0221] Those skilled in the art will understand that, unless otherwise specified, the singular forms "a," "an," "the," and "the" used herein may also include the plural forms. The terms "including" and "comprising" used in the embodiments of this application mean that the corresponding features can be implemented as the presented features, information, data, steps, and operations, but do not exclude the implementation of other features, information, data, steps, operations, etc. supported by the technical field.
[0222] The terms "first," "second," "third," "fourth," "1," "2," and the like (if any) in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than that shown or described in the drawings.
[0223] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times respectively. Under different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0224] The above description is only an optional implementation method for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, the use of other similar implementation methods based on the technical ideas of this application also falls within the protection scope of the embodiments of this application.
Claims
1. An image processing method, characterized in that: The method comprises: In response to a received face-swapping request, obtaining attribute parameters of the image to be face-swapping, wherein the face-swapping request is used to request that a face in the image to be face-swapping be replaced with a target face, the attribute parameters of the image to be face-swapping indicating three-dimensional attributes of the face in the image to be face-swapping, the three-dimensional attributes including at least one of a posture attribute or a spatial environment attribute in a three-dimensional space; Determining the attribute parameters of the image to be face-swapped and the pre-stored attribute parameters of the target face as target attribute parameters; splicing the target attribute parameters with pre-stored facial features of the target face to obtain target comprehensive features, wherein the facial features are features of the target face in a two-dimensional plane and are used to characterize the identity of the object having the facial features; Encoding the image to be face-swapped to obtain image coding features of the image to be face-swapped; Migrating the target comprehensive features to the image coding features of the face-swapped image through an adaptive instance regularization method to obtain a fused coding feature; The fusion coding feature is decoded to obtain a target face-swapped image including a fused face, where the fused face is a fusion of the face in the image to be face-swapped and the target face.
2. The image processing method according to claim 1, wherein: The attribute parameters of the image include at least one of a shape coefficient, an expression coefficient, an angle coefficient, a texture coefficient or an illumination coefficient; The step of determining the attribute parameters of the image to be face-swapped and the pre-stored attribute parameters of the target face as target attribute parameters includes: The shape coefficient of the target face and the pre-configured parameters of the image to be face-swapped are determined as the target attribute parameters, where the pre-configured parameters include at least one of an expression coefficient, an angle coefficient, a texture coefficient, or an illumination coefficient.
3. The image processing method according to claim 1, wherein: The step of migrating the target comprehensive features to the image coding features of the face-swapped image by using an adaptive instance regularization method to obtain fused coding features includes: Obtaining a mean value and a standard deviation of the image coding feature in at least one feature channel, and obtaining a mean value and a standard deviation of the target comprehensive feature in at least one feature channel; The mean and standard deviation of the image coding feature in each feature channel are aligned with the mean and standard deviation of the target comprehensive feature in the corresponding feature channel to obtain the fused coding feature.
4. The image processing method according to claim 1, wherein: The target face-swap image is obtained by a pre-trained face-swap model; the face-swap model is used to swap the target face into any facial image based on pre-stored attribute data and facial features of the target face; The training method of the face-changing model includes: Acquiring facial features and attribute parameters of a first sample image, and acquiring attribute parameters of a second sample image, wherein the first sample image includes the target face and the second sample image includes the face to be replaced; determining the attribute parameters of the first sample image and the attribute parameters of the second sample image as sample attribute parameters, wherein the sample attribute parameters are used to indicate desired attributes of the face in the sample face-swapped image to be generated; splicing the sample attribute parameters and the facial features of the first sample image to obtain a sample comprehensive feature; Inputting the second sample image into the encoder of the initial model for encoding to obtain a sample encoding feature; Migrating the sample comprehensive features to the sample encoding features of the second sample image through an adaptive instance regularization method to obtain a sample fusion feature; Inputting the sample fusion features into the decoder of the initial model for decoding to obtain a sample face-swapped image; Based on the differences between the sample face-swapped image and the sample attribute parameters, the facial features of the first sample image, and the second sample image, the total loss of the initial model is determined, and the initial model is trained based on the total loss until the training is stopped when the target conditions are met, thereby obtaining the face-swapped model.
5. The image processing method according to claim 4, characterized in that The determining of the total loss of the initial model based on the differences between the sample face-swapped image and the sample attribute parameters, the facial features of the first sample image, and the second sample image, comprises: Obtaining a first similarity between the attribute parameters of the sample face-swapped image and the attribute parameters of the sample; Obtaining a second similarity between facial features of the sample face-swapped image and facial features of the first sample image; obtaining, by the discriminator of the initial model, a third similarity between the second sample image and the sample face-swapped image; The total loss is determined based on the first similarity, the second similarity, and the third similarity.
6. The image processing method according to claim 5, characterized in that The obtaining, by the discriminator of the initial model, a third similarity between the second sample image and the sample face-swapped image includes: Inputting the second sample image as a real image into the discriminator, and inputting the sample face-swapped image into the discriminator; Obtaining, by the discriminator, a scaled image of the second sample image at at least one scale and a scaled image of the sample face-swapped image at the corresponding at least one scale; Obtaining a discrimination probability corresponding to at least one scaled image of the second sample image, and obtaining a discrimination probability corresponding to at least one scaled image of the sample face-swapped image, the discrimination probability of an image being used to indicate a probability that the image is a real image; The third similarity is determined based on at least one discrimination probability corresponding to the second sample image and at least one discrimination probability corresponding to the sample face-swapped image.
7. The image processing method according to claim 4, wherein: The acquiring of facial features and attribute parameters of the first sample image includes: Acquire at least two posture images as the first sample images, wherein the at least two posture images include at least two facial postures of the target face; Based on the at least two posture images, acquiring facial features and attribute parameters corresponding to the at least two facial postures; using an average of the facial features corresponding to the at least two facial postures as the facial features of the first sample image, and using an average of the attribute parameters corresponding to the at least two facial postures as the attribute parameters of the first sample image; Accordingly, after obtaining the facial features and attribute parameters of the first sample image, the method further includes: The facial features and attribute parameters of the first sample image are stored.
8. The image processing method according to claim 7, wherein: The acquiring at least two posture images as the first sample images includes: Performing facial recognition on at least two image frames included in the video of the target object to obtain at least two image frames including the target face, where the target face is the face of the target object; Perform face cropping on the at least two image frames to obtain the at least two posture images, and use the at least two posture images as the first sample images.
9. An image processing device, characterized in that: The device comprises: an attribute parameter acquisition module, configured to acquire attribute parameters of the image to be face-swapped in response to a received face-swap request, wherein the face-swap request is used to request that a face in the image to be face-swapped be replaced with a target face, and the attribute parameters of the image to be face-swapped are used to indicate three-dimensional attributes of the face in the image to be face-swapped, wherein the three-dimensional attributes include at least one of a posture attribute or a spatial environment attribute in a three-dimensional space; a target attribute parameter determination module, configured to determine the attribute parameters of the image to be face-swapped and the pre-stored attribute parameters of the target face as target attribute parameters; a comprehensive feature determination module, configured to combine the target attribute parameters with pre-stored facial features of a target face to obtain a target comprehensive feature, wherein the facial feature is a feature of the target face in a two-dimensional plane and is used to characterize the identity of an object having the facial feature; An encoding module, configured to encode the image to be face-swapped to obtain image encoding features of the image to be face-swapped; A migration module, configured to migrate the target comprehensive features to the image coding features of the image to be face-swapped by using an adaptive instance regularization method to obtain a fused coding feature; A decoding module is used to decode the fusion coding features to obtain a target face-changing image including a fused face, where the fused face is a fusion of the face in the image to be face-changed and the target face.
10. The image processing device according to claim 9, wherein The attribute parameters of the image include at least one of a shape coefficient, an expression coefficient, an angle coefficient, a texture coefficient or an illumination coefficient; The target attribute parameter determination module is specifically configured to: The shape coefficient of the target face and the pre-configured parameters of the image to be face-swapped are determined as the target attribute parameters, where the pre-configured parameters include at least one of an expression coefficient, an angle coefficient, a texture coefficient, or an illumination coefficient.
11. The image processing device according to claim 9, wherein When the migration module is used to migrate the target comprehensive features to the image coding features of the face-swapped image through an adaptive instance regularization method to obtain a fused coding feature, it is specifically used to: Obtaining a mean value and a standard deviation of the image coding feature in at least one feature channel, and obtaining a mean value and a standard deviation of the target comprehensive feature in at least one feature channel; The mean and standard deviation of the image coding feature in each feature channel are aligned with the mean and standard deviation of the target comprehensive feature in the corresponding feature channel to obtain the fused coding feature.
12. The image processing device according to claim 9, wherein The target face-swap image is obtained by a pre-trained face-swap model; the face-swap model is used to swap the target face into any facial image based on pre-stored attribute data and facial features of the target face; The device further includes a model training module, which includes: an acquiring unit, configured to acquire facial features and attribute parameters of a first sample image, and acquire attribute parameters of a second sample image, wherein the first sample image includes the target face and the second sample image includes the face to be replaced; a sample attribute parameter determining unit, configured to determine the attribute parameters of the first sample image and the attribute parameters of the second sample image as sample attribute parameters, wherein the sample attribute parameters are used to indicate desired attributes of the face in the sample face-swapped image to be generated; a sample comprehensive feature acquisition unit, configured to combine the sample attribute parameters with the facial features of the first sample image to obtain a sample comprehensive feature; an encoding unit, configured to input the second sample image into an encoder of an initial model for encoding to obtain a sample encoding feature; a migration unit, which migrates the sample comprehensive features to the sample encoding features of the second sample image through an adaptive instance regularization method to obtain a sample fusion feature; A decoding unit, configured to input the sample fusion features into a decoder of the initial model for decoding to obtain a sample face-swapped image; A training unit is used to determine the total loss of the initial model based on the differences between the sample face-swapped image and the sample attribute parameters, the facial features of the first sample image, and the second sample image, and to train the initial model based on the total loss until the training is stopped when the target conditions are met, thereby obtaining the face-swapped model.
13. The image processing device according to claim 12, wherein: When determining the total loss of the initial model based on the differences between the sample face-swapped image and the sample attribute parameters, the facial features of the first sample image, and the second sample image, the training unit is specifically configured to: Obtaining a first similarity between the attribute parameters of the sample face-swapped image and the attribute parameters of the sample; Obtaining a second similarity between facial features of the sample face-swapped image and facial features of the first sample image; obtaining, by the discriminator of the initial model, a third similarity between the second sample image and the sample face-swapped image; The total loss is determined based on the first similarity, the second similarity, and the third similarity.
14. The device according to claim 13, characterized in that When the training unit obtains the third similarity between the second sample image and the sample face-swapped image through the discriminator of the initial model, the training unit is specifically configured to: Inputting the second sample image as a real image into the discriminator, and inputting the sample face-swapped image into the discriminator; Obtaining, by the discriminator, a scaled image of the second sample image at at least one scale and a scaled image of the sample face-swapped image at the corresponding at least one scale; Obtaining a discrimination probability corresponding to at least one scaled image of the second sample image, and obtaining a discrimination probability corresponding to at least one scaled image of the sample face-swapped image, the discrimination probability of an image being used to indicate a probability that the image is a real image; The third similarity is determined based on at least one discrimination probability corresponding to the second sample image and at least one discrimination probability corresponding to the sample face-swapped image.
15. The image processing device according to claim 14, wherein When acquiring the facial features and attribute parameters of the first sample image, the acquiring unit is specifically configured to: Acquire at least two posture images as the first sample images, wherein the at least two posture images include at least two facial postures of the target face; Based on the at least two posture images, acquiring facial features and attribute parameters corresponding to the at least two facial postures; using an average of the facial features corresponding to the at least two facial postures as the facial features of the first sample image, and using an average of the attribute parameters corresponding to the at least two facial postures as the attribute parameters of the first sample image; Correspondingly, the device further includes a storage unit, which is used to store the facial features and attribute parameters of the first sample image.
16. The image processing device according to claim 15, wherein: When the acquisition unit acquires at least two posture images as the first sample images, the acquisition unit is specifically configured to: Performing facial recognition on at least two image frames included in the video of the target object to obtain at least two image frames including the target face, where the target face is the face of the target object; Perform face cropping on the at least two image frames to obtain the at least two posture images, and use the at least two posture images as the first sample images.
17. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the image processing method according to any one of claims 1 to 8.
18. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 8 is implemented.
19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and computer storage medium
CN111508050A
Generative adversarial network training method, image face changing and video face changing method and device
CN111783603A