Image Processing Method, Apparatus, Computer Device, and Storage Medium

By extracting additional image and identity features of the training source facial image, and using the encoder and decoder of the facial replacement model to optimize the model parameters, the problem of large images in the face swap technology is solved, and higher consistency and better face swap effect are achieved.

CN113705290BActive Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110216698.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-26
Publication Date
2025-07-04
Estimated Expiration
2041-02-26

AI Technical Summary

Technical Problem

There is a big difference between the images obtained by the existing face-changing technology and the ideal face-changing image, resulting in poor face-changing effect.

Method used

By obtaining the training source facial image and template facial image, additional image features and identity features are extracted, and image processing is performed using the encoder and decoder of the face replacement model, the additional image difference between the decoded facial image and the contrasting facial image is calculated, and the model parameters of the encoder and decoder are adjusted based on the target model loss value, and the facial replacement model is optimized.

Benefits of technology

It improves the identity consistency, attribute consistency and additional image consistency between the image after face change and the source facial image, and improves the face change effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113705290B_ABST
    Figure CN113705290B_ABST
Patent Text Reader

Abstract

The present application relates to an image processing method, apparatus, computer device, and storage medium. The method includes: extracting source additional image features from a training source facial image; extracting source identity features from the training source facial image; inputting a training template facial image into an encoder of a facial replacement model to obtain facial attribute features; inputting the source additional image features, the source identity features, and the facial attribute features into a decoder of the facial replacement model to obtain a decoded facial image; obtaining a target model loss value based on the additional image difference between the decoded facial image and a comparison facial image; and adjusting the model parameters of the encoder and the decoder based on the target model loss value to obtain a trained facial replacement model. Using this method can improve the facial replacement effect. The image processing method in the present application can be based on artificial intelligence, and this solution is applied in the field of video face swapping, which can improve the face swapping effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and particularly to an image processing method, apparatus, computer device, and storage medium. Background Art

[0002] With the development of computer technologies and artificial intelligence technologies, face swapping technology has emerged. Face swapping technology refers to replacing the face of a person in an image with another face. Face swapping technology has many application scenarios, such as in the production of film and television characters, game character design, virtual avatars, and privacy protection.

[0003] Currently, face swapping technology can be implemented through a neural network model based on artificial intelligence. For example, an image can be input into a neural network model for face swapping, and the neural network model can output an image obtained by swapping the face of the input image.

[0004] However, there is a large difference between the image obtained by existing face swapping technology and the ideal face-swapped image, resulting in a poor face swapping effect. Summary of the Invention

[0005] Based on this, it is necessary to provide an image processing method, apparatus, computer device, and storage medium that can improve the face swapping effect for the above technical problems.

[0006] An image processing method, the method includes: obtaining a training source facial image and a training template facial image; extracting additional image features from the training source facial image to obtain source additional image features corresponding to the training source facial image; extracting identity features from the training source facial image to obtain source identity features corresponding to the training source facial image; inputting the training template facial image into an encoder in a face replacement model for encoding to obtain facial attribute features; inputting the source additional image features, source identity features, and the facial attribute features into a decoder in the face replacement model for decoding to obtain a decoded facial image; obtaining an additional image difference between the decoded facial image and a comparison facial image, and obtaining a target model loss value based on the additional image difference; the target model loss value has a positive correlation with the additional image difference; the comparison facial image includes at least one of the training source facial image or a standard facial image corresponding to the decoded facial image; adjusting model parameters of the encoder and the decoder based on the target model loss value to obtain a trained face replacement model for performing image processing according to the face replacement model.

[0007] An image processing device, the device comprising: a first facial image acquisition module for acquiring a training source facial image and a training template facial image; a training source additional image feature obtaining module for extracting additional image features from the training source facial image to obtain source additional image features corresponding to the training source facial image; a source identity feature obtaining module for extracting identity features from the training source facial image to obtain source identity features corresponding to the training source facial image; a facial attribute feature obtaining module for inputting the training template facial image into an encoder in a facial replacement model for encoding to obtain facial attribute features; a decoded facial image obtaining module for inputting the source additional image features, the source identity features, and the facial attribute features into a decoder in the facial replacement model for decoding to obtain a decoded facial image; a target model loss value obtaining module for obtaining an additional image difference between the decoded facial image and a comparison facial image, and obtaining a target model loss value based on the additional image difference; the target model loss value being in a positive correlation relationship with the additional image difference; the comparison facial image including at least one of the training source facial image or a standard facial image corresponding to the decoded facial image; a trained facial replacement model obtaining module for adjusting model parameters of the encoder and the decoder based on the target model loss value to obtain a trained facial replacement model for performing image processing according to the facial replacement model.

[0008] In some embodiments, the additional image difference includes a first image feature difference, and the target model loss value obtaining module includes: a target additional image feature obtaining unit for extracting additional image features from the decoded facial image to obtain target additional image features corresponding to the decoded facial image; a first image feature difference obtaining unit for determining an image feature difference between the source additional image features and the target additional image features as the first image feature difference; a first target model loss value obtaining unit for obtaining a target model loss value based on the first image feature difference.

[0009] In some embodiments, the target model loss value obtaining module includes: an additional image region obtaining unit for identifying an additional image of the comparison facial image to obtain an additional image region corresponding to the comparison facial image; an additional image enhancement value obtaining unit for obtaining an additional image enhancement value corresponding to the additional image region; an additional image difference obtaining unit for determining an image difference between the additional image region and an image region at a corresponding position in the decoded facial image as an additional image difference; a second target model loss value obtaining unit for obtaining an additional image loss value based on the additional image difference, and performing enhancement processing on the additional image loss value by using the additional image enhancement value to obtain a target model loss value.

[0010] In some embodiments, the additional image difference obtaining unit is further configured to obtain additional pixel points in the additional image region, and obtain decoded pixel points matching the positions of the additional pixel points from the decoded facial image; calculate the pixel value difference between the additional pixel points and the decoded pixel points; perform statistics on the pixel value differences corresponding to the additional image region to obtain a difference statistic value, and use the difference statistic value as the additional image difference.

[0011] In some embodiments, the additional image difference obtaining unit is further configured to perform feature extraction on the additional image region to obtain extracted additional image features; perform feature extraction on the image region corresponding to the additional image region in the decoded facial image to obtain decoded image features; calculate the image feature difference between the extracted additional image features and the decoded image features as the second image feature difference; and obtain the additional image difference based on the second image feature difference.

[0012] In some embodiments, the second target model loss value obtaining unit is further configured to obtain an additional image loss value based on the additional image difference, perform enhancement processing on the additional image loss value by using the additional image enhancement value to obtain an enhanced additional image loss value; obtain the non-additional image region corresponding to the comparison facial image, determine the image difference between the non-additional image region and the image region at the corresponding position in the decoded facial image as the non-additional image difference; obtain a non-additional image loss value based on the non-additional image difference, where the non-additional image enhancement value corresponding to the non-additional image loss value is less than the additional image enhancement value; and obtain the target model loss value according to the enhanced additional image loss value and the non-additional image loss value.

[0013] In some embodiments, the target model loss value obtaining module includes: an additional image loss value obtaining unit configured to obtain an additional image loss value based on the additional image difference; a target identity feature obtaining unit configured to perform identity feature extraction on the decoded facial image to obtain the target identity feature corresponding to the decoded facial image; an identity loss value obtaining unit configured to obtain an identity loss value based on the identity feature difference between the source identity feature and the target identity feature; and a third target model loss value obtaining unit configured to obtain the target model loss value according to the additional image loss value and the identity loss value.

[0014] In some embodiments, the module for training source additional image feature obtaining is further configured to input the training source facial image into a trained additional image feature extraction network to extract additional image features, so as to obtain source additional image features corresponding to the training source facial image; the module for obtaining the trained facial replacement model is further configured to keep the network parameters of the additional image feature extraction network unchanged, and adjust the model parameters of the encoder and the decoder based on the target model loss value, so as to obtain a trained facial replacement model for image processing according to the facial replacement model.

[0015] In some embodiments, the apparatus further includes a module for obtaining an additional image feature extraction network, and the module for obtaining an additional image feature extraction network includes: an additional object image obtaining unit configured to obtain an additional object image, where the additional object image includes an additional object corresponding to the additional image feature; a training unit configured to use the additional object image to train a to-be-trained additional image recognition model to obtain a trained additional image recognition model; and a feature extraction layer extraction unit configured to extract a feature extraction layer before an image recognition layer from the trained additional image recognition model as the additional image feature extraction network.

[0016] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented: obtaining a training source facial image and a training template facial image; performing additional image feature extraction on the training source facial image to obtain source additional image features corresponding to the training source facial image; performing identity feature extraction on the training source facial image to obtain source identity features corresponding to the training source facial image; inputting the training template facial image into an encoder in a facial replacement model for encoding to obtain facial attribute features; inputting the source additional image features, the source identity features, and the facial attribute features into a decoder in the facial replacement model for decoding to obtain a decoded facial image; obtaining an additional image difference between the decoded facial image and a comparison facial image, and obtaining a target model loss value based on the additional image difference; the target model loss value has a positive correlation with the additional image difference; the comparison facial image includes at least one of the training source facial image or a standard facial image corresponding to the decoded facial image; adjusting the model parameters of the encoder and the decoder based on the target model loss value to obtain a trained facial replacement model for image processing according to the facial replacement model.

[0017] A computer-readable storage medium has a computer program stored thereon. When the computer program is executed by a processor, the following steps are implemented: obtaining a training source facial image and a training template facial image; extracting additional image features from the training source facial image to obtain source additional image features corresponding to the training source facial image; extracting identity features from the training source facial image to obtain source identity features corresponding to the training source facial image; inputting the training template facial image into an encoder in a facial replacement model for encoding to obtain facial attribute features; inputting the source additional image features, the source identity features, and the facial attribute features into a decoder in the facial replacement model for decoding to obtain a decoded facial image; obtaining an additional image difference between the decoded facial image and a comparison facial image, and obtaining a target model loss value based on the additional image difference; the target model loss value has a positive correlation with the additional image difference; the comparison facial image includes at least one of the training source facial image or a standard facial image corresponding to the decoded facial image; adjusting model parameters of the encoder and the decoder based on the target model loss value to obtain a trained facial replacement model for performing image processing according to the facial replacement model.

[0018] In some embodiments, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0019] The above image processing method, device, computer device and storage medium obtain a training source facial image and a training template facial image, extract additional image features from the training source facial image to obtain source additional image features corresponding to the training source facial image, extract identity features from the training source facial image to obtain source identity features corresponding to the training source facial image, input the training template facial image into the encoder in the face replacement model for encoding to obtain facial attribute features, input the source additional image features, source identity features and facial attribute features into the decoder in the face replacement model for decoding to obtain a decoded facial image, obtain the additional image difference between the decoded facial image and a comparison facial image, obtain a target model loss value based on the additional image difference, and the target model loss value has a positive correlation with the additional image difference. The comparison facial image includes at least one of the training source facial image or the standard facial image corresponding to the decoded facial image. Adjust the model parameters of the encoder and decoder based on the target model loss value to obtain a trained face replacement model for image processing according to the face replacement model. Since the decoded facial image is obtained by inputting the source additional image features, source identity features and facial attribute features into the decoder in the face replacement model, and the target model loss value is obtained based on the additional image difference between the decoded facial image and the comparison facial image, and the target model loss value has a positive correlation with the additional image difference. Therefore, the trained face replacement model can improve the consistency of the identity of the face in the face-swapped image with the identity of the source facial image, improve the consistency of the attributes of the face-swapped image with the attributes of the template facial image, and improve the consistency of the additional image of the face-swapped image with the additional image of the source facial image. When using this face replacement model for image processing, the face replacement effect, that is, the face-swapping effect, can be improved.

[0020] An image processing method, the method includes: obtaining a target source facial image and a target template facial image; extracting additional image features from the target source facial image to obtain target source additional image features corresponding to the target source facial image; extracting identity features from the target source facial image to obtain target identity features corresponding to the target source facial image; inputting the target template facial image into the encoder in the trained face replacement model for encoding to obtain target facial attribute features; inputting the target source additional image features, the target identity features and the target facial attribute features into the decoder in the face replacement model for decoding to obtain a face replacement image, the face in the face replacement image matching the face of the target source facial image, and the attributes in the face replacement image matching the attributes of the target template facial image.

[0021] An image processing device, the device comprising: a second facial image acquisition module, configured to acquire a target source facial image and a target template facial image; a target source additional image feature obtaining module, configured to perform additional image feature extraction on the target source facial image to obtain a target source additional image feature corresponding to the target source facial image; a target identity feature obtaining module, configured to perform identity feature extraction on the target source facial image to obtain a target identity feature corresponding to the target source facial image; a target facial attribute feature obtaining module, configured to input the target template facial image into an encoder in a trained facial replacement model for encoding to obtain a target facial attribute feature; a facial replacement image obtaining module, configured to input the target source additional image feature, the target identity feature, and the target facial attribute feature into a decoder in the facial replacement model for decoding to obtain a facial replacement image, wherein the face in the facial replacement image matches the face in the target source facial image, and the attributes in the facial replacement image match the attributes in the target template facial image.

[0022] In some embodiments, the second facial image acquisition module includes: a target object image acquisition unit, configured to acquire a target object image corresponding to a target object with a face to be replaced; a face comparison unit, configured to determine a current video frame in a target video, and compare the current object face in the current video frame with the target object face in the target object image; a target source facial image obtaining unit, configured to, when the current object face matches the target object face, segment a matching target template facial image from the current video frame, and use a reference facial image corresponding to a reference object of the target object as the target source facial image; the device is further configured to use the facial replacement image to replace the target template facial image in the current video frame to obtain an updated current video frame.

[0023] A computer device, comprising a memory and a processor, the memory storing a computer program, and the processor, when executing the computer program, implementing the following steps: acquiring a target source facial image and a target template facial image; performing additional image feature extraction on the target source facial image to obtain a target source additional image feature corresponding to the target source facial image; performing identity feature extraction on the target source facial image to obtain a target identity feature corresponding to the target source facial image; inputting the target template facial image into an encoder in a trained facial replacement model for encoding to obtain a target facial attribute feature; inputting the target source additional image feature, the target identity feature, and the target facial attribute feature into a decoder in the facial replacement model for decoding to obtain a facial replacement image, wherein the face in the facial replacement image matches the face in the target source facial image, and the attributes in the facial replacement image match the attributes in the target template facial image.

[0024] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: obtaining a target source face image and a target template face image; extracting additional image features from the target source face image to obtain target source additional image features corresponding to the target source face image; extracting identity features from the target source face image to obtain target identity features corresponding to the target source face image; inputting the target template face image into an encoder in a trained face replacement model for encoding to obtain target face attribute features; decoding the target source additional image features, the target identity features, and the target face attribute features into a decoder in the face replacement model to obtain a face replacement image, where the face in the face replacement image matches the face of the target source face image, and the attributes in the face replacement image match the attributes of the target template face image.

[0025] In some embodiments, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0026] For the above image processing method, device, computer device, and storage medium, a target source face image and a target template face image are obtained, additional image features are extracted from the target source face image to obtain target source additional image features corresponding to the target source face image, identity features are extracted from the target source face image to obtain target identity features corresponding to the target source face image, the target template face image is input into an encoder in a trained face replacement model for encoding to obtain target face attribute features, and the target source additional image features, target identity features, and target face attribute features are decoded into a decoder in the face replacement model to obtain a face replacement image. Since the face in the face replacement image matches the face of the target source face image and the attributes in the face replacement image match the attributes of the target template face image, the consistency of the identity and additional image of the face replacement image and the target source face image is improved, and the consistency of the attribute information of the face replacement image and the target template face image is ensured, thereby improving the face replacement effect. Description of the Drawings

[0027] Figure 1 It is an application environment diagram of the image processing method in some embodiments;

[0028] Figure 2 It is a flowchart of the image processing method in some embodiments;

[0029] Figure 3 It is a schematic flowchart of an image processing method in some embodiments;

[0030] Figure 4 It is an effect diagram of face swapping in some embodiments;

[0031] Figure 5 It is a schematic diagram of obtaining a decoded face image in some embodiments;

[0032] Figure 6 It is a schematic diagram of obtaining an additional image loss value in some embodiments;

[0033] Figure 7 It is a schematic diagram of obtaining a discrimination loss value in some embodiments;

[0034] Figure 8 It is a structural block diagram of an image processing apparatus in some embodiments;

[0035] Figure 9 It is a structural block diagram of an image processing apparatus in some embodiments;

[0036] Figure 10 It is an internal structural diagram of a computer device in some embodiments. Detailed implementation manners

[0037] In order to make the objectives, technical solutions, and advantages of this application clearer, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0038] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0039] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0040] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as target recognition, tracking, and measurement, and further performs image processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision research related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0041] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0042] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0043] The solution provided in the embodiments of this application involves technologies such as computer vision technology and machine learning in artificial intelligence, and is specifically described through the following embodiments:

[0044] The image processing method provided in this application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. Specifically, the server 104 can obtain the training source facial image and the training template facial image, extract additional image features from the training source facial image to obtain the source additional image features corresponding to the training source facial image, extract identity features from the training source facial image to obtain the source identity features corresponding to the training source facial image, input the training template facial image into the encoder in the face replacement model for encoding to obtain facial attribute features, input the source additional image features, source identity features, and facial attribute features into the decoder in the face replacement model for decoding to obtain the decoded facial image, obtain the additional image difference between the decoded facial image and the comparison facial image, and obtain the target model loss value based on the additional image difference; the target model loss value is positively correlated with the additional image difference; the comparison facial image includes at least one of the training source facial image or the standard facial image corresponding to the decoded facial image, and adjust the model parameters of the encoder and decoder based on the target model loss value to obtain the trained face replacement model for image processing according to the face replacement model. For example, the server can obtain the target source facial image and the target template facial image. For example, it can obtain the target source facial image and the target template facial image from the terminal 102. For example, the terminal 102 can send a face replacement request to the server, and the face replacement request can carry the target source facial image and the target template facial image. The server responds to the face replacement request, extracts additional image features from the target source facial image to obtain the target source additional image features corresponding to the target source facial image, extracts identity features from the target source facial image to obtain the target identity features corresponding to the target source facial image, inputs the target template facial image into the encoder in the trained face replacement model for encoding to obtain the target facial attribute features, inputs the target source additional image features, target identity features, and target facial attribute features into the decoder in the face replacement model for decoding to obtain the face replacement image, the face in the face replacement image matches the face of the target source facial image, and the attributes in the face replacement image match the attributes of the target template facial image.

[0045] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0046] In some embodiments, as Figure 2 shown, a method for image processing is provided. Taking the server 104 in Figure 1 as an example for illustration, the method includes the following steps:

[0047] S202, obtain the training source facial image and the training template facial image.

[0048] Among them, the source facial image is an image providing a face, and the face of the image after face swapping is derived from the source facial image. The template facial image is an image providing a template for face swapping. The face of the source image is transplanted into the template facial image, that is, the face in the template facial image is replaced by the face in the source facial image, thereby forming the image after face swapping. The training source facial image is the source facial image for training the model, and the training template facial image is the template facial image for training the model. The training source facial image and the training template facial image are facial images of different objects. The object can be a person or an animal. For example, the training source facial image is the facial image of object A, and the training template facial image is the facial image of object B. When the object is a person, the training source facial image and the training template facial image can be the facial images of different people. The training source facial image and the training template facial image have different attribute information. The attribute information refers to the information related to the image, including but not limited to at least one of facial expression, makeup, pose, or background. For example, the training source facial image has a happy expression, and the training template facial image has a sad expression. The training source facial image is a real image, and the training template facial image can be a real image or a synthetic image. The training source facial image and the training template facial image can be video frames in a video. The training source facial image is a facial image wearing glasses, that is, the training source facial image includes glasses, and the training template facial image may or may not include glasses.

[0049] In the embodiment of the present application, when it is necessary to train the facial replacement model to be trained, the training source facial image and the training template facial image can be obtained. For example, the server can obtain the facial image of the first object as the training source facial image, and obtain the facial image of the second object as the training template facial image, where the facial image of the second object can be a real image or a synthetic image. Among them, the facial replacement model is a model for performing facial replacement. The facial replacement model to be trained can be a model that has not been trained at all, or a model that has been trained and needs further optimization. It can be a neural network model based on artificial intelligence. For example, it can be a convolutional neural network model. When the object is a person, that is, when the facial image is a human face image, the facial replacement model can also be called a face replacement model, which is used to replace the human face in the image.

[0050] In some embodiments, the training template facial image is synthesized based on the standard facial image corresponding to the training source facial image. The standard facial image is an expected real image according to the training source facial image and the model facial image. The standard facial image and the training source facial image are facial images of the same object. The standard facial image and the training source facial image have different attribute information, such as different expressions. For example, a reference facial image can be obtained. The reference facial image and the training template facial image are facial images of the same object, and the reference facial image and the training template facial image have different attribute information, such as different postures. The face in the reference facial image can be replaced into the standard facial image to obtain the training template facial image.

[0051] S204. Extract additional image features from the training source facial image to obtain the source additional image features corresponding to the training source facial image.

[0052] Among them, the additional image feature refers to the feature of the additional image on the face. The image of the object includes the additional image and the original image. The original image is the original image of the object, that is, a real image. For example, when the object is a person, the original image is determined by the person's appearance, and can be reflected by at least one of hairstyle, surface color, birthmark or mole.

[0053] The additional image does not belong to the original image of the object and can be reflected by the accessories attached to the face, such as glasses and earrings worn on the face. The accessories on the face can be called additional objects or appendages. Appendages have a greater impact on the overall image. For example, glasses have a greater impact on a person's overall image. When the additional object includes glasses and earrings, the additional image feature can include at least one of the glasses feature or the earrings feature. By extracting additional image features from different images, the additional image features corresponding to the images can be obtained. The source additional image features corresponding to the training source facial image are the additional image features obtained by extracting additional features from the training source facial image.

[0054] Specifically, the source additional image features may include source glasses features. For example, the server can extract the glasses features from the training source face image to obtain the source glasses features corresponding to the training source face image. For example, the server can use the trained additional image feature extraction network to extract additional features from the training source face image to obtain the source additional image features corresponding to the training source face image. There are various additional image feature extraction networks for extracting additional image features. For example, it may include a glasses feature extraction network for extracting glasses features, an earring feature extraction network for extracting earring features, or a pimple feature extraction network for extracting pimple features. The server can obtain multiple types of additional image feature extraction networks, input the source face image into each additional image feature extraction network respectively, and obtain the additional features output by each additional image feature extraction network to form the source additional image features corresponding to the source face image. For example, when the training source face image is a face image wearing glasses, the server can input the face image wearing glasses into the glasses feature extraction network to obtain the source glasses features corresponding to the face image wearing glasses. The trained additional image feature extraction network can be extracted from the trained additional image recognition model. For example, the feature extraction layer before the image recognition layer can be extracted from the trained additional image recognition model as the additional image feature extraction network. The additional image recognition model is used to recognize additional objects. There are various additional image recognition models. For example, it may include a glasses recognition model for recognizing glasses. The glasses recognition model can recognize the type of glasses. The additional image recognition model can be based on an artificial intelligence neural network. For example, it can be based on the resnet50 (Residual Network 50) network. Among them, the model based on resnet50 can perform convolution operations on the input image, input the features obtained by convolution into one or more residual blocks for further feature extraction, and perform image recognition based on the extracted features. For example, the additional image recognition model may include a residual feature extraction network. The additional image recognition model can use the feature extraction network before the residual feature extraction network to extract features from the training source face image, use the extracted features as the input features of the residual feature extraction network, obtain the processed features obtained by processing the input features based on the residual feature extraction network, fuse the processed features with the input features, and obtain the features finally output by the residual feature extraction network. The additional image recognition model based on the resnet50 network can improve the expression ability of features, improve the network convergence speed, and improve the accuracy of the model when the number of network layers increases.

[0055] In some embodiments, when training the additional image feature extraction network to be trained, the server can also obtain video segments that do not include the specific additional object as negative samples, and use the negative samples and positive samples to train the additional image feature extraction network. The positive samples refer to video segments that include the specific additional object. After training, the trained additional image feature extraction network is obtained. For example, the glasses feature extraction network is trained using video segments that do not include glasses and video segments that include glasses, and the trained glasses feature extraction network is obtained.

[0056] S206. Extract the identity features from the training source facial image to obtain the source identity features corresponding to the training source facial image.

[0057] Among them, the identity features refer to the features used to identify the identity of an object, and can include at least one of the facial feature features or facial contour features of the object. The facial feature features refer to the features corresponding to the facial features on the face, and the facial contour features refer to the features corresponding to the facial contour. The source identity features corresponding to the training source facial image refer to the identity features obtained by extracting the identity features from the training source facial image.

[0058] Specifically, the server can obtain a trained identity recognition model, input the training source facial image into the trained identity recognition model, and obtain the source identity features corresponding to the training source facial image. The identity recognition model is used to identify the identity of an object. When the training source facial image is a face image, the identity recognition model can be a face recognition model, and the face recognition model is used to identify the person to whom the face belongs.

[0059] S208. Input the training template facial image into the encoder of the face replacement model for encoding to obtain facial attribute features.

[0060] Among them, the face replacement model is used to replace the face, that is, to replace the face in one image (denoted as the first image) with the face in another image (denoted as the second image) to obtain the replaced facial image. By training the face replacement model, the replaced facial image can be made to have the same identity as the first image and the same facial attributes as the second image. Facial attributes have nothing to do with identity and can include at least one of posture, expression, surface color, brightness, illumination, or texture. The facial attribute features refer to the features corresponding to the facial attributes. The face replacement model can include an encoder, and the encoder is used to encode the image to obtain facial attribute features.

[0061] Specifically, the server can input the training template facial image into the encoder of the face replacement model for encoding to obtain the facial attribute features corresponding to the training template facial image.

[0062] S210. Input the source additional image features, source identity features, and facial attribute features into the decoder in the face replacement model for decoding to obtain a decoded face image.

[0063] Among them, the face replacement model may further include a decoder, which is used to generate a decoded face image by using the source additional image features, source identity features, and facial attribute features. The decoded face image is an image obtained by the decoder decoding based on the source additional image features, source identity features, and facial attribute features. Among them, the encoder and the decoder may be neural networks based on artificial intelligence. For example, they may be based on the resnet network model and may be a model composed of several network layers in resnet. For example, they may be a model composed of 8 network layers in resnet.

[0064] Specifically, the server may perform feature fusion on the source additional image features, source identity features, and facial attribute features to obtain a target fusion feature. For example, the source additional image features, source identity features, and facial attribute features may be formed into a feature triple as the target fusion feature, and the target fusion feature is input into the decoder in the face replacement model for decoding to obtain a decoded face image.

[0065] In some embodiments, the server may input the training template face image and the training source face image into the encoder in the face replacement model for encoding to obtain the facial attribute features corresponding to the training template face image and the source face features corresponding to the training source face image, and perform feature fusion on the source face features, facial attribute features, source additional image features, and source identity features to obtain a target fusion feature. Among them, the source face features may include at least one of the features corresponding to the attributes or identities of the training source face image.

[0066] In some embodiments, the encoder may include an attribute feature extraction model and a facial feature extraction model. The attribute feature extraction model is used to extract the facial attribute features corresponding to the training template face image, and the facial feature extraction model is used to extract the source face features from the training source face image.

[0067] In some embodiments, the encoder and the decoder can be neural networks in the generative network of a Generative Adversarial Network (GAN). A generative adversarial network includes a Generator model and a Discriminator model. The generative adversarial network learns by having the Generator model and the Discriminator model play against each other to obtain a desired machine learning model, which is a method of unsupervised learning. The goal of the Generator model is to obtain the desired output based on the input. The goal of the Discriminator model is to distinguish the output of the Generator model from real images as much as possible. The input of the Discriminator model includes the output of the Generator model and real images. The two network models learn by competing against each other and continuously adjust their parameters. The ultimate goal is for the Generator model to deceive the Discriminator model as much as possible so that the Discriminator model cannot determine whether the output result of the Generator model is real.

[0068] S212, obtain the additional image difference between the decoded facial image and the comparison facial image, and obtain the target model loss value based on the additional image difference; the target model loss value is positively correlated with the additional image difference; the comparison facial image includes at least one of the training source facial image or the standard facial image corresponding to the decoded facial image.

[0069] Among them, the comparison facial image can include at least one of the training source facial image or the standard facial image corresponding to the decoded facial image. The standard facial image corresponding to the decoded facial image is consistent with the training source facial image in terms of identity and additional objects, and is consistent with the training template facial image in terms of attributes. The standard facial image corresponding to the decoded facial image is the image that is expected to be generated corresponding to the decoded facial image, and can also be understood as a label, that is, it is hoped that the decoded facial image is as consistent as possible with the standard facial image. It can be an image obtained by real shooting or a synthesized image.

[0070] The additional image difference is used to reflect the difference between the decoded facial image and the additional objects included in the comparison facial image. It can include the difference between features or the difference between pixel values. For example, it can include the difference between the target additional image features and the comparison additional image features. The target additional image features are the additional image features obtained by extracting the additional image features from the decoded facial image, and the comparison additional image features are the additional image features obtained by extracting the additional image features from the comparison facial image. The additional image difference can also include the difference in pixel values between the additional image region and the matching image region. The additional image region refers to the region where the additional object is located in the comparison facial image, and the matching image region refers to the region in the decoded facial image that matches the position of the additional image region, such as the region with the same position as the additional image region. The comparison additional image features can include at least one of the source additional image features or the standard additional image features. The standard additional image features are the additional image features obtained by extracting the additional image features from the standard facial image. The additional image difference can include the difference between the target additional image features and the source additional image features, and can also include the difference between the target additional image features and the standard additional image features.

[0071] The target model loss value is calculated based on the additional image difference and is positively correlated with the additional image difference. For example, the target model loss value can be the additional image difference, or it can be obtained by performing a linear operation or a non-linear operation on the additional image difference. The linear operation can include at least one of addition, subtraction, multiplication, or division. The non-linear operation can include at least one of logarithmic operation, square root operation, exponential operation, or trigonometric function operation. Among them, the loss value is obtained based on the loss, and the loss function is a function used to represent the "risk" or "loss" of an event. Among them, the positive correlation relationship means that: under the condition that other conditions remain unchanged, the change directions of two variables are the same. When one variable changes from large to small, the other variable also changes from large to small. It can be understood that the positive correlation relationship here means that the change directions are the same, but it does not require that when one variable changes a little, the other variable must also change. For example, it can be set that when variable a is from 10 to 20, variable b is 100, and when variable a is from 20 to 30, variable b is 120. In this way, the change directions of a and b are both that when a becomes larger, b also becomes larger. However, within the range of a from 10 to 20, b can remain unchanged.

[0072] Specifically, the server can input the decoded facial image into the trained additional image feature extraction network to obtain the target additional image features corresponding to the decoded facial image, and input the comparison facial image into the trained additional image feature extraction network to obtain the comparison additional image features corresponding to the comparison facial image.

[0073] In some embodiments, the comparison facial image is a standard facial image. The server can obtain the non-attachment image area from the standard facial image, obtain the area in the decoded facial image that matches the position of the non-attachment image area to get the matching non-attachment area, obtain the non-attachment image difference between the non-attachment image area and the matching non-attachment area, and obtain the target model loss value based on the attachment image difference and the non-attachment image difference. The process of obtaining the non-attachment image difference can refer to the relevant content of obtaining the attachment image difference, which will not be elaborated here. Among them, the non-attachment image area refers to the area in the standard facial image except the attachment image area.

[0074] In some embodiments, the target model loss value may further include an identity loss value. The server can obtain the target identity feature corresponding to the decoded facial image, calculate the difference between the target identity feature and the comparison identity feature to obtain the identity loss value. The difference between the target identity feature and the comparison identity feature is positively correlated with the identity loss value. The target identity feature is the identity feature obtained by extracting the identity feature from the decoded facial image. The comparison identity feature is the identity feature obtained by extracting the identity feature from the comparison facial image. The comparison identity feature may include at least one of the source identity feature or the standard identity feature. The standard identity feature is the identity feature obtained by extracting the identity feature from the standard facial image.

[0075] In some embodiments, the server can calculate the similarity between the target identity feature and the comparison identity feature to obtain the identity feature similarity, and obtain the identity loss value based on the identity feature similarity. The difference between the target identity feature and the comparison identity feature is negatively correlated with the identity feature similarity, and the identity loss value is negatively correlated with the identity feature similarity. For example, when the comparison identity feature is the source identity feature, the server can use formula (1) to calculate the identity loss value id_loss, where id_loss represents the identity loss value, result_id_feature represents the target identity feature, and src_id_feature represents the source identity feature. cosine_similarity(result_id_feature,src_id_feature) represents the cosine similarity (cosine similarity) between result_id_feature and src_id_feature, that is, the identity feature similarity.

[0076] id_loss = 1 - cosine_similarity(result_id_feature,src_id_feature) (1)

[0077] Among them, the calculation formula of cosine similarity can be expressed as formula (2) for example, where A and B are respectively a vector, and A i and B i respectively represent the components of vector A and vector B. similarity and cos(θ) represent cosine similarity. θ represents the included angle between vector A and B.

[0078]

[0079] The negative correlation relationship means that: under the condition that other conditions remain unchanged, the changing directions of two variables are opposite. When one variable changes from large to small, the other variable changes from small to large. It can be understood that the negative correlation relationship here refers to the opposite changing directions, but it does not require that when one variable changes a little, the other variable must also change.

[0080] S214. Adjust the model parameters of the encoder and the decoder based on the target model loss value to obtain a trained face replacement model, so as to perform image processing according to the face replacement model.

[0081] Among them, the model parameters refer to the variable parameters inside the model. For a neural network model, it can also be called neural network weights.

[0082] Specifically, the server can adjust the model parameters of the encoder and the decoder based on the target model loss value to jointly train the encoder and the decoder, obtain a trained encoder and a trained decoder, and based on the trained encoder and the trained decoder, obtain a trained face update model. The trained face update model includes the trained encoder and the trained decoder.

[0083] In some embodiments, a gradient descent method can be adopted, such as the gradient descent method based on Adam, to adjust the model parameters in the face update model in the direction of decreasing the target model loss value to obtain a trained face update model. After obtaining the trained face update model, the trained face update model can be used to perform image processing. The facial features and contours in the first face image can be replaced into the second face image to achieve face replacement for the second face image.

[0084] In some embodiments, the image processing method provided in this application can be used and applied to video face replacement to replace the face of the first person in the video with the face of the second person, so as to achieve face replacement for the person in the video, that is, to achieve video face replacement.

[0085] Video face swapping refers to replacing the face in a face image (denoted as the original face image) with the face in another face image (denoted as the face image before replacement), obtaining the face image after replacement, and ensuring that the face image after replacement is consistent with the face image before replacement in terms of expression, angle, and background, and that the identity corresponding to the face image after replacement is the same as that of the original face image. The original face image is face image a, the face image before replacement is face image b, and the face image after replacement is face image c. It is obvious that face image c is consistent with face image b in terms of expression, angle, and background, and face image c is consistent with face image a in terms of identity, that is, they are the faces of the same person.

[0086] Video face swapping can be applied in various scenarios, such as film and television portrait production, game character design, virtual avatars, and privacy protection. In film and television production, some professional actions are usually completed by professional personnel. After shooting a video of a professional person, face swapping technology can be used to replace the face of the professional person in the video with the face of a real actor, or replace the face of an actor with misdeeds in the video, thus saving the cost of film and television production. In game production, the generation and transformation of game art images and the production of art resources require a lot of costs. Face swapping technology can be used to generate characters in a specific style, thus helping to save costs for the art department. In live streaming, users can replace their faces with the faces of virtual characters, thus protecting the privacy of users and also increasing the fun of live streaming.

[0087] It can be understood that the training of the model can be iterative, that is, the trained face replacement model can be iteratively trained, and the training stops when the model convergence condition is met. The model convergence condition can include that the change in the model loss value is less than the preset loss value change, or it can also include that the change in the model parameters is less than the preset parameter change value. For example, when there are multiple training samples composed of source face images and template face images, training can be carried out multiple times, and each time one or more training samples are used for model training. Multiple means at least two.

[0088] In the above image processing method, a training source facial image and a training template facial image are obtained. The training source facial image is subjected to additional image feature extraction to obtain the source additional image features corresponding to the training source facial image. The training source facial image is subjected to identity feature extraction to obtain the source identity features corresponding to the training source facial image. The training template facial image is input into the encoder in the face replacement model for encoding to obtain facial attribute features. The source additional image features, the source identity features, and the facial attribute features are input into the decoder in the face replacement model for decoding to obtain a decoded facial image. The additional image difference between the decoded facial image and the comparison facial image is obtained. Based on the additional image difference, a target model loss value is obtained. The target model loss value is positively correlated with the additional image difference. The comparison facial image includes at least one of the training source facial image or the standard facial image corresponding to the decoded facial image. Based on the target model loss value, the model parameters of the encoder and the decoder are adjusted to obtain a trained face replacement model for image processing according to the face replacement model. Since the decoded facial image is obtained by inputting the source additional image features, the source identity features, and the facial attribute features into the decoder in the face replacement model, and the target model loss value is obtained based on the additional image difference between the decoded facial image and the comparison facial image, and the target model loss value is positively correlated with the additional image difference. Therefore, the trained face replacement model can improve the consistency between the identity of the face-swapped image and the identity of the source facial image, improve the consistency between the attributes of the face-swapped image and the template facial image, and improve the consistency between the additional image of the face-swapped image and the additional image of the source facial image. When using this face replacement model for image processing, the face replacement effect, that is, the face-swapping effect, can be improved.

[0089] In some embodiments, the additional image difference includes a first image feature difference. Obtaining the additional image difference between the decoded facial image and the comparison facial image and obtaining the target model loss value based on the additional image difference includes: performing additional image feature extraction on the decoded facial image to obtain the target additional image features corresponding to the decoded facial image; determining the image feature difference between the source additional image features and the target additional image features as the first image feature difference; and obtaining the target model loss value based on the first image feature difference.

[0090] Among them, the image feature difference refers to the difference between image features, and the first image feature difference refers to the difference between the source additional image features and the target additional image features. The target model loss value is positively correlated with the first image feature difference.

[0091] Specifically, the server can input the decoded facial image into the trained additional image feature extraction network to obtain the target additional image feature corresponding to the decoded facial image, calculate the similarity between the source additional image feature and the target additional image feature to obtain the additional feature similarity, and obtain the first image feature difference based on the additional feature similarity. The additional feature similarity and the first image feature difference are negatively correlated. For example, the opposite number or reciprocal of the additional feature similarity can be used as the first image feature difference. The additional feature similarity can be the cosine similarity.

[0092] In some embodiments, the server can obtain the additional feature loss value based on the first image feature difference. The additional feature loss value is positively correlated with the first image feature difference and negatively correlated with the additional feature similarity. The first image feature difference can be used as the additional feature loss value, or the first image feature difference can be linearly or non-linearly calculated to obtain the additional feature loss value. The target model loss value is obtained based on the additional feature loss value, and the target model loss value is positively correlated with the additional feature loss value.

[0093] In some embodiments, the server can use the result of subtracting the preset value from the additional feature similarity as the first image feature difference and use the first image feature difference as the additional feature loss value. The preset value can be set in advance as needed. For example, it can be 1. When the additional image feature is the glasses feature, the source additional image feature can also be called the source glasses feature, the target additional image feature can also be called the target glasses feature, the additional feature similarity can also be called the glasses similarity, and the additional feature loss value can also be called the glasses loss value. For example, the glasses loss value can be calculated using formula (3). Among them, glass_loss represents the glasses loss value, result_glass_feature represents the target glasses feature, and src_glass_feature represents the source glasses feature. cosine_similarity(result_glass_feature,src_glass_feature) represents the cosine similarity between result_glass_feature and src_glass_feature, that is, the additional feature similarity.

[0094] glass_loss = 1 - cosine_similarity(result_glass_feature,src_glass_feature) (3)

[0095] In this embodiment, the image feature difference between the source additional image feature and the target additional image feature is determined as the first image feature difference, and the target model loss value is obtained based on the first image feature difference. Since the target model loss value is negatively correlated with the first image feature difference, the smaller the difference between the source additional image feature and the target additional image feature, the smaller the target model loss value. Therefore, when adjusting the model parameters to reduce the target model loss value, the difference between the source additional image feature and the target additional image feature will be reduced, thereby improving the similarity in appearance between the decoded facial image and the training source facial image and enhancing the facial replacement effect.

[0096] In some embodiments, obtaining the additional image difference between the decoded facial image and the comparison facial image and obtaining the target model loss value based on the additional image difference includes: identifying the additional image of the comparison facial image to obtain the additional image region corresponding to the comparison facial image; obtaining the additional image enhancement value corresponding to the additional image region; determining the image difference between the additional image region and the image region at the corresponding position in the decoded facial image as the additional image difference; obtaining the additional image loss value based on the additional image difference, and using the additional image enhancement value to perform enhancement processing on the additional image loss value to obtain the target model loss value.

[0097] Among them, the additional image region refers to the region where the additional object is located. Identifying the additional image of the comparison facial image means determining the region of the additional object in the comparison facial image. The additional image difference may include the image difference between the additional image region and the image region at the corresponding position in the decoded facial image, and the image region at the corresponding position in the decoded facial image refers to the matching image region.

[0098] The additional image loss value is obtained according to the additional image difference, and the additional image loss value is positively correlated with the additional image difference. For example, the additional image difference can be used as the additional image loss value, or the additional image loss value can be obtained by performing linear or non-linear operations on the additional image difference. The additional image enhancement value can be a pre-set value, such as 6, for performing enhancement processing on the additional image loss value. Performing enhancement processing on the additional image loss value using the additional image enhancement value may include: performing at least one of addition operation or multiplication operation using the additional image enhancement value and the additional image loss value.

[0099] Specifically, the server can identify additional objects in the comparison facial image, determine the area where the additional object is located from the comparison facial image to obtain an additional image area. For example, a trained additional area recognition model can be obtained. The additional area recognition model can identify the area where the additional object is located from the image. Of course, the additional area recognition model can also identify other areas of the face from the image, such as the area where the mouth is located. When the comparison facial image is a human face image, the additional area recognition model can be a face segmentation model, which is used to segment the human face to obtain areas in the human face, such as obtaining the glasses area in the human face, that is, the area where the glasses are located. The server can determine the area position information of the additional image area from the comparison facial image, determine the area corresponding to the area position information from the decoded facial image to obtain a matching image area, calculate the difference between the additional image area and the matching image area to obtain an additional image difference.

[0100] In some embodiments, the image difference can include differences in pixels or differences in features. For example, the statistical value of the differences in pixel values at corresponding positions between the additional image area and the matching image area can be calculated to obtain a difference statistical value, which reflects the differences in pixels. The server can also extract features from the additional image area to obtain extracted additional image features, extract features from the matching image area to obtain decoded image features, calculate the difference between the extracted additional image features and the decoded image features to obtain a second image feature difference, which reflects the differences in features. The additional image difference can also include at least one of the second image feature difference or the difference statistical value.

[0101] In some embodiments, the server can obtain an additional feature-level loss value based on the second image feature difference and an additional pixel-level loss value based on the difference statistical value. The additional feature-level loss value is positively correlated with the second image feature difference, and the additional pixel-level loss value is positively correlated with the difference statistical value. The additional image loss value includes at least one of the additional pixel-level loss value or the additional feature-level loss value.

[0102] In some embodiments, the comparison facial image is a standard facial image. The server can use an additional image enhancement value to enhance the additional image loss value to obtain an enhanced additional image loss value, can obtain a non-additional image loss value based on the non-additional image difference, obtain a non-additional image enhancement value corresponding to the non-additional image area, use the non-additional image enhancement value to enhance the non-additional image loss value to obtain an enhanced non-additional image loss value, and perform weighted calculation according to the enhanced additional image loss value and the enhanced non-additional image loss value to obtain a target model loss value. Among them, the non-additional image loss value is positively correlated with the non-additional image difference.

[0103] In some embodiments, the server can obtain a mask of the additional image region, such as the mask corresponding to the glasses region, by performing masking processing on the image, so as to segment the additional image region from the image. The mask of the additional image region may include the mask values corresponding to the pixel points in the additional image region (denoted as the first mask value), and may also include the mask values corresponding to the pixel points in the non-additional image region (denoted as the second mask value). The first mask value is greater than the second mask value. The first mask value and the second mask value can be set as needed. For example, the first mask value is 1 and the second mask value is 0. The non-additional image region refers to the region in the image other than the additional image region. For example, when the standard facial image is a facial image including glasses, the standard facial image can be subjected to masking processing to obtain the mask of the glasses region. The mask of the glasses region can be expressed as glass_mask = segmentation(gt_img) (4), where glass_mask represents the mask of the glasses region, gt_img represents the standard facial image, and segmentation(·) represents performing masking processing on the image to obtain the mask of the glasses region. Glass_mask may include the mask values corresponding to the pixel points in the glasses region and the mask values corresponding to the pixel points in the non-glasses region.

[0104] In some embodiments, the comparison facial image is the standard facial image. The server can calculate based on the mask of the additional image region to obtain the additional image enhancement value and the non-additional image enhancement value. For example, the first mask value corresponding to the additional image region can be determined from the mask of the additional region, and the additional image enhancement value can be obtained according to the first mask value. The additional image enhancement value is positively correlated with the first mask value. Similarly, the non-additional image enhancement value can be obtained, and the non-additional image enhancement value is positively correlated with the second mask value corresponding to the non-additional image region. When the additional image region is the glasses region, the additional image enhancement value can be referred to as the glasses enhancement value, and the non-additional image enhancement value can be referred to as the non-glasses enhancement value. For example, the glasses enhancement value and the non-glasses enhancement value can be calculated using formula (5), where glass_weight is a preset value that can be set as needed, for example, it can be 5. mask_weight includes the glasses enhancement value and the non-glasses enhancement value. The glasses enhancement value in mask_weight is calculated according to the mask value corresponding to the glasses region in glass_mask, and the non-glasses enhancement value in mask_weight is calculated according to the mask value corresponding to the non-glasses region in glass_mask. For example, when the mask value corresponding to the glasses region is 1, the mask value corresponding to the non-glasses region is 0, and glass_weight is 5, the glasses enhancement value is 6 and the non-glasses enhancement value is 1. mask_weight = (1 + glass_weight * glass_mask) (5).

[0105] For example, the enhanced non - additional image loss value and the enhanced additional image loss value can be calculated using formula (6). Among them, result represents the decoded facial image, gt_img represents the standard facial image, |result - gt_img| represents the pixel difference between the decoded facial image and the standard facial image, including the non - additional image loss value and the additional image loss value. The Reconstruction_loss includes the enhanced non - additional image loss value and the enhanced additional image loss value. The enhanced non - additional image loss value in Reconstruction_loss is the product of the non - additional image enhancement value (i.e., the non - glasses enhancement value) in mask_weight and the non - additional image loss value in |result - gt_img|. The enhanced additional image loss value in Reconstruction_loss is the product of the additional image enhancement value (i.e., the glasses enhancement value) in mask_weight and the additional image loss value in |result - gt_img|. Reconstruction_loss can be called the reconstruction loss function. Reconstruction_loss = mask_weight * |result - gt_img| (6).

[0106] In this embodiment, by using the additional image enhancement value to enhance the additional image loss value to obtain the target model loss value, the amplification of the additional image loss value can be achieved. When adjusting the model parameters through the target model loss value, it is beneficial to enhance the additional image area, such as enhancing the glasses area, and can better obtain the retention effect of the glasses, that is, improve the similarity between the additional object in the decoded facial image and the additional object in the training source facial image.

[0107] In some embodiments, determining the image difference between the additional image area and the image area at the corresponding position in the decoded facial image as the additional image difference includes: obtaining the additional pixel points in the additional image area, and obtaining the decoded pixel points in the decoded facial image that match the positions of the additional pixel points; calculating the pixel value difference between the additional pixel points and the decoded pixel points; statistically analyzing the pixel value difference corresponding to the additional image area to obtain a difference statistical value, and taking the difference statistical value as the additional image difference.

[0108] Among them, the additional pixel points are the pixel points in the additional image area. The decoded pixel points are the pixel points in the decoded facial image that match the positions of the additional pixel points. The pixel value difference refers to the difference in pixel values between the additional pixel points and the decoded pixel points. The difference statistical value is the statistical value corresponding to each pixel value difference, such as the sum result or the average value.

[0109] Specifically, the server can obtain the additional pixel values corresponding to the additional pixel points, obtain the decoded pixel values corresponding to the decoded pixel points that match the positions of the additional pixel points from the decoded facial image, calculate the difference between the additional pixel values and the decoded pixel values to obtain the pixel value difference, perform statistical operations on the respective pixel value differences, such as summation operations or mean operations, to obtain the difference statistical value, and obtain the additional image difference based on the difference statistical value. For example, the difference statistical value can be used as the additional image difference, or the difference statistical value and the first image feature difference can be used as the additional image difference.

[0110] In this embodiment, by performing statistics on the pixel value differences corresponding to the additional image area to obtain the difference statistical value, it is possible to determine the difference in pixel values between the additional image area in the decoded facial image and the corresponding area in the comparison facial image. Using the difference statistical value as the additional image difference can accurately reflect the difference in appearance and improve the accuracy of the additional image difference.

[0111] In some embodiments, determining the image difference between the additional image area and the image area at the corresponding position in the decoded facial image as the additional image difference includes: extracting features from the additional image area to obtain the extracted additional image features; extracting features from the image area corresponding to the additional image area in the decoded facial image to obtain the decoded image features; calculating the image feature difference between the extracted additional image features and the decoded image features as the second image feature difference; and obtaining the additional image difference based on the second image feature difference.

[0112] Among them, the extracted additional image features are the features obtained by extracting features from the additional image area. The decoded image features are the features obtained by extracting features from the image area corresponding to the additional image area in the decoded facial image. The image area corresponding to the additional image area in the decoded facial image refers to the area in the decoded facial image that matches the position of the additional image area. The second image feature difference refers to the difference between the extracted additional image features and the decoded image features. The additional image difference can also include the second image feature difference.

[0113] Specifically, the steps of extracting features from the additional image area to obtain the extracted additional image features can include: extracting additional image features from the comparison facial image to obtain the comparison additional image features corresponding to the comparison facial image, and using the comparison additional image features as the extracted additional image features. The extracted additional image features can include at least one of the source additional image features or the standard additional image features. The steps of extracting features from the image area corresponding to the additional image area in the decoded facial image to obtain the decoded image features can include: extracting additional image features from the decoded facial image to obtain the decoded additional image features corresponding to the decoded facial image as the decoded image features.

[0114] In some embodiments, there may be multiple additional image features to be extracted and multiple decoded image features. For example, the server may input the decoded facial image or the matching image region into a preset neural network model, and use one or more feature extraction layers of the preset neural network model to extract features from the matching image region, obtaining the decoded image features output by each feature extraction layer. Similarly, the standard facial image or the additional image region may be input into the preset neural network model, and one or more feature extraction layers of the preset neural network model are used to extract features from the additional image region, obtaining the extracted additional image features output by each feature extraction layer. The server may calculate the difference between the decoded image features and the extracted additional image features output by the same feature extraction layer, and use the statistical value of each difference as the second image feature difference. Among them, the preset neural network model may be any trained model, which may be a model based on a convolutional neural network. For example, it may be a pre-trained alexnet network model.

[0115] In this embodiment, the image feature difference between the extracted additional image features and the decoded image features is calculated as the second image feature difference, and the additional image difference is obtained based on the second image feature difference, so that the additional image difference can accurately reflect the feature difference between the additional object in the decoded facial image and the comparison facial image, improving the accuracy of the additional image difference.

[0116] In some embodiments, obtaining the additional image loss value based on the additional image difference and performing enhancement processing on the additional image loss value using the additional image enhancement value to obtain the target model loss value includes: obtaining the additional image loss value based on the additional image difference and performing enhancement processing on the additional image loss value using the additional image enhancement value to obtain the enhanced additional image loss value; obtaining the non-additional image region corresponding to the comparison facial image, and determining the image difference between the non-additional image region and the image region at the corresponding position in the decoded facial image as the non-additional image difference; obtaining the non-additional image loss value based on the non-additional image difference, and the non-additional image enhancement value corresponding to the non-additional image loss value is less than the additional image enhancement value; obtaining the target model loss value according to the enhanced additional image loss value and the non-additional image loss value.

[0117] Among them, the enhanced additional image loss value is obtained by enhancing the additional image loss value using the additional image enhancement value. The non-additional image region refers to the region in the comparison facial image except the additional image region. The non-additional image difference refers to the image difference between the non-additional image region and the corresponding image region in the decoded facial image. The process of obtaining the non-additional image difference can refer to the relevant steps of obtaining the additional image difference, which will not be elaborated here. The non-additional image loss value is positively correlated with the non-additional image difference. The non-additional image loss value can correspond to a non-additional image enhancement value, which is used to enhance the non-additional image loss value. The non-additional image enhancement value can be preset.

[0118] Specifically, the server can use the non-additional image enhancement value to enhance the non-additional image loss value to obtain an enhanced non-additional image loss value, and perform weighted calculation based on the enhanced additional image loss value and the enhanced non-additional image loss value to obtain the target model loss value.

[0119] In some embodiments, the server can input the decoded facial image into a preset neural network model, and use one or more feature extraction layers in the preset neural network model to extract features from the decoded facial image to obtain each decoded facial feature. Since the decoded facial image includes a matching image region, the decoded facial features include decoded image features obtained by extracting features from the matching image region, and also include non-matching image features, which are features obtained by extracting features from the region outside the matching image region. Similarly, the server can input the standard facial image into the preset neural network model to obtain each standard facial feature. Since the standard facial image includes an additional image region, the standard facial features include the extracted additional image features obtained by extracting features from the additional image region. Since the standard facial image also includes a non-additional image region, the standard facial features also include non-additional image features obtained by extracting features from the non-additional image region. The server can calculate the difference between the non-additional image features and the non-matching image features obtained by the same layer of feature extraction layer, perform statistical calculation based on each difference to obtain the non-image feature difference, and obtain the non-additional image features based on the non-image feature difference. The non-additional image difference can include the non-image feature difference.

[0120] For example, each decoded facial feature can be obtained according to formula (7), and each standard facial feature can be obtained according to formula (8), where alexnet_feature(result) represents inputting the decoded facial image result into the alexnet network model and outputting the features output by result in the four feature extraction layers of the alexnet network model. result_fea1, result_fea2, result_fea3, and result_fea4 are the decoded facial features of the decoded facial image result output by each feature extraction layer in the four feature extraction layers respectively. alexnet_feature(gt_img) represents inputting the standard facial image gt_img into the alexnet network model and outputting the features output by gt_img in the four feature extraction layers of the alexnet network model. gt_img_fea1, gt_img_fea2, gt_img_fea3, and gt_img_fea4 are the standard facial features of the standard facial image gt_img output by each feature extraction layer in the four feature extraction layers respectively.

[0121] result_fea1,result_fea2,result_fea3,result_fea4=alexnet_feature(result) (7)gt_img_fea1,gt_img_fea2,gt_img_fea3,gt_img_fea4=alexnet_feature(gt_img) (8)

[0122] In some embodiments, the server may calculate the difference between the decoded facial features and the standard facial features to obtain the facial feature difference, determine the facial difference loss value based on the facial feature difference. The facial feature difference includes the second image feature difference and the non-image feature difference. The facial difference loss value includes the loss value determined based on the second image feature difference and the loss value determined based on the non-image feature difference. By performing enhancement processing on the facial feature difference, the enhanced non-attachment image loss value and the enhanced attachment image loss value can be obtained. For example, the enhanced non-attachment image loss value and the enhanced attachment image loss value can be calculated using formula (9). Among them, |result_fea1 - gt_img_fea1|, |result_fea2 - gt_img_fea2|, |result_fea3 - gt_img_fea3|, and |result_fea4 - gt_img_fea4| represent the extracted feature differences. LPIPS_loss includes the enhanced non-attachment image loss value and the enhanced attachment image loss value. LPIPS_loss can be referred to as the LPIPS (Learned Perceptual Image Patch Similarity) loss function. The LPIPS loss can also be called the perceptual loss, which is used to evaluate the difference between two images. The larger the LPIPS loss, the greater the difference between the two images; the smaller the LPIPS loss, the smaller the difference between the two images. The LPIPS loss function is a loss function at the feature level and can compare the differences between two images.

[0123]

[0124] In this embodiment, obtaining the target model loss value according to the enhanced attachment image loss value and the non-attachment image loss value can make the target model loss value include both the loss value obtained from the attachment image area and the loss value obtained from the non-attachment image area. It can not only improve the effect of the area where the attached object is located after facial replacement, but also improve the effect of the area outside the attached object, thereby improving the facial replacement effect.

[0125] In some embodiments, obtaining the target model loss value based on the attachment image difference includes: obtaining the attachment image loss value based on the attachment image difference; extracting the identity features of the decoded facial image to obtain the target identity features corresponding to the decoded facial image; obtaining the identity loss value based on the identity feature difference between the source identity features and the target identity features; and obtaining the target model loss value according to the attachment image loss value and the identity loss value.

[0126] Specifically, the target identity feature is the identity feature obtained by extracting the identity feature from the decoded facial image. The identity feature difference refers to the difference between the source identity feature and the target identity feature, and the identity loss value is positively correlated with the identity feature difference.

[0127] In this embodiment, the identity loss value is obtained based on the identity feature difference between the source identity feature and the target identity feature, and the target model loss value is obtained according to the additional image loss value and the identity loss value, which can make the identity of the decoded facial image consistent with the identity of the training source facial image, and make the additional image of the decoded facial image consistent with the additional image of the training source facial image, thereby improving the facial replacement effect.

[0128] In some embodiments, extracting the additional image feature from the training source facial image to obtain the source additional image feature corresponding to the training source facial image includes: inputting the training source facial image into the trained additional image feature extraction network to perform additional image feature extraction to obtain the source additional image feature corresponding to the training source facial image; adjusting the model parameters of the encoder and the decoder based on the target model loss value to obtain the trained facial replacement model for image processing according to the facial replacement model includes: keeping the network parameters of the additional image feature extraction network unchanged, adjusting the model parameters of the encoder and the decoder based on the target model loss value to obtain the trained facial replacement model for image processing according to the facial replacement model.

[0129] Among them, the additional image feature extraction network is used to extract the additional image feature, which can be the feature extraction layer in the additional image feature extraction network. The network parameter value refers to the variable parameter inside the network. For a neural network, the network parameter can be called the weight.

[0130] Specifically, the server can jointly train the encoder and the decoder, keep the network parameters of the trained additional image feature extraction network unchanged, and adjust the model parameters in the encoder and the decoder based on the target model loss value, so that the target model loss value continuously decreases, thereby reducing the difference between the identity feature of the decoded facial image and the identity feature of the training source facial image, and reducing the difference between the attribute feature of the decoded facial image and the training template facial image, thereby improving the facial replacement effect.

[0131] In this embodiment, keeping the network parameters of the additional image feature extraction network unchanged, adjusting the model parameters of the encoder and the decoder based on the target model loss value to obtain the trained facial replacement model for image processing according to the facial replacement model improves the facial replacement effect.

[0132] In some embodiments, the steps of obtaining the additional image feature extraction network include: acquiring an additional object image, where the additional object image includes an additional object corresponding to the additional image feature; training a to-be-trained additional image recognition model using the additional object image to obtain a trained additional image recognition model; and extracting a feature extraction layer before the image recognition layer from the trained additional image recognition model as the additional image feature extraction network.

[0133] Among them, the additional object image refers to an image including an additional object, such as an image including glasses. The additional image recognition model may include a feature extraction layer and an image recognition layer. The image recognition layer is used to recognize the additional object image according to the features extracted by the feature extraction layer, such as for recognizing the type of the additional object in the additional object image.

[0134] Specifically, the server may acquire a to-be-trained additional image recognition model and train the to-be-trained additional image recognition model using a video segment including an additional object to obtain a trained additional image recognition model. The type of the additional object that the additional image recognition model can recognize may be the type of the additional object corresponding to the video segment during training. For example, the server may detect a specific additional object in multiple collected video segments. When it is detected that the video segment includes the specific additional object, the video segment is marked to obtain one or more video segments with marking information. "Multiple" means at least two. The same additional object may correspond to multiple categories. For example, glasses may correspond to multiple types of glasses, such as sunglasses and myopia glasses. The marking information may be the category information corresponding to the specific additional object and is used to identify the category corresponding to the specific additional object. The categories of the specific additional objects corresponding to different video segments may be the same or different. The marked video segments are used to train the to-be-trained additional image recognition model to obtain a trained additional image recognition model. The additional image recognition model may be a binary classification network or a multi-classification network. "Multi-" means at least three. The specific additional object may be any additional object. The objects in the collected video segments may be the same. For example, video segments may be collected for the same person. For example, when the specific additional object is glasses, the server may acquire multiple video segments including glasses, mark the video segments according to the type of glasses in the video segments to obtain video segments with marking information. The marking information is, for example, video_1, video_2, video_3,..., video_n. Different marking information corresponds to different glasses categories. The video segments with marking information are used to train the glasses recognition model to obtain a trained glasses recognition model, and the feature extraction layer is extracted from the glasses recognition model to obtain the glasses feature extraction network.

[0135] In some embodiments, when training an additional image recognition model to be trained, the server may also obtain video segments that do not include the specific additional object as negative samples, and use the negative samples and positive samples to train the additional image recognition model. The positive samples refer to video segments that include the specific additional object, and the trained additional image recognition model is obtained. For example, the glasses feature extraction network is trained using video segments that do not include glasses and video segments that include glasses, and the trained glasses feature extraction network is obtained.

[0136] In this embodiment, the feature extraction layer before the image recognition layer is extracted from the trained additional image recognition model as the additional image feature extraction network, so that a network capable of accurately extracting additional image features can be obtained, improving the accuracy of feature extraction.

[0137] In some embodiments, the server may input the decoded facial image into a discriminator. The discriminator is used to determine whether the decoded facial image belongs to a real image, and obtain the decoded discrimination result of the discriminator. The decoded discrimination result may include a decoded discrimination probability, and the decoded discrimination probability refers to the probability that the decoded facial image belongs to a real image. Based on the decoded discrimination probability, a generation loss value is obtained. The target model loss value may also be positively correlated with the generation loss value, and the generation loss value is negatively correlated with the decoded discrimination probability. For example, the generation loss value can be expressed as formula (10), where D(result) represents the decoded discrimination probability and G_loss represents the generation loss value. G_loss = log(1 - D(result)) (10)

[0138] In some embodiments, the target model loss value can be expressed as formula (11), where id_loss represents the identity loss value, glass_loss represents the additional feature loss value (i.e., the glasses loss value), Reconstruction_loss includes the enhanced non-additional image loss value and the enhanced additional image loss value, and LPIPS_loss includes the enhanced non-additional image loss value and the enhanced additional image loss value. G_loss represents the generation loss value.

[0139] loss = id_loss + glass_loss + Reconstruction_loss + LPIPS_loss + G_loss (11)

[0140] In some embodiments, the server may input a standard facial image into a discriminator to obtain a standard discrimination result. The standard discrimination result may include a standard discrimination probability, which refers to the probability that the standard facial image belongs to a real image. Based on the standard discrimination probability and the decoded discrimination probability, a discrimination loss value is obtained. The discrimination loss value is negatively correlated with the standard discrimination probability and positively correlated with the decoded discrimination probability. The server may adjust the model parameters of the discriminator based on the discrimination loss value to obtain a trained discriminator, so that the discriminator can correctly determine whether an image is a real image. The generation loss value and the discrimination loss value may be used to obtain a generative adversarial loss value, that is, the generative adversarial loss value may include the generation loss value and the discrimination loss value.

[0141] In some embodiments, the discriminator may be a multi-scale discriminator. The server may perform scale transformation on the decoded facial image to obtain decoded facial images of multiple scales, such as obtaining a decoded facial image of the first scale, a decoded facial image of the second scale, and a decoded facial image of the third scale. Among them, the first scale, the second scale, and the third scale may be set as needed. For example, the first scale may be the original scale of the decoded facial image, the second scale may be 1 / 2 of the first scale, and the third scale may be 1 / 2 of the second scale. Similarly, a standard facial image of the first scale, a standard facial image of the second scale, and a standard facial image of the third scale may be obtained. The server may input the decoded facial images of each scale into the multi-scale discriminator to obtain the decoded discrimination probabilities corresponding to the decoded facial images of each scale. Similarly, the standard discrimination probabilities corresponding to the standard facial images of each scale may be obtained. The discrimination loss value is obtained from the respective decoded discrimination probabilities and the respective standard discrimination probabilities. For example, the discrimination loss value may be expressed as formula (11), where D(gt_img), D(gt_img_1 / 2), and D(gt_img_1 / 4) respectively represent the discrimination probabilities obtained by inputting the standard facial image of the original size, the standard facial image of 1 / 2 the original size, and the standard facial image of 1 / 4 the original size into the multi-scale discriminator. D(result), D(result_1 / 2), and D(result_1 / 4) respectively represent the discrimination probabilities obtained by inputting the decoded facial image of the original size, the decoded facial image of 1 / 2 the original size, and the decoded facial image of 1 / 4 the original size into the discriminator. The original size of the standard facial image and the original size of the decoded facial image may be the same.

[0142]

[0143] In some embodiments, as Figure 3 shown, an image processing method is provided, and this method is applied to Figure 1Taking the server 104 in [description] as an example, it includes the following steps: S302, obtaining a target source facial image and a target template facial image. S304, performing additional image feature extraction on the target source facial image to obtain a target source additional image feature corresponding to the target source facial image. S306, performing identity feature extraction on the target source facial image to obtain a target identity feature corresponding to the target source facial image. S308, inputting the target template facial image into the encoder in the trained facial replacement model for encoding to obtain a target facial attribute feature. S310, inputting the target source additional image feature, the target identity feature, and the target facial attribute feature into the decoder in the facial replacement model for decoding to obtain a facial replacement image, where the face in the facial replacement image matches the face in the target source facial image, and the attributes in the facial replacement image match the attributes in the target template facial image.

[0144] Among them, the target source facial image and the target template facial image are facial images of different objects, such as face images of different people. The target source additional image feature is the feature obtained by performing additional image feature extraction on the target source facial image, and the target identity feature is the feature obtained by performing identity feature extraction on the target source facial image. The target facial attribute feature is the attribute feature obtained by encoding the target template facial image using the encoder. The identity of the facial replacement image is the same as the identity of the target source facial image. The face in the facial replacement image matching the face in the target source facial image means that the identity of the facial replacement image is the same as the identity of the target source facial image. The facial replacement image is consistent with the target source facial image in terms of additional image.

[0145] Specifically, the server can input the target source additional image feature, the target identity feature, and the target facial attribute feature into the decoder in the facial replacement model for decoding, so as to replace the face of the target source facial image into the target template facial image to obtain a facial replacement image, and make the facial replacement image consistent with the target source facial image in terms of identity and additional image, and consistent with the target template facial image in terms of attributes.

[0146] As Figure 4 shown, the target source facial image can be, for example, the face image (a) in Figure 4 , and the target template facial image can be, for example, the face image (b) in Figure 4 , and the facial replacement image can be, for example, the one in Figure 4The face image (c) therein is obtained by replacing the face in the face image (a) with the face in the face image (b). It can be seen from the face image (c) that the identity and additional image of the face image (c) are consistent with those of the face image (b), that is, the face image (c) and the face image (b) are the faces of the same person, and the face image (c) includes the same glasses as the face image (b). The attributes of the face image (c) are consistent with those of the face image (a). For example, it can be seen from the face image (c) that the hairstyles of the face image (c) and the face image (a) are consistent, and the opening angle of the mouth in the face image (c) is larger than that in the face image (b), thus conforming to the opening angle of the mouth in the face image (a).

[0147] In the above image processing method, a target source facial image and a target template facial image are obtained. The additional image features of the target source facial image are extracted to obtain the target source additional image features corresponding to the target source facial image. The identity features of the target source facial image are extracted to obtain the target identity features corresponding to the target source facial image. The target template facial image is input into the encoder of the trained face replacement model for encoding to obtain the target facial attribute features. The target source additional image features, the target identity features, and the target facial attribute features are input into the decoder of the face replacement model for decoding to obtain the face replacement image. Since the face in the face replacement image matches the face in the target source facial image, and the attributes in the face replacement image match the attributes in the target template facial image, the consistency of the identity and additional image between the face replacement image and the target source facial image is improved, and the consistency of the attribute information between the face replacement image and the target template facial image is ensured, thereby improving the face replacement effect.

[0148] In some embodiments, obtaining the target source facial image and the target template facial image includes: obtaining the target object image corresponding to the target object of the face to be replaced; determining the current video frame in the target video, and comparing the current object face in the current video frame with the target object face in the target object image; when the current object face matches the target object face, segmenting the matching target template facial image from the current video frame, and using the reference facial image corresponding to the reference object of the target object as the target source facial image; the method further includes: replacing the target template facial image in the current video frame with the face replacement image to obtain the updated current video frame.

[0149] Among them, the target object refers to the object whose face is to be replaced. For example, when the face of person A needs to be replaced, the target object can be person A. The target object image refers to an image including the facial region of the target object. The target video refers to a video including the target object, that is, the target object to be face-replaced is included in the target video. The target video can be, for example, a movie clip in which an actor needs to have their face replaced. The current video frame can be any video frame in the target video, and the current object face refers to the facial region of the current object included in the current video frame, and the current object refers to the object included in the current video frame. The reference object of the target object refers to the object to which the face used for face-replacing the target object belongs. For example, when the face of person A needs to be replaced with the face of person B, then person B is the reference object of person A. The reference face image refers to an image including the facial region of the reference object.

[0150] Specifically, the server can extract the identity features of the current object face to obtain the current object identity features, determine the identity of the current object based on the current object identity features, extract the identity features of the target object face to obtain the target object identity features, determine the identity of the target object based on the target object identity features, and when the identity of the current object is consistent with the identity of the target object, determine that the current object face matches the target object face.

[0151] In some embodiments, when it is determined that the current object face matches the target object face, the server can segment the facial region corresponding to the current object from the current video frame as the target template face image, and can segment the facial region of the reference object from the image including the reference object to obtain the reference face image as the target source face image.

[0152] In some embodiments, the server can use the face replacement image to replace the target template face image in the current video frame to obtain the updated current video frame, and based on each replaced current video frame, obtain the replaced video corresponding to the target video, so that the face of the target object in the target video is replaced with the face of the reference object, realizing video face replacement.

[0153] In this embodiment, using the face replacement image to replace the target template face image in the current video frame to obtain the updated current video frame, so that the face of the target object in the target video is replaced with the face of the reference object. The face replacement image is consistent with the target source face image in terms of identity and additional image, and is consistent with the target template face image in terms of attributes, so the effect of video face replacement is improved.

[0154] This application also provides an application scenario, which is film and television production, and this application scenario applies the above image processing method. Specifically, the application of this image processing method in this application scenario is as follows:

[0155] 1. The face replacement model to be trained can be trained using the image processing method provided by the present application to obtain a trained face replacement model. The face replacement model includes an encoder and a decoder.

[0156] Among them, a first face image of a first person can be obtained from a captured video as a training source face image, and a second face image of a second person can be obtained as a training template face image. The second person image can be a synthesized image or a real captured image. A third face image of the first person is obtained from the captured video as a standard face image. The third face image has the same attributes as the second face image and the same additional image as the first face image, that is, the additional objects on the face in the third face image are the same as those on the face in the first face image.

[0157] Specifically, obtain the source additional image features and source identity features of the training source face image, input the training source face image and the training template face image into the encoder in the face replacement model for encoding to obtain encoded features. The encoded features may include the face attribute features corresponding to the training template face image and the image features of the training source face image. Input the feature triple composed of the source additional image features, source identity features, and encoded features into the decoder in the face replacement model for decoding to obtain a decoded face image. Obtain the identity features of the decoded face image to get the target identity features, and obtain the additional image features of the decoded face image to get the target additional image features. Based on the difference between the target identity features and the source identity features, obtain an identity loss value. Based on the difference between the target additional image features and the source additional image features, obtain a first additional image loss value. Segment the additional image region from the standard face image, obtain the region that matches the position of the additional image region from the decoded face image to get a matching image region, and based on the difference in pixel values between the additional image region and the matching image region, obtain a second additional image loss value. Based on the difference in features between the additional image region and the matching image region, obtain a third additional image loss value. Perform enhancement processing on the second additional image loss value to obtain an enhanced second additional image loss value, and perform enhancement processing on the third additional image loss value to obtain an enhanced third additional image loss value. Input the decoded face image into the discriminator to obtain a decoded discrimination result, and based on the decoded discrimination result, obtain a generation loss value. Input the standard face image into the discriminator to obtain a standard discrimination result, and based on the decoded discrimination result and the standard discrimination result, obtain a discrimination loss value. Use the identity loss value, the first additional image loss value, the enhanced second additional image loss value, the enhanced third additional image loss value, and the generation loss value to adjust the model parameters of the decoder and the encoder to obtain a trained decoder and a trained encoder, and use the discrimination loss value to adjust the model parameters of the discriminator to obtain a trained discriminator. Based on the trained decoder and the trained encoder, obtain a trained face replacement model.

[0158] For example, taking the additional object as glasses for illustration, the first additional image loss value can be, for example, Figure 5 the glasses loss value in Figure 5As shown in the figure, the training source face image is input into the glasses feature extraction network G to obtain the source glasses feature. The training source face image is input into the face feature extraction network F to obtain the source identity feature. The training source face image and the training template face image are input into the encoder E to obtain the encoded feature. The source glasses feature, the source identity feature, and the encoded feature are input into the decoder D1 to obtain the decoded face image. The decoded face image is input into the glasses feature extraction network G to obtain the target glasses feature. The decoded face image is input into the face feature extraction network F to obtain the target identity feature. The source identity feature and the target identity feature are input into the identity loss value calculation module. The identity loss value calculation module calculates the source identity feature and the target identity feature based on the identity feature loss function included in itself to obtain the identity loss value. The source glasses feature and the target glasses feature are input into the glasses loss value calculation module. The glasses loss value calculation module can calculate the target glasses feature and the source glasses feature based on the glasses feature loss function it includes to obtain the glasses loss value, that is, the first additional image loss value.

[0159] As Figure 6 shown in the figure, the decoded face image is segmented by the face segmentation network to obtain the matching image area, that is, the glasses area. The standard face image is segmented by the face segmentation network to obtain the additional image area. The second additional image loss value and the third additional image loss value are obtained based on the difference between the glasses area and the additional image area. As Figure 5 shown in the figure, the additional image area and the glasses area are input into the pixel difference calculation module. The pixel difference calculation module calculates the difference in pixel values between the additional image area and the glasses area to obtain the second additional image loss value. The second additional image loss value is enhanced by the first enhancement processing module to obtain the enhanced second additional image loss value. The additional image area is extracted to obtain the additional image feature, and the glasses area is feature-extracted to obtain the glasses feature. The glasses feature and the additional image feature are input into the feature difference calculation module. The feature difference calculation module calculates the difference between the glasses feature and the additional image feature to obtain the third image loss value. The third additional image loss value is enhanced by the second enhancement processing module to obtain the enhanced third additional image loss value.

[0160] As Figure 5 shown in the figure, the decoded face image is input into the discriminator D2 to obtain the decoded discrimination result. The decoded discrimination result is input into the generation loss value calculation module. The generation loss value calculation module calculates the decoded discrimination result based on the generation loss function included in itself to obtain the generation loss value. As Figure 7 shown in the figure, the decoded face image is input into the discriminator D2 to obtain the decoded discrimination result. The standard face image is input into the discriminator D2 to obtain the standard discrimination result. AsFigure 5 As shown, the decoded discrimination result and the standard discrimination result are input into the discrimination loss value calculation module. The discrimination loss value calculation module uses the discrimination loss function included in itself to calculate the decoded discrimination result and the standard discrimination result. The decoded discrimination result may include the decoded discrimination probability, and the standard discrimination result may include the standard discrimination probability. The discrimination loss value calculation module uses the discrimination loss function included in itself to calculate the decoded discrimination probability and the standard discrimination probability, and obtains the discrimination loss value.

[0161] 2. Obtain the film and television video that needs to be face-swapped, determine the video frame where the target person whose face needs to be replaced is located in the film and television video as the target template face image, and obtain the reference face image corresponding to the reference person corresponding to the target person as the target source face image. The face of the target person in the film and television video can be replaced with the face of the reference person by using the trained face replacement model to obtain the face-swapped film and television video.

[0162] Specifically, the target identity feature and the target source additional image feature of the target source face image can be obtained. The target template face image and the target source face image are input into the encoder of the trained face replacement model to obtain the target encoded feature. The triple composed of the target encoded feature, the target identity feature, and the target source additional image feature is input into the decoder of the trained face replacement model to obtain the face-swapped image corresponding to the target template face image. The face-swapped image is consistent with the target source face image in terms of identity and additional image, and the face-swapped image is consistent with the target template face image in terms of attributes, thereby realizing the replacement of the face of the target person in the film and television video.

[0163] The present application further provides an application scenario, which is for game character design and applies the above image processing method. Specifically, the application of the image processing method in this application scenario is as follows: The face image corresponding to the first game character can be obtained to get the first game face image, and the face image corresponding to the second game character can be obtained to get the second game face image. A game character refers to a character designed in game design. The first game character is different from the second game character. Using the trained face replacement model, the face in the first game face image can be replaced with the face of the second game character, or the face in the second game face image can be replaced with the face of the first game character. For example, the identity features of the first game face image can be extracted to obtain the first identity features, the additional image features of the first game face image can be extracted to obtain the first additional image features, the second game face image is input into the encoder of the trained face replacement model for encoding to obtain the game face attribute features, the game face attribute features, the first identity features, and the first additional image features are combined to form a game feature triple, and the game feature triple is input into the decoder of the trained face replacement model for decoding to obtain the face image after face replacement. The identity and additional image in the face image after face replacement are the same as those in the first game face image, and the attributes are the same as those in the second game face image, thus realizing the replacement of the face in the second game face image with the face of the first game character.

[0164] Applying the trained face replacement model to game character design can quickly generate characters with a specific style, improve the efficiency of game character design, and reduce the cost of game character design.

[0165] The image processing method provided by the present application can also be applied to scenarios of virtual avatars or live broadcasts. The application of the image processing method in this application scenario is as follows: The live face image of the current live broadcast person can be obtained, and the virtual face image of the virtual character can be obtained. Using the trained face replacement model, the face in the live face image is replaced with the face in the virtual character image. A virtual character refers to a non-real person, which can be drawn manually or by computer. Specifically, the identity features of the virtual face image can be extracted to obtain the virtual identity features, the additional image features of the virtual face image can be extracted to obtain the virtual image features, the live face image is input into the encoder of the trained face replacement model for encoding to obtain the live face attribute features, and the virtual identity features, the virtual image features, and the live face attribute features are input into the decoder of the trained face replacement model for decoding to obtain the face image after replacement. The face in the face image after replacement is the face of the virtual character.

[0166] In an avatar, by using a trained face replacement model to replace the face of a real person with the face of a virtual person, the privacy of the user can be protected.

[0167] It should be understood that although Figure 2-7 the steps in the flowchart of Figure 2-7 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0168] In some embodiments, as Figure 8 shown, an image processing apparatus is provided. This apparatus can be a software module, a hardware module, or a combination of both to form a part of a computer device. Specifically, the apparatus includes: a first facial image acquisition module 802, a training source additional image feature obtaining module 804, a source identity feature obtaining module 806, a facial attribute feature obtaining module 808, a decoded facial image obtaining module 810, a target model loss value obtaining module 812, and a trained face replacement model obtaining module 814, where:

[0169] The first facial image acquisition module 802 is configured to acquire a training source facial image and a training template facial image.

[0170] The training source additional image feature obtaining module 804 is configured to extract additional image features from the training source facial image to obtain source additional image features corresponding to the training source facial image.

[0171] The source identity feature obtaining module 806 is configured to extract identity features from the training source facial image to obtain source identity features corresponding to the training source facial image.

[0172] The facial attribute feature obtaining module 808 is configured to input the training template facial image into the encoder in the face replacement model for encoding to obtain facial attribute features.

[0173] The decoded facial image obtaining module 810 is configured to input the source additional image features, the source identity features, and the facial attribute features into the decoder in the face replacement model for decoding to obtain a decoded facial image.

[0174] The target model loss value obtaining module 812 is configured to obtain the additional image difference between the decoded facial image and the comparison facial image, and obtain the target model loss value based on the additional image difference; the target model loss value is positively correlated with the additional image difference; the comparison facial image includes at least one of the training source facial image or the standard facial image corresponding to the decoded facial image.

[0175] The trained facial replacement model obtaining module 814 is configured to adjust the model parameters of the encoder and the decoder based on the target model loss value to obtain a trained facial replacement model for performing image processing according to the facial replacement model.

[0176] In some embodiments, the additional image difference includes a first image feature difference, and the target model loss value obtaining module 812 includes:

[0177] The target additional image feature obtaining unit is configured to perform additional image feature extraction on the decoded facial image to obtain the target additional image feature corresponding to the decoded facial image.

[0178] The first image feature difference obtaining unit is configured to determine the image feature difference between the source additional image feature and the target additional image feature as the first image feature difference.

[0179] The first target model loss value obtaining unit is configured to obtain the target model loss value based on the first image feature difference.

[0180] In some embodiments, the target model loss value obtaining module 812 includes:

[0181] The additional image region obtaining unit is configured to identify the additional image of the comparison facial image to obtain the additional image region corresponding to the comparison facial image.

[0182] The additional image enhancement value obtaining unit is configured to obtain the additional image enhancement value corresponding to the additional image region.

[0183] The additional image difference obtaining unit is configured to determine the image difference between the additional image region and the image region at the corresponding position in the decoded facial image as the additional image difference.

[0184] The second target model loss value obtaining unit is configured to obtain the additional image loss value based on the additional image difference, and perform enhancement processing on the additional image loss value by using the additional image enhancement value to obtain the target model loss value.

[0185] In some embodiments, the additional image difference obtaining unit is further configured to obtain additional pixel points in the additional image region, and obtain decoded pixel points matching the positions of the additional pixel points from the decoded face image; calculate the pixel value difference between the additional pixel points and the decoded pixel points; perform statistics on the pixel value differences corresponding to the additional image region to obtain a difference statistic value, and use the difference statistic value as the additional image difference.

[0186] In some embodiments, the additional image difference obtaining unit is further configured to perform feature extraction on the additional image region to obtain the extracted additional image features; perform feature extraction on the image region corresponding to the additional image region in the decoded face image to obtain the decoded image features; calculate the image feature difference between the extracted additional image features and the decoded image features as the second image feature difference; and obtain the additional image difference based on the second image feature difference.

[0187] In some embodiments, the second target model loss value obtaining unit is further configured to obtain an additional image loss value based on the additional image difference, perform enhancement processing on the additional image loss value by using the additional image enhancement value to obtain an enhanced additional image loss value; obtain the non-additional image region corresponding to the comparison face image, and determine the image difference between the non-additional image region and the image region at the corresponding position in the decoded face image as the non-additional image difference; obtain the non-additional image loss value based on the non-additional image difference, where the non-additional image enhancement value corresponding to the non-additional image loss value is less than the additional image enhancement value; and obtain the target model loss value according to the enhanced additional image loss value and the non-additional image loss value.

[0188] In some embodiments, the target model loss value obtaining module 812 includes:

[0189] The additional image loss value obtaining unit is configured to obtain an additional image loss value based on the additional image difference.

[0190] The target identity feature obtaining unit is configured to perform identity feature extraction on the decoded face image to obtain the target identity feature corresponding to the decoded face image.

[0191] The identity loss value obtaining unit is configured to obtain an identity loss value based on the identity feature difference between the source identity feature and the target identity feature.

[0192] The third target model loss value obtaining unit is configured to obtain the target model loss value according to the additional image loss value and the identity loss value.

[0193] In some embodiments, the module 804 for training source additional image features is further configured to input the training source facial image into the trained additional image feature extraction network to extract additional image features, and obtain the source additional image features corresponding to the training source facial image; the module 814 for obtaining the trained facial replacement model is further configured to keep the network parameters of the additional image feature extraction network unchanged, adjust the model parameters of the encoder and the decoder based on the target model loss value, and obtain the trained facial replacement model for image processing according to the facial replacement model.

[0194] In some embodiments, the apparatus further includes a module for obtaining an additional image feature extraction network, and the module for obtaining an additional image feature extraction network includes:

[0195] An additional object image obtaining unit, configured to obtain an additional object image, where the additional object image includes an additional object corresponding to the additional image feature.

[0196] A training unit, configured to use the additional object image to train the to-be-trained additional image recognition model, and obtain the trained additional image recognition model.

[0197] A feature extraction layer extraction unit, configured to extract the feature extraction layer before the image recognition layer from the trained additional image recognition model as the additional image feature extraction network.

[0198] In some embodiments, as Figure 9 shown, an image processing apparatus is provided. The apparatus may be a software module, a hardware module, or a combination of both to form a part of a computer device. The apparatus specifically includes: a second facial image obtaining module 902, a target source additional image feature obtaining module 904, a target identity feature obtaining module 906, a target facial attribute feature obtaining module 908, and a facial replacement image obtaining module 910, where:

[0199] The second facial image obtaining module 902 is configured to obtain a target source facial image and a target template facial image.

[0200] The target source additional image feature obtaining module 904 is configured to extract additional image features from the target source facial image, and obtain the target source additional image features corresponding to the target source facial image.

[0201] The target identity feature obtaining module 906 is configured to extract identity features from the target source facial image, and obtain the target identity features corresponding to the target source facial image.

[0202] The target facial attribute feature obtaining module 908 is configured to input the target template facial image into the encoder of the trained facial replacement model for encoding, and obtain the target facial attribute features.

[0203] The face replacement image generation module 910 is configured to decode the target source additional image features, target identity features, and target face attribute features into a decoder in the face replacement model to obtain a face replacement image. The face in the face replacement image matches the face in the target source face image, and the attributes in the face replacement image match the attributes in the target template face image.

[0204] In some embodiments, the second face image acquisition module 902 includes: a target object image acquisition unit configured to acquire a target object image corresponding to a target object with a face to be replaced; a face comparison unit configured to determine a current video frame in the target video and compare the current object face in the current video frame with the target object face in the target object image; a target source face image generation unit configured to, when the current object face matches the target object face, segment a matching target template face image from the current video frame and use the reference face image corresponding to the reference object of the target object as the target source face image. The apparatus is further configured to replace the target template face image in the current video frame with the face replacement image to obtain an updated current video frame.

[0205] For the specific definition of the image processing apparatus, reference may be made to the definition of the image processing method above, which will not be elaborated here. Each module in the above image processing apparatus may be implemented in whole or in part by software, hardware, or a combination thereof. The above modules may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0206] In some embodiments, a computer device is provided. The computer device may be a server, and its internal structural diagram may be as Figure 10 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Wherein, the processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store relevant data involved in the image processing method. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements an image processing method.

[0207] Those skilled in the art can understand that Figure 10The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0208] In some embodiments, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0209] In some embodiments, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0210] In some embodiments, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0211] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0212] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0213] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining a training source facial image and a training template facial image; Performing additional image feature extraction on the training source facial image to obtain the source additional image features corresponding to the training source facial image; Performing identity feature extraction on the training source facial image to obtain the source identity features corresponding to the training source facial image; Inputting the training template facial image into the encoder in the facial replacement model for encoding to obtain facial attribute features; Inputting the source additional image features, the source identity features, and the facial attribute features into the decoder in the facial replacement model for decoding to obtain a decoded facial image; Obtaining the additional image difference between the decoded facial image and a comparison facial image, and obtaining a target model loss value based on the additional image difference, including: using the additional image enhancement value corresponding to the additional image area in the comparison facial image to enhance the additional image loss value to obtain the target model loss value, where the additional image loss value is the image difference between the additional image area and the image area at the corresponding position in the decoded facial image; the target model loss value is positively correlated with the additional image difference; the comparison facial image includes at least one of the training source facial image or the standard facial image corresponding to the decoded facial image; Adjusting the model parameters of the encoder and the decoder based on the target model loss value to obtain a trained facial replacement model for performing image processing according to the facial replacement model.

2. The method according to claim 1, wherein The additional image difference includes a first image feature difference, and obtaining the additional image difference between the decoded facial image and a comparison facial image, and obtaining a target model loss value based on the additional image difference includes: Performing additional image feature extraction on the decoded facial image to obtain the target additional image features corresponding to the decoded facial image; Determining the image feature difference between the source additional image features and the target additional image features as the first image feature difference; Obtaining a target model loss value based on the first image feature difference.

3. The method according to claim 1, wherein The using the additional image enhancement value corresponding to the additional image area in the comparison facial image to enhance the additional image loss value to obtain the target model loss value includes: Identifying the additional image of the comparison facial image to obtain the additional image area corresponding to the comparison facial image; Obtaining the additional image enhancement value corresponding to the additional image area; Determining the image difference between the additional image area and the image area at the corresponding position in the decoded facial image as the additional image difference; Obtaining an additional image loss value based on the additional image difference, and using the additional image enhancement value to enhance the additional image loss value to obtain the target model loss value.

4. The method according to claim 3, characterized in that, The determining the image difference between the additional image area and the image area at the corresponding position in the decoded facial image as the additional image difference includes: Obtaining the additional pixel points in the additional image area, and obtaining the decoded pixel points that match the positions of the additional pixel points from the decoded facial image; Calculate the pixel value difference between the additional pixel points and the decoded pixel points; Statistically analyze the pixel value differences corresponding to the additional image area to obtain a difference statistical value, and use the difference statistical value as the additional image difference.

5. The method according to claim 3, characterized in that, Determining the image difference between the additional image area and the image area at the corresponding position in the decoded facial image as the additional image difference includes: Extract features from the additional image area to obtain extracted additional image features; Extract features from the image area corresponding to the additional image area in the decoded facial image to obtain decoded image features; Calculate the image feature difference between the extracted additional image features and the decoded image features as the second image feature difference; Obtain the additional image difference based on the second image feature difference.

6. The method according to claim 3, characterized in that The obtaining the target model loss value by obtaining the additional image loss value based on the additional image difference and enhancing the additional image loss value using the additional image enhancement value includes: Obtain the additional image loss value based on the additional image difference, and enhance the additional image loss value using the additional image enhancement value to obtain the enhanced additional image loss value; Obtain the non-additional image area corresponding to the comparison facial image, and determine the image difference between the non-additional image area and the image area at the corresponding position in the decoded facial image as the non-additional image difference; Obtain the non-additional image loss value based on the non-additional image difference, and the non-additional image enhancement value corresponding to the non-additional image loss value is less than the additional image enhancement value; Obtain the target model loss value according to the enhanced additional image loss value and the non-additional image loss value.

7. The method according to claim 1, wherein The obtaining the target model loss value based on the additional image difference includes: Obtain the additional image loss value based on the additional image difference; Extract identity features from the decoded facial image to obtain the target identity features corresponding to the decoded facial image; Obtain the identity loss value based on the identity feature difference between the source identity feature and the target identity feature; Obtain the target model loss value according to the additional image loss value and the identity loss value.

8. The method according to claim 1, characterized in that The extracting the source additional image features corresponding to the training source facial image by extracting additional image features from the training source facial image includes: Input the training source facial image into the trained additional image feature extraction network for additional image feature extraction to obtain the source additional image features corresponding to the training source facial image; The adjusting the model parameters of the encoder and the decoder based on the target model loss value to obtain the trained facial replacement model for image processing according to the facial replacement model includes: Keep the network parameters of the additional image feature extraction network unchanged, adjust the model parameters of the encoder and the decoder based on the target model loss value to obtain the trained facial replacement model for image processing according to the facial replacement model.

9. The method according to claim 8, characterized in that, The steps of obtaining the additional image feature extraction network include: Obtain an additional object image, where the additional object image includes the additional object corresponding to the additional image features; Train an additional image recognition model to be trained using the additional object image to obtain a trained additional image recognition model; Extract the feature extraction layer before the image recognition layer from the trained additional image recognition model as the additional image feature extraction network.

10. An image processing method, characterized in that, The method includes: Obtain a target source face image and a target template face image; Extract additional image features from the target source face image to obtain target source additional image features corresponding to the target source face image; Extract identity features from the target source face image to obtain target identity features corresponding to the target source face image; Input the target template face image into the encoder in the trained face replacement model for encoding to obtain target face attribute features; Input the target source additional image features, the target identity features, and the target face attribute features into the decoder in the face replacement model for decoding to obtain a face replacement image, where the face in the face replacement image matches the face in the target source face image, and the attributes in the face replacement image match the attributes in the target template face image; Among them, the encoder and the decoder are obtained by adjusting model parameters based on a target model loss value, and the target model loss value is obtained by enhancing the additional image loss value using the additional image enhancement value corresponding to the additional image area in the comparison face image. The additional image loss value is the image difference between the additional image area and the image area at the corresponding position in the decoded face image.

11. The method according to claim 10, wherein The obtaining of the target source face image and the target template face image includes: Obtain a target object image corresponding to a target object with a face to be replaced; Determine the current video frame in the target video, and compare the current object face in the current video frame with the target object face in the target object image; When the current object face matches the target object face, segment the matching target template face image from the current video frame, and use the reference face image of the reference object of the target object as the target source face image; The method further includes: Use the face replacement image to replace the target template face image in the current video frame to obtain an updated current video frame.

12. An image processing apparatus, characterized in that, The device includes: A first face image acquisition module for acquiring a training source face image and a training template face image; A module for obtaining trained source additional image features, which is used to extract additional image features from the training source face image to obtain source additional image features corresponding to the training source face image; A module for obtaining source identity features, which is used to extract identity features from the training source face image to obtain source identity features corresponding to the training source face image; A module for obtaining face attribute features, which is used to input the training template face image into the encoder in the face replacement model for encoding to obtain face attribute features; A module for obtaining a decoded face image, which is used to input the source additional image features, source identity features, and the face attribute features into the decoder in the face replacement model for decoding to obtain a decoded face image; The target model loss value obtaining module is used to obtain the additional image difference between the decoded facial image and the comparison facial image, and obtain the target model loss value based on the additional image difference, including: using the additional image enhancement value corresponding to the additional image area in the comparison facial image to enhance the additional image loss value to obtain the target model loss value, where the additional image loss value is the image difference between the additional image area and the image area at the corresponding position in the decoded facial image; the target model loss value is positively correlated with the additional image difference; the comparison facial image includes at least one of the training source facial image or the standard facial image corresponding to the decoded facial image; The trained facial replacement model obtaining module is used to adjust the model parameters of the encoder and the decoder based on the target model loss value to obtain a trained facial replacement model for image processing according to the facial replacement model.

13. The image processing apparatus according to claim 12, wherein The additional image difference includes a first image feature difference, and the target model loss value obtaining module includes: The target additional image feature obtaining unit is used to extract additional image features from the decoded facial image to obtain the target additional image features corresponding to the decoded facial image; The first image feature difference obtaining unit is used to determine the image feature difference between the source additional image features and the target additional image features as the first image feature difference; The first target model loss value obtaining unit is used to obtain the target model loss value based on the first image feature difference.

14. The image processing apparatus according to claim 12, wherein The target model loss value obtaining module includes: The additional image area obtaining unit is used to identify the additional image of the comparison facial image to obtain the additional image area corresponding to the comparison facial image; The additional image enhancement value obtaining unit is used to obtain the additional image enhancement value corresponding to the additional image area; The additional image difference obtaining unit is used to determine the image difference between the additional image area and the image area at the corresponding position in the decoded facial image as the additional image difference; The second target model loss value obtaining unit is used to obtain the additional image loss value based on the additional image difference, and use the additional image enhancement value to enhance the additional image loss value to obtain the target model loss value.

15. The image processing apparatus according to claim 14, wherein The additional image difference obtaining unit is further used to obtain the additional pixel points in the additional image area, and obtain the decoded pixel points matching the positions of the additional pixel points from the decoded facial image; calculate the pixel value difference between the additional pixel points and the decoded pixel points; perform statistics on the pixel value difference corresponding to the additional image area to obtain a difference statistic value, and use the difference statistic value as the additional image difference.

16. The image processing apparatus according to claim 14, wherein The additional image difference obtaining unit is further used to extract features from the additional image area to obtain the extracted additional image features; extract features from the image area corresponding to the additional image area in the decoded facial image to obtain the decoded image features; Calculate the image feature difference between the extracted additional image features and the decoded image features as the second image feature difference; obtain the additional image difference based on the second image feature difference.

17. The image processing apparatus according to claim 14, wherein The second target model loss value obtaining unit is further configured to obtain an additional image loss value based on the additional image difference, and perform enhancement processing on the additional image loss value by using the additional image enhancement value to obtain an enhanced additional image loss value. Obtain the non-additional image region corresponding to the comparison facial image, determine the image difference between the non-additional image region and the image region at the corresponding position in the decoded facial image as the non-additional image difference; obtain the non-additional image loss value based on the non-additional image difference, and the non-additional image enhancement value corresponding to the non-additional image loss value is less than the additional image enhancement value. Obtain the target model loss value according to the enhanced additional image loss value and the non-additional image loss value.

18. The image processing apparatus according to claim 12, wherein The obtaining the target model loss value based on the additional image difference, the target model loss value obtaining module includes: An additional image loss value obtaining unit, configured to obtain an additional image loss value based on the additional image difference. A target identity feature obtaining unit, configured to extract identity features from the decoded facial image to obtain the target identity features corresponding to the decoded facial image. An identity loss value obtaining unit, configured to obtain an identity loss value based on the identity feature difference between the source identity feature and the target identity feature. A third target model loss value obtaining unit, configured to obtain the target model loss value according to the additional image loss value and the identity loss value.

19. The image processing apparatus according to claim 12, wherein The training source additional image feature obtaining module is further configured to input the training source facial image into the trained additional image feature extraction network to extract additional image features, and obtain the source additional image features corresponding to the training source facial image; the trained facial replacement model obtaining module is further configured to keep the network parameters of the additional image feature extraction network unchanged, and adjust the model parameters of the encoder and the decoder based on the target model loss value to obtain the trained facial replacement model for image processing according to the facial replacement model.

20. The image processing apparatus according to claim 19, wherein The apparatus further includes an additional image feature extraction network obtaining module, and the additional image feature extraction network obtaining module includes: An additional object image obtaining unit, configured to obtain an additional object image, where the additional object image includes an additional object corresponding to the additional image features. A training unit, configured to train the to-be-trained additional image recognition model by using the additional object image to obtain the trained additional image recognition model. A feature extraction layer extraction unit, configured to extract the feature extraction layer before the image recognition layer from the trained additional image recognition model as the additional image feature extraction network.

21. An image processing apparatus, characterized in that, The apparatus includes: A second facial image obtaining module, configured to obtain a target source facial image and a target template facial image. A target source additional image feature obtaining module, configured to extract additional image features from the target source facial image to obtain the target source additional image features corresponding to the target source facial image. A target identity feature obtaining module, configured to extract identity features from the target source face image to obtain target identity features corresponding to the target source face image; A target face attribute feature obtaining module, configured to input the target template face image into an encoder in a trained face replacement model for encoding to obtain target face attribute features; A face replacement image obtaining module, configured to input the target source additional image features, the target identity features, and the target face attribute features into a decoder in the face replacement model for decoding to obtain a face replacement image, where the face in the face replacement image matches the face in the target source face image, and the attributes in the face replacement image match the attributes in the target template face image; Wherein, the encoder and the decoder are obtained by adjusting model parameters based on a target model loss value, and the target model loss value is obtained by enhancing an additional image loss value by using an additional image enhancement value corresponding to an additional image area in a comparison face image, and the additional image loss value is an image difference between the additional image area and an image area at a corresponding position in a decoded face image.

22. The image processing apparatus according to claim 21, wherein The second face image obtaining module includes: a target object image obtaining unit, configured to obtain a target object image corresponding to a target object with a face to be replaced; a face comparison unit, configured to determine a current video frame in a target video, and compare a current object face in the current video frame with a target object face in the target object image; a target source face image obtaining unit, configured to, when the current object face matches the target object face, segment a matching target template face image from the current video frame, and use a reference face image corresponding to a reference object of the target object as the target source face image; the apparatus is further configured to use the face replacement image to replace the target template face image in the current video frame to obtain an updated current video frame.

23. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

24. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

25. A computer program product comprising computer instructions, characterized in that, When the computer instruction is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Image processing method and device, model training method and device, computer equipment and storage medium

    CN111401216A