A facial image processing method, apparatus, and computer-readable storage medium

By extracting and fusing three-dimensional modeling parameters in facial image processing, the accuracy problem of facial image processing in the prior art is solved, and a more realistic facial replacement effect is achieved.

CN114943799BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110648608.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-10
Publication Date
2025-07-11
Estimated Expiration
2041-06-10

AI Technical Summary

Technical Problem

The existing facial image processing methods cannot process complex scenes and texture details through three-dimensional modeling. The method of generating an adversarial network results in large shape errors after facial replacement, reducing the accuracy of facial image processing.

Method used

By obtaining images of the source face and template face, performing feature extraction, facial modeling is performed, fusion of three-dimensional modeling parameters, building a three-dimensional facial image, and replacing the objects in the face template based on image texture features, three-dimensional facial features and attribute features to improve the accuracy of facial image processing.

Benefits of technology

By identifying and fusing three-dimensional modeling parameters, three-dimensional facial features are constructed, which improves the accuracy of facial image processing and makes the result of replacement facials more realistic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114943799B_ABST
    Figure CN114943799B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a facial image processing method, apparatus, and computer-readable storage medium. After obtaining a facial image of a source face and a facial template image of a template face, the embodiment of the present invention extracts features from the facial image and the facial template image to obtain the image texture features of the source object and the attribute features of the object in the facial template image. Then, facial modeling is performed on the source object and the object in the facial template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object, and the first three-dimensional modeling parameters and the second three-dimensional modeling parameters are fused to obtain target three-dimensional modeling parameters. A three-dimensional facial image is constructed according to the target three-dimensional modeling parameters to obtain the three-dimensional facial features of the three-dimensional facial image. Based on the image texture features, three-dimensional facial features, and attribute features, the object in the facial template image is replaced with the source object to obtain a target facial image. This solution can improve the accuracy of facial image processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and in particular, to a method and apparatus for facial image processing and a computer-readable storage medium. Background Art

[0002] In recent years, with the development of technologies, in applications such as movie special effects and Internet social networking, there is a need to replace the face of an object in a facial image with the face of another object while maintaining the style of the object in the facial image. In response to this need, facial images need to be processed. Existing facial image processing methods mainly achieve the replacement of the face of an object through three-dimensional modeling or generative adversarial networks.

[0003] In the process of researching and practicing the existing technologies, the inventors of the present invention found that the method based on three-dimensional modeling cannot handle complex scenes and texture details, and the shape error of the replaced face obtained by using the generative adversarial network method is relatively large. Therefore, the accuracy of facial image processing is greatly reduced. Summary of the Invention

[0004] Embodiments of the present invention provide a method and apparatus for facial image processing and a computer-readable storage medium, which can improve the accuracy of facial image processing.

[0005] A method for facial image processing includes:

[0006] Obtaining a facial image of a source face and a facial template image of a template face, where the facial image includes a source object;

[0007] Performing feature extraction on the facial image and the facial template image to obtain the image texture feature of the source object and the attribute feature of the object in the facial template image;

[0008] According to the facial image and the facial template image, performing facial modeling on the source object and the object in the facial template image to obtain first three-dimensional modeling parameters of the source object and second three-dimensional modeling parameters of the object in the facial template image, and fusing the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain target three-dimensional modeling parameters;

[0009] Constructing a three-dimensional facial image according to the target three-dimensional modeling parameters to obtain three-dimensional facial features of the three-dimensional facial image;

[0010] Based on the image texture feature, the three-dimensional facial feature, and the attribute feature, replacing the object in the facial template image with the source object to obtain a target facial image.

[0011] Correspondingly, an embodiment of the present invention provides a facial image processing apparatus, including:

[0012] An acquisition unit for acquiring a facial image of a source face and a facial template image of a template face, the facial image including a source object;

[0013] An extraction unit for extracting features from the facial image and the facial template image to obtain the image texture features of the source object and the attribute features of the object in the facial template image;

[0014] A fusion unit for performing facial modeling on the source object and the object in the facial template image according to the facial image and the facial template image to obtain a first three-dimensional modeling parameter of the source object and a second three-dimensional modeling parameter of the object in the facial template image, and fusing the first three-dimensional modeling parameter and the second three-dimensional modeling parameter to obtain a target three-dimensional modeling parameter;

[0015] A construction unit for constructing a three-dimensional facial image according to the target three-dimensional modeling parameter to obtain the three-dimensional facial features of the three-dimensional facial image;

[0016] A replacement unit for replacing the object in the facial template image with the source object based on the image texture features, the three-dimensional facial features, and the attribute features to obtain a target facial image.

[0017] Optionally, in some embodiments, the fusion unit may specifically be configured to extract the facial shape parameters corresponding to the facial image from the first three-dimensional modeling parameter; extract the facial action parameters corresponding to the facial template image from the second three-dimensional modeling parameter; and fuse the facial shape parameters and the facial action parameters to obtain a target three-dimensional modeling parameter.

[0018] Optionally, in some embodiments, the replacement unit may specifically be configured to splice the image texture features and the three-dimensional facial features to obtain facial features; use a trained post-processing image model to fuse the facial features and the attribute features to obtain fused facial features; and generate a target facial image based on the fused facial features, where the target facial image is an image obtained by replacing the object in the facial template image with the source object.

[0019] Optionally, in some embodiments, the replacement unit may specifically be configured to construct a facial mask corresponding to the fused facial features to obtain an initial facial mask; fuse the initial facial mask, the fused facial features, and the attribute features to obtain target facial features; adjust the target facial features, and construct a facial mask corresponding to the adjusted facial features to obtain a target facial mask; and generate a target facial image based on the target facial mask and the adjusted facial features.

[0020] Optionally, in some embodiments, the replacement unit may specifically be configured to perform feature transformation on the attribute features to obtain the target attribute features of the object in the facial template image; determine the weighting parameters of the fused facial features and the target attribute features according to the initial facial mask; perform weighting on the fused facial features and the target attribute features according to the weighting parameters, and fuse the weighted facial features and the weighted attribute features to obtain the target facial features.

[0021] Optionally, in some embodiments, the replacement unit may specifically be configured to generate an initial facial image according to the adjusted facial features, and screen out the image within the target facial mask in the initial facial image to obtain a basic facial image; identify the image outside the target facial mask in the facial template image to obtain a background image; and fuse the basic facial image and the background image to obtain a target facial image.

[0022] Optionally, in some embodiments, the facial image processing device may further include a training unit, and the training unit may specifically be configured to obtain a set of facial image samples, and screen out at least one pair of image samples from the set of facial image samples, where the pair of image samples includes a facial image sample and a facial template image sample; use a preset image processing model to replace the object in the facial template image sample with the object in the facial image sample to obtain a predicted facial image; and converge the preset image processing model based on the pair of image samples and the predicted facial image to obtain a trained image processing model.

[0023] Optionally, in some embodiments, the training unit may specifically be configured to perform feature extraction on the facial image sample and the facial template image by using a preset image processing model to obtain the image sample texture features of the facial image sample and the sample attribute features of the object in the facial template image sample; respectively identify the sample three-dimensional modeling parameters in the facial image sample and the facial template image sample, and fuse the identified sample three-dimensional modeling parameters to obtain target sample three-dimensional modeling parameters; construct a sample three-dimensional facial image according to the target sample three-dimensional modeling parameters to obtain the sample three-dimensional facial features of the sample three-dimensional facial image, and fuse the sample three-dimensional facial features, the image sample texture features, and the sample attribute features to obtain the fused sample facial features; construct a facial mask corresponding to the fused sample facial features to obtain an initial sample facial mask, and generate a predicted facial image based on the initial sample facial mask, the fused sample facial features, and the sample attribute features.

[0024] Optionally, in some embodiments, the training unit may specifically be configured to fuse the initial sample face mask, the fused sample face features, and the sample attribute features to obtain target sample face features; adjust the target sample face features, and construct a face mask corresponding to the adjusted sample face features to obtain a target sample face mask; generate a predicted face image based on the target sample face mask and the adjusted sample face features.

[0025] Optionally, in some embodiments, the training unit may specifically be configured to generate an initial predicted face image based on the target sample face features and the initial sample face mask; determine the shape loss information of the image sample pair according to the sample three-dimensional face image, the initial predicted face image, and the predicted face image; calculate the face similarity between the face image samples in the image sample pair and the predicted face image and the initial predicted face image respectively to obtain the image loss information of the image sample pair; determine the segmentation loss information of the image sample pair based on the initial sample face mask and the target sample face mask; determine the face loss information of the image sample pair according to the image sample pair, the predicted face image, and the initial predicted face image; fuse the shape loss information, the image loss information, the segmentation loss information, and the face loss information, and converge a preset image processing model based on the fused loss information to obtain a trained image processing model.

[0026] Optionally, in some embodiments, the training unit may specifically be configured to obtain first projection information of the sample three-dimensional face image, and extract first position information of the face contour from the first projection information; construct a target three-dimensional face image corresponding to the initial predicted face image and the predicted face image, and obtain second projection information of the target three-dimensional face image; extract second position information of the face contour in the initial predicted face image and third position information of the face contour in the predicted face image from the second projection information, and calculate the distance between the face contours respectively according to the first position information, the second position information, and the third position information to obtain the shape loss information of the image sample pair.

[0027] Optionally, in some embodiments, the training unit may specifically be configured to obtain a template map mask of the face template image sample in the image sample pair; adjust the size of the template map mask to obtain an adjusted template map mask; calculate the size differences between the adjusted template map mask and the initial sample face mask and the target sample face mask respectively, and fuse the size differences to obtain the segmentation loss information of the image sample pair.

[0028] Optionally, in some embodiments, the training unit may be specifically configured to calculate the similarity between the image sample pair and the predicted facial image and the initial predicted facial image respectively, so as to obtain the similarity loss information of the image sample pair; determine the adversarial loss information and the cycle loss information of the image sample pair according to the facial template image sample, the predicted facial image and the initial predicted facial image in the image sample pair; and use the similarity loss information, the adversarial loss information and the cycle loss information as the facial loss information of the image sample pair.

[0029] Optionally, in some embodiments, when the objects in the facial image sample and the facial template image sample are the same object, the training unit may be specifically configured to calculate the spatial similarity between the facial template image sample and the predicted facial image and the initial predicted facial image respectively, so as to obtain the spatial similarity loss information of the image sample pair; extract the image features of the facial template image sample, the predicted facial image and the initial predicted facial image, and calculate the feature similarity between the image features, so as to obtain the feature similarity loss information of the image sample pair; and use the spatial similarity loss information and the feature similarity loss information as the similarity loss information of the image sample pair.

[0030] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory, where the memory stores an application program, and the processor is configured to run the application program in the memory to implement the facial image processing method provided by the embodiment of the present invention.

[0031] In addition, an embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the facial image processing methods provided by the embodiment of the present invention.

[0032] After obtaining the facial image of the source face and the facial template image of the template face, the embodiment of the present application extracts features from the facial image and the facial template image to obtain the image texture features of the source object and the attribute features of the object in the facial template image. Then, based on the facial image and the facial template image, facial modeling is performed on the source object and the object in the facial template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image, and the first three-dimensional modeling parameters and the second three-dimensional modeling parameters are fused to obtain the target three-dimensional modeling parameters. A three-dimensional facial image is constructed according to the target three-dimensional modeling parameters to obtain the three-dimensional facial features of the three-dimensional facial image; based on the image texture features, three-dimensional facial features, and attribute features, the object in the facial template image is replaced with the source object to obtain the target facial image; since this solution can identify the three-dimensional modeling parameters in the facial image and the facial template image, and construct a three-dimensional facial image based on the three-dimensional modeling parameters, thereby obtaining three-dimensional facial features, enabling the facial features to obtain geometric constraints, and moreover, fusing the three-dimensional facial features, image texture features, and attribute features can make the result of the replaced face more realistic. Therefore, the accuracy of facial image processing can be improved. Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0034] Figure 1 is a schematic diagram of the scenario of the image processing method provided by the embodiment of the present invention;

[0035] Figure 2 is a schematic flowchart of the image processing method provided by the embodiment of the present invention;

[0036] Figure 3 is a schematic diagram of the recombination of the target three-dimensional modeling parameters provided by the embodiment of the present invention;

[0037] Figure 4 is a schematic diagram of the splicing of facial features provided by the embodiment of the present invention;

[0038] Figure 5 is a schematic diagram of generating the target facial image provided by the embodiment of the present invention;

[0039] Figure 6 is a schematic diagram of determining the shape loss information provided by the embodiment of the present invention;

[0040] Figure 7 is a schematic diagram of the overall image processing model after training provided by the embodiment of the present invention;

[0041] Figure 8 It is a schematic diagram for comparing the effects of the target face-swapped image provided by the embodiment of the present invention with the face-swapped image of the prior art;

[0042] Figure 9 It is another schematic flowchart of the image processing method provided by the embodiment of the present invention;

[0043] Figure 10 It is a schematic structural diagram of the image processing method provided by the embodiment of the present invention;

[0044] Figure 11 It is another schematic structural diagram of the image processing method provided by the embodiment of the present invention;

[0045] Figure 12 It is a schematic structural diagram of the electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.

[0047] The embodiment of the present invention provides a facial image processing method, device, and computer-readable storage medium. Among them, the facial image processing device can be integrated in an electronic device, and the electronic device can be a server or a terminal device, etc.

[0048] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.

[0049] For example, refer to Figure 1, taking the case where the facial image processing device is integrated in an electronic device as an example, after the electronic device acquires the facial image of the source face and the facial template image of the template face, it extracts features from the facial image and the facial template image to obtain the image texture features of the source object and the attribute features of the object in the facial template image. Then, based on the facial image and the facial template image, it performs facial modeling on the source object and the object in the facial template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image, and fuses the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain the target three-dimensional modeling parameters. According to the target three-dimensional modeling parameters, it constructs a three-dimensional facial image to obtain the three-dimensional facial features of the three-dimensional facial image; based on the image texture features, the three-dimensional facial features and the attribute features, it replaces the object in the facial template image with the source object to obtain the target facial image, thereby improving the accuracy of facial image processing.

[0050] Among them, the process of facial image processing can be regarded as replacing the object in the facial template image with the source face. It can be understood as performing face swapping on the face object in the facial template image. Taking the face as a human face as an example, face swapping means replacing the face identity in the facial template image with the person in the source image while keeping elements such as the pose, expression, makeup and background of the face in the facial template image unchanged. It can usually be applied in scenarios such as film and television production, game entertainment and e-commerce sales.

[0051] It should be noted that the image processing method provided in the embodiments of this application involves computer vision technology in the field of artificial intelligence. That is, in the embodiments of this application, the computer vision technology of artificial intelligence can be used to replace the object in the facial template image with the source object of the facial image to obtain the target facial image.

[0052] The so-called artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, a theory, method, technology and application system that can perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to make the machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0053] Among them, computer vision technology (CV). Computer vision is a science that studies how to enable machines to "see". Further, it refers to using cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0054] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0055] This embodiment will be described from the perspective of a facial image processing device. The facial image processing device can be specifically integrated in an electronic device, which can be a server or a terminal device, etc.; among them, the terminal can include devices such as a tablet computer, a laptop computer, a personal computer (PC, Personal Computer), a wearable device, a virtual reality device, or other intelligent devices that can perform facial image processing.

[0056] A facial image processing method includes:

[0057] Obtain a facial image of a source face and a facial template image of a template face. The facial image includes a source object. Extract features from the facial image and the facial template image to obtain the image texture features of the source object and the attribute features of the object in the facial template image. Based on the facial image and the facial template image, perform facial modeling on the source object and the object in the facial template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image, and fuse the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain target three-dimensional modeling parameters; construct a three-dimensional facial image based on the target three-dimensional modeling parameters to obtain the three-dimensional facial features of the three-dimensional facial image; based on the image texture features, three-dimensional facial features, and attribute features, replace the object in the facial template image with the source object to obtain a target facial image.

[0058] As Figure 2 shown, the specific process of this image processing method is as follows:

[0059] 101. Obtain the facial image of the source face and the facial template image of the template face.

[0060] Among them, the facial image includes the source object. The so-called source object can be the object contained in the facial image. Taking the facial image as a human face image as an example, the source object can be the person corresponding to the human face image. The source face is the source of the facial object for providing facial object replacement. Corresponding to it is the template face, which contains the facial object to be replaced and other elements such as the facial background to be maintained.

[0061] Among them, there are various ways to obtain the facial image and the facial template image. For example, the facial image and the facial template image can be directly obtained, or when the number of facial images and facial template images is large or the memory is large, the facial image and the facial template image can also be obtained indirectly. Specifically, it can be as follows:

[0062] (1) Directly obtain the facial image and the facial template image

[0063] For example, the original facial image uploaded by the user and the image processing information corresponding to the original facial image can be directly received. According to the image processing information, the facial image of the source face and the facial template image of the template face are screened out from the original facial image, or a pair of facial images is obtained from the image database or the network, and any one of the pair of facial images is randomly selected as the facial image of the source face, and the other facial image in the pair of facial images is used as the facial template image.

[0064] (2) Indirectly obtain the facial image and the facial template image

[0065] For example, an image processing request sent by the terminal can be received. The image processing request carries the storage address of the original facial image and the image processing information. According to the storage address, the original facial image is obtained from the memory, cache or third-party database. According to the image processing information, the facial image of the source face and the facial template image of the template face are screened out from the original facial image.

[0066] Optionally, after successfully obtaining the original facial image, a prompt message can also be sent to the terminal to prompt the terminal that the original facial image has been successfully obtained.

[0067] Optionally, after successfully obtaining the original facial image, preprocessing can also be performed on the original facial image to obtain the facial image and the facial template image. There are various preprocessing methods. For example, the size of the original facial image can be adjusted to a preset size, or facial key point registration can also be used to align the facial objects in the original facial image to a unified position.

[0068] 102. Extract features from the facial image and the facial template image to obtain the image texture features of the source object and the attribute features of the object in the facial template image.

[0069] Among them, the image texture features can be identity features on the facial image texture.

[0070] Among them, there can be various ways of feature extraction, specifically as follows:

[0071] For example, the encoder network of the trained image processing model can be used to perform feature encoding on the facial template image to obtain the attribute features of the object in the facial template image, and the face recognition network of the trained image processing model can be used to extract features from the facial image to obtain the image texture features of the source object.

[0072] Among them, the structure of the encoder network can be various. For example, it can be a residual network stacked by multiple residual blocks (Res-Block). The specific stacking quantity can be set according to the actual application. For example, it can be 8 or any value, or it can also be an encoding block composed of multiple convolutional layers and activation layers.

[0073] 103. According to the facial image and the facial template image, perform facial modeling on the source object and the object in the facial template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image, and fuse the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain the target three-dimensional modeling parameters.

[0074] Among them, the three-dimensional modeling parameters can be the parameters for constructing a three-dimensional facial image (model). The three-dimensional modeling parameters include facial shape parameters and facial motion parameters. The facial motion parameters can include expression parameters and pose parameters. Based on these three-dimensional modeling parameters, a three-dimensional model of the face can be constructed through 3DMM (three-dimensional facial reconstruction model) or other three-dimensional facial reconstruction models.

[0075] Among them, there can be various ways of performing facial modeling on the source object and the object in the facial template image and fusing the three-dimensional modeling parameters, specifically as follows:

[0076] For example, a three-dimensional facial reconstruction model can be used to perform regression on the facial image and the facial template image, so as to perform facial modeling on the source object and the object in the facial template image, and obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image in the constructed facial model. Extract the facial shape parameters corresponding to the facial image from the first three-dimensional modeling parameters, extract the facial motion parameters corresponding to the facial template image from the second three-dimensional modeling parameters, and fuse the facial shape parameters and the facial motion parameters to obtain the target three-dimensional modeling parameters.

[0077] Among them, there are various ways to obtain the first 3D modeling parameter and the second 3D modeling parameter in the constructed facial model. For example, the first 3D modeling parameter can be directly extracted from the facial model of the source object, and the second 3D modeling parameter can be extracted from the facial model of the object in the facial template image. Or, the facial model can be converted into 3D modeling parameters to obtain the first 3D modeling parameter and the second 3D modeling parameter.

[0078] After obtaining the first 3D modeling parameter and the second 3D modeling parameter, the first 3D modeling parameter and the second 3D modeling parameter can be fused. The fusion process can actually be regarded as reorganizing the 3D modeling parameters. The facial shape parameter is extracted from the first 3D modeling parameter of the facial image, and the facial action parameter is extracted from the second 3D modeling parameter of the facial template image. The fusion of the facial shape parameter and the facial action parameter can be carried out in various ways. For example, it can be directly spliced and combined. Or, the weighted parameter and the basic modeling parameter of the facial shape parameter and the facial action parameter can be obtained, and according to the weighted parameter, the facial shape parameter and the facial action parameter are weighted, and the weighted facial shape parameter and facial action parameter are fused with the basic modeling parameter to obtain the target 3D modeling parameter, which can be specifically shown in formula (1):

[0079]

[0080] Among them, S is the target 3D modeling parameter, and α and β are the weighted parameters. is the basic modeling parameter.

[0081] Among them, the target 3D modeling parameter can be regarded as reorganizing the facial shape of the source object in the facial image and the expression and posture of the object in the facial template image, which can be specifically shown in Figure 3 As shown, the reorganized target 3D modeling parameter takes into account the facial shape of the source object in the facial image and the expression and posture of the object in the facial template image in the geometric features, so as to improve the similarity of the facial shape.

[0082] 104. Construct a 3D facial image according to the target 3D modeling parameter to obtain the 3D facial features of the 3D facial image.

[0083] Among them, the 3D facial image can be the facial image of the 3D model constructed based on the target 3D modeling parameter.

[0084] Among them, the 3D facial features can be the shape features of the facial image and the action (expression and posture) features of the facial template image included in the 3D facial image.

[0085] Among them, there are various ways to construct the 3D facial image, which can be specifically as follows:

[0086] For example, a three-dimensional facial model corresponding to target three-dimensional modeling parameters can be constructed through a three-dimensional facial reconstruction model to obtain a three-dimensional facial image, and the three-dimensional features of the three-dimensional facial model can be acquired and used as the three-dimensional facial features of the three-dimensional facial image. Alternatively, a three-dimensional facial model corresponding to target three-dimensional modeling parameters can be constructed through a three-dimensional facial reconstruction model, and the three-dimensional facial model can be adjusted. For instance, local fitting or optimization can be performed, etc., to obtain a three-dimensional facial image, and the three-dimensional features of the adjusted three-dimensional facial model can be acquired and used as the three-dimensional facial features of the three-dimensional facial image.

[0087] 105. Based on the image texture features, three-dimensional facial features, and attribute features, replace the object in the facial template image with the source object to obtain the target facial image.

[0088] For example, the image texture features and three-dimensional facial features can be spliced to obtain facial features. The trained image processing model is used to fuse the facial features and attribute features to obtain the fused facial features, and the target facial image is generated based on the fused facial features. The target facial image is an image in which the object in the facial template image is replaced with the source object. Specifically, it can be as follows:

[0089] S1. Splice the image texture features and three-dimensional facial features to obtain facial features.

[0090] For example, the image texture features and three-dimensional facial features can be directly spliced to obtain facial features. Alternatively, the weighted parameters of the image texture features and three-dimensional facial features can be acquired, and the image texture features and three-dimensional facial features are weighted according to the weighted parameters, and the weighted image texture features and three-dimensional facial features are fused to obtain facial features. Or, the feature depths of the image texture features and three-dimensional facial features can be acquired, the feature depths are adjusted to a unified target feature depth, and based on the target feature depth, the image texture features and three-dimensional facial features are spliced to obtain facial features.

[0091] Among them, the facial features obtained by splicing can be identity features sensitive to the facial shape (face shape), mainly because geometric constraints of the facial shape are added to the three-dimensional facial features, such as Figure 4 shown, the constraint of the facial shape in the three-dimensional facial features is the constraint of the facial shape in the facial image of the source face.

[0092] S2. Use the trained image processing model to fuse the facial features and attribute features to obtain the fused facial features.

[0093] For example, the decoder network of the trained image processing model can be used to decode the facial features and attribute features, and the decoded facial features and attribute features are fused to obtain the fused facial features.

[0094] Among them, the structure of the decoder network can be various. For example, it can be a residual network stacked with multiple Res-Blocks, or it can be a network stacked with multiple Res-Blocks including AdaIN (style transfer) layers, or it can also be a decoding block composed of multiple convolutional layers and activation layers.

[0095] Among them, the stacking number of Res-Blocks or Res-Blocks including AdaIN layers can be set according to the actual application. For example, it can be 5 or any number.

[0096] S3. Generate a target facial image based on the fused facial features.

[0097] For example, a facial mask corresponding to the fused facial features can be constructed to obtain an initial facial mask. The initial facial mask, the fused facial features, and the attribute features are fused to obtain target facial features. The target facial features are adjusted, and a facial mask corresponding to the adjusted facial features is constructed to obtain a target facial mask. The target facial image is generated based on the target facial mask and the adjusted facial features. Specifically, it can be as follows:

[0098] (1) Construct a facial mask corresponding to the fused facial features to obtain an initial facial mask.

[0099] For example, a semantic segmentation network can be used to extract features from the fused facial features, and based on the extracted semantic features, the segmentation region is identified in the facial image or facial template image. Based on the segmentation region, the facial image or facial template image is segmented, and the segmented facial image is occluded to obtain the initial facial mask. Or, a semantic segmentation network can be used to extract features from the fused facial features, and based on the extracted semantic features, the initial facial mask is directly generated.

[0100] (2) Fuse the initial facial mask, the fused facial features, and the attribute features to obtain target facial features.

[0101] For example, the attribute features can be feature-transformed to obtain the target attribute features of the object in the facial template image. According to the initial facial mask, the weighting parameters of the fused facial features and the target attribute features are determined. According to the weighting parameters, the fused facial features and the target attribute features are weighted, and the weighted facial features and the weighted attribute features are fused to obtain the target facial features.

[0102] Among them, there are various ways to perform feature transformation on the attribute features. For example, one or more Res-Blocks can be used to perform feature transformation on the attribute features to obtain the target attribute features of the object in the facial template image.

[0103] After performing feature transformation on the attribute features, the weighted parameters can be determined in various ways. For example, in a facial image or a facial template image, regions outside the initial facial mask are selected to obtain a target image region, and the weighted parameters are determined based on the initial facial mask and the target image region.

[0104] Among them, the fused facial features and the fused attribute features are fused, and the fusion process can be various. For example, the fused facial features and the fused attribute features can be directly added to obtain the target facial features, which can be specifically shown in formula (2):

[0105] z fuse = M low ⊙ z dec +(1 - M low ) ⊙ σ(z enc ) (2)

[0106] Among them, M low is the initial facial mask, z dec is the fused facial feature, z enc is the attribute feature, and σ is a Res - Block structure.

[0107] (3) Adjust the target facial features and construct a facial mask corresponding to the adjusted facial features to obtain the target facial mask.

[0108] For example, an upsampling structure can be used to adjust the target facial features to obtain the adjusted facial features, and a semantic segmentation network can be used to construct a facial mask corresponding to the adjusted facial features to obtain the target facial mask.

[0109] Among them, there are various ways to adjust the target facial features using an upsampling structure. For example, an upsampling structure composed of one or more Res - Blocks can be used to enlarge the size of the target facial features to a preset size to obtain the adjusted facial features, thereby improving the resolution of the target facial features.

[0110] After adjusting the target facial features, a facial mask corresponding to the adjusted facial features can be constructed in various ways. For example, a semantic segmentation network can be used to extract features from the adjusted facial features, and based on the extracted semantic features, the segmentation region can be identified in the facial image or the facial template image, and based on the segmentation region, the facial image or the facial template image can be segmented, and the segmented facial image can be blocked to obtain the target facial mask. Or, a semantic segmentation network can be used to extract features from the adjusted facial features, and based on the extracted semantic features, the target facial mask can be directly generated.

[0111] (4)Generate a target facial image based on the target facial mask and the adjusted facial features.

[0112] For example, an initial facial image can be generated based on the adjusted facial features, and the image within the target facial mask can be screened out from the initial facial image to obtain a basic facial image. The image outside the target facial mask can be recognized from the facial template image to obtain a background image, and the basic facial image and the background image can be fused to obtain the target facial image.

[0113] Among them, there are various processes for fusing the basic facial image and the background image. For example, the basic facial image and the background image can be spliced to obtain the target facial image, specifically as shown in formula (3):

[0114] I r = M r ⊙ I out +(1 - M r )⊙ I t (3)

[0115] Among them, I r is the target facial image, M r is the target facial mask, I out is the initial facial image, and I t is the facial template image.

[0116] Optionally, the sizes of the basic facial image and the background image can also be adjusted, and the adjusted basic facial image and the background image can be superimposed to obtain the target facial image.

[0117] Among them, the process of generating the target facial image based on the fused facial features can be as Figure 5 shown. Through the facial semantic fusion module in the trained image processing model, the fused facial features output by the decoder network, the shallow encoding features (attribute features), and the initial facial mask corresponding to the fused facial features are semantically fused, so as to obtain the target facial features. The size of the target facial features is adjusted by upsampling, the target facial mask is generated based on the adjusted facial features, and the target facial image can be generated based on the target facial mask and the adjusted facial features.

[0118] Among them, the trained image processing model can be set according to the actual application requirements. In addition, it should be noted that the trained image processing model can be preset by the maintenance personnel or can be trained by the image processing device itself. That is, before the step of "using the trained image processing model to fuse the facial features and the attribute features to obtain the fused facial features", the image processing method can also include:

[0119] Obtain a set of facial image samples, and screen out at least one pair of image samples from the set of facial image samples. The pair of image samples includes a facial image sample and a facial template image sample. Use a preset image processing model to replace the object in the facial template image sample with the object in the facial image sample to obtain a predicted facial image. Converge the preset image processing model based on the pair of image samples and the predicted facial image to obtain a trained image processing model. Specifically, it can be as follows:

[0120] (1) Obtain a set of facial image samples, and screen out at least one pair of image samples from the set of facial image samples.

[0121] Among them, the pair of image samples includes a facial image sample and a facial template image sample.

[0122] Among them, there are various ways to obtain the set of facial image samples. Specifically, it can be as follows:

[0123] For example, multiple original facial image samples can be obtained, preprocessed to obtain a set of facial image samples. Arbitrarily screen out two facial image samples from the set of facial image samples, and specify any one of the images in the facial image samples as the facial image sample, then the other image sample can be the facial template image sample, so as to obtain a pair of image samples.

[0124] Among them, there are various preprocessing methods. For example, facial key point registration can be used to align the faces in the original prominent samples to a unified position, and the size of the aligned original image samples can be adjusted to a preset size, so as to obtain a set of image samples. The preset size can be set according to actual applications. For example, it can be 256×256 or other sizes.

[0125] (2) Use a preset image processing model to replace the object in the facial template image sample with the object in the facial image sample to obtain a predicted facial image.

[0126] For example, a preset image processing model can be used to extract features from the facial image sample and the facial template image, obtain the image sample texture features of the facial image sample and the sample attribute features of the object in the facial template image sample, respectively identify the sample three-dimensional modeling parameters in the facial image sample and the facial template image sample, fuse the identified sample three-dimensional modeling parameters to obtain the target sample three-dimensional detection parameters, construct a sample three-dimensional facial image according to the target sample three-dimensional modeling parameters, obtain the template three-dimensional facial features of the sample three-dimensional facial image, and fuse the sample three-dimensional facial features, the image sample texture features and the sample attribute features to obtain the fused sample facial features, construct a facial mask corresponding to the fused sample facial features to obtain the initial sample facial mask, and generate a predicted facial image based on the initial sample facial mask, the fused sample facial features and the sample attribute features.

[0127] Among them, steps such as extracting features from the facial image sample and the facial template image, identifying the sample three-dimensional modeling parameters, fusing the template three-dimensional modeling parameters, fusing the sample three-dimensional facial features, the image sample texture features and the sample attribute features, and constructing a facial mask corresponding to the fused sample facial features can be referred to above and will not be elaborated here one by one.

[0128] Among them, there are various ways to generate the predicted facial image. For example, the initial sample facial mask, the fused sample facial features and the sample attribute features can be fused to obtain the target sample facial features, the target sample facial features are adjusted, and a facial mask corresponding to the adjusted sample facial features is constructed to obtain the target sample facial mask, and a predicted facial image is generated based on the target sample facial mask and the adjusted sample facial features. Specifically, it can be referred to above and will not be elaborated here one by one.

[0129] (3) Converge the preset image processing model based on the image sample pair and the predicted facial image pair to obtain the trained image processing model.

[0130] For example, an initial predicted facial image can be generated based on the target sample facial features and the initial sample facial mask, the shape loss information of the image sample pair is determined according to the sample three-dimensional facial image, the initial predicted facial image and the predicted facial image, the facial similarity between the facial image sample in the image sample pair and the predicted facial image and the initial predicted facial image is calculated respectively to obtain the image loss information of the image sample pair, the segmentation loss information of the image sample pair is determined based on the initial sample facial mask and the target sample facial mask, the facial loss information of the image sample pair is determined according to the image sample pair, the predicted facial image and the initial predicted facial image, the shape loss information, the image loss information, the segmentation loss information and the facial loss information are fused, and the preset image processing model is converged based on the fused loss information to obtain the trained image processing model. Specifically, it can be as follows:

[0131] C1. Generate an initial predicted facial image based on the facial features of the target sample and the initial sample facial mask.

[0132] For example, an initial facial sample image can be generated according to the facial features of the target sample, and the images within the initial sample facial mask are screened out from the candidate facial images to obtain a basic facial sample image. The images outside the initial sample facial mask are screened out from the facial template image samples to obtain a background sample image. The background sample image and the basic facial sample image are fused to obtain the initial predicted facial image. For details, please refer to the above text and will not be elaborated here one by one.

[0133] C2. Determine the shape loss information of the image sample pair according to the sample three-dimensional facial image, the initial predicted facial image, and the predicted facial image.

[0134] For example, the first projection information of the sample three-dimensional facial image can be obtained, and the first position information of the facial contour is extracted from the first projection information. The target three-dimensional facial images corresponding to the initial predicted facial image and the predicted facial image are constructed, and the second projection information of the target three-dimensional facial image is obtained. The second position information of the facial contour in the initial predicted facial image and the third position information of the facial contour in the predicted facial image are extracted from the second projection information. The distances between the facial contours are calculated respectively according to the first position information, the second position information, and the third position information to obtain the shape loss information of the image sample pair.

[0135] Among them, there are various ways to obtain the first projection information of the sample three-dimensional facial image. For example, the 2D projection of the sample three-dimensional facial image under a preset angle coefficient can be obtained by using a three-dimensional renderer to obtain the first projection information. There are various three-dimensional renderers, such as pytorch3d (a three-dimensional renderer) or other three-dimensional renderers.

[0136] After obtaining the first projection information, the first position information of the facial contour can be extracted from the first projection information. There are various extraction methods. For example, the positions of multiple contour points of the facial contour can be obtained in the 2D projection, so as to obtain the first position information of the facial contour. The number of contour points can be 18 or other numbers.

[0137] Among them, there are various ways to construct the target three-dimensional facial image. For example, a facial three-dimensional reconstruction model can be used to reconstruct the target three-dimensional facial images corresponding to the initial predicted facial image and the predicted facial image. The reconstruction method can be referred to the above text and will not be elaborated here one by one.

[0138] After constructing the target three-dimensional facial image, the second projection information of the target three-dimensional facial image can be obtained, and the second position information of the facial contour in the initial predicted facial image can be extracted from the second projection information, and the third position information of the facial contour in the predicted facial image can be extracted from the third projection information. For specific details, please refer to the above text and will not be elaborated here one by one.

[0139] After extracting the second position information and the third position information of the facial contour, the distance between the facial contours can be calculated, so as to obtain the shape loss information of the image sample pair. There are various calculation methods. For example, the distance between the facial contours can be calculated according to the first position information, the second position information and the third position information. For instance, calculate the position difference between the contour points of the same facial contour in the first position information and the second position information, and calculate the position difference between the contour points of the same facial contour in the first position information and the third position information. Then, calculate the average value of the position differences to obtain the shape loss information of the image sample pair. Specifically, it can be as shown in formula (4):

[0140]

[0141] where L shape is the first shape loss information, N is the number of contour points of the facial contour, is the position of the contour point of the facial contour in the first position information, is the position of the same contour point in the second position information, is the position of the same contour point in the third position information.

[0142] Among them, the shape loss information of the image sample pair can be regarded as the distance between the three-dimensional facial image and the facial contour corresponding to the three-dimensional facial model corresponding to the initial predicted face and the predicted facial image. Specifically, it can be as Figure 6 shown.

[0143] C3. Calculate the facial similarity between the facial image sample in the image sample pair and the predicted facial image and the initial predicted facial image respectively to obtain the image loss information of the image sample pair.

[0144] For example, a pre-trained facial recognition feature extractor can be used to extract the facial features in the facial image sample, the predicted facial image and the initial predicted facial image, and calculate the cosine similarity between the facial features of the facial image sample and the predicted facial image, so as to obtain the first facial similarity between the facial image sample and the predicted facial image. Specifically, it can be as shown in formula (5):

[0145] L id1 = 1 - cos(v id (I s ), v id(I r )) (5)

[0146] Among them, L id1 is the first facial similarity, v id is a pre-trained facial recognition feature extractor, I s is a facial image sample, and I r is the predicted facial image.

[0147] Calculate the second facial similarity between the facial image sample and the initial predicted facial image. For the specific calculation process, please refer to the above text and will not be elaborated here one by one. Then, fuse the first facial similarity and the second facial similarity to obtain the image loss information. There are various fusion methods. For example, the first facial similarity and the second facial similarity can be directly added to obtain the image loss information. Specifically, it can be as shown in formula (6):

[0148] L id = (1 - cos(v id (I s ), v id (I r ))) + (1 - cos(v id (I s ), v id (I low ))) (6)

[0149] Among them, L id is the image loss information, v id is a pre-trained facial recognition feature extractor, I s is a facial image sample, I r is the predicted facial image, and I low is the initial predicted image.

[0150] Optionally, the fusion method can also be to obtain the weighting parameters of the first facial similarity and the second facial similarity, weight the first facial similarity and the second facial similarity according to the weighting parameters, and fuse the weighted first facial similarity and the second facial similarity to obtain the image loss information.

[0151] C4. Determine the segmentation loss information of the image sample pair based on the initial sample facial mask and the target sample facial mask.

[0152] For example, obtain the template map mask of the facial template image sample in the image sample pair, adjust the size of the template map mask to obtain the adjusted template map mask, calculate the size differences between the adjusted template map mask and the initial sample facial mask and the target sample facial mask respectively, and fuse the size differences to obtain the segmentation loss information of the image sample pair.

[0153] Among them, there are various ways to obtain the template map mask of the facial template image sample in the image sample pair. For example, a trained semantic segmentation network can be used to predict the template map mask of the facial template image sample in the image sample pair. Alternatively, mask features can be extracted from the facial template image, and based on the mask features, the mask corresponding to the facial template image sample can be screened out from a preset mask set to obtain the template map mask.

[0154] After obtaining the template map mask, the size of the template map mask can be adjusted in various ways. For example, in order to meet the changing requirements of the facial shape, the template map mask can be expanded outward by a preset number of pixels. For example, it can be 15 pixels or other numbers of pixels.

[0155] After obtaining the adjusted template map mask, the size differences between the adjusted template map mask and the initial sample facial mask and the target sample facial mask can be calculated respectively, and the size differences can be fused to obtain the segmentation loss information of the image sample pair, which can be specifically shown as in formula (7):

[0156] L seg =‖R(M tar )-M low ‖1+‖M tar -M r ‖1 (7)

[0157] Among them, L seg is the segmentation loss information, R(M tar ) is the adjusted template map mask, M tar is the template map mask, M low is the initial sample facial mask, and M r is the target sample facial mask.

[0158] C5. Determine the facial loss information of the image sample pair according to the image sample pair, the predicted facial image, and the initial predicted facial image.

[0159] For example, the similarities between the image sample pair and the predicted facial image and the initial predicted facial image can be calculated respectively to obtain the similarity loss information of the image sample pair. According to the facial template image sample, the predicted facial image, and the initial predicted facial image in the image sample pair, the adversarial loss information and the cycle loss information of the image sample pair can be determined, and the similarity loss information, the adversarial loss information, and the cycle loss information are used as the facial loss information of the image sample pair.

[0160] Among them, there are various ways to calculate the similarity loss information of image sample pairs. For example, when the objects in the facial image sample and the facial template image sample are the same object, the spatial similarities between the facial template image sample and the predicted facial image and the initial predicted facial image are calculated respectively to obtain the spatial loss information of the image sample pair. The image features of the facial template image sample, the predicted facial image, and the initial predicted facial image are extracted, and the feature similarities between the image features are calculated to obtain the feature similarity loss information of the image sample pair. The spatial similarity loss information and the feature similarity loss information are used as the similarity loss information of the image sample pair.

[0161] Among them, the spatial similarity can be understood as the similarity constraint between the facial image sample and the predicted facial image and the initial predicted facial image in the image RGB (color channel) space. There are various ways to calculate this spatial similarity. For example, the L1 norm can be used to calculate the first spatial similarity between the facial image sample and the predicted facial image, and the second spatial similarity between the facial image sample and the initial predicted facial image. The first spatial similarity and the second spatial similarity are fused to obtain the spatial similarity loss information of the image sample pair, which can be specifically shown in formula (8) as follows:

[0162] L rec =||I r -I t ||1+||I low -R(I t )||1 (8)

[0163] Among them, L rec is the spatial similarity loss information, I r is the predicted facial image, I t is the facial image sample, I low is the initial predicted facial image, R(I t ) is the scaled facial image, whose size (resolution) is the same as that of the initial predicted facial image.

[0164] Among them, the feature similarity can also be called the perceptual similarity, which is mainly used to represent the similarity between the features input at each feature layer of the same object in different images. There are various ways to calculate this feature similarity. For example, a feature extraction network can be used to extract the image features of the facial template image sample, the predicted facial image, and the initial predicted facial image, obtain the image features output by each feature layer of the feature extraction network, calculate the feature difference between the image features output by the facial template image sample and the predicted facial image at the same feature layer, and then fuse the feature difference and the feature size of the image features to obtain the first feature similarity, which can be specifically shown in formula (9) as follows:

[0165]

[0166] Among them, I perceptual is the first feature similarity, C i H i W i is the size of the image features of the i-th layer, F i (I r ) are the image features of the predicted facial image at the i-th layer, F i (I t ) are the image features of the facial template image sample at the i-th layer.

[0167] Calculate the feature difference between the image features output by the facial template image sample and the initial predicted facial image at the same feature layer, so as to obtain the second feature similarity. The specific process can be referred to above and will not be elaborated here one by one. Fuse the first feature similarity and the second feature similarity to obtain the feature similarity loss information of the image sample pair. There are various ways of fusion. For example, the first feature similarity and the second feature similarity can be directly added to obtain the feature similarity loss information. Or, the weighting coefficients of the first feature similarity and the second feature similarity can also be obtained, and the first feature similarity and the second feature similarity are weighted according to the weighting coefficients, and the weighted first feature similarity and the second feature similarity are fused to obtain the feature similarity loss information.

[0168] Among them, there are various ways to calculate the adversarial loss information. For example, an adversarial network can be used to calculate the first adversarial parameters of the facial template image and the predicted facial image respectively. Based on the first adversarial parameters, the first adversarial loss information is determined. Specifically, it can be as shown in formula (10):

[0169]

[0170] Among them, L adv is the first adversarial loss information, is the expected value of the facial template image, I t is the facial template image sample, is the expected value of the generated predicted facial image, G(I t , I s ) is the predicted facial image, D(I t ) and D(G(I t , I s ) are the first adversarial parameters.

[0171] The adversarial network is used to calculate the second adversarial parameters of the facial template image and the initial predicted facial image respectively. Based on the second adversarial parameters, the second adversarial loss information is determined. The determination process is as described above. The first adversarial loss information and the second adversarial loss information are fused to obtain the adversarial loss information. There are various fusion processes. For example, the first adversarial loss information and the second adversarial loss information can be directly added to obtain the adversarial loss information. Or, the weighted parameters of the first adversarial loss information and the second adversarial loss information can be obtained, and based on the weighted parameters, the first adversarial loss information and the second adversarial loss information are weighted, and the weighted first adversarial loss information and the second adversarial loss information are fused to obtain the adversarial loss information.

[0172] Among them, there are also various calculation methods for the cycle loss information. For example, the L1 norm can be used to calculate the norm of the facial template image sample and the predicted facial image obtained after image processing, so as to obtain the cycle loss information of the image sample pair. Specifically, it can be as shown in formula (11):

[0173] L cyc = ||I t - G(I r , I t )||1 (11)

[0174] Among them, L cyc is the cycle loss information, I t is the facial template image sample, I r is the predicted facial image, and G is the image processing function.

[0175] C6. The shape loss information, the image loss information, the segmentation loss information, and the facial loss information are fused, and based on the fused loss information, the preset image processing model is converged to obtain the trained image processing model.

[0176] For example, the shape loss information and the image loss information are fused to obtain the first loss information, the segmentation loss information and the facial loss information are fused to obtain the second loss information, and the first loss information and the second loss information are fused to obtain the fused loss information. Based on the fused loss information, the preset image processing model is converged to obtain the trained image processing model.

[0177] Among them, there are various ways to fuse the shape loss information and the image loss information. For example, the first weighted parameters of the shape loss information and the image loss information can be obtained, and based on the first weighted parameters, the shape loss information and the image loss information are weighted, and the weighted shape loss information and the image loss information are fused to obtain the first loss information. Specifically, it can be as shown in formula (12):

[0178] Lsid = λ shape L shape + λ id L id (12)

[0179] where L sid is the first loss information, λ shape and λ id are the first weighting parameters, L shape is the shape loss information, and L id is the image loss information. The first weighting parameters can be set according to actual applications. For example, the first weighting parameter for the shape loss information can be 5 or any other value, and the first weighting parameter for the image loss information can be 0.5 or other values.

[0180] Among them, there are also multiple ways to fuse the segmentation loss information and the facial loss information. For example, the second weighting parameters of the segmentation loss information and the facial loss information can be obtained, and the segmentation loss information and the facial loss information are weighted according to the second weighting parameters, and the weighted segmentation loss information and facial loss information are fused to obtain the second loss information. Specifically, it can be as shown in formula (13):

[0181] L real = L adv + λ0L seg + λ1L rec + λ2L cyc + λ3L lpips (13)

[0182] where L real is the second loss information, L adv is the adversarial loss information, L seg is the segmentation loss information, L rec is the spatial similarity loss information, L cyc is the periodic loss information, L lpips is the feature similarity loss information, and λ0, λ1, λ2, and λ3 are the second weighting parameters respectively. The second weighting parameters can be set according to actual applications. For example, λ0 can be 100, λ1 can be 20, λ2 can be 1, and λ3 can be 5, or other arbitrary values.

[0183] After obtaining the first loss information and the second loss information, the first loss information and the second loss information can be fused to obtain the fused loss information. There are also multiple ways of fusion. For example, the first loss information and the second loss information can be directly added to obtain the fused loss information. Specifically, it can be as shown in formula (14):

[0184] L = L sid + Lreal (14)

[0185] Among them, L is the fused loss information, L sid is the first loss information, L real is the second loss information.

[0186] After obtaining the fused loss information, the preset image processing model can be converged based on the fused loss information to obtain a trained image processing model. There are various ways of convergence. For example, the gradient descent algorithm can be used to update the network parameters of the preset image processing model according to the fused loss information to converge the preset image processing model and obtain a trained image processing model. Or, other convergence algorithms can also be used to update the network parameters of the preset image processing model according to the fused loss information to converge the preset image processing model and obtain a trained image processing model.

[0187] It should be noted that the overall trained image processing model adopts a scheme based on a generative adversarial network, and two new modules are proposed for the generator on this basis, using three-dimensional methods and semantic segmentation to enhance face shape changes and realism. As Figure 7 shown, the generator of this model is mainly composed of four parts, namely an encoder, a decoder, a face shape-sensitive facial feature extractor, and a facial semantic fusion module. The encoder is used to extract the attribute features of the facial template image, and the decoder is used to fuse the facial features and attribute features. The face shape-sensitive facial feature extractor is used to generate facial features, and the facial semantic fusion module is used to generate the target facial image, completing the face swapping process of replacing the object in the facial template image with the source object.

[0188] Taking the facial image as a human face as an example, through the processing of the facial image and the facial template image by this scheme, the target face-swapped image is obtained. The image effects of the target face-swapped image and the face-swapped images obtained by the existing technologies (FaceSwap, FaceShifter, and SimSwap) can be directly compared, and the comparison results can be as Figure 8 shown. A facial identification quantitative experiment is carried out on the target face-swapped image and the face-swapped images of the existing technologies. The experimental data is shown in Table (1). Through the experiment, it can be found that the target face-swapped image obtained by this scheme can well retain the shape and target attributes of the source face and has higher image quality.

[0189] Table (1)

[0190] Method ID Attitude Shape Prior Art 1 54.19 2.51 0.610 Prior Art 2 97.38 2.96 0.511 Prior Art 3 92.83 1.53 0.540 This Solution 98.48 2.63 0.540

[0191] As described above, after obtaining the facial image of the source face and the facial template image of the template face in the embodiments of the present application, feature extraction is performed on the facial image and the facial template image to obtain the image texture features of the source object and the attribute features of the object in the facial template image. Then, based on the facial image and the facial template image, facial modeling is performed on the source object and the object in the facial template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image, and the first three-dimensional modeling parameters and the second three-dimensional modeling parameters are fused to obtain the target three-dimensional modeling parameters. According to the target three-dimensional modeling parameters, a three-dimensional facial image is constructed to obtain the three-dimensional facial features of the three-dimensional facial image. Then, based on the image texture features, the three-dimensional facial features, and the attribute features, the object in the facial template image is replaced with the source object to obtain the target facial image. Since this solution can identify the three-dimensional modeling parameters in the facial image and the facial template image and construct a three-dimensional facial image based on the three-dimensional modeling parameters, thereby obtaining the three-dimensional facial features, enabling the facial features to obtain geometric constraints. Moreover, by fusing the three-dimensional facial features, the image texture features, and the attribute features, the result of the replaced face can be made more realistic. Therefore, the accuracy of facial image processing can be improved.

[0192] According to the method described in the above embodiments, the following will be further described in detail with examples.

[0193] In this embodiment, it will be described by taking the image processing device being specifically integrated in an electronic device, the electronic device being a server, and the facial image being a human face image as an example.

[0194] (1) The server trains a preset image processing model to obtain a trained image processing model

[0195] (1) The server obtains a set of human face image samples and filters out at least one pair of image samples from the set of human face image samples.

[0196] For example, multiple original human face image samples can be obtained, the faces in the original prominent samples are aligned to a unified position by using facial key point registration, and the size of the aligned original human face image samples is adjusted to 256×256, thereby obtaining the set of human face image samples. Arbitrarily filter out two image samples from the set of human face image samples, and specify any one of the images as a human face image sample in the image samples, then the other image sample can be a human face template image sample, thereby obtaining a pair of image samples.

[0197] (2) The server uses the preset image processing model to replace the object in the human face template image sample with the object in the human face image sample to obtain a predicted human face image.

[0198] For example, the server can use a preset image processing model to extract features from the face image sample and the face template image sample, obtain the image sample texture features of the face image sample and the sample attribute features of the object in the face template image sample, respectively identify the sample three-dimensional modeling parameters in the face image sample and the face template image sample, fuse the identified sample three-dimensional modeling parameters to obtain the target sample three-dimensional detection parameters, construct a sample three-dimensional face image according to the target sample three-dimensional modeling parameters, obtain the template three-dimensional facial features of the sample three-dimensional face image, and fuse the sample three-dimensional facial features, the image sample texture features and the sample attribute features to obtain the fused sample facial features, construct a facial mask corresponding to the fused sample facial features to obtain the initial sample facial mask. Fuse the initial sample facial mask, the fused sample facial features and the sample attribute features to obtain the target sample facial features, adjust the target sample facial features, and construct a facial mask corresponding to the adjusted sample facial features to obtain the target sample facial mask, and generate a predicted face image based on the target sample facial mask and the adjusted sample facial features.

[0199] (3) The server converges the preset image processing model based on the image sample pair and the predicted face image pair to obtain a trained image processing model.

[0200] For example, the server can generate an initial predicted face image based on the target sample facial features and the initial sample facial mask, determine the shape loss information of the image sample pair according to the sample three-dimensional face image, the initial predicted face image and the predicted face image, determine the segmentation loss information of the image sample pair based on the initial sample facial mask and the target sample facial mask, determine the facial loss information of the image sample pair according to the image sample pair, the predicted face image and the initial predicted face image, fuse the shape loss information, the segmentation loss information and the facial loss information, and converge the preset image processing model based on the fused loss information to obtain a trained image processing model, which can be specifically as follows:

[0201] D1. The server generates an initial predicted face image based on the target sample facial features and the initial sample facial mask.

[0202] For example, the server can generate an initial facial sample image according to the target sample facial features, screen out the images within the initial sample facial mask from the candidate face images to obtain a basic facial sample image, screen out the images outside the initial sample facial mask from the face template image sample to obtain a background sample image, and fuse the background sample image and the basic facial sample image to obtain the initial predicted face image.

[0203] D2. The server determines the shape loss information of the image sample pair according to the sample three-dimensional face image, the initial predicted face image and the predicted face image.

[0204] For example, the server can use the 3D renderer pytorch3d to obtain the 2D projection of the sample 3D face image under the preset angle coefficient, obtain the first projection information, and obtain the positions of 18 contour points of the facial contour in the 2D projection, so as to obtain the first position information of the facial contour.

[0205] The server reconstructs the initial predicted face image and the target 3D face image corresponding to the predicted face image by using the facial 3D reconstruction model, obtains the second projection information of the target 3D face image, and extracts the second position information of the facial contour in the initial predicted face image from the second projection information and extracts the third position information of the facial contour in the predicted face image from the third projection information.

[0206] The server calculates the position differences between the contour points of the same facial contour in the first position information and the second position information, and calculates the position differences between the contour points of the same facial contour in the first position information and the third position information. Then, it calculates the average value of the position differences, so as to obtain the shape loss information of the image sample pair, which can be specifically shown in formula (4).

[0207] D3. The server calculates the facial similarities between the facial image sample in the image sample pair and the predicted facial image and the initial predicted facial image respectively to obtain the image loss information of the image sample pair.

[0208] For example, the server can use the pre-trained facial recognition feature extractor to extract the facial features in the facial image sample, the predicted facial image and the initial predicted facial image, calculate the cosine similarity between the facial features of the facial image sample and the predicted facial image, so as to obtain the first facial similarity between the facial image sample and the predicted facial image, and calculate the second facial similarity between the facial image sample and the initial predicted facial image, and directly add the first facial similarity and the second facial similarity to obtain the image loss information, which can be specifically shown in formula (6).

[0209] D4. The server determines the segmentation loss information of the image sample pair based on the initial sample facial mask and the target sample facial mask.

[0210] For example, the server can use the trained semantic segmentation network to predict the template map mask of the face template image sample in the image sample pair. Alternatively, the server can also extract the mask feature from the face template image sample, and based on the mask feature, filter out the mask corresponding to the face template image sample from the preset mask set to obtain the template map mask. Expand the template map mask outward by 15 pixels to obtain the adjusted template map mask. Calculate the size differences between the adjusted template map mask and the initial sample face mask and the target sample face mask respectively, and fuse the size differences to obtain the segmentation loss information of the image sample pair, which can be specifically shown in formula (7).

[0211] D5. The server determines the face loss information of the image sample pair according to the image sample pair, the predicted face image, and the initial predicted face image.

[0212] For example, when the objects in the face image sample and the face template image sample are the same object, the server can use the L1 norm to calculate the first spatial similarity between the face image sample and the predicted face image, and the second spatial similarity between the face image sample and the initial predicted face image, and fuse the first spatial similarity and the second spatial similarity to obtain the spatial similarity loss information of the image sample pair, which can be specifically shown in formula (8).

[0213] The server can also use the feature extraction network to extract the image features of the face template image sample, the predicted face image, and the initial predicted face image, obtain the image features output by each feature layer of the feature extraction network, calculate the feature differences between the image features output by the face template image sample and the predicted face image in the same feature layer, and then fuse the feature differences and the feature sizes of the image features to obtain the first feature similarity, which can be specifically shown in formula (9), and calculate the feature differences between the image features output by the face template image sample and the initial predicted face image in the same feature layer to obtain the second feature similarity. Directly add the first feature similarity and the second feature similarity to obtain the feature similarity loss information. Alternatively, the server can also obtain the weighting coefficients of the first feature similarity and the second feature similarity, weight the first feature similarity and the second feature similarity according to the weighting coefficients, and fuse the weighted first feature similarity and the second feature similarity to obtain the feature similarity loss information. Fuse the face similarity loss information, the spatial similarity loss information, and the feature similarity loss information to obtain the similarity loss information of the image sample pair.

[0214] The server can use a generative adversarial network to calculate the first adversarial parameters of the facial template image and the predicted face image respectively. Based on the first adversarial parameters, the first adversarial loss information can be determined, which can be specifically shown in formula (10). And the server can use the generative adversarial network to calculate the second adversarial parameters of the face template image sample and the initial predicted face image respectively. Based on the second adversarial parameters, the second adversarial loss information can be determined. Add the first adversarial loss information and the second adversarial loss information to obtain the adversarial loss information. Or, the weighted parameters of the first adversarial loss information and the second adversarial loss information can also be obtained. According to the weighted parameters, the first adversarial loss information and the second adversarial loss information are weighted, and the weighted first adversarial loss information and the second adversarial loss information are fused to obtain the adversarial loss information.

[0215] The server can also use the L1 norm to calculate the norm of the facial template image sample and the predicted facial image obtained after image processing, so as to obtain the periodic loss information of the image sample pair, which can be specifically shown in formula (11).

[0216] After obtaining the spatial similarity loss information, the feature similarity loss information, the adversarial loss information and the periodic loss information, the spatial similarity loss information, the feature similarity loss information, the adversarial loss information and the periodic loss information can be used as the facial loss information of the image sample pair.

[0217] D6. The server fuses the shape loss information, the image loss information, the segmentation loss information and the facial loss information, and converges the preset image processing model based on the fused loss information to obtain the trained image processing model.

[0218] For example, the server obtains the first weighted parameter of the shape loss information and the image loss information, weights the shape loss information and the image loss information according to the first weighted parameter, and fuses the weighted shape loss information and the image loss information to obtain the first loss information, which can be specifically shown in formula (12).

[0219] The server obtains the second weighted parameter of the segmentation loss information and the facial loss information, weights the segmentation loss information and the facial loss information according to the second weighted parameter, and fuses the weighted segmentation loss information and the facial loss information to obtain the second loss information, which can be specifically shown in formula (13).

[0220] Add the first loss information and the second loss information to obtain the fused loss information, which can be specifically shown in formula (14).

[0221] After obtaining the merged loss information, the server can use the gradient descent algorithm to update the network parameters of the preset image processing model according to the merged loss information to converge the preset image processing model and obtain the trained image processing model. Alternatively, other convergence algorithms can also be used to update the network parameters of the preset image processing model according to the merged loss information to converge the preset image processing model and obtain the trained image processing model.

[0222] As Figure 9 shown, an image processing method has the following specific process:

[0223] 201. The server acquires the face image of the source face and the face template image of the template face.

[0224] For example, the server can directly receive the original face image uploaded by the user and the image processing information corresponding to the original face image, and according to the image processing information, screen out the face image of the source face and the face template image of the template face from the original face image. Alternatively, the server can obtain a pair of face images in the image database or on the network, arbitrarily select one face image in the pair of face images as the face image of the source face, and use the other face image in the pair of face images as the face template image of the template face.

[0225] When the number of face images and face template images is large or the memory is large, the server can receive the image processing request sent by the terminal. The image processing request carries the storage address of the original face image and the image processing information. According to the storage address, the server obtains the original face image in the memory, cache or third-party database, adjusts the size of the original face image to the preset size, or can also use facial key point registration to align the facial objects in the original face image to the same position to obtain the processed face image, and according to the image processing information, screen out the face image of the source face and the face template image of the template face from the processed face image.

[0226] 202. The server extracts features from the face image and the face template image to obtain the image texture features of the source object and the attribute features of the object in the face template image.

[0227] For example, the server can use the encoder network (stacked by 8 Res-Blocks) of the trained image processing model to perform feature encoding on the face template image to obtain the attribute features of the object in the face template image, and use the face recognition network of the trained image processing model to extract features from the face image to obtain the image texture features of the source object.

[0228] 203. The server performs facial modeling on the source object and the object in the facial template image based on the facial image and the facial template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image, and fuses the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain the target three-dimensional modeling parameters.

[0229] For example, the server uses a three-dimensional facial reconstruction model to regress the facial image and the facial template image, thereby performing facial modeling on the source object and the object in the facial template image, directly extracting the first three-dimensional modeling parameters from the facial model of the source object, and extracting the second three-dimensional modeling parameters from the facial model of the object in the facial template image. Alternatively, the facial model can also be converted into three-dimensional modeling parameters to obtain the first three-dimensional modeling parameters and the second three-dimensional modeling parameters.

[0230] The server extracts the facial shape parameters corresponding to the facial image from the first three-dimensional modeling parameters, extracts the facial action parameters corresponding to the facial template image from the second three-dimensional modeling parameters, and directly concatenates and combines the facial shape parameters and the facial action parameters. Alternatively, the weighted parameters and the basic modeling parameters of the facial shape parameters and the facial action parameters can also be obtained, the facial shape parameters and the facial action parameters are weighted according to the weighted parameters, and the weighted facial shape parameters and the facial action parameters are fused with the basic modeling parameters to obtain the target three-dimensional modeling parameters, which can be specifically shown in formula (1).

[0231] 204. The server constructs a three-dimensional face image based on the target three-dimensional modeling parameters to obtain the three-dimensional facial features of the three-dimensional face image.

[0232] For example, the server can construct a three-dimensional facial model corresponding to the target three-dimensional modeling parameters through a three-dimensional facial reconstruction model to obtain a three-dimensional face image, obtain the three-dimensional features of the three-dimensional facial model, and use the three-dimensional features as the three-dimensional facial features of the three-dimensional face image. Alternatively, a three-dimensional facial model corresponding to the target three-dimensional modeling parameters can be constructed through a three-dimensional facial reconstruction model, the three-dimensional facial model can be locally fitted or optimized, etc. to obtain a three-dimensional face image, obtain the three-dimensional features of the adjusted three-dimensional facial model, and use the three-dimensional features as the three-dimensional facial features of the three-dimensional face image.

[0233] 205. The server concatenates the image texture features and the three-dimensional facial features to obtain facial features.

[0234] For example, the server can directly splice the image texture features and the three-dimensional facial features to obtain the facial features. Or, it can obtain the weighting parameters of the image texture features and the three-dimensional facial features, weight the image texture features and the three-dimensional facial features according to the weighting parameters, and fuse the weighted image texture features and the three-dimensional facial features to obtain the facial features. Or, it can also obtain the feature depths of the image texture features and the three-dimensional facial features, adjust the feature depths to a unified target feature depth, and based on the target feature depth, splice the image texture features and the three-dimensional facial features to obtain the facial features.

[0235] 206. The server uses the trained image processing model to fuse the facial features and the attribute features to obtain the fused facial features.

[0236] For example, the server can use the decoder network (consisting of 5 Res-Blocks with AdaIN stacked) of the trained image processing model to decode the facial features and the attribute features, and fuse the decoded facial features and the attribute features to obtain the fused facial features.

[0237] 207. The server generates the target face image based on the fused facial features.

[0238] For example, the server can construct a facial mask corresponding to the fused facial features to obtain the initial facial mask, fuse the initial facial mask, the fused facial features and the attribute features to obtain the target facial features, adjust the target facial features, and construct a facial mask corresponding to the adjusted facial features to obtain the target facial mask, and generate the target face image based on the target facial mask and the adjusted facial features. Specifically, it can be as follows:

[0239] (1) The server constructs a facial mask corresponding to the fused facial features to obtain the initial facial mask.

[0240] For example, the server can use the semantic segmentation network to extract the features of the fused facial features, and based on the extracted semantic features, identify the segmentation area in the face image or the face template image, and based on the segmentation area, segment the face image or the face template image, and perform occlusion processing on the segmented face image to obtain the initial facial mask. Or, it can use the semantic segmentation network to extract the features of the fused facial features, and based on the extracted semantic features, directly generate the initial facial mask.

[0241] (2) The server fuses the initial facial mask, the fused facial features and the attribute features to obtain the target facial features.

[0242] For example, the server can use one or more Res-Blocks to perform feature transformation on the attribute features to obtain the target attribute features of the object in the face template image. Filter out the area outside the initial facial mask in the face image or face template image to obtain the target image area, and determine the weighting parameter according to the initial facial mask and the target image area. Directly add the weighted facial features and the weighted attribute features to obtain the target facial features, which can be specifically shown in formula (2).

[0243] (3) The server adjusts the target facial features and constructs a facial mask corresponding to the adjusted facial features to obtain the target facial mask.

[0244] For example, the server can use an upsampling structure composed of one or more Res-Blocks to enlarge the size of the target facial features to a preset size to obtain the adjusted facial features. Use a semantic segmentation network to extract features from the adjusted facial features, and based on the extracted semantic features, identify the segmentation area in the face image or face template image, and based on the segmentation area, segment the face image or face template image, and perform occlusion processing on the segmented face image to obtain the target facial mask. Or, a semantic segmentation network can be used to extract features from the adjusted facial features, and based on the extracted semantic features, directly generate the target facial mask.

[0245] (4) The server generates a target face image based on the target facial mask and the adjusted facial features.

[0246] For example, the server can generate an initial face image according to the adjusted facial features, and filter out the image within the target facial mask in the initial face image to obtain the basic face image, and identify the image outside the target facial mask in the face template image to obtain the background image. Stitch the basic face image and the background image to obtain the target face image. Or, the sizes of the basic face image and the background image can also be adjusted, and the adjusted-size basic face image and background image are superimposed to obtain the target face image, which can be specifically shown in formula (3).

[0247] As can be seen from the above, after obtaining the face image of the source face and the face template image of the template face in the embodiments of the present application, feature extraction is performed on the face image and the face template image to obtain the image texture features of the source object and the attribute features of the object in the face template image. Then, based on the face image and the face template image, facial modeling is performed on the source object and the object in the face template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the face template image, and the first three-dimensional modeling parameters and the second three-dimensional modeling parameters are fused to obtain the target three-dimensional modeling parameters. According to the target three-dimensional modeling parameters, a three-dimensional face image is constructed to obtain the three-dimensional face features of the three-dimensional face image. Then, based on the image texture features, three-dimensional face features, and attribute features, the object in the face template image is replaced with the source object to obtain the target face image. Since this solution can identify the three-dimensional modeling parameters in the face image and the face template image, and construct a three-dimensional face image based on the three-dimensional modeling parameters, thereby obtaining three-dimensional face features, making the face features obtain geometric constraints. Moreover, by fusing the three-dimensional face features, image texture features, and attribute features, the result of the replaced face can be made more realistic. Therefore, the accuracy of face image processing can be improved.

[0248] To better implement the above method, an embodiment of the present invention further provides an image processing apparatus, which can be integrated in an electronic device, such as a server or a terminal. The terminal may include a tablet computer, a notebook computer, and / or a personal computer, etc.

[0249] For example, as Figure 10 shown, the image processing apparatus may include an acquisition unit 301, an extraction unit 302, a fusion unit 303, a construction unit 304, and a replacement unit 305, as follows:

[0250] (1) Acquisition unit 301;

[0251] The acquisition unit 301 is configured to acquire the face image of the source face and the face template image of the template face, and the face image includes the source object.

[0252] For example, the acquisition unit 301 may specifically be configured to directly acquire the face image and the face template image, or indirectly acquire the face image and the face template image when the number of face images and face template images is large or the memory is large.

[0253] (2) Extraction unit 302;

[0254] The extraction unit 302 is configured to perform feature extraction on the face image and the face template image to obtain the image texture features of the source object and the attribute features of the object in the face template image.

[0255] For example, the extraction unit 302 can be specifically used to perform feature encoding on the facial template image by using the encoder network of the trained image processing model to obtain the attribute features of the object in the facial template image, and perform feature extraction on the facial image by using the facial recognition network of the trained image processing model to obtain the image texture features of the source object.

[0256] (3) The fusion unit 303;

[0257] The fusion unit 303 is used to perform facial modeling on the source object and the object in the facial template image according to the facial image and the facial template image, so as to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image, and fuse the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain the target three-dimensional modeling parameters.

[0258] For example, the fusion unit 303 can be specifically used to perform regression on the facial image and the facial template image by using a three-dimensional facial reconstruction model, so as to perform facial modeling on the source object and the object in the facial template image, and obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image in the constructed facial model. Extract the facial shape parameters corresponding to the facial image from the first three-dimensional modeling parameters, extract the facial action parameters corresponding to the facial template image from the second three-dimensional modeling parameters, and fuse the facial shape parameters and the facial action parameters to obtain the target three-dimensional modeling parameters.

[0259] (4) The construction unit 304;

[0260] The construction unit 304 is used to construct a three-dimensional facial image according to the target three-dimensional modeling parameters to obtain the three-dimensional facial features of the three-dimensional facial image.

[0261] For example, the construction unit 304 can be specifically used to construct a three-dimensional facial model corresponding to the target three-dimensional modeling parameters through a three-dimensional facial reconstruction model to obtain a three-dimensional facial image, obtain the three-dimensional features of the three-dimensional facial model, and use the three-dimensional features as the three-dimensional facial features of the three-dimensional facial image. Or, a three-dimensional facial model corresponding to the target three-dimensional modeling parameters can be constructed through a three-dimensional facial reconstruction model, and the three-dimensional facial model can be adjusted, for example, local fitting or optimization can be performed, etc., to obtain a three-dimensional facial image, obtain the three-dimensional features of the adjusted three-dimensional facial model, and use the three-dimensional features as the three-dimensional facial features of the three-dimensional facial image.

[0262] (5) The replacement unit 305;

[0263] The replacement unit 305 is used to replace the object in the facial template image with the source object based on the image texture features, three-dimensional facial features, and attribute features to obtain a target facial image.

[0264] For example, the replacement unit 305 can be specifically used to splice the image texture feature and the three-dimensional facial feature to obtain a facial feature, and use the trained image processing model to fuse the facial feature and the attribute feature to obtain a fused facial feature, and generate a target facial image based on the fused facial feature. The target facial image is an image obtained by replacing the object in the facial template image with the source object.

[0265] Optionally, the image processing apparatus may further include a training unit 306, as Figure 11 shown, specifically as follows:

[0266] The training unit 306 can be specifically used to train a preset image processing model to obtain a trained image processing model.

[0267] For example, the training unit 306 can be specifically used to obtain a set of facial image samples, and screen out at least one pair of image samples from the set of facial image samples. The pair of image samples includes a facial image sample and a facial template image sample. Use the preset image processing model to replace the object in the facial template image sample with the object in the facial image sample to obtain a predicted facial image, and converge the preset image processing model based on the pair of image samples and the predicted facial image to obtain a trained image processing model.

[0268] In specific implementation, each of the above units can be implemented as an independent entity, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the foregoing method embodiments, which will not be elaborated herein.

[0269] As can be seen from the above, after the obtaining unit 301 obtains the facial image of the source face and the facial template image of the template face in this embodiment, the extraction unit 302 extracts features from the facial image and the facial template image to obtain the image texture features of the source object and the attribute features of the object in the facial template image. Then, the fusion unit 303 performs facial modeling on the source object and the object in the facial template image according to the facial image and the facial template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image, and fuses the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain the target three-dimensional modeling parameters. The construction unit 304 constructs a three-dimensional facial image according to the target three-dimensional modeling parameters to obtain the three-dimensional facial features of the three-dimensional facial image. Then, the replacement unit 305 replaces the object in the facial template image with the source object based on the image texture features, the three-dimensional facial features, and the attribute features to obtain the target facial image. Since this solution can identify the three-dimensional modeling parameters in the facial image and the facial template image, and construct a three-dimensional facial image based on the three-dimensional modeling parameters, thereby obtaining three-dimensional facial features, enabling the facial features to obtain geometric constraints. Moreover, by fusing the three-dimensional facial features, the image texture features, and the attribute features, the result of the replaced face can be made more realistic. Therefore, the accuracy of facial image processing can be improved.

[0270] An embodiment of the present invention also provides an electronic device, as Figure 12 shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of the present invention. Specifically:

[0271] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404, and other components. Those skilled in the art can understand that Figure 12 the structure of the electronic device shown in

[0272] does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Among them:

[0273] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device. In addition, the memory 402 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 can also include a memory controller to provide the processor 401 with access to the memory 402.

[0274] The electronic device further includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0275] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0276] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:

[0277] Obtain the facial image of the source face and the facial template image of the template face. The facial image includes the source object. Extract features from the facial image and the facial template image to obtain the image texture features of the source object and the attribute features of the object in the facial template image. According to the facial image and the facial template image, perform facial modeling on the source object and the object in the facial template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image, and fuse the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain the target three-dimensional modeling parameters; construct a three-dimensional facial image according to the target three-dimensional modeling parameters to obtain the three-dimensional facial features of the three-dimensional facial image; based on the image texture features, three-dimensional facial features and attribute features, replace the object in the facial template image with the source object to obtain the target facial image.

[0278] For example, an electronic device obtains a facial image and a facial template image, or, when the number of facial images and facial template images is large or the memory is large, indirectly obtains the facial image and the facial template image. Use the encoder network of the trained image processing model to perform feature encoding on the facial template image to obtain the attribute features of the object in the facial template image, and use the facial recognition network of the trained image processing model to perform feature extraction on the facial image to obtain the image texture features of the source object. Use a three-dimensional facial reconstruction model to perform regression on the facial image and the facial template image, so as to perform facial modeling on the source object and the object in the facial template image, and obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image in the constructed facial model. Extract the facial shape parameters corresponding to the facial image from the first three-dimensional modeling parameters, and extract the facial action parameters corresponding to the facial template image from the second three-dimensional modeling parameters. Fuse the facial shape parameters and the facial action parameters to obtain the target three-dimensional modeling parameters. Construct a three-dimensional facial model corresponding to the target three-dimensional modeling parameters through the three-dimensional facial reconstruction model to obtain a three-dimensional facial image, and obtain the three-dimensional features of the three-dimensional facial model. Use the three-dimensional features as the three-dimensional facial features of the three-dimensional facial image, or, a three-dimensional facial model corresponding to the target three-dimensional modeling parameters can be constructed through the three-dimensional facial reconstruction model, and the three-dimensional facial model can be adjusted, for example, local fitting or optimization can be performed, etc., to obtain a three-dimensional facial image, and obtain the three-dimensional features of the adjusted three-dimensional facial model. Use the three-dimensional features as the three-dimensional facial features of the three-dimensional facial image. Concatenate the image texture features and the three-dimensional facial features to obtain facial features, and use the trained image processing model to fuse the facial features and the attribute features to obtain the fused facial features. Generate a target facial image based on the fused facial features. The target facial image is an image in which the object in the facial template image is replaced with the source object.

[0279] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.

[0280] As can be seen from the above, after obtaining the face image of the source face and the face template image of the template face in the embodiments of the present application, feature extraction is performed on the face image and the face template image to obtain the image texture features of the source object and the attribute features of the object in the face template image. Then, based on the face image and the face template image, face modeling is performed on the source object and the object in the face template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the face template image, and the first three-dimensional modeling parameters and the second three-dimensional modeling parameters are fused to obtain the target three-dimensional modeling parameters. A three-dimensional face image is constructed according to the target three-dimensional modeling parameters to obtain the three-dimensional face features of the three-dimensional face image. Then, based on the image texture features, the three-dimensional face features, and the attribute features, the object in the face template image is replaced with the source object to obtain the target face image. Since this solution can identify the three-dimensional modeling parameters in the face image and the face template image, and construct a three-dimensional face image based on the three-dimensional modeling parameters, thereby obtaining the three-dimensional face features, the face features are geometrically constrained. Moreover, by fusing the three-dimensional face features, the image texture features, and the attribute features, the result of the replaced face can be made more realistic. Therefore, the accuracy of face image processing can be improved.

[0281] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by instructions controlling related hardware. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0282] For this reason, an embodiment of the present invention provides a computer-readable storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps in any one of the image processing methods provided by the embodiments of the present invention. For example, the instructions can execute the following steps:

[0283] Obtain the face image of the source face and the face template image of the template face. The face image includes the source object. Perform feature extraction on the face image and the face template image to obtain the image texture features of the source object and the attribute features of the object in the face template image. Based on the face image and the face template image, perform face modeling on the source object and the object in the face template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the face template image, and fuse the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain the target three-dimensional modeling parameters; construct a three-dimensional face image according to the target three-dimensional modeling parameters to obtain the three-dimensional face features of the three-dimensional face image; based on the image texture features, the three-dimensional face features, and the attribute features, replace the object in the face template image with the source object to obtain the target face image.

[0284] For example, an electronic device acquires a facial image and a facial template image, or, when the number of facial images and facial template images is large or the memory is large, indirectly acquires the facial image and the facial template image. The encoder network of the trained image processing model is used to perform feature encoding on the facial template image to obtain the attribute features of the object in the facial template image, and the facial recognition network of the trained image processing model is used to extract the image texture features of the source object from the facial image. A three-dimensional facial reconstruction model is used to perform regression on the facial image and the facial template image, so as to perform facial modeling on the source object and the object in the facial template image, and obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image in the constructed facial model. The facial shape parameters corresponding to the facial image are extracted from the first three-dimensional modeling parameters, and the facial motion parameters corresponding to the facial template image are extracted from the second three-dimensional modeling parameters. The facial shape parameters and the facial motion parameters are fused to obtain the target three-dimensional modeling parameters. A three-dimensional facial model corresponding to the target three-dimensional modeling parameters is constructed through the three-dimensional facial reconstruction model to obtain a three-dimensional facial image, and the three-dimensional features of the three-dimensional facial model are acquired. The three-dimensional features are used as the three-dimensional facial features of the three-dimensional facial image. Alternatively, a three-dimensional facial model corresponding to the target three-dimensional modeling parameters can be constructed through the three-dimensional facial reconstruction model, and the three-dimensional facial model can be adjusted, for example, local fitting or optimization can be performed, etc., to obtain a three-dimensional facial image, and the three-dimensional features of the adjusted three-dimensional facial model are acquired. The three-dimensional features are used as the three-dimensional facial features of the three-dimensional facial image. The image texture features and the three-dimensional facial features are spliced to obtain facial features, and the trained image processing model is used to fuse the facial features and the attribute features to obtain the fused facial features. A target facial image is generated based on the fused facial features, and the target facial image is an image obtained by replacing the object in the facial template image with the source object.

[0285] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated herein.

[0286] Wherein, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0287] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the image processing methods provided in the embodiments of the present invention, the beneficial effects that can be achieved by any of the image processing methods provided in the embodiments of the present invention can be realized. For details, reference may be made to the previous embodiments and will not be elaborated herein.

[0288] Wherein, according to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various alternative implementations of the above-mentioned image processing aspect or the image face swapping aspect.

[0289] The above has introduced in detail an image processing method, apparatus and computer-readable storage medium provided by the embodiments of the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A facial image processing method, characterized in that, Including: Obtaining a facial image of a source face and a facial template image of a template face, where the facial image includes a source object; Performing feature extraction on the facial image and the facial template image to obtain the image texture feature of the source object and the attribute feature of the object in the facial template image; Performing facial modeling on the source object and the object in the facial template image according to the facial image and the facial template image to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image, and fusing the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain target three-dimensional modeling parameters; Constructing a three-dimensional facial image according to the target three-dimensional modeling parameters to obtain the three-dimensional facial feature of the three-dimensional facial image; Using a trained image processing model to fuse the facial feature and the attribute feature to obtain a fused facial feature, where the facial feature is obtained by splicing the image texture feature and the three-dimensional facial feature. The trained image processing model is a model obtained by converging a preset image processing model based on facial image samples, facial template image samples, and predicted facial images. The predicted facial image is generated based on the fused sample facial feature, the initial sample facial mask corresponding to the fused sample facial feature, and the sample attribute feature of the object in the facial template image sample. The fused sample facial feature is the target sample three-dimensional modeling parameter obtained by fusing the sample three-dimensional modeling parameter of the facial image sample and the sample three-dimensional modeling parameter of the facial template image sample, and is obtained by fusing the sample three-dimensional facial feature of the constructed sample three-dimensional facial image with the image sample texture feature and the sample attribute feature of the facial image sample; Generating a target facial image based on the fused facial feature, where the target facial image is an image obtained by replacing the object in the facial template image with the source object.

2. The facial image processing method according to claim 1, wherein, The fusing the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain target three-dimensional modeling parameters includes: Extracting the facial shape parameter corresponding to the facial image from the first three-dimensional modeling parameters; Extracting the facial action parameter corresponding to the facial template image from the second three-dimensional modeling parameters; Fusing the facial shape parameter and the facial action parameter to obtain target three-dimensional modeling parameters.

3. The facial image processing method according to claim 1, wherein The generating a target facial image based on the fused facial feature includes: Constructing a facial mask corresponding to the fused facial feature to obtain an initial facial mask; Fusing the initial facial mask, the fused facial feature, and the attribute feature to obtain a target facial feature; Adjusting the target facial feature and constructing a facial mask corresponding to the adjusted facial feature to obtain a target facial mask; Generating a target facial image based on the target facial mask and the adjusted facial feature.

4. The facial image processing method according to claim 3, wherein The fusing the initial facial mask, the fused facial feature, and the attribute feature to obtain a target facial feature includes: Performing feature transformation on the attribute feature to obtain the target attribute feature of the object in the facial template image; Determining the weighting parameter of the fused facial feature and the target attribute feature according to the initial facial mask; According to the weighted parameters, the fused facial features and target attribute features are weighted, and the weighted facial features and weighted attribute features are fused to obtain the target facial features.

5. The facial image processing method according to claim 3, wherein Generating the target facial image based on the target facial mask and the adjusted facial features includes: Generating an initial facial image according to the adjusted facial features, and screening out the image within the target facial mask in the initial facial image to obtain a basic facial image; Identifying the image outside the target facial mask in the facial template image to obtain a background image; Fusing the basic facial image and the background image to obtain the target facial image.

6. The facial image processing method according to claim 1, wherein, Before fusing the facial features and attribute features by using the trained image processing model to obtain the fused facial features, it further includes: Obtaining a set of facial image samples, and screening out at least one pair of image samples from the set of facial image samples, where the pair of image samples includes a facial image sample and a facial template image sample.

7. The facial image processing method according to claim 6, wherein Before fusing the facial features and attribute features by using the trained image processing model to obtain the fused facial features, it further includes: Using a preset image processing model to extract features from the facial image sample and the facial template image sample to obtain the image sample texture features of the facial image sample and the sample attribute features of the object in the facial template image sample; Identifying the sample three-dimensional modeling parameters in the facial image sample and the facial template image sample respectively, and fusing the identified sample three-dimensional modeling parameters to obtain the target sample three-dimensional modeling parameters; Constructing a sample three-dimensional facial image according to the target sample three-dimensional modeling parameters to obtain the sample three-dimensional facial features of the sample three-dimensional facial image, and fusing the sample three-dimensional facial features, the image sample texture features and the sample attribute features to obtain the fused sample facial features; Constructing a facial mask corresponding to the fused sample facial features to obtain an initial sample facial mask, and generating a predicted facial image based on the initial sample facial mask, the fused sample facial features and the sample attribute features.

8. The facial image processing method according to claim 7, characterized in that Generating a predicted image based on the initial sample facial mask, the fused sample facial features and the sample attribute features includes: Fusing the initial sample facial mask, the fused sample facial features and the sample attribute features to obtain the target sample facial features; Adjusting the target sample facial features, and constructing a facial mask corresponding to the adjusted sample facial features to obtain the target sample facial mask; Generating a predicted facial image based on the target sample facial mask and the adjusted sample facial features.

9. The facial image processing method according to claim 8, wherein Before fusing the facial features and attribute features by using the trained image processing model to obtain the fused facial features, it further includes: Generating an initial predicted facial image based on the target sample facial features and the initial sample facial mask; Determining the shape loss information of the pair of image samples according to the sample three-dimensional facial image, the initial predicted facial image and the predicted facial image. Calculate the facial similarity between the facial image samples in the image sample pair and the predicted facial image and the initial predicted facial image respectively, so as to obtain the image loss information of the image sample pair; Determine the segmentation loss information of the image sample pair based on the initial sample facial mask and the target sample facial mask; Determine the facial loss information of the image sample pair according to the image sample pair, the predicted facial image and the initial predicted facial image; Fuse the shape loss information, the image loss information, the segmentation loss information and the facial loss information, and converge the preset image processing model based on the fused loss information to obtain the trained image processing model.

10. The facial image processing method according to claim 9, wherein The determining the shape loss information of the image sample pair according to the sample three-dimensional facial image, the initial predicted facial image and the predicted facial image includes: Obtain the first projection information of the sample three-dimensional facial image, and extract the first position information of the facial contour from the first projection information; Construct a target three-dimensional facial image corresponding to the initial predicted facial image and the predicted facial image, and obtain the second projection information of the target three-dimensional facial image; Extract the second position information of the facial contour in the initial predicted facial image and the third position information of the facial contour in the predicted facial image from the second projection information, and calculate the distance between the facial contours respectively according to the first position information, the second position information and the third position information, so as to obtain the shape loss information of the image pair.

11. The facial image processing method according to claim 9, wherein The determining the segmentation loss information of the image sample pair based on the initial sample facial mask and the target sample facial mask includes: Obtain the template map mask of the facial template image sample in the image sample pair; Adjust the size of the template map mask to obtain the adjusted template map mask; Calculate the size differences between the adjusted template map mask and the initial sample facial mask and the target sample facial mask respectively, and fuse the size differences to obtain the segmentation loss information of the image sample pair.

12. The facial image processing method according to claim 9, wherein The determining the facial loss information of the image sample pair according to the image sample pair, the predicted facial image and the initial predicted facial image includes: Calculate the similarities between the image sample pair and the predicted facial image and the initial predicted facial image respectively, so as to obtain the similarity loss information of the image sample pair; Determine the adversarial loss information and the cycle loss information of the image sample pair according to the facial template image sample, the predicted facial image and the initial predicted facial image in the image sample pair; Use the similarity loss information, the adversarial loss information and the cycle loss information as the facial loss information of the image sample pair.

13. The facial image processing method according to claim 12, wherein, The calculating the similarities between the image sample pair and the predicted facial image and the initial predicted facial image respectively, so as to obtain the similarity loss information of the image sample pair includes: When the objects in the facial image sample and the facial template image sample are the same object, calculate the spatial similarities between the facial template image sample and the predicted facial image and the initial predicted facial image respectively, so as to obtain the spatial similarity loss information of the image sample pair; Extract the image features of the facial template image sample, the predicted facial image, and the initial predicted facial image, and calculate the feature similarity between the image features to obtain the feature similarity loss information of the image sample pair; Use the spatial similarity loss information and the feature similarity loss information as the similarity loss information of the image sample pair.

14. A facial image processing device, characterized in that, Comprising: An acquisition unit, configured to acquire a facial image of a source face and a facial template image of a template face, where the facial image includes a source object; An extraction unit, configured to perform feature extraction on the facial image and the facial template image to obtain the image texture features of the source object and the attribute features of the object in the facial template image; A fusion unit, configured to perform facial modeling on the source object and the object in the facial template image according to the facial image and the facial template image, to obtain the first three-dimensional modeling parameters of the source object and the second three-dimensional modeling parameters of the object in the facial template image, and fuse the first three-dimensional modeling parameters and the second three-dimensional modeling parameters to obtain target three-dimensional modeling parameters; A construction unit, configured to construct a three-dimensional facial image according to the target three-dimensional modeling parameters to obtain the three-dimensional facial features of the three-dimensional facial image; A replacement unit, configured to fuse the facial features and the attribute features by using a trained image processing model to obtain the fused facial features, where the facial features are obtained by splicing the image texture features and the three-dimensional facial features, the trained image processing model is a model obtained by converging a preset image processing model based on a facial image sample, a facial template image sample, and a predicted facial image, the predicted facial image is generated based on the fused sample facial features, the initial sample facial mask corresponding to the fused sample facial features, and the sample attribute features of the object in the facial template image sample, the fused sample facial features are the target sample three-dimensional modeling parameters obtained by fusing the sample three-dimensional modeling parameters of the facial image sample and the sample three-dimensional modeling parameters of the facial template image sample, and are obtained by fusing the sample three-dimensional facial features of the constructed sample three-dimensional facial image with the image sample texture features and the sample attribute features of the facial image sample; Generate a target facial image based on the fused facial features, where the target facial image is an image obtained by replacing the object in the facial template image with the source object.

15. The facial image processing device according to claim 14, wherein The fusion unit is configured to extract the facial shape parameters corresponding to the facial image from the first three-dimensional modeling parameters; extract the facial action parameters corresponding to the facial template image from the second three-dimensional modeling parameters; and fuse the facial shape parameters and the facial action parameters to obtain target three-dimensional modeling parameters.

16. The facial image processing device according to claim 14, wherein The replacement unit is configured to construct a facial mask corresponding to the fused facial features to obtain an initial facial mask; fuse the initial facial mask, the fused facial features, and the attribute features to obtain target facial features; adjust the target facial features, and construct a facial mask corresponding to the adjusted facial features to obtain a target facial mask; and generate a target facial image based on the target facial mask and the adjusted facial features.

17. The facial image processing device according to claim 16, wherein The replacement unit is specifically configured to perform feature transformation on the attribute features to obtain the target attribute features of the object in the facial template image; determine the weighting parameters of the fused facial features and the target attribute features according to the initial facial mask; weight the fused facial features and the target attribute features according to the weighting parameters, and fuse the weighted facial features and the weighted attribute features to obtain the target facial features.

18. The facial image processing apparatus according to claim 16, wherein The replacement unit is configured to generate an initial facial image according to the adjusted facial features, and screen out the image within the target facial mask in the initial facial image to obtain a basic facial image; identify the image outside the target facial mask in the facial template image to obtain a background image; fuse the basic facial image and the background image to obtain a target facial image.

19. The facial image processing apparatus according to claim 14, wherein The facial image processing device further includes a training unit, and the training unit is configured to obtain a set of facial image samples, and screen out at least one pair of image samples from the set of facial image samples, and the pair of image samples includes a facial image sample and a facial template image sample.

20. The facial image processing apparatus according to claim 19, wherein The facial image processing device further includes a training unit, and the training unit is configured to extract features from the facial image sample and the facial template image by using a preset image processing model to obtain the image sample texture features of the facial image sample and the sample attribute features of the object in the facial template image sample. Identify the sample three-dimensional modeling parameters in the facial image sample and the facial template image sample respectively, and fuse the identified sample three-dimensional modeling parameters to obtain the target sample three-dimensional modeling parameters. Construct a sample three-dimensional facial image according to the target sample three-dimensional modeling parameters to obtain the sample three-dimensional facial features of the sample three-dimensional facial image, and fuse the sample three-dimensional facial features, the image sample texture features and the sample attribute features to obtain the fused sample facial features; construct a facial mask corresponding to the fused sample facial features to obtain an initial sample facial mask, and generate a predicted facial image based on the initial sample facial mask, the fused sample facial features and the sample attribute features.

21. The facial image processing device according to claim 20, wherein The training unit is configured to fuse the initial sample facial mask, the fused sample facial features and the sample attribute features to obtain the target sample facial features; adjust the target sample facial features, and construct a facial mask corresponding to the adjusted sample facial features to obtain the target sample facial mask; generate a predicted facial image based on the target sample facial mask and the adjusted sample facial features.

22. The facial image processing apparatus according to claim 21, wherein The training unit is used to generate an initial predicted facial image based on the target sample facial features and the initial sample facial mask; determine the shape loss information of the image sample pair according to the sample three-dimensional facial image, the initial predicted facial image and the predicted facial image; calculate the facial similarity between the facial image samples in the image sample pair and the predicted facial image and the initial predicted facial image respectively to obtain the image loss information of the image sample pair; determine the segmentation loss information of the image sample pair based on the initial sample facial mask and the target sample facial mask; determine the facial loss information of the image sample pair according to the image sample pair, the predicted facial image and the initial predicted facial image; fuse the shape loss information, the image loss information, the segmentation loss information and the facial loss information, and converge the preset image processing model based on the fused loss information to obtain the trained image processing model.

23. The facial image processing device according to claim 22, wherein The training unit is used to obtain the first projection information of the sample three-dimensional facial image and extract the first position information of the facial contour from the first projection information; construct the target three-dimensional facial image corresponding to the initial predicted facial image and the predicted facial image, and obtain the second projection information of the target three-dimensional facial image; extract the second position information of the facial contour in the initial predicted facial image and the third position information of the facial contour in the predicted facial image from the second projection information, and calculate the distance between the facial contours respectively according to the first position information, the second position information and the third position information to obtain the shape loss information of the image sample pair.

24. The facial image processing device according to claim 22, wherein The training unit is used to obtain the template map mask of the facial template image sample in the image sample pair; adjust the size of the template map mask to obtain the adjusted template map mask. Calculate the size differences between the adjusted template map mask and the initial sample facial mask and the target sample facial mask respectively, and fuse the size differences to obtain the segmentation loss information of the image sample pair.

25. The facial image processing device according to claim 22, wherein The training unit is used to calculate the similarity between the image sample pair and the predicted facial image and the initial predicted facial image respectively to obtain the similarity loss information of the image sample pair; determine the adversarial loss information and the cycle loss information of the image sample pair according to the facial template image sample, the predicted facial image and the initial predicted facial image in the image sample pair; use the similarity loss information, the adversarial loss information and the cycle loss information as the facial loss information of the image sample pair.

26. The facial image processing device according to claim 25, wherein, When the objects in the facial image sample and the facial template image sample are the same object, the training unit is configured to calculate the spatial similarity between the facial template image sample and the predicted facial image and the initial predicted facial image respectively, so as to obtain the spatial similarity loss information of the image sample pair; extract the image features of the facial template image sample, the predicted facial image and the initial predicted facial image, and calculate the feature similarity between the image features, so as to obtain the feature similarity loss information of the image sample pair; and use the spatial similarity loss information and the feature similarity loss information as the similarity loss information of the image sample pair.

27. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the facial image processing method according to any one of claims 1 to 13.

28. An electronic device, comprising a processor and a memory, where the memory stores an application program, and the processor is configured to run the application program in the memory to implement the steps in the facial image processing method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Face image processing method and device, storage medium and electronic equipment

    CN111967397A

  • Face image fusion method and device, storage medium and electronic equipment

    CN112257657A