Facial image processing method and training method of facial image processing model

The proposed method enhances face image processing accuracy by employing three-dimensional modeling and feature fusion to preserve facial contours during face swapping, addressing the neglect of contours in existing methods.

CN114973349BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110963370.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-20
Publication Date
2025-07-15
Estimated Expiration
2041-08-20

AI Technical Summary

Technical Problem

When performing facial image processing methods, existing facial image processing methods often ignore facial contours, resulting in reduced processing accuracy.

Method used

Through three-dimensional facial modeling and generation adversarial networks, the three-dimensional features of facial images and template images are obtained, fusion processing is performed, and facial replacement features are extracted and transformed by training the facial image processing model to ensure the accuracy of facial contours.

Benefits of technology

Improve the accuracy of facial image processing, maintain the expression, texture, angle and lighting information of the facial template image, and replace it with the identity of the source face.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973349B_ABST
    Figure CN114973349B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method for facial image processing and a method for training a facial image processing model. It is possible to obtain a facial image of a source face and a facial template image of a template face; perform three-dimensional facial modeling on the facial image and the facial template image to obtain three-dimensional facial image features of the facial image and three-dimensional facial template image features of the facial template image; perform fusion processing on the three-dimensional facial image features and the three-dimensional facial template image features to obtain fused three-dimensional facial image features; perform facial replacement feature extraction processing on the facial image based on the facial template image to obtain initial facial replacement features; perform conversion processing on the initial facial replacement features based on the fused three-dimensional facial image features to obtain target facial replacement features; and replace the template face in the facial template image with the source face based on the target facial replacement features and the facial features of the facial image to obtain a replaced facial image, thereby improving the accuracy of facial image processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method for facial image processing and a method for training a facial image processing model. Background Art

[0002] In recent years, with the development of technology, in applications such as movie special effects and Internet social networking, there is a need to replace the face of an object in a facial image with the face of another object while maintaining the style of the object in the facial image. To meet this need, facial images need to be processed. Existing facial image processing methods mainly achieve the facial replacement of an object through three-dimensional modeling or generative adversarial networks.

[0003] In the research and practice of the prior art, the inventors of the present application found that the existing methods for facial replacement only consider the preservation of the identity of the face, while the facial contour, as the edge region of the face, is often ignored, which will reduce the accuracy of facial image processing. Summary of the Invention

[0004] Embodiments of the present application propose a method for facial image processing and a method for training a facial image processing model, which improve the accuracy of facial image processing.

[0005] Embodiments of the present application provide a method for facial image processing, including:

[0006] Obtaining a facial image of a source face and a facial template image of a template face;

[0007] Performing three-dimensional facial modeling on the facial image and the facial template image to obtain three-dimensional facial image features of the facial image and three-dimensional facial template image features of the facial template image;

[0008] Fusing the three-dimensional facial image features and the three-dimensional facial template image features to obtain fused three-dimensional facial image features;

[0009] Based on the facial template image, performing facial replacement feature extraction processing on the facial image to obtain initial facial replacement features;

[0010] Performing conversion processing on the initial facial replacement features based on the fused three-dimensional facial image features to obtain target facial replacement features;

[0011] Using a trained facial image processing model, based on the target facial replacement features and the facial features of the facial image, replacing the template face in the facial template image with the source face to obtain a replaced facial image.

[0012] Correspondingly, embodiments of the present application further provide a facial image processing apparatus, including:

[0013] A first acquisition unit, configured to acquire a facial image of a source face and a facial template image of a template face;

[0014] A three-dimensional facial modeling unit, configured to perform three-dimensional facial modeling on the facial image and the facial template image to obtain three-dimensional facial image features of the facial image and three-dimensional facial template image features of the facial template image;

[0015] A first fusion unit, configured to perform a fusion process on the three-dimensional facial image features and the three-dimensional facial template image features to obtain fused three-dimensional facial image features;

[0016] A feature extraction unit, configured to perform facial replacement feature extraction processing on the facial image based on the facial template image to obtain initial facial replacement features;

[0017] A conversion unit, configured to perform a conversion process on the initial facial replacement features based on the fused three-dimensional facial image features to obtain target facial replacement features;

[0018] A first replacement unit, configured to use a trained facial image processing model to replace the template face in the facial template image with the source face based on the target facial replacement features and the facial features of the facial image to obtain a replaced facial image.

[0019] In one embodiment, the first fusion unit includes:

[0020] A first extraction subunit, configured to extract source face identity parameters corresponding to the facial image from the three-dimensional facial image features;

[0021] A second extraction subunit, configured to extract template facial image parameters corresponding to the facial template image from the three-dimensional facial template image features;

[0022] A first fusion subunit, configured to fuse the source face identity parameters and the template facial image parameters to obtain the fused three-dimensional facial image features.

[0023] In one embodiment, the feature extraction unit includes:

[0024] A first encoding subunit, configured to perform encoding processing on the facial template image to obtain first encoding features of the facial template image;

[0025] A second encoding subunit, configured to perform encoding processing on the facial image to obtain second encoding features of the facial image;

[0026] A first adjustment subunit, configured to adjust the first encoded feature based on the second encoded feature to obtain the initial facial replacement feature.

[0027] In one embodiment, the conversion unit includes:

[0028] A first statistical subunit, configured to perform statistical processing on the fused three-dimensional facial image feature to obtain a statistically processed three-dimensional facial image feature, and perform statistical processing on the initial facial replacement feature to obtain a statistically processed facial replacement feature;

[0029] A second statistical subunit, configured to perform logical operation processing on the initial facial replacement feature and the statistically processed facial replacement feature to obtain an operationally processed facial replacement feature;

[0030] A logical operation processing subunit, configured to perform logical operation processing on the operationally processed facial replacement feature and the statistically processed three-dimensional facial image feature to obtain the target facial replacement feature.

[0031] An embodiment of the present application further provides a method for training a facial image processing model, including:

[0032] Obtaining a training image sample group, where the training image sample group includes facial image samples, facial template image samples, and facial reference image samples;

[0033] Using a preset facial image processing model to replace the template face in the facial template image sample with the source face in the facial image sample to obtain a predicted facial image;

[0034] Performing three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image, and performing three-dimensional facial contour point detection on the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample;

[0035] Calculating the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain the facial contour loss information between the predicted facial image and the facial reference image sample;

[0036] Adjusting the preset facial image processing model based on the facial contour loss information to obtain a trained facial image processing model.

[0037] Correspondingly, an embodiment of the present application further provides a training device for a facial image processing model, including:

[0038] A second obtaining unit, configured to obtain a training image sample group, where the training image sample group includes facial image samples, facial template image samples, and facial reference image samples;

[0039] A second replacement unit, configured to replace the template face in the template face image sample with the source face in the face image sample by using a preset face image processing model, so as to obtain a predicted face image;

[0040] A three-dimensional face contour point detection unit, configured to perform three-dimensional face contour point detection on the predicted face image to obtain the three-dimensional face contour points of the predicted face image, and perform three-dimensional face contour point detection on the face reference image sample to obtain the three-dimensional face contour points of the face reference image sample;

[0041] A calculation unit, configured to calculate the difference between the three-dimensional face contour points of the predicted face image and the three-dimensional face contour points of the face reference image sample, so as to obtain the face contour loss information between the predicted face image and the face reference image sample;

[0042] An adjustment unit, configured to adjust the preset face image processing model based on the face contour loss information to obtain a trained face image processing model.

[0043] In one embodiment, the three-dimensional face contour point detection unit includes:

[0044] A three-dimensional face modeling subunit, configured to perform three-dimensional face modeling on the predicted face image to obtain the three-dimensional predicted face image features of the predicted face image;

[0045] A three-dimensional key point projection subunit, configured to perform three-dimensional key point projection on the three-dimensional predicted face image features to obtain the three-dimensional face key points of the predicted face image;

[0046] A screening subunit, configured to screen out the three-dimensional face contour points from the three-dimensional face key points based on the three-dimensional face key points.

[0047] In one embodiment, the three-dimensional key point projection subunit includes:

[0048] An extraction module, configured to extract the predicted face identity parameters and predicted face expression parameters of the predicted face image from the three-dimensional predicted face image features;

[0049] A three-dimensional key point projection module, configured to perform three-dimensional key point projection on the predicted face identity parameters and predicted face expression parameters by using a preset transfer parameter to obtain the three-dimensional face key points of the predicted face image.

[0050] In one embodiment, the training device for the face image processing model further includes:

[0051] A first calculation unit, configured to calculate the difference between the facial reference image sample and the predicted facial image except for the three-dimensional facial contour points, so as to obtain first loss information of the facial reference image sample and the predicted facial image, where the first loss information includes loss information other than the facial contour loss information;

[0052] A second calculation unit, configured to calculate the facial feature loss information between the facial image sample and the predicted facial image;

[0053] A second fusion unit, configured to perform a fusion process on the first loss information and the facial feature loss information to obtain second loss information.

[0054] In one embodiment, the first calculation unit includes:

[0055] A first calculation subunit, configured to calculate the pixel difference between the facial reference image sample and the predicted facial image to obtain pixel loss information;

[0056] A second calculation subunit, configured to calculate the feature difference between the facial reference image sample and the predicted facial image to obtain feature loss information;

[0057] A third calculation subunit, configured to calculate the discrimination difference between the facial reference image sample and the predicted facial image to obtain discrimination loss information.

[0058] In one embodiment, the second calculation subunit includes:

[0059] A first calculation module, configured to calculate the two-dimensional feature difference between the facial reference image sample and the predicted facial image to obtain two-dimensional feature loss information;

[0060] A second calculation module, configured to calculate the three-dimensional feature difference between the facial reference image sample and the predicted facial image to obtain three-dimensional feature loss information;

[0061] A first fusion module, configured to fuse the two-dimensional feature loss information and the three-dimensional feature loss information to obtain the feature loss information.

[0062] In one embodiment, the third calculation subunit includes:

[0063] A scale transformation module, configured to perform scale transformation processing on the facial reference image sample and the predicted facial image respectively to obtain at least one facial reference image sample after scale transformation and at least one predicted facial image after scale transformation;

[0064] A discrimination module, configured to perform discrimination processing on the at least one face reference image sample after the at least one scale transformation and the at least one predicted face image after the at least one scale transformation respectively, to obtain a first discrimination parameter of the face reference image sample after the scale transformation and a second discrimination parameter of the predicted face image after the scale transformation;

[0065] A third calculation module, configured to calculate the discrimination loss information based on the first discrimination parameter and the second discrimination parameter.

[0066] In one embodiment, the adjustment unit includes:

[0067] An acquisition subunit, configured to acquire the model parameters of the trained face image processing model;

[0068] A second fusion subunit, configured to perform fusion processing on the face contour information and the second loss information to obtain third loss information;

[0069] A parameter adjustment unit, configured to adjust the model parameters by using the third loss information to obtain the trained face image processing model.

[0070] An embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional manners in the above-mentioned aspect.

[0071] Correspondingly, an embodiment of the present application also provides a storage medium. The storage medium stores instructions, and when the instructions are executed by a processor, the automatic reply processing method or the training method of the face image processing model provided in any one of the embodiments of the present application is implemented.

[0072] An embodiment of the present application can acquire a face image of a source face and a face template image of a template face; perform three-dimensional face modeling on the face image and the face template image to obtain three-dimensional face image features of the face image and three-dimensional face template image features of the face template image; perform fusion processing on the three-dimensional face image features and the three-dimensional face template image features to obtain fused three-dimensional face image features; perform face replacement feature extraction processing on the face image based on the face template image to obtain initial face replacement features; perform conversion processing on the initial face replacement features based on the fused three-dimensional face image features to obtain target face replacement features; use the trained face image processing model to replace the template face in the face template image with the source face based on the target face replacement features and the face features of the face image, so as to obtain a replaced face image, thereby improving the accuracy of face image processing. Brief Description of the Drawings

[0073] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0074] Figure 1 is a schematic diagram of the scenario of the method provided by the embodiment of the present application;

[0075] Figure 2 is a flowchart of the facial image processing method provided by the embodiment of the present application;

[0076] Figure 3 is a schematic diagram of the scenario of the facial image processing method provided by the embodiment of the present application;

[0077] Figure 4 is another schematic diagram of the scenario of the facial image processing method provided by the embodiment of the present application;

[0078] Figure 5 is a flowchart of the training method of the facial image processing model provided by the embodiment of the present application;

[0079] Figure 6 is a schematic diagram of the scenario of the training method of the facial image processing model provided by the embodiment of the present application;

[0080] Figure 7 is another flowchart of the facial image processing method provided by the embodiment of the present application;

[0081] Figure 8 is another flowchart of the training method of the facial image processing model provided by the embodiment of the present application;

[0082] Figure 9 is a schematic diagram of the structure of the facial image processing device provided by the embodiment of the present application;

[0083] Figure 10 is a schematic diagram of the structure of the training device of the facial image processing model provided by the embodiment of the present application;

[0084] Figure 11 is a schematic diagram of the structure of the terminal provided by the embodiment of the present application. Detailed Description of the Embodiments

[0085] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. However, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0086] An embodiment of the present application provides a facial image processing method. This facial image processing method can be executed by a facial image processing device, which can be integrated into a computer device. Among them, the computer device can include at least one of a terminal, a server, and the like.

[0087] In one embodiment, to better implement the facial image processing method proposed in the embodiments of the present application, correspondingly, the embodiments of the present application also provide a training method for a facial image processing model, so that the facial image processing method proposed in the embodiments of the present application can be executed using the facial image processing model. Specifically, the embodiments of the present application provide a training method for a facial image processing model. This training method for a facial image processing model can be executed by a training device for the facial image processing model, which can be integrated into a computer device. Among them, the computer device can include at least one of a terminal, a server, and the like.

[0088] Among them, the terminal can be a smart phone, a tablet computer, a laptop computer, a personal computer (PC), a smart home, a wearable electronic device, a VR / AR device, an in-vehicle computer, and so on. The server can be an interoperability server between multiple heterogeneous systems or a background server, and can also be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms, and so on.

[0089] In one embodiment, as Figure 1As described above, the facial image processing device can be integrated into computer devices such as terminals or servers to implement the facial image processing method proposed in the embodiments of the present application. Specifically, the computer device can obtain the facial image of the source face and the facial template image of the template face; perform three-dimensional facial modeling on the facial image and the facial template image to obtain the three-dimensional facial image features of the facial image and the three-dimensional facial template image features of the facial template image; perform fusion processing on the three-dimensional facial image features and the three-dimensional facial template image features to obtain the fused three-dimensional facial image features; perform facial replacement feature extraction processing on the facial image based on the facial template image to obtain the initial facial replacement features; perform conversion processing on the initial facial replacement features based on the fused three-dimensional facial image features to obtain the target facial replacement features; use the trained facial image processing model to replace the template face in the facial template image with the source face based on the target facial replacement features and the facial features of the facial image, and obtain the replaced facial image.

[0090] Correspondingly, the training device of the facial image processing model can be integrated into computer devices such as terminals or servers to implement the training method of the facial image processing model proposed in the embodiments of the present application. Specifically, the computer device can obtain a training image sample group, which includes facial image samples, facial template image samples, and facial reference image samples; use a preset facial image processing model to replace the template face in the facial template image sample with the source face in the facial image sample to obtain a predicted facial image; perform three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image, and perform three-dimensional facial contour point detection on the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample; calculate the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain the facial contour loss information between the predicted facial image and the facial reference image sample; adjust the preset facial image processing model based on the facial contour loss information to obtain the trained facial image processing model.

[0091] Among them, the process of facial image processing can be regarded as replacing the object in the facial template image with the source. It can be understood as performing face swapping on the face object in the facial template image. Taking the face as an example of a human face, face swapping means replacing the face identity in the facial template image with the person in the source image, while keeping elements such as the pose, expression, makeup, and background of the face in the facial template image unchanged. The facial image processing method proposed in the embodiments of the present application can generally be applied in scenarios such as ID photo production, film and television portrait production, game character design, virtual avatars, and privacy protection.

[0092] Among them, the training process of the facial image processing model can be regarded as inputting multiple groups of training image samples into the facial image processing model, so that the facial image processing model can continuously learn from multiple groups of training image samples, continuously summarize rules, and finally accurately replace the template face in the facial template image with the source face of the facial image.

[0093] It should be noted that the image processing method and the training method of the facial image processing model provided in the embodiments of the present application involve computer vision technology in the field of artificial intelligence. That is, in the embodiments of the present application, the object in the facial template image can be replaced with the source object of the facial image by using the computer vision technology of artificial intelligence to obtain the replaced facial image.

[0094] The so-called artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, and is a theory, method, technology and application system that can perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0095] Among them, computer vision technology (CV): Computer vision is a science that studies how to make machines "see". Further speaking, it refers to using cameras and computers to replace human eyes to identify and measure targets, etc., which is machine vision, and further perform graphic processing to make the computer process into an image that is more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0096] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0097] The embodiments of the present application will be described from the perspective of a facial image processing device, which can be integrated in a computer device, and the computer device can be a server or a terminal device, etc.

[0098] As Figure 2 described, a facial image processing method is provided, and the specific process includes:

[0099] 101. Obtain the facial image of the source face and the facial template image of the template face.

[0100] Among them, the facial image includes a source object. The so-called source object can be an object contained in the facial image. Taking the facial image as a human face image as an example, the source object can be the person corresponding to the human face image. The source face is the source of the facial object for providing facial object replacement. Corresponding to it is the template face, which contains other elements such as the facial object to be replaced and the facial background to be maintained.

[0101] For example, as Figure 3 shown, Figure 3 001 in Figure 3 can be a facial image, Figure 3 002 in

[0102] can be a facial template image, and 003 in the figure can be the facial image after replacement. It can be seen from

[0103] (1) Directly obtain the facial image and the facial template image

[0104] For example, the original facial image uploaded by the user and the image processing information corresponding to the original facial image can be directly received, and according to the image processing information, the facial image of the source face and the facial template image of the template face are screened out from the original facial image, or a pair of facial images is obtained from an image database or the network, and any one of the pair of facial images is randomly selected as the facial image of the source face, and the other facial image in the pair of facial images is used as the facial template image.

[0105] (2) Indirectly obtain the facial image and the facial template image

[0106] For example, an image processing request sent by a terminal can be received. The storage address of the original facial image and image processing information are carried in the image processing request. According to the storage address, the original facial image is obtained from the memory, cache or third-party database. According to the image processing information, the facial image of the source face and the facial template image of the template face are screened out from the original facial image.

[0107] Optionally, after successfully obtaining the original facial image, a prompt message can also be sent to the terminal to prompt the terminal that the original facial image has been successfully obtained.

[0108] Optionally, after successfully obtaining the original facial image, preprocessing can also be performed on the original facial image to obtain the facial image and the facial template image. There are various ways of preprocessing. For example, the size of the original facial image can be adjusted to a preset size, or the facial object in the original facial image can be aligned to a unified position using facial key point registration.

[0109] 102. Perform three-dimensional facial modeling on the facial image and the facial template image to obtain the three-dimensional facial image features of the facial image and the three-dimensional facial template image features of the facial template image.

[0110] Among them, the three-dimensional facial image features may include information describing the characteristics of the facial image in three-dimensional space. For example, through the three-dimensional facial image features, information such as the characteristics of the facial features of the source face in the facial image, facial contour, texture, angle, and illumination can be known.

[0111] Among them, the three-dimensional facial template image features may include information describing the characteristics of the facial template image in three-dimensional space. For example, through the three-dimensional facial template image, information such as the expression, texture, angle, and illumination of the template face in the facial template image can be known.

[0112] In one embodiment, the three-dimensional facial image features may include multiple parameters, and these parameters together constitute the three-dimensional facial image features. For example, the three-dimensional facial image features may include source face identity parameters, source face expression parameters, source face texture parameters, source face angle parameters, and source face illumination parameters, etc.

[0113] Among them, the source face identity parameters include parameters that can describe the identity of the source face in the three-dimensional facial image. For example, through the identity parameters, the source face of the facial image can be distinguished from the source faces of other facial images, and who the source face of the facial image is can be known through the identity parameters.

[0114] Among them, the source face expression parameters include parameters that can describe the expression of the source face in the facial image.

[0115] Among them, the source facial texture parameters include parameters that can describe the texture of the facial image.

[0116] Among them, the source facial angle parameters include parameters that can describe the angle of the source face in the facial image. For example, through the angle parameters, it can be known whether the source face is looking left or right, or facing straight ahead.

[0117] Among them, the source facial illumination parameters include parameters that can describe the brightness and darkness of the facial image.

[0118] In one embodiment, the three-dimensional facial template image features may also include multiple parameters, and these parameters together constitute the three-dimensional facial template image features. For example, the three-dimensional facial template image features may also include template facial identity parameters, template facial expression parameters, template facial texture parameters, template facial angle parameters, and template facial illumination parameters, and so on.

[0119] In one embodiment, when replacing the template face in the facial template image with the source face of the facial image, we often hope to retain information such as the expression, texture, angle, and illumination of the template face, while replacing the identity in the template face with the identity of the source face. For example, as Figure 3 shown, assume that the person in facial image 001 is Zhang San, and the person in facial template image 002 is Li Si. The facial replacement only replaces Li Si with Zhang San, but the information such as the expression, texture, angle, and illumination of Li Si in facial template image 002 is still retained. Therefore, the template facial expression parameters, template facial texture parameters, template facial angle parameters, and template facial illumination parameters can together constitute the template facial image parameters.

[0120] In one embodiment, the three-dimensional facial image features and the three-dimensional facial template image features can have various forms of expression. For example, the three-dimensional facial image features and the three-dimensional facial template image features can be vectors. Another example is that the three-dimensional facial image features and the three-dimensional facial template image features can be matrices, and so on.

[0121] In one embodiment, various methods can be used to perform three-dimensional facial modeling on the facial image and the facial template image to obtain the three-dimensional facial image features of the facial image and the three-dimensional facial template image features of the facial template image.

[0122] For example, a three-dimensional facial modeling model can be used to perform three-dimensional facial modeling on the facial image and the facial template image respectively, so as to obtain the three-dimensional facial image features of the facial image and the three-dimensional facial template image features of the facial template image.

[0123] Among them, the three-dimensional facial modeling model can include a model that performs three-dimensional modeling on the image and obtains the three-dimensional features of the image from it.

[0124] For example, the 3D facial modeling model may include at least one of Convolutional Neural Networks (CNN), Deep residual network (ResNet), 3D Facial Reconstruction Model (3DMM), etc.

[0125] For another example, a 3D facial reconstruction model can be used to perform regression on the facial image and the facial template image respectively, so as to perform facial modeling on the source face and the template face, and obtain multiple source facial parameters of the facial image and multiple template facial parameters of the template facial image in the 3D facial reconstruction model. Then, multiple source facial parameters can be fused to obtain 3D facial image features, and multiple template facial parameters can be fused to obtain 3D facial template image features.

[0126] 103. Fuse the 3D facial image features and the 3D facial template image features to obtain the fused 3D facial image features.

[0127] In one embodiment, since the replaced facial image has both the features of the facial image and the features of the facial template image, the 3D facial image features and the 3D facial template image features can be fused to obtain the fused 3D facial image features.

[0128] Among them, when replacing the template face in the facial template image with the source face of the facial image, we often hope to retain information such as the expression, texture, angle, and illumination of the template face, while replacing the identity in the template face with the identity of the source face. Therefore, the source facial identity parameters and the template facial image parameters can be fused to obtain the fused 3D facial image features. Specifically, the step "fuse the 3D facial image features and the 3D facial template image features to obtain the fused 3D facial image features" may include:

[0129] Extract the source facial identity parameters corresponding to the facial image from the 3D facial image features;

[0130] Extract the template facial image parameters corresponding to the facial template image from the 3D facial template image features;

[0131] Fuse the source facial identity parameters and the template facial image parameters to obtain the fused 3D facial image features.

[0132] Among them, the source facial identity parameters include parameters that can describe the identity of the source face in the 3D facial image.

[0133] Among them, the template facial image parameters may include template facial expression parameters, template facial texture parameters, template facial angle parameters, and template facial lighting parameters.

[0134] In one embodiment, the source facial identity parameters and the template facial image parameters can be fused in various ways to obtain the fused three-dimensional facial image features.

[0135] For example, the source facial identity parameters and the template facial image parameters can be concatenated. Another example is that the source facial identity parameters and the template facial image parameters can be added. Another example is that the source facial identity parameters and the template facial image parameters can be weighted and summed, and so on.

[0136] In one embodiment, the source facial identity parameters and the template facial image parameters can be added to obtain the fused three-dimensional facial image features.

[0137] For example, the source facial identity parameters, the template facial expression parameters, the template facial texture parameters, the template facial angle parameters, and the template facial lighting parameters can be added to obtain the fused three-dimensional facial image features. Among them, the addition process can be shown as the following formula:

[0138]

[0139] Among them, the symbol can represent the source facial identity parameters, the symbol can represent the template facial expression parameters, the symbol can represent the template facial texture parameters, the symbol can represent the template facial angle parameters, the symbol can represent

[0140] 104. Based on the facial template image, perform facial replacement feature extraction processing on the facial image to obtain the initial facial replacement features.

[0141] In one embodiment, performing three-dimensional facial modeling on the facial image and the facial template image is equivalent to processing the facial image and the facial template image from the perspective of three-dimensional space. And performing facial replacement feature extraction processing on the facial image based on the facial template image is equivalent to processing the facial template image and the facial image from the perspective of two-dimensional space.

[0142] By starting from the perspectives of two-dimensional space and three-dimensional space respectively, more features of the facial image and the facial template image can be obtained, so that when performing facial replacement processing, there are more information bases, thereby improving the accuracy of the facial image processing method.

[0143] Among them, the initial facial replacement feature may include a feature that forms a mapping relationship between the facial image and the facial template image.

[0144] In one embodiment, various methods can be used to perform facial replacement feature extraction processing on the facial image based on the facial template image to obtain the initial facial replacement feature.

[0145] For example, machine learning networks such as Convolutional Neural Networks (CNN) and Generative Adversarial Networks (GAN) can be used to perform facial replacement feature extraction processing on the facial image based on the facial template image to obtain the initial facial replacement feature.

[0146] As one of the methods of deep learning, different from ordinary neural networks, GAN consists of two main networks, one is the generator or the generative network, and the other is the discriminator or the discriminative network. The core logic of GAN is that the generator and the discriminator confront and play against each other.

[0147] Among them, the generator can be a neural network, and its function is to generate content. For example, the generator can generate a picture, a piece of text, a video, etc.

[0148] Among them, the discriminator can also be a neural network, and its function is to discriminate the content input into the discriminator. For example, taking a picture as an example, the goal of the discriminator is to determine whether the picture input into the discriminator is a picture generated by the generator or a real picture.

[0149] Among them, the generator may include a decoder and an encoder. The function of the encoder is to compress real-world data and compress high-dimensional data into low-dimensional data. The function of the decoder is to restore the compressed data to the original data.

[0150] Another example is that a machine learning network composed of multiple convolutional neural networks can be used to perform facial replacement feature extraction processing on the facial image based on the facial template image to obtain the initial facial replacement feature.

[0151] For example, the facial image and the facial replacement image can be input into a machine learning network composed of multiple convolutional neural networks. Then, the machine learning network will continuously reduce the resolution of the facial image and the facial replacement image and encode them into the initial facial replacement feature in the latent space.

[0152] Among them, the latent space includes the space formed by the structure of the machine learning network. For example, if the machine learning network includes an input layer, an output layer, and several convolutional layers in the middle of the input layer and the output layer, then these several convolutional layers can constitute the latent space.

[0153] In one embodiment, the step of "performing facial replacement feature extraction processing on the facial image based on the facial template image to obtain the initial facial replacement feature" may include:

[0154] Performing encoding processing on the facial template image to obtain the first encoding feature of the facial template image;

[0155] Performing encoding processing on the facial image to obtain the second encoding feature of the facial image;

[0156] Adjusting the first encoding feature based on the second encoding feature to obtain the initial facial replacement feature.

[0157] 105. Performing transformation processing on the initial facial replacement feature based on the fused three-dimensional facial image feature to obtain the target facial replacement feature.

[0158] In one embodiment, the fused three-dimensional facial image feature can illustrate the relationship between the facial image and the facial template image in the three-dimensional space, while the initial facial replacement feature can illustrate the relationship between the facial image and the facial template image in the two-dimensional space. To improve the accuracy of facial replacement, the fused three-dimensional facial image feature and the initial facial replacement feature can be mapped to the same space. Specifically, performing transformation processing on the initial facial replacement feature based on the fused three-dimensional facial image feature to obtain the target facial replacement feature.

[0159] In one embodiment, various methods can be used to perform transformation processing on the initial facial replacement feature based on the fused three-dimensional facial image feature to obtain the target facial replacement feature.

[0160] In one embodiment, the norm method can be used to perform transformation processing on the initial facial replacement feature based on the fused three-dimensional facial image feature to obtain the target facial replacement feature. For example, the L1 norm or the L2 norm can be used to perform transformation processing on the initial facial replacement feature based on the fused three-dimensional facial image feature to obtain the target facial replacement feature.

[0161] In one embodiment, the step of "performing transformation processing on the initial facial replacement feature based on the fused three-dimensional facial image feature to obtain the target facial replacement feature" may include:

[0162] Performing statistical processing on the fused three-dimensional facial image feature to obtain the statistically processed three-dimensional facial image feature, and performing statistical processing on the initial facial replacement feature to obtain the statistically processed facial replacement feature;

[0163] Perform a logical operation on the initial facial replacement feature and the statistical post-facial replacement feature to obtain the post-operation facial replacement feature;

[0164] Perform a logical operation on the post-operation facial replacement feature and the statistical post-three-dimensional facial image feature to obtain the target facial replacement feature.

[0165] Among them, the statistical processing includes the way of processing data using methods of mathematical statistics. For example, the statistical processing may include calculating the mean of the data, calculating the variance of the data, calculating the standard deviation of the data, or calculating the covariance of the data, and so on.

[0166] Among them, the logical operation processing method may include the way of processing data using basic operations such as addition, subtraction, multiplication, and division. For example, the logical operation processing may include dividing the data. Another example is that the logical operation processing may include subtracting the data, and so on.

[0167] In one embodiment, Adaptive Instance Normalization (AdaIN) can also be used to perform a conversion process on the initial facial replacement feature based on the fused post-three-dimensional facial image feature to obtain the target facial replacement feature.

[0168] Among them, AdaIN is a method that can align the mean and variance of the fused post-three-dimensional facial image feature to the mean and variance of the initial facial replacement feature, thereby realizing the feature conversion of the image. For example, AdaIN can align the mean and variance of the fused post-three-dimensional facial image feature to the mean and variance of the initial facial replacement feature, thereby obtaining the replacement of the feature.

[0169] Among them, using AdaIN to perform a conversion process on the initial facial replacement feature based on the fused post-three-dimensional facial image feature to obtain the target facial replacement feature can be described as follows:

[0170]

[0171] Among them, x can represent the initial facial replacement feature, y can represent the fused post-three-dimensional facial image feature. σ() and μ() can represent the mean and standard deviation respectively. AdaIN(x, y) can represent the target facial replacement feature.

[0172] In one embodiment, statistical processing can be performed on the fused post-three-dimensional facial image feature according to AdaIN to obtain the statistical post-three-dimensional facial image feature, and statistical processing can be performed on the initial facial replacement feature to obtain the statistical post-facial replacement feature. For example, the mean and standard deviation of the fused post-three-dimensional facial image feature can be calculated according to AdaIN, and the mean and standard deviation of the initial facial replacement feature can be calculated.

[0173] Then, according to AdaIN, logical operation processing can be performed on the initial facial replacement feature and the statistical posterior facial replacement feature to obtain the operation posterior facial replacement feature. For example, the standard deviation of the initial facial replacement feature can be subtracted from the initial facial replacement feature and then divided to obtain the operation posterior facial replacement feature.

[0174] Next, according to AdaIN, logical operation processing can be performed on the operation posterior facial replacement feature and the statistical posterior three-dimensional facial image feature to obtain the target facial replacement feature. For example, the mean value of the operation posterior panel replacement feature and the statistical posterior three-dimensional facial image feature can be multiplied and then added to obtain the target facial replacement feature.

[0175] 106. Using the trained facial image processing model based on the target facial replacement feature and the facial feature of the facial image, replace the template face in the facial template image with the source face to obtain the replaced facial image.

[0176] In one embodiment, in order to improve the similarity between the face in the replaced facial image and the original face, and thus improve the accuracy of the facial image processing method, feature extraction can be performed on the facial image to obtain the facial feature of the facial image. Then, based on the target facial replacement feature and the facial feature of the facial image, replace the template face in the facial template image with the source face to obtain the replaced facial image.

[0177] Among them, performing feature extraction on the facial image may include representing facial information by some numbers, and these data may refer to the facial features of the facial image.

[0178] In one embodiment, the facial feature in the embodiment of the present application may be the geometric feature of the source face or the characterization feature of the source face.

[0179] Among them, the geometric feature refers to the geometric relationship between facial features such as eyes, nose and mouth, such as distance, area and angle.

[0180] Among them, the characterization feature uses the gray information of the facial image to extract global or local features through some algorithms.

[0181] In one embodiment, multiple methods can be used to perform feature extraction on the facial image to obtain the facial feature of the facial image.

[0182] For example, when the facial feature is the geometric feature of the source face, point generation processing can be performed on the facial image to obtain the position information of the facial part feature points in the facial image; then, the spatial difference information between the position information of the facial part feature points can be calculated respectively, and the spatial difference information can be used as the facial feature.

[0183] For another example, when the facial feature is a representative feature of the source face, a convolution kernel can be used to perform convolution extraction on the facial image to obtain the convolution information of the facial image. Then, the convolution information of the facial image is subjected to regression processing to obtain the facial feature.

[0184] For another example, a machine learning network capable of extracting facial features can also be used to extract features from the facial image. For example, a CNN, Deep Neural Networks (DNN), etc. can be used to extract features from the facial image to obtain the facial features of the facial image.

[0185] In one embodiment, based on the target face replacement feature and the facial feature of the facial image, the template face in the facial template image can be replaced with the source face to obtain the replaced facial image.

[0186] Among them, there are various ways to replace the template face in the facial template image with the source face based on the target face replacement feature and the facial feature of the facial image to obtain the replaced facial image.

[0187] For example, using GAN, based on the target face replacement feature and the facial feature of the facial image, the template face in the facial template image is replaced with the source face to obtain the replaced facial image.

[0188] For another example, a trained facial image processing model can be used to replace the template face in the facial template image with the source face based on the target face replacement feature and the facial feature of the facial image to obtain the replaced facial image.

[0189] Among them, the trained facial image processing model is proposed in the embodiments of the present application and can implement the facial image processing method proposed in the embodiments of the present application.

[0190] In one embodiment, the model structure of the trained image processing model may include a three-dimensional facial modeling network, a facial feature extraction network, and an adversarial generation network. Among them, the adversarial generation network may include a generator and a discriminator. Among them, the generator may include an encoder and a decoder.

[0191] Among them, the three-dimensional facial modeling network may be a ResNet network. The three-dimensional facial modeling network can be used to perform three-dimensional facial modeling on the facial image and the facial template image to obtain the three-dimensional facial image features of the facial image and the three-dimensional facial template image features of the facial template image. In addition, the three-dimensional facial modeling network can also be used to perform fusion processing on the three-dimensional facial image features and the three-dimensional facial template image features to obtain the fused three-dimensional facial image features.

[0192] Among them, the facial feature extraction network can be a CNN network. The facial feature extraction network can be used to extract features from a facial image to obtain facial features.

[0193] Among them, the encoder can perform facial replacement feature extraction processing on the facial image based on the facial template image to obtain an initial facial replacement feature.

[0194] Among them, the decoder can replace the template face in the facial template image with the source face based on the target facial replacement feature and the facial features of the facial image to obtain the replaced facial image. For example, the decoder can decode the target facial replacement feature and the facial features of the facial image to obtain the decoded facial replacement feature and the decoded facial features. Then, the decoded panel replacement feature and the decoded panel feature are fused to obtain the fused facial feature. Next, the fused facial feature can be mapped into a preset probability distribution space to obtain the probability distribution of the fused facial feature. And the replaced facial image is generated based on the probability distribution of the fused facial feature.

[0195] Among them, the preset probability distribution space is a mathematical space in the decoder. This preset probability distribution space is a space that is continuously formed during the training of the facial image processing model and can generate content that conforms to the feature generation and training purposes.

[0196] In the example of this application, three-dimensional facial contour points are used to train the facial image processing model, so that the replaced facial image obtained through the trained facial image processing model can maintain the facial contour of the source face in the facial image, making the effect of facial image processing more realistic and improving the accuracy of facial image processing.

[0197] For example, as Figure 4 described. Among them, Figure 4 011 in Figure 4 can be a facial image, Figure 4 012 in Figure 4 can be a facial template image, Figure 4 013 in

[0198] An embodiment of the present application proposes a facial image processing method, which includes: obtaining a facial image of a source face and a facial template image of a template face; performing three-dimensional facial modeling on the facial image and the facial template image to obtain three-dimensional facial image features of the facial image and three-dimensional facial template image features of the facial template image; performing fusion processing on the three-dimensional facial image features and the three-dimensional facial template image features to obtain fused three-dimensional facial image features; based on the facial template image, performing facial replacement feature extraction processing on the facial image to obtain initial facial replacement features; performing conversion processing on the initial facial replacement features based on the fused three-dimensional facial image features to obtain target facial replacement features; using a trained facial image processing model to replace the template face in the facial template image with the source face based on the target facial replacement features and the facial features of the facial image, to obtain a replaced facial image. By starting from the perspectives of the two-dimensional space and the three-dimensional space respectively, the embodiment of the present application can obtain more features of the facial image and the facial template image, so that when performing facial replacement processing, there are more information bases, thereby improving the accuracy of the facial image processing method.

[0199] In one embodiment, in order to better implement the facial image processing method proposed in the embodiment of the present application, the embodiment of the present application also correspondingly proposes a training method for a facial image processing model.

[0200] Next, the embodiment of the present application will be described from the perspective of a training device for a facial image processing model. The training device for the facial image processing model can be integrated in a computer device, and the computer device can be a server or a terminal device, etc.

[0201] As Figure 5 described, a training method for a facial image processing model is provided, and the specific process includes:

[0202] 201. Obtain a training image sample group, which includes facial image samples, facial template image samples, and facial reference image samples.

[0203] Among them, the training image sample group includes data used for training a preset facial image processing model.

[0204] In one embodiment, the training image sample group includes facial image samples, facial template image samples, and facial reference image samples.

[0205] Among them, the facial image samples can correspond to facial images. The facial template image samples can correspond to facial template images. Among them, the facial image samples also include the source face, and the facial template image samples include the template face. For example, as Figure 6 shown, Figure 6The 015 therein can be a facial image sample, and the 016 can be a facial template image sample.

[0206] The facial reference image sample includes a reference image synthesized using the facial image and the facial template image sample. This facial reference image sample not only has the information of the source face in the facial image but also has the image information such as the texture, angle, lighting, and expression of the template face in the facial template image sample. This facial reference image sample can be equivalent to replacing the subsequent facial image, and the difference is that the facial reference image sample is artificially synthesized. In addition, the facial reference image sample is equivalent to the training purpose of the developer, and its role is to play a reference role for the image processing model during the training process of the image processing model, so that the trained image processing model can generate content that meets the requirements.

[0207] In one embodiment, there can be various ways to obtain the training image sample group. For example, the training image sample group can be directly obtained from an open-source website. Another example is that by collecting facial image samples and facial template image samples and synthesizing facial reference image samples using the facial image samples and the facial template image samples, the training image sample group can be obtained.

[0208] 202. Use a preset facial image processing model to replace the template face in the facial template image sample with the source face in the facial image sample to obtain a predicted facial image.

[0209] Among them, the predicted facial image can include the image obtained after replacing the template face in the facial template image sample with the source face in the facial image sample. For example, as Figure 6 shown, Figure 6 the 017 therein can be a predicted facial image.

[0210] Among them, the preset facial image processing model includes the facial image processing model to be trained.

[0211] Among them, the structure of the preset facial image processing model is the same as that of the trained facial image processing model in step 106, but the predicted facial image generated by the preset facial image processing model does not yet meet the training purpose.

[0212] In one embodiment, in the process of using a preset facial image processing model to replace the template face in the facial template image sample with the source face in the facial image sample to obtain a predicted facial image, it can be specifically as follows:

[0213] Use a preset facial image processing model to perform three-dimensional facial modeling on the facial image sample and the facial template image sample to obtain the three-dimensional facial image sample features of the facial image sample and the three-dimensional facial template image sample features of the facial template image sample;

[0214] Use a preset facial image processing model to fuse the three-dimensional facial image sample features and the three-dimensional facial template image sample features to obtain the fused three-dimensional facial image sample features;

[0215] Use a preset facial image processing model to perform facial replacement feature extraction processing on the facial image sample based on the facial template image sample to obtain the initial facial replacement sample features;

[0216] Use a preset facial image processing model to perform conversion processing on the initial facial replacement sample features based on the fused three-dimensional facial image sample features to obtain the target facial replacement sample features;

[0217] Use a preset facial image processing model to replace the template face in the facial template image sample with the source face of the facial image sample based on the target facial replacement sample features and the facial features of the facial image sample to obtain a predicted facial image.

[0218] For example, as Figure 6 shown, three-dimensional facial modeling can be performed on the facial template image sample and the facial image sample to obtain the three-dimensional facial image sample features of the facial image sample and the three-dimensional facial template image sample features of the facial template image sample. Then, the preset facial image processing model can fuse the three-dimensional facial image sample features and the three-dimensional facial template image sample features to obtain the fused three-dimensional facial image sample features. In addition, the preset facial image processing model performs facial replacement feature extraction processing on the facial image sample based on the facial template image sample to obtain the initial facial replacement sample features. Then, the preset facial image processing model can use AdaIN to inject the fused three-dimensional facial image sample features into the initial facial replacement sample features. Among them, Figure 6 the three-dimensional features can include the fused three-dimensional facial image sample features.

[0219] Among them, more detailed steps can refer to Step 101 to Step 106, which will not be repeated here.

[0220] 203. Detect the three-dimensional facial contour points of the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image, and detect the three-dimensional facial contour points of the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample.

[0221] Among them, the three-dimensional facial contour points include information points that describe the facial contour in the image from three-dimensional space. For example, through the three-dimensional facial contour points of the predicted facial image, the contour information of the predicted face in the predicted facial image can be known. For another example, through the three-dimensional facial contour points of the facial reference image sample, the contour information of the face in the facial reference image sample can be known.

[0222] In one embodiment, in order to improve the accuracy of facial image processing, in the embodiments of the present application, a preset facial image processing model can be trained using three-dimensional facial contour points. Therefore, three-dimensional facial contour point detection can be performed on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image, and three-dimensional facial contour point detection can be performed on the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample.

[0223] In one embodiment, various methods can be used to perform three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image, and three-dimensional facial contour point detection can be performed on the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample.

[0224] For example, the predicted facial image can be projected into a three-dimensional space, and the three-dimensional facial contour points of the predicted facial image can be searched in the three-dimensional space.

[0225] Again, for example, three-dimensional facial modeling can be performed on the predicted facial image to obtain the three-dimensional predicted facial image features of the predicted facial image; three-dimensional key point projection can be performed on the three-dimensional predicted facial image features to obtain the three-dimensional facial key points of the predicted facial image; based on the three-dimensional facial key points, the three-dimensional facial contour points can be selected from the three-dimensional facial key points. Specifically, the step of "performing three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image" may include:

[0226] Performing three-dimensional facial modeling on the predicted facial image to obtain the three-dimensional predicted facial image features of the predicted facial image;

[0227] Performing three-dimensional key point projection on the three-dimensional predicted facial image features to obtain the three-dimensional facial key points of the predicted facial image;

[0228] Based on the three-dimensional facial key points, the three-dimensional facial contour points are selected from the three-dimensional facial key points.

[0229] Among them, the three-dimensional predicted facial image features may include information describing the characteristics of the predicted facial image in a three-dimensional space. For example, through the three-dimensional predicted facial image features, information such as the characteristics of the predicted facial features, facial contour, texture, angle, and lighting in the predicted facial image can be known.

[0230] Among them, the three-dimensional predicted facial image features include multiple parameters, and these parameters together constitute the three-dimensional predicted facial image features. For example, the three-dimensional predicted facial image features may include predicted facial identity parameters, predicted facial expression parameters, predicted facial texture parameters, predicted facial angle parameters, and predicted facial lighting parameters, etc.

[0231] Among them, the three-dimensional facial key points include all key points with information of the predicted face.

[0232] In one embodiment, there are various ways to perform 3D facial modeling on the predicted facial image to obtain the 3D predicted facial image features of the predicted facial image. Among them, the method of performing 3D modeling on the predicted facial image can refer to step 102, which will not be elaborated here again.

[0233] In one embodiment, as Figure 6 shown, 3D facial modeling can be performed on the predicted facial image as Figure 6 in 018 to obtain the 3D predicted facial image features of the predicted facial image; then, 3D key point projection is performed on the 3D predicted facial image features to obtain the 3D facial key points of the predicted facial image. Then, as Figure 6 in 019, 3D facial contour points are selected from the 3D facial key points based on the 3D facial key points.

[0234] In one embodiment, various methods can be used to perform 3D key point projection on the 3D predicted facial image features to obtain the 3D facial key points of the predicted facial image.

[0235] In one embodiment, a projection function can be used to perform 3D key point projection on the 3D predicted facial image features to obtain the 3D facial key points of the predicted facial image. For example, the 3D predicted facial image features can be projected according to the following formula to obtain the 3D facial key points of the predicted facial image:

[0236] result_3d_points = reconstruction_without_tex(result_3d_feature)

[0237] where result_3d_points can be the 3D facial key points; result_3d_feature can be the 3D predicted facial image features; reconstruction_without_tex() can be the projection function.

[0238] Among them, the projection function can be various types of functions. For example, the projection function can be glOrtho(), glFrustum(), gluPerspective(), etc. in the Open Graphics Library (OpenGL).

[0239] In one embodiment, the 3D facial key points of the predicted facial image can be obtained based on the predicted facial identity parameters and the predicted facial expression parameters. Specifically, the step of "performing 3D key point projection on the 3D predicted facial image features to obtain the 3D facial key points of the predicted facial image" can include:

[0240] Extract the predicted facial identity parameters and predicted facial expression parameters of the predicted facial image from the three-dimensional predicted facial image features;

[0241] Use preset transfer parameters to perform three-dimensional key point projection on the predicted facial identity parameters and predicted facial expression parameters to obtain the three-dimensional facial key points of the predicted facial image.

[0242] Among them, the preset transfer parameters include pre-set parameters that can achieve information transfer.

[0243] In one embodiment, the predicted facial identity parameters and predicted facial expression parameters can be projected into three-dimensional key points according to the following formula to obtain the three-dimensional facial key points of the predicted facial image:

[0244] result_3d_points = idBase * id_coeff + exBase * ex_coeff + meanshape

[0245] Among them, id_coeff can be the predicted facial identity parameter, ex_coeff can be the predicted facial expression parameter, and idBase, exBase, and meanshape can be preset transfer parameters.

[0246] In one embodiment, after obtaining the three-dimensional facial key points, three-dimensional facial contour points can be screened out from the three-dimensional facial key points based on the three-dimensional facial key points.

[0247] For example, the position information of the three-dimensional facial key points can be obtained, and then the three-dimensional facial contour points can be screened out from the three-dimensional facial key points based on the position information of the three-dimensional facial key points. For example, the three-dimensional facial key points with position information at the edge can be determined as the three-dimensional facial contour points.

[0248] For another example, the three-dimensional facial contour points can be screened out according to the output order of the three-dimensional facial key points. In one embodiment, the output order of the three-dimensional facial key points is specified according to the pre-set settings. For example, there are 68 three-dimensional facial key points, and the first 17 of them can be the three-dimensional facial contour points, so the three-dimensional facial key points with the output order in the first 17 positions can be determined as the three-dimensional facial contour points.

[0249] In one embodiment, the step of "performing three-dimensional facial contour point detection on the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample" can refer to the step of "performing three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image", which will not be repeated here.

[0250] 204. Calculate the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain the facial contour loss information between the predicted facial image and the facial reference image sample.

[0251] In one embodiment, after obtaining the three-dimensional facial contour points, the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample can be calculated to obtain the facial contour loss information between the predicted facial image and the facial reference image sample.

[0252] Among them, there can be various ways to calculate the facial contour loss information.

[0253] For example, the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample can be calculated to obtain the facial contour loss information. Another example is to calculate the Yu Xuan similarity between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain the facial contour loss information.

[0254] For example, when calculating the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample, it can be carried out according to the following formula:

[0255] 3d_point_loss = abs(gt_3d_OutlookPoint - result_3d_OutlookPoint)

[0256] Among them, gt_3d_OutlookPoint can be the three-dimensional facial contour points of the facial reference image sample, result_3d_OutlookPoint can be the three-dimensional facial contour points of the predicted facial image, 3d_point_loss can be the facial contour loss information, and abs() can be the absolute value symbol.

[0257] In one embodiment, in order to improve the performance of the facial image processing model after training and enable the facial image processing model after training to generate images that meet the requirements, the loss information can be calculated from other multiple dimensions, and the preset facial image processing model can be adjusted by using the facial contour loss information and the loss information in other dimensions.

[0258] For example, as Figure 6 shown, the facial feature loss information between the facial image sample and the predicted facial image can be calculated, so that in addition to using the facial contour loss information to adjust the preset facial image processing model, other loss information can also be used to adjust the preset facial image processing model.

[0259] Specifically, after calculating the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain the facial contour loss information between the predicted facial image and the facial reference image sample, the following steps may be included:

[0260] Calculate the difference between the facial reference image sample and the predicted facial image except for the three-dimensional facial contour points to obtain the first loss information between the facial reference image sample and the predicted facial image, where the first loss information includes loss information other than the facial contour loss information;

[0261] Calculate the facial feature loss information between the facial image sample and the predicted facial image;

[0262] Fuse the first loss information and the facial feature loss information to obtain the second loss information.

[0263] Among them, the loss information other than the facial contour loss information may include loss information in other dimensions. For example, the loss information in other dimensions may include pixel loss information, feature loss information, discriminative loss information, and so on.

[0264] Among them, the pixel loss information may include the loss information between the facial reference image sample and the predicted facial image at the pixel level.

[0265] Among them, the feature loss information may include the loss information between the facial reference image sample and the predicted facial image at the feature level. For example, the feature loss information may refer to the difference between the face of the facial reference image sample and the predicted face of the predicted facial image.

[0266] In one embodiment, the preset facial image processing model has a discriminator, and the function of the discriminator is to identify whether the image generated by the generator is a real image. Therefore, the discriminative loss information may include the information generated by the discriminator after discriminating the facial reference image sample and the predicted facial image.

[0267] In one embodiment, when the first loss information includes pixel loss information, feature loss information, and discriminative loss information, the step of "calculating the difference between the facial reference image sample and the predicted facial image except for the three-dimensional facial contour points to obtain the first loss information between the facial reference image sample and the predicted facial image" may include:

[0268] Calculate the pixel difference between the facial reference image sample and the predicted facial image to obtain the pixel loss information;

[0269] Calculate the feature difference between the facial reference image sample and the predicted facial image to obtain the feature loss information;

[0270] Calculate the discriminative difference between the facial reference image sample and the predicted facial image to obtain discriminative loss information.

[0271] In one embodiment, when calculating the pixel difference between the facial reference image sample and the predicted facial image, the pixel information of the facial reference image and the predicted facial image can be extracted, and then the difference between the pixel information can be calculated to obtain pixel loss information.

[0272] For example, the values of the facial reference image sample on the color channels can be extracted, and the values of the predicted facial image on the color channels can be extracted, and the absolute value of the difference between the two can be obtained to obtain pixel loss information. For example, it can be as shown in the following formula:

[0273] Resconstruction loss = abs(result - gt_img)

[0274] Where result can be the pixel information of the predicted facial image, gt_img can be the pixel information of the facial reference image sample, and Resconstruction loss can be the pixel loss information.

[0275] In one embodiment, since the preset facial image processing model can perform feature extraction on the image in three-dimensional space and continue feature extraction in two-dimensional space, the feature loss information can include two-dimensional feature loss information and three-dimensional feature loss information. Then the step of "calculating the feature difference between the facial reference image sample and the predicted facial image to obtain feature loss information" can include:

[0276] Calculate the two-dimensional feature difference between the facial reference image sample and the predicted facial image to obtain two-dimensional feature loss information;

[0277] Calculate the three-dimensional feature difference between the facial reference image sample and the predicted facial image to obtain three-dimensional feature loss information;

[0278] Fuse the two-dimensional feature loss information and the three-dimensional feature loss information to obtain feature loss information.

[0279] Among them, the two-dimensional feature difference can include the difference in features between the facial reference image sample and the predicted facial image in two-dimensional space. For example, the two-dimensional feature difference can include the difference in the image features of the facial reference image sample and the predicted facial image.

[0280] Among them, the three-dimensional feature difference may include the difference in features between the facial reference image sample and the predicted facial image in the three-dimensional space. For example, the three-dimensional feature difference may include the difference between the three-dimensional facial reference image sample features of the facial reference image sample and the three-dimensional predicted facial image features of the predicted facial image.

[0281] In one embodiment, when calculating the two-dimensional feature difference between the facial reference image sample and the predicted facial image, feature extraction can be performed on the facial reference image sample and the predicted facial image respectively to obtain the image feature information of the facial reference image sample and the image feature information of the predicted facial image. Then, the difference between the image feature information of the facial reference image sample and the image feature information of the predicted facial image is calculated.

[0282] For example, the Alexnet network can be used to perform feature extraction on the facial reference image sample and the predicted facial image to obtain the image feature information of the facial reference image sample and the image feature information of the predicted facial image.

[0283] Among them, the Alexnet network consists of 5 convolutional layers and 3 fully connected layers. Among them, the 5 convolutional layers can be used to perform feature extraction on the image, and there is an information transfer relationship between each layer. For example, after the first convolutional layer performs feature extraction on the image, the extracted information will be transferred to the second convolutional layer. Then, the second convolutional layer will continue to perform feature extraction on the information extracted by the first convolutional layer and transfer the extracted information to the third convolutional layer. And so on, finally, the fifth convolutional layer will transfer the extracted information to the fully connected layer.

[0284] In one embodiment, various methods can be used to calculate the difference between the image feature information of the facial reference image sample and the image feature information of the predicted facial image.

[0285] For example, the image perception similarity metric (LPIPS) can be used to calculate the difference between the image feature information of the facial reference image sample and the image feature information of the predicted facial image. Another example is that the difference method can be used to calculate the two-dimensional feature loss information. Another example is that the cosine similarity method can be used to calculate the two-dimensional feature loss information, and so on.

[0286] In one embodiment, when using the Alexnet network to perform feature extraction on the image, the difference between the image feature information of the facial reference image sample and the image feature information of the predicted facial image can be calculated in each convolutional layer.

[0287] For example, use the Alexnet network to extract features from the facial reference image samples, obtaining gt_img_feal1, gt_img_feal2, gt_img_feal3, and gt_img_feal4. Among them, gt_img_feal1, gt_img_feal2, gt_img_feal3, and gt_img_feal4 can respectively refer to the feature information of the facial reference image samples output by 4 convolutional layers in the Alexnet network.

[0288] Similarly, the Alexnet network can be used to extract features from the predicted facial image, obtaining result_feal1, result_feal2, result_feal3, and result_feal4. Among them, result_feal1, result_feal2, result_feal3, and result_feal4 can respectively refer to the feature information of the predicted facial image output by 4 convolutional layers in the Alexnet network.

[0289] Then, the two-dimensional feature loss information can be calculated according to the following formula:

[0290] Two_loss =

[0291] abs(result_feal1 - gt_img_feal1) + abs(result_feal2 - gt_img_feal2) +

[0292] abs(result_feal3 - gt_img_feal3) + abs(result_feal4 - gt_img_feal4)

[0293] Among them, Two_loss can be the two-dimensional feature loss information.

[0294] In one embodiment, facial modeling can be performed on the facial reference image samples and the predicted facial image to obtain the three-dimensional facial reference image sample features of the facial reference image samples and the three-dimensional predicted facial image features of the predicted facial image. Then, the difference between the three-dimensional facial reference image sample features and the three-dimensional predicted facial image features can be calculated.

[0295] In one embodiment, various methods can also be used to calculate the difference between the three-dimensional facial reference image sample features and the three-dimensional predicted facial image features. For example, the image perception similarity metric (LPIPS) can be used to calculate the three-dimensional feature loss information. Another example is that the difference calculation method can be used to calculate the three-dimensional feature loss information. Another example is that the cosine similarity method can be used to calculate the three-dimensional feature loss information, and so on.

[0296] In one embodiment, the three-dimensional feature loss information can be calculated according to the following formula:

[0297]

[0298] where, can represent the three-dimensional feature loss information, can represent the three-dimensional facial reference image sample features, can represent the three-dimensional predicted facial image features.

[0299] In one embodiment, after obtaining the two-dimensional feature loss information and the three-dimensional feature loss information, the two-dimensional feature loss information and the three-dimensional feature loss information can be fused to obtain the feature loss information. For example, the two-dimensional feature loss information and the three-dimensional feature loss information can be added to obtain the feature loss information. Another example is that the two-dimensional feature loss information and the three-dimensional feature loss information can be weighted and then summed to obtain the feature loss information.

[0300] In one embodiment, when calculating the discriminant difference between the facial reference image sample and the predicted facial image, the facial reference image sample and the predicted facial image can be subjected to scale transformation, and a discriminator is used to discriminate the scaled image, thereby improving the richness of the discriminant loss information. Specifically, the step of "calculating the discriminant difference between the facial reference image sample and the predicted facial image to obtain the discriminant loss information" can include:

[0301] Perform scale transformation processing on the facial reference image sample and the predicted facial image respectively to obtain at least one scaled facial reference image sample and at least one scaled predicted facial image;

[0302] Perform discriminant processing on at least one scaled facial reference image sample and at least one scaled predicted facial image respectively to obtain a first discriminant parameter of the scaled facial reference image sample and a second discriminant parameter of the scaled predicted facial image;

[0303] Calculate the discriminant loss information based on the first discriminant parameter and the second discriminant parameter.

[0304] Among them, the scale transformation can refer to changing the size of the image. For example, if the size of the image is 256 in length × 256 in width, through scale transformation, the size of the image can be changed to 128 in length × 128 in width.

[0305] For example, if the original size of the facial reference image sample is a, through scale transformation processing, a facial reference image sample with a size of 1 / 2a and a facial reference image sample with a size of 1 / 4a can be obtained.

[0306] Similarly, assuming that the original size of the predicted facial image is b, through scale transformation processing, a predicted facial image with a size of 1 / 2b and a predicted facial image with a size of 1 / 4b can be obtained.

[0307] Next, discriminant processing can be performed on at least one facial reference image sample after scale transformation and at least one predicted facial image after scale transformation respectively to obtain a first discriminant parameter of the facial reference image sample after scale transformation and a second discriminant parameter of the predicted facial image after scale transformation.

[0308] For example, facial reference image samples with an original size of a, facial reference image samples with a size of 1 / 2a, and facial reference image samples with a size of 1 / 4a can be input into the discriminator to obtain discriminant results.

[0309] For instance, after inputting facial reference image samples with an original size of a, facial reference image samples with a size of 1 / 2a, and facial reference image samples with a size of 1 / 4a into the discriminator, the obtained discriminant results are D(gt_img), D(gt_img_1 / 2), and D(gt_img_1 / 4) respectively. Among them, the symbol D() can represent the discriminant result of the discriminator. gt_img can refer to the facial reference image sample with an original size of a, gt_img_1 / 2 can refer to the facial reference image sample with a size of 1 / 2a, and gt_img_1 / 4 can refer to the facial reference image sample with a size of 1 / 4a.

[0310] In one embodiment, the discriminant result is generally represented by a parameter. For example, D() is generally a value between 0 and 1. Among them, when the discriminant result is 1, it indicates that the image passes the discrimination, and when the discriminant result is 0, it indicates that the image fails to pass the discrimination.

[0311] For example, the first discriminant parameter can include D(gt_img), D(gt_img_1 / 2), and D(gt_img_1 / 4).

[0312] Again, for example, predicted facial images with an original size of b, predicted facial images with a size of 1 / 2b, and predicted facial images with a size of 1 / 4b can be input into the discriminator to obtain discriminant results.

[0313] For instance, after inputting predicted facial images with an original size of b, predicted facial images with a size of 1 / 2b, and predicted facial images with a size of 1 / 4b into the discriminator, the obtained discriminant results are D(result), D(result_1 / 2), and D(result_1 / 4) respectively. Among them, result can refer to the predicted facial image with an original size of b, result_1 / 2 can refer to the predicted facial image with a size of 1 / 2a, and result_1 / 4 can refer to the predicted facial image with a size of 1 / 4a.

[0314] Among them, the second discrimination parameter may include the discrimination result of the discriminator. For example, the second discrimination parameter may include D(result), D(result_1 / 2), and D(result_1 / 4).

[0315] In one embodiment, the discrimination loss information may be calculated based on the first discrimination parameter and the second discrimination parameter in various ways. For example, the discrimination loss information may be calculated by taking the difference. For another example, the discrimination loss information may be calculated by using the cosine similarity, and so on.

[0316] In one embodiment, the discrimination loss information may be calculated according to the following method:

[0317] D_loss = 1 / 3 * (-logD(gt_img) - logD(result)

[0318] -logD(gt_img_1 / 2) - logD(result_1 / 2)

[0319] -logD(gt_img_1 / 4) - logD(result_1 / 4))

[0320] Among them, D_loss may be the discrimination loss information.

[0321] In one embodiment, in order to make the identity parameter of the predicted face in the predicted face image as similar as possible to the identity parameter of the source face in the face image sample, the face feature loss information between the face image sample and the predicted face image sample may also be calculated.

[0322] For example, face feature extraction may be performed on the predicted face image and the face image sample to obtain the face features of the predicted face image and the face features of the face image sample. Then, the face feature loss information between the face features of the predicted face image and the face features of the face image sample is calculated.

[0323] In one embodiment, the face feature loss information between the face features of the predicted face image and the face features of the face image sample may be calculated in various ways. For example, the two-dimensional feature loss information may be calculated by taking the difference. For another example, the two-dimensional feature loss information may be calculated by using the cosine similarity, and so on.

[0324] In one embodiment, the two-dimensional feature loss information may be calculated according to the following formula:

[0325]

[0326] Among them, id loss may be the two-dimensional feature loss information, It can be the facial features for predicting a facial image, or the facial features of a facial image sample. cosine similarity It can be the calculation method of cosine similarity, where cosine similarity can be expressed as follows:

[0327]

[0328] where A and B can be vectors, A i can be the components in vector A, and B i can be the components in vector B. i can refer to the i-th component, and n can refer to the total number of components in the vector.

[0329] In one embodiment, after obtaining the first loss information and the facial feature information, the first loss information and the facial feature loss information can be fused to obtain the second loss information. For example, the first loss information and the facial feature information can be added to obtain the second loss information. Or, for another example, the first loss information and the facial feature information can be weighted and summed to obtain the second loss information.

[0330] 205. Adjust the preset facial image processing model based on the facial contour loss information to obtain the trained facial image processing model.

[0331] In one embodiment, the preset facial image processing model can be adjusted based on the facial contour loss information to obtain the trained image processing model.

[0332] For example, as Figure 6 shown, after obtaining the facial contour loss information, the facial contour loss information can be used to constrain the three-dimensional facial contour points of the predicted facial image to be consistent with the three-dimensional facial contour points of the facial reference image sample.

[0333] For example, the model parameters in the preset facial image processing model can be adjusted based on the facial contour loss information to obtain the adjusted facial image processing model. Then, the adjusted facial image processing model is trained using the training image sample group. Through the above repeated operations, when the facial contour loss information is less than a certain degree or meets the requirements, it indicates that the training has achieved the purpose. At this time, a trained image processing model with performance meeting the requirements can also be obtained.

[0334] In one embodiment, the facial contour loss information and the second loss information can also be fused to obtain third loss information, and then the preset image processing model is adjusted using the third loss information to obtain a trained facial image processing model. Specifically, the step of "adjusting the preset facial image processing model based on the facial contour loss information to obtain a trained facial image processing model" may include:

[0335] Obtain the model parameters of the trained facial image processing model;

[0336] Fuse the facial contour information and the second loss information to obtain third loss information;

[0337] Adjust the model parameters using the third loss information to obtain a trained facial image processing model.

[0338] For example, the facial contour information and the second loss information can be added to obtain third loss information. Then, the model parameters of the preset facial image processing model are adjusted using the third loss information to obtain a trained facial image processing model.

[0339] In one embodiment, during the training of the preset facial image processing model, in addition to learning how to perform facial replacement, it will also learn three-dimensional features to predict the three-dimensional features of the image. For example, as Figure 6 shown.

[0340] In the embodiments of the present application, a training image sample group can be obtained. The training image sample group includes facial image samples, facial template image samples, and facial reference image samples; the template face in the facial template image sample is replaced with the source face in the facial image sample using the preset facial image processing model to obtain a predicted facial image; three-dimensional facial contour points of the predicted facial image are detected, and three-dimensional facial contour points of the predicted facial image are obtained, and three-dimensional facial contour points of the facial reference image sample are detected for the facial reference image sample to obtain three-dimensional facial contour points of the facial reference image sample; the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample is calculated to obtain the facial contour loss information between the predicted facial image and the facial reference image sample; the preset facial image processing model is adjusted based on the facial contour loss information to obtain a trained facial image processing model. By training the preset facial image processing model using three-dimensional facial contour points, when the trained facial image processing model replaces the template face in the facial template image with the source face, the replaced facial image can maintain the facial contour of the source face, thereby improving the accuracy of the facial image processing method.

[0341] In addition, in the embodiments of the present application, by calculating loss information in multiple different dimensions and using the loss information in multiple dimensions to adjust a preset facial image processing model, the preset facial image processing model can use the loss information in multiple dimensions to adjust parameters in different dimensions, so that the trained facial image processing model has better performance.

[0342] According to the method described in the above embodiments, the following will give further detailed examples.

[0343] The embodiments of the present application will take the training of a facial image processing model integrated on a computer device as an example to introduce the method of the embodiments of the present application.

[0344] In one embodiment, as Figure 7 shown, a method for training a facial image processing model is as follows:

[0345] 301. The computer device obtains a training image sample group, which includes facial image samples, facial template image samples, and facial reference image samples.

[0346] For example, the facial image sample can be represented as source, the facial template image sample can be represented as target, and the facial reference image sample can be represented as gt_img.

[0347] 302. The computer device uses the preset facial image processing model to replace the template face in the facial template image sample with the source face in the facial image sample to obtain a predicted facial image.

[0348] For example, the computer device can input source and target into the encoder in the preset facial image processing model. The encoder will continuously reduce the resolutions of source and target and encode them into initial facial replacement sample features in the latent space.

[0349] In addition, the computer device can use a facial feature extraction network to extract features from source to obtain the facial feature source_id_feature of source.

[0350] In addition, the computer device can also perform three-dimensional facial modeling on source and target to obtain the three-dimensional facial image sample features of the facial image sample and the three-dimensional facial template image sample features of the facial template image sample.

[0351] Then the computer device can use the preset facial image processing model to perform conversion processing on the initial facial replacement sample features based on the fused three-dimensional facial image sample features to obtain target facial replacement sample features.

[0352] Next, the computer device can use a preset facial image processing model to replace the template face in the facial template image sample with the source face of the facial image sample based on the target facial replacement sample features and the facial features of the facial image sample, obtaining a predicted facial image.

[0353] Among them, the predicted facial image can be represented as result.

[0354] 303. The computer device performs three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image, and performs three-dimensional facial contour point detection on the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample.

[0355] For example, the computer device can calculate the three-dimensional predicted facial image features of result (which can be represented as result_3d_feature).

[0356] Then, the computer device can perform three-dimensional key point projection on the three-dimensional predicted facial image features to obtain the three-dimensional facial key points of the predicted facial image. For example, as shown in the following formula:

[0357] result_3d_points = reconstruction_without_tex(result_3d_feature)

[0358] Next, the computer device can screen out the three-dimensional facial contour points from the three-dimensional facial key points based on the three-dimensional facial key points.

[0359] Similarly, the computer device can calculate the three-dimensional facial template image sample features of gt_img (which can be represented as gt_3d_feature).

[0360] Then, the computer device can perform three-dimensional key point projection on the three-dimensional facial template image sample features to obtain the three-dimensional facial key points of the facial reference image sample. For example, as shown in the following formula:

[0361] gt_3d_points = reconstruction_without_tex(gt_3d_feature)

[0362] Among them, gt_3d_points can be the three-dimensional facial key points of the facial reference image sample.

[0363] 304. The computer device calculates the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample, obtaining the facial contour loss information between the predicted facial image and the facial reference image sample.

[0364] For example, the facial contour loss information can be calculated according to the following formula:

[0365] 3d_point_loss = abs(gt_3d_OutlookPoint - result_3d_OutlookPoint)

[0366] In one embodiment, other loss information can also be calculated, and the preset facial image processing model can be adjusted jointly using the facial contour loss information and the other loss information to obtain the trained facial image processing model.

[0367] For example, the facial feature loss information, pixel loss information, feature loss information, discriminant loss information, and facial contour loss information can be added together. Then, the obtained loss information after addition is used to adjust the preset facial image processing model to obtain the trained facial image processing model.

[0368] 305. The computer device adjusts the preset facial image processing model based on the facial contour loss information to obtain the trained facial image processing model.

[0369] In the embodiment of the present application, the computer device can obtain a training image sample group, where the training image sample group includes facial image samples, facial template image samples, and facial reference image samples; the computer device can use the preset facial image processing model to replace the template face in the facial template image sample with the source face in the facial image sample to obtain a predicted facial image; the computer device can perform three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image, and perform three-dimensional facial contour point detection on the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample; the computer device calculates the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain the facial contour loss information between the predicted facial image and the facial reference image sample; the computer device adjusts the preset facial image processing model based on the facial contour loss information to obtain the trained facial image processing model. By training the preset facial image processing model using the three-dimensional facial contour points, when the trained facial image processing model replaces the template face in the facial template image with the source face, the replaced facial image can maintain the facial contour of the source face, thereby improving the accuracy of the facial image processing method.

[0370] The embodiment of the present application will take the training of the facial image processing model integrated on the computer device as an example to introduce the method of the embodiment of the present application.

[0371] In one embodiment, as Figure 8 shown, a facial image processing method has the following specific process:

[0372] 401. The computer device acquires the facial image of the source face and the facial template image of the template face.

[0373] 402. The computer device performs three-dimensional facial modeling on the facial image and the facial template image to obtain the three-dimensional facial image features of the facial image and the three-dimensional facial template image features of the facial template image.

[0374] 403. The computer device performs fusion processing on the three-dimensional facial image features and the three-dimensional facial template image features to obtain the fused three-dimensional facial image features.

[0375] 404. The computer device performs facial replacement feature extraction processing on the facial image based on the facial template image to obtain the initial facial replacement features.

[0376] 405. The computer device performs conversion processing on the initial facial replacement features based on the fused three-dimensional facial image features to obtain the target facial replacement features.

[0377] 406. The computer device uses the trained facial image processing model to replace the template face in the facial template image with the source face based on the target facial replacement features and the facial features of the facial image, and obtains the replaced facial image.

[0378] In the embodiments of the present application, the computer device acquires the facial image of the source face and the facial template image of the template face; the computer device performs three-dimensional facial modeling on the facial image and the facial template image to obtain the three-dimensional facial image features of the facial image and the three-dimensional facial template image features of the facial template image; the computer device performs fusion processing on the three-dimensional facial image features and the three-dimensional facial template image features to obtain the fused three-dimensional facial image features; the computer device performs facial replacement feature extraction processing on the facial image based on the facial template image to obtain the initial facial replacement features; the computer device performs conversion processing on the initial facial replacement features based on the fused three-dimensional facial image features to obtain the target facial replacement features; the computer device uses the trained facial image processing model to replace the template face in the facial template image with the source face based on the target facial replacement features and the facial features of the facial image, and obtains the replaced facial image. By starting from the perspectives of two-dimensional space and three-dimensional space respectively in the embodiments of the present application, more features of the facial image and the facial template image can be obtained, so that there are more information bases when performing facial replacement processing, thereby improving the accuracy of the facial image processing method.

[0379] To better implement the facial image processing method provided by the embodiments of the present application, in one embodiment, a facial image processing device is further provided. This facial image processing device can be integrated into a computer device. The meanings of the nouns are the same as those in the above facial image processing method, and the specific implementation details can be referred to the descriptions in the method embodiments.

[0380] In one embodiment, a facial image processing device is provided. This facial image processing device can be specifically integrated in a computer device, such as Figure 9 shown. The facial image processing device includes: a first acquisition unit 501, a three-dimensional facial modeling unit 502, a first fusion unit 503, a feature extraction unit 504, a conversion unit 505, and a first replacement unit 506, specifically as follows:

[0381] The first acquisition unit 501 is configured to acquire a facial image of a source face and a facial template image of a template face;

[0382] The three-dimensional facial modeling unit 502 is configured to perform three-dimensional facial modeling on the facial image and the facial template image to obtain three-dimensional facial image features of the facial image and three-dimensional facial template image features of the facial template image;

[0383] The first fusion unit 503 is configured to perform fusion processing on the three-dimensional facial image features and the three-dimensional facial template image features to obtain fused three-dimensional facial image features;

[0384] The feature extraction unit 504 is configured to perform facial replacement feature extraction processing on the facial image based on the facial template image to obtain initial facial replacement features;

[0385] The conversion unit 505 is configured to perform conversion processing on the initial facial replacement features based on the fused three-dimensional facial image features to obtain target facial replacement features;

[0386] The first replacement unit 506 is configured to use the trained facial image processing model to replace the template face in the facial template image with the source face based on the target facial replacement features and the facial features of the facial image to obtain a replaced facial image.

[0387] In one embodiment, the first fusion unit 503 includes:

[0388] The first extraction subunit is configured to extract the source face identity parameters corresponding to the facial image from the three-dimensional facial image features;

[0389] The second extraction subunit is configured to extract the template facial image parameters corresponding to the facial template image from the three-dimensional facial template image features;

[0390] The first fusion subunit is configured to fuse the source facial identity parameters and the template facial image parameters to obtain the fused three-dimensional facial image features.

[0391] In one embodiment, the feature extraction unit 504 includes:

[0392] The first encoding subunit is configured to perform encoding processing on the facial template image to obtain the first encoded feature of the facial template image;

[0393] The second encoding subunit is configured to perform encoding processing on the facial image to obtain the second encoded feature of the facial image;

[0394] The first adjustment subunit is configured to adjust the first encoded feature based on the second encoded feature to obtain the initial facial replacement feature.

[0395] In one embodiment, the conversion unit 505 includes:

[0396] The first statistical subunit is configured to perform statistical processing on the fused three-dimensional facial image features to obtain the statistically processed three-dimensional facial image features, and perform statistical processing on the initial facial replacement feature to obtain the statistically processed facial replacement feature;

[0397] The second statistical subunit is configured to perform logical operation processing on the initial facial replacement feature and the statistically processed facial replacement feature to obtain the operationally processed facial replacement feature;

[0398] The logical operation processing subunit is configured to perform logical operation processing on the operationally processed facial replacement feature and the statistically processed three-dimensional facial image features to obtain the target facial replacement feature.

[0399] In specific implementation, each of the above units can be implemented as an independent entity, or can be combined arbitrarily and implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the foregoing method embodiments, which will not be elaborated herein.

[0400] Through the above-mentioned facial image processing device, the accuracy of replacing the facial image can be improved.

[0401] In addition, in one embodiment, a training device for a facial image processing model is further provided. The training device for the facial image processing model can be integrated into a computer device. The meanings of the nouns are the same as those in the foregoing training method for the facial image processing model, and the specific implementation details can refer to the description in the method embodiments.

[0402] In one embodiment, a training device for a facial image processing model is provided. The training device for the facial image processing model can be specifically integrated in a computer device, such as Figure 10 As shown, the training device for the facial image processing model includes: a second acquisition unit 601, a second replacement unit 602, a three-dimensional facial contour point detection unit 603, a calculation unit 604, and an adjustment unit 605, specifically as follows:

[0403] The second acquisition unit 601 is configured to acquire a training image sample group, where the training image sample group includes facial image samples, facial template image samples, and facial reference image samples;

[0404] The second replacement unit 602 is configured to use a preset facial image processing model to replace the template face in the facial template image sample with the source face in the facial image sample to obtain a predicted facial image;

[0405] The three-dimensional facial contour point detection unit 603 is configured to perform three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image, and perform three-dimensional facial contour point detection on the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample;

[0406] The calculation unit 604 is configured to calculate the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain the facial contour loss information between the predicted facial image and the facial reference image sample;

[0407] The adjustment unit 605 is configured to adjust the preset facial image processing model based on the facial contour loss information to obtain a trained facial image processing model.

[0408] In one embodiment, the three-dimensional facial contour point detection unit 603 includes:

[0409] A three-dimensional facial modeling subunit, configured to perform three-dimensional facial modeling on the predicted facial image to obtain three-dimensional predicted facial image features of the predicted facial image;

[0410] A three-dimensional key point projection subunit, configured to perform three-dimensional key point projection on the three-dimensional predicted facial image features to obtain the three-dimensional facial key points of the predicted facial image;

[0411] A screening subunit, configured to screen out the three-dimensional facial contour points from the three-dimensional facial key points based on the three-dimensional facial key points.

[0412] In one embodiment, the three-dimensional key point projection subunit includes:

[0413] An extraction module, configured to extract the predicted facial identity parameter and the predicted facial expression parameter of the predicted facial image from the three-dimensional predicted facial image features;

[0414] A three-dimensional key point projection module, configured to perform three-dimensional key point projection on the predicted facial identity parameter and the predicted facial expression parameter by using a preset transfer parameter, so as to obtain three-dimensional facial key points of the predicted facial image.

[0415] In one embodiment, the training device of the facial image processing model further includes:

[0416] A first calculation unit, configured to calculate a difference between the facial reference image sample and the predicted facial image except for three-dimensional facial contour points, so as to obtain first loss information between the facial reference image sample and the predicted facial image, where the first loss information includes loss information other than facial contour loss information;

[0417] A second calculation unit, configured to calculate facial feature loss information between the facial image sample and the predicted facial image;

[0418] A second fusion unit, configured to perform fusion processing on the first loss information and the facial feature loss information to obtain second loss information.

[0419] In one embodiment, the first calculation unit includes:

[0420] A first calculation subunit, configured to calculate a pixel difference between the facial reference image sample and the predicted facial image to obtain pixel loss information;

[0421] A second calculation subunit, configured to calculate a feature difference between the facial reference image sample and the predicted facial image to obtain feature loss information;

[0422] A third calculation subunit, configured to calculate a discrimination difference between the facial reference image sample and the predicted facial image to obtain discrimination loss information.

[0423] In one embodiment, the second calculation subunit includes:

[0424] A first calculation module, configured to calculate a two-dimensional feature difference between the facial reference image sample and the predicted facial image to obtain two-dimensional feature loss information;

[0425] A second calculation module, configured to calculate a three-dimensional feature difference between the facial reference image sample and the predicted facial image to obtain three-dimensional feature loss information;

[0426] A first fusion module, configured to fuse the two-dimensional feature loss information and the three-dimensional feature loss information to obtain the feature loss information.

[0427] In one embodiment, the third calculation subunit includes:

[0428] A scale transformation module, configured to perform scale transformation processing on the facial reference image sample and the predicted facial image respectively, to obtain at least one facial reference image sample after scale transformation and at least one predicted facial image after scale transformation;

[0429] A discrimination module, configured to perform discrimination processing on the at least one facial reference image sample after scale transformation and the at least one predicted facial image after scale transformation respectively, to obtain a first discrimination parameter of the facial reference image sample after scale transformation and a second discrimination parameter of the predicted facial image after scale transformation;

[0430] A third calculation module, configured to calculate the discrimination loss information based on the first discrimination parameter and the second discrimination parameter.

[0431] In one embodiment, the adjustment unit 605 includes:

[0432] An acquisition subunit, configured to acquire the model parameters of the trained facial image processing model;

[0433] A second fusion subunit, configured to fuse the facial contour information and the second loss information to obtain third loss information;

[0434] A parameter adjustment unit, configured to adjust the model parameters by using the third loss information to obtain the trained facial image processing model.

[0435] An embodiment of the present application further provides a computer device, which may include a terminal or a server. For example, the terminal may be a mobile phone, a tablet computer, etc.; or the computer device may be a server, etc. As Figure 11 shown, it shows a schematic structural diagram of the terminal involved in the embodiment of the present application. Specifically:

[0436] The computer device may include a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, a power supply 703, an input unit 704 and other components. Those skilled in the art can understand that Figure 11 the structural diagram of the computer device shown in does not constitute a limitation on the computer device, and may include more or fewer components than shown, or combine certain components, or arrange different components. Among them:

[0437] The processor 701 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 702, and by invoking the data stored in the memory 702, it executes various functions of the computer device and processes data, thereby performing an overall detection of the computer device. Optionally, the processor 701 may include one or more processing cores; preferably, the processor 701 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 701 either.

[0438] The memory 702 can be used to store software programs and modules. The processor 701 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store data created according to the use of the computer device. In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702.

[0439] The computer device also includes a power supply 703 that powers each component. Preferably, the power supply 703 can be logically connected to the processor 701 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 703 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or an inverter, and a power status indicator.

[0440] The computer device may also include an input unit 704, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0441] Although not shown, the computer device may also include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 701 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 702 according to the following instructions, and the processor 701 will run the application programs stored in the memory 702 to achieve various functions as follows:

[0442] Obtain a facial image of a source face and a facial template image of a template face;

[0443] Perform three-dimensional facial modeling on the facial image and the facial template image to obtain three-dimensional facial image features of the facial image and three-dimensional facial template image features of the facial template image;

[0444] Fuse the three-dimensional facial image features and the three-dimensional facial template image features to obtain fused three-dimensional facial image features;

[0445] Based on the facial template image, perform facial replacement feature extraction processing on the facial image to obtain initial facial replacement features;

[0446] Perform conversion processing on the initial facial replacement features based on the fused three-dimensional facial image features to obtain target facial replacement features;

[0447] Use a trained facial image processing model to replace the template face in the facial template image with the source face based on the target facial replacement features and the facial features of the facial image to obtain a replaced facial image.

[0448] Or

[0449] Obtain a training image sample group, where the training image sample group includes facial image samples, facial template image samples, and facial reference image samples;

[0450] Use a preset facial image processing model to replace the template face in the facial template image sample with the source face in the facial image sample to obtain a predicted facial image;

[0451] Perform three-dimensional facial contour point detection on the predicted facial image to obtain three-dimensional facial contour points of the predicted facial image, and perform three-dimensional facial contour point detection on the facial reference image sample to obtain three-dimensional facial contour points of the facial reference image sample;

[0452] Calculate the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain facial contour loss information between the predicted facial image and the facial reference image sample;

[0453] Adjust the preset facial image processing model based on the facial contour loss information to obtain a trained facial image processing model.

[0454] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.

[0455] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations in the above embodiments.

[0456] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a computer program, or by controlling related hardware through a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0457] Therefore, an embodiment of the present application further provides a storage medium, in which a computer program is stored, and the computer program can be loaded by a processor to execute the steps in any one of the automatic reply processing methods provided by the embodiments of the present application. For example, the computer program can execute the following steps:

[0458] Obtain a facial image of a source face and a facial template image of a template face;

[0459] Perform three-dimensional facial modeling on the facial image and the facial template image to obtain three-dimensional facial image features of the facial image and three-dimensional facial template image features of the facial template image;

[0460] Fuse the three-dimensional facial image features and the three-dimensional facial template image features to obtain fused three-dimensional facial image features;

[0461] Based on the facial template image, perform facial replacement feature extraction processing on the facial image to obtain initial facial replacement features;

[0462] Perform conversion processing on the initial facial replacement features based on the fused three-dimensional facial image features to obtain target facial replacement features;

[0463] Use a trained facial image processing model to replace the template face in the facial template image with the source face based on the target facial replacement features and the facial features of the facial image to obtain a replaced facial image.

[0464] Or

[0465] Obtain a training image sample group, where the training image sample group includes facial image samples, facial template image samples, and facial reference image samples;

[0466] Using a preset facial image processing model, replace the template face in the template facial image sample with the source face in the facial image sample to obtain a predicted facial image;

[0467] Perform three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image, and perform three-dimensional facial contour point detection on the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample;

[0468] Calculate the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain the facial contour loss information between the predicted facial image and the facial reference image sample;

[0469] Adjust the preset facial image processing model based on the facial contour loss information to obtain a trained facial image processing model.

[0470] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0471] Since the computer program stored in this storage medium can execute the steps in any of the facial image processing methods provided in the embodiments of the present application, the beneficial effects achievable by any of the facial image processing methods provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments, which will not be elaborated here.

[0472] The above has introduced in detail a facial image processing method and a facial image processing model provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A training method for a facial image processing model, characterized in that, Including: Obtaining a group of training image samples, where the group of training image samples includes facial image samples, facial template image samples, and facial reference image samples; Using a preset facial image processing model to replace the template face in the facial template image sample with the source face in the facial image sample to obtain a predicted facial image; Performing three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image, and performing three-dimensional facial contour point detection on the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample; Calculating the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain the facial contour loss information between the predicted facial image and the facial reference image sample; Adjusting the preset facial image processing model based on the facial contour loss information to obtain a trained facial image processing model.

2. The training method according to claim 1, wherein The performing three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image includes: Performing three-dimensional facial modeling on the predicted facial image to obtain the three-dimensional predicted facial image features of the predicted facial image; Performing three-dimensional key point projection on the three-dimensional predicted facial image features to obtain the three-dimensional facial key points of the predicted facial image; Based on the three-dimensional facial key points, screening out the three-dimensional facial contour points from the three-dimensional facial key points.

3. The training method according to claim 2, characterized in that The performing three-dimensional key point projection on the three-dimensional predicted facial image features to obtain the three-dimensional facial key points of the predicted facial image includes: Extracting the predicted facial identity parameters and predicted facial expression parameters of the predicted facial image from the three-dimensional predicted facial image features; Performing three-dimensional key point projection on the predicted facial identity parameters and predicted facial expression parameters by using a preset transfer parameter to obtain the three-dimensional facial key points of the predicted facial image.

4. The training method according to claim 1, wherein After the calculating the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain the facial contour loss information between the predicted facial image and the facial reference image sample, it further includes: Calculating the difference between the facial reference image sample and the predicted facial image except for the three-dimensional facial contour points to obtain the first loss information between the facial reference image sample and the predicted facial image, where the first loss information includes loss information other than the facial contour loss information; Calculating the facial feature loss information between the facial image sample and the predicted facial image; Performing fusion processing on the first loss information and the facial feature loss information to obtain second loss information.

5. The training method according to claim 4, wherein The first loss information includes pixel loss information, feature loss information, and discriminant loss information; The calculating the difference between the facial reference image sample and the predicted facial image except for the three-dimensional facial contour points to obtain the first loss information between the facial reference image sample and the predicted facial image includes: Calculate the pixel difference between the facial reference image sample and the predicted facial image to obtain pixel loss information; Calculate the feature difference between the facial reference image sample and the predicted facial image to obtain feature loss information; Calculate the discriminant difference between the facial reference image sample and the predicted facial image to obtain discriminant loss information.

6. The training method according to claim 5, wherein The feature loss information includes three-dimensional feature loss information and two-dimensional feature loss information; The calculating the feature difference between the facial reference image sample and the predicted facial image to obtain feature loss information includes: Calculate the two-dimensional feature difference between the facial reference image sample and the predicted facial image to obtain two-dimensional feature loss information; Calculate the three-dimensional feature difference between the facial reference image sample and the predicted facial image to obtain three-dimensional feature loss information; Fuse the two-dimensional feature loss information and the three-dimensional feature loss information to obtain the feature loss information.

7. The training method according to claim 5, wherein The calculating the discriminant difference between the facial reference image sample and the predicted facial image to obtain discriminant loss information includes: Perform scale transformation processing on the facial reference image sample and the predicted facial image respectively to obtain at least one scaled facial reference image sample and at least one scaled predicted facial image; Perform discriminant processing on the at least one scaled facial reference image sample and the at least one scaled predicted facial image respectively to obtain a first discriminant parameter of the scaled facial reference image sample and a second discriminant parameter of the scaled predicted facial image; Calculate the discriminant loss information based on the first discriminant parameter and the second discriminant parameter.

8. The training method according to claim 4, wherein The adjusting the preset facial image processing model based on the facial contour loss information to obtain a trained facial image processing model includes: Obtain the model parameters of the trained facial image processing model; Fuse the facial contour information and the second loss information to obtain third loss information; Use the third loss information to adjust the model parameters to obtain the trained facial image processing model.

9. A facial image processing method, characterized in that, Includes: Obtain the facial image of the source face and the facial template image of the template face; Perform three-dimensional facial modeling on the facial image and the facial template image to obtain the three-dimensional facial image features of the facial image and the three-dimensional facial template image features of the facial template image; Fuse the three-dimensional facial image features and the three-dimensional facial template image features to obtain fused three-dimensional facial image features; Based on the facial template image, perform facial replacement feature extraction processing on the facial image to obtain initial facial replacement features; Perform conversion processing on the initial facial replacement features based on the fused three-dimensional facial image features to obtain target facial replacement features; Using the trained facial image processing model, based on the target facial replacement feature and the facial feature of the facial image, replace the template face in the facial template image with the source face to obtain a replaced facial image, where the trained facial image processing model is obtained by using the training method of the facial image processing model according to any one of claims 1 to 8.

10. The facial image processing method according to claim 9, wherein The fusing process of the three-dimensional facial image feature and the three-dimensional facial template image feature to obtain a fused three-dimensional facial image feature includes: Extracting the source face identity parameter corresponding to the facial image from the three-dimensional facial image feature; Extracting the template facial image parameter corresponding to the facial template image from the three-dimensional facial template image feature; Fusing the source face identity parameter and the template facial image parameter to obtain the fused three-dimensional facial image feature.

11. The facial image processing method according to claim 9, wherein The process of extracting the facial replacement feature from the facial image based on the facial template image to obtain an initial facial replacement feature includes: Encoding the facial template image to obtain a first encoded feature of the facial template image; Encoding the facial image to obtain a second encoded feature of the facial image; Adjusting the first encoded feature based on the second encoded feature to obtain the initial facial replacement feature.

12. The facial image processing method according to claim 9, wherein, The process of converting the initial facial replacement feature based on the fused three-dimensional facial image feature to obtain a target facial replacement feature includes: Performing statistical processing on the fused three-dimensional facial image feature to obtain a statistically processed three-dimensional facial image feature, and performing statistical processing on the initial facial replacement feature to obtain a statistically processed facial replacement feature; Performing a logical operation on the initial facial replacement feature and the statistically processed facial replacement feature to obtain an operation facial replacement feature; Performing a logical operation on the operation facial replacement feature and the statistically processed three-dimensional facial image feature to obtain the target facial replacement feature.

13. The facial image processing method according to claim 9, wherein The fusing process of the three-dimensional facial image feature and the three-dimensional facial template image feature to obtain a fused three-dimensional facial image feature includes: Using the trained facial image processing model to fuse the three-dimensional facial image feature and the three-dimensional facial template image feature to obtain a fused three-dimensional facial image feature; The process of extracting the facial replacement feature from the facial image based on the facial template image to obtain an initial facial replacement feature includes: Using the trained facial image processing model to extract the facial replacement feature from the facial image based on the facial template image to obtain an initial facial replacement feature; The process of converting the initial facial replacement feature based on the fused three-dimensional facial image feature to obtain a target facial replacement feature includes: Using the trained facial image processing model to convert the initial facial replacement feature based on the fused three-dimensional facial image feature to obtain a target facial replacement feature.

14. A training device for a facial image processing model, characterized in that, Including: A second acquisition unit, configured to acquire a group of training image samples, where the group of training image samples includes facial image samples, facial template image samples, and facial reference image samples; A first replacement unit, configured to use a preset facial image processing model to replace the template face in the facial template image sample with the source face in the facial image sample to obtain a predicted facial image; A three-dimensional facial contour point detection unit, configured to perform three-dimensional facial contour point detection on the predicted facial image to obtain the three-dimensional facial contour points of the predicted facial image, and perform three-dimensional facial contour point detection on the facial reference image sample to obtain the three-dimensional facial contour points of the facial reference image sample; A calculation unit, configured to calculate the difference between the three-dimensional facial contour points of the predicted facial image and the three-dimensional facial contour points of the facial reference image sample to obtain the facial contour loss information between the predicted facial image and the facial reference image sample; An adjustment unit, configured to adjust the preset facial image processing model based on the facial contour loss information to obtain a trained facial image processing model.

15. A facial image processing device, characterized in that, Comprising: A first acquisition unit, configured to acquire a facial image of a source face and a facial template image of a template face; A three-dimensional facial modeling unit, configured to perform three-dimensional facial modeling on the facial image and the facial template image to obtain three-dimensional facial image features of the facial image and three-dimensional facial template image features of the facial template image; A fusion unit, configured to perform a fusion process on the three-dimensional facial image features and the three-dimensional facial template image features to obtain fused three-dimensional facial image features; A feature extraction unit, configured to perform facial replacement feature extraction processing on the facial image based on the facial template image to obtain initial facial replacement features; a conversion unit, configured to perform conversion processing on the initial facial replacement features based on the fused three-dimensional facial image features to obtain target facial replacement features; A first replacement unit, configured to use the trained facial image processing model to replace the template face in the facial template image with the source face based on the target facial replacement features and the facial features of the facial image to obtain a replaced facial image; wherein, the trained facial image processing model is obtained by using the facial image processing model training device as described in claim 14.

16. A computer program product, characterized in that, A computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the method as described in any one of claims 1 to 13.

17. A storage medium, characterized in that, The storage medium stores instructions, and when the instructions are executed by a processor, the method as described in any one of claims 1 to 13 is implemented.

18. A computer device, characterized in that, Comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the method as described in any one of claims 1 to 13 is implemented.

Citation Information

Patent Citations

  • Image processing method, image processing model training method and equipment

    CN111553267A

  • Image fusion method, apparatus, and storage medium

    WO2019233229A1