Face fusion model training method, face fusion method and device
Through a training method of the face fusion model, the dual training process of source face and reference face image is used to solve the problem that face fusion is difficult to decouple identity and attribute information in the prior art, and the face fusion effect with high similarity and high consistency is achieved.
Patent Information
- Application Number
- CN202510253429.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-24
AI Technical Summary
Existing face fusion technology is difficult to effectively decouple the identity information and attribute information of face images, resulting in low similarity between faces and stiff expressions after fusing.
A training method for face fusion model is proposed. By obtaining the source face image and reference face image, using the feature extraction network and decoding network in the initial face fusion model for training, and gradually obtaining the trained face fusion model. The method includes a two-step training process: first, use the source face image to train to obtain the second decoding network to ensure that it is not affected by the identity of the reference face image; then, when the parameters of the second decoding network are fixed, the reference face image is used to train the second feature extraction network to fully decouple the attribute information.
Through this method, the similarity between the fused face and the source face is improved, and the attribute consistency between the fused face and the reference face is enhanced, and the quality of the face fusion image is improved.
Smart Images

Figure CN120198754A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, particularly to technical fields such as computer vision, deep learning, large models, and augmented reality. Specifically, it relates to a method for training a face fusion model, a face fusion method, and a device. Background Art
[0002] AIGC (Artificial Intelligence Generated Content) refers to a technical method based on artificial intelligence. Through the learning and recognition of existing data, it generates relevant content with appropriate generalization ability. AIGC can be widely applied to fields such as text generation and image processing. In the field of image processing, face fusion technology refers to the fusion processing of two or more face images to generate a fused face image. With the continuous development of face fusion technology, the application fields of face fusion technology are becoming more and more extensive. Summary of the Invention
[0003] This application provides a method for training a face fusion model, a face fusion method, and a device. The specific solutions are as follows:
[0004] According to one aspect of this application, there is provided a method for training a face fusion model, including:
[0005] Obtain a source face image and a reference face image;
[0006] Use the source face image to train the first feature extraction network and the first decoding network in the initial face fusion model to obtain a second feature extraction network and a second decoding network;
[0007] Use the reference face image and the second decoding network to train the second feature extraction network to obtain a trained face fusion model.
[0008] According to another aspect of this application, there is provided a face fusion method, including:
[0009] Obtain a reference face image;
[0010] Input the reference face image into the face fusion model to obtain a face fusion image; wherein, the face fusion model is trained by the method described in the embodiment of one aspect.
[0011] According to another aspect of this application, there is provided a device for training a face fusion model, including:
[0012] An acquisition module, configured to obtain a source face image and a reference face image;
[0013] The first training module is configured to train the first feature extraction network and the first decoding network in the initial face fusion model by using the source face image, so as to obtain a second feature extraction network and a second decoding network;
[0014] The second training module is configured to train the second feature extraction network by using the reference face image and the second decoding network, so as to obtain a trained face fusion model.
[0015] According to another aspect of the present application, there is provided a face fusion device, including:
[0016] An acquisition module, configured to acquire a reference face image;
[0017] A face fusion module, configured to input the reference face image into a face fusion model to obtain a face fusion image; wherein, the face fusion model is trained by using the method described in the embodiment of one aspect.
[0018] According to another aspect of the present application, there is provided an electronic device, including:
[0019] At least one processor; and
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the method described in the above embodiment.
[0022] According to another aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the above embodiment.
[0023] According to another aspect of the present application, there is provided a computer program product, including a computer program, and the computer program realizes the steps of the method described in the above embodiment when being executed by a processor.
[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings are used to better understand the solution and do not constitute a limitation to the present application. Among them:
[0026] Figure 1 is a schematic flowchart of a method for training a face fusion model provided by an embodiment of the present application;
[0027] Figure 2 Flow schematic diagram of the training method of the face fusion model provided in another embodiment of the present application;
[0028] Figure 3 Flow schematic diagram of the training method of the face fusion model provided in another embodiment of the present application;
[0029] Figure 4 Structural schematic diagram of the initial face fusion model provided in an embodiment of the present application;
[0030] Figure 5 Flow schematic diagram of the face fusion method provided in an embodiment of the present application;
[0031] Figure 6 Schematic diagram of the training process of the face fusion model provided in an embodiment of the present application;
[0032] Figure 7 Structural schematic diagram of the training device of the face fusion model provided in an embodiment of the present application;
[0033] Figure 8 Structural schematic diagram of the face fusion device provided in an embodiment of the present application;
[0034] Figure 9 It is a block diagram of an electronic device for implementing the training method of the face fusion model in an embodiment of the present application. Detailed implementation manners
[0035] The following describes exemplary embodiments of the present application with reference to the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0036] It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solution of the present application all comply with the relevant regulations of national laws and regulations and do not violate public order and good customs.
[0037] The following describes the training method, face fusion method, device, electronic device, and storage medium of the face fusion model in the embodiments of the present application with reference to the accompanying drawings.
[0038] In some scenarios where face fusion technology is applied, such as in scenarios of film and television works processing, etc., a source face image providing identity information and a reference face image providing reference information can be fused according to certain rules to generate a fused face image. The goal is to make the fused face image have the identity of the source face image while retaining other attribute information of the reference face image except for the identity, such as color, pose, expression, etc.
[0039] In some embodiments, a face recognition network can be trained first and its parameters can be fixed. The source face graph is input into the face recognition network, and the feature vector of the source face image is output. Subsequently, the reference face image is input into an encoder, and the vector of the reference face image in the latent space of the encoder is output. It is fused with the feature vector of the source face image, and then the fused vector is input into a decoder to decode the image after face fusion. By training the encoder and the decoder, the purpose of face fusion is achieved.
[0040] However, there is no direct supervision when training the neural network in this method. Therefore, it is challenging to control the network to accurately decouple the identity information and attribute information of the image. In actual usage scenarios, these two types of information often cannot be well decoupled, which may cause the fused face to have the identity of part of the reference face image or the expression of the source face image, resulting in a relatively low similarity between the fused face and the source face and a rigid expression.
[0041] Based on this, an embodiment of the present application proposes a training method for a face fusion model. Figure 1 It is a schematic flowchart of the training method for the face fusion model provided by an embodiment of the present application.
[0042] The training method for the face fusion model in the embodiment of the present application can be executed by the training device for the face fusion model in the embodiment of the present application, and this device can be configured in an electronic device.
[0043] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc., which are hardware devices with various operating systems, touch screens, and / or display screens.
[0044] As Figure 1 shown, the training method for the face fusion model includes:
[0045] Step 101, obtain a source face image and a reference face image.
[0046] Among them, the source face image can refer to the face image of the source face, and the source face image is used to provide identity information; the reference face image can refer to the face image of the reference face, and the reference face image is used to provide attribute information. For example, the attribute information can include pose, expression, etc.
[0047] In addition, the source face images are face images belonging to the same identity, and the reference face images are also face images belonging to the same identity. For example, the source face image is the face image of object a, and the reference face image can be the face image of object b.
[0048] Exemplarily, the reference face image can be extracted from the reference video that needs face replacement, and the source face image can be extracted from the source video, so as to replace the reference face with the source face.
[0049] For example, for a reference video that needs face replacement, the original uncolored film can be provided, with a resolution not exceeding 2k and a frame rate of 25FPS. At the same time, a video of a person with the source identity is shot for training. The shooting should include multiple scenes and simulate the lighting, color, etc. in the reference video. The person with the source identity can show their forehead and make expressions such as smiling, angry, etc. A total of about 500 minutes of data is collected, with a resolution not exceeding 2k and a frame rate of 25FPS. Then all the videos are parsed into video frame pictures at 25FPS, and the face information in the video frames is extracted through a face detection and key point detection model, and the source face image and the reference face image are filtered out through identity recognition.
[0050] Exemplarily, the video frames containing the reference face can be screened out from the reference video, and then the reference face image can be extracted from the screened video frames by using an alignment matrix. The source face image can also be extracted from the source video by using the same method.
[0051] Step 102: Use the source face image to train the first feature extraction network and the first decoding network in the initial face fusion model to obtain a second feature extraction network and a second decoding network.
[0052] In this application, the initial face fusion model can include a first feature extraction network, a first decoding network, etc. Among them, the first feature extraction network can be an initial feature extraction network, and the first decoding network can be an initial decoding network corresponding to the source face image.
[0053] In this application, the source face image is used for feature extraction by the first feature extraction network, and the extracted feature vector is input into the first decoding network for decoding. Based on the decoding result, the parameters of the first feature extraction network and the parameters of the first decoding network are adjusted. The goal is to make the image output by the first decoding network after parameter adjustment the same as the source face image, so as to decouple and fuse the identity information of the source face image into the neural network parameters, achieving a higher similarity.
[0054] Exemplarily, the first feature extraction network may include an encoder, an MLP (Multilayer Perceptron), etc. For example, the source face image can be input into the encoder to obtain a latent vector, the latent vector is input into the MLP for inference to obtain a new latent vector, and then the reconstructed source face image is obtained through the first decoding network.
[0055] Step 103: Use the reference face image and the second decoding network to train the second feature extraction network to obtain a trained face fusion model.
[0056] In this application, the parameters of the second decoding network can be fixed, and the second feature extraction network is continuously trained using the reference face image and the second decoding network to obtain a trained face fusion model.
[0057] Among them, the trained face fusion model may include a second decoding network, a feature extraction network obtained by training the second feature extraction network, etc.
[0058] It can be understood that the feature extraction network obtained by training the second feature extraction network can extract the information of the source face image and the reference face image.
[0059] Exemplarily, the reference face image can be input into the second feature extraction network for feature extraction. Based on the extracted feature vector of the reference face image, the second decoding network is used for decoding to obtain the first face fusion image, and the second feature extraction network is adjusted in parameters using the first face fusion image and the source face image.
[0060] Among them, the first face fusion image can be understood as the face image after the source face image and the reference face image are fused. The fused face image retains the identity information provided by the source face image and the attribute information provided by the reference face image.
[0061] The trained face fusion model of the embodiments of this application can be applied to scenarios such as film and television work processing. For example, the face image of object A is the source face image, and the face image of object B is the reference face image. The face fusion model is trained using the face images of object A and object B by the above method. Thus, the face fusion model can be used to replace the face of object B in the film and television work with the face of object A, so that the replaced face retains the identity information of object A and the attribute information of object B.
[0062] In the embodiments of the present application, by using the source face image to train the first feature extraction network and the first decoding network in the initial face fusion model, the second feature extraction network and the second decoding network are obtained. Then, using the reference face image and the second decoding network, the second feature extraction network is trained to obtain the trained face fusion model. Thus, by using the source face image to train the second decoding network, the second decoding network is not affected by the identity of the reference face image, thereby improving the similarity between the fused face and the source face. In addition, with the parameters of the second decoding network fixed, the second feature extraction network is continuously trained using the reference face image, so that the trained second feature extraction network can fully decouple the attribute information of the reference face image, making the fused face have the attribute information of the reference face image, improving the consistency of the expression, pose, etc. between the fused face and the reference face image, and thus improving the quality of the face fusion image.
[0063] Figure 2 It is a schematic flowchart of a method for training a face fusion model provided in another embodiment of the present application.
[0064] As Figure 2 shown, the method for training the face fusion model includes:
[0065] Step 201, obtain a source face image and a reference face image.
[0066] In the present application, step 201 can adopt any implementation manner in the embodiments of the present application, so it will not be elaborated here.
[0067] Step 202, use the first feature extraction network to extract features from the source face image to obtain a first feature vector.
[0068] In the present application, the source face image can be input into the first feature extraction network for feature extraction, such as extracting facial contour features, facial feature details, face spatial structure features, etc., to obtain a first feature vector.
[0069] Step 203, use the first decoding network to decode the first feature vector to obtain a first reconstructed image.
[0070] In the present application, the first feature vector can be input into the first decoding network for decoding to obtain a first reconstructed image. Among them, the first reconstructed image can be understood as the reconstructed source face image.
[0071] Step 204, according to the first reconstructed image and the source face image, train the first feature extraction network and the first decoding network to obtain a second feature extraction network and a second decoding network.
[0072] As a possible implementation, the first reconstruction loss can be determined according to the difference between the first reconstructed image and the source face image. Based on the first reconstruction loss, the first feature extraction network and the first decoding network are trained with the aim of making the image output by the first decoding network the same as the source face image.
[0073] Exemplarily, according to the first reconstruction loss, the parameters of the first feature extraction network and the first decoding network can be adjusted, and the source face image is used to continue training the first feature extraction network and the first decoding network with adjusted parameters until the training end condition is met, obtaining the second feature extraction network and the second decoding network.
[0074] Exemplarily, the mean square error between the first reconstructed image and the source face image can be calculated, and the mean square error is used as the first reconstruction loss. For example, the first reconstruction loss can be calculated using the following formula (1):
[0075]
[0076] where L1 represents the first reconstruction loss, N represents the total number of pixels in the source face image or the first reconstructed image, I s (i) represents the value of the i-th pixel in the source face image, and I1(i) represents the value of the i-th pixel in the first reconstructed image.
[0077] Thus, by determining the first reconstruction loss according to the first reconstructed image and the source face image, and using the first reconstruction loss to train the first feature extraction network and the first decoding network, the trained first decoding network can restore the source face image, improving the image reconstruction ability.
[0078] Optionally, the initial face fusion model may further include a first discriminative network. Also, according to the first reconstructed image and the source face image, using the first discriminative network, the first discriminative loss can be determined. Based on the first discriminative loss and the first reconstruction loss, the first feature extraction network and the first decoding network are trained to obtain the second feature extraction network and the second decoding network.
[0079] Exemplarily, the first reconstructed image and the source face image can be input into the first discriminative network, where the source face image is a real image and the first reconstructed image is a generated image. The first discriminative network can output the probability that the first reconstructed image is a real image, and the first discriminative loss is determined according to this probability. For example, the greater the probability that the first reconstructed image is a real image, the greater the first discriminative loss.
[0080] Exemplarily, the first discrimination loss and the first reconstruction loss can be weighted to obtain a weighted sum, and the first feature extraction network and the first decoding network can be trained according to the weighted sum. Among them, the weights of the first discrimination loss and the first reconstruction loss can be determined according to actual needs, and no limitation is imposed thereon.
[0081] Thus, according to the first reconstructed image and the source face image, the first discrimination loss can be determined through the first discrimination network. According to the first reconstruction loss combined with the first discrimination loss, the first feature extraction network and the first decoding network can be trained, so that the first reconstructed image approaches the source face image through the first discrimination network, and the quality of the image generated by the second decoding network can be improved.
[0082] Optionally, the source face image and the reconstructed source face image output by the second decoding network can also be used to train the first discrimination network to obtain a second discrimination network. Among them, the first feature extraction network and the first decoding network can be regarded as a generation network, and the generation network and the first discrimination network can be alternately trained until the training end condition is met.
[0083] Step 205: Use the reference face image and the second decoding network to train the second feature extraction network to obtain a trained face fusion model.
[0084] In this application, step 205 can adopt any implementation manner in the embodiments of this application, so it will not be elaborated herein.
[0085] In the embodiments of this application, the source face image is subjected to feature extraction through the first feature extraction network to obtain a first feature vector, and the first decoding network is used to decode the first feature vector to obtain a first reconstructed image. Based on the first reconstructed image and the source face image, the first feature extraction network and the first decoding network are trained, which can significantly improve the accuracy of face image feature extraction and enhance the reconstruction ability of the decoding network, etc.
[0086] Figure 3 It is a schematic flowchart of a method for training a face fusion model provided in another embodiment of this application.
[0087] As Figure 3 shown, the method for training the face fusion model includes:
[0088] Step 301: Obtain a source face image and a reference face image.
[0089] Step 302: Use the source face image to train the first feature extraction network and the first decoding network in the initial face fusion model to obtain a second feature extraction network and a second decoding network.
[0090] In this application, steps 301 - 302 can adopt any implementation manner in the embodiments of this application, so details are not described herein again.
[0091] Step 303: Use the second feature extraction network to extract features from the reference face image to obtain a second feature vector, and use the second decoding network to decode the second feature vector to obtain a first face fusion image.
[0092] In this application, the reference face image can be input into the second feature extraction network for feature extraction to obtain a second feature vector, and the second feature vector is input into the second decoding network for decoding. The second decoding network outputs a first face fusion image, and the obtained first face fusion image contains the identity information of the source face image. Thus, the purpose of replacing the reference face with the source face can be achieved through the second feature extraction network and the second decoding network.
[0093] Step 304: Use the second feature extraction network to extract features from the source face image to obtain a third feature vector, and use the second decoding network to decode the third feature vector to obtain a second reconstructed image.
[0094] In this application, the source face image can be input into the second feature extraction network for feature extraction to obtain a third feature vector, and the third feature vector is input into the second decoding network for decoding to obtain the second reconstructed image output by the second decoding network.
[0095] Step 305: Train the second feature extraction network according to the first face fusion image and the second reconstructed image to obtain a target feature extraction network.
[0096] In this application, the initial face fusion model may further include a third decoding network. The third decoding network can be a decoding network for reconstructing the reference face image. The second feature vector can also be input into the third decoding network for decoding to obtain a third reconstructed image output by the third decoding network. Then, train the second feature extraction network according to the first face fusion image, the second reconstructed image, and in combination with the third reconstructed image to obtain a target feature extraction network. Thus, the target feature extraction network can also extract features from the reference face image, improving the feature extraction ability of the target feature extraction network.
[0097] As a possible implementation, the initial face fusion model may further include a third discriminative network. The first face fusion image, the second reconstructed image, and the source face image may also be input into the second discriminative network. The second discriminative loss may be determined according to the output of the second discriminative network. The third reconstructed image and the reference face image may be input into the third discriminative network. The third discriminative loss may be determined according to the output of the third discriminative network. The second feature extraction network may be trained according to the second discriminative loss and the third discriminative loss to obtain the target feature extraction network.
[0098] Among them, the first face fusion image and the second reconstructed image may be generated images, and the source face image may be a real image. The second discriminative network outputs the probability that the first face fusion image is a real image and the probability that the second reconstructed image is a real image. Then, the second discriminative loss may be determined according to the probability that the first face fusion image output by the second discriminative network is a real image and the probability that the second reconstructed image is a real image.
[0099] Among them, the third reconstructed image may be a generated image, and the reference face image may be used as a real image. The third discriminative network outputs the probability that the third reconstructed image is a real image. The third discriminative loss may be determined according to the probability that the third reconstructed image is a real image.
[0100] Exemplarily, the second feature extraction network may be trained according to the second discriminative loss first, and then the third decoding network and the feature extraction network trained based on the second discriminative loss may be trained using the third discriminative loss. Alternatively, the second feature extraction network and the third decoding network may be trained according to the third discriminative loss first, and then the feature extraction network trained based on the third discriminative loss may be continuously trained according to the second discriminative loss.
[0101] Thus, based on the second discriminative network and the third discriminative network, training the second feature extraction network can improve the feature extraction ability of the model, and further improve the similarity between the fused face and the source face and the consistency of the attribute information between the fused face and the reference face.
[0102] Optionally, the second reconstruction loss may also be determined according to the difference between the second reconstructed image and the source face image, and the third reconstruction loss may be determined according to the difference between the third reconstructed image and the reference face image. The second feature extraction network may be trained according to the second discriminative loss and the third discriminative loss, in combination with the second reconstruction loss and the third reconstruction loss, to obtain the target feature extraction network.
[0103] For example, the second feature extraction network may be trained first according to the sum of the second discriminative loss and the second reconstruction loss, and then the third decoding network and the trained feature extraction network may be trained according to the third reconstruction loss.
[0104] For another example, the second feature extraction network and the third decoding network can be trained according to the third discrimination loss, and then the obtained feature extraction network can be further trained according to the second discrimination loss and the second reconstruction loss.
[0105] Thus, based on the second discrimination network and the third discrimination network, combined with the reconstruction loss of the source face image by the second decoding network and the reconstruction loss of the reference face image by the third decoding network, the second feature extraction network is trained, thereby further improving the feature extraction ability of the second feature extraction network.
[0106] As another possible implementation, the generation loss can be determined according to the difference between the first face fusion image and the source face image, and the second feature extraction network can be trained according to the generation loss, the second reconstruction loss and the third reconstruction loss to obtain the target feature extraction network.
[0107] For example, the second feature extraction network can be trained according to the generation loss and the second reconstruction loss, and then the third reconstruction loss can be used to further train the third decoding network and the obtained feature extraction network. For another example, the second feature extraction network and the third decoding network can be trained first according to the third reconstruction loss, and then the obtained feature extraction network can be further trained according to the generation loss and the second reconstruction loss.
[0108] Thus, based on the generation loss, the reconstruction loss of the source face image by the second decoding network and the reconstruction loss of the reference face image by the third decoding network, the second feature extraction network is trained, which can improve the feature extraction ability of the second feature extraction network while reducing the training complexity.
[0109] In the embodiments of the present application, by using the second feature extraction network and the second decoding network, the reference face image is replaced to obtain the first face fusion image, and the source face image is reconstructed by using the second feature extraction network and the second decoding network to obtain the second reconstruction image. Based on the first face fusion image and the second reconstruction image, the second feature extraction network is trained to obtain the target feature extraction network, which can improve the feature extraction ability of the target feature extraction network, and further improve the similarity between the face fusion image output by the second decoding network and the source face image.
[0110] In one embodiment of the present application, the initial face fusion model may include a first feature extraction network and a first decoding network, and may also include a third decoding network. First, the second feature extraction network may be trained according to the first face fusion image and the second reconstructed image to obtain a third feature extraction network. Then, the third feature extraction network may be used to extract features from the reference face image to obtain a fourth feature vector, and the third decoding network may be used to decode the fourth feature vector to obtain a fourth reconstructed image. Based on the fourth reconstructed image and the reference face image, the third feature extraction network and the third decoding network may be trained. On this basis, the reference image and the source face image may be used to continue the training, and finally, a target feature extraction network may be obtained.
[0111] Exemplarily, according to the first face fusion image, the second reconstructed image, and the source face image, a second discrimination loss may be obtained through a second discrimination network, and a second reconstruction loss may be determined according to the second reconstructed image and the source face image. Then, based on the sum of the second discrimination loss and the second reconstruction loss, the second feature extraction network may be trained to obtain a third feature extraction network.
[0112] Exemplarily, according to the fourth reconstructed image and the reference face image, a fourth discrimination loss may be obtained through a third discrimination network, and a fourth reconstruction loss may be obtained according to the difference between the fourth reconstructed image and the reference face image. Based on the sum of the fourth discrimination loss and the fourth reconstruction loss, the third feature extraction network and the third decoding network may be trained.
[0113] Thus, first, according to the face fusion image and the reconstructed source face image obtained based on the reference face image, the second feature extraction network may be trained, and then the reconstructed reference face image may be used to continue the training, thereby continuously training the shared feature extraction network and improving the feature extraction ability of the feature extraction network.
[0114] In one embodiment of the present application, the initial face fusion model may include a first feature extraction network and a first decoding network, and may also include a third decoding network. First, the second feature extraction network may be used to extract features from the reference face image to obtain a second feature vector, and the third decoding network may be used to decode the reference face image to obtain a third reconstructed image. Based on the third reconstructed image and the reference face image, the second feature extraction network and the third decoding network may be trained. After training, the second feature extraction network may obtain a fourth feature extraction network. Then, the fourth feature extraction network may be used to extract features from the reference face image to obtain a fifth feature vector, and the second decoding network may be used to decode the fifth feature vector to obtain a second face fusion image. Then, based on the second face fusion image, the fifth reconstructed image, and the source face image, the fourth feature extraction network may be trained. On this basis, the reference image and the source face image may be used to continue the training, and finally, a target feature extraction network may be obtained.
[0115] Exemplarily, a third reconstruction loss may be obtained according to the difference between the third reconstructed image and the reference face image. According to the third reconstructed image and the reference face image, a third discrimination loss may be obtained through a third discrimination network. Based on the third reconstruction loss and the third discrimination loss, the second feature extraction network and the third decoding network are trained.
[0116] Exemplarily, according to the second face fusion image, the fifth reconstructed image, and the source face image, a fifth discrimination loss may be obtained through a second discrimination network. According to the difference between the fifth reconstructed image and the source face image, a sixth reconstruction loss may be obtained. Then, based on the sum of the fifth discrimination loss and the sixth reconstruction loss, the fourth feature extraction network is trained.
[0117] Thus, the second feature extraction network may be trained first according to the reconstructed reference face image, and then the face fusion image and the reconstructed source face image obtained by using the reference face image are used to continue the training, so as to continuously train the shared feature extraction network and improve the feature extraction ability of the feature extraction network.
[0118] It should be noted that in this application, there is no limitation on the execution order of training the shared feature extraction network by using the face fusion image obtained based on the reference face image and the source face image reconstructed based on the second decoding network, and training the decoder corresponding to the reference face image and the common feature extraction network by using the reconstructed reference face image.
[0119] In an embodiment of this application, as Figure 4 shown, the initial face fusion model may include a shared first feature extraction network 401, a first decoding network 402 corresponding to the source face image, a first discrimination network 403, a third decoding network 404 corresponding to the reference face image, and a third discrimination network 405. That is to say, the source face image and the reference face image share a feature extraction network, and the source face image and the reference face image each have a corresponding decoding network and discrimination network.
[0120] When training the initial face fusion model, the first feature extraction network 401, the first decoding network 402, and the first discrimination network 403 may be trained first by using the source face image. Among them, the first feature extraction network 401 and the first decoding network 402 may be used as a generation network, and the generation network and the first discrimination network 403 are alternately trained to finally obtain a second feature extraction network, a second decoding network, and a second discrimination network.
[0121] After that, with the parameters of the second decoding network fixed, the second feature extraction network, the second discriminative network, the third decoding network, and the third discriminative network are trained using the reference face image, the source face image, and the second decoding network, where the second discriminative network and the third discriminative network can be trained separately.
[0122] It should be noted that this application does not limit the execution order of training the shared feature extraction network using the face fusion image obtained from the reference face image and the source face image reconstructed based on the second decoding network, and training the decoder corresponding to the reference face image and the shared feature extraction network using the reconstructed reference face image.
[0123] Finally, the second feature extraction network, the second discriminative network, the third decoding network, and the third discriminative network are correspondingly trained to obtain the target feature extraction network, the fourth discriminative network, the fourth decoding network, and the fifth discriminative network after training. That is to say, the trained face fusion model includes the target feature extraction network, the second decoding network corresponding to the source face image and the fourth discriminative network, and the fourth decoding network and the fifth discriminative network corresponding to the reference face image.
[0124] To implement the above embodiments, this application's embodiments also propose a face fusion method. Figure 5 It is a schematic flowchart of the face fusion method provided by an embodiment of this application.
[0125] As Figure 5 described, the face fusion method includes:
[0126] Step 501, obtain a reference face image.
[0127] Exemplarily, the reference face image can be extracted from a single captured image, or from different frame images in a video, or obtained through other means, and this is not limited.
[0128] Step 502, input the reference face image into the face fusion model to obtain a face fusion image.
[0129] In this application, the face fusion model can be trained using the training method of any of the above embodiments.
[0130] Exemplarily, the face fusion model can include a target feature extraction network and a second decoding network, where the second decoding network is trained using the source face image. The target feature extraction network can be used to extract features from the reference face image to obtain a feature vector, and the feature vector is input into the second decoding network for decoding to obtain a face fusion image.
[0131] Among them, the face fusion image includes the identity information of the source face image and the attribute information of the reference face image.
[0132] In the embodiment of the present application, by using the face fusion model trained by the above training method for face fusion, the influence of the identity information of the reference face image can be removed, so that the fused face has a high similarity with the source face.
[0133] In an embodiment of the present application, the reference face image can be extracted from the original reference image based on the alignment matrix. For example, multiplying the original reference image by the alignment matrix to obtain the reference face image. After obtaining the face fusion image, the face fusion image can also be masked to obtain a face mask image, and the face mask image is fused with the original reference image to obtain an initial fusion image, and then the initial fusion image is unaligned according to the face alignment matrix to obtain the target fusion image.
[0134] Exemplarily, a segmentation model can be used to mask the face fusion image to obtain a face mask image. Among them, the face mask image can be a binary image or a grayscale image, which is not limited thereto. For example, the face area in the face mask image can be marked as 1, and other areas can be marked as 0.
[0135] Exemplarily, the alignment matrix can be inversely transformed to obtain an inverse matrix, and according to the inverse matrix, an image transformation operation is performed on the face fusion image, so as to unalign the initial fusion image to the original reference image to obtain the target fusion image.
[0136] Thus, by fusing the face mask image and the original reference image, the face area of the face fusion image and the background area of the original reference image can be retained, and the initial fusion image is unaligned to the original reference image by using the alignment matrix to achieve the final face fusion and improve the face fusion quality.
[0137] To facilitate understanding of the solution of the embodiment of the present application, the following is combined with Figure 6 for illustration. Figure 6 This is a schematic diagram of the training process of the face fusion model provided by the embodiment of the present application.
[0138] As Figure 6 shown, the face fusion model includes an encoder 601, an MLP 602, a source image decoder 603, a source image discriminator 604, a reference image decoder 605, and a reference image discriminator 606. That is to say, Figure 6 the face fusion model shown includes an encoder, a multi-layer perceptron, two decoders, and two discriminators.
[0139] Among them, the source image decoder 603 and the source image discriminator 604 are the decoder and discriminator corresponding to the source face image, and the reference image decoder 605 and the reference image discriminator 606 are the decoder and discriminator corresponding to the reference face image.
[0140] Among them, the encoder 601 can be composed of multiple downsampled convolutional neural networks, which are responsible for compressing the source face image and the reference face image into a high-dimensional latent space to obtain a latent vector that is more conducive to subsequent network processing. The MLP 602 can be composed of multiple fully connected layers, which are responsible for understanding the information such as the face position, pose, and expression in the latent vector and enhancing the learning ability of the neural network. The two decoders, the source image decoder 603 and the reference image decoder 605, are composed of multiple upsampled convolutional neural networks, which decode the latent vector into the image space to obtain the fused face image. The two discriminators, the source image discriminator 604 and the reference image discriminator 606, are responsible for improving the quality of the face-fused image.
[0141] The training of the face fusion model includes two steps:
[0142] Step 1: First, train a decoder that only generates source identity images. The purpose of Step 1 is to obtain a network that only generates the identity of the source face image. The detailed training method can include: preparing about 500,000 processed source face images, inputting 16 images per batch into the encoder 601 to obtain a latent vector, inputting the latent vector into the MLP 602 for inference to obtain a new latent vector, and then obtaining the reconstructed source face image through the source image decoder 603. Subsequently, the source face image and the reconstructed source face image are input into the source image discriminator 604, and the source image discriminator 604 is responsible for making the quality of the reconstructed image approach that of the source face image. Since the source image decoder 603 has not received the information of the reference image and can only output the image of the source face, it will not be affected by the identity of the reference face image, thus ensuring the similarity of the fused face.
[0143] Step 2: Freeze the parameters of the source image decoder 603 trained in Step 1. Subsequently, the reference face image is also added to the training, and the encoder 601, the MLP 602, the reference image decoder 605, and the reference image discriminator 606 are continued to be trained. The purpose of Step 2 is to enable the encoder 601 to fully decouple the attribute information of the reference image, so that the fused face has the pose and expression of the reference image without affecting its similarity.
[0144] The detailed training process of step 2 may include: The reference face image can be input into the encoder 601. After passing through the encoder 601 and the MLP 602, the reference face image is input into the reference image decoder 605 to output the reconstructed reference face image. The reference face image and the reconstructed reference face image are input into the reference image discriminator 606, and the reference image discriminator 606 improves the quality of the reconstructed image. The source face image is input into the encoder 601. After passing through the encoder 601 and the MLP 602, it is input into the source image decoder 603 to output the reconstructed source face image, and the reconstructed source face image is input into the source image discriminator 604.
[0145] Meanwhile, the latent vector of the reference face image output by the MLP 602 can also be input into the source image decoder 603 for decoding to obtain the face fusion image. The face fusion image, the source face image, and the reconstructed source face image are all input into the source image discriminator 604. Thus, inputting the face fusion image into the source image discriminator 604 reduces the distance between it and the source face image in the space of the source image discriminator 604, further improving the face fusion similarity.
[0146] It should be noted that this application does not limit the execution order of training the encoder 601 and the MLP 602 using the first total loss obtained from the face fusion image, the source face image, and the reconstructed source face image in step 2, and training the encoder 601, the MLP 602, and the reference image decoder 605 using the second total loss obtained from the reference face image and the reconstructed reference face image.
[0147] For example, the network parameters can be adjusted according to the order of obtaining the two total losses. If the two total losses are obtained simultaneously, the shared encoder 601 and MLP 602 can be first adjusted using the first total loss. On this basis, the reference image decoder 605 and the shared encoder 601 and MLP 602 can be further adjusted using the second total loss. Or, the reference image decoder 605 and the shared encoder 601 and MLP 602 can be first adjusted using the second total loss. On this basis, the shared encoder 601 and MLP 602 can be further adjusted using the first total loss.
[0148] Inference process: After inputting the reference face image into the encoder 601 and the MLP 602, it is then input into the source image decoder 603 for decoding to obtain the face fusion image.
[0149] For an original reference image, first perform data alignment and filtering to obtain a reference face image and store the data alignment matrix. Subsequently, input the reference face image into the encoder 601 to obtain the latent vector of the reference face image, and then input it into the MLP 602 for inference to obtain a new latent vector. Input the new latent vector into the source image decoder 603 to obtain the face fusion image output by the source image decoder 603. After that, perform masking on the face fusion image to obtain a face mask image, and fuse the face mask image with the original reference image, that is, retain the face area of the face fusion image and the background area of the original reference image. Finally, use the stored alignment matrix to align the fused image back to the original reference image to achieve the final face fusion.
[0150] The fused face obtained by the face fusion solution of this application has a high similarity with the source face, can remove the influence of the identity information of the reference face image, and improves the quality of the face fusion image.
[0151] It should be noted that Figure 6 The decoder in corresponds to the decoding network in the above-mentioned embodiment, the discriminator corresponds to the discriminant network in the above-mentioned embodiment, and the encoder and MLP belong to the feature extraction network.
[0152] To implement the above-mentioned embodiment, the embodiment of this application also proposes a training device for a face fusion model. Figure 7 It is a schematic structural diagram of the training device for the face fusion model provided by an embodiment of this application.
[0153] As Figure 7 shown, the training device 700 for the face fusion model includes:
[0154] An acquisition module 710, configured to acquire a source face image and a reference face image;
[0155] A first training module 720, configured to train the first feature extraction network and the first decoding network in the initial face fusion model by using the source face image to obtain a second feature extraction network and a second decoding network;
[0156] A second training module 730, configured to train the second feature extraction network by using the reference face image and the second decoding network to obtain a trained face fusion model.
[0157] Optionally, the first training module 720 is configured to:
[0158] Extract features from the source face image by using the first feature extraction network to obtain a first feature vector;
[0159] Decode the first feature vector by using the first decoding network to obtain a first reconstructed image;
[0160] According to the first reconstructed image and the source face image, train the first feature extraction network and the first decoding network to obtain the second feature extraction network and the second decoding network.
[0161] Optionally, the first training module 720 is configured to:
[0162] Determine a first reconstruction loss according to the first reconstructed image and the source face image;
[0163] Train the first feature extraction network and the first decoding network according to the first reconstruction loss to obtain the second feature extraction network and the second decoding network.
[0164] Optionally, the initial face fusion model further includes a first discriminant network, and the first training module 720 is configured to:
[0165] Determine a first discriminant loss according to the first reconstructed image and the source face image by using the first discriminant network;
[0166] Train the first feature extraction network and the first decoding network according to the first reconstruction loss and the first discriminant loss to obtain the second feature extraction network and the second decoding network.
[0167] Optionally, the second training module 730 is configured to:
[0168] Extract features from the reference face image by using the second feature extraction network to obtain a second feature vector, and decode the second feature vector by using the second decoding network to obtain a face fusion image;
[0169] Extract features from the source face image by using the second feature extraction network to obtain a third feature vector, and decode the third feature vector by using the second decoding network to obtain a second reconstructed image;
[0170] Train the second feature extraction network according to the face fusion image and the second reconstructed image to obtain a target feature extraction network.
[0171] Optionally, the initial face fusion model further includes a third decoding network, and the second training module 730 is configured to:
[0172] Decode the second feature vector by using the third decoding network to obtain a third reconstructed image;
[0173] Training the second feature extraction network according to the face fusion image, the second reconstructed image, and the third reconstructed image to obtain the target feature extraction network.
[0174] Optionally, the second training module 730 is configured to:
[0175] Optionally, determining a second discrimination loss according to the face fusion image, the second reconstructed image, and the source face image by using a second discrimination network;
[0176] Determining a third discrimination loss according to the third reconstructed image and the reference face image through a third discrimination network in the initial face fusion model;
[0177] Training the second feature extraction network according to the second discrimination loss and the third discrimination loss to obtain the target feature extraction network.
[0178] Optionally, the second training module 730 is configured to:
[0179] Determining a second reconstruction loss according to the second reconstructed image and the source face image;
[0180] Determining a third reconstruction loss according to the third reconstructed image and the reference face image;
[0181] Training the second feature extraction network according to the second discrimination loss, the third discrimination loss, the second reconstruction loss, and the third reconstruction loss to obtain the target feature extraction network.
[0182] Optionally, the second training module 730 is configured to:
[0183] Determining a generation loss according to the face fusion image and the source face image;
[0184] Determining a second reconstruction loss according to the second reconstructed image and the source face image;
[0185] Determining a third reconstruction loss according to the third reconstructed image and the reference face image;
[0186] Training the second feature extraction network according to the generation loss, the second reconstruction loss, and the third reconstruction loss to obtain the target feature extraction network.
[0187] Optionally, the initial face fusion model further includes a third decoding network, and the second training module 730 is configured to:
[0188] Training the second feature extraction network according to the first face fusion image and the second reconstructed image to obtain a third feature extraction network;
[0189] Use the third feature extraction network to extract features from the reference face image to obtain a fourth feature vector, and use the third decoding network to decode the fourth feature vector to obtain a fourth reconstructed image;
[0190] Train the third feature extraction network according to the fourth reconstructed image and the reference face image to obtain the target feature extraction network.
[0191] Optionally, the initial face fusion model further includes a third decoding network, and the second training module 730 is used for:
[0192] Use the second feature extraction network to extract features from the reference face image to obtain a second feature vector, and use the third decoding network to decode the reference face image to obtain a third reconstructed image;
[0193] Train the second feature extraction network according to the third reconstructed image and the reference face image to obtain a fourth feature extraction network;
[0194] Use the fourth feature extraction network to extract features from the reference face image to obtain a fifth feature vector, and use the second decoding network to decode the fifth feature vector to obtain a second face fusion image;
[0195] Use the fourth feature extraction network to extract features from the source face image to obtain a sixth feature vector, and use the second decoding network to decode the sixth feature vector to obtain a fifth reconstructed image;
[0196] Train the fourth feature extraction network according to the second face fusion image, the fifth reconstructed image and the source face image to obtain the target feature extraction network.
[0197] It should be noted that the explanations of the foregoing embodiments of the training method of the face fusion model also apply to the training device of the face fusion model in this embodiment, so details are not described herein again.
[0198] In the embodiments of the present application, by using the source face image to train the first feature extraction network and the first decoding network in the initial face fusion model, a second feature extraction network and a second decoding network are obtained. Then, using the reference face image and the second decoding network, the second feature extraction network is trained to obtain a trained face fusion model. Thus, by training the second decoding network using the source face image, the second decoding network is not affected by the identity of the reference face image, thereby improving the similarity between the fused face and the source face. Additionally, with the parameters of the second decoding network fixed, the second feature extraction network is continuously trained using the reference face image, enabling the trained second feature extraction network to fully decouple the attribute information of the reference face image, such that the fused face has the attribute information of the reference face image, improving the consistency of the fused face with the expression, pose, etc. of the reference face image, and thus enhancing the quality of the face fusion image.
[0199] To implement the above embodiments, the embodiments of the present application also propose a face fusion device. Figure 8 The following is a schematic structural diagram of the face fusion device provided in an embodiment of the present application.
[0200] As Figure 8 shown, the face fusion device 800 includes:
[0201] An acquisition module 810, configured to acquire a reference face image;
[0202] A face fusion module 820, configured to input the reference face image into a face fusion model to obtain a face fusion image; wherein, the face fusion model is trained by using the training method described in any of the above embodiments.
[0203] It should be noted that the explanations of the foregoing face fusion method embodiments also apply to the face fusion device of this embodiment, and thus will not be elaborated herein.
[0204] In the embodiments of the present application, by performing face fusion using the face fusion model trained by the above training method, the influence of the identity information of the reference face image can be removed, such that the fused face has a high similarity to the source face.
[0205] According to the embodiments of the present application, the present application also provides an electronic device, a readable storage medium, and a computer program product.
[0206] Figure 9FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present application described and / or claimed herein.
[0207] As Figure 9 shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 902 or a computer program loaded from a storage unit 908 into a RAM (Random Access Memory) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An I / O (Input / Output) interface 905 is also connected to the bus 904.
[0208] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0209] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the training method of the face fusion model. For example, in some embodiments, the training method of the face fusion model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the training method of the face fusion model described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the training method of the face fusion model in any other suitable way (e.g., by means of firmware).
[0210] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0211] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0212] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0213] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0214] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0215] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (Virtual Private Server). The server can also be a server of a distributed system, or a server combined with blockchain.
[0216] It should be noted that the electronic device used to implement the face fusion method of the embodiments of the present application is similar in structure to the above-mentioned electronic device, so it will not be elaborated herein.
[0217] According to an embodiment of the present application, the present application also provides a computer program product, which, when executed by an instruction processor in the computer program product, executes the training method of the face fusion model or the face fusion method proposed in the above embodiments of the present application.
[0218] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present application can be achieved, and no limitation is imposed herein.
[0219] The above specific embodiments do not constitute a limitation to the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A training method for a face fusion model, comprising: Obtaining a source face image and a reference face image; Using the source face image to train a first feature extraction network and a first decoding network in an initial face fusion model to obtain a second feature extraction network and a second decoding network; The second feature extraction network is trained using the reference face image and the second decoding network to obtain a trained face fusion model.
2. The method of claim 1, wherein: The method of using the source face image to train the first feature extraction network and the first decoding network in the initial face fusion model to obtain the second feature extraction network and the second decoding network includes: Using the first feature extraction network to extract features from the source face image to obtain a first feature vector; Decoding the first feature vector using the first decoding network to obtain a first reconstructed image; The first feature extraction network and the first decoding network are trained according to the first reconstructed image and the source face image to obtain the second feature extraction network and the second decoding network.
3. The method of claim 2, wherein: The step of training the first feature extraction network and the first decoding network according to the first reconstructed image and the source face image to obtain the second feature extraction network and the second decoding network comprises: Determining a first reconstruction loss according to the first reconstructed image and the source face image; The first feature extraction network and the first decoding network are trained according to the first reconstruction loss to obtain the second feature extraction network and the second decoding network.
4. The method of claim 3, wherein: The initial face fusion model also includes a first discriminant network, and the first feature extraction network and the first decoding network are trained according to the first reconstruction loss to obtain the second feature extraction network and the second decoding network, including: Determine a first discriminant loss using the first discriminant network according to the first reconstructed image and the source face image; The first feature extraction network and the first decoding network are trained according to the first reconstruction loss and the first discrimination loss to obtain the second feature extraction network and the second decoding network.
5. The method of claim 1, wherein: The step of training the second feature extraction network by using the reference face image and the second decoding network includes: Using the second feature extraction network to extract features from the reference face image to obtain a second feature vector, and using the second decoding network to decode the second feature vector to obtain a first face fusion image; Using the second feature extraction network to extract features from the source face image to obtain a third feature vector, and using the second decoding network to decode the third feature vector to obtain a second reconstructed image; The second feature extraction network is trained according to the first face fusion image and the second reconstructed image to obtain a target feature extraction network.
6. The method of claim 5, wherein: The initial face fusion model further includes a third decoding network, and the second feature extraction network is trained according to the first face fusion image and the second reconstructed image to obtain a target feature extraction network, including: Decoding the second feature vector using the third decoding network to obtain a third reconstructed image; The second feature extraction network is trained according to the first face fusion image, the second reconstructed image and the third reconstructed image to obtain the target feature extraction network.
7. The method of claim 6, wherein: The step of training the second feature extraction network according to the first face fusion image, the second reconstructed image, and the third reconstructed image to obtain the target feature extraction network includes: Determine a second discriminant loss using a second discriminant network according to the first face fusion image, the second reconstructed image and the source face image; Determining a third discriminant loss through a third discriminant network in the initial face fusion model according to the third reconstructed image and the reference face image; The second feature extraction network is trained according to the second discriminant loss and the third discriminant loss to obtain the target feature extraction network.
8. The method of claim 7, wherein: The step of training the second feature extraction network according to the second discriminant loss and the third discriminant loss to obtain the target feature extraction network includes: determining a second reconstruction loss according to the second reconstructed image and the source face image; Determining a third reconstruction loss according to the third reconstructed image and the reference face image; The second feature extraction network is trained according to the second discriminant loss, the third discriminant loss, the second reconstruction loss and the third reconstruction loss to obtain the target feature extraction network.
9. The method of claim 6, wherein: The step of training the second feature extraction network according to the first face fusion image, the second reconstructed image, and the third reconstructed image to obtain the target feature extraction network includes: Determining a generation loss according to the first face fusion image and the source face image; determining a second reconstruction loss according to the second reconstructed image and the source face image; Determining a third reconstruction loss according to the third reconstructed image and the reference face image; The second feature extraction network is trained according to the generation loss, the second reconstruction loss and the third reconstruction loss to obtain the target feature extraction network.
10. The method of claim 5, wherein: The initial face fusion model further includes a third decoding network, and the second feature extraction network is trained according to the first face fusion image and the second reconstructed image to obtain a target feature extraction network, including: Training the second feature extraction network according to the first face fusion image and the second reconstructed image to obtain a third feature extraction network; Using the third feature extraction network to extract features from the reference face image to obtain a fourth feature vector, and using the third decoding network to decode the fourth feature vector to obtain a fourth reconstructed image; The third feature extraction network is trained according to the fourth reconstructed image and the reference face image to obtain the target feature extraction network.
11. The method of claim 1, wherein: The initial face fusion model further includes a third decoding network, and the training of the second feature extraction network using the reference face image and the second decoding network includes: Using the second feature extraction network to extract features from the reference face image to obtain a second feature vector, and using the third decoding network to decode the reference face image to obtain a third reconstructed image; Training the second feature extraction network according to the third reconstructed image and the reference face image to obtain a fourth feature extraction network; Using the fourth feature extraction network to extract features from the reference face image to obtain a fifth feature vector, and using the second decoding network to decode the fifth feature vector to obtain a second face fusion image; Using the fourth feature extraction network to extract features from the source face image to obtain a sixth feature vector, and using the second decoding network to decode the sixth feature vector to obtain a fifth reconstructed image; The fourth feature extraction network is trained according to the second face fusion image, the fifth reconstructed image and the source face image to obtain the target feature extraction network.
12. A face fusion method, comprising: Obtain a reference face image; The reference face image is input into a face fusion model to obtain a face fusion image; wherein the face fusion model is trained using the method described in any one of claims 1-11.
13. The method of claim 12, wherein: The reference face image is extracted from the original reference image based on the alignment matrix, and the method further comprises: Performing mask processing on the face fusion image to obtain a face mask image; Fusing the face mask image with the original reference image to obtain an initial fused image; The initial fused image is de-aligned according to the alignment matrix to obtain a target fused image.
14. A training device for a face fusion model, comprising: An acquisition module, used for acquiring a source face image and a reference face image; A first training module is used to train a first feature extraction network and a first decoding network in an initial face fusion model using the source face image to obtain a second feature extraction network and a second decoding network; The second training module is used to train the second feature extraction network using the reference face image and the second decoding network to obtain a trained face fusion model.
15. A face fusion device, comprising: An acquisition module, used for acquiring a reference face image; A face fusion module is used to input the reference face image into a face fusion model to obtain a face fusion image; wherein the face fusion model is trained using the method described in any one of claims 1-13.
16. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.
17. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-13.
18. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13.