Model determination method, image processing method, device, equipment and storage medium
By training the target face-changing model through feature alignment processing, the problem of introducing texture and skin color attribute features of real face images into virtual face images is solved, ensuring the style consistency of the face-changing images and improving the image quality and authenticity.
Patent Information
- Application Number
- CN202210975057.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-15
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-08-15
AI Technical Summary
When existing technologies use a target face-changing model to replace the identity features of a real face image with a virtual face image, it is easy to introduce attribute features such as texture and skin color, resulting in a facial style inconsistency between the face-changing image and the target virtual face image, creating a sense of disharmony.
By constructing a training sample set, training the initial face-changing model to obtain an intermediate face-changing model, and inputting the virtual face image into the initial virtual face reconstruction model, performing feature alignment processing, correcting the decoder in the initial virtual face reconstruction model, obtaining the target virtual face reconstruction model, and finally obtaining the target face-changing model to ensure that the style of the face-changing image is consistent with the style of the virtual face image.
The facial style consistency between the face-swapped image and the target virtual face image is achieved, avoiding the sense of incongruity after the real face image is migrated to the virtual face image, and improving the authenticity and image quality of the face-swapped image.
Smart Images

Figure CN115294423B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a model determination method, image processing method, apparatus, device and storage medium. Background Art
[0002] With the development of the Internet and computer technology, face-swapping has become a new trend in social entertainment. For example, in gaming, players can replace a target image (e.g., a game character) with a source image (e.g., a player's image or an image of a favorite celebrity), thereby changing the character's identity while preserving its attributes.
[0003] At present, the initial face-changing model is trained based on the training samples constructed based on the real face images as the source image samples and the virtual face images as the target image samples, and the target face-changing model is directly obtained. The target face-changing model is used to replace the identity features of the target real face image with the target virtual face image to obtain the face-changing image.
[0004] However, when using the target face-changing model to replace the identity features of the target real face image with the identity features of the target virtual face image, the attribute features of the target real face image, such as texture and skin color, may be introduced. This will cause a sense of incongruity after the target real person image is migrated to the virtual person image, and it is difficult to ensure the consistency of the facial style of the face-changing image and the target virtual face image. Summary of the Invention
[0005] The purpose of this application is to address the deficiencies in the above-mentioned prior art and provide a model determination method, image processing method, device, equipment and storage medium, which can avoid the sense of incongruity after the real person image is migrated to the virtual human image, and thus ensure the consistency of the facial style of the face-swapped image and the target virtual human face image.
[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of the present application are as follows:
[0007] In a first aspect, an embodiment of the present application provides a model determination method, the method comprising:
[0008] Constructing a first training sample according to a training sample set, wherein the training sample set includes real face image samples and virtual face image samples, and the first training sample includes a source image sample and a target image sample;
[0009] training an initial face-swapping model based on the first training sample to obtain an intermediate face-swapping model, wherein the intermediate face-swapping model is used to process a fused facial feature vector to obtain a predicted face-swapping image, wherein the fused facial feature vector includes an identity feature vector of a source image sample and an attribute feature vector of a target image sample;
[0010] Inputting the virtual facial image samples in the training sample set into the initial virtual facial reconstruction model to obtain a virtual facial feature vector;
[0011] performing feature alignment on the virtual facial feature vector and the fused facial feature vector generated by the intermediate face-swap model, and correcting an initial virtual facial reconstruction decoder in the initial virtual facial reconstruction model based on a loss in the feature alignment to obtain a target virtual facial reconstruction model, wherein the target virtual facial reconstruction model is used to process the virtual facial feature vector to obtain a reconstructed virtual facial image;
[0012] A target face-changing model is obtained according to the target virtual face reconstruction model and the intermediate face-changing model.
[0013] In a second aspect, an embodiment of the present application further provides an image processing method, the method comprising:
[0014] Obtaining a target real face image and a target virtual face image;
[0015] The target real facial image and the target virtual facial image are respectively input into a target face-changing model to obtain a face-changing image, wherein the face-changing image includes the identity features of the target real facial image and the attribute features of the target virtual facial image, wherein the target face-changing model is obtained by the model determination method of the first aspect mentioned above.
[0016] In a third aspect, an embodiment of the present application further provides a model determination device, the device comprising:
[0017] A construction module, configured to construct a first training sample based on a training sample set, wherein the training sample set includes real facial image samples and virtual facial image samples, and the first training sample includes a source image sample and a target image sample;
[0018] a first determination module, configured to train an initial face-swapping model based on the first training sample to obtain an intermediate face-swapping model, wherein the intermediate face-swapping model is configured to process a fused facial feature vector to obtain a predicted face-swapping image, wherein the fused facial feature vector includes an identity feature vector of a source image sample and an attribute feature vector of a target image sample;
[0019] A first input module is configured to input the virtual facial image samples in the training sample set into an initial virtual facial reconstruction model to obtain a virtual facial feature vector;
[0020] a feature alignment module, configured to perform feature alignment on the virtual facial feature vector and the fused facial feature vector generated by the intermediate face-swap model, and to correct an initial virtual facial reconstruction decoder in the initial virtual facial reconstruction model based on a loss in the feature alignment to obtain a target virtual facial reconstruction model, wherein the target virtual facial reconstruction model is used to process the virtual facial feature vector to obtain a reconstructed virtual facial image;
[0021] The second determining module is used to obtain a target face-changing model according to the target virtual face reconstruction model and the intermediate face-changing model.
[0022] In a fourth aspect, an embodiment of the present application further provides an image processing device, comprising:
[0023] An acquisition module, used to acquire a target real face image and a target virtual face image;
[0024] The second input module is used to input the target real facial image and the target virtual facial image into the target face-changing model respectively to obtain a face-changing image, wherein the face-changing image includes the identity features of the target real facial image and the attribute features of the target virtual facial image, wherein the target face-changing model is obtained by the model determination device of the third aspect mentioned above.
[0025] In a fifth aspect, an embodiment of the present application provides an electronic device comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the model determination method of the first aspect or the steps of the image processing method of the second aspect.
[0026] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the model determination method of the first aspect or the steps of the image processing method of the second aspect are executed.
[0027] The beneficial effects of this application are:
[0028] An embodiment of the present application provides a model determination method, an image processing method, an apparatus, a device, and a storage medium, the method comprising: constructing a first training sample based on a training sample set; training an initial face-changing model based on the first training sample to obtain an intermediate face-changing model; inputting a virtual facial image sample in the training sample set into the initial virtual face reconstruction model to obtain a virtual facial feature vector; performing feature alignment on the virtual facial feature vector and a fused facial feature vector generated by the intermediate face-changing model, and correcting an initial virtual face reconstruction decoder in the initial virtual face reconstruction model based on the loss of feature alignment to obtain a target virtual face reconstruction model; and obtaining a target face-changing model based on the target virtual face reconstruction model and the intermediate face-changing model.
[0029] Using the model determination method provided in the embodiments of the present application, after the intermediate face-swapping model is trained using the first training sample, the initial virtual face reconstruction model including the initial virtual face reconstruction decoder can be trained based on the intermediate face-swapping model. Since the initial virtual face reconstruction model is fed with virtual face image samples, during the training of the initial virtual face reconstruction model, the virtual face feature vector and the fused face feature vector generated by the intermediate face-swapping model can be feature aligned to ensure that the distribution of the virtual face feature vector input to the initial virtual face reconstruction decoder is consistent with the distribution of the fused face feature vector. This allows the target virtual face reconstruction model obtained by the final training to not only focus on the attribute features of the virtual face image, such as texture and skin color, but also to normally decode the fused face feature vector containing real face image information. In other words, the target face-swapping model determined based on the target virtual face reconstruction model and the intermediate face-swapping model can ensure that the style of the generated face-swapping image is consistent with the style of the virtual face image, avoiding the sense of disharmony that occurs after the real face image is transferred to the virtual face image, thereby improving the authenticity and image quality of the face-swapping image. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0031] Figure 1 A flow chart of a model determination method provided in an embodiment of the present application;
[0032] Figure 2 A schematic diagram of the structure of an initial face-changing model provided in an embodiment of the present application;
[0033] Figure 3A schematic diagram of the structure of an intermediate face-swapping model combined with an initial virtual face reconstruction model provided in an embodiment of the present application;
[0034] Figure 4 A schematic diagram of the structure of a target face-changing model provided in an embodiment of the present application;
[0035] Figure 5 A flow chart of another model determination method provided in an embodiment of the present application;
[0036] Figure 6 A flowchart of another model determination method provided in an embodiment of the present application;
[0037] Figure 7 A flowchart of another model determination method provided in an embodiment of the present application;
[0038] Figure 8 A flowchart of an image processing method provided in an embodiment of the present application;
[0039] Figure 9 A schematic diagram of the structure of a model determination device provided in an embodiment of the present application;
[0040] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.
[0042] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without creative work are within the scope of protection of the present application.
[0043] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0044] In recent years, with the development of face synthesis technology, face swapping has been widely used. Face swapping refers to replacing the facial area of a target image with the facial area of a source image to change the target image's identity characteristics (such as face shape and eyebrow direction) while preserving the target image's attribute characteristics (such as head posture and facial expression).
[0045] However, the applicant has found through research that the identity features and attribute features in the image have a certain correlation. When the source image and the target image are swapped to obtain a face-swapped image, some attribute features in the source image (such as the player image), such as texture and skin color, may be brought into the face-swapped image. This will cause the texture, skin color and other attributes in the face-swapped image to deviate from the texture and skin color in the target image (such as the game character image), making the facial style of the face-swapped image inconsistent with the facial style corresponding to the target image.
[0046] In response to the above-mentioned problems, the present application uses the following embodiments to solve the problems. Before explaining the embodiments of the present application in detail, the application scenarios of the present application are first introduced. The application scenario can specifically be a scenario for personalizing the image of a game character. For example, the image of a game character in a game CG (Computer Graphics) video can be set to a favorite celebrity image, the player's own image, etc. The specific setting process can refer to the following examples of the present application. It should be noted that the technical solution provided by the present application can be applied not only in the field of games, but also in the fields of cultural tourism, film and television production, etc., without being limited thereto.
[0047] The embodiments mentioned below in this application can be divided into two parts, the first part is the training model stage, and the second part is the application model stage. For the first part, this application combines the pre-built initial face-changing model and the initial virtual face reconstruction model. In an exemplary manner, the initial face-changing model can be trained using the first training sample to obtain an intermediate face-changing model, and then the virtual face image is input into the initial virtual face reconstruction model to obtain virtual face features. Based on the virtual face features and the fused face features obtained by inputting the first training sample into the intermediate face-changing model again, feature alignment processing is performed, and then the target virtual face reconstruction model is obtained by training. Finally, the target face-changing model is obtained based on the intermediate face-changing model and the target virtual face reconstruction model. Alternatively, the initial face-swapping model and the initial virtual face reconstruction model may be trained together. Specifically, the first training sample is input into the initial face-swapping model, and the virtual face image sample is input into the initial virtual face reconstruction model. Feature alignment is performed based on the fused facial features generated during the training of the initial face-swapping model and the virtual facial features generated during the training of the initial virtual face reconstruction model. After a training stop condition is met, an intermediate face-swapping model and a target virtual face reconstruction model may be obtained through training. Finally, a target face-swapping model may be obtained based on the intermediate face-swapping model and the target virtual face reconstruction model. To clearly illustrate the present application, the first example mentioned above is used for explanation, but this is not intended to be limiting.
[0048] Regarding the second part, after obtaining the target face-swapping model, the acquired target real face image and target virtual face image can be face-swapped to obtain a face-swapped image whose texture and skin color more closely resemble those of the target virtual face image, thus avoiding any sense of incompatibility between the face-swapped image and the target image in the corresponding application scenario. In other words, face-swapping based on the target face-swapping model obtained above can improve the authenticity and image quality of the face-swapped image after migrating the real face image to the virtual face image.
[0049] The method mentioned in this application is illustrated below with reference to the accompanying drawings. Figure 1 This is a flow chart of a model determination method provided in an embodiment of the present application. Figure 1 As shown, the method may include:
[0050] S101: Construct a first training sample according to a training sample set.
[0051] The training sample set includes real face image samples and virtual face image samples, and the first training sample includes source image samples and target image samples.
[0052] As an exemplary example, in a game scenario, the real facial image sample can be an image including a real facial area, such as a player image, a celebrity image, etc., and the virtual facial image sample can be an image including a virtual facial area, such as a game character image, etc. It should be noted that this application does not limit real facial image samples and virtual facial image samples.
[0053] It is understood that the source image and target image in the first training sample have the following relationship, which allows the identity features in the source image to be transferred to the target image. The source image and target image in the first training sample can have the following relationship with the real facial image samples and virtual facial image samples in the training sample set: the source image and target image in the first training sample can both be real facial image samples in the training sample set, or the source image can be a real facial image sample in the training sample set, and the target image can be a virtual facial image sample in the training sample set. This application does not limit this.
[0054] S102: Train an initial face-changing model according to the first training sample to obtain an intermediate face-changing model.
[0055] The intermediate face-changing model is used to process the fused facial feature vector to obtain a predicted face-changing image. The fused facial feature vector includes the identity feature vector of the source image sample and the attribute feature vector of the target image sample.
[0056] Combine Figure 2 To explain, Figure 2 This is a structural diagram of an initial face-changing model provided in an embodiment of the present application, such as Figure 2 As shown, the initial face-changing model 200 includes an initial identity encoder 201, an initial attribute encoder 202, and an initial fusion unit 203. The initial identity encoder 201 and the initial attribute encoder 202 are respectively connected to the initial fusion unit 203. The initial identity encoder 201 is used to encode the input source image sample to obtain an identity feature vector. The initial attribute encoder 202 is used to encode the input target image sample to obtain an attribute feature vector. The initial fusion unit 203 is used to fuse the identity feature vector and the attribute feature vector to obtain a fused facial feature vector, and then obtain a predicted face-changing image by fusing the facial feature vector.
[0057] It can be understood that training the initial face-changing model 200 is essentially to correct the learning parameters in the initial identity encoder 201, the initial attribute encoder 202, the initial fuser 203, and the initial face-changing decoder 204 according to the preset loss function.
[0058] Using the first training sample to train the initial face-changing model can be expressed as:
[0059] res=Ghuman (E id (src), D(tgt))
[0060] Among them, G human represents the initial face-changing model, src, tgt, and res represent the source image sample, target image sample, and predicted face-changing image, respectively. id represents the identity encoder and D represents the attribute encoder.
[0061] The training process of the initial face-swapping model 200 can be carried out in the following supervised manner, specifically based on the identity similarity between the predicted face-swapping image res and the source image sample src and the attribute similarity between the predicted face-swapping image res and the target image sample src, wherein the identity similarity loss is defined as: L id =1-cos(E id (src), E id (res)), where cos represents the calculation of cosine similarity.
[0062] The attribute similarity loss is defined as: L attr =||D(tgt)-D(res)||2, where ||*||2 represents the Euclidean distance.
[0063] That is to say, the total loss L1 corresponding to the training initial face-changing model is equal to the identity similarity loss L id And the attribute similarity loss L attr Relatedly, when the total loss function L1 meets the preset training stop condition, the intermediate face-changing model can be trained.
[0064] Combine Figure 3 To explain, Figure 3 This is a schematic diagram of the structure of an intermediate face-changing model combined with an initial virtual face reconstruction model provided in an embodiment of the present application. Figure 3 As shown, the intermediate face-swapping model 300 includes an identity encoder 301, an attribute encoder 302, and a fuser 303. It is understood that the initial fuser 203 mentioned above corresponds to the fuser 303, that is, the fuser 303 is used to fuse the identity feature vector and the attribute feature vector to obtain a fused facial feature vector. Then, the intermediate face-swapping model 300 can obtain a predicted face-swapping image based on the fused facial feature vector.
[0065] S103: Input the virtual facial image samples in the training sample set into the initial virtual facial reconstruction model to obtain a virtual facial feature vector.
[0066] Continue to combine Figure 3To clearly describe the training process of the initial virtual facial reconstruction model 30, a single training iteration is used as the dimension for explanation. The initial virtual facial reconstruction model 30 includes an initial virtual facial reconstruction encoder 305, which is used to encode virtual facial image samples (such as game character image samples) in received virtual facial images to obtain virtual facial feature vectors. The initial virtual facial reconstruction model 30 can then reconstruct the virtual facial image based on the virtual facial feature vectors.
[0067] S104: aligning the virtual facial feature vector and the fused facial feature vector generated by the intermediate face-changing model, and correcting the initial virtual facial reconstruction decoder in the initial virtual facial reconstruction model according to the loss of the feature alignment to obtain a target virtual facial reconstruction model.
[0068] The target virtual face reconstruction model is used to process the virtual face feature vector to obtain a reconstructed virtual face image.
[0069] Continue to combine Figure 3 To explain, from Figure 3 As can be seen, the initial virtual face reconstruction encoder 305 is not only connected to the initial virtual face reconstruction decoder 306, but also to the output of the fuser 303 in the intermediate face-swapping model 300. Based on this connection, the constructed first training sample is again input into the intermediate face-swapping model 300. The fused facial feature vector output by the fuser 303 in the intermediate face-swapping model 300 and the virtual facial feature vector output by the initial virtual face reconstruction encoder 305 in the initial virtual face reconstruction model 300 are feature aligned to obtain the feature alignment loss. It is understandable that as the training process progresses, the initial virtual face reconstruction encoder 305 may output multiple virtual facial feature vectors, and the fuser 303 in the intermediate face-swapping model 300 will also output multiple fused facial feature vectors, i.e., features having both the virtual facial feature vector distribution and the fused facial feature vector distribution.
[0070] The aforementioned feature alignment loss can be represented by a decision value, which can be used to characterize the probability that the distribution of the virtual facial feature vectors is consistent with the distribution of the fused facial feature vectors. The process of changing this decision value is the process of correcting the learning parameters in the initial virtual facial reconstruction decoder 306 within the initial virtual facial reconstruction model 30. When the decision value changes to meet the training stop condition, training is continued to obtain the target virtual facial reconstruction model. As can be seen from the above description, the initial virtual facial reconstruction decoder 306 is used to process the virtual facial feature vectors to obtain a reconstructed virtual facial image. In other words, the target virtual facial reconstruction decoder is used to process the virtual facial feature vectors to obtain a reconstructed virtual facial image.
[0071] It is understood that the purpose of feature alignment mentioned here is to ensure that the distribution of the virtual facial feature vectors output by the initial virtual facial reconstruction encoder 305 in the initial virtual facial reconstruction model 30 is consistent with that of the fused facial feature vectors output by the fuser 303 in the intermediate face-swap model 300. The target virtual facial reconstruction model thus trained can, on the one hand, focus on the attribute features of the virtual facial image, such as texture and skin color, and on the other hand, can normally decode the fused facial feature vectors containing real facial image information to output the face-swap image.
[0072] S105: Obtain a target face-changing model based on the target virtual face reconstruction model and the intermediate face-changing model.
[0073] After training to obtain the target virtual face reconstruction model, the intermediate face-swapping model can be modified based on the target virtual face reconstruction model to obtain a modified intermediate face-swapping model. In one exemplary embodiment, the modified intermediate face-swapping model can be directly used as the target face-swapping model. In another exemplary embodiment, after obtaining the modified intermediate face-swapping model, the modified intermediate face-swapping model can be first trained using real facial image samples and virtual facial image samples in the training sample set as source image samples and target image samples, respectively. After the training stop condition is met, the target face-swapping model can be obtained.
[0074] In summary, in the model determination method provided by the present application, after the intermediate face-swapping model is obtained by training using the first training sample, the initial virtual face reconstruction model including the initial virtual face reconstruction decoder can be trained based on the intermediate face-swapping model. Since the initial virtual face reconstruction model is input with a virtual face image sample, during the training of the initial virtual face reconstruction model, the virtual face feature vector and the fused face feature vector generated by the intermediate face-swapping model can be feature aligned so that the distribution of the virtual face feature vector input to the initial virtual face reconstruction decoder is consistent with the distribution of the fused face feature vector. In this way, the target virtual face reconstruction model obtained by the final training can not only focus on the attribute features of the virtual face image, such as texture and skin color, but also can normally decode the fused face feature vector containing real face image information. In other words, the target face-swapping model finally determined based on the target virtual face reconstruction model and the intermediate face-swapping model can ensure that the style of the generated face-swapping image is consistent with the style of the virtual face image, avoiding the sense of disharmony that occurs after the real face image is transferred to the virtual face image, thereby improving the authenticity and image quality of the face-swapping image.
[0075] Optionally, the intermediate face-changing model includes an intermediate face-changing decoder; and the target virtual face reconstruction model includes a target virtual face reconstruction decoder.
[0076] Combine Figure 2 as well as Figure 3 To explain, from Figure 2 As can be seen from the figure, the initial face-changing model 200 also includes an initial face-changing decoder 204. The initial fusion unit 203 is connected to the initial face-changing decoder 204. The initial fusion unit 203 is used to fuse the identity feature vector and the attribute feature vector to obtain a fused facial feature vector. The initial face-changing decoder 204 can be used to decode the fused facial feature vector to obtain a predicted face-changing image. Figure 3 It can be seen that the intermediate face-changing model 300 also includes an intermediate face-changing decoder 304 corresponding to the initial face-changing decoder 204. It can be understood that the intermediate face-changing decoder 304 is used to decode the fused facial feature vector to obtain a predicted face-changing image.
[0077] According to the above description, the initial virtual face reconstruction model 30 includes an initial virtual face reconstruction encoder 305, which is connected to an initial virtual face reconstruction decoder 306. The initial virtual face reconstruction encoder 305 is used to decode the virtual face image samples (such as game character image samples) in the received virtual face image to obtain a virtual face feature vector, and the initial virtual face reconstruction decoder 306 is used to decode the virtual face feature vector to obtain a reconstructed virtual face image. Based on this, when the judgment value corresponding to the feature alignment loss changes to meet the training stop condition, the target virtual face reconstruction model is further trained, that is, the initial virtual face reconstruction decoder 306 becomes the target virtual face reconstruction decoder in the target virtual face reconstruction model.
[0078] Furthermore, the target face-changing model is obtained according to the target virtual face reconstruction model and the intermediate face-changing model, including: replacing the intermediate face-changing decoder in the intermediate face-changing model with the target virtual face reconstruction decoder to obtain the target face-changing model.
[0079] You can Figure 3 The intermediate face swap decoder 304 in the intermediate face swap model 300 is replaced by the target virtual face reconstruction decoder 401 in the target virtual face reconstruction model, as shown in FIG. Figure 4 As shown. In one exemplary embodiment, after the replacement, the intermediate face-swapping model after the replacement can be directly used as the target face-swapping model. In another exemplary embodiment, after the replacement, the intermediate face-swapping model after the replacement can be trained using real face image samples and virtual face image samples in the training sample set as source image samples and target image samples, respectively. After the training stop condition is met, the target face-swapping model is obtained.
[0080] like Figure 3As shown, the initial virtual face reconstruction model 30 includes an initial discriminator 307. The initial discriminator 307 inputs the fused facial feature vector output by the fuser 303 in the intermediate face swapping model 300 and the virtual facial feature vector output by the initial virtual face reconstruction encoder 305 in the initial virtual face reconstruction model 30, and then performs feature alignment on the virtual facial feature vector and the fused facial feature vector. Furthermore, the target virtual face reconstruction model includes a target virtual face reconstruction decoder.
[0081] Figure 5 This is a flow chart of another model determination method provided in the embodiment of the present application. Figure 5 As shown, optionally, the above-mentioned virtual facial feature vector and the fused facial feature vector generated by the intermediate face-changing model are feature-aligned, and the initial virtual facial reconstruction decoder in the initial virtual facial reconstruction model is corrected according to the loss of feature alignment to obtain the target virtual facial reconstruction model, including:
[0082] S501: Input the virtual facial feature vector and the fused facial feature vector into an initial discriminator, and the initial discriminator performs feature alignment processing to obtain a feature alignment loss.
[0083] In an exemplary embodiment, the virtual facial feature vector and the fused facial feature vector are input into an initial discriminator, and the initial discriminator obtains a feature alignment loss based on a difference between first distribution information of the fused facial feature vector and second distribution information of the virtual facial feature vector.
[0084] The initial discriminator can determine the distribution information of the feature vector. In an exemplary embodiment, the initial discriminator can use an identifier (e.g., 0 or 1) to represent the distribution of the fused facial feature vector and the distribution of the virtual facial feature vector. For example, if the distribution information of the fused facial feature vector is the first distribution information, then the first distribution information of the fused facial feature vector can be represented as 0. As long as the difference between the second distribution information of the virtual facial feature vector and the first distribution information does not meet the difference condition corresponding to the preset training stop condition, the initial discriminator can represent the distribution information of the virtual facial feature vector as 1. The initial discriminator obtains the loss of feature alignment based on the difference between the first distribution information and the second distribution information.
[0085] S502: Correcting an initial virtual face reconstruction decoder in the initial virtual face reconstruction model according to the loss of feature alignment, the virtual face image samples, the pixel loss corresponding to the reconstructed virtual face image, and the perceptual loss to obtain a target virtual face reconstruction model.
[0086] Based on the virtual facial image sample and the reconstructed virtual facial image, a pixel loss and a perceptual loss can be determined. Based on the pixel loss, the perceptual loss, and the feature alignment loss obtained above, the learning parameters of the initial virtual facial reconstruction encoder 305, the initial virtual facial reconstruction decoder 306, and the initial discriminator in the initial virtual facial reconstruction model 30 can be modified.
[0087] Assuming that the virtual face image sample is I and the reconstructed virtual face image is R, the pixel loss is L rec Defined as:
[0088] L rec =||RI||2
[0089] Among them, ||*||2 represents the Euclidean distance.
[0090] Perceptual loss L p Defined as: L p =||F(R)-F(I)||2, where F represents the initial virtual face reconstruction encoder, and the specific structure can be a VGG (Visual Geometry Group) network structure.
[0091] That is to say, the total loss L2 corresponding to the initial virtual face reconstruction model is not only related to the pixel loss L rec , Perceptual Loss L p The target virtual face reconstruction model is trained when the total loss L2 meets the preset stopping condition. The target virtual face reconstruction model includes a target virtual face reconstruction encoder, a target virtual face reconstruction decoder, and a target discriminator. When the target virtual face reconstruction decoder decodes the fused facial feature vector, it can ensure that the generated virtual face image has the corresponding style, such as a game-style face-swapped image.
[0092] The following example describes the relationship between the source image samples and the target image samples in the first training sample and the real face image samples and the virtual face image samples in the training sample set.
[0093] Optionally, constructing the first training sample based on the training sample set includes: constructing source image samples and target image samples in the first training sample based on real facial image samples in the training sample set, respectively, wherein the source image samples and the target image samples are different real facial image samples.
[0094] As can be seen from the above description, the training sample set includes both real facial image samples and virtual facial image samples. In one exemplary embodiment, two different real facial images in the training sample set can be combined into multiple real facial image sample groups. Multiple real facial image sample groups are obtained based on a preset number of training samples. Each real facial image sample group is then input into the initial face-swapping model for training, resulting in an intermediate face-swapping model. The first training sample is any real facial image sample group.
[0095] Combine Figure 2 as well as Figure 3 For example, any real face image sample group (first training sample) is used as an example for explanation. The real face image sample group includes real face image sample 1 and real face image sample 2. The real face image sample 1 (i.e., source image sample) is input into the initial identity encoder 201 in the initial face-changing model 200, and the real face image sample 2 (target image sample) is input into the initial attribute encoder 202 in the initial face-changing model 200 for encoding to train the initial face-changing model 200. When the training stop condition is met, an intermediate face-changing model is obtained through training. The intermediate face-changing model can be, for example, Figure 3 The specific structures of the intermediate face-changing model 300, the initial face-changing model 200, and the intermediate face-changing model 300 shown in the figure can be referred to the above-mentioned relevant parts and will not be described again here.
[0096] It is understandable that the number of real facial image samples in the training sample set far exceeds the number of virtual facial image samples. Using real facial image samples as source image samples and target image samples in the training samples for training the initial face-changing model can greatly improve the accuracy and robustness of the intermediate face-changing model obtained through training, and thus improve the accuracy and robustness of the target face-changing model obtained later.
[0097] Optionally, the relationship between the source and target image samples in the first training sample and the real and virtual facial image samples in the training sample set can be as follows. In one exemplary embodiment, the real facial image samples in the training sample set can be used as the source image samples in the first training sample, and the virtual facial image samples can be used as the target image samples in the first training sample. In other words, the initial face-swapping model can be trained using training samples constructed from the real and virtual facial image samples. When a training stop condition is met, an intermediate face-swapping model is obtained through training.
[0098] The following is a specific example of the step of replacing the intermediate face-swapping decoder in the intermediate face-swapping model with the target virtual face reconstruction decoder to obtain the target face-swapping model, when the source image sample and the target image sample in the first training sample are both real facial image samples in the training sample set.
[0099] Figure 6 This is a flow chart of another model determination method provided by this application. Figure 6 As shown, optionally, the intermediate face-swapping decoder in the intermediate face-swapping model is replaced with a target virtual face reconstruction decoder to obtain a target face-swapping model, including:
[0100] S601. Replace the intermediate face-swapping decoder in the intermediate face-swapping model with the target virtual face reconstruction decoder to obtain the replaced intermediate face-swapping model.
[0101] Combine Figure 2 as well as Figure 3 To illustrate, using the source image samples and target image samples that are both real face image samples Figure 2 The initial face-changing model 200 in is trained and the results are as follows Figure 3 After the intermediate face-changing model 300 is built, the initial virtual face reconstruction model 30 can be trained based on the intermediate face-changing model 300. After the training of the initial virtual face reconstruction model 30 is completed, the intermediate face-changing decoder 304 in the intermediate face-changing model 300 can be replaced by the target virtual face reconstruction decoder in the target virtual face reconstruction model to obtain the replaced intermediate face-changing model.
[0102] S602: Construct source image samples and target image samples in the second training sample according to the real face image samples and the virtual face image samples in the training sample set, respectively, using the real face image samples as the source image samples and the virtual face image samples as the target image samples.
[0103] S603: Input the second training sample into the replaced intermediate face-changing model to obtain a target face-changing model through training.
[0104] In one exemplary embodiment, the training sample set can be first divided into multiple real-virtual facial image sample groups, each of which includes a real facial image sample and a virtual facial image sample. Multiple real-virtual facial image sample groups are obtained based on a preset number of training samples. Each real-virtual facial image sample group is input into a replaced intermediate face-swapping model for training to obtain a target face-swapping model, wherein the second training sample is any real-virtual facial image sample group.
[0105] refer to Figure 4Assuming that a real-virtual facial image sample group includes a real facial image sample 1 and a virtual facial image sample 2, the real facial image sample 1 is input as the source facial image sample into the identity encoder 301 in the replaced intermediate face-changing model, and the virtual facial image sample 2 as the virtual facial image sample is input into the attribute encoder 302 in the replaced intermediate face-changing model for encoding to train the replaced intermediate face-changing model. When the training stop condition is met, the target face-changing model is obtained through training.
[0106] It can be understood that since the source image samples and the target image samples in the first training samples corresponding to the trained intermediate face-changing model are both real facial image samples, and in actual scenarios, face-changing is performed on real facial images and virtual facial images, after obtaining the replaced intermediate face-changing model, the second training samples of the source image samples as real facial image samples and the target image samples as virtual facial image samples are used to fine-tune the replaced intermediate face-changing model. In this way, the final target face-changing model will be suitable for actual application scenarios, thereby improving the accuracy of the target face-changing model.
[0107] Figure 7 This is a flow chart of another model determination method provided by this application. Figure 7 As shown, optionally, before constructing the first training sample according to the training sample set, the method further includes:
[0108] S701: Acquire a plurality of preset initial virtual facial image samples and a plurality of real facial image samples.
[0109] Among them, multiple real face image samples can be obtained from the real face image database, and the real face images include real face areas, such as player face areas, celebrity face areas, etc.
[0110] Taking a game scene as an example, a plurality of preset different initial game character image samples (ie, initial virtual face image samples), such as 30, 3D models corresponding to the samples may be obtained.
[0111] S702: Generate multiple virtual facial image samples according to the expression parameters and head posture parameters of each initial virtual facial image sample.
[0112] S703: Build a training sample set based on the multiple virtual facial image samples and the multiple real facial image samples.
[0113] Continuing with the above example, each 3D model has corresponding expression parameters and head posture parameters. These expression parameters and head posture parameters can be modified according to a preset modification strategy to generate multiple virtual facial image samples corresponding to each initial game character image sample. It should be noted that this application does not limit the number of virtual facial image samples. After obtaining multiple virtual facial image samples and multiple real facial image samples, a training sample set can be obtained.
[0114] In this way, multiple virtual facial image samples can be quickly obtained by simply obtaining multiple preset initial virtual facial image samples, which can improve the efficiency of obtaining the target face-changing model.
[0115] The following is an example of applying the target face-swapping model after obtaining the target face-swapping model.
[0116] Figure 8 This is a flow chart of an image processing method provided in an embodiment of the present application. Figure 8 As shown, the method may include:
[0117] S801: Acquire a target real face image and a target virtual face image.
[0118] The target real facial image and the target virtual facial image can be any picture containing facial information. The target real facial image can be a picture containing the facial information of a game player or a picture containing the facial information of another person. The target virtual facial image can be a picture containing the facial information of a game character. The target real facial image and the target virtual facial image can be pre-stored pictures or pictures captured by a camera device. The target real face image and the target virtual facial image can be a single picture or one of the continuous video frames containing facial information in the video data. It should be noted that this application does not limit them.
[0119] S802: Input the target real face image and the target virtual face image into the target face-changing model respectively to obtain a face-changing image.
[0120] Among them, the face-changing image includes the identity features of the target real face image and the attribute features of the target virtual face image. The training process of the target face-changing model can refer to the description of the relevant parts above.
[0121] Combine Figure 4To illustrate, the target real facial image is input into the identity encoder 301 in the target face-changing model, and the target virtual facial image is input into the attribute encoder 302 in the target face-changing model. The fusion device 303 fuses the identity feature vector output by the identity encoder 301 and the attribute feature vector output by the attribute encoder 302 to obtain a fused facial feature vector. The target virtual face reconstruction decoder 400 in the target face-changing model decodes the fused facial feature vector to obtain a face-changing image.
[0122] The facial style of the face-swapped image obtained in the above manner matches the facial style of the target virtual facial image. If the target virtual facial image is a game character image, the facial style of the face-swapped image is the game style, that is, the attribute characteristics of the face-swapped image are more matched with the attribute characteristics of the target virtual facial image.
[0123] Figure 9 This is a schematic diagram of the structure of a model determination device provided in an embodiment of the present application. Figure 9 As shown, the device includes:
[0124] The construction module 901 is used to construct a first training sample according to the training sample set, where the training sample set includes real face image samples and virtual face image samples, and the first training sample includes source image samples and target image samples.
[0125] A first determining module 902 is configured to train an initial face-swapping model based on the first training sample to obtain an intermediate face-swapping model, wherein the intermediate face-swapping model is configured to process a fused facial feature vector to obtain a predicted face-swapping image, wherein the fused facial feature vector includes an identity feature vector of a source image sample and an attribute feature vector of a target image sample;
[0126] A first input module 903 is configured to input a virtual facial image sample in a training sample set into an initial virtual facial reconstruction model to obtain a virtual facial feature vector;
[0127] A feature alignment module 904 is configured to perform feature alignment on the virtual facial feature vector and the fused facial feature vector generated by the intermediate face-swap model, and to modify the initial virtual facial reconstruction decoder in the initial virtual facial reconstruction model based on the loss of the feature alignment to obtain a target virtual facial reconstruction model. The target virtual facial reconstruction model is used to process the virtual facial feature vector to obtain a reconstructed virtual facial image.
[0128] The second determining module 905 is configured to obtain a target face-changing model based on the target virtual face reconstruction model and the intermediate face-changing model.
[0129] Optionally, the intermediate face-swapping model includes an intermediate face-swapping decoder; the target virtual face reconstruction model includes a target virtual face reconstruction decoder;
[0130] Correspondingly, the second determining module 905 is specifically configured to replace the intermediate face-changing decoder in the intermediate face-changing model with the target virtual face reconstruction decoder to obtain the target face-changing model.
[0131] Optionally, the target virtual face reconstruction model includes a target virtual face reconstruction decoder; the initial virtual face reconstruction model includes an initial discriminator;
[0132] Correspondingly, the feature alignment module 904 is specifically configured to input the virtual facial feature vector and the fused facial feature vector into the initial discriminator, which performs feature alignment processing to obtain a feature alignment loss; and to correct the initial virtual facial reconstruction decoder in the initial virtual facial reconstruction model based on the feature alignment loss, the virtual facial image samples, and the pixel loss and perceptual loss corresponding to the reconstructed virtual facial image to obtain a target virtual facial reconstruction model.
[0133] Optionally, the feature alignment module 904 is further specifically configured to input the virtual facial feature vector and the fused facial feature vector into an initial discriminator, and the initial discriminator obtains a feature alignment loss based on a difference between the first distribution information of the fused facial feature vector and the second distribution information of the virtual facial feature vector.
[0134] Optionally, the construction module 901 is specifically configured to construct source image samples and target image samples in the first training sample according to the real facial image samples in the training sample set, wherein the source image samples and the target image samples are different real facial image samples.
[0135] Optionally, the second determination module 905 is further specifically used to replace the intermediate face-changing decoder in the intermediate face-changing model with the target virtual face reconstruction decoder to obtain the replaced intermediate face-changing model; according to the real face image samples and the virtual face image samples in the training sample set, the source image samples and the target image samples in the second training samples are respectively constructed, and the real face image samples are used as the source image samples and the virtual face image samples are used as the target image samples; the second training samples are input into the replaced intermediate face-changing model to train and obtain the target face-changing model.
[0136] Optionally, the device further comprises: a building module;
[0137] The assembly module is used to obtain a plurality of preset initial virtual facial image samples and a plurality of real facial image samples; generate a plurality of virtual facial image samples based on the expression parameters and head posture parameters of each initial virtual facial image sample; and assemble a training sample set based on the plurality of virtual facial image samples and the plurality of real facial image samples.
[0138] Optionally, an embodiment of the present application further provides an image processing device, comprising:
[0139] An acquisition module, used to acquire a target real face image and a target virtual face image;
[0140] The second input module is used to input the target real facial image and the target virtual facial image into the target face-changing model respectively to obtain a face-changing image, which includes the identity features of the target real facial image and the attribute features of the target virtual facial image, wherein the target face-changing model is obtained by the model determination device mentioned in the above example.
[0141] The above-mentioned device is used to execute the method provided in the above-mentioned embodiment. Its implementation principle and technical effect are similar and will not be repeated here.
[0142] The above modules can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more microprocessors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, when a module is implemented by scheduling program code through a processing element, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0143] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 10 As shown, the electronic device may include: a processor 1001, a storage medium 1002, and a bus 1003. The storage medium 1002 stores machine-readable instructions executable by the processor 1001. When the electronic device is running, the processor 1001 communicates with the storage medium 1002 via the bus 1003. The processor 1001 executes the machine-readable instructions to perform the following steps:
[0144] In a feasible implementation scheme, when executing the model determination method, the processor 1001 is specifically used to: construct a first training sample based on a training sample set, the training sample set including real facial image samples and virtual facial image samples, and the first training sample including source image samples and target image samples; train an initial face-changing model based on the first training sample to obtain an intermediate face-changing model, the intermediate face-changing model is used to process the fused facial feature vector to obtain a predicted face-changing image, the fused facial feature vector including the identity feature vector of the source image sample and the attribute feature vector of the target image sample; input the virtual facial image sample in the training sample set into the initial virtual face reconstruction model to obtain a virtual facial feature vector; perform feature alignment on the virtual facial feature vector and the fused facial feature vector generated by the intermediate face-changing model, and correct the initial virtual face reconstruction decoder in the initial virtual face reconstruction model according to the loss of feature alignment to obtain a target virtual face reconstruction model, the target virtual face reconstruction model is used to process the virtual facial feature vector to obtain a reconstructed virtual facial image; obtain a target face-changing model based on the target virtual face reconstruction model and the intermediate face-changing model.
[0145] In a feasible implementation scheme, when executing the model determination method, the processor 1001 is specifically used to: replace the intermediate face-changing decoder in the intermediate face-changing model with the target virtual face reconstruction decoder to obtain the target face-changing model.
[0146] In a feasible implementation scheme, when executing the model determination method, the processor 1001 is specifically used to: input the virtual facial feature vector and the fused facial feature vector into the initial discriminator, and the initial discriminator performs feature alignment processing to obtain the feature alignment loss; according to the feature alignment loss, the virtual facial image sample and the pixel loss corresponding to the reconstructed virtual facial image, and the perceptual loss, the initial virtual facial reconstruction decoder in the initial virtual facial reconstruction model is corrected to obtain the target virtual facial reconstruction model.
[0147] In a feasible implementation scheme, when executing the model determination method, the processor 1001 is specifically configured to: input the virtual facial feature vector and the fused facial feature vector into an initial discriminator, and the initial discriminator obtains the feature alignment loss based on the difference between the first distribution information of the fused facial feature vector and the second distribution information of the virtual facial feature vector.
[0148] In a feasible implementation scheme, when executing the model determination method, the processor 1001 is specifically used to: construct the source image samples and target image samples in the first training sample according to the real facial image samples in the training sample set, and the source image samples and the target image samples are different real facial image samples.
[0149] In a feasible implementation scheme, when executing the model determination method, the processor 1001 is specifically used to: replace the intermediate face-changing decoder in the intermediate face-changing model with the target virtual face reconstruction decoder to obtain the replaced intermediate face-changing model; construct the source image samples and the target image samples in the second training sample according to the real face image samples and the virtual face image samples in the training sample set, respectively, use the real face image samples as the source image samples, and use the virtual face image samples as the target image samples; input the second training sample into the replaced intermediate face-changing model, and train to obtain the target face-changing model.
[0150] In a feasible implementation scheme, when executing the model determination method, the processor 1001 is specifically used to: obtain a preset plurality of initial virtual facial image samples and a plurality of real facial image samples; generate a plurality of virtual facial image samples based on the expression parameters and head posture parameters of each initial virtual facial image sample; and form a training sample set based on the plurality of virtual facial image samples and the plurality of real facial image samples.
[0151] In a feasible implementation scheme, when executing the image processing method, the processor 1001 is specifically used to obtain a target real facial image and a target virtual facial image; the target real facial image and the target virtual facial image are respectively input into the target face-changing model to obtain a face-changing image, which includes the identity features of the target real facial image and the attribute features of the target virtual facial image, wherein the target face-changing model is obtained by the above-mentioned model determination method.
[0152] Optionally, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor performs the following steps:
[0153] In a feasible implementation scheme, when executing the model determination method, the processor is specifically used to: construct a first training sample based on a training sample set, the training sample set including real facial image samples and virtual facial image samples, and the first training sample including source image samples and target image samples; train an initial face-changing model based on the first training sample to obtain an intermediate face-changing model, the intermediate face-changing model is used to process the fused facial feature vector to obtain a predicted face-changing image, the fused facial feature vector including the identity feature vector of the source image sample and the attribute feature vector of the target image sample; input the virtual facial image sample in the training sample set into the initial virtual face reconstruction model to obtain a virtual facial feature vector; perform feature alignment on the virtual facial feature vector and the fused facial feature vector generated by the intermediate face-changing model, and correct the initial virtual face reconstruction decoder in the initial virtual face reconstruction model according to the loss of feature alignment to obtain a target virtual face reconstruction model, the target virtual face reconstruction model is used to process the virtual facial feature vector to obtain a reconstructed virtual facial image; obtain a target face-changing model based on the target virtual face reconstruction model and the intermediate face-changing model.
[0154] In a feasible implementation scheme, when executing the model determination method, the processor is specifically used to: replace the intermediate face-changing decoder in the intermediate face-changing model with the target virtual face reconstruction decoder to obtain the target face-changing model.
[0155] In a feasible implementation scheme, when executing the model determination method, the processor is specifically used to: input the virtual facial feature vector and the fused facial feature vector into the initial discriminator, and the initial discriminator performs feature alignment processing to obtain the feature alignment loss; according to the feature alignment loss, the virtual facial image sample and the pixel loss corresponding to the reconstructed virtual facial image, and the perceptual loss, correct the initial virtual facial reconstruction decoder in the initial virtual facial reconstruction model to obtain the target virtual facial reconstruction model.
[0156] In one feasible implementation, when executing the model determination method, the processor is specifically configured to: input the virtual facial feature vector and the fused facial feature vector into an initial discriminator, and the initial discriminator obtains a feature alignment loss based on a difference between first distribution information of the fused facial feature vector and second distribution information of the virtual facial feature vector.
[0157] In a feasible implementation scheme, when executing the model determination method, the processor is specifically used to: construct source image samples and target image samples in the first training sample based on real facial image samples in the training sample set, where the source image samples and target image samples are different real facial image samples.
[0158] In a feasible implementation scheme, when the processor executes the model determination method, it is specifically used to: replace the intermediate face-changing decoder in the intermediate face-changing model with the target virtual face reconstruction decoder to obtain the replaced intermediate face-changing model; construct the source image samples and target image samples in the second training sample according to the real face image samples and the virtual face image samples in the training sample set, respectively, use the real face image samples as the source image samples, and use the virtual face image samples as the target image samples; input the second training sample into the replaced intermediate face-changing model, and train to obtain the target face-changing model.
[0159] In a feasible implementation scheme, when executing the model determination method, the processor is specifically used to: obtain a preset plurality of initial virtual facial image samples and a plurality of real facial image samples; generate a plurality of virtual facial image samples based on the expression parameters and head posture parameters of each initial virtual facial image sample; and form a training sample set based on the plurality of virtual facial image samples and the plurality of real facial image samples.
[0160] In a feasible implementation scheme, when executing the image processing method, the processor is specifically used to obtain a target real facial image and a target virtual facial image; the target real facial image and the target virtual facial image are respectively input into the target face-changing model to obtain a face-changing image, which includes the identity features of the target real facial image and the attribute features of the target virtual facial image, wherein the target face-changing model is obtained by the above-mentioned model determination method.
[0161] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0162] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0163] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0164] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor (English: processor) to perform some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (English: Read-Only Memory, abbreviated: ROM), a random access memory (English: Random Access Memory, abbreviated: RAM), a disk or an optical disk, and other media that can store program code.
[0165] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0166] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application. It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A model determination method, characterized in that: The method comprises: Constructing a first training sample according to a training sample set, wherein the training sample set includes real face image samples and virtual face image samples, and the first training sample includes a source image sample and a target image sample; training an initial face-swapping model based on the first training sample to obtain an intermediate face-swapping model, wherein the intermediate face-swapping model is used to process a fused facial feature vector to obtain a predicted face-swapping image, wherein the fused facial feature vector includes an identity feature vector of a source image sample and an attribute feature vector of a target image sample; Inputting the virtual facial image samples in the training sample set into the initial virtual facial reconstruction model to obtain a virtual facial feature vector; performing feature alignment on the virtual facial feature vector and the fused facial feature vector generated by the intermediate face-swap model, and correcting an initial virtual facial reconstruction decoder in the initial virtual facial reconstruction model based on a loss in the feature alignment to obtain a target virtual facial reconstruction model, wherein the target virtual facial reconstruction model is used to process the virtual facial feature vector to obtain a reconstructed virtual facial image; Obtaining a target face-changing model according to the target virtual face reconstruction model and the intermediate face-changing model; The intermediate face-swapping model includes an intermediate face-swapping decoder; the target virtual face reconstruction model includes a target virtual face reconstruction decoder; The step of obtaining a target face-changing model based on the target virtual face reconstruction model and the intermediate face-changing model comprises: The intermediate face-swapping decoder in the intermediate face-swapping model is replaced with the target virtual face reconstruction decoder to obtain a target face-swapping model.
2. The method according to claim 1, characterized in that The target virtual face reconstruction model includes a target virtual face reconstruction decoder; the initial virtual face reconstruction model includes an initial discriminator; the virtual face feature vector and the fused face feature vector generated by the intermediate face-changing model are feature-aligned, and the initial virtual face reconstruction decoder in the initial virtual face reconstruction model is corrected according to the loss of the feature alignment to obtain the target virtual face reconstruction model, including: Inputting the virtual facial feature vector and the fused facial feature vector into the initial discriminator, and performing feature alignment processing by the initial discriminator to obtain a feature alignment loss; An initial virtual face reconstruction decoder in the initial virtual face reconstruction model is corrected according to the feature alignment loss, the virtual face image sample, and the pixel loss and perceptual loss corresponding to the reconstructed virtual face image to obtain the target virtual face reconstruction model.
3. The method according to claim 2, characterized in that The virtual facial feature vector and the fused facial feature vector are input into the initial discriminator, and the initial discriminator performs feature alignment processing, The loss of feature alignment is obtained, including: The virtual facial feature vector and the fused facial feature vector are input into the initial discriminator, and the initial discriminator obtains the feature alignment loss according to the difference between the first distribution information of the fused facial feature vector and the second distribution information of the virtual facial feature vector.
4. The method according to claim 1, wherein The step of constructing a first training sample according to the training sample set includes: According to the real facial image samples in the training sample set, source image samples and target image samples in the first training samples are respectively constructed, and the source image samples and the target image samples are different real facial image samples.
5. The method according to claim 4, characterized in that The step of replacing the intermediate face-swapping decoder in the intermediate face-swapping model with the target virtual face reconstruction decoder to obtain the target face-swapping model comprises: Replacing the intermediate face-swapping decoder in the intermediate face-swapping model with the target virtual face reconstruction decoder to obtain a replaced intermediate face-swapping model; constructing source image samples and target image samples in the second training sample according to the real face image samples and the virtual face image samples in the training sample set, respectively, using the real face image samples as the source image samples and the virtual face image samples as the target image samples; The second training sample is input into the replaced intermediate face-changing model to obtain the target face-changing model through training.
6. The method according to any one of claims 1 to 5, characterized in that Before constructing the first training sample according to the training sample set, the method further includes: Acquire a plurality of preset initial virtual facial image samples and a plurality of real facial image samples; generating a plurality of virtual facial image samples according to the expression parameters and head posture parameters of each of the initial virtual facial image samples; The training sample set is constructed according to the multiple virtual facial image samples and the multiple real facial image samples.
7. An image processing method, characterized in that: The method comprises: Obtaining a target real face image and a target virtual face image; The target real facial image and the target virtual facial image are respectively input into a target face-changing model to obtain a face-changing image, wherein the face-changing image includes the identity features of the target real facial image and the attribute features of the target virtual facial image, wherein the target face-changing model is obtained by the model determination method described in any one of claims 1 to 6 above.
8. A model determination device, characterized in that: The device comprises: A construction module, configured to construct a first training sample based on a training sample set, wherein the training sample set includes real facial image samples and virtual facial image samples, and the first training sample includes a source image sample and a target image sample; a first determination module, configured to train an initial face-swapping model based on the first training sample to obtain an intermediate face-swapping model, wherein the intermediate face-swapping model is configured to process a fused facial feature vector to obtain a predicted face-swapping image, wherein the fused facial feature vector includes an identity feature vector of a source image sample and an attribute feature vector of a target image sample; a first input module, configured to input the virtual facial image samples in the training sample set into the initial virtual facial reconstruction model to obtain a virtual facial feature vector; a feature alignment module, configured to perform feature alignment on the virtual facial feature vector and the fused facial feature vector generated by the intermediate face-swap model, and to correct an initial virtual facial reconstruction decoder in the initial virtual facial reconstruction model based on a loss in the feature alignment to obtain a target virtual facial reconstruction model, wherein the target virtual facial reconstruction model is used to process the virtual facial feature vector to obtain a reconstructed virtual facial image; a second determining module, configured to obtain a target face-changing model based on the target virtual face reconstruction model and the intermediate face-changing model; The intermediate face-swapping model includes an intermediate face-swapping decoder; the target virtual face reconstruction model includes a target virtual face reconstruction decoder; The second determination module is specifically configured to replace the intermediate face-changing decoder in the intermediate face-changing model with the target virtual face reconstruction decoder to obtain a target face-changing model.
9. An image processing device, characterized in that: The device comprises: An acquisition module, used to acquire a target real face image and a target virtual face image; The second input module is used to input the target real facial image and the target virtual facial image into the target face-changing model respectively to obtain a face-changing image, wherein the face-changing image includes the identity features of the target real facial image and the attribute features of the target virtual facial image, wherein the target face-changing model is obtained by the model determination device described in claim 8 above.
10. An electronic device, characterized in that: include: A processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the storage medium communicate via the bus, and the processor executes the machine-readable instructions to perform the steps of the model determination method according to any one of claims 1 to 6 or the steps of the image processing method according to claim 7.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the model determination method according to any one of claims 1 to 6 or the steps of the image processing method according to claim 7.
Citation Information
Patent Citations
Generative adversarial network training method, image face changing and video face changing method and device
CN111783603A