An Image Face-Swapping Method Based on Feature Marking Training Strategy
By feature marking the source image and target image during the training stage of the generative adversarial network and using triangle mark supervision, the problem of the generative adversarial network paying too much attention to the appearance information of the target image is solved, and the quality and reliability of the face-changing image are improved.
Patent Information
- Application Number
- CN202210436120.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-22
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-04-22
AI Technical Summary
The existing generative adversarial networks tend to pay too much attention to the appearance information of the target image during training, resulting in the introduction of the appearance information of the target image during face change, interfering with the appearance information of the source image, and reducing the quality and reliability of the face change.
A feature marking training strategy is adopted, by feature marking the source image and the target image in the training stage, and using triangle markings to supervise the key points, the generative adversarial network is guided to obtain appearance information from the source image, reducing the appearance information interference of the target image, and a loss function is used to train it in combination with generation adversarial loss and mark L1 loss.
The quality and reliability of the generative adversarial network during the face change process is improved, ensuring that the appearance information of the source image is correctly expressed in the face change results, and reducing the appearance information interference of the target image.
Smart Images

Figure CN115019223B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more specifically, to an image face swapping method based on a feature marking training strategy. Background Art
[0002] With the popularity of video applications represented by live streaming and short videos, computer vision technology has been increasingly applied. Among them, AI face swapping is an important technology and is used in fields such as photo editing, film and television production, and AI digital human production. Currently, a generative adversarial network is generally used to implement the face swapping function. The network accepts the input of two images, the source image Xs and the target image Xt, and outputs a face swapped image Y with the appearance of the person in the source image Xs and the expression and facial movements of the target image Xt (the background is from the source image Xs).
[0003] In the training of generative adversarial networks, in order to generate clear and realistic images, it is generally necessary to use pictures to supervise the generated images. When the portraits in the source image Xs and the target image Xt are not of the same person, it is generally impossible to obtain a real image with the appearance of the person in the source image Xs and the expression and action of the target image Xt, so the training cannot proceed. Therefore, in the training of the network, usually two frames are intercepted from a video containing a single-person portrait. One frame is used as the source image Xs, and the other frame is used as the target image Xt, and the portraits in the two frames are of the same person. In this way, the target image Xt contains both the appearance information and the action and expression information of the person. Therefore, only the target image Xt needs to be used for supervision. Therefore, the loss functions used in the training are generally the generative adversarial loss, the reconstruction loss of the face-swapped image Y to the target image Xt, and the local losses of some sub-modules. However, when the portraits in the source image Xs and the target image Xt are of the same person, the above loss functions cannot constrain whether the appearance information of the network is obtained from the source image Xs or the target image Xt. Even when the generative adversarial network ignores the source image Xs and only obtains the appearance information and the action and expression information from the target image Xt for reconstruction, the above loss functions can also be reduced to a very low level. Although some networks have made distinctions between the source image Xs and the target image Xt in terms of model design. For example, in the paper "One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing", it designs a total of four feature extraction modules for the source image Xs, namely an appearance extraction module, a neutral expression 3D keypoint extraction module, a head pose detection module, and an expression offset module, while only designs two feature extraction modules, namely a head pose detection module and an expression offset module, for the target image Xt. However, when training with pictures of different portraits, it can be seen in the experiment that the images generated by the model are obviously marked with the appearance features of the target image Xt. Therefore, even if different feature extraction paths are designed for the source image Xs and the target image Xt, and only a small number of features are extracted from the target image Xt, the convolutional neural network still has the ability to introduce the appearance information of the target image Xt into the face-swapped result. In the current conventional training method, regardless of the source image Xs, the generated face-swapped image Y is only identical to the target image Xt, which is likely to make the generative adversarial network overly focus on the target image Xt. When the generative adversarial network overly focuses on the target image Xt, the above phenomenon of introducing the appearance information of the target image into the face-swapped result will be more likely to occur.
[0004] Due to the above situation, the network training process may introduce the appearance information in the target image Xt into the face-swapped image Y. During the training process, since the source image Xs and the target image Xt are of the same person, it will not be manifested. However, in the inference application stage, when using the source image Xs and the target image Xt of different people for face swapping, the appearance information of the target image Xt introduced by the network will interfere with the appearance information of the source image Xs, reducing the quality of face swapping; even during the training process, the network only learns to reconstruct by obtaining information from the target image and completely ignores the appearance information of the source image Xs in the inference stage, failing to achieve the face-swapping effect at all.
[0005] The reason for the above problems lies in the differences between the network training method and the inference application method: During training, due to the need for supervision, the people in the source image Xs and the target image Xt input to the network can only be the same person, making it difficult to define the source of appearance information, and the face-swapped result Y can only be exactly the same as the target image Xt, resulting in the network paying too much attention to the target image Xt and being more likely to introduce its appearance information; while in the inference application stage, the appearance information can only come from the source image Xs. After careful analysis, it can be found that under the existing training method, no matter how the source image Xs changes, one target image Xt will only correspond to one face-swapped image, and changing the source image Xs will not change the face-swapped result, which will make the model pay too much attention to the target image Xt, thus making it easier for the model to introduce the appearance information of the target image Xt. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides an image face-swapping method based on a feature marking training strategy.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0008] An image face-swapping method based on a feature marking training strategy includes the following steps:
[0009] S1: In the training stage, randomly intercept two frames of images from a video containing only single-person headshots, arbitrarily select one frame as the source image Xs and the other frame as the target image Xt, and perform standardization processing on both of them.
[0010] S2: Use a face detection network to detect and crop the faces in the two images.
[0011] S3: Use a face key point detection network to detect the face key points in the cropped images.
[0012] S4: Calculate the yaw angle among the three Euler angles representing the face orientation according to the face key points of the two images detected in step S3.
[0013] S5: According to the yaw angle calculated in step S4, select a side face with more exposed parts in the two images, and try to avoid the side face being blocked by the nose;
[0014] S6: Arbitrarily select three key points from the set of human face key points of the four parts of the cheek contour line, the lower edge of the eyes, the upper edge of the outer lip, and the nose of the selected side face, and arbitrarily select a color to paint triangular marks of the same color at the corresponding key points of the source image Xs and the target image Xt to obtain the marked source image Xs' and the marked target image Xt';
[0015] S7: In each iteration of training, input the source image Xs and the target image Xt into the generative adversarial network to generate a face-swapped image Y, and use the target image Xt for supervision; input the source image Xs, the marked source image Xs' and the target image Xt into the generative adversarial network to generate a marked face-swapped image Y', and use the marked target image Xt' for supervision, and stop training after the loss function drops to the set threshold;
[0016] S8: In the inference application stage, input a source image Xt containing the head portrait of person A and a driving video Vd containing the head portrait of person B, then a generated video Vo containing the appearance of person A and having the expressions and facial movements of person B can be obtained.
[0017] Preferably, the specific operation of the normalization process in step S1 includes:
[0018] During model training, select a video from the video dataset containing single-person head portraits, arbitrarily intercept two frames from it, one frame as the source image Xs, and the other frame as the target image Xt. Perform normalization processing on the two images and scale them to a fixed size.
[0019] Preferably, the specific operation of step S3 includes: sequentially input the cropped source image Xs and target image Xt into the face key point network RFID to detect 68 face key points in the two images.
[0020] Preferably, the specific operation of step S4 includes: select 10 key points from the 68 face key points and the corresponding position key point coordinates of the preset 3D face model, input them into the solvePnP function of OpenCV to calculate the rotation vector of each image's face, and finally convert the rotation vector into Euler angles.
[0021] Preferably, the specific operation of step S5 includes: add the yaw angles of the faces of the source image Xs and the target image Xt, if it is a positive value, select the right face, otherwise select the left face.
[0022] Preferably, the specific operation of step S6 includes: making the face key points of four parts, namely the cheek contour lines of the left and right side faces, the lower edges of the eyes, the upper edges of the outer sides of the lips, and the nose, into two sets. According to the side face selected in step S5, three key points are selected from the corresponding set, and any one color is selected to paint the same random color in the triangular area formed by connecting the corresponding three key points in the two images, obtaining the marked source image Xs’ and the marked target image Xt.
[0023] Preferably, the network model is included in step S7. The network model is composed of an appearance extraction module, a neutral expression three-dimensional key point extraction module, a head pose detection module, an expression offset module, and a video generation module. In each training iteration, four features are extracted from the source image Xs by the appearance extraction module, the neutral expression three-dimensional key point extraction module, the head pose detection module, and the expression offset module, and two features, namely the head pose and the expression offset, are extracted from the target image Xt by the head pose detection module and the expression offset module. The six extracted features are input into the video generation module to generate a face-swapped image Y, and it is supervised by the target image Xt. Then, the appearance feature in the six features is replaced with the marked appearance feature extracted from the marked source image Xs’ with triangular marks by the appearance extraction module, and the recombined six features are input into the image generation module to generate a marked face-swapped image Y’ with triangular marks, and it is supervised by the marked target image Xt’ with triangular marks. Stop when the loss function decreases to the set threshold.
[0024] The expression of the loss function is:
[0025] Loss=W1×LGan+W2×LmarkL1
[0026] Where LGan is the loss function of the original face-swapping generative adversarial model, LmarkL1 is the L1 loss between the marked face-swapped image Y’ and the marked target image Xt’, and W1 and W2 are weight coefficients for balancing the importance of different loss functions.
[0027] Preferably, the 10 key points are respectively the four corner key points of the two eyes, the key points at the left and right ends of the two eyebrows, and the key points on the left and right sides of the nose wings.
[0028] Compared with the prior art, the beneficial effects of the present invention are:
[0029] The present invention can break the training mode in which one target image Xt corresponds to one face-swapped image Y commonly used in the face-swapping generative adversarial network, guide the network to obtain appearance information from the source image Xs during the training stage, reduce the interference of the appearance information in the target image Xt, and thus improve the face-swapping quality and reliability of the network model. Description of the Drawings
[0030] Figure 1 Flow chart for training the generative adversarial network of the present invention;
[0031] Figure 2 Structure diagram of the face swapping image generation network of the present invention;
[0032] Figure 3 Structure diagram of the marked face swapping image generation network of the present invention. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] As Figures 1-3 shown, S1: In the training stage, arbitrarily intercept two frames of images from a video containing only a single-person head portrait, select one frame as the source image Xs and the other frame as the target image Xt, perform normalization processing on both of them, and scale them to a fixed size;
[0035] S2: Input the source image Xs and the target image Xt into the face detection network RetinaFace in sequence to obtain the coordinates of the face detection boxes in the two images, expand the face detection boxes by a certain proportion, and crop the source image Xs and the target image Xt;
[0036] S3: Input the cropped source image Xs and the target image Xt into the face key point network RFID in sequence to detect 68 face key points in the two images;
[0037] S4: Select 10 key points from the 68 face key points (respectively the four corner key points of the two eyes, the key points at the left and right ends of the two eyebrows, and the key points on the left and right sides of the nose wings) and the corresponding position key point coordinates of the preset 3D face model, input them into the solvePnP function of OpenCV to calculate the rotation vectors of the faces in each image, and finally convert the rotation vectors into Euler angles;
[0038] S5: Select one side of the side face that shows more relative parts in the two images according to the yaw angle, and try to avoid the side face being blocked by the nose. Add the yaw angles of the faces of the source image Xs and the target image Xt. If it is a positive value, select the right side face, otherwise select the left side face;
[0039] S6: To minimize the occlusion of facial features by the triangle markings, two sets of facial key points are created from the following four parts: the cheek contour lines on the left and right side faces, the lower edges of the eyes, the upper outer edges of the lips, and the nose. Based on the selected side face, three key points are selected from the corresponding set, and a random color is chosen. The same color is then applied to the triangular regions formed by connecting the corresponding three key points in both images, thus applying the same triangle marking to both images, resulting in the marked source image Xs’ and the marked target image Xt’.
[0040] S7: The network model consists of four feature extraction modules, namely, the appearance extraction module, the neutral expression 3D key point extraction module, the head pose detection module, and the expression offset module, and a video generation module. In each training iteration, the four feature extraction modules, namely, the appearance extraction module, the neutral expression 3D key point extraction module, the head pose detection module, and the expression offset module, are used to extract four features from the source image Xs. The head pose detection module and the expression offset module are used to extract two features, namely, the head pose and the expression offset, from the target image Xt. The above six features are input into the image generation module to generate the face-swapped image Y, which is supervised by the target image Xt. Then, the appearance feature among the above six features is replaced with the marked appearance feature extracted from the marked source image Xs’ with the triangle marking by the appearance extraction module. The re-combined six features are input into the image generation module to generate the marked face-swapped image Y’ with the triangle marking, which is supervised by the marked target image Xt’ with the triangle marking. Training stops when the loss function drops to the set threshold.
[0041] The expression of the loss function is:
[0042] Loss = W1 × LGan + W2 × LmarkL1,
[0043] where LGan is the loss function of the original face-swapping generative adversarial model, and LmarkL1 is the L1 loss between the marked face-swapped image Y’ and the marked target image Xt’. W1 and W2 are weight coefficients for balancing the importance of different loss functions, which can be set according to the actual situation during application.
[0044] The video generation model consists of four feature extraction modules, namely an appearance extraction module, a neutral expression 3D key point extraction module, a head pose detection module, and an expression deviation module, and a video generation module. The model training process is as follows: In each iteration, the dataset provides a source image Xs and a target image Xt. According to the data augmentation strategy, the labeled source image Xs’ and the labeled target image Xt’ are generated. Four features are extracted from the source image Xs by the four feature extraction modules, namely the appearance extraction module, the neutral expression 3D key point extraction module, the head pose detection module, and the expression deviation module. Two features, namely the head pose and the expression deviation, are extracted from the target image Xt by the head pose detection module and the expression deviation module. The above six features are input into the image generation module to generate a face-swapped image Y, which is supervised by the target image Xt. Then, the appearance feature in the above six features is replaced with the labeled appearance feature extracted from the labeled source image by the appearance extraction module, and the re-combined six features are input into the image generation module to generate a labeled face-swapped image Y’, which is supervised by the labeled target image Xt’.
[0045] The above only elaborates on the preferred embodiments of the present invention in detail. However, the present invention is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the purpose of the present invention, and all such changes should be included within the protection scope of the present invention.
Claims
1. An image face swapping method based on a feature marking training strategy, characterized in that: It includes the following steps: S1: In the training phase, arbitrarily intercept two frames of images from a video containing only a single-person head. Optionally select one frame as the source image Xs and the other frame as the target image Xt, and perform normalization processing on both of them. S2: Use a face detection network to perform face detection and cropping on the two images. S3: Use a face key point detection network to detect face key points on the cropped images. S4: Calculate the yaw angle among the three Euler angles representing the face orientation based on the face key points of the two images detected in step S3. S5: According to the yaw angle calculated in step S4, select a side face with more exposed parts relative to the two images, and try to avoid the side face being blocked by the nose. S6: Arbitrarily select three key points from the set of face key points of the four parts: the cheek contour line, the lower edge of the eyes, the upper edge of the outer lip, and the nose of the selected side face. Arbitrarily select a color and apply triangular marks of the same color to the corresponding key points of the source image Xs and the target image Xt to obtain the marked source image Xs' and the marked target image Xt'. S7: In each iteration of training, input the source image Xs and the target image Xt into the generative adversarial network to generate a face-swapped image Y, and supervise it with the target image Xt; input the source image Xs, the marked source image Xs' and the target image Xt into the generative adversarial network to generate a marked face-swapped image Y', and supervise it with the marked target image Xt'. Stop training after the loss function drops to the set threshold. S8: In the inference application phase, input a source image Xt containing the head of person A and a driving video Vd containing the head of person B, then a generated video Vo containing the appearance of person A and the expressions and facial movements of person B can be obtained. The step S7 includes a network model, which is composed of an appearance extraction module, a neutral expression three-dimensional key point extraction module, a head pose detection module, an expression offset module, and a video generation module; in each training iteration, use the appearance extraction module, the neutral expression three-dimensional key point extraction module, the head pose detection module, and the expression offset module to extract four features from the source image Xs, and use the head pose detection module and the expression offset module to extract two features of the head pose and expression offset from the target image Xt; input the six extracted features into the video generation module to generate a face-swapped image Y, and supervise it with the target image Xt; then replace the appearance feature among the six features with the marked appearance feature extracted from the marked source image Xs' with triangular marks by the appearance extraction module, input the recombined six features into the image generation module to generate a marked face-swapped image Y' with triangular marks, and supervise it with the marked target image Xt' with triangular marks. Stop after the loss function drops to the set threshold. The expression of the loss function is: Loss = W1 × LGan + W2 × LmarkL1 Where LGan is the loss function of the original face-swapping generative adversarial model, LmarkL1 is the L1 loss between the marked face-swapped image Y' and the marked target image Xt', and W1, W2 are weight coefficients for balancing the importance of different loss functions.
2. The image face swapping method based on a feature marking training strategy according to claim 1, wherein: The specific operations of the normalization process in step S1 are as follows: During model training, select a video from the video dataset containing single-person head images, arbitrarily intercept two frames from it, one frame as the source image Xs, and the other frame as the target image Xt. Normalize the two images and scale them to a fixed size.
3. A face swapping method based on a feature marking training strategy according to claim 1, characterized in that: The specific operations of step S3 are as follows: Sequentially input the cropped source image Xs and target image Xt into the face key point network RFID to detect 68 face key points in the two images.
4. A face-swapping method based on a feature marking training strategy according to claim 1, characterized in that: The specific operations of step S4 are as follows: Select 10 key points from the 68 face key points and input the key point coordinates corresponding to the preset 3D face model into the solvePnP function of OpenCV to calculate the rotation vector of the face in each image, and finally convert the rotation vector into Euler angles.
5. A face-swapping method based on a feature marking training strategy according to claim 1, characterized in that: The specific operations of step S5 are as follows: Add the face yaw angles of the source image Xs and the target image Xt. If the result is positive, select the right face; otherwise, select the left face.
6. The image face swapping method based on a feature marking training strategy according to claim 1, wherein: The specific operations of step S6 are as follows: Make two sets of the face key points of the four parts, namely the cheek contour line, the lower edge of the eyes, the upper edge of the outer lip, and the nose, of the left and right side faces. According to the side face selected in step S5, select 3 key points from the corresponding set, choose any color, and paint the same random color in the triangular area formed by connecting the corresponding three key points in the two images to obtain the marked source image Xs’ and the marked target image Xt.
7. A face swapping method based on a feature marking training strategy according to claim 4, characterized in that: The 10 key points are respectively the four corner key points of the two eyes, the key points at the left and right ends of the two eyebrows, and the key points on the left and right sides of the nose wings.
Citation Information
Patent Citations
Real-time facial expression migration method based on generative adversarial
CN113343761A
High-fidelity face privacy protection method and system based on generative adversarial network
CN113343878A