Face recognition face changing method and system based on deep learning

Through the facial recognition face swap method based on deep learning, using technologies such as face key point detection, deep feature extraction and adversarial networks, the problems of low efficiency, low degree of automation and privacy leakage are solved, and efficient, automated, natural and safe face swap effect is achieved.

CN120047984AInactive Publication Date: 2025-05-27SHANDONG GUANG TELECOMM NETWORK OPERATION CO LTD

Patent Information

Application Number
CN202510114683.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional face swap technology has low efficiency, low degree of automation, high professionalism, unnatural effects and risks of privacy leakage.

Method used

The face recognition and face change method based on deep learning is adopted, and face change images are generated through face key point detection, deep feature extraction, geometric alignment, adversarial network, attention mechanism refinement and face segmentation model fusion to generate natural and realistic face change effects.

Benefits of technology

It significantly improves the efficiency and automation of face change, lowers the technical threshold, and the generated face change images are natural, realistic and safe, avoiding the risk of privacy leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047984A_ABST
    Figure CN120047984A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a face recognition face changing method and system based on deep learning. The method comprises the steps of obtaining a source image and a target image input by a user; performing face key point detection on the source image and the target image; performing deep feature extraction on the source image and the target image after key point detection based on a deep convolutional network; carrying out geometric alignment on the source image and the target image after depth feature extraction based on affine transformation; according to the source image and the target image after geometric alignment, generating a face changing image by using an adversarial network GAN; carrying out key region refinement on the face changing image based on an attention mechanism; fusing the refined face changing image with a background region based on a face segmentation model; according to the method, the GAN technology is combined with a plurality of core modules such as key point detection, region segmentation, style coding and semantic consistency correction, so that high-precision face changing of the target portrait image and the source portrait image is effectively realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a face recognition face swapping method and system based on deep learning. Background Art

[0002] In the current fields of image processing and video editing, traditional face swapping technologies have many deficiencies. The technical defects existing in the existing traditional face swapping technologies mainly include:

[0003] In terms of efficiency: Manual operations are cumbersome. Traditional face swapping technologies usually require users to manually perform complex operations such as facial feature alignment, color correction, and edge fusion. This is not only time-consuming and laborious, but also error-prone. Low degree of automation: The lack of efficient automated tools makes the face swapping process cumbersome and inefficient.

[0004] In terms of professionalism: High technical requirements. Traditional face swapping technologies require users to have certain professional knowledge of image processing and video editing. For non-professionals, the learning cost is relatively high. The effect is not natural: Due to the limitations of manual operations, traditional face swapping technologies often fail to achieve natural and realistic face swapping effects, especially in terms of facial expressions and lighting effects.

[0005] In terms of security: Risk of privacy leakage. When processing images or videos containing face information, traditional face swapping technologies may have the risk of privacy leakage, especially during data transmission and storage.

[0006] Therefore, there is an urgent need for an artificial intelligence-based face swapping technology that is efficient, automated, easy to operate, and secure to meet the needs of the majority of users. Summary of the Invention

[0007] To solve the above-mentioned problems, the present invention provides a face recognition face swapping method and system based on deep learning. Through the intelligent recognition and matching of face features by AI and image processing algorithms, the present invention aims to greatly simplify the face swapping process, lower the technical threshold, and provide users with a fast, efficient, and natural face swapping experience.

[0008] In a first aspect, a face recognition face swapping method based on deep learning provided by the present invention adopts the following technical solution:

[0009] A face recognition face swapping method based on deep learning includes:

[0010] Obtain a source image and a target image input by a user;

[0011] Perform face key point detection on the source image and the target image;

[0012] Based on a deep convolutional network, perform deep feature extraction on the source image and the target image after key point detection;

[0013] Geometrically align the source image and the target image after depth feature extraction based on affine transformation;

[0014] Generate a face-swapped image using the adversarial network GAN based on the geometrically aligned source image and target image;

[0015] Refine the key regions of the face-swapped image based on the attention mechanism;

[0016] Fuse the refined face-swapped image with the background region based on the face segmentation model;

[0017] Generate the final face-swapped image through optimization processing.

[0018] Furthermore, the face key point detection for the source image and the target image includes the operation of using the Mediapipe detection algorithm to detect the key points of the source image and the target image to obtain a key point set.

[0019] Furthermore, the depth feature extraction for the source image and the target image after key point detection based on the deep convolutional network includes using the deep convolutional network to extract the depth features of the source image and the target image to obtain high-dimensional features. The depth feature extraction process is:

[0020] F = ResNet(I) = Conv(W, I) + b

[0021] where W is the convolutional kernel weight, b is the bias, and I is the input image.

[0022] Furthermore, the geometric alignment of the source image and the target image after depth feature extraction based on affine transformation includes using the key point sets K source and K target of the source image and the target image to calculate the affine transformation matrix T affine , and align the source image to the geometric structure of the target image through affine transformation so that the facial feature points of the two images are spatially aligned. The affine transformation formula is:

[0023]

[0024] where s is the scaling factor, θ is the rotation angle, and t x , t y are the translation amounts.

[0025] Furthermore, the generation of the face-swapped image using the adversarial network GAN based on the geometrically aligned source image and target image includes using the depth features F aligned of the geometrically aligned source image I target and the target image I targetInput the generative adversarial network GAN and set the optimization objective, which is expressed as:

[0026]

[0027] where D is the discriminator, G is the generator, Pdata is the real image distribution, p z is the noise distribution, and the generated face-swapped image I GAN contains the facial style of the target image while retaining the structural information of the source image.

[0028] Furthermore, the key areas of the eyes and mouth are refined through the attention mechanism, and the calculation formula for the refinement by the attention mechanism is:

[0029]

[0030] The refined feature is expressed as:

[0031] F refined = M·V

[0032] where Q, K, and V represent the query, key, and value matrices respectively.

[0033] Furthermore, the refined face-swapped image is fused with the background area based on the face segmentation model, including using the face segmentation model BiSeNet to generate the segmentation mask of the facial area and fusing the face-swapped result with the background area. The segmentation mask ensures that the face-swapped result is applied only in the face area, while the background and other non-facial areas remain the same as the target image, which is expressed as:

[0034] I final = M seg ·I GAN +(1 - M seg )·I target where Ifinal is the final face-swapped result and Itarget is the target image.

[0035] In the second aspect, a face recognition and face-swapping system based on deep learning includes:

[0036] A data acquisition module, configured to acquire the source image and the target image input by the user;

[0037] A detection module, configured to perform face key point detection on the source image and the target image;

[0038] A feature extraction module, configured to perform deep feature extraction on the source image and the target image after key point detection based on a deep convolutional network;

[0039] An alignment module, configured to geometrically align a source image and a target image after depth feature extraction based on an affine transformation;

[0040] A face swapping module, configured to generate a face swapped image using an adversarial network GAN based on the geometrically aligned source image and target image;

[0041] A refinement module, configured to refine key regions of the face swapped image based on an attention mechanism;

[0042] A fusion module, configured to fuse the refined face swapped image with a background region based on a face segmentation model;

[0043] An optimization module, configured to generate a final face swapped image through optimization processing.

[0044] In a third aspect, the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the described face recognition and face swapping method based on deep learning.

[0045] In a fourth aspect, the present invention provides a terminal device, including a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the described face recognition and face swapping method based on deep learning.

[0046] In summary, the present invention has the following beneficial technical effects:

[0047] By adopting generative adversarial network (GAN) technology combined with multiple core modules such as key point detection, region segmentation, style encoding, and semantic consistency correction, the present invention effectively realizes high-precision face swapping between a target portrait and a source portrait. Compared with traditional face swapping methods, the present invention can generate natural, realistic, and highly consistent face swapping results while maintaining the identity characteristics of the target portrait subject. Specifically, the method of the present invention significantly improves the face swapping effect and practicality in the following aspects:

[0048] 1. The improved key point calibration algorithm proposed by the present invention significantly improves the accuracy of face key point detection, ensuring the correctness of the face structure during face swapping. By optimizing the error of key points and geometric correction, the problem of face swapping deviation caused by inaccurate key point positioning is effectively solved.

[0049] 2. In terms of face region segmentation, the BiSeNet model is used to accurately segment the target image and the source image, and the segmentation accuracy is significantly better than traditional methods based on skin color detection. This improvement not only effectively avoids the problems of segmentation errors and blurred boundaries, but also greatly improves the accuracy of the face swapping region, providing a solid foundation for generating high-quality face swapped images.

[0050] 3. Through the style encoder based on StyleGAN2, the present invention realizes the deep style fusion of the target image and the source image. During the fusion process, by using the weighted average strategy of the style vectors, the fusion ratio between the target image and the source image can be flexibly adjusted to meet the personalized needs of various application scenarios. This method avoids the problem of abrupt transition caused by directly replacing the face in traditional face swapping techniques, ensuring that the generated results are more natural.

[0051] 4. In terms of optimizing the face swapping result, the present invention proposes a method for optimizing the generated image based on multi-scale losses. Through the joint optimization of pixel-level loss, perceptual loss, and semantic consistency correction loss, the face swapping result not only maintains consistency in the overall structure but also shows a sense of reality in local details (such as areas like eyes, mouth, nose, etc.). Especially through semantic consistency correction, the present invention can ensure that the semantic features of the face swapped image highly match those of the target image, thereby further improving the visual quality of the result.

[0052] 5. Finally, the present invention realizes the seamless fusion of the generated image and the target image through a face mask. This method avoids the problems of abrupt boundary transition or color mismatch that may occur in traditional face swapping techniques. The finally generated face swapped image not only has a natural style but also has authentic details.

[0053] In addition, the method of the present invention also has remarkable efficiency and applicability. Different from traditional face swapping techniques that rely on large-scale deep learning training, the present invention does not require high-performance GPU computing devices and can complete the face swapping operation only relying on a small amount of data, greatly reducing the computing cost. At the same time, the execution process of this method is simple and can be quickly applied to various face swapping scenarios, including film and television production, virtual image generation, game character replacement, etc. Description of the Drawings

[0054] Figure 1 is a schematic diagram of the model training process of Embodiment 1 of the present invention;

[0055] Figure 2 is a schematic diagram of the face swapping method process of Embodiment 1 of the present invention. Detailed Embodiments

[0056] The present invention will be further described in detail below with reference to the accompanying drawings.

[0057] Embodiment 1

[0058] Refer to Figure 1 and Figure 2 , a face recognition face swapping method based on deep learning in this embodiment includes:

[0059] S1. Receive the source image material and the target face image input by the user;

[0060] Among them, a large number of image datasets containing different facial features, expressions, angles, and lighting conditions are collected, as well as corresponding annotation data (such as facial key points, facial features, etc.).

[0061] S2. Facial key point detection and calibration.

[0062] Perform facial key point detection on the post-makeup portrait and the target portrait respectively to obtain the coordinates of 68 key points. The key point detection is implemented using the Dlib framework. Among them, first, for the input target face image I target and the post-makeup face image I style perform face detection and cropping. Using the OpenCV or MTCNN detection framework, accurately locate the face area and crop it to a unified size (such as 256x256). This process can eliminate background interference in the image and ensure that the input to the GAN is a high-quality face area.

[0063] (1) Calibration method:

[0064] Detect the key point coordinates: The detected key point coordinates are denoted as (x i , y i ), i = 1, 2,..., 68;

[0065] Color difference calculation: Select the reference point coordinates of the upper lip key point and calculate the color difference between other points and the reference point:

[0066]

[0067] The extracted key points are used to precisely define the geometric shape of the face and ensure the accuracy of subsequent geometric alignment. The key point detection process can adopt heat map regression technology to improve the detection accuracy.

[0068] Expressed as:

[0069]

[0070] (2) Facial area segmentation

[0071] Use the key points combined with the skin detection algorithm and the BiSeNet model to segment the post-makeup portrait and the target portrait:

[0072] Obtain the preliminary contour through key point detection;

[0073] Use the skin detection algorithm to detect the forehead area;

[0074] Obtain the first contour according to the intersection of the two;

[0075] The second contour is obtained by segmenting with the BiSeNet model, and the eye and mouth regions are combined to obtain the accurate face mask M(x, y).

[0076] Among them, first, the skin color model is used: (1) Convert the image to the HSV or YCrCb color space; HSV space: The skin area often falls within 0 < H < 50, 50 < S < 150. YCrCb space: The skin area often falls within 77 < Cr < 127, 133 < Cb < 173, and create a skin detection mask.

[0077] (2) Detect the forehead area: Extract the possible forehead area from the upper face area defined by the key points. Perform an intersection operation on the skin detection mask and the forehead area to obtain the preliminary forehead area mask.

[0078] (3) Optimize the skin mask: Use morphological operations (such as dilation and erosion) to eliminate small noise points. Smooth the boundary to improve the accuracy of the mask.

[0079] (4) BiSeNet model segmentation, BiSeNet (Bilateral Segmentation Network) is an efficient semantic segmentation network suitable for face segmentation tasks. Use a bilateral architecture: Spatial branch: Capture high-resolution details; Semantic branch: Extract global context information.

[0080] 2. Model training, use a publicly available face segmentation dataset (such as CelebAMask-HQ) for pre-training, label categories: facial parts (such as skin, eyes, mouth, nose, hair, etc.).

[0081] 3. Model inference, output a pixel-level segmentation mask, including facial parts; Extract the regions related to the human face in the segmentation result to form the second contour.

[0082] S3. Deep feature extraction After completing the key point detection, perform deep feature extraction on the source image and the target image through a deep convolutional neural network (such as ResNet or MobileNet) to obtain high-dimensional features. After completing the key point detection, perform deep feature extraction on the source image and the target image through a deep convolutional neural network (such as ResNet or MobileNet) to obtain high-dimensional features F source and F target Deep feature extraction can capture the global semantic information and local texture details of the human face, providing semantic support for subsequent face swapping generation.

[0083] The process of deep feature extraction can be expressed as:

[0084] F = ResNet(I) = Conv(W, I) + b

[0085] Among them, W is the convolution kernel weight, b is the bias, and I is the input image. This step ensures that the extracted features not only contain the geometric shape of the face but also can capture texture and light and shadow information.

[0086] Among them, the core formula of the convolution operation is F = f(W * I + b), where W is the convolution kernel weight, b is the bias, I is the input image, and f is the activation function (such as ReLU). The convolution layer sequentially extracts low-level features (such as edges), intermediate features (such as textures and shapes), and high-level features (such as semantic information), and reduces the dimension through pooling layers (such as max pooling and average pooling) to retain key information. Batch Normalization is commonly used in the network to standardize the features, ensuring the stability and convergence speed of training. In addition, local and global features are captured through multi-scale convolution kernels, and shallow layer detail information is retained by combining skip connections, further enhancing the expression ability of light and shadow and texture. Finally, the feature map is converted into a fixed-size feature vector through global average pooling (GAP) and fully connected layers, providing input for subsequent tasks (such as classification, matching, or generation). This process extracts rich semantic multi-level features while ensuring efficient computation, taking into account both local details and global characteristics.

[0087] S4. Geometric alignment and affine transformation are based on the target face and the key point set P of the face target and P style , calculate the affine transformation matrix T. Through affine transformation, the geometric structure of the post-makeup face is aligned to the shape of the target face, thus eliminating the structural differences between the two images.

[0088] The transformation matrix T is calculated using Procrustes analysis or the least squares method. Among them,

[0089] The source point set is The target point set is The transformation matrix is expressed as T = sR + t;

[0090] (1) First, perform centering processing:

[0091] Calculate the centers of the two point sets:

[0092]

[0093] Decentralize the two point sets:

[0094] (2) Calculation of the rotation matrix R:

[0095] Calculate the covariance matrix: H = X′ T Y′

[0096] Perform singular value decomposition (SVD) on H: H = U∑VT

[0097] Calculate the rotation matrix: R = VU T

[0098] (3) Calculation of the scaling factor s:

[0099] (4) Calculation of the translation vector t:

[0100] Calculate the translation vector by combining scaling and rotation:

[0101] (5) Composite transformation matrix T:

[0102] The final transformation is: T = sR, t.

[0103] The aligned face image can be better fused with the target face, which is expressed by the formula:

[0104]

[0105] The affine transformation formula is:

[0106]

[0107] where s is the scaling factor, θ is the rotation angle, and t x ,t y is the translation amount.

[0108] Apply the affine transformation matrix T to the source image to obtain the aligned image. After affine transformation, the core of the image alignment lies in that it realizes the geometric position consistency of key points by using linear mapping. This not only includes the adjustment of rotation, scaling and translation, but also involves the optimization of the overall structure, so that the source image and the target image are spatially aligned.

[0109] The aligned image can be expressed as:

[0110] I aligned = T affine (I source )

[0111] Geometric alignment is a crucial step in face swapping technology, which ensures the structural compatibility between the source image and the target image.

[0112] S5. Build a GAN model

[0113] Build a face swapping model based on GAN, including a generator (G) and a discriminator (D):

[0114] Among them, the generator: learns the style of portrait pictures and generates face swapped images

[0115] Discriminator: Evaluate the authenticity and semantic consistency of the generated images.

[0116] The optimization objective of the generator is:

[0117]

[0118] The optimization objective of the discriminator is:

[0119]

[0120] The geometrically aligned images and the target face image Itarget are input into a pre-trained generative adversarial network (GAN). The generator G generates an initial face-swapped image after fusion.

[0121] The generator G is responsible for generating natural face-swapped images, and the discriminator D optimizes G by evaluating true and false image pairs.

[0122] Use the StyleGAN or GANimation model to achieve style transfer and image synthesis.

[0123] The formula is expressed as:

[0124]

[0125] where, P data : The distribution of real data. P z : The distribution of random noise (usually Gaussian distribution or uniform distribution). G(z): The image generated by the generator G from the random noise z. D(x) The output of the discriminator D for the input x (the value range is between [0,1], indicating the probability that x belongs to real data).

[0126] The optimization objective of GAN is as follows:

[0127]

[0128] where, D is the discriminator, G is the generator, pdata is the real image distribution, p z is the noise distribution.

[0129] S6. Style Encoding

[0130] Among them, use the latent space encoder of StyleGAN2 to map the post-makeup portrait image and the target portrait image to style vectors respectively:

[0131] w makeup =E(I makeup ), w target =E(I target )

[0132] Generate a new style vector through weighted fusion:

[0133] w combined = α·w makeup +(1 - α)·w target

[0134] where α ∈ [0, 1] is the style fusion weight.

[0135] S7. Key Region Refinement

[0136] Semantic Consistency Correction

[0137] Among them, using a pre - trained semantic segmentation network, the portrait image and the target portrait image are segmented, including regions such as eyebrows, eyes and mouths. Based on the key region alignment rule:

[0138]

[0139] Optimize the style vector to minimize the semantic difference.

[0140] Detail Optimization, using a multi - scale loss function to optimize the generated image:

[0141] (1) Pixel Loss:

[0142]

[0143] (2) Perceptual Loss:

[0144]

[0145] where V i is the feature map of the i - th layer of the pre - trained VGG network.

[0146] S8. Fusing the Face - Swapped Image

[0147] Perform pixel - level and feature - level fusion on the initially generated face - swapped image to enhance the authenticity of local details. Adopt a regional mask M face to optimize key regions such as eyes and mouths.

[0148] Adopt a multi - resolution fusion strategy for local regions to improve the detail effect. Expressed as:

[0149]

[0150] where F(x, y): the value of the fused image at the pixel point (x, y)

[0151] M i (x, y): the weight of the i - th regional mask at the pixel point (x, y),

[0152] M i (x, y) ∈ [0, 1]

[0153] I src,i The pixel value of the (x, y) source image in the i-th region.

[0154] I tgt,i The pixel value of the (x, y) target image in the i-th region.

[0155] Apply different fusion weights to different regions using the facial region mask.

[0156] The formula is expressed as:

[0157] I final (x, y) = αM face ·I fused (x, y) + (1 - α)·I target (x, y)

[0158] Where, I final (x, y): The value of the final fused image at the pixel point (x, y)

[0159] M face The facial region mask, representing the fusion weight within the facial region, M face (x, y) ∈ [0, 1], the weight of the non-facial region is 0.

[0160] I fused (x, y): The value of the image after local region fusion at the pixel point (x, y).

[0161] L target (x, y) The value of the target image at the pixel point (x, y).

[0162] Fuse the generated image with the target portrait using the face mask M(x, y):

[0163] I final (x, y) = M(x, y)·I gen (x, y) + (1 - M(x, y))·I target (x, y)

[0164] Where, I final (x, y): The value of the final fused image at the pixel position (x, y);

[0165] The mask weight at the pixel position (x, y), representing the fusion ratio of the generated image and the target image, M(x, y) ∈ [0, 1]; M(x, y) = 1 means completely using the value of the generated image. M(x, y) = 0 means completely using the value of the target image. I gen (x, y) The value of the generated image at the pixel position (x, y); I target(x, y) The value of the target image at the pixel position (x, y); Use a face segmentation model (such as BiSeNet) to generate a segmentation mask for the facial region, and fuse the face-swapping result with the background region:

[0166] I final = M seg ·IGAN+(1 - M seg )·I target

[0167] The segmentation mask ensures that the face-swapping result is only applied to the face region, while the background and other non-facial regions remain consistent with the target image.

[0168] S9. After obtaining the face-swapped image, use VGGFace or ArcFace to extract the identity features of the target face, and compare the features with the face-swapping result to ensure identity consistency.

[0169] Among them, an Euclidean distance constraint is imposed on the target face features and the generated result features.

[0170] Add an identity preservation loss function to optimize the generator: expressed as:

[0171]

[0172] Among them, Cid is an output variable; Ifused represents the value actually used or the current state value; Itarget represents the target value to be achieved or the reference standard.

[0173] Step 8: Result output. Output the final face-swapped image. The face-swapping result is natural and real, and retains the main features of the target portrait. Among them, the resolution of the optimized image is enhanced to generate a high-resolution face-swapped image, which supports multiple output formats to meet different application requirements. Use super-resolution reconstruction technology (such as ESRGAN) to enhance the image resolution. The output image format supports PNG, JPEG, etc. for subsequent storage or display.

[0174] Embodiment 2

[0175] This embodiment provides a deep learning-based face recognition and face-swapping system, including:

[0176] A data acquisition module, configured to

[0177] A computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device to perform the deep learning-based face recognition and face-swapping method described above.

[0178] A terminal device includes a processor and a computer-readable storage medium. The processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the described face recognition face-swapping method based on deep learning.

[0179] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.

Claims

1. A face recognition face-changing method based on deep learning, characterized in that: include: Get the source image and target image input by the user; Perform facial key point detection on the source image and the target image; Based on the deep convolutional network, deep features are extracted from the source image and target image after key point detection; Geometrically align the source image and target image after deep feature extraction based on affine transformation; Generate a face-swapped image using the adversarial network GAN based on the geometrically aligned source image and target image; Refine the key areas of face-swapped images based on the attention mechanism; Based on the face segmentation model, the refined face-changing image is fused with the background area; The final face-changing image is generated through optimization processing.

2. A face recognition and face-changing method based on deep learning according to claim 1, characterized in that: The detecting of facial key points on the source image and the target image includes performing key point detection on the source image and the target image using a Mediapipe detection algorithm to obtain a key point set.

3. The face recognition and face-changing method based on deep learning according to claim 2, characterized in that: The method of performing deep feature extraction on the source image and the target image after key point detection based on the deep convolutional network includes using the deep convolutional network to perform deep feature extraction on the source image and the target image to obtain high-dimensional features. The deep feature extraction process is: F = ResNet(I) = Conv(W, I) + b Among them, W is the convolution kernel weight, b is the bias, and I is the input image.

4. The face recognition and face-changing method based on deep learning according to claim 3, characterized in that: The method geometrically aligns the source image and the target image after the deep feature extraction based on the affine transformation, including using the key point set K of the source image and the target image source and K target , calculate the affine transformation matrix T affine , align the source image to the geometric structure of the target image through affine transformation, so that the facial feature points of the two images are aligned in space. The affine transformation formula is: Where s is the scaling factor, θ is the rotation angle, and t x ,t y is the translation amount.

5. The face recognition and face-changing method based on deep learning according to claim 4, characterized in that: The method generates a face-changing image by using a confrontation network GAN according to the geometrically aligned source image and the target image, comprising: aligned and the target image I target The deep feature F target Input the Generative Adversarial Network (GAN) and set the optimization target. The optimization target of GAN is expressed as: Among them, D is the discriminator, G is the generator, and pda t a is the real image distribution, p z is the noise distribution, and the generated face-changing image I GAN Contains the facial style of the target image while preserving the structural information of the source image.

6. The face recognition and face-changing method based on deep learning according to claim 5, characterized in that: The eye and mouth key areas are refined by the attention mechanism, and the calculation formula of the attention mechanism refinement is: The refined features are expressed as: F refined =M·V Among them, Q, K, and V represent query, key, and value matrices respectively.

7. The face recognition and face-changing method based on deep learning according to claim 6, characterized in that: The face-swapping image after refinement is fused with the background area based on the face segmentation model, including using the face segmentation model BiSeNet to generate a segmentation mask of the face area, and fusing the face-swapping result with the background area. The segmentation mask ensures that the face-swapping result is applied only to the face area, while the background and other non-face areas remain consistent with the target image, which is expressed as: I final =M seg I GAN +(1-M seg )·I target Among them, Ifinal is the final face-changing result, and Itarget is the target image.

8. A facial recognition and face-changing system based on deep learning, characterized in that: include: The data acquisition module is configured to acquire a source image and a target image input by a user; The detection module is configured to detect facial key points on the source image and the target image; The feature extraction module is configured to perform deep feature extraction on the source image and the target image after key point detection based on a deep convolutional network; An alignment module is configured to geometrically align the source image and the target image after the deep feature extraction based on an affine transformation; The face-changing module is configured to generate a face-changing image using a GAN adversarial network according to the geometrically aligned source image and the target image; The refinement module is configured to refine the key areas of the face-swapped image based on the attention mechanism; The fusion module is configured to fuse the refined face-swapped image with the background area based on the face segmentation model; The optimization module is configured to generate a final face-changing image through optimization processing.

9. A computer-readable storage medium storing a plurality of instructions, characterized in that: The instructions are suitable for being loaded by a processor of a terminal device and executing the method according to claim 1 .

10. A terminal device, comprising a processor and a computer-readable storage medium, wherein the processor is used to implement each instruction; and the computer-readable storage medium is used to store multiple instructions, characterized in that: The instructions are suitable for being loaded by a processor and executing the method as claimed in claim 1 .

Citation Information

Patent Citations

  • Face replacement method based on multistage attribute encoder and attention mechanism

    CN112766160A

  • Face model training method and device, face changing method and device and electronic equipment

    CN115083000A

  • Makeup migration method based on generative adversarial network

    CN115496650A

  • Method for changing face after makeup based on deep convolutional neural network and face key point detection

    CN116052240A

  • High-definition face changing method based on stylized generative adversarial network

    CN117391929A

Cited By

  • Face recognition method based on deep learning in cloud computing scene

    CN120954068A