Facial disguise method and system based on expression motion transfer

By acquiring the target face image and displaying it on the display for detection by the face recognition system, a third party makes facial expressions and movements according to random instructions. The system uses a conditional GAN ​​model to generate a video of the target face's movements, and performs environmental flashing and image enhancement. This solves the security risks of existing face recognition systems and achieves an effective camouflage effect.

CN115565217BActive Publication Date: 2026-05-08AEROSPACE SCI & IND SHENZHEN GROUP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AEROSPACE SCI & IND SHENZHEN GROUP
Filing Date
2022-08-15
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing facial recognition systems have security vulnerabilities, and third-party facial spoofing technologies threaten the security and credibility of these systems. Therefore, it is necessary to provide an effective spoofing method within the scope permitted by laws and regulations.

Method used

By acquiring the target face image and displaying it on the display for detection by the face recognition system, a third party makes facial expressions and movements according to random instructions, captures driving video, extracts key point information, uses a conditional GAN ​​model to generate the target face's motion video, and performs environmental flash processing and image enhancement to achieve facial expression and movement transfer.

Benefits of technology

Within the scope permitted by laws and regulations, it enables the transfer of facial expressions and movements from a third party to the target face, thereby completing the spoofing of the facial recognition system and enhancing the system's security and credibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115565217B_ABST
    Figure CN115565217B_ABST
Patent Text Reader

Abstract

The application provides a face camouflage method and system based on expression action migration. The method comprises the following steps: extracting a driving video made by a third person according to a random instruction from a face recognition system, taking the difference between the key points of each frame image in the driving video and the key points of a target face, generating an action video of the target face, migrating the action of the third person in the driving video to the target face, retaining the appearance of the target face, adding the motion characteristics of the face of the third person in the driving video, obtaining an action video of the target face, and completing face recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of facial recognition technology, and in particular relates to a facial spoofing method and system based on facial expression and motion transfer. Background Technology

[0002] Today, facial recognition technology is widely used in people's daily lives, such as for unlocking mobile phones, verifying accounts, access control systems, financial payments, and police fugitive apprehension, providing important support for the construction of smart cities and safe cities. However, existing facial recognition systems still have many security vulnerabilities. Facial spoofing technology seriously threatens the security and credibility of facial recognition technology, posing significant security risks not only to users' property and privacy but also to public safety management. Researching how to use third-party facial spoofing techniques can better eliminate the security vulnerabilities of facial recognition. Summary of the Invention

[0003] The technical problem to be solved by this invention is how to use third-party facial expressions to disguise the face of a target within the scope permitted by laws and regulations. A facial disguise method and system based on facial expression and motion transfer is proposed.

[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0005] A face spoofing method based on facial expression and motion transfer includes the following steps:

[0006] Step 1: Acquire the target face image and display the target face image on the display for detection by the face recognition system;

[0007] Step 2: The third party makes corresponding facial expressions and actions according to the random instructions issued by the facial recognition system, and the driving video of the third party making the facial expressions and actions is captured;

[0008] Step 3: Extract the key point information of the target face, and sequentially extract the key point information of the third-party face in each frame of the driving video;

[0009] Step 4: Calculate the difference between the key point information of the third party's face and the key point information of the target face in all frames of the driving video, and calculate the frame images of the target face one by one according to the difference to form the action video of the target face according to the command.

[0010] Step 5: Display the target action video on the display for detection by the face recognition system.

[0011] Furthermore, the key point information extraction method in step 3 uses a key point extraction network model for extraction.

[0012] Furthermore, the training method for the key point extraction network model is as follows:

[0013] Step 3.1: Using the same face motion video, extract key point information from two different frames in the motion video. Let the key point information extracted from the previous frame X be H, and the key point information extracted from the next frame X' be H'. Then the sparse motion representation of the key points between the two frames is H^, H^ = H-H'.

[0014] Step 3.2: Based on the sparse motion representation of key points X in the previous frame and the sparse motion representation H^ between the two frames, input the dense motion representation model to obtain the feature vector of the dense motion representation of key points.

[0015] Step 3.3: Train the key point extraction network model using a conditional GAN. In the generator of the conditional GAN, input the previous frame X into the encoder to obtain the feature vector of the previous frame X. Concatenate the feature vector of the previous frame X with the feature vector of the dense motion representation of the key points in each key point channel direction. Decode the concatenated feature vector using the decoder to generate image X^.

[0016] Step 3.4: Input (X^,H') and the true label (X',H') into the discriminator of the conditional GAN ​​for discrimination. When training the generator, the discriminator should try to make (X^,H') judged as true, and when training the discriminator, the discriminator should try to make (X^,H') judged as false and (X',H') judged as true.

[0017] Step 3.5: Return to step 3.1 until the key point extraction network model is trained.

[0018] Furthermore, the method for calculating a frame of the target face image based on the difference in step 4 is as follows: the key point information of the target face and the difference are spliced ​​together on the key point channel to obtain the spliced ​​key point information, and the image is obtained by decoding based on the spliced ​​key point information.

[0019] Furthermore, the system collects changes in ambient light in front of the face recognition system and applies flashing technology to the target face in real time based on these changes.

[0020] Furthermore, image enhancement of the target face is performed using a contrast enhancement method.

[0021] Furthermore, the method for enhancing contrast is histogram equalization.

[0022] The present invention also provides a face spoofing system based on facial expression and motion transfer, comprising the following modules:

[0023] Target face acquisition module: used to acquire target face images and display the target face images on the display end for detection by the face recognition system;

[0024] Drive video acquisition module: used to enable a third party to make corresponding facial expressions and actions according to random instructions issued by the face recognition system, and to acquire drive video of the third party making the facial expressions and actions;

[0025] Key point information extraction module: used to extract key point information of the target face, and to extract key point information of the third party's face in each frame of the driving video in sequence;

[0026] Target motion video production module: It is used to calculate the difference between the key point information of the third party's face and the key point information of the target face in all frames of the driving video, and to calculate the frame images of the target face one by one according to the difference to form the motion video of the target face according to the instructions.

[0027] Recognition module: Used to display the target action video on the display end for detection by the face recognition system.

[0028] By adopting the above technical solution, the present invention has the following beneficial effects:

[0029] This invention provides a face spoofing method and system based on facial expression and motion transfer. By extracting a driving video of a third party performing random instructions issued by a face recognition system, the key points of each frame in the driving video are compared with the key points of the target face to generate a motion video of the target face. This transfers the actions of the third party in the driving video to the target face, thus preserving the appearance of the target face. By adding the motion characteristics of the third party's face in the driving video, the motion video of the target face is obtained, thus completing face recognition. Attached Figure Description

[0030] Figure 1 This is a system flowchart of the present invention;

[0031] Figure 2 This is a schematic diagram of a specific implementation of action transfer;

[0032] Figure 3 This is the overall network diagram;

[0033] Figure 4 This is a schematic diagram of the overall process; Detailed Implementation

[0034] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] Figures 1 to 4 This invention illustrates a specific embodiment of a face spoofing method based on facial expression and motion transfer, comprising the following steps:

[0036] Step 1: Acquire the target face image and display it on the display for detection by the face recognition system; in order to better disguise, it is necessary to first acquire the target face and display it on the display for detection by the face recognition system.

[0037] Step 2: A third party performs corresponding facial expressions based on random instructions from the face recognition system, and the driving video of the third party's facial expressions is captured. Facial expressions include blinking, opening the mouth, and shaking the head. The difficulty in faking this part lies in the fact that it is a random action and cannot be prepared in advance. Therefore, this invention uses a third party to perform actions based on random instructions, extracts the facial expressions made by the third party, and transfers them to the target face to complete the facial expression transfer. In this embodiment, the third party's face is captured, and the face is detected using the lightweight detection network FaceDetectorYN built into OpenCV, scaled to 512. 512 is used to subsequently transfer facial expressions and head movements onto the face of the target person to be attacked.

[0038] Step 3: Extract key point information of the target face, and sequentially extract key point information of the third-party face in each frame of the driving video; in this embodiment, a key point extraction network model is used for extraction.

[0039] Specifically, the training method of the key point extraction network model is as follows: Figure 3 As shown:

[0040] Step 3.1: Using the same facial motion video, extract keypoint information from two different frames of the video. Let the keypoint information extracted from the first frame X be H, and the keypoint information extracted from the second frame X' be H'. Then, the sparse motion representation of the keypoints between the two frames is H^, where H^ = H - H'. In this embodiment, the keypoint information extracted from the two images is a multi-channel Gaussian distribution heatmap, with one keypoint per channel. This embodiment uses a self-supervised keypoint extraction method, which can be found in the reference "Unsupervised Learning of Object Landmarks through Conditional Image Generation, http: / / www.robots.ox.ac.uk / ~vgg / research / unsupervised_landmarks / ".

[0041] Step 3.2: Based on the sparse motion representation of keypoints X from the previous frame and H^ between the two frames, input the dense motion representation model to obtain the feature vector of the dense motion representation of keypoints. The conversion from sparse motion representation to dense motion representation is accomplished using a dense motion transformation network M. The dense motion transformation network M converts the input X and H^ to R^ through the U-Net network (《U-Net: Convolutional Networks for Biomedical Image Segmentation》). H×W×2 The format is as follows: R represents a vector space, H×W×2 is a vector space, H and W are the width and height of the feature map, and 2 represents the shift value in the X and Y directions.

[0042] Step 3.3: Train the keypoint extraction network model using a conditional GAN. In the generator of the conditional GAN, the previous frame X is input into the encoder to obtain the feature vector of the previous frame X. The feature vector of the previous frame X is concatenated with the feature vector of the dense motion representation of the keypoints in each keypoint channel direction. The concatenated feature vector is decoded using the decoder to generate image X^. The reason why the sparse motion representation needs to be converted into a dense motion representation is that the conditional GAN ​​in step 3.3 can only process aligned images. In the generator, there is a certain offset between X and X^. The sparse motion representation needs to be converted into a dense motion representation before it can be input into the generator to align the image.

[0043] Step 3.4: Input (X^,H') and the true label (X',H') into the discriminator of the conditional GAN ​​for discrimination. When training the generator, the discriminator should try to make (X^,H') judged as true, and when training the discriminator, the discriminator should try to make (X^,H') judged as false and (X',H') judged as true.

[0044] Step 3.5: Return to step 3.1 until the key point extraction network model is trained.

[0045] In this embodiment, by continuously extracting two different frames from the same facial action video as training samples, the key point extraction network model is trained, thereby effectively capturing the key point information of a third-party face in the driving video. Then, the action difference between the target face and the third-party face driving video is combined to form the action video of the target face.

[0046] Step 4: Calculate the difference between the key point information of the third party's face and the key point information of the target face in all frames of the driving video, and calculate the target face frame by frame based on the difference to form the action video of the target face according to the instructions.

[0047] In this embodiment, the method for calculating a frame of the target face based on the difference is as follows: the key point information of the target face and the difference are spliced ​​together on the key point channel to obtain the spliced ​​key point information, and the image is obtained by decoding based on the spliced ​​key point information.

[0048] Step 5: Display the target action video on the display for detection by the face recognition system.

[0049] To more vividly simulate the lighting conditions of a target face, changes in ambient flashing in front of the face recognition system are collected, and the target face is flashed accordingly. In this embodiment, the method of using a front-facing camera to sense ambient flashing in front of the face recognition system employs a face liveness detection method as described in patent document CN109376608A, which uses different colored flashing to change the target face in real time based on the flashing color, achieving a camouflage effect.

[0050] In this embodiment, a contrast enhancement method is also used to enhance the image of the target face, making the brightness value distribution of the image more even to avoid it being too bright or too dark. Specifically, the contrast enhancement method is histogram equalization. Histogram equalization is a method to enhance image contrast. Its main idea is to transform the histogram distribution of an image into an approximately uniform distribution through a cumulative distribution function, thereby enhancing the image contrast. In order to expand the brightness range of the original image, a mapping function is needed to evenly map the pixel values ​​of the original image to the new histogram. This mapping function has two conditions:

[0051] ① The original pixel value order must not be disrupted, and the relationship between bright and dark values ​​must not be changed after mapping;

[0052] ② After mapping, the value must be within the original range, that is, the value range of the pixel mapping function should be between 0 and 255.

[0053] The steps of histogram equalization:

[0054] ① Scan each pixel of the original grayscale image sequentially and calculate the grayscale histogram of the image;

[0055] ② Calculate the cumulative distribution function of the gray-level histogram;

[0056] ③ Obtain the mapping relationship between input and output based on the cumulative distribution function and histogram equalization principle.

[0057] ④ Finally, perform image transformation based on the results obtained from the mapping relationship.

[0058] The present invention also provides a face spoofing system based on facial expression and motion transfer, comprising the following modules:

[0059] Target face acquisition module: used to acquire target face images and display the target face images on the display end for detection by the face recognition system;

[0060] Drive video acquisition module: used to enable a third party to make corresponding facial expressions and actions according to random instructions issued by the face recognition system, and to acquire drive video of the third party making the facial expressions and actions;

[0061] Key point information extraction module: used to extract key point information of the target face, and to extract key point information of the third party's face in each frame of the driving video in sequence;

[0062] Target motion video production module: It is used to calculate the difference between the key point information of the third party's face and the key point information of the target face in all frames of the driving video, and to calculate the frame images of the target face one by one according to the difference to form the motion video of the target face according to the instructions.

[0063] Recognition module: Used to display the target action video on the display end for detection by the face recognition system.

[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A face spoofing method based on facial expression and motion transfer, characterized in that, Includes the following steps: Step 1: Acquire the target face image and display the target face image on the display for detection by the face recognition system; Step 2: The third party makes corresponding facial expressions and actions according to the random instructions issued by the facial recognition system, and the driving video of the third party making the facial expressions and actions is captured; Step 3: Extract the key point information of the target face, and sequentially extract the key point information of the third-party face in each frame of the driving video; The key point information extraction method uses a key point extraction network model for extraction; The training method for the key point extraction network model is as follows: Step 3.1: Using the same face motion video, extract key point information from two different frames in the motion video. Let the key point information extracted from the previous frame X be H, and the key point information extracted from the next frame X' be H'. Then the sparse motion representation of the key points between the two frames is H^, H^ = H-H'. Step 3.2: Based on the sparse motion representation H^ of the keypoints in the previous frame X and the keypoints between the two frames, input the dense motion representation model to obtain the feature vector of the dense motion representation of keypoints. The sparse motion representation is then converted to a dense motion representation using a dense motion transformation network M. The dense motion transformation network M converts the input X and H^ into R^ through a Unet network. H×W×2 The format is: R represents a vector space, H×W×2 is a vector space, H and W are the width and height of the feature map, and 2 represents the shift value in the X and Y directions. Step 3.3: Train the key point extraction network model using a conditional GAN. In the generator of the conditional GAN, input the previous frame X into the encoder to obtain the feature vector of the previous frame X. Concatenate the feature vector of the previous frame X with the feature vector of the dense motion representation of the key points in each key point channel direction. Decode the concatenated feature vector using the decoder to generate image X^. Step 3.4: Input (X^,H') and the true label (X',H') into the discriminator of the conditional GAN ​​for discrimination. When training the generator, the discriminator should try to make (X^,H') judged as true, and when training the discriminator, the discriminator should try to make (X^,H') judged as false and (X',H') judged as true. Step 3.5: Return to Step 3.1 until the key point extraction network model is trained; Step 4: Calculate the difference between the key point information of the third party's face and the key point information of the target face in all frames of the driving video, and calculate the frame images of the target face one by one according to the difference to form the action video of the target face according to the command. Step 5: Display the target action video on the display for detection by the face recognition system.

2. The face camouflage method according to claim 1, characterized in that, The method for calculating a frame of the target face image based on the difference in step 4 is as follows: the key point information of the target face and the difference are concatenated on the key point channel to obtain the concatenated key point information, and the image is obtained by decoding based on the concatenated key point information.

3. The face camouflage method according to claim 1 or 2, characterized in that, The system collects changes in ambient light in front of the face recognition system and applies flashing technology to the target face in real time based on these changes.

4. The face camouflage method according to claim 3, characterized in that, Image enhancement of the target face is performed using a method that increases contrast.

5. The face camouflage method according to claim 4, characterized in that, The method for enhancing contrast is histogram equalization.

6. A face spoofing system based on facial expression and motion transfer, characterized in that, Includes the following modules: Target face acquisition module: used to acquire target face images and display the target face images on the display end for detection by the face recognition system; Drive video acquisition module: used to enable a third party to make corresponding facial expressions and actions according to random instructions issued by the face recognition system, and to acquire drive video of the third party making the facial expressions and actions; Key point information extraction module: used to extract key point information of the target face, and to extract key point information of the third party's face in each frame of the driving video in sequence; The key point information extraction method uses a key point extraction network model for extraction; The training method for the key point extraction network model is as follows: 1) Using the same face motion video, extract key point information from two different frames in the motion video. Let the key point information extracted from the previous frame X be H, and the key point information extracted from the next frame X' be H'. Then the sparse motion of key points between the two frames is represented by H^, H^ = H-H'. 2) Based on the sparse motion representation H^ of the keypoints in the previous frame X and the keypoints between the two frames, the dense motion representation model is used to obtain the feature vector of the dense motion representation of keypoints. The sparse motion representation is then converted into a dense motion representation using a dense motion transformation network M. The dense motion transformation network M converts the input X and H^ into R through a Unet network. H×W×2 The format is: R represents a vector space, H×W×2 is a vector space, H and W are the width and height of the feature map, and 2 represents the shift value in the X and Y directions. 3) The key point extraction network model is trained using a conditional GAN. In the generator of the conditional GAN, the previous frame X is input into the encoder to obtain the feature vector of the previous frame X. The feature vector of the previous frame X is concatenated with the feature vector of the dense motion representation of the key points in each key point channel direction. The concatenated feature vector is decoded using the decoder to generate the image X^. 4) Input (X^,H') and the true label (X',H') into the discriminator of the conditional GAN ​​for discrimination. When training the generator, the discriminator should try to make (X^,H') judged as true, and when training the discriminator, the discriminator should try to make (X^,H') judged as false and (X',H') judged as true. 5) Return to step 1) until the key point extraction network model is trained; Target motion video production module: It is used to calculate the difference between the key point information of the third party's face and the key point information of the target face in all frames of the driving video, and to calculate the frame images of the target face one by one according to the difference to form the motion video of the target face according to the instructions. Recognition module: Used to display the target action video on the display end for detection by the face recognition system.

Citation Information

Patent Citations

  • Method for detecting human face in vivo

    CN109376608A

  • Facial expression migration method and device, electronic equipment and storage medium

    CN112541445A

  • Face conversion model training method, storage medium, and terminal device

    WO2021023003A1