An Active Defense Method and System Against Deepfakes

The proactive defense system generates and embeds a model-agnostic watermark to distort and detect deepfake content, addressing the limitations of passive detection methods by ensuring high detection accuracy across diverse deepfake models.

CN115273247BActive Publication Date: 2025-07-15PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210845845.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-07-15
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

The existing technology can only passively detect deep forged content, cannot prevent its generation and propagation, and requires continuous update of detectors, which is expensive.

Method used

Generate the active defense watermarks for the model, embed and detect the watermarks by training the encoder-decoder to embed and detect the watermarks, use the gradient sequence of multiple depth forgery models to generate a distorted watermark, and detect the watermark changes through the decoder to determine whether the depth forgery has been experienced.

Benefits of technology

Active defense against a variety of deep forgery models is achieved, which can completely prevent tampering, and there is no need for deep forgery model structure information, and the detection accuracy is high.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273247B_ABST
    Figure CN115273247B_ABST
Patent Text Reader

Abstract

The present invention discloses an active defense method and system against deep fakes, belonging to the field of artificial intelligence security. The present invention generates an active defense watermark that is common to models. After embedding the watermark into a medium containing face information, it can distort the generation of deep fake models, and it can be detected whether the media content has undergone deep fakes through this watermark, completely preventing deep fake tampering. The present invention has a defensive ability against a variety of deep fake models, and can achieve the defensive effect without the structural information of the deep fake models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence security and relates to deep learning technologies such as computer vision, deepfake, and active defense. Background Art

[0002] With the continuous development of deep learning technologies, the technology for modifying face images and videos: Deepfake has become extremely popular on the Internet. Generally, deepfake technologies can modify faces through attribute modification or face replacement, can modify external features such as hair color and face shape, and can also replace a face onto other videos and images, making a person perform actions inconsistent with their identity or convey false information. For example, StarGAN (StarGAN: Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation) can generate face tampering images with different facial features and expressions from an original face image; InterfaceGAN (Interpreting the Latent Space of GANs for Semantic Face Editing) can generate face images with controllable photographing angles through latent variable editing.

[0003] Many short video platforms have begun to take measures to regulate and prohibit face-swapping videos. However, the current measures taken by platforms against deepfakes are mainly passive detection, that is, training detectors to detect videos that have been produced and released to determine whether they are deepfake content. Such detection can only provide passive defense and post-event evidence collection, and cannot prevent the generation and spread of deepfake content, and there is no way to cut off the adverse effects caused by false content; moreover, in the face of rapidly evolving deepfake models, detectors need to be continuously trained and updated, and the cost is extremely high. Summary of the Invention

[0004] In order to cut off the adverse effects brought by deepfakes, the present invention proposes an active defense method and system for deepfakes.

[0005] The technical solution provided by the present invention is as follows:

[0006] An active defense method for deepfakes, characterized in that its steps include:

[0007] 1) Obtain an active defense watermark: Prepare multiple deepfake models and the parameters of the already trained deepfake models. Specifically, it includes:

[0008] 1-1) Input any original training image and the image with a defensive watermark (if it is the first training, initialize the watermark as random noise) into the deepfake model to obtain the tampered images of the original image and the image with the watermark added.

[0009] 1-2) Backpropagate the loss on different deepfake models to obtain the gradient sequence on the image.

[0010] 1-3) Synthesize the gradient sequences of each image and each model, and after imposing upper and lower bounds on them, obtain a defensive watermark.

[0011] 1-4) Update the watermark based on the defensive watermark obtained in the previous training during each training. Specifically, the watermark obtained in this training needs to be multiplied by a coefficient α (usually 0.01) and the previous watermark multiplied by the coefficient 1 - α to obtain a new defensive watermark.

[0012] 1-5) Repeat until the training times limit is reached to obtain an active defensive watermark that can distort the generation of multiple deepfake models.

[0013] 2) Train watermark embedding and detection, specifically including:

[0014] 2-1) Prepare a certain number of face images.

[0015] 2-2) Train a training encoder-decoder. Among them, the encoder embeds the active defensive watermark obtained in the previous step into the input image, and ensures the invisibility of the embedded information through the loss function. Then, the decoder reads the embedded image and decodes the encoded watermark, and ensures the accuracy of the decoded information through the loss function. When the training is completed, the corresponding encoder and decoder weights are generated.

[0016] 3) Deepfake detection, specifically including:

[0017] 3-1) Prepare the face images to be protected (or split the video to be protected frame by frame), and the deepfake model to be defended against.

[0018] 3-2) After using the encoder obtained in the previous step to embed the active defensive watermark into the face image, input the face image into the deepfake model to obtain the forged image.

[0019] 3-3) Decode the encoded watermark from the forged image through the decoder obtained in the previous step, and compare it with the original embedded watermark. When the bit difference between the two is greater than or equal to the set threshold (usually 0.4), it is considered that the image has been deepfaked.

[0020] An active defense system against deepfakes, characterized in that the system includes:

[0021] 1) Deepfake model interface module: It includes functions for inputting pictures into the deepfake model and obtaining the generation results;

[0022] 2) Active defense watermark generation module: It is used to generate defense watermarks for protecting human faces from multiple deepfake models; Specifically, this module first completes the access of the deepfake model, calls the basic watermark generation algorithm, and combines the watermark fusion technology to generate the active defense watermark common to the model.

[0023] 3) Active defense watermark embedding module: This module trains an encoder-decoder, and uses the encoder to embed the common watermark generated by the active defense watermark generation module into the face picture.

[0024] 4) Watermark defense effect evaluation module: It is used to evaluate the degree of distortion of the output of the deepfake model caused by the watermark;

[0025] 5) Deepfake detection module: Through the decoder provided by the active defense watermark embedding module, it detects the pictures embedded with the watermark to determine whether a deepfake model has modified these pictures.

[0026] Advantages of the present invention:

[0027] The present invention generates an active defense watermark common to the model. After embedding this watermark into the media containing human face information, it can cause distortion in the generation of the deepfake model, and it can detect whether the media content has undergone deepfake through this watermark, completely preventing deepfake tampering. The present invention has a defense ability against multiple deepfake models, and can achieve the defense effect without the structural information of the deepfake model. Description of the drawings

[0028] Figure 1 It is a schematic diagram of the generation of the active defense watermark of the present invention;

[0029] Figure 2 It is a schematic diagram of the embedding of the active defense watermark and the deepfake detection of the present invention. Detailed implementation manners

[0030] The present invention designs an active defense system against deepfakes. This system includes five modules: a deepfake model interface, watermark generation, watermark embedding, defense effect evaluation, and deepfake detection. Among them:

[0031] 1) Deepfake model interface module: It includes functions for inputting pictures into the deepfake model and obtaining the generation results;

[0032] 2) Active Defense Watermark Generation Module: Used to generate defense watermarks for protecting human faces from multiple deepfake models; specifically, this module first completes the access of the deepfake model and calls the basic watermark generation algorithm, and combines the watermark fusion technology to generate the active defense watermark common to the models.

[0033] 3) Active Defense Watermark Embedding Module: This module trains an encoder-decoder, and uses the encoder to embed the common watermark generated by the active defense watermark generation module into the face image.

[0034] 4) Watermark Defense Effect Evaluation Module: Used to evaluate the degree of distortion of the output of the deepfake model caused by the watermark;

[0035] 5) Deepfake Detection Module: Through the decoder provided by the active defense watermark embedding module, it detects the images embedded with the watermark to determine whether there is a deepfake model that has modified these images.

[0036] To further illustrate the present invention, its specific implementation manners are described below through examples, but the applicable scope of the method is not limited in any way.

[0037] Taking the large-scale face attribute dataset CelebA (CelebFaces Attributes Dataset: http: / / mmlab.ie.cuhk.edu.hk / projects / CelebA.html) and the deepfake models HiSD, Stargan, AttGAN, Attentiongan trained on this dataset as the attack targets, and using the PGD attack algorithm as the basic attack algorithm to illustrate how to generate active defense watermarks, how to perform watermark embedding, and how to perform deepfake detection.

[0038] Prepare the already encapsulated Deepfake model; read in the clean CelebA dataset, scale it to the size of 256×256 and perform standardized preprocessing, and divide the CelebaA dataset into a training set, a validation set, and a test set.

[0039] The first step is to obtain the active defense watermark, as Figure 1 shown:

[0040] 1) Input any batch of original training images and the images with the defense watermark added (if it is the first training, initialize the watermark as random noise) into the deepfake model to obtain the tampered images of the original images and the images with the watermark added.

[0041] 2) Backpropagate the loss on different deepfake models to obtain the gradient sequence on the input images. Where the loss is the loss function of the output of the deepfake model obtained from the original images and the images with the watermark added:

[0042] Loss generation = MSE(G(I), G(I + W))

[0043] where I is the original image, W is the watermark, and G is the deepfake model.

[0044] 3) Combine the gradient sequences of each image and each model, and after imposing upper and lower bounds on them, obtain a defensive watermark. Specifically, when combining the gradient sequences of each image, after obtaining the gradients of a batch of images (8 images) on one model, the gradients are averaged to obtain g avg , and the PGD algorithm is used to iteratively update 10 times in the positive direction of the gradient to obtain the adversarial perturbation P:

[0045]

[0046]

[0047]

[0048] When combining the gradient sequences of each model, the adversarial perturbation P obtained on this model needs to be multiplied by the coefficient α (usually 0.01) and the previous watermark is multiplied by 1 - α to obtain a new defensive watermark.

[0049] W′ ← (1 - α)W + αP

[0050] 4) Repeat until 128 images are trained to obtain an active defensive watermark that can distort the generation of the deepfake model.

[0051] The second step is to train watermark embedding and detection:

[0052] 1) Use the training set of CelebA to train a pair of encoder - decoders based on convolutional neural networks. Among them, the encoder embeds the active defensive watermark obtained in the previous step into the input image, and through the loss function, the embedded image and the original image are constrained to be close enough, that is, minimize the mean square error to ensure that the embedded information is invisible.

[0053] Loss encoding = MSE(E(I), E(I, W))

[0054] where E is the encoder and W is the active defensive watermark obtained in the previous step.

[0055] 2) Then, the decoder reads the embedded image and decodes the encoded watermark, and through the loss function, the bit error between the decoded result and the original watermark is constrained, that is, minimize the BCE error function with logits.

[0056] Loss decoding= BCEwithLogitsLoss9W, D(E(I, W)))

[0057] where D is the decoder.

[0058] 3) After the training is completed, generate the corresponding encoder E and decoder weights.

[0059] The third step, deepfake detection, is as follows Figure 2 shown:[[]]

[0060] 1) Select the CelebA test set for input;

[0061] 2) Use the encoder obtained in the previous step to embed the active defense watermark into the face image, and then input the face image into each deepfake model to obtain the forged image;

[0062] Decode the encoded watermark from the forged image through the decoder obtained in the previous step, and compare it with the original embedded watermark. When the bit difference between the two is greater than or equal to the set threshold (0.4), it is considered that the image has been deepfaked. On the full CelebA test set, the minimum encoding change rate of the deepfake model encoding and the non-forged one is 41.0%, which can be detected.

[0063] In the deepfake model attack test with unknown model structure, the present invention obtains a 100% deepfake defense rate.

[0064] The present invention has been described above through detailed implementation cases. Researchers and technicians in the field can make non-substantive changes in form or content according to the above steps without departing from the scope of the substantive protection of the present invention. Therefore, the present invention is not limited to the content disclosed in the above embodiments, and the protection scope of the present invention shall be subject to the claims.

Claims

1. An active defense method against deepfakes, characterized in that, The steps include: 1) Obtain an active defense watermark; specifically including: 1-1) Input any original training image and the image with the defense watermark into the deepfake model to obtain the tampered images of the original image and the watermarked image; 1-2) Backpropagate the loss on different deepfake models to obtain the gradient sequence on the image; the loss is the loss function output by the deepfake model obtained from the original image and the watermarked image: Loss generation = MSE(G(I), G(I + W)) where I is the original image, W is the watermark, and G is the deepfake model; 1-3) Synthesize the gradient sequences of each image and each model, and after imposing upper and lower bounds on them, obtain an active defense watermark; 2) Train watermark embedding and detection: specifically including: 2-1) Prepare a certain number of face images; 2-2) Train an encoder-decoder. Among them, the encoder embeds the active defense watermark obtained in the previous step into the input image, and ensures the invisibility of the embedded information through the loss function; the decoder reads the embedded image and decodes the encoded watermark, and ensures the accuracy of the decoded information through the loss function; when the training is completed, generate the corresponding encoder and decoder weights; 3) Deepfake detection: specifically including: 3-1) Prepare the face images to be protected and the deepfake models to be defended against; 3-2) Use the encoder obtained in the previous step to embed the active defense watermark into the face image, and then input the face image into the deepfake model to obtain the forged image; 3-3) Decode the encoded watermark from the forged image through the decoder obtained in the previous step, and compare it with the original embedded watermark. When the bit difference between the two is greater than or equal to the set threshold, it is considered that the image has been deepfaked.

2. The active defense method against deepfakes according to claim 1, characterized in that Update the watermark based on the defense watermark obtained in the previous training during each training. Specifically, the watermark obtained in this training needs to be multiplied by the coefficient α and the previous watermark multiplied by the coefficient 1-α to obtain a new defense watermark.

3. The active defense method against deepfakes according to claim 1, wherein In step 2-2), train a pair of encoder-decoders based on convolutional neural networks. Among them, the encoder embeds the active defense watermark obtained in the previous step into the input image, and through the loss function, constrains the embedded image to be close enough to the original image, that is, minimizes the mean square error to ensure the invisibility of the embedded information.

4. The active defense method against deep fakes according to claim 1, characterized in that, The threshold set in step 3-3) is 0.

4.

5. An active defense system against deepfakes, which is used to execute the active defense method against deepfakes described in any one of claims 1-4, characterized in that, The system includes: 1) Deepfake model interface module: includes functions for inputting images into the deepfake model and obtaining the generation results; 2) Active defense watermark generation module: used to generate defense watermarks for protecting faces from multiple deepfake models; 3) Active defense watermark embedding module: this module trains an encoder-decoder, and uses the encoder to embed the general watermark generated by the active defense watermark generation module into the face image; 4) Watermark defense effect evaluation module: used to evaluate the degree of distortion of the output of the deepfake model caused by the watermark; 5) Deepfake detection module: through the decoder provided by the active defense watermark embedding module, detect the images embedded with the watermark to determine whether any deepfake model has modified these images.

6. The active defense system against deep fakes as claimed in claim 5, wherein, The active defense watermark generation module. This module first completes the access to the deepfake model, calls the basic watermark generation algorithm, and combines the watermark fusion technology to generate the active defense watermark common to the model.

Citation Information

Patent Citations

  • Watermark adding method and device based on deep learning, watermark extracting method and device based on deep learning and storage medium

    CN111768327A

  • Deep semi-fragile watermarking method for image authentication and confrontation sample defense

    CN113689318A