A Zero-Shot Deepfake Attribution Method Based on Bimodal Guidance

By constructing a dual-modal-guided zero-sample depth forgery attribution model, combined with the characteristics of multiple visual perspectives such as images, noise and edges, the problem of insufficient generalization ability caused by the single visual mode in the prior art is solved, and higher accuracy and robustness are achieved.

CN120220216BActive Publication Date: 2025-07-29QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510694479.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-29
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing research on deep falsification attribution mainly focuses on visual modes, and the failure to fully explore text and face analysis limits the generalization ability of the model and fails to effectively evaluate the generalization performance when facing the unseen generator.

Method used

Using the zero-sample depth forgery attribution method based on dual-modal guidance, a model including fine-grained text generator, face parser, multi-view vision encoder, language encoder, analytical encoder and multi-layer perceptron was constructed. The total loss function and Adam optimizer were trained, and multi-modal feature extraction and attribution were performed by combining the depth forgery features of multi-visual perspectives such as images, noise and edges.

Benefits of technology

It improves the accuracy and robustness of the deep falsification attribution model in the zero-sample situation, can trace the unknown generator more effectively, and improves the accuracy and scalability of the attribution process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220216B_ABST
    Figure CN120220216B_ABST
Patent Text Reader

Abstract

The present invention relates to a zero-shot deepfake attribution method based on bimodal guidance, belonging to the technical field of zero-shot deepfake attribution methods. It includes the following steps: constructing a zero-shot deepfake attribution dataset, dividing it into a training set and a test set, and preprocessing the training set; constructing a zero-shot deepfake attribution model, which includes a fine-grained text generator, a face parser, a multi-view visual encoder module, a language encoder, a parsing encoder, a predictor, and a multi-layer perceptron; inputting the preprocessed data into the model to obtain predicted language features and predicted image forgery attribution category features; training and optimizing the model through a total loss function and an optimizer to obtain a trained model; inputting the images in the test set into the multi-view visual encoder module of the trained model to obtain the final deepfake attribution discrimination result. The present invention can improve the accuracy and robustness of the deepfake attribution model in the zero-shot scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of zero-shot deepfake attribution methods, and particularly relates to a zero-shot deepfake attribution method based on bimodal guidance. Background Art

[0002] With the popularization of advanced generation frameworks such as generative adversarial networks and diffusion models, maliciously forged deepfake content on social platforms has raised major concerns about the trustworthiness of facial images and personal privacy. Therefore, tracing the source of deepfakes has become a crucial and urgent issue. Deepfake attribution technology has emerged, aiming to identify and trace forged human faces through deep learning-based methods.

[0003] Existing deepfake attribution research mainly focuses on the interaction in various fields of the visual modality, and other modalities such as text and face parsing have not been fully explored, which limits the generalization ability of the model; in addition, they often fail to carefully evaluate the generalization performance of deepfake attribution models when facing unseen generators. Therefore, the present invention provides a zero-shot deepfake attribution method based on bimodal guidance to solve the above problems. Summary of the Invention

[0004] In order to solve the above problems, the present invention provides a zero-shot deepfake attribution method based on bimodal guidance.

[0005] To achieve the above object, the present invention is realized through the following technical solutions:

[0006] The present invention provides a zero-shot deepfake attribution method based on bimodal guidance, including the following steps:

[0007] S1. Construct a zero-shot deepfake attribution dataset and divide it into a training set and a test set; preprocess the training set to obtain a preprocessed face image tensor and a forged attribution label tensor;

[0008] S2. Construct a zero-shot deepfake attribution model, which includes a fine-grained text generator, a face parser, a multi-view visual encoder module, a language encoder, a parsing encoder, a predictor, and a multi-layer perceptron;

[0009] S3. Input the preprocessed forged attribution label tensor into the fine-grained text generator to obtain a fine-grained attribution text ; input the preprocessed face image tensor into the face parser to obtain a face parsing image ; the fine-grained attribution text and the face parsing image and the preprocessed face image tensor constitute a face image text parsing pair ;

[0010] S4. Input the pre - processed face image tensor into the multi - view visual encoder module to obtain multi - view visual features ;

[0011] S5. Input the fine - grained attribution text into the language encoder to obtain the language attribution global features ;

[0012] S6. Input the face parsing image into the parsing encoder to obtain the face parsing features ;

[0013] S7. Input the multi - view visual features into the predictor and the multi - layer perceptron respectively, and obtain the predicted language features and the predicted image forgery attribution category features ;

[0014] S8. Optimize and train the parameters in each module through the total loss function and the Adam optimizer to obtain the optimized and trained zero - shot deep - fake attribution model;

[0015] S9. Input the image to be detected in the test set into the multi - view visual encoder module of the optimized and trained zero - shot deep - fake attribution model, and then pass through the multi - layer perceptron and the softmax function to obtain the final deep - fake attribution discrimination result.

[0016] Further, step S1 specifically includes:

[0017] The data set contains several face images, and each face image has a corresponding forgery attribution label;

[0018] Uniformly adjust the width×height of each face image in the deep - fake attribution face data set to , and normalize the face image by dividing the pixel value of each face image by 255, and encapsulate the normalized face image as a tensor representation , represents the vector space, represents the number of images in each batch, represents that the number of channels of a face image tensor is 3;

[0019] The forgery attribution label corresponding to each face image is processed through the torch.tensor function in PyTorch to obtain the pre - processed forgery attribution label tensor.

[0020] Further, step S4 specifically includes:

[0021] Construct a multi - view visual coding module, which includes an image encoder, an edge encoder, and a noise encoder; the image encoder is a convolutional vision transformer model CviT; the edge encoder includes an edge backbone module and an edge transformer block, and the edge backbone module includes a number of stacked convolutional layers, where the output channels of the first convolutional layer with 1 input channel is 32, and the edge transformer block is the same as the transformer block of CviT; the noise encoder includes an image patch selector, a steganalysis enrichment model SRM, and a convolutional vision transformer model CviT;

[0022] S41. Face image text parsing pair The pre - processed face image tensor is input into the multi - view visual coding module. After passing through the image encoder, a global manipulation image of the face appearance image with an output dimension of is output, and the formula is as follows: , , the formula is as follows:

[0023] ,

[0024] where, represents the feature dimension of the image , represents the number of images in each batch, which is the same as the quantity of the face image tensor , represents the operation of the convolutional vision transformer model CviT;

[0025] S42. The pre - processed face image tensor is input into the edge encoder. After passing through the edge backbone module, a local edge feature map of the face image with an output dimension of is output, and the formula is as follows:

[0026] ,

[0027] where, represents the convolutional neural network backbone operation, represents the parameters of the convolutional neural network backbone, represents the number of channels of the face image edge feature map, represents the height of the face image edge feature map, represents the width of the face image edge feature map; the local edge feature map of the face image is input into the edge transformer block, and the output dimension is Face edge global manipulation image , ;

[0028] S43. Preprocessed face image tensor Input to the noise encoder, and through the image block selector, the richest image blocks are obtained , , represents the number of images in each batch, represents the richest image blocks The number of channels is 3, represents the richest image blocks The width × height of; The richest image blocks Input to the steganalysis rich model SRM for processing to obtain noise , and the formula is as follows:

[0029] ,

[0030] where, represents the operation of the steganalysis rich model; The noise Input to the convolutional vision transformer model CviT, and the output dimension is Face noise image global manipulation image , , and the formula is as follows:

[0031] ,

[0032] S44. The face appearance image global manipulation image , the face edge global manipulation image and the face noise image global manipulation image Are fused to obtain the multi-view visual features of the global image visual fusion , and the formula is as follows:

[0033] ,

[0034] where, represents the element-wise addition operation.

[0035] Furthermore, step S5 specifically includes:

[0036] S51. Fine-grained attribution text Pass through the tokenizer to obtain a sequence of word tokens, and the word tokens in the sequence of word tokens are mapped to word embedding tensors through the word embedding layer , according to the word embedding tensor and the position of the automatically generated word embedding tensor Obtain a fine-grained attribution text sequence vector with position information , which is expressed by the formula as follows:

[0037] ;

[0038] S52. Construct a language encoder, which includes consecutive Transformer blocks. Each Transformer block includes a multi-head attention module and a feed-forward neural network module. The layer above the multi-head attention module and the feed-forward neural network module is a normalization layer, and the layer below is a residual layer;

[0039] Input the fine-grained attribution text sequence vector with position information into the language encoder. After being normalized by the normalization layer, it is input into the multi-head attention module of the first consecutive Transformer block for global multi-head attention calculation, and then the text global semantic feature is obtained through the residual layer. The formula is expressed as follows:

[0040] ,

[0041] wherein, represents the normalization operation, represents the operation of the multi-head attention module; The text global semantic feature is normalized by the normalization layer and then input into the feed-forward neural network module, and then the refined global language feature is obtained through the residual layer. The formula is expressed as follows:

[0042] ,

[0043] wherein, represents the operation of the feed-forward neural network module; Take the output of the first Transformer block of the language encoder as the input of the second Transformer block, and the output of the second Transformer block as the input of the third Transformer block, and iterate multiple times until the operation of the th Transformer block is completed to obtain the fine-grained language feature ; Take the last word from the fine-grained language feature to obtain the language attribution global feature .

[0044] Furthermore, step S6 specifically includes:

[0045] Construct a parsing encoder module, which includes a parsing backbone module and a parsing Transformer encoder; the parsing backbone module includes a number of stacked convolutional layers, and the structure of the parsing backbone module is the same as that of the edge backbone module; the parsing Transformer encoder includes Q Transformer blocks;

[0046] S61. Input the face parsing image into the parsing encoder module. After passing through the parsing backbone module, output a face image parsing local feature map with a dimension of . The formula is as follows: .

[0047] .

[0048] Among them, represents the operation of the parsing encoder module; the dimension of the face parsing local feature map is . represents the number of channels of the face image parsing feature map, represents the height of the face image parsing feature map, represents the width of the face image parsing feature map;

[0049] S62. Flatten the face image parsing local feature map along the channels using the reshaping function in the library into a two-dimensional noise block sequence . . . represents the number of patches, represents the th two-dimensional face parsing image patch; calculate the two-dimensional face parsing image patch sequence with position information . The formula is as follows:

[0050] .

[0051] Among them, represents a learnable class tensor generated automatically, represents the mapping hidden vector, . represents the th mapping hidden vector of the two-dimensional face parsing image patch , represents the position of the automatically generated two-dimensional face parsing image patch sequence;

[0052] S63. The two-dimensional face parsing image patch sequence Input into the Q personal face parsing Transformer blocks. After passing through the first Transformer block, it sequentially passes through the multi-head self-attention module and the multi-layer perceptron in the first Transformer block module, and outputs the first two-dimensional face parsing global feature map , which is expressed by the formula as follows:

[0053] ,

[0054] ,

[0055] wherein, represents the sequence of two-dimensional face parsing image blocks output after passing through the multi-head self-attention module, represents the operation of the normalization layer, represents the operation of the multi-head self-attention module, represents the multi-layer perceptron module operation; take the output of the first Transformer block as the input of the second Transformer block, and then take the output of the second Transformer block as the input of the third Transformer block, and iterate multiple times until the operation of the Qth Transformer block is completed, obtaining the Qth two-dimensional face parsing global feature map ; take out the learnable class tensor in the Qth two-dimensional face parsing global feature map to obtain the face parsing feature .

[0056] Furthermore, step S7 specifically includes:

[0057] Input the multi-view visual features into a predictor composed of fully connected layers and a multi-layer perceptron composed of fully connected layers respectively, to obtain the predicted language feature and the predicted image forgery attribution category feature , which is expressed by the formula as follows:

[0058] ,

[0059] ,

[0060] wherein, represents the predictor parameter, represents the multi-layer perceptron parameter, represents the softmax function.

[0061] Furthermore, step S8 specifically includes:

[0062] Through divergence loss function optimize the language attribution global features and the predicted language features The formula is as follows:

[0063] ,

[0064] where represents the index of , represents the transpose, represents the language attribution global features of the th image,

[0065] Through the cross-modal contrast loss function to optimize the language attribution global features and the multi-view visual features The formula is as follows:

[0066] ,

[0067] ,

[0068] ,

[0069] ,

[0070] ,

[0071] where represents the visual-to-language contrast loss, represents the language-to-video contrast loss, represents the label of the sample pair of the th input face image, represents the th language-to-visual similarity matrix, represents the cosine similarity function, represents the trainable temperature parameter, represents the th multi-view visual feature, represents the A multi-view visual feature represents the th language attribution global feature;

[0072] The face parsing feature and the multi-view visual feature are optimized by the cross-view contrast loss function as follows:

[0073] ,

[0074] ,

[0075] ,

[0076] ,

[0077] ,

[0078] where represents the visual-to-parsing contrast loss, represents the parsing-to-visual contrast loss, represents the th parsing-to-visual similarity matrix, represents the th visual-to-parsing similarity matrix, represents the th multi-view visual feature, represents the th face parsing feature, represents the th multi-view visual feature, represents the th face parsing feature;

[0079] The predicted image forgery attribution category feature and the image generator-level label are optimized by the deepfake attribution loss function as follows:

[0080] ,

[0081] where represents the th image generator-level label, represents the th predicted image forgery attribution category feature;

[0082] According to the divergence loss function , cross-vision contrast loss , cross-modal contrast loss and the depth attribution loss function to obtain the total loss , .

[0083] Furthermore, step S9 specifically includes

[0084] inputting the face images to be detected in the test set into the optimized and trained multi-view visual encoder module to obtain the global visual features of the images , and then using the global visual features of the images to calculate the final true / false discrimination result through a multi-layer perceptron and a softmax function , and the formula is as follows:

[0085] ,

[0086] wherein, represents the multi-view visual encoder

[0087] The advantages of the present invention are as follows:

[0088] By introducing the strategies of multi-view learning and dual-modal guidance, the present invention improves the accuracy and robustness of the deepfake attribution model in the zero-shot scenario. By combining the deepfake features of multiple visual perspectives such as images, noises, and edges, the present invention can more effectively trace the unseen generators. Meanwhile, with the support of face parsing and language modality, the accuracy and scalability of the attribution process are further improved, thereby significantly enhancing the application value and practical effect of the deepfake attribution technology BRIEF DESCRIPTION OF THE DRAWINGS

[0089] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention

[0090] Figure 1 is the flowchart of the steps of the method of the present invention

[0091] Figure 2 is the face parsing diagram of true images, GAN face images, and diffusion face images DETAILED DESCRIPTION OF THE EMBODIMENTS

[0092] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0093] Embodiment 1

[0094] In this embodiment, as Figure 1 shown, the present invention provides a zero-shot deepfake attribution method based on bimodal guidance. The specific steps include:

[0095] S1. Construct a zero-shot deepfake attribution dataset and divide it into a training set and a test set; preprocess the training set to obtain preprocessed face image tensors and forgery attribution label tensors;

[0096] Specifically, the dataset contains a number of face images, and each face image has a corresponding forgery attribution label;

[0097] Unify the width × height of each face image in the deepfake attribution face dataset to , and normalize the face image by dividing the pixel value of each face image by 255. Enclose the normalized face image as a tensor representation , represents the vector space, represents the number of images in each batch, represents that the number of channels of a face image tensor is 3;

[0098] The forgery attribution label corresponding to each face image is processed by the torch.tensor function in PyTorch to obtain a preprocessed forgery attribution label tensor.

[0099] The forgery attribution labels corresponding to the face images in the training set are as follows. The first layer is the coarse-grained layer, and the labels are denoted as [b,0] or [b,1], Indicates the number of images in each batch. 0 represents real images, and 1 represents fake images; The second layer is the manipulation layer, and the labels are denoted as [b,0] or [b,1] or [b,2] or [b,3]. 0 represents real images, 1 represents identity-swapped images, 2 represents attribute-manipulated images, and 3 represents full-face synthesized images; The third layer is the fine-grained layer, and the labels are denoted as [b,0] or [b,1] or [b,2] or [b,3] or [b,4] or [b,5] or [b,6]. 0 represents real images, 1 represents identity-swapped GAN images, 2 represents identity-swapped diffusion images, 3 represents attribute-manipulated GAN images, 4 represents attribute-manipulated diffusion images, 5 represents full-face synthesized GAN images, and 6 represents full-face synthesized diffusion images; The fourth layer is the generator layer, and the labels are denoted as [b,0] or [b,1] or [b,2] or [b,3] or [b,4] or [b,5] or [b,6] or [b,7] or [b,8] or [b,9] or [b,10]. 0 represents real images, 1 represents DiffFace generator images, 2 represents Diffae generator images, 3 represents LatentDiffusion generator images, 4 represents CollaborativeDiffusion generator images, 5 represents DDPM generator images, 6 represents FSLSD generator images, 7 represents FaceSwapper generator images, 8 represents LatentTransformer generator images, 9 represents StyleGAN2 generator images, and 10 represents StyleGAN3 generator images.

[0100] S2. Construct a zero-shot deepfake attribution model, which includes a fine-grained text generator, a face parser, a multi-view visual encoder module, a language encoder, a parsing encoder, a predictor, and a multi-layer perceptron;

[0101] S3. Input the preprocessed forged attribution label tensor into the fine-grained text generator to obtain fine-grained attribution text ; Input the preprocessed face image tensor into the face parser to obtain a face parsing image ; The fine-grained attribution text , the face parsing image , and the preprocessed face image tensor constitute a face image-text parsing pair ;

[0102] Figure 2 Demonstrates the distinguishable face parsing prior features between real and fake face images; It can be seen that there are obvious differences between the real parsing image and the parsing images of GAN-generated faces and the face parsing images generated by the diffusion model. Therefore, this parsing feature can be incorporated into the model as a valuable prior feature.

[0103] S4. Input the preprocessed face image tensor into the multi-view visual encoder module to obtain multi-view visual features ;

[0104] Specifically, construct a multi-field visual encoding module, which includes an image encoder, an edge encoder, and a noise encoder; the image encoder is a convolutional vision transformer model CviT; the edge encoder includes an edge backbone module and an edge transformer block. The edge backbone module includes a stack of several convolutional layers, where the output channel of the first convolutional layer with an input channel of 1 is 32, and the edge transformer block is the same as the transformer block of CviT; the noise encoder includes an image patch selector, a steganalysis enrichment model SRM, and a convolutional vision transformer model CviT;

[0105] S41. Face image text parsing pair The preprocessed face image tensor in is input into the multi-field visual encoding module. After passing through the image encoder, a face appearance image global manipulation image with a dimension of is output , , and the formula is as follows:

[0106] ,

[0107] where represents the feature dimension of the image , represents the number of images in each batch, which is the same as the quantity of the face image tensor , represents the operation of the convolutional vision transformer model CviT;

[0108] S42. The preprocessed face image tensor is input into the edge encoder. After passing through the edge backbone module, a face image edge local feature map with a dimension of is output , and the formula is as follows:

[0109] ,

[0110] where represents the convolutional neural network backbone operation, represents the parameters of the convolutional neural network backbone, represents the number of channels of the face image edge feature map, Denote the height of the edge feature map of the face image, and the width of the edge feature map of the face image; input the local edge feature map of the face image into the edge transformer block, and output a face edge global manipulation image with a dimension of ; , ;

[0111] S43. The preprocessed face image tensor is input into the noise encoder, and through the image patch selector, the richest image patches are obtained, , where represents the number of images in each batch, and represents that the richest image patch has 3 channels, and represents the width × height of the richest image patch ; input the richest image patch into the steganalysis rich model SRM for processing to obtain noise , and the formula is as follows:

[0112] ,

[0113] where, represents the operation of the steganalysis rich model; input the noise into the convolutional vision transformer model CviT, and output a face noise image global manipulation image with a dimension of ; , , and the formula is as follows:

[0114] ,

[0115] S44. Fuse the face appearance image global manipulation image , the face edge global manipulation image and the face noise image global manipulation image to obtain the multi-view visual features of the global image visual fusion , and the formula is as follows:

[0116] ,

[0117] where, represents the element-wise addition operation.

[0118] S5. Input the fine-grained attribution text into the language encoder to obtain the language attribution global feature[[ID=8]] ;

[0119] Specifically, S51. Fine-grained attribution text The word token sequence is obtained through a tokenizer, and the word tokens in the word token sequence are mapped into word embedding tensors through a word embedding layer , according to the word embedding tensor and the positions of the automatically generated word embedding tensors to obtain a fine-grained attribution text sequence vector with position information , which is expressed by the formula as follows:

[0120] ;

[0121] S52. Construct a language encoder, the language encoder includes consecutive Transformer blocks, each Transformer block includes a multi-head attention module and a feed-forward neural network module, and the upper layer of the multi-head attention module and the feed-forward neural network module is a normalization layer, and the lower layer is a residual layer;

[0122] The fine-grained attribution text sequence vector with position information is input into the language encoder, and after being normalized by the normalization layer, it is input into the multi-head attention module of the first consecutive Transformer block for global multi-head attention calculation, and then the text global semantic features are obtained through the residual layer , which is expressed by the formula as follows:

[0123] ,

[0124] wherein, represents the normalization operation, represents the operation of the multi-head attention module; the text global semantic features after being normalized by the normalization layer are input into the feed-forward neural network module, and then the refined global language features are obtained through the residual layer , which is expressed by the formula as follows:

[0125] ,

[0126] wherein, represents the operation of the feed-forward neural network module; the output of the first Transformer block of the language encoder is used as the input of the second Transformer block, and the output of the second Transformer block is used as the input of the third Transformer block, and the iteration is performed multiple times until the operation of the th Transformer block is completed to obtain the fine-grained language features ; Take the last word from the fine-grained language features to obtain the global language attribution feature . .

[0127] S6. Input the face parsing image into the parsing encoder to obtain the face parsing feature ;

[0128] Specifically, construct a parsing encoder module, which includes a parsing backbone module and a parsing Transformer encoder; the parsing backbone module includes a number of stacked convolutional layers, and the structure of the parsing backbone module is the same as that of the edge backbone module; the parsing Transformer encoder includes Q Transformer blocks;

[0129] S61. Input the face parsing image into the parsing encoder module, pass through the parsing backbone module, and output a face image parsing local feature map with a dimension of , and the formula is as follows:

[0130] ,

[0131] where represents the operation of the parsing encoder module; the dimension of the face parsing local feature map is , represents the number of channels of the face image parsing feature map, represents the height of the face image parsing feature map, represents the width of the face image parsing feature map;

[0132] S62. Flatten the face image parsing local feature map along the channels using the reshaping function in the library into a two-dimensional noise block sequence , , where[[ID=6*]] represents the number of patches, represents the th two-dimensional face parsing image patch; calculate the two-dimensional face parsing image patch sequence with position information, and the formula is as follows:

[0133] ,

[0134] where represents a learnable class tensor generated automatically, Denote the mapping hidden vector, , Denote the -th two-dimensional face parsing image patch 's mapping hidden vector, Denote the position of the automatically generated sequence of two-dimensional face parsing image patches;

[0135] S63. Input the sequence of two-dimensional face parsing image patches into Q face parsing Transformer blocks. After passing through the first Transformer block, sequentially pass through the multi-head self-attention module and the multi-layer perceptron module in the first Transformer block, and output the first two-dimensional face parsing global feature map , and the formula is as follows:

[0136] ,

[0137] ,

[0138] where, Denote the output of the sequence of two-dimensional face parsing image patches after passing through the multi-head self-attention module, Denote the operation of the normalization layer, Denote the operation of the multi-head self-attention module, Denote the operation of the multi-layer perceptron module; Take the output of the first Transformer block as the input of the second Transformer block, and then take the output of the second Transformer block as the input of the third Transformer block, and iterate multiple times until the operation of the Q-th Transformer block is completed, obtaining the Q-th two-dimensional face parsing global feature map ; Take out the learnable class tensor in the Q-th two-dimensional face parsing global feature map to obtain the face parsing feature .

[0139] S7. Input the multi-view visual features into the predictor and the multi-layer perceptron respectively, and obtain the predicted language feature and the predicted image forgery attribution category feature respectively;

[0140] Specifically, input the multi-view visual features into a predictor composed of fully connected layers and a multi-layer perceptron composed of fully connected layers respectively, and obtain the predicted language feature And the predicted image forgery attribution category features , which is expressed by the following formula:

[0141] ,

[0142] ,

[0143] Among them, represents the predictor parameter, represents the multi-layer perceptron parameter, represents the softmax function.

[0144] S8. Optimize and train the parameters in each module through the total loss function and the Adam optimizer to obtain the optimized and trained zero-shot deep forgery attribution model;

[0145] Specifically, through the divergence loss function optimize the language attribution global feature and the predicted language feature , which is expressed by the following formula:

[0146] ,

[0147] Among them, represents the index of represents the transpose, represents the language attribution global feature of the th image,

[0148] optimize the language attribution global feature and the multi-view visual feature through the cross-modal contrast loss function, which is expressed by the following formula:

[0149] ,

[0150] ,

[0151] ,

[0152] ,

[0153] ,

[0154] Among them, represents the visual-to-language contrast loss, represents the language-to-video contrast loss, Indicates the label of the sample pair of the input face images, Indicates the th language-to-vision similarity matrix, Indicates the cosine similarity function, Indicates the trainable temperature parameter, Indicates the th vision-to-language similarity matrix, Indicates the th multi-view visual feature, Indicates the th language attribution global feature, Indicates the th multi-view visual feature, Indicates the th language attribution global feature;

[0155] The face parsing feature and the multi-view visual feature are optimized through the cross-view contrast loss function as follows:

[0156] ,

[0157] ,

[0158] ,

[0159] ,

[0160] ,

[0161] where Indicates the vision-to-parsing contrast loss, Indicates the parsing-to-vision contrast loss, Indicates the th parsing-to-vision similarity matrix, Indicates the th vision-to-parsing similarity matrix, Indicates the th multi-view visual feature, Indicates the th face parsing feature, Indicates the th multi-view visual feature, Indicates the th face parsing feature;

[0162] Through the deepfake attribution loss function Predict the category features of forged attribution for the predicted image and the image Generator-level label are optimized, and the formula is as follows:

[0163] ,

[0164] wherein, represents the th image generator-level label, represents the th category feature of forged attribution for the predicted image;

[0165] According to the divergence loss function , cross-vision contrast loss , cross-modal contrast loss and depth attribution loss function to obtain the total loss , .

[0166] S9. Input the image to be detected in the test set into the multi-view visual encoder module of the zero-shot deep forgery attribution model optimized by training, and then through the multi-layer perceptron and the softmax function, obtain the final deep forgery attribution discrimination result.

[0167] Specifically, input the face image to be detected in the test set into the optimized and trained multi-view visual encoder module to obtain the global visual feature of the image , and then calculate the final true / false discrimination result through the multi-layer perceptron and the softmax function , and the formula is as follows:

[0168] ,

[0169] wherein, represents the multi-view visual encoder.

[0170] Example 2

[0171] In this embodiment, in the scenario of cross-generator evaluation of zero-shot deepfake attribution performance, the seen generators in the GenFace dataset are used for training, and then the model is tested on various unseen generators in the DF40 dataset, which simulates the situation where the model detects and attributes deepfake content generated by unknown generators in actual applications. The existing dataset GenFace is input into the full-face synthesis generative adversarial network generator StyleGAN2, the full-face synthesis generative adversarial network generator StyleGAN3, the full-face synthesis diffusion model generator CollDiff, the full-face synthesis diffusion model generator DDPM, the full-face synthesis diffusion model generator LatDiff, the identity swapping generative adversarial network generator FSLSD, the identity swapping generative adversarial network generator FaceSwapper, the identity swapping diffusion model generator DiffFace, the attribute editing generative adversarial network generator LatTrans, and the attribute editing diffusion model generator Diffae to generate image data as the training set of the zero-shot deepfake dataset; the existing dataset DF40 is input into the full-face synthesis generative adversarial network generator VQGAN, the full-face synthesis diffusion model generator SD-2.1, the full-face synthesis diffusion model generator DiT, the identity swapping generative adversarial network generator SimSwap, the identity swapping generative adversarial network generator UniFace, the identity swapping generative adversarial network generator InSwapper, the identity swapping generative adversarial network generator e4s, and the identity swapping diffusion model generator REFace to generate image data as the test set of the zero-shot deepfake dataset. Table 1 shows the comparison of zero-shot deepfake attribution performance of different deepfake attribution methods, which is the accuracy ACC score of different deepfake attribution methods on the unseen generators in the DF40 dataset after training with the seen generators in the GenFace dataset. Bold and underlined indicate the best result and the second-best result respectively.

[0172] Table 1 Comparison of zero-shot deepfake attribution performance of different deepfake attribution methods

[0173]

[0174] The MFCLIP model is a multi-modal fine-grained CLIP model for general Diffusion Face Forgery Detection (DFFD), which conducts fine-grained language-guided image noise forgery representation learning. A threshold is defined to evaluate whether the deepfake attribution model has seen the generator. Specifically, when the maximum probability output by the model is lower than the threshold, the generator is marked as unseen. The threshold is set to 0.9. When the threshold is 0.9, the average accuracy ACC of the method of the present invention is about 28% higher than that of the MFCLIP model. The reason is that the model of the present invention introduces global face parsing features and performs dual-modal (text and parsing)-guided multi-view (image, noise, and edge) deepfake attribution representation learning. Compared with the single language-guided method of MFCLIP, it can extract features from more dimensions, has stronger detection and attribution capabilities for deepfake content, and thus performs better in the test scenario of unknown generators.

[0175] The average accuracy ACC of the method of the present invention is nearly 8%, 34%, and 59% higher than that of the Comparative Pseudo-Learning Model CPL, the Deepfake Network Architecture Detector DNA-Det, and the Detection Attribution Fake Image Model DE-FAKE, respectively. The Deepfake Network Architecture Detector DNA-Det only explores global attribution representations in the image modality through block-based contrastive learning, while the method of the present invention focuses on multiple modalities (text and face parsing) and multiple views (image, noise, and edge forgery embeddings). This enables the method of the present invention to capture features more comprehensively and be more adaptable when facing deepfake content generated by different generators, thus showing a higher detection accuracy in the cross-generator evaluation scenario, demonstrating the effectiveness of the model of the present invention in the zero-shot deepfake attribution ZS-DFA performance.

[0176] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A zero - sample deepfake attribution method based on bimodal guidance, characterized in that, It includes the following steps: S1. Construct a zero-shot deepfake attribution dataset and divide it into a training set and a test set; Preprocess the training set to obtain the preprocessed face image tensor and the fake attribution label tensor; S2. Construct a zero-shot deepfake attribution model, which includes a fine-grained text generator, a face parser, a multi-view visual encoder module, a language encoder, a parsing encoder, a predictor, and a multi-layer perceptron; S3. Input the preprocessed forged attribution label tensor into the fine-grained text generator to obtain the fine-grained attribution text ; Input the preprocessed face image tensor into a face parser to obtain a face parsing image ; The fine-grained attribution text , the face parsing image , and the preprocessed face image tensor constitute a face image text parsing pair ; S4. Input the preprocessed face image tensor into the multi-view visual encoder module to obtain multi-view visual features ; S5. Input the fine-grained attribution text into the language encoder to obtain the global language attribution feature ; S6. Input the face parsing image into the parsing encoder to obtain face parsing features ; S7. Input the multi-view visual features into the predictor and the multi-layer perceptron respectively, and obtain the predicted language features and the predicted image forgery attribution category features ; S8. Optimize and train the parameters in each module through the total loss function and the Adam optimizer to obtain an optimized and trained zero-shot deepfake attribution model; S9. Input the image to be detected in the test set into the multi-view visual encoder module of the optimized and trained zero-shot deepfake attribution model, and then pass through the multi-layer perceptron and the softmax function to obtain the final deepfake attribution discrimination result.

2. The zero-shot deepfake attribution method based on bimodal guidance according to claim 1, wherein Step S1 specifically includes: The dataset contains several face images, and each face image has a corresponding fake attribution label; Uniformly adjust the width × height of each face image in the deepfake attribution face dataset to , and normalize the face image by dividing the pixel value of each face image by 255, and encapsulate the normalized face image as a tensor representation , represents the vector space, represents the number of images in each batch, represents that the number of channels of a face image tensor is 3; The fake attribution label corresponding to each face image is processed by the torch.tensor function in PyTorch to obtain the preprocessed fake attribution label tensor.

3. The zero-shot deepfake attribution method based on bimodal guidance according to claim 2, wherein Step S4 specifically includes: Construct a multi-view visual encoding module, which includes an image encoder, an edge encoder, and a noise encoder; the image encoder is the convolutional vision transformer model CviT; the edge encoder includes an edge backbone module and an edge transformer block, and the edge backbone module includes several stacked convolutional layers, where the output channel of the first convolutional layer with an input channel of 1 is 32, and the edge transformer block is the same as the transformer block of CviT; the noise encoder includes an image patch selector, a steganalysis enrichment model SRM, and a convolutional vision transformer model CviT; S41. Face image text parsing pair The preprocessed face image tensor in is input into the multi-view visual encoding module. After passing through the image encoder, the output is a global manipulation image of the face appearance image with a dimension of , , , which is expressed by the formula as follows: , Among them, represents the feature dimension of the image , represents the number of images in each batch, which is consistent with the number of human face image tensors ; represents the operation of the convolutional vision transformer model CviT; Preprocessed face image tensor is input into the edge encoder, passes through the edge backbone module, and outputs a face image edge local feature map with a dimension of , which is expressed by the formula as follows: , as shown in the following formula: , Among them, represents the backbone operation of the convolutional neural network, represents the parameters of the backbone of the convolutional neural network, represents the number of channels of the edge feature map of the face image, represents the height of the edge feature map of the face image, represents the width of the edge feature map of the face image; the local edge feature map of the face image is input into the edge transformer block, and the output is a face edge global manipulation image with a dimension of , ; Preprocessed face image tensor Input into the noise encoder, and through the image patch selector, the richest image patches are obtained , , denotes the number of images per batch, denotes the richest image patches whose number of channels is 3, denotes the richest image patches width × height; Input the richest image patches into the steganalysis richness model SRM for processing to obtain noise , and the formula is as follows: , Among them, represents the operation of the steganalysis enrichment model; adding noise to the convolutional vision transformer model CviT, and outputting a global manipulation image of a face noise image with a dimension of , , , and the formula is expressed as follows: , S44. Global manipulation of the facial appearance image, the global manipulation image of the facial edge , and the global manipulation image of the facial noise image are fused to obtain the multi-view visual features of the global image visual fusion , which is expressed by the following formula: ​ , Among them, represents an element-wise addition operation.

4. A zero - sample deepfake attribution method based on dual - modality guidance according to claim 3, characterized in that, Step S5 specifically includes: S51. Fine-grained attribution text The word token sequence is obtained through a tokenizer, and the word tokens in the word token sequence are mapped into word embedding tensors through a word embedding layer , according to the word embedding tensor and the position of the automatically generated word embedding tensor A fine-grained attribution text sequence vector with position information is obtained , which is expressed by the formula as follows: ; S52. Construct a language encoder, the language encoder includes consecutive Transformer blocks, each Transformer block includes a multi-head attention module and a feed-forward neural network module, and the upper layer of the multi-head attention module and the feed-forward neural network module is a normalization layer, and the lower layer of each is a residual layer; The fine-grained attribution text sequence vector with location information is input into the language encoder and, after going through the normalization operation by the normalization layer, is input into the multi-head attention module of the first consecutive Transformer block for global multi-head attention calculation, and then the global semantic features of the text are obtained through the residual layer , and the formula is as follows: , Among them, represents the normalization operation, represents the operation of the multi-head attention module; the global semantic features of the text after passing through the normalization layer for normalization and then input into the feed-forward neural network module, and then refined global language features are obtained through the residual layer , which is expressed by the formula as follows: , Among them, represents the operation of the feed-forward neural network module; the output of the first transformer block of the language encoder is used as the input of the second transformer block, and the output of the second transformer block is used as the input of the third transformer block, and the iteration is performed multiple times until the operation of the th transformer block is completed to obtain the fine-grained language features ; the last word is taken from the fine-grained language features to obtain the global language attribution feature .​ 5. A zero-shot deepfake attribution method based on dual-modal guidance according to claim 4, characterized in that Step S6 specifically includes: Construct a parsing encoder module, which includes a parsing backbone module and a parsing transformer encoder; the parsing backbone module includes several stacked convolutional layers, and the structure of the parsing backbone module is the same as that of the edge backbone module; the parsing transformer encoder includes Q transformer blocks; S61. Input the face parsing image into the parsing encoder module. After passing through the parsing backbone module, output a face image parsing local feature map with a dimension of . The formula is expressed as follows: , Among them, represents the operation of the parsing encoder module; the local feature map of face parsing has a dimension of , where represents the number of channels of the face image parsing feature map, represents the height of the face image parsing feature map, represents the width of the face image parsing feature map; S62. Parse the local feature map of the face image Along the channel, use in the library the reshape function to flatten it into a sequence of two-dimensional noise blocks , , denote the number of patches, denote the th two-dimensional face parsing image patch; calculate the sequence of two-dimensional face parsing image patches with position information , the formula is as follows: , Among them, represents an automatically generated learnable tensor-like object, represents the mapped latent vector, , denotes the th two-dimensional face parsing image patch 's mapped latent vector, represents the position of the automatically generated sequence of two-dimensional face parsing image patches; S63. Input the two-dimensional face parsing image patch sequence into Q face parsing Transformer blocks. After passing through the first Transformer block, it successively passes through the multi-head self-attention module and the multi-layer perceptron module in the first Transformer block, and outputs the first two-dimensional face parsing global feature map , which is expressed by the following formula: , , Among them, represents a sequence of two-dimensional face parsing image patches which is the output of the multi-head self-attention module, represents the operation of the normalization layer, represents the operation of the multi-head self-attention module, represents a multi-layer perceptron module operation; taking the output of the first transformer block as the input of the second transformer block, and then taking the output of the second transformer block as the input of the third transformer block, and iterating multiple times until the operation of the Qth transformer block is completed, obtaining the Qth two-dimensional face parsing global feature map ; taking out the learnable class tensor in the Qth two-dimensional face parsing global feature map to obtain the face parsing feature .

6. The zero-shot deepfake attribution method based on bimodal guidance according to claim 5, characterized in that, Step S7 specifically includes: Input multi-view visual features into a predictor composed of fully connected layers and a multi-layer perceptron composed of fully connected layers respectively to obtain predicted language features and predicted image forgery attribution category features , which is expressed by the formula as follows: , , Among them, represents the predictor parameter, represents the multi-layer perceptron parameter, represents the softmax function.

7. A zero-shot deepfake attribution method based on bimodal guidance according to claim 6, characterized in that, Step S8 specifically includes: By divergence loss function optimize the global language attribution features and the predicted language features The formula is as follows: , Among them, represents the index of denotes transpose, represents the language attribution global feature of the represents the predicted language attribution global feature of the Through cross-modal contrastive loss The function attributes global features to language And multi-view visual features For optimization, the formula is as follows: , , , , , Among them, represents the visual-to-language contrast loss, represents the language-to-video contrast loss, represents the label of the sample pair of the th input face image, represents the th language-to-visual similarity matrix, represents the cosine similarity function, represents the trainable temperature parameter, represents the th visual-to-language similarity matrix, represents the th multi-view visual feature, represents the th language attribution global feature, represents the th multi-view visual feature, represents the th language attribution global feature; By cross-vision contrast loss function Optimize the face parsing features And multi-view visual features The specific calculation is as follows: , , , , , Among them, represents the visual-to-parsing contrast loss, represents the parsing-to-visual contrast loss, represents the th parsing-to-visual similarity matrix, represents the th visual-to-parsing similarity matrix, represents the th multi-view visual feature, represents the th face parsing feature, represents the th multi-view visual feature, represents the th face parsing feature; Attribution Loss Function for Deepfakes Attribution category features of the predicted image forgery With the image Generator-level label Is optimized, and the formula is as follows: , Among them, represents the th image generator-level label, represents the th predicted image forgery attribution category feature; According to the divergence loss function , cross-field contrast loss , cross-modal contrast loss and depth attribution loss function to obtain the total loss , .

8. A zero-shot deepfake attribution method based on bimodal guidance according to claim 7, characterized in that, Step S9 specifically includes Input the face image to be detected in the test set into the optimized and trained multi-view visual encoder module to obtain the global visual features of the image , and then use the global visual features of the image to calculate the final true / false discrimination result through a multi-layer perceptron and a softmax function . The formula is as follows: , Among them, represents a multi-view visual encoder.

Citation Information

Patent Citations

  • False face recognition method and device and computer readable storage medium

    CN111783505A

  • Forensics method for synthesized face image based on local binary pattern and deep learning

    WO2021134871A1