Zero-sample deep-forgery attribution method based on bimodal guidance

By introducing a method of dual-modal guidance and multi-view feature fusion in the deep falsification attribution model, the problem of insufficient visual modal singularity and generalization capabilities in the prior art is solved, and a higher accuracy and robustness of deep falsification attribution is achieved.

CN120220216AActive Publication Date: 2025-06-27QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +3
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510694479.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing research on deep falsification attribution mainly focuses on visual modes. Failure to fully explore text and face analytical modes has limited the generalization ability of the model and failed to effectively evaluate the generalization performance of the model when facing the unseen generator.

Method used

Using a zero-sample depth forgery attribution method based on dual-modal guidance, multi-modal feature fusion and optimization training are carried out by constructing models including fine-grained text generators, face parsers, multi-view vision encoder modules, language encoders, analytical encoders, predictors and multi-layer perceptrons, combined with multi-view depth forgery features such as images, noise and edges, multi-modal feature fusion and optimization training are carried out.

Benefits of technology

It improves the accuracy and robustness of the deep forgery attribution model in the zero-sample situation, can trace the unknown generator more effectively, improves the accuracy and scalability of the attribution process, and significantly improves the application value and actual effect of the deep forgery attribution technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220216A_ABST
    Figure CN120220216A_ABST
Patent Text Reader

Abstract

The invention relates to a zero-sample deep-forgery attribution method based on bimodal guidance, and belongs to the technical field of zero-sample deep-forgery attribution methods. The method comprises the following steps: constructing a zero-sample deep-counterfeit attribution data set, dividing the zero-sample deep-counterfeit attribution data set into a training set and a test set, and preprocessing the training set; constructing a zero-sample deep counterfeit attribution model, wherein the model comprises a fine-grained text generator, a face analyzer, a multi-view visual encoder module, a language encoder, an analysis encoder, a predictor and a multi-layer perceptron; inputting the preprocessed data into the model to obtain predicted language features and predicted image forgery attribution category features; performing training optimization on the model through a total loss function and an optimizer to obtain a trained model; and inputting the images in the test set into the multi-view visual encoder module of the trained model to obtain a final deep counterfeit attribution discrimination result. According to the invention, the accuracy and robustness of the deep counterfeit attribution model in a zero sample situation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of zero-shot deepfake attribution methods, and particularly relates to a zero-shot deepfake attribution method based on bimodal guidance. Background Art

[0002] With the popularization of advanced generation frameworks such as generative adversarial networks and diffusion models, maliciously forged deepfake content on social platforms has raised major concerns about the trustworthiness of facial images and personal privacy. Therefore, tracing the source of deepfakes has become a crucial and urgent issue. Deepfake attribution technology has emerged, aiming to identify and trace forged human faces through deep learning-based methods.

[0003] Existing deepfake attribution research mainly focuses on the interaction of various fields in the visual modality, and other modalities such as text and face parsing have not been fully explored, limiting the generalization ability of the model; in addition, they often fail to carefully evaluate the generalization performance of deepfake attribution models when facing unseen generators. Therefore, the present invention provides a zero-shot deepfake attribution method based on bimodal guidance to solve the above problems. Summary of the Invention

[0004] In order to solve the above problems, the present invention provides a zero-shot deepfake attribution method based on bimodal guidance.

[0005] To achieve the above object, the present invention is realized through the following technical solutions: The present invention provides a zero-shot deepfake attribution method based on bimodal guidance, including the following steps: S1. Construct a zero-shot deepfake attribution dataset and divide it into a training set and a test set; preprocess the training set to obtain a preprocessed face image tensor and a forged attribution label tensor; S2. Construct a zero-shot deepfake attribution model, which includes a fine-grained text generator, a face parser, a multi-view visual encoder module, a language encoder, a parsing encoder, a predictor, and a multi-layer perceptron; S3. Input the preprocessed forged attribution label tensor into the fine-grained text generator to obtain a fine-grained attribution text ; input the preprocessed face image tensor into the face parser to obtain a face parsing image ; the fine-grained attribution text and the face parsing image and the preprocessed face image tensor constitute a face image text parsing pair ; S4. Input the preprocessed face image tensor Obtain multi-view visual features by inputting into the multi-view visual encoder module ; S5. Input the fine-grained attribution text into the language encoder to obtain the language attribution global features ; S6. Input the face parsing image into the parsing encoder to obtain the face parsing features ; S7. Input the multi-view visual features into the predictor and the multi-layer perceptron respectively to obtain the predicted language features and the predicted image forgery attribution category features ; S8. Optimize and train the parameters in each module through the total loss function and the Adam optimizer to obtain the optimized and trained zero-shot deep forgery attribution model; S9. Input the images to be detected in the test set into the multi-view visual encoder module of the optimized and trained zero-shot deep forgery attribution model, and then pass through the multi-layer perceptron and the softmax function to obtain the final deep forgery attribution discrimination result.

[0006] Furthermore, step S1 specifically includes: The data set contains a number of face images, and each face image has a corresponding forgery attribution label; Uniformly adjust the width × height of each face image in the deep forgery attribution face data set to , and normalize the face image by dividing the pixel value of each face image by 255, and encapsulate the normalized face image as a tensor representation , representing the vector space, representing the number of images in each batch, representing that the number of channels of a face image tensor is 3; The forgery attribution label corresponding to each face image is processed through the torch.tensor function in PyTorch to obtain the preprocessed forgery attribution label tensor.

[0007] Furthermore, step S4 specifically includes: Construct a multi-view visual coding module, which includes an image encoder, an edge encoder, and a noise encoder; the image encoder is a convolutional vision transformer model CviT; the edge encoder includes an edge backbone module and an edge transformer block, and the edge backbone module includes a number of stacked convolutional layers, where the output channel of the first convolutional layer with an input channel of 1 is 32, and the edge transformer block is the same as the transformer block of CviT; the noise encoder includes an image patch selector, a steganalysis enrichment model SRM, and a convolutional vision transformer model CviT; S41. Face image text parsing pair The preprocessed face image tensor in is input into the multi-view visual coding module, and after passing through the image encoder, the output dimension is of the global manipulation image of the face appearance image , , and the formula is expressed as follows: , where represents the feature dimension of the image , represents the number of images in each batch, which is the same as the number of of the face image tensor , represents the operation of the convolutional vision transformer model CviT; S42. The preprocessed face image tensor is input into the edge encoder, and after passing through the edge backbone module, the output dimension is of the local edge feature map of the face image , and the formula is expressed as follows: , where represents the convolutional neural network backbone operation, represents the parameters of the convolutional neural network backbone, represents the number of channels of the face image edge feature map, represents the height of the face image edge feature map, represents the width of the face image edge feature map; the local edge feature map of the face image is input into the edge transformer block, and the output dimension is of the global manipulation image of the face edge , ; S43. The preprocessed face image tensor Input to the noise encoder, and through the image block selector, the richest image blocks are obtained , , represents the number of images in each batch, represents the richest image blocks has 3 channels, represents the richest image blocks width × height; Input the richest image blocks to the steganalysis rich model SRM for processing to obtain noise , and the formula is as follows: , where, represents the operation of the steganalysis rich model; Input the noise to the convolutional vision transformer model CviT, and output a face noise image global manipulation image with a dimension of , , , and the formula is as follows: , S44. Fuse the face appearance image global manipulation image , the face edge global manipulation image and the face noise image global manipulation image to obtain the multi-view visual features of the global image visual fusion , and the formula is as follows: , where, represents the element-wise addition operation.

[0008] Furthermore, step S5 specifically includes: S51. The fine-grained attribution text passes through the tokenizer to obtain a sequence of word tokens, and the word tokens in the sequence of word tokens are mapped to word embedding tensors through the word embedding layer . According to the word embedding tensor and the automatically generated position of the word embedding tensor , a fine-grained attribution text sequence vector with position information is obtained , and the formula is as follows: ; S52. Construct a language encoder, and the language encoder includes consecutive transformer blocks, and each transformer block includes a multi-head attention module and a feed-forward neural network module. The layer above the multi-head attention module and the feed-forward neural network module is The normalization layer, and the next layer is a residual layer for all; Input the fine-grained attribution text sequence vector with position information into the language encoder. After being normalized by the normalization layer, it is input into the multi-head attention module of the first consecutive Transformer block for global multi-head attention calculation, and then the text global semantic features are obtained through the residual layer , and the formula is expressed as follows: , where, represents the normalization operation, represents the operation of the multi-head attention module; the text global semantic features After being normalized by the normalization layer, it is input into the feed-forward neural network module, and then the refined global language features are obtained through the residual layer , and the formula is expressed as follows: , where, represents the operation of the feed-forward neural network module; the output of the first Transformer block of the language encoder is used as the input of the second Transformer block, and the output of the second Transformer block is used as the input of the third Transformer block, and the operation is iterated multiple times until the operation of the th Transformer block is completed to obtain the fine-grained language features ; Take the last word from the fine-grained language features to obtain the global language attribution feature .

[0009] Furthermore, step S6 specifically includes: Construct a parsing encoder module, which includes a parsing backbone module and a parsing Transformer encoder; the parsing backbone module includes a stack of several convolutional layers, and the structure of the parsing backbone module is the same as that of the edge backbone module; the parsing Transformer encoder includes Q Transformer blocks; S61. Input the face parsing image into the parsing encoder module. After passing through the parsing backbone module, output the face image parsing local feature map with a dimension of , and the formula is expressed as follows: , where, Indicates the operation of the parsing encoder module; the local feature map of face parsing has a dimension of , where represents the number of channels of the face image parsing feature map, represents the height of the face image parsing feature map, and represents the width of the face image parsing feature map; S62. Flatten the local feature map of face image parsing along the channels using the reshaping function in the library into a sequence of two-dimensional noise blocks , , , where represents the number of patches, and represents the -th two-dimensional face parsing image patch; calculate the sequence of two-dimensional face parsing image patches with position information , and the formula is as follows: , where represents a learnable class tensor generated automatically, represents the mapping hidden vector, , where represents the -th two-dimensional face parsing image patch and represents the mapping hidden vector of the -th two-dimensional face parsing image patch, and represents the position of the automatically generated sequence of two-dimensional face parsing image patches; S63. Input the sequence of two-dimensional face parsing image patches into Q face parsing Transformer blocks. After passing through the first Transformer block, in the first Transformer block, pass through the multi-head self-attention module and the multi-layer perceptron module in sequence, and output the first two-dimensional face parsing global feature map , and the formula is as follows: , , where represents the output of the sequence of two-dimensional face parsing image patches after passing through the multi-head self-attention module, represents the operation of the normalization layer, represents the operation of the multi-head self-attention module, and represents the operation of the multi-layer perceptron module; output the result of the first Transformer block As the input of the second Transformer block, and then the output of the second Transformer block As the input of the third Transformer block, iterate multiple times until the operation of the Q-th Transformer block is completed, and obtain the Q-th two-dimensional face parsing global feature map ; Take out the learnable class tensor in the Q-th two-dimensional face parsing global feature map to obtain the face parsing feature .

[0010] Furthermore, step S7 specifically includes: Input the multi-view visual features into a predictor composed of fully connected layers and a multi-layer perceptron composed of fully connected layers respectively, to obtain the predicted language feature and the predicted image forgery attribution category feature , and the formula is expressed as follows: , , where, represents the predictor parameter, represents the multi-layer perceptron parameter, represents the softmax function.

[0011] Furthermore, step S8 specifically includes: Through the divergence loss function optimize the language attribution global feature and the predicted language feature , and the formula is expressed as follows: , where, represents the index of, represents the transpose, represents the language attribution global feature of the th image, represents the predicted language attribution global feature of the th image; Optimize the language attribution global feature and the multi-view visual features through the cross-modal contrast loss , , , , , wherein, represents the visual-to-language contrast loss, represents the language-to-video contrast loss, represents the th label of the sample pair of the input face images, label, represents the th language-to-visual similarity matrix, represents the cosine similarity function, represents the trainable temperature parameter, represents the th visual-to-language similarity matrix, represents the th multi-view visual feature, represents the th language attribution global feature, represents the th multi-view visual feature, represents the th language attribution global feature; The face parsing feature and the multi-view visual feature are optimized by the cross-view contrast loss function as follows: , , , , , wherein, represents the visual-to-parsing contrast loss, represents the parsing-to-visual contrast loss, represents the th parsing-to-visual similarity matrix, represents the th visual-to-parsing similarity matrix, represents the th multi-view visual feature, represents the th face parsing feature, represents the th multi-view visual feature, represents the th face parsing feature; By the deepfake attribution loss function Predict the category features of forged attribution for the predicted image and the image Generator-level label is optimized, and the formula is as follows: , wherein, represents the th image generator-level label, represents the th predicted image forged attribution category feature; According to the described divergence loss function , cross-view contrast loss , cross-modal contrast loss and depth attribution loss function to obtain the total loss , .

[0012] Furthermore, step S9 specifically includes Input the face image to be detected in the test set into the optimized and trained multi-view visual encoder module to obtain the global visual feature of the image , and then use the global visual feature of the image to calculate the final true / false discrimination result through a multi-layer perceptron and a softmax function , and the formula is as follows: , wherein, represents the multi-view visual encoder.

[0013] The advantages of the present invention are as follows: By introducing a multi-view learning and dual-modal guidance strategy, the present invention improves the accuracy and robustness of the deep forgery attribution model in the zero-shot scenario. By combining the deep forgery features of multiple visual perspectives such as images, noises, and edges, the present invention can more effectively trace the unseen generator. At the same time, with the support of face parsing and language modality, the accuracy and scalability of the attribution process are further improved, thus significantly enhancing the application value and actual effect of the deep forgery attribution technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention.

[0015] Figure 1 is the step flow chart of the method of the present invention; Figure 2 Face parsing diagrams of real images, GAN face images, and diffusion face images. Detailed implementation

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0017] Embodiment 1 In this embodiment, as Figure 1 shown, the present invention provides a zero-shot deepfake attribution method based on bimodal guidance. The specific steps include: S1. Construct a zero-shot deepfake attribution dataset and divide it into a training set and a test set; preprocess the training set to obtain preprocessed face image tensors and forgery attribution label tensors; Specifically, the dataset contains several face images, and each face image has a corresponding forgery attribution label; The width × height of each face image in the deepfake attribution face dataset is uniformly adjusted to , and the pixel values of each face image are divided by 255 to normalize the face image. The normalized face image is encapsulated as a tensor representation , represents the vector space, represents the number of images in each batch, represents that the number of channels of a face image tensor is 3; The forgery attribution label corresponding to each face image is processed by the torch.tensor function in PyTorch to obtain a preprocessed forgery attribution label tensor.

[0018] The forgery attribution labels corresponding to the face images in the training set are as follows. The first layer is the coarse-grained layer, and the labels are denoted as [b, 0] or [b, 1]. Indicates the number of images in each batch. 0 represents real images, and 1 represents fake images. The second layer is the manipulation layer, and the labels are denoted as [b,0] or [b,1] or [b,2] or [b,3]. 0 represents real images, 1 represents identity-swapped images, 2 represents attribute-manipulated images, and 3 represents full-face synthesized images. The third layer is the fine-grained layer, and the labels are denoted as [b,0] or [b,1] or [b,2] or [b,3] or [b,4] or [b,5] or [b,6]. 0 represents real images, 1 represents identity-swapped GAN images, 2 represents identity-swapped diffusion images, 3 represents attribute-manipulated GAN images, 4 represents attribute-manipulated diffusion images, 5 represents full-face synthesized GAN images, and 6 represents full-face synthesized diffusion images. The fourth layer is the generator layer, and the labels are denoted as [b,0] or [b,1] or [b,2] or [b,3] or [b,4] or [b,5] or [b,6] or [b,7] or [b,8] or [b,9] or [b,10]. 0 represents real images, 1 represents DiffFace generator images, 2 represents Diffae generator images, 3 represents LatentDiffusion generator images, 4 represents CollaborativeDiffusion generator images, 5 represents DDPM generator images, 6 represents FSLSD generator images, 7 represents FaceSwapper generator images, 8 represents LatentTransformer generator images, 9 represents StyleGAN2 generator images, and 10 represents StyleGAN3 generator images.

[0019] S2. Construct a zero-shot deepfake attribution model, which includes a fine-grained text generator, a face parser, a multi-view visual encoder module, a language encoder, a parsing encoder, a predictor, and a multi-layer perceptron; S3. Input the preprocessed forged attribution label tensor into the fine-grained text generator to obtain fine-grained attribution text ; Input the preprocessed face image tensor into the face parser to obtain a face parsing image ; The fine-grained attribution text , the face parsing image and the preprocessed face image tensor constitute a face image text parsing pair ; Figure 2 shows the prior face parsing features that distinguish between real and fake face images; it can be seen that there are obvious differences between the real parsing images and the parsing images of GAN-generated faces and the face parsing images generated by the diffusion model. Therefore, this parsing feature can be incorporated into the model as a valuable prior feature.

[0020] S4. Input the preprocessed face image tensor The multi-view visual features are obtained by inputting into the multi-view visual encoder module ; Specifically, a multi-field-of-view visual encoding module is constructed, which includes an image encoder, an edge encoder, and a noise encoder; the image encoder is the convolutional vision transformer model CviT; the edge encoder includes an edge backbone module and an edge transformer block, and the edge backbone module includes a stack of several convolutional layers, where the output channel of the first convolutional layer with an input channel of 1 is 32, and the edge transformer block is the same as the transformer block of CviT; the noise encoder includes an image patch selector, a steganalysis enrichment model SRM, and a convolutional vision transformer model CviT; S41. Face image text parsing pair The preprocessed face image tensor in is input into the multi-field-of-view visual encoding module, and after passing through the image encoder, the face appearance image global manipulation image with an output dimension of is obtained , , and the formula is as follows: , where represents the feature dimension of the image , represents the number of images in each batch, which is the same as the quantity of the face image tensor , represents the operation of the convolutional vision transformer model CviT; S42. The preprocessed face image tensor is input into the edge encoder. After passing through the edge backbone module, the face image edge local feature map with an output dimension of is obtained , and the formula is as follows: , where represents the convolutional neural network backbone operation, represents the parameters of the convolutional neural network backbone, represents the number of channels of the face image edge feature map, represents the height of the face image edge feature map, represents the width of the face image edge feature map; the face image edge local feature map is input into the edge transformer block, and the face edge global manipulation image with an output dimension of is obtained , ; S43. The preprocessed face image tensor is input into the noise encoder, and through the image block selector, the richest image blocks are obtained , , represents the number of images in each batch, represents the richest image blocks has 3 channels, represents the richest image blocks width × height; The richest image blocks are input into the steganalysis rich model SRM for processing to obtain noise , and the formula is as follows: , where, represents the operation of the steganalysis rich model; The noise is input into the convolutional vision transformer model CviT, and the output dimension is the global manipulation image of the face noise image , , and the formula is as follows: , S44. The global manipulation image of the face appearance image , the global manipulation image of the face edge and the global manipulation image of the face noise image are fused to obtain the multi-view visual features of the global image visual fusion , and the formula is as follows: , where, represents the element-wise addition operation.

[0021] S5. The fine-grained attribution text is input into the language encoder to obtain the language attribution global feature ; Specifically, S51. The fine-grained attribution text passes through the tokenizer to obtain the word token sequence, and the word tokens in the word token sequence are mapped to the word embedding tensor through the word embedding layer . According to the word embedding tensor and the position of the automatically generated word embedding tensor , the fine-grained attribution text sequence vector with position information is obtained , and the formula is as follows: ; S52. Construct a language encoder, which includes consecutive Transformer blocks, each Transformer block includes a multi-head attention module and a feed-forward neural network module, and the upper layer of both the multi-head attention module and the feed-forward neural network module is a normalization layer, and the lower layer of both is a residual layer; Input the fine-grained attribution text sequence vector with position information into the language encoder. After being normalized by the normalization layer, it is input into the multi-head attention module of the first consecutive Transformer block for global multi-head attention calculation, and then the text global semantic feature is obtained through the residual layer , and the formula is as follows: , where represents the normalization operation, represents the operation of the multi-head attention module; the text global semantic feature After being normalized by the normalization layer, it is input into the feed-forward neural network module, and then the refined global language feature is obtained through the residual layer , and the formula is as follows: , where represents the operation of the feed-forward neural network module; take the output of the first Transformer block of the language encoder as the input of the second Transformer block, and the output of the second Transformer block as the input of the third Transformer block, and iterate multiple times until the operation of the th Transformer block is completed to obtain the fine-grained language feature ; take the last word from the fine-grained language feature to obtain the language attribution global feature .

[0022] S6. Input the face parsing image into the parsing encoder to obtain the face parsing feature ; Specifically, construct a parsing encoder module, which includes a parsing backbone module and a parsing Transformer encoder; the parsing backbone module includes a stack of several convolutional layers, and the structure of the parsing backbone module is the same as that of the edge backbone module; the parsing Transformer encoder includes Q Transformer blocks; S61. Input the face parsing image Input into the parsing encoder module, passing through the parsing backbone module, and outputting a local face image parsing feature map with a dimension of The formula is as follows: , , where represents the operation of the parsing encoder module; the local face parsing feature map has a dimension of , represents the number of channels of the face image parsing feature map, represents the height of the face image parsing feature map, represents the width of the face image parsing feature map; S62. Flatten the local face image parsing feature map along the channels using the reshaping function in the library into a two-dimensional noise block sequence , , , represents the number of patches, represents the th two-dimensional face parsing image patch; calculate the two-dimensional face parsing image patch sequence with position information , and the formula is as follows: , where represents a learnable class tensor generated automatically, represents the mapping hidden vector, , represents the th two-dimensional face parsing image patch 's mapping hidden vector, represents the position of the automatically generated two-dimensional face parsing image patch sequence; S63. Input the two-dimensional face parsing image patch sequence into Q face parsing Transformer blocks. After passing through the first Transformer block, it successively passes through the multi-head self-attention module and the multi-layer perceptron module in the first Transformer block, and outputs the first two-dimensional face parsing global feature map , and the formula is as follows: , , where represents the two-dimensional face parsing image patch sequence Output after the multi-head self-attention module, represents the operation of the normalization layer, represents the operation of the multi-head self-attention module, represents the multi-layer perceptron module operation; Take the output of the first transformer block as the input of the second transformer block, and then take the output of the second transformer block as the input of the third transformer block, and iterate multiple times until the operation of the Qth transformer block is completed to obtain the Qth two-dimensional face parsing global feature map ; Take out the learnable class tensor in the Qth two-dimensional face parsing global feature map to obtain the face parsing feature .

[0023] S7. Input the multi-view visual features into the predictor and the multi-layer perceptron respectively to obtain the predicted language features and the predicted image forgery attribution category features ; Specifically, input the multi-view visual features into a predictor composed of fully connected layers and a multi-layer perceptron composed of fully connected layers respectively to obtain the predicted language features and the predicted image forgery attribution category features , and the formula is as follows: , , where, represents the predictor parameter, represents the multi-layer perceptron parameter, represents the softmax function.

[0024] S8. Optimize and train the parameters in each module through the total loss function and the Adam optimizer to obtain an optimized and trained zero-shot deep forgery attribution model; Specifically, through the divergence loss function optimize the language attribution global feature and the predicted language feature , and the formula is as follows: , where, represents the index of, represents the transpose, Denote the global language attribution features of the th image, and denote the predicted global language attribution features of the th image; Optimize the global language attribution features and the multi-view visual features through the cross-modal contrast loss function, and the formula is as follows: , , , , , where denotes the visual-to-language contrast loss, denotes the language-to-video contrast loss, denotes the th label of the sample pair of the input face image, and denotes the th language-to-visual similarity matrix, denotes the cosine similarity function, denotes the trainable temperature parameter, denotes the th visual-to-language similarity matrix, denotes the th multi-view visual feature, denotes the th global language attribution feature, denotes the th multi-view visual feature, denotes the th global language attribution feature; Optimize the face parsing features and the multi-view visual features through the cross-field contrast loss function, and the specific calculation is as follows: , , , , , where denotes the visual-to-parsing contrast loss, denotes the parsing-to-visual contrast loss Denote the th visual-similarity matrix parsed, Denote the th similarity matrix from visual to parsed, Denote the th multi-view visual feature, Denote the th face parsing feature, Denote the th multi-view visual feature, Denote the th face parsing feature; Optimize the predicted image forgery attribution category feature and the generator-level label of the image through the deepfake attribution loss function, and the formula is as follows: , wherein, Denote the th generator-level label of the image , Denote the th predicted image forgery attribution category feature; According to the divergence loss function , cross-view contrast loss , cross-modal contrast loss and deep attribution loss function to obtain the total loss , .

[0025] S9. Input the image to be detected in the test set into the multi-view visual encoder module of the zero-shot deepfake attribution model optimized by training, and then obtain the final deepfake attribution discrimination result through the multi-layer perceptron and the softmax function.

[0026] Specifically, input the face image to be detected in the test set into the optimized and trained multi-view visual encoder module to obtain the global visual feature of the image, and then calculate the final true / false discrimination result through the multi-layer perceptron and the softmax function, and the formula is as follows: , wherein, denotes the multi-view visual encoder.

[0027] Example 2 In this example, in the scenario of cross-generator evaluation of zero-shot deepfake attribution performance, the seen generators in the GenFace dataset are used for training, and then the model is tested on various unseen generators in the DF40 dataset, which simulates the detection and attribution of the model when facing deepfake content generated by unknown generators in actual applications. The existing dataset GenFace is input into the full-face synthesis generative adversarial network generator StyleGAN2, the full-face synthesis generative adversarial network generator StyleGAN3, the full-face synthesis diffusion model generator CollDiff, the full-face synthesis diffusion model generator DDPM, the full-face synthesis diffusion model generator LatDiff, the identity swapping generative adversarial network generator FSLSD, the identity swapping generative adversarial network generator FaceSwapper, the identity swapping diffusion model generator DiffFace, the attribute editing generative adversarial network generator LatTrans, and the attribute editing diffusion model generator Diffae to generate image data as the training set of the zero-shot deepfake dataset; the existing dataset DF40 is input into the full-face synthesis generative adversarial network generator VQGAN, the full-face synthesis diffusion model generator SD-2.1, the full-face synthesis diffusion model generator DiT, the identity swapping generative adversarial network generator SimSwap, the identity swapping generative adversarial network generator UniFace, the identity swapping generative adversarial network generator InSwapper, the identity swapping generative adversarial network generator e4s, and the identity swapping diffusion model generator REFace to generate image data as the test set of the zero-shot deepfake dataset. Table 1 shows the comparison of the zero-shot deepfake attribution performance of different deepfake attribution methods, which is the accuracy ACC score of different deepfake attribution methods on the unseen generators in the DF40 dataset after training with the seen generators in the GenFace dataset. Bold and underlined indicate the best result and the second-best result respectively.

[0028] Table 1 Comparison of zero-shot deepfake attribution performance of different deepfake attribution methods The MFCLIP model is a multimodal fine-grained CLIP model for general Diffusion Face Forgery Detection (DFFD), performing fine-grained language-guided image noise forgery representation learning. A threshold is defined to evaluate whether the deepfake attribution model has seen the generator. Specifically, when the maximum probability output by the model is lower than the threshold, the generator is marked as unseen. The threshold is set to 0.9. At a threshold of 0.9, the average accuracy ACC of the method of the present invention is about 28% higher than that of the MFCLIP model. The reason is that the model of the present invention introduces global face parsing features and performs dual-modal (text and parsing)-guided multi-view (image, noise, and edge) deepfake attribution representation learning. Compared with the single language-guided method of MFCLIP, it can extract features from more dimensions, has stronger detection and attribution capabilities for deepfake content, and thus performs better in test scenarios with unknown generators.

[0029] The average accuracy ACC of the method of the present invention is nearly 8%, 34%, and 59% higher than that of the Comparative Pseudo-Learning Model CPL, the Deepfake Network Architecture Detector DNA-Det, and the Detection and Attribution of Fake Images Model DE-FAKE, respectively. The Deepfake Network Architecture Detector DNA-Det only explores global attribution representations in the image modality through block-based contrastive learning, while the method of the present invention focuses on multiple modalities (text and face parsing) and multiple perspectives (image, noise, and edge forgery embeddings). This enables the method of the present invention to capture features more comprehensively and be more adaptable when facing deepfake content generated by different generators, thus showing a higher detection accuracy in the cross-generator evaluation scenario, demonstrating the effectiveness of the model of the present invention in the zero-shot deepfake attribution ZS-DFA performance.

[0030] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A zero-shot deepfake attribution method based on bimodal guidance, characterized in that Including the following steps: S1. Construct a zero-shot deepfake attribution dataset and divide it into a training set and a test set; Preprocess the training set to obtain preprocessed face image tensors and forged attribution label tensors; S2. Construct a zero-shot deepfake attribution model, which includes a fine-grained text generator, a face parser, a multi-view visual encoder module, a language encoder, a parsing encoder, a predictor, and a multi-layer perceptron; S3. Input the preprocessed forged attribution label tensor into the fine-grained text generator to obtain the fine-grained attribution text ; Input the preprocessed face image tensor into a face parser to obtain a face parsing image ; The fine-grained attribution text , the face parsing image and the preprocessed face image tensor constitute a face image text parsing pair ; S4. Input the preprocessed face image tensor into the multi-view visual encoder module to obtain multi-view visual features ; S5. Input the fine-grained attribution text into the language encoder to obtain the global language attribution feature ; S6. Input the face parsing image into the parsing encoder to obtain face parsing features ; S7. Input the multi-view visual features into the predictor and the multi-layer perceptron respectively, and obtain the predicted language features and the predicted image forgery attribution category features ; S8. Optimize and train the parameters in each module through a total loss function and an Adam optimizer to obtain an optimized and trained zero-shot deepfake attribution model; S9. Input the image to be detected in the test set into the multi-view visual encoder module of the optimized and trained zero-shot deepfake attribution model, and then pass through a multi-layer perceptron and a softmax function to obtain the final deepfake attribution discrimination result.

2. The zero-shot deepfake attribution method based on dual-modal guidance according to claim 1, wherein Step S1 specifically includes: The dataset contains a number of face images, and each face image has a corresponding forged attribution label; Uniformly adjust the width × height of each face image in the deepfake attribution face dataset to , and normalize the face image by dividing the pixel value of each face image by 255, and encapsulate the normalized face image as a tensor representation , represents the vector space, represents the number of images in each batch, represents that the number of channels of a face image tensor is 3; The forged attribution label corresponding to each face image is processed through the torch.tensor function in PyTorch to obtain a preprocessed forged attribution label tensor.

3. The zero-shot deepfake attribution method based on bimodal guidance according to claim 2, characterized in that, Step S4 specifically includes: Construct a multi-view visual encoding module, which includes an image encoder, an edge encoder, and a noise encoder; the image encoder is a convolutional vision transformer model CviT; the edge encoder includes an edge backbone module and an edge transformer block, and the edge backbone module includes a stack of several convolutional layers, where the output channel of the first convolutional layer with an input channel of 1 is 32, and the edge transformer block is the same as the transformer block of CviT; the noise encoder includes an image patch selector, a steganalysis enrichment model SRM, and a convolutional vision transformer model CviT; S41. Face image text parsing pair The preprocessed face image tensor in is input into the at most field-of-view visual coding module. After passing through the image encoder, the output is a global manipulation image of the face appearance image with a dimension of , , , which is expressed by the formula as follows: , Among them, represents the feature dimension of the image , represents the number of images in each batch, which is the same as the number of human face image tensors ; represents the operation of the convolutional vision transformer model CviT. Preprocessed face image tensor is input into the edge encoder, passes through the edge backbone module, and outputs a face image edge local feature map with a dimension of , which is expressed by the following formula: , as shown below: , Among them, represents the backbone operation of the convolutional neural network, represents the parameters of the backbone of the convolutional neural network, represents the number of channels of the edge feature map of the face image, represents the height of the edge feature map of the face image, represents the width of the edge feature map of the face image; input the local edge feature map of the face image into the edge transformer block, and output the face edge global manipulation image with a dimension of , ;​ Preprocessed face image tensor Input it into the noise encoder, and through the image patch selector, the richest image patches are obtained , , represents the number of images in each batch, represents the richest image patches whose number of channels is 3, represents the richest image patches width × height; Input the richest image patches into the steganalysis rich model SRM for processing to obtain noise , and the formula is as follows: , Among them, represents the operation of the steganalysis enrichment model; the noise is input into the convolutional vision transformer model CviT, and the output dimension is the global manipulation image of the face noise image , , and the formula is expressed as follows: , S44. Global manipulation of the face appearance image, the global manipulation image of the face edge , and the global manipulation image of the face noise image are fused to obtain the multi-view visual features of the global image visual fusion , which is expressed by the formula as follows: ​ , Among them, represents an element-wise addition operation.

4. A zero-shot deepfake attribution method based on bimodal guidance according to claim 3, characterized in that Step S5 specifically includes: S51. Fine-grained attribution text After passing through a tokenizer to obtain a sequence of word tokens, the word tokens in the sequence of word tokens are mapped to word embedding tensors through a word embedding layer , according to the word embedding tensor and the positions of the automatically generated word embedding tensors obtain a fine-grained attribution text sequence vector with position information , which is expressed by the formula as follows: ; S52. Construct a language encoder, which includes consecutive Transformer blocks, each of which includes a multi-head attention module and a feed-forward neural network module. The upper layer of both the multi-head attention module and the feed-forward neural network module is a normalization layer, and the lower layer of both is a residual layer; Input the fine-grained attribution text sequence vector with position information into the language encoder. After passing through the normalization layer for normalization, it is input into the multi-head attention module of the first consecutive Transformer block for global multi-head attention calculation, and then the global semantic features of the text are obtained through the residual layer , and the formula is as follows: , Among them, represents the normalization operation, represents the operation of the multi-head attention module; the global semantic features of the text after passing through the normalization layer for normalization and then input into the feed-forward neural network module, and then refined global language features are obtained through the residual layer , which is expressed by the formula as follows: , Among them, represents the operation of the feed-forward neural network module; the output of the first transformer block of the language encoder is used as the input of the second transformer block, and the output of the second transformer block is used as the input of the third transformer block, and the iteration is performed multiple times until the operation of the th transformer block is completed to obtain fine-grained language features ; the last word is taken from the fine-grained language features to obtain the global language attribution feature . .

5. A zero-shot deepfake attribution method based on bimodal guidance according to claim 4, characterized in that, Step S6 specifically includes: Construct a parsing encoder module, which includes a parsing backbone module and a parsing transformer encoder; the parsing backbone module includes a stack of several convolutional layers, and the structure of the parsing backbone module is the same as that of the edge backbone module; the parsing transformer encoder includes Q transformer blocks; S61. Input the face parsing image into the parsing encoder module, and after passing through the parsing backbone module, output a face image parsing local feature map with a dimension of , which is expressed by the following formula: , Among them, represents the operation of the parsing encoder module; the local feature map of face parsing has a dimension of , represents the number of channels of the face image parsing feature map, represents the height of the face image parsing feature map, represents the width of the face image parsing feature map; S62. Parse the local feature map of the face image Along the channel, use in the library the reshape function to flatten it into a sequence of two-dimensional noise blocks , , denote the number of patches, denote the th two-dimensional face parsing image patch; calculate the sequence of two-dimensional face parsing image patches with position information , and the formula is as follows: , Among them, represents an automatically generated learnable tensor-like object, represents a mapping hidden vector, , represents the th two-dimensional face parsing image patch 's mapping hidden vector, represents the position of the automatically generated sequence of two-dimensional face parsing image patches; S63. Input the two-dimensional face parsing image patch sequence into Q face parsing Transformer blocks. After passing through the first Transformer block, it successively passes through the multi-head self-attention module and the multi-layer perceptron module in the first Transformer block, and outputs the first two-dimensional face parsing global feature map , which is expressed by the formula as follows: , , Among them, represents a sequence of two-dimensional face parsing image blocks which is the output of the multi-head self-attention module, represents the operation of the normalization layer, represents the operation of the multi-head self-attention module, represents a multi-layer perceptron module operation; taking the output of the first transformer block as the input of the second transformer block, and then taking the output of the second transformer block as the input of the third transformer block, and iterating multiple times until the operation of the Qth transformer block is completed to obtain the Qth two-dimensional face parsing global feature map ; taking out the learnable class tensor in the Qth two-dimensional face parsing global feature map to obtain the face parsing feature .

6. A zero-shot deepfake attribution method based on dual-modal guidance according to claim 5, characterized in that Step S7 specifically includes: Input multi-view visual features into a predictor composed of fully connected layers and a multi-layer perceptron composed of fully connected layers respectively to obtain predicted language features and predicted image forgery attribution category features , which is expressed by the formula as follows: , , Among them, represents the predictor parameter, represents the multi-layer perceptron parameter, represents the softmax function.

7. A zero-shot deepfake attribution method based on bimodal guidance according to claim 6, characterized in that, Step S8 specifically includes: Via Divergence loss function Optimize the global language attribution features And the predicted language features The formula is as follows: , in, express The index of represents transpose, Indicates The language attributed global features of the image, Indicates Predicted language attribution global features for each image; Through cross-modal contrastive loss The function attributes global features to language And multi-view visual features For optimization, the formula is as follows: , , , , , Among them, represents the visual-to-language contrast loss, represents the language-to-video contrast loss, represents the label of the sample pair of the th input face image, represents the th language-to-visual similarity matrix, represents the cosine similarity function, represents the trainable temperature parameter, represents the th visual-to-language similarity matrix, represents the th multi-view visual feature, represents the th language attribution global feature, represents the th multi-view visual feature, represents the th language attribution global feature; By means of a cross-field contrast loss function Optimize the face parsing features And multi-view visual features The specific calculation is as follows: , , , , , Among them, represents the visual-to-parsing contrast loss, represents the parsing-to-visual contrast loss, represents the th parsing-to-visual similarity matrix, represents the th visual-to-parsing similarity matrix, represents the th multi-view visual feature, represents the th face parsing feature, represents the th multi-view visual feature, represents the th face parsing feature; Attribution Loss Function for Deepfakes Attribution category features of the predicted image forgery And the image Generator-level label Are optimized, and the formula is as follows: , Among them, represents the th image generator-level label, represents the th predicted image forgery attribution category feature; According to the divergence loss function , cross-field contrast loss , cross-modal contrast loss and depth attribution loss function to obtain the total loss , .

8. A zero-shot deepfake attribution method based on dual-modal guidance according to claim 7, characterized in that, Step S9 specifically includes Input the face images to be detected in the test set into the optimized and trained multi-view visual encoder module to obtain the global visual features of the images , and then use the global visual features of the images to calculate the final true / false discrimination result through a multi-layer perceptron and a softmax function . The formula is as follows: , Among them, represents a multi-view visual encoder.

Citation Information

Patent Citations

  • False face recognition method and device and computer readable storage medium

    CN111783505A

  • Counterfeit face detection method and device based on difference perception element learning

    CN112784781A

  • Counterfeit information detection method combining bimodal understanding and large language model

    CN119003767A

  • Facial expression-based detection method for deepfake by generative artificial intelligence (AI)

    US20240378921A1

  • Forensics method for synthesized face image based on local binary pattern and deep learning

    WO2021134871A1