Diffusion counterfeited face detection method based on multi-modal fine-grained CLIP

Through the multimodal fine-grained CLIP method, the visual and language features of face images are extracted, and the problem of difficult detection of fake face images synthesized by diffusion models is solved in the prior art, and more efficient and accurate fake face detection is achieved.

CN119942616APending Publication Date: 2025-05-06QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +3
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510060439.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect highly realistic forged face images based on diffusion model synthesis, and insufficient exploration of fine-grained noise and text modalities limits the generalization ability of the model.

Method used

The diffusion forged face detection method based on multimodal fine-grained CLIP is adopted. By constructing a layered fine-grained face data set, visual and language features are extracted using multimodal vision encoder and fine-grained language encoder, and the attention module and multi-layer perceptron are optimized through samples to achieve cross-modal detection.

Benefits of technology

It improves the detection ability of diffusion synthetic facial images, enhances the generalization ability and detection accuracy of the model, and can effectively identify fake facial images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942616A_ABST
    Figure CN119942616A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of diffusion counterfeit face detection methods, in particular to a diffusion counterfeit face detection method based on multi-modal fine-grained CLIP, and the method specifically comprises the following steps: constructing a layered fine-grained face data set, and carrying out the preprocessing of a face image in the data set and a corresponding layered fine-grained label; a layered fine-grained label tensor obtained after preprocessing is input into a fine-grained text generator to obtain a fine-grained text # imgabs0 #, and the fine-grained text # imgabs0 # and a face image tensor # imgabs1 # form a face image text pair # imgabs2 #; respectively processing the face image text pair # imgabs3 # to obtain a predicted language feature # imgabs4 # and a predicted image category feature # imgabs5 #; parameters in all the modules are optimized and trained through a loss function and an Adam optimizer, and a CLIP multi-mode fine-grained model after optimization training is obtained; and inputting a to-be-detected image into the trained and optimized CLIP multi-mode fine-grained model to obtain a final true and false discrimination result. According to the invention, the efficiency and the cross-modal detection capability of the diffusion counterfeit face detection model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of diffusion forged face detection methods, and in particular to a diffusion forged face detection method based on multimodal fine-grained CLIP. Background Art

[0002] With the rapid development of advanced face generation technologies such as diffusion models, forged face images are becoming more and more realistic, causing great concern in society and academia. Therefore, there is an urgent need to develop a robust and universal face forgery detection technology.

[0003] Current methods mainly rely on image modalities to capture the characteristics of facial forgery, and have not fully explored other important modalities such as fine-grained noise and text, which limits the generalization ability of the model. In addition, most facial forgery detection methods mainly identify facial images generated by generative adversarial networks (GANs), but have difficulty detecting emerging facial images synthesized based on diffusion models. Moreover, the contrastive language image pre-training model CLIP tends to maximize the distance between related negative sample pairs, which is considered a suboptimal feature alignment scheme because related negative sample pairs should be closer to each other in the feature space.

[0004] Therefore, the present invention provides a diffusion forged face detection method based on multimodal fine-grained CLIP to solve the above problems. Summary of the invention

[0005] In view of the deficiencies in the prior art, the present invention develops a diffusion forged face detection method based on multimodal fine-grained CLIP. The main purpose of the present invention is to improve the efficiency of the diffusion forged face detection model and the cross-modal detection capability.

[0006] The technical solution to the technical problem solved by the present invention is a diffusion forged face detection method based on multimodal fine-grained CLIP, which comprises the following steps: S1. Constructing a hierarchical and fine-grained face dataset , The dataset contains several face images, each of which has a corresponding hierarchical fine-grained label. The face images and the corresponding hierarchical fine-grained labels are preprocessed respectively to obtain the preprocessed face images and the corresponding hierarchical fine-grained label tensors. S2: Input the hierarchical fine-grained label tensor obtained after preprocessing into the fine-grained text generator to obtain fine-grained text , and the preprocessed face image tensor Forming face image text pairs ; S3, the preprocessed face image tensor Input into the multimodal visual encoder module to obtain multimodal visual features ; S4. Fine-grained text Input into the fine-grained language encoder to obtain fine-grained language features ; S5. Multimodal visual features and fine-grained language features They are input into the sample pair attention module to obtain the visual-language similarity matrix and the language-visual similarity matrix respectively; S6. Multimodal visual features of faces Input into the predictor and multi-layer perceptron respectively to obtain the predicted language features and predicted image category features ; S7. Utilization Divergence loss function Fine-grained language features and predictive language features Optimize using cross-modal contrast loss Function-to-sample similarity matrix Optimize and then use the manipulation detection loss function For predicting image category features Optimize with the real image labels and summarize to get the total loss function. Use the total loss function and Adam optimizer to optimize and train the parameters in each module to get the optimized and trained CLIP multimodal fine-grained model. S8. Input the image to be detected into the trained and optimized CLIP multimodal fine-grained model to obtain the final true or false discrimination result.

[0007] S1 is as follows: Hierarchical Fine-Grained Face Dataset The width × height of each face image in is uniformly adjusted to , and divide each face image pixel value by 255 to normalize the face image, and encapsulate the normalized face image as Tensor representation of , represents a vector space, represents any face image tensor in the preprocessed hierarchical fine-grained face dataset, and 3 represents the face image tensor The number of channels is 3; The hierarchical fine-grained labels corresponding to face images are as follows. The first layer is the coarse-grained layer, and the labels are recorded as or , represents the number of images in each batch, 0 represents a true image, and 1 represents a false image; the second layer is the manipulation layer, and the label is recorded as or or or , 0 represents the real image, 1 represents the identity exchange image, 2 represents the attribute manipulation image, and 3 represents the full face synthesis image; the third layer is the fine-grained layer, and the label is recorded as or or or or or or , 0 represents the real image, 1 represents the identity swap GAN image, 2 represents the identity swap diffusion image, 3 represents the attribute manipulation GAN image, 4 represents the attribute manipulation diffusion image, 5 represents the full face synthesis GAN image, and 6 represents the full face synthesis diffusion image; the fourth layer is the generator layer, and the label is recorded as or or or or or or or or or or , 0 represents the real image, 1 represents the image of the DiffFace face swapping framework, 2 represents the image of the Diffae diffusion model, 3 represents the image of the generated model in the LatentDiffusion deep learning, 4 represents the image of the CollaborativeDiffusion diffusion model framework, 5 represents the image of the DDPM denoising diffusion probability model, 6 represents the image of the FSLSD limited slip differential, 7 represents the FaceSwapper face swapping model image, 8 represents the image of the LatentTransformer image generation model, 9 represents the image of the StyleGAN2 model improved based on the generative adversarial network, and 10 represents the image of the StyleGAN3 model improved based on the generative adversarial network.

[0008] S2 is as follows: Design a fine-grained text generator module, which is based on The hierarchical fine-grained labels provided by the dataset generate corresponding text hints for each face image , to construct image-text pairs; The fine-grained text generator module corresponds to the hierarchical fine-grained labels. The first layer of the fine-grained text generator module is the coarse-grained layer, the second layer is the manipulation layer, the third layer is the fine-grained layer, and the fourth layer is the generator layer. The fine-grained text generator module analyzes the label content layer by layer according to the input hierarchical fine-grained labels, and then generates the corresponding text prompts. , and then with the preprocessed face image tensor Together they form a face image text pair .

[0009] S3 is as follows: S3.1. Construct a multimodal visual encoding module, which includes an image encoder, an image block selector and a noise encoder, wherein the image encoder is composed of a convolutional visual attention model CviT, and the noise encoder is composed of a convolutional neural network backbone and L noise attention blocks; S3.2. Pairing face image with text The face image tensor in Input to the multimodal visual encoding module, after the convolutional visual attention model CviT, the output dimension is Global manipulation of face images , , the calculation process is as follows: , in, Representing images The characteristic dimension of Represents the number of images in each batch, and the face image tensor of The number is consistent, Represents the operation of the convolutional visual attention model CviT; Global manipulation of face images After the image block selector, the image Divided into non-overlapping image blocks, through the gray-level co-occurrence matrix calculate The homogeneity scores of the image patches are calculated, and the image patch with the lowest score is selected as the richest image patch. , , Indicates the number of images in each batch, Represents the richest image patch The number of channels is 3, Represents the richest image patch Width × height; The richest image patch Output noise of input to steganalysis enrichment model SRM , the calculation formula is as follows: , in, Represents the operations of the steganalysis enrichment model; Then the noise The input to the noise encoder module first passes through the convolutional neural network backbone, and the output dimension is Local noise map of face image , the calculation process is as follows: , in, represents the backbone operation of the convolutional neural network, Represents the parameters of the CNN backbone and the local noise map of the face image The dimension is , Indicates the number of channels of the local feature map of face noise, Indicates the height of the local feature map of face noise, Indicates the width of the local feature map of face noise; Then the local feature map of face noise Utilize along the channel Curry's The reshape function flattens the sequence into two-dimensional noise blocks. , , express The number of patches, Indicates A two-dimensional noise image patch; Then calculate the two-dimensional noise block sequence with position information , the calculation process is as follows: , in, represents the automatically generated learnable class tensor, represents the mapping latent vector, , Indicates 2D noisy image patch The mapping hidden vector of represents the position of the automatically generated two-dimensional noise image block sequence; Then the two-dimensional noise block sequence Input to L noise attention blocks, First, it passes through the first noise attention block, and then passes through the multi-head self-attention module and Module, and finally output the first two-dimensional noise global feature map , the calculation process is as follows: , , in, represents the operation of the normalization layer, represents the operation of the multi-head self-attention module, express Operation of the module; Then the output of the first noise attention block is As the input of the second noise attention block, the output of the second noise attention block is As the input of the third noise attention block, iterate multiple times until the Lth noise attention block outputs , represents the L-th two-dimensional noise global feature map; Finally, the L-th two-dimensional noise global feature map The learnable class tensor in is taken out to obtain the global noise feature ,Will and Fusion is performed to obtain the multimodal visual features of global image noise fusion , the calculation process is as follows: .

[0010] S4 is as follows: S4.1. Fine-grained text First, the word tag sequence is obtained through the word segmenter, and then the word tag is mapped into a word embedding tensor through the word embedding layer , and then according to the word embedding tensor and the position of the automatically generated word embedding tensor Get fine-grained class-aware text sequence vector with position information , the calculation process is as follows: ; S4.2. Construct a fine-grained language encoder. The fine-grained language encoder includes There are continuous attention modules, each of which includes a multi-head attention module and a feedforward neural network module. The previous layer of the multi-head attention module and the feedforward neural network module are Normalization layer, the next layer is the residual layer; S4.3. Fine-grained class-aware text sequence vector with position information Input into the fine-grained language encoder, First, after the first continuous attention module, go through After the normalization operation, the normalization layer is input into the multi-head attention module of the first continuous attention module for global multi-head attention calculation, and then the global semantic features of the text are obtained through the residual layer. , the calculation process is as follows: , in, represents the normalization operation, Represents the operation of the multi-head attention module; Then Go through The normalization layer is normalized and then input into the feedforward neural network module, and then passes through the residual layer to obtain the refined global visual language fusion features. , the calculation process is as follows: , in, Represents the operation of a feedforward neural network module; Then the output of the first layer attention module of the language encoder is used as the input of the second attention module, and the output of the second layer attention module is used as the input of the third attention module, and iterates multiple times until the first The output of the attention module The final output of the fine-grained language encoder is the fine-grained language feature. Take the last word Get fine-grained language features .

[0011] S5 is as follows: Construct a sample pair attention module, which consists of a learnable sample pair attention parameter matrix to convert multimodal visual features and fine-grained language features They are input into the sample pair attention module respectively to obtain the visual language similarity matrix and the language visual similarity matrix. The calculation process is as follows: , , in, represents the visual language similarity matrix, represents the language-visual similarity matrix, represents element-wise dot product, represents the sigmoid activation function, represents the learnable weight matrix.

[0012] S6 is as follows: Will Input to a predictor composed of fully connected layers to obtain predicted language features , the calculation process is as follows: , in, represents the predictor parameters; Then Input into a multi-layer perceptron composed of fully connected layers to obtain the predicted image category features , the calculation process is as follows: , in, Represents the parameters of the multilayer perceptron.

[0013] S7 is as follows: pass Divergence loss function Fine-grained language features and predictive language features Optimize, the specific calculation is as follows: , in, Indicates the number of face images involved in training, express The index of represents transpose, Indicates Fine-grained language features of an image, Indicates The predicted language features of the image, Represents the Softmax activation function; Through cross-modal contrast loss Function-to-sample similarity matrix Optimize the sample pair similarity matrix , the specific calculation is as follows: , , , , , in, represents the contrastive loss from vision to language, represents the contrastive loss from language to video, Indicates The sample pairs of input face images Label, Indicates A language-to-video similarity matrix, express Another index of represents the cosine similarity function, represents a trainable temperature parameter, Indicates The video-to-visual language similarity matrix, Indicates Multimodal visual features , Indicates Fine-grained language features , Indicates Multimodal visual features , Indicates Fine-grained language features ; Then, by manipulating the detection loss function For predicting image category features With real image Coarse-grained labeling Optimize, the specific calculation is as follows: , in, Indicates Real images Coarse-grained labels, Indicates Predicted image category features; According to Divergence loss function , cross-modal contrast loss and manipulation detection loss function Get the total loss , the calculation process is as follows: .

[0014] S8 is as follows: The face image to be detected is input into the optimized and trained CLIP multimodal fine-grained model to obtain the global visual features of the image. , and then The final true and false discrimination result is obtained by multi-layer perceptron calculation , the calculation process is as follows: , in, represents the multimodal visual encoder, Represents a multilayer perceptron.

[0015] The effects provided in the summary of the invention are only the effects of the embodiments, rather than all the effects of the invention. The above technical solution has the following advantages or beneficial effects: (1) Improved generalization ability: Based on the advanced contrastive language-image pre-training model, the designed multimodal fine-grained CLIP technology can detect diffuse synthetic facial images, not just those generated by GAN, thus improving the generalization ability of the model; (2) Fine-grained noise encoder: The fine-grained noise encoder extracts fine-grained noise forgery patterns from the richest areas of the image, improving the accuracy and generalization of the model in detecting facial forgeries; (3) Innovative plug-and-play sample pair attention method: It can emphasize relevant negative sample pairs and suppress irrelevant negative sample pairs, so that cross-modal sample pairs can be aligned more flexibly, which further enhances the detection performance of the model. This module can be plugged and played into the visual language model, which can significantly improve the performance of the visual language model by adding only a small amount of parameters and computational complexity.

[0016] In summary, the present invention uses the visual manipulation features learned by the multimodal visual encoder to detect diffuse forged faces, which can improve the forgery detection efficiency and cross-modal detection capability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0018] Figure 1 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION

[0019] In order to clearly illustrate the technical features of the present invention, the present invention is described in detail below through specific implementation methods and in conjunction with the accompanying drawings.

[0020] Example 1 A diffusion forged face detection method based on multimodal fine-grained CLIP includes the following steps: S1. Constructing a hierarchical and fine-grained face dataset , The dataset contains several face images, each of which has a corresponding hierarchical fine-grained label. The face images and the corresponding hierarchical fine-grained labels are preprocessed respectively to obtain the preprocessed face images and the corresponding hierarchical fine-grained label tensors. S2: Input the hierarchical fine-grained label tensor obtained after preprocessing into the fine-grained text generator to obtain fine-grained text , and the preprocessed face image tensor Forming face image text pairs ; S3, the preprocessed face image tensor Input into the multimodal visual encoder module to obtain multimodal visual features ; S4. Fine-grained text Input into the fine-grained language encoder to obtain fine-grained language features ; S5. Multimodal visual features and fine-grained language features They are input into the sample pair attention module to obtain the visual-language similarity matrix and the language-visual similarity matrix respectively; S6. Multimodal visual features of faces Input into the predictor and multi-layer perceptron respectively to obtain the predicted language features and predicted image category features ; S7. Utilization Divergence loss function Fine-grained language features and predictive language features Optimize using cross-modal contrast loss Function-to-sample similarity matrix Optimize and then use the manipulation detection loss function For predicting image category features Optimize with the real image labels and summarize to get the total loss function. Use the total loss function and Adam optimizer to optimize and train the parameters in each module to get the optimized and trained CLIP multimodal fine-grained model. S8. Input the image to be detected into the trained and optimized CLIP multimodal fine-grained model to obtain the final true or false discrimination result.

[0021] S1 is as follows: Hierarchical Fine-Grained Face Dataset The width × height of each face image in is uniformly adjusted to , and divide each face image pixel value by 255 to normalize the face image, and encapsulate the normalized face image as Tensor representation of , represents a vector space, represents any face image tensor in the preprocessed hierarchical fine-grained face dataset, and 3 represents the face image tensor The number of channels is 3; The hierarchical fine-grained labels corresponding to face images are as follows. The first layer is the coarse-grained layer, and the labels are recorded as or , represents the number of images in each batch, 0 represents a true image, and 1 represents a false image; the second layer is the manipulation layer, and the label is recorded as or or or , 0 represents the real image, 1 represents the identity exchange image, 2 represents the attribute manipulation image, and 3 represents the full face synthesis image; the third layer is the fine-grained layer, and the label is recorded as or or or or or or , 0 represents the real image, 1 represents the identity swap GAN image, 2 represents the identity swap diffusion image, 3 represents the attribute manipulation GAN image, 4 represents the attribute manipulation diffusion image, 5 represents the full face synthesis GAN image, and 6 represents the full face synthesis diffusion image; the fourth layer is the generator layer, and the label is recorded as or or or or or or or or or or , 0 represents the real image, 1 represents the image of the DiffFace face swapping framework, 2 represents the image of the Diffae diffusion model, 3 represents the image of the generated model in the LatentDiffusion deep learning, 4 represents the image of the CollaborativeDiffusion diffusion model framework, 5 represents the image of the DDPM denoising diffusion probability model, 6 represents the image of the FSLSD limited slip differential, 7 represents the FaceSwapper face swapping model image, 8 represents the image of the LatentTransformer image generation model, 9 represents the image of the StyleGAN2 model improved based on the generative adversarial network, and 10 represents the image of the StyleGAN3 model improved based on the generative adversarial network.

[0022] S2 is as follows: Design a fine-grained text generator module, which is based on The hierarchical fine-grained labels provided by the dataset generate corresponding text hints for each face image , to construct image-text pairs; The fine-grained text generator module corresponds to the hierarchical fine-grained labels. The first layer of the fine-grained text generator module is the coarse-grained layer, the second layer is the manipulation layer, the third layer is the fine-grained layer, and the fourth layer is the generator layer. The fine-grained text generator module analyzes the label content layer by layer according to the input hierarchical fine-grained labels, and then generates the corresponding text prompts. , and then with the preprocessed face image tensor Together they form a face image text pair .

[0023] Specifically, a corresponding text description is created for each level label of the image. In the first layer, the coarse-grained label layer, a text description is generated, label 0 is "a photo of a real face", label 1 is "a photo of a fake face"; the second layer is the manipulation layer, which divides the forged face images into three types, label 1 corresponds to "a photo of a face with swapped identities", label 2 corresponds to "a photo of a face with attribute manipulation", and label 3 corresponds to "a photo of a face synthesized with the whole face"; the third layer generates text prompts, label 1 corresponds to "a photo generated by a diffusion model", label 2 corresponds to "a photo generated by a GAN-based Model-generated photos"; in the fourth generator layer, label 3 corresponds to the description "the source generation model of this photo is LatentDiffusion", label 4 corresponds to "the source generation model of this photo is CollaborativeDiffusion", label 5 corresponds to "the source generation model of this photo is DDPM", label 1 corresponds to "the source generation model of this photo is DiffFace", label 9 corresponds to "the source generation model of this photo is StyleGAN2", label 0 corresponds to "the source generation model of this photo is StyleGAN3", label 6 corresponds to "the source generation model of this photo is FSLSD", label 7 corresponds to "the source generation model of this photo is FaceSwapper", label 8 corresponds to "the source generation model of this photo is LatentTransformer", label 2 corresponds to "the source generation model of this photo is Diffae", and the image-text pairs are input into the multimodal visual encoder and the fine-grained language encoder respectively.

[0024] S3 is as follows: S3.1. Construct a multimodal visual encoding module, which includes an image encoder, an image block selector and a noise encoder, wherein the image encoder is composed of a convolutional visual attention model CviT, and the noise encoder is composed of a convolutional neural network backbone and L noise attention blocks; S3.2. Pairing face image with text The face image tensor in Input to the multimodal visual encoding module, after the convolutional visual attention model CviT, the output dimension is Global manipulation of face images , , the calculation process is as follows: , in, Representing images The characteristic dimension of Represents the number of images in each batch, and the face image tensor of The number is consistent, Represents the operation of the convolutional visual attention model CviT; Global manipulation of face images After the image block selector, the image Divided into non-overlapping image blocks, through the gray-level co-occurrence matrix calculate The homogeneity scores of the image patches are calculated, and the image patch with the lowest score is selected as the richest image patch. , , Indicates the number of images in each batch, Represents the richest image patch The number of channels is 3, Represents the richest image patch Width × height; The richest image patch Output noise of input to steganalysis enrichment model SRM , the calculation formula is as follows: , in, Represents the operations of the steganalysis enrichment model; Then the noise The input to the noise encoder module first passes through the convolutional neural network backbone, and the output dimension is Local noise map of face image , the calculation process is as follows: , in, represents the backbone operation of the convolutional neural network, Represents the parameters of the CNN backbone and the local noise map of the face image The dimension is , Indicates the number of channels of the local feature map of face noise, Indicates the height of the local feature map of face noise, Indicates the width of the local feature map of face noise; Then the local feature map of face noise Utilize along the channel Curry's The reshape function flattens the sequence into two-dimensional noise blocks. , , express The number of patches, Indicates A two-dimensional noise image patch; Then calculate the two-dimensional noise block sequence with position information , the calculation process is as follows: , in, represents the automatically generated learnable class tensor, represents the mapping latent vector, , Indicates 2D noisy image patch The mapping hidden vector of represents the position of the automatically generated two-dimensional noise image block sequence; Then the two-dimensional noise block sequence Input to L noise attention blocks, First, it passes through the first noise attention block, and then passes through the multi-head self-attention module and Module, and finally output the first two-dimensional noise global feature map , the calculation process is as follows: , , in, represents the operation of the normalization layer, represents the operation of the multi-head self-attention module, express Operation of the module; Then the output of the first noise attention block is As the input of the second noise attention block, the output of the second noise attention block is As the input of the third noise attention block, iterate multiple times until the Lth noise attention block outputs , represents the L-th two-dimensional noise global feature map; Finally, the L-th two-dimensional noise global feature map The learnable class tensor in is taken out to obtain the global noise feature ,Will and Fusion is performed to obtain the multimodal visual features of global image noise fusion , the calculation process is as follows: .

[0025] S4 is as follows: S4.1. Fine-grained text First, the word tag sequence is obtained through the word segmenter, and then the word tag is mapped into a word embedding tensor through the word embedding layer , and then according to the word embedding tensor and the position of the automatically generated word embedding tensor Get fine-grained class-aware text sequence vector with position information , the calculation process is as follows: ; S4.2. Construct a fine-grained language encoder. The fine-grained language encoder includes There are continuous attention modules, each of which includes a multi-head attention module and a feedforward neural network module. The previous layer of the multi-head attention module and the feedforward neural network module are Normalization layer, the next layer is the residual layer; S4.3. Fine-grained class-aware text sequence vector with position information Input into the fine-grained language encoder, First, after the first continuous attention module, go through After the normalization operation, the normalization layer is input into the multi-head attention module of the first continuous attention module for global multi-head attention calculation, and then the global semantic features of the text are obtained through the residual layer. , the calculation process is as follows: , in, represents the normalization operation, Represents the operation of the multi-head attention module; Then Go through The normalization layer is normalized and then input into the feedforward neural network module, and then passes through the residual layer to obtain the refined global visual language fusion features. , the calculation process is as follows: , in, Represents the operation of a feedforward neural network module; Then the output of the first layer attention module of the language encoder is used as the input of the second attention module, and the output of the second layer attention module is used as the input of the third attention module, and iterates multiple times until the first The output of the attention module The final output of the fine-grained language encoder is the fine-grained language feature. Take the last word Get fine-grained language features .

[0026] S5 is as follows: Construct a sample pair attention module, which consists of a learnable sample pair attention parameter matrix to convert multimodal visual features and fine-grained language features They are input into the sample pair attention module respectively to obtain the visual language similarity matrix and the language visual similarity matrix. The calculation process is as follows: , , in, represents the visual language similarity matrix, represents the language-visual similarity matrix, represents element-wise dot product, represents the sigmoid activation function, represents the learnable weight matrix.

[0027] S6 is as follows: Will Input to a predictor composed of fully connected layers to obtain predicted language features , the calculation process is as follows: , in, represents the predictor parameters; Then Input into a multi-layer perceptron composed of fully connected layers to obtain the predicted image category features , the calculation process is as follows: , in, Represents the parameters of the multilayer perceptron.

[0028] S7 is as follows: pass Divergence loss function Fine-grained language features and predictive language features Optimize, the specific calculation is as follows: , in, Indicates the number of face images involved in training, express The index of represents transpose, Indicates Fine-grained language features of an image, Indicates The predicted language features of the image, Represents the Softmax activation function; Through cross-modal contrast loss Function-to-sample similarity matrix Optimize the sample pair similarity matrix , the specific calculation is as follows: , , , , , in, represents the contrastive loss from vision to language, represents the contrastive loss from language to video, Indicates The sample pairs of input face images Label, Indicates A language-to-video similarity matrix, express Another index of represents the cosine similarity function, represents a trainable temperature parameter, Indicates The video-to-visual language similarity matrix, Indicates Multimodal visual features , Indicates Fine-grained language features , Indicates Multimodal visual features , Indicates Fine-grained language features ; Then, by manipulating the detection loss function For predicting image category features With real image Coarse-grained labeling Optimize, the specific calculation is as follows: , in, Indicates Real images Coarse-grained labels, Indicates Predicted image category features; According to Divergence loss function , cross-modal contrast loss and manipulation detection loss function Get the total loss , the calculation process is as follows: .

[0029] S8 is as follows: The face image to be detected is input into the optimized and trained CLIP multimodal fine-grained model to obtain the global visual features of the image. , and then The final true and false discrimination result is obtained by multi-layer perceptron calculation , the calculation process is as follows: , in, represents the multimodal visual encoder, Represents a multilayer perceptron.

[0030] Example 2 Taking this case background as an example, the method of the present invention is used as follows: Case background: On a well-known international news website, a report claimed to have captured a video clip of a political leader B making extreme remarks in private. The video quickly fermented on the Internet, causing widespread controversy and public opinion storm. Political leader B and his team strongly denied this and pointed out that the video footage was likely carefully forged. After preliminary analysis by the technical team, the video used the latest deep forgery technology, especially combined with advanced diffusion models, making the forged footage almost indistinguishable from the real thing. Even professionals find it difficult to distinguish the authenticity with the naked eye.

[0031] Application process and situation: (1) Preliminary screening and difficulties: The news website first used the existing deep learning-based face forgery detection tool to conduct a preliminary screening of the video. However, due to the fact that the tool was designed mainly for traditional GAN ​​forgery technology, the detection effect was not ideal when faced with the highly realistic images generated by the diffusion model, and it was unable to effectively distinguish the authenticity.

[0032] (2) Introducing multimodal fine-grained analysis technology: In view of the limitations of the initial screening, the news website decided to introduce the MFCLIP model proposed in this paper, which combines multimodal fine-grained analysis of image-noise modalities, advanced contrastive language-image pre-training (CLIP) model and innovative sample pair attention (SPA) method. By introducing language-guided facial forgery representation learning, the MFCLIP model can mine comprehensive and fine-grained forgery traces across modalities, effectively dealing with forged images generated by the diffusion model.

[0033] (3) Accurate positioning and verification: With the help of multimodal fine-grained CLIP analysis technology, the system successfully identified and accurately located the diffuse forged face area in the video. The news website then submitted the detection results to an independent third-party verification agency for review to ensure the accuracy and fairness of the detection.

[0034] (4) Quick response and follow-up measures: Once the video was confirmed to be fake, the news website immediately removed it from the shelves, issued a clarification statement to the public, and initiated an investigation into the rumor maker. In addition, the website also strengthened cooperation with the technical team and continued to optimize detection technology, especially the ability to identify new fake technologies, to ensure that it can respond more effectively to similar incidents in the future.

[0035] (5) Social impact and reflection: This incident not only successfully exposed a serious online rumor, but also triggered a deep reflection on the authenticity of information among the public. By introducing multimodal fine-grained CLIP analysis technology, news websites not only improved their content review capabilities, but also set an example for maintaining a healthy cyberspace and promoting the true dissemination of information. More importantly, this incident prompted all sectors of society to pay more attention to the harm of online rumors and strengthen the supervision and crackdown on online information.

[0036] This example once again proves the important role of multimodal fine-grained CLIP analysis technology in dealing with complex and ever-changing face forgery problems, and is a powerful tool for maintaining a healthy network ecosystem and improving the authenticity of information. At the same time, it also reminds us that in the information age, it is crucial to maintain critical thinking and vigilance about information to avoid being misled by false information.

[0037] Although the above describes the specific implementation mode of the invention in conjunction with the drawings, it is not intended to limit the scope of protection of the invention. Based on the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present invention.

Claims

1. A diffusion forged face detection method based on multimodal fine-grained CLIP, characterized in that: The following steps are involved: S1. Constructing a hierarchical and fine-grained face dataset , The dataset contains several face images, each of which has a corresponding hierarchical fine-grained label. The face images and the corresponding hierarchical fine-grained labels are preprocessed respectively to obtain the preprocessed face images and the corresponding hierarchical fine-grained label tensors. S2: Input the hierarchical fine-grained label tensor obtained after preprocessing into the fine-grained text generator to obtain fine-grained text , and the preprocessed face image tensor Forming face image text pairs ; S3, the preprocessed face image tensor Input into the multimodal visual encoder module to obtain multimodal visual features ; S4. Fine-grained text Input into the fine-grained language encoder to obtain fine-grained language features ; S5. Multimodal visual features and fine-grained language features They are input into the sample pair attention module to obtain the visual-language similarity matrix and the language-visual similarity matrix respectively; S6. Multimodal visual features of faces Input into the predictor and multi-layer perceptron respectively to obtain the predicted language features and predicted image category features ; S7. Utilization Divergence loss function Fine-grained language features and predictive language features Optimize using cross-modal contrast loss Function-to-sample similarity matrix Optimize and then use the manipulation detection loss function For predicting image category features Optimize with the real image labels and summarize to get the total loss function. Use the total loss function and Adam optimizer to optimize and train the parameters in each module to get the optimized and trained CLIP multimodal fine-grained model. S8. Input the image to be detected into the trained and optimized CLIP multimodal fine-grained model to obtain the final true or false discrimination result.

2. The method for detecting a diffuse forged face based on multimodal fine-grained CLIP according to claim 1, characterized in that: S1 is as follows: Hierarchical Fine-Grained Face Dataset The width × height of each face image in is uniformly adjusted to , and divide each face image pixel value by 255 to normalize the face image, and encapsulate the normalized face image as Tensor representation of , represents a vector space, represents any face image tensor in the preprocessed hierarchical fine-grained face dataset, and 3 represents the face image tensor The number of channels is 3; The hierarchical fine-grained labels corresponding to face images are as follows. The first layer is the coarse-grained layer, and the labels are recorded as or , represents the number of images in each batch, 0 represents a true image, and 1 represents a false image; the second layer is the manipulation layer, and the label is recorded as or or or , 0 represents the real image, 1 represents the identity exchange image, 2 represents the attribute manipulation image, and 3 represents the full face synthesis image; the third layer is the fine-grained layer, and the label is recorded as or or or or or or , 0 represents the real image, 1 represents the identity swap GAN image, 2 represents the identity swap diffusion image, 3 represents the attribute manipulation GAN image, 4 represents the attribute manipulation diffusion image, 5 represents the full face synthesis GAN image, and 6 represents the full face synthesis diffusion image; The fourth layer is the generator layer, and the label is recorded as or or or or or or or or or or , 0 represents the real image, 1 represents the image of the DiffFace face swapping framework, 2 represents the image of the Diffae diffusion model, 3 represents the image of the generated model in the LatentDiffusion deep learning, 4 represents the image of the CollaborativeDiffusion diffusion model framework, 5 represents the image of the DDPM denoising diffusion probability model, 6 represents the image of the FSLSD limited slip differential, 7 represents the FaceSwapper face swapping model image, 8 represents the image of the LatentTransformer image generation model, 9 represents the image of the StyleGAN2 model improved based on the generative adversarial network, and 10 represents the image of the StyleGAN3 model improved based on the generative adversarial network.

3. The method for detecting diffuse forged faces based on multimodal fine-grained CLIP according to claim 2, characterized in that: S2 is as follows: Design a fine-grained text generator module. The fine-grained text generator module is based on The hierarchical fine-grained labels provided by the dataset generate corresponding text hints for each face image , to construct image-text pairs; The fine-grained text generator module corresponds to the hierarchical fine-grained labels. The first layer of the fine-grained text generator module is the coarse-grained layer, the second layer is the manipulation layer, the third layer is the fine-grained layer, and the fourth layer is the generator layer. The fine-grained text generator module analyzes the label content layer by layer according to the input hierarchical fine-grained labels, and then generates the corresponding text prompts. , and then with the preprocessed face image tensor Together they form a face image text pair .

4. The method for detecting diffuse forged faces based on multimodal fine-grained CLIP according to claim 3, characterized in that: S3 is as follows: S3.

1. Construct a multimodal visual encoding module, which includes an image encoder, an image block selector, and a noise encoder, wherein the image encoder is composed of a convolutional visual attention model CviT, and the noise encoder is composed of a convolutional neural network backbone and L noise attention blocks; S3.

2. Pairing face image with text The face image tensor in Input to the multimodal visual encoding module, after the convolutional visual attention model CviT, the output dimension is Global manipulation of face images , , the calculation process is as follows: , in, Representing images The characteristic dimension of Represents the number of images in each batch, and the face image tensor of The same number, Represents the operation of the convolutional visual attention model CviT; Global manipulation of face images After the image block selector, the image Divided into non-overlapping image blocks, through the gray-level co-occurrence matrix calculate The homogeneity scores of the image patches are calculated, and the image patch with the lowest score is selected as the richest image patch. , , Indicates the number of images in each batch, Represents the richest image patch The number of channels is 3, Represents the richest image patch Width × height; The richest image patch Output noise of input to steganalysis enrichment model SRM , the calculation formula is as follows: , in, Represents the operations of the steganalysis enrichment model; Then the noise The input to the noise encoder module first passes through the convolutional neural network backbone, and the output dimension is Local noise map of face image , the calculation process is as follows: , in, represents the backbone operation of the convolutional neural network, Represents the parameters of the CNN backbone and the local noise map of the face image The dimension is , Indicates the number of channels of the local feature map of face noise, Indicates the height of the local feature map of face noise, Indicates the width of the local feature map of face noise; Then the local feature map of face noise Utilize along the channel Curry's The reshape function flattens the sequence into two-dimensional noise blocks. , , express The number of patches, Indicates A two-dimensional noise image patch; Then calculate the two-dimensional noise block sequence with position information , the calculation process is as follows: , in, represents the automatically generated learnable class tensor, represents the mapping latent vector, , Indicates 2D noisy image patch The mapping hidden vector of represents the position of the automatically generated two-dimensional noise image block sequence; Then the two-dimensional noise block sequence Input to L noise attention blocks, First, it passes through the first noise attention block, and then passes through the multi-head self-attention module and Module, and finally output the first two-dimensional noise global feature map , the calculation process is as follows: , , in, represents the operation of the normalization layer, represents the operation of the multi-head self-attention module, express Operation of the module; Then the output of the first noise attention block is As the input of the second noise attention block, the output of the second noise attention block is As the input of the third noise attention block, iterate multiple times until the Lth noise attention block outputs , represents the L-th two-dimensional noise global feature map; Finally, the L-th two-dimensional noise global feature map The learnable class tensor in is taken out to obtain the global noise feature ,Will and Fusion is performed to obtain the multimodal visual features of global image noise fusion , the calculation process is as follows: 。 5. The method for detecting diffuse forged faces based on multimodal fine-grained CLIP according to claim 4, characterized in that: S4 is as follows: S4.1 Fine-grained text First, the word tag sequence is obtained through the word segmenter, and then the word tag is mapped into a word embedding tensor through the word embedding layer , and then according to the word embedding tensor and the position of the automatically generated word embedding tensor Get fine-grained class-aware text sequence vector with position information , the calculation process is as follows: ; S4.

2. Construct a fine-grained language encoder. The fine-grained language encoder includes There are continuous attention modules, each of which includes a multi-head attention module and a feedforward neural network module. The previous layer of the multi-head attention module and the feedforward neural network module are Normalization layer, the next layer is the residual layer; S4.

3. Fine-grained class-aware text sequence vector with position information Input into a fine-grained language encoder, First, after the first continuous attention module, go through After the normalization operation, the normalization layer is input into the multi-head attention module of the first continuous attention module for global multi-head attention calculation, and then the global semantic features of the text are obtained through the residual layer. , the calculation process is as follows: , in, represents the normalization operation, Represents the operation of the multi-head attention module; Then Go through The normalization layer is normalized and then input into the feedforward neural network module, and then passes through the residual layer to obtain the refined global visual language fusion features. , the calculation process is as follows: , in, Represents the operation of a feedforward neural network module; Then the output of the first layer attention module of the language encoder is used as the input of the second attention module, and the output of the second layer attention module is used as the input of the third attention module, and iterates multiple times until the first The output of the attention module The final output of the fine-grained language encoder is the fine-grained language feature. Take the last word Get fine-grained language features .

6. The method for detecting diffuse forged faces based on multimodal fine-grained CLIP according to claim 5, characterized in that: S5 is as follows: Construct a sample pair attention module, which consists of a learnable sample pair attention parameter matrix to convert multimodal visual features and fine-grained language features They are input into the sample pair attention module respectively to obtain the visual language similarity matrix and the language visual similarity matrix. The calculation process is as follows: , , in, represents the visual language similarity matrix, represents the language-visual similarity matrix, represents element-wise dot product, represents the sigmoid activation function, represents the learnable weight matrix.

7. The method for detecting a diffuse forged face based on multimodal fine-grained CLIP according to claim 6, wherein S6 The details are as follows: Will Input to a predictor composed of fully connected layers to obtain predicted language features , the calculation process is as follows: , in, represents the predictor parameters; Then Input into a multi-layer perceptron composed of fully connected layers to obtain the predicted image category features , the calculation process is as follows: , in, Represents the parameters of the multilayer perceptron.

8. The method for detecting diffuse forged faces based on multimodal fine-grained CLIP according to claim 7, characterized in that: S7 is as follows: pass Divergence loss function Fine-grained language features and predictive language features Optimize, the specific calculation is as follows: , in, Indicates the number of face images involved in training, express The index of represents transpose, Indicates Fine-grained language features of an image, Indicates The predicted language features of the image, Represents the Softmax activation function; Through cross-modal contrast loss Function-to-sample similarity matrix Optimize the sample pair similarity matrix , the specific calculation is as follows: , , , , , in, represents the contrastive loss from vision to language, represents the contrastive loss from language to video, Indicates The sample pairs of input face images Label, Indicates A language-to-video similarity matrix, express Another index of represents the cosine similarity function, represents a trainable temperature parameter, Indicates The video-to-visual language similarity matrix, Indicates Multimodal visual features , Indicates Fine-grained language features , Indicates Multimodal visual features , Indicates Fine-grained language features ; Then by manipulating the detection loss function For predicting image category features With real image Coarse-grained labeling Optimize, the specific calculation is as follows: , in, Indicates Real images Coarse-grained labels, Indicates Predicted image category features; According to Divergence loss function , cross-modal contrast loss and manipulation detection loss function Get the total loss , the calculation process is as follows: 。 9. The method for detecting diffuse forged faces based on multimodal fine-grained CLIP according to claim 8, characterized in that: S8 is as follows: The face image to be detected is input into the optimized and trained CLIP multimodal fine-grained model to obtain the global visual features of the image. , and then The final true and false discrimination result is obtained by multi-layer perceptron calculation , the calculation process is as follows: , in, represents the multimodal visual encoder, Represents a multilayer perceptron.

Citation Information

Cited By

  • Image forgery detection method and device based on CLIP model

    CN120997613A

  • Image forgery detection method and device based on CLIP model

    CN120997613B

  • Face forgery detection method, system and device and medium

    CN122200821A