Method for detecting and positioning multi-mode media image-text synchronous forgery
By using forgery perception comparison learning in the graphic and text encoding part and semantic interaction and feature fusion between the graphic and text and audio-visual modality, the problem of poor detection and positioning performance of graphic and text synchronization forgery in the existing technology is solved, and accurate detection and positioning of graphic and text synchronization forgery is achieved, which significantly improves detection accuracy and positioning capabilities.
Patent Information
- Application Number
- CN202510197027.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The existing multimodal depth forgery detection model has poor performance when detecting synchronous forgery of graphics and texts, and cannot perform accurate detection and positioning.
By using forgery perception comparison to learn the overall semantic embedding of the aligned image and text in the graphic encoding part, and perform fine-grained semantic interactions and deep-level feature fusion between the graphic and text modes, the graphic and text context information provided by the audio and video are used to enhance the image features and text features, thereby realizing accurate detection and positioning of graphic and text synchronous forgery.
The accuracy and positioning capability of multimodal media graphic and text synchronization forgery detection have been significantly improved, so that the precise detection and positioning performance of graphic and text synchronization forgery has been significantly improved.
Smart Images

Figure CN120125979A_ABST
Abstract
Description
Technical Field
[0001] The present invention specifically relates to a method for detecting and locating synchronous forgery of multi-modal media graphics and text. Background Art
[0002] In recent years, with the rapid development of generative artificial intelligence technology, it has become a reality to generate realistic images and texts. At the same time, this technological progress has also made malicious forgery of multi-modal media content more complex and concealed. Current multi-modal deep forgery detection models mainly rely on the semantic consistency between graphic and text modalities to identify forgery traces;
[0003] However, when detecting synchronous forgery of graphics and text, multi-modal deep forgery detection models cannot accurately detect and locate synchronous forgery of graphics and text, and the performance of detecting and locating synchronous forgery of graphics and text is poor.
[0004] Therefore, it is necessary to invent a method for detecting and locating synchronous forgery of multi-modal media graphics and text to solve the above problems. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for detecting and locating synchronous forgery of multi-modal media graphics and text. By using forgery-aware contrastive learning in the graphic and text encoding part, the overall semantic embeddings of images and texts are aligned, so as to better capture the semantic correlation and potential inconsistent information between images and texts. Through fine-grained semantic interaction and deep feature fusion between graphic and text and video and audio modalities, the graphic and text context information provided by video and audio is used to enhance image features and text features, so as to facilitate the deeper revelation of graphic and text forgery traces. At the same time, through the detection of synchronous forgery of graphics and text, accurate judgment of the authenticity of graphic and text pairs and effective identification of the types of synchronous forgery of graphics and text are realized. Through the localization of synchronous forgery of graphics and text, high-precision identification of image forgery regions and text forgery tokens is realized, so that the performance of accurate detection and localization of synchronous forgery of graphics and text is significantly improved to solve the above deficiencies in the technology.
[0006] To achieve the above purpose, the present invention provides the following technical solutions: A method for detecting and locating synchronous forgery of multi-modal media graphics and text, comprising the following steps:
[0007] Step 1, create a multi-modal encoding module for multi-modal encoding;
[0008] Step 2, create a multi-modal hierarchical fusion module for multi-modal fusion;
[0009] Step 3, perform detection and location of synchronous forgery of graphics and text;
[0010] Step 4, train the model for detecting and locating synchronous forgery of multi-modal media graphics and text.
[0011] For the aforementioned method for detecting and locating synchronous forgery of multimodal media graphics and text, in step 1, a multimodal encoding module is created for multimodal encoding, including the following steps:
[0012] 1.1. Encode the graphics and text;
[0013] 1.2. Encode the video and audio.
[0014] For the aforementioned method for detecting and locating synchronous forgery of multimodal media graphics and text, in step 1.1, the graphics and text are encoded, and the specific steps are as follows:
[0015] 1.1.1. For any graphics and text (I, T), use the image encoder ViT to encode the image I into E i (I) = {I cls , I tok ;
[0016] where I cls is the global semantic embedding of I;
[0017] I tok = {I tok1 ,..., I tokN} is the N image patch embeddings of I;
[0018] 1.1.2. Use the text encoder BERT to encode the text T into E t (T) = {T cls , T tok};
[0019] where T cls is the global semantic embedding of T;
[0020] T tok = {T tok1 ,..., T tokM} is the token embeddings of M word tokens in T;
[0021] 1.1.3. Based on the InfoNCE loss, use forgery-aware contrastive learning to align the overall semantic embeddings of the image I and the text T. The specific formula is as follows:
[0022]
[0023] where, is the contrastive loss from image to text;
[0024] is the contrastive loss from text to image;
[0025] τ is the temperature hyperparameter;
[0026] is a negative sample set composed of K texts with inconsistent semantic embeddings with I;
[0027] is a negative sample set composed of K images with inconsistent semantic embeddings with T;
[0028] is the negative sample T k 's global semantic embedding;
[0029] is the negative sample I k 's global semantic embedding;
[0030] is the text momentum encoder 's global semantic encoding of T;
[0031] is the text momentum encoder 's global semantic encoding of T k ;
[0032] is the text momentum encoder 's global semantic encoding of I;
[0033] is the text momentum encoder 's global semantic encoding of I k ;
[0034] h i is the image projection head;
[0035] h t is the text projection head;
[0036] is the image momentum projection head, is the text-image momentum projection head, used to map the text-image semantics to a low-dimensional space for similarity calculation;
[0037] E p(I,T) is the mathematical expectation of the joint probability distribution of the image and text when the image is aligned with the text;
[0038] E p(T,I) is the mathematical expectation of the joint probability distribution of the text and image when the text is aligned with the image;
[0039] 1.1.4. Perform in-modal embedding alignment, then the total text-image alignment loss of the entire text-image forgery-aware contrastive learning Its calculation formula is as follows:
[0040]
[0041] Among them, is the image-to-image contrast loss;
[0042] is the text-to-text contrast loss.
[0043] For the aforementioned method for detecting and locating multi-modal media graphic and text synchronous forgery, in step 1.2, the audio and video are encoded, and the specific steps are as follows:
[0044] 1.2.1. For any audio-video pair (V, A), the video is divided into T segments, and t key frames are taken from each segment as the video-level input;
[0045] 1.2.2. Use the short-time Fourier transform to process the T * t key frames, and use the obtained Mel spectrogram as the audio-level input;
[0046] 1.2.3. Use the transformer to process the audio-video level input to form an initial audio-video embedding sequence, specifically:
[0047] V 0 =[V cls , V 1|1 , V 2|1 ,..., V t|T and A 0 =[A cls , A 1 , A 2 ,..., A T ;
[0048] Among them, V cls is the global video semantics;
[0049] A cls is the global audio semantics;
[0050] V i|j is the i-th key frame in the j-th video segment;
[0051] A i is the audio segment corresponding to the i-th video segment;
[0052] 1.2.4. Use average pooling to update the initial video embedding sequence V 0 , and the specific formula is as follows:
[0053]
[0054] Among them, is element-wise addition;
[0055] 1.2.5. Use the time encoder E of the spatio-temporal encoder TSE tm to obtain from V0 and A 0 Extract the temporal features of the video and audio respectively, and the specific formulas are as follows:
[0056]
[0057] where m ∈ {V, A} is the modality marker for video (V) and audio (A);
[0058] is E tm The temporal feature of modality m extracted at the l-th layer, and 1 ≤ l ≤ L;
[0059] and is E tm The input of the first layer;
[0060] MSA and LN are the multi-head attention and normalization layers respectively;
[0061] SI (1 ≤ SI ≤ T) is the video-audio temporal segment number;
[0062] 1.2.6. Use the spatial encoder E in TSE spa Extract the spatio-temporal features of the video and audio from the temporal features and respectively, and the specific formulas are as follows:
[0063]
[0064] where is the embedding of n patches of the key frame or Mel spectrogram;
[0065] POS is a learnable vector used to mark the positions of n patches in, with an initial value of zero;
[0066] is a broadcast operation;
[0067] is E spa The spatio-temporal feature of modality m extracted at the l-th layer, and 1 ≤ l ≤ L - 1;
[0068] is E spa The input of the first layer;
[0069] FF is the feed-forward layer.
[0070] For the aforementioned method for detecting and locating multimodal media graphic-text synchronous forgery, in step 2, create a multimodal hierarchical fusion module for multimodal fusion, including the following steps:
[0071] 2.1. The bidirectional cross-attention mechanism is adopted to perform pairwise fusion on images and texts, conduct fine-grained semantic interaction between images and texts, and make full use of the complementary information between images and texts. The specific formula is as follows:
[0072]
[0073]
[0074] W Q 、W K and W V are learnable parameters shared by images and texts, used to align the entity or sentiment attribute embeddings in images and texts;
[0075] E i (I)W Q is the attention scoring matrix of I;
[0076] E t (T)W Q is the attention scoring matrix of T;
[0077] BiCroAtt() is the bidirectional cross-attention function;
[0078] U I ={I cls ,I tok} is the embedding of I;
[0079] U T ={T cls ,T tok} is the embedding of T;
[0080] I tok ={I 1 ,...,I N} are N image patches in I;
[0081] T tok ={T 1 ,...,T M} are the embeddings of M tokens in T;
[0082] 2.2. The cross-attention mechanism is adopted to perform pairwise fusion on videos and audios, and utilize the complementary features between videos and audios. The specific formula is as follows:
[0083]
[0084] Among them, P V and P A are the spatio-temporal encodings of videos and audios respectively;
[0085] D is the encoding dimension;
[0086] CroAtt() is the cross-attention function;
[0087] U V ={V cls ,V tok} is the fused information of the video-audio pair;
[0088] V cls is the global semantic embedding of video-audio;
[0089] V tok is the local spatio-temporal feature embedding;
[0090] 2.3. Use cross-modal contrastive learning to compare the similarity of the embeddings of video-audio pairs and image-text pairs, and calculate the contrastive loss The specific formula is as follows:
[0091]
[0092] where y I =y V and y I ≠y V are the positive sample (true image-text pair) and the negative sample (synchronously forged image-text pair) respectively;
[0093] N is the sample size of each batch;
[0094] sim() is the similarity function;
[0095] α is the margin value, which is used to adjust the similarity of the negative sample to prevent the model from relying too much on the easily separable negative sample;
[0096] 2.4. Cross-fuse the image-text-video-audio.
[0097] For the method of detecting and locating the synchronous forgery of multi-modal media image-text mentioned above, in step 2.4, cross-fuse the image-text-video-audio, and the specific steps are as follows:
[0098] 2.4.1. Based on cross-attention, incorporate video-audio information into the image and text respectively, so that both the image and text contain context information. The specific formula is as follows:
[0099]
[0100] where LN() is the layer normalization operation;
[0101] is the image embedding after the image incorporates video-audio information;
[0102] is the text embedding after the text incorporates video-audio information;
[0103] 2.4.2. Based on the cross-fusion information of images and audio-visuals, as well as the cross-fusion information of text and audio-visuals, use self-attention and residual networks to obtain image enhancement features and text enhancement features. The specific formulas are as follows:
[0104]
[0105] where SAT() represents the self-attention mechanism;
[0106] is the enhanced image embedding;
[0107] is the enhanced text embedding;
[0108] and are the overall semantic embeddings of the enhanced image and text respectively;
[0109] and are the enhanced image patch and text token embeddings respectively.
[0110] For the aforementioned method for detecting and locating synchronous forged graphics and texts in multimodal media, in step 3, perform synchronous forged graphics and texts detection and forged location, including the following steps:
[0111] 3.1. Synchronous forged graphics and texts detection;
[0112] 3.2. Synchronous forged graphics and texts location.
[0113] For the aforementioned method for detecting and locating synchronous forged graphics and texts in multimodal media, in step 3.1, the synchronous forged graphics and texts detection includes graphic and text authenticity identification and synchronous forged graphics and texts type identification. And the specific steps of the synchronous forged graphics and texts detection are as follows:
[0114] 3.1.1. Establish an understanding of the global semantics of graphic and text pairs. The specific formula is as follows:
[0115]
[0116] where and are the enhanced global semantic embeddings of images and texts formed by multi-modal hierarchical fusion respectively;
[0117] concat() is the concatenation operation;
[0118] FC is the fully connected layer;
[0119] cls I,T is the global semantic embedding of graphic and text pairs;
[0120] 3.1.2. Implement text and image authenticity recognition and text and image synchronous forgery type recognition using a binary classifier and a multi-label classifier. The specific formulas are as follows:
[0121]
[0122] Among them, C bin () and C mul () are the binary classifier and the multi-label classifier respectively. Both are constructed based on the multi-layer perceptron (MLP), and the Gaussian error linear unit (GELU) is used as the activation function at the same time;
[0123] y bin and y mul are the text and image authenticity label and the synchronous forgery type label respectively;
[0124] H() is the cross entropy;
[0125] E (I,T)~P is the mathematical expectation of the entropy distribution;
[0126] and are the optimization objectives of the binary classifier and the multi-label classifier respectively.
[0127] For the aforementioned method for detecting and locating text and image synchronous forgery in multi-modal media, in step 3.2, text and image synchronous forgery location includes image forgery area location and text forgery token location. The specific steps for text and image synchronous forgery location are as follows:
[0128] 3.2.1. Select a BBox detector for image forgery area location, and combine norm and generalized
[0129] intersection over union to calculate the location loss The specific formula is as follows:
[0130]
[0131] Among them, is the image patch embedding;
[0132] D i () is the BBox detector that uses MLP and GELU as the architecture and activation function respectively;
[0133] Sigmoid is the activation function;
[0134] y box is the normalized image forgery area location label;
[0135] || || is norm;
[0136] is the generalized intersection - union ratio;
[0137] E IF is the mathematical expectation of the localization deviation distribution of the image forgery area;
[0138] 3.2.2. Select a token anomaly detector to locate text forgery tokens and calculate the localization loss in combination with cross - entropy
[0139] and KL divergence The specific formula is as follows:
[0140]
[0141] where is the text token embedding;
[0142] D t () is a token anomaly detector built based on BERT, used to predict the anomaly probability of each token
[0143] in;
[0144] y tok is the text forgery token localization label;
[0145] is the momentum version of
[0146] is the momentum version of the token anomaly detector;
[0147] H() is the cross - entropy, which is used here to measure the difference between the detection result of D t () and the true label;
[0148] KL[||] is the Kullback - Leibler divergence, used to evaluate the probability distribution difference between the detection results of D t () and ;
[0149] probability distribution difference;
[0150] α ∈ (0,1) is a hyperparameter used to balance the localization accuracy and stability of the model;
[0151] E TF is the mathematical expectation of the localization deviation distribution of text anomaly tokens.
[0152] For the aforementioned method of detecting and locating multi - modal media graphic - text synchronous forgery, in step 4, the model for detecting and locating multi - modal media graphic - text synchronous forgery is trained. The specific formula is as follows:
[0153]
[0154] Among them, is the global optimization objective for the training of the entire system.
[0155] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0156] In the present invention, in the graphic and text encoding part, the overall semantic embedding of the image and text is aligned by using forged perception contrast learning, so as to better capture the semantic correlation and potential inconsistent information between the image and text. When performing multimodal fusion, through fine-grained semantic interaction and deep feature fusion between the graphic and text and the video and audio modalities, the graphic and text context information provided by the video and audio is used to enhance the image features and text features, which is convenient for revealing the graphic and text forgery traces more deeply. Through graphic and text synchronous forgery detection, accurate judgment of the authenticity of the graphic and text pair and effective identification of the graphic and text synchronous forgery type are realized. Through graphic and text forgery localization, high-precision identification of the image forgery area and text forgery tokens is realized, significantly improving the accuracy of graphic and text synchronous forgery detection in multimodal media and enhancing its interpretability. Thus, high-precision identification of the forgery area and forgery details is realized, significantly improving the localization ability of the forgery content, and significantly enhancing the accurate detection and localization performance of graphic and text synchronous forgery. Brief Description of the Drawings
[0157] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0158] Figure 1 is the flow chart of the present invention;
[0159] Figure 2 is the statistical information diagram of the DLSF dataset of the present invention;
[0160] Figure 3 is the schematic diagram of Mel spectrum segmentation of the present invention;
[0161] Figure 4 is the performance comparison diagram of different models in the multi-classification task of identifying graphic and text synchronous forgery types;
[0162] Figure 5 is the visualization demonstration diagram of the graphic and text synchronous forgery detection and localization results of the present invention;
[0163] Figure 6 is the attention comparison diagram of different models for graphic and text synchronous forgery. Detailed Embodiments
[0164] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0165] The present invention provides a method for detecting and locating synchronous forgery of multimodal media graphics and text as shown in Figure 1-6 , comprising the following steps:
[0166] Step 1: Create a multimodal encoding module for multimodal encoding, including the following steps;
[0167] 1.1 Encode the graphics and text. The specific steps are as follows:
[0168] 1.1.1 For any graphics and text (I, T), use the image encoder ViT to encode the image I into E i (I) = {I cls , I to ;
[0169] where I cls is the global semantic embedding of I;
[0170] I tok = {I tok1 ,..., I tokN} is the N image patch embeddings of I;
[0171] 1.1.2 Use the text encoder BERT to encode the text T into E t (T) = {T cls , T tok};
[0172] where T cls is the global semantic embedding of T;
[0173] T tok = {T tok1 ,..., T tokM} is the token embeddings of M word tokens in T;
[0174] 1.1.3 Based on the InfoNCE loss, use forgery-aware contrastive learning to align the overall semantic embeddings of the image I and the text T. The specific formula is as follows:
[0175]
[0176] where is the contrastive loss from image to text;
[0177] is the contrastive loss from text to image;
[0178] τ is the temperature hyperparameter;
[0179] A negative sample set consisting of K texts with semantic embeddings inconsistent with I;
[0180] A negative sample set consisting of K images with semantic embeddings inconsistent with T;
[0181] Is the negative sample T k 's global semantic embedding;
[0182] Is the negative sample I k 's global semantic embedding;
[0183] Is the text momentum encoder 's global semantic encoding of T;
[0184] Is the text momentum encoder 's global semantic encoding of T k ;
[0185] Is the text momentum encoder 's global semantic encoding of I;
[0186] Is the text momentum encoder 's global semantic encoding of I k ;
[0187] h i Is the image projection head;
[0188] h t Is the text projection head;
[0189] Is the image momentum projection head, Is the text-image momentum projection head, used to map text-image semantics to a low-dimensional space for similarity calculation;
[0190] E p(I,T) Is the mathematical expectation of the joint probability distribution of the image and text when the image is aligned with the text;
[0191] E p(T,I) Is the mathematical expectation of the joint probability distribution of the text and image when the text is aligned with the image;
[0192] 1.1.4. Perform in-modal embedding alignment, then the total text-image alignment loss of the entire text-image forgery-aware contrastive learning Its calculation formula is as follows:
[0193]
[0194] Among them, is the image-to-image contrast loss;
[0195] is the text-to-text contrast loss;
[0196] This step 1.1.4 can make the semantics within the modality consistent;
[0197] 1.2. Encode the video and audio, and the specific steps are as follows:
[0198] 1.2.1. For any video-audio pair (V, A), divide the video into T segments, and take t key frames from each segment as the video-level input;
[0199] 1.2.2. Process the T * t key frames using the short-time Fourier transform, and take the obtained Mel spectrogram as the audio-level input;
[0200] 1.2.3. Use the transformer to process the video-audio-level input to form an initial video-audio embedding sequence, specifically:
[0201] V 0 = [V cls , V 1|1 , V 2|1 ,..., V t|T and A0 = [A cls , A 1 , A 2 ,..., A T ;
[0202] Among them, V cls is the global video semantics;
[0203] A cls is the global audio semantics;
[0204] V i|j is the i-th key frame in the j-th video segment;
[0205] A i is the audio segment corresponding to the i-th video segment;
[0206] 1.2.4. Use average pooling to update the initial video embedding sequence V 0 , and the specific formula is as follows:
[0207]
[0208] Among them, is element-wise addition;
[0209] This step 1.2.4 can ensure the temporal synchronization of the video-audio modality;
[0210] 1.2.5. Use the time encoder E of the spatio-temporal encoder TSE tm Extract the temporal features of the video and audio from V 0 and A 0 respectively, and the specific formula is as follows:
[0211]
[0212] where m ∈ {V, A} is the modality marker for video (V) and audio (A);
[0213] is the temporal feature of modality m extracted by the l-th layer of E, and 1 ≤ l ≤ L; tm
[0214] and are the inputs of the first layer of E; tm MSA and LN are the multi-head attention and normalization layers respectively;
[0216] SI (1 ≤ SI ≤ T) is the video-audio temporal segment number;
[0217] 1.2.6. Use the spatial encoder E in TSE spa Extract the spatio-temporal features of the video and audio from the temporal features and respectively, and the specific formula is as follows:
[0218]
[0219] where is the embedding of n patches of the key frame or Mel spectrogram;
[0220] POS is a learnable vector used to mark the position of n patches in with an initial value of zero;
[0221] is the broadcast operation;
[0222] is the spatio-temporal feature of modality m extracted by the l-th layer of E, and 1 ≤ l ≤ L - 1; spa
[0223] is the input of the first layer of E; spa FF is the feed-forward layer;
[0224]
[0225] In this step 1, by setting up multi-modal encoding, a more comprehensive and accurate feature representation is provided for multi-modal forgery detection. Among them, in the image-text encoding part, forgery-aware contrastive learning is used to align the overall semantic embeddings of images and texts, so as to better capture the semantic correlations and potential inconsistent information between images and texts;
[0226] By aligning the within-modal embeddings, the within-modal semantic consistency is further ensured, and the accuracy of semantic expression is improved. In the video-audio encoding part, the short-time Fourier transform is combined to process audio, and the spatio-temporal features of video key frames are extracted by multi-head attention, effectively capturing the temporal synchronization and spatial features of video and audio modalities;
[0227] Overall, multi-modal encoding not only enhances the semantic correlations between image and text modalities, and between video and audio modalities, but also improves the robustness and generalization ability of forgery detection.
[0228] Step 2: Create a multi-modal hierarchical fusion module for multi-modal fusion, including the following steps;
[0229] 2.1. Adopt a bidirectional cross-attention mechanism to perform pairwise fusion of images and texts, conduct fine-grained semantic interaction between images and texts, and make full use of the complementary information of images and texts. The specific formula is as follows:
[0230]
[0231] W Q 、W K and W V are learnable parameters shared by images and texts, used to align the entity or sentiment attribute embeddings in images and texts;
[0232] E i (I)W Q is the attention scoring matrix of I;
[0233] E t (T)W Q is the attention scoring matrix of T;
[0234] BiCroAtt() is the bidirectional cross-attention function;
[0235] U I ={I cls ,I tok} is the embedding of I;
[0236] U T ={T cls ,T tok} is the embedding of T;
[0237] I tok ={I 1,...,I N} are N image patches in I;
[0238] T tok = {T 1 ,...,T M} are the embeddings of M tokens in T;
[0239] 2.2. Pairwise fusion of video and audio is performed using the cross-attention mechanism, leveraging the complementary features between video and audio. The specific formula is as follows:
[0240]
[0241] where P V and P A are the spatio-temporal encodings of video and audio respectively;
[0242] D is the encoding dimension;
[0243] CroAtt() is the cross-attention function;
[0244] U V = {V cls ,V tok} is the fused information of the video-audio pair;
[0245] V cls is the global semantic embedding of the video-audio;
[0246] V tok is the local spatio-temporal feature embedding;
[0247] 2.3. Cross-modal contrastive learning is used to compare the similarities of the embeddings of video-audio pairs and image-text pairs, and the contrastive loss L CON is calculated. The specific formula is as follows:
[0248]
[0249] where y I = y V and y I ≠ y V are the positive sample (true image-text pair) and negative sample (synchronized forged image-text pair) respectively;
[0250] N is the sample size per batch;
[0251] sim() is the similarity function;
[0252] α is the margin value, used to adjust the similarity of negative samples to prevent the model from relying too much on easily separable negative samples;
[0253] 2.4. Cross-fusion of image-text and video-audio is performed. The specific steps are as follows:
[0254] 2.4.1. Incorporate audio-visual information into images and texts respectively based on cross-attention, so that both images and texts contain context information. The specific formula is as follows:
[0255]
[0256] Among them, LN() is the layer normalization operation;
[0257] is the image embedding after the image incorporates audio-visual information;
[0258] is the text embedding after the text incorporates audio-visual information;
[0259] 2.4.2. Based on the cross-fusion information between images and audio-visuals, and the cross-fusion information between texts and audio-visuals, use self-attention and residual networks to obtain image enhancement features and text enhancement features. The specific formula is as follows:
[0260]
[0261] Among them, SAT() represents the self-attention mechanism;
[0262] is the enhanced image embedding;
[0263] is the enhanced text embedding;
[0264] and are the overall semantic embeddings of the enhanced image and text respectively;
[0265] and are the enhanced image patch and text token embeddings respectively;
[0266] In this step 2, through multi-modal fusion, fine-grained semantic interaction and deep feature fusion are carried out between the image-text and audio-visual modalities, so as to enhance the image features and text features by using the image-text context information provided by the audio-visuals, facilitate the deeper revelation of image-text forgery traces, and thus significantly improve the ability to capture multi-modal forgery traces;
[0267] Among them, through pairwise modal fusion, the text-image pair is fused using bidirectional cross-attention and the audio-visual pair is fused using cross-attention. This not only makes full use of the complementary information between the image-text modalities, but also provides key anchoring information for the authenticity detection of image-text through the logic and temporal sequence coherence of the audio-visual modality;
[0268] Through cross-modal contrastive learning, the embeddings of real and fake images and texts are aligned, effectively distinguishing real and fake samples;
[0269] Furthermore, through cross-fusion and feature enhancement, the cross-self-attention mechanism is used to perform deep interactive fusion of text, image, video and audio modalities, revealing the cross-modal inconsistency caused by the simultaneous forgery of text and image.
[0270] At the same time, the image and text features are enhanced based on self-attention and residual networks, which strengthens the model's sensitivity to potential forgery traces in images and texts.
[0271] Overall, multimodal fusion effectively improves the model's global semantic understanding ability of multimodal media content and the robustness of forgery detection.
[0272] Step 3: Performing image and text synchronization forgery detection and forgery location, including the following steps:
[0273] 3.1. Detection of synchronous forgery of images and texts, including identification of true and false images and identification of synchronous forgery types of images and texts. The specific steps of synchronous forgery detection of images and texts are as follows:
[0274] 3.1.1. Establish the understanding of global semantics of images and texts. The specific formula is as follows:
[0275]
[0276] in, and They are respectively the enhanced global semantic embeddings of images and texts formed by multimodal hierarchical fusion;
[0277] concat() is a concatenation operation;
[0278] FC is the fully connected layer;
[0279] cls I,T Global semantic embedding for image and text pairs;
[0280] 3.1.2. Use binary classifiers and multi-label classifiers to realize the recognition of true and false images and texts and the recognition of the type of simultaneous forgery of images and texts. The specific formula is as follows:
[0281]
[0282] Among them, C bin () and C mul () are binary classifiers and multi-label classifiers, respectively. Both are built based on multi-layer perceptron (MLP) and use Gaussian error linear unit (GELU) as the activation function;
[0283] y bin and mulThey are the true / false labels for images and texts and the synchronization forgery type labels respectively;
[0284] H() is the cross entropy;
[0285] E (I,T)~P is the mathematical expectation of the entropy distribution;
[0286] and are the optimization objectives of the binary classifier and the multi-label classifier respectively;
[0287] 3.2. Image and text synchronization forgery localization, including image forgery area localization and text forgery token localization, and the specific steps of image and text synchronization forgery localization are as follows:
[0288] 3.2.1. Select a BBox detector for image forgery area localization, and calculate the localization loss by combining the norm and the generalized intersection over union The specific formula is as follows:
[0289]
[0290] Among them, is the image patch embedding;
[0291] D i () is the BBox detector that uses MLP and GELU as the architecture and activation function respectively;
[0292] Sigmoid is the activation function;
[0293] y box is the normalized image forgery area localization label;
[0294] || || is the norm;
[0295] is the generalized intersection over union;
[0296] E IF is the mathematical expectation of the image forgery area localization deviation distribution;
[0297] 3.2.2. Select a token anomaly detector for text forgery token localization, and calculate the localization loss by combining the cross entropy and the KL divergence The specific formula is as follows:
[0298]
[0299] Among them, is the text token embedding;
[0300] D t() is a token anomaly detector built based on BERT, used to predict the anomaly probability of each token in
[0301] y tok is the text forgery token localization label;
[0302] is the momentum version of
[0303] is the momentum version of the token anomaly detector;
[0304] H() is the cross entropy, used here to measure the difference between the detection result of D t () and the true label;
[0305] KL[||] is the Kullback-Leibler divergence, used to evaluate the probability distribution difference between the detection result of D t () and ;
[0306] α∈(0,1) is a hyperparameter, used to balance the localization accuracy and stability of the model;
[0307] E TF is the mathematical expectation of the text anomaly token localization deviation distribution;
[0308] In this step 3, through the synchronized text and image forgery detection, the accurate judgment of the overall authenticity of the multi-modal media content and the effective identification of the forgery type are realized;
[0309] Among them, using a binary classifier to classify the authenticity of text and images effectively enhances the model's ability to detect synchronized text and image forgeries;
[0310] Identifying the forgery type through a multi-label classifier can accurately distinguish different forgery methods and provide more fine-grained detection results;
[0311] Based on the fusion representation of the global semantic embedding of text and images, this method significantly improves the comprehensiveness and robustness of synchronized text and image forgery detection;
[0312] Through the synchronized text and image forgery localization, the high-precision identification of the image forgery area and the text forgery token is realized, enhancing the interpretability of the synchronized text and image forgery detection result;
[0313] Among them, using image patch embedding and a BBox detector to localize the image forgery area, and optimizing the localization loss by combining the norm and the generalized intersection over union (IoU), improves the accuracy of the image forgery area localization;
[0314] Meanwhile, with the help of text token embeddings and a BERT-based token anomaly detector to locate forged tokens, the localization loss is optimized by combining cross-entropy and KL divergence, effectively capturing the forged tokens in the text;
[0315] Overall, this method can accurately identify the specific regions and tokens of synchronized text-image forgery, providing interpretability and practicality for multi-modal synchronized forgery detection.
[0316] Step 4: Train the multi-modal media text-image synchronized forgery detection and localization model, and the specific formula is as follows:
[0317]
[0318] Among them, is the global optimization objective for the training of the entire system;
[0319] In this Step 4, backpropagation training is performed based on supervised learning, and through it can ensure the collaborative optimization of text-image synchronized forgery detection and localization.
[0320] Verification experiment
[0321] The first step: Establish the DLSF dataset
[0322] The present invention uses the DLSF dataset to evaluate the performance of the multi-modal media text-image synchronized forgery detection and localization model, that is, the MHFRT model. The specific process of establishing the DLSF dataset is as follows:
[0323] I. Dataset overview
[0324] The present invention uses the text, video, and audio data captured from Toutiao to construct a multi-modal media resource pool, that is, O = {p o |p o =(I o , T o , V o , A o )}, through text-image synchronized forgery of the image I o and the text T o , as well as editing and encoding of the video V o and the audio A o , the first multi-modal text-image synchronized forgery dataset DLSF is released. The DLSF dataset is compared with multiple existing datasets in terms of modality, TISF samples, and localization labels. The specific results are shown in Table 1 below:
[0325] Table 1
[0326]
[0327]
[0328] As can be seen from Table 1: Compared with the existing datasets, the newly released DLSF dataset contains more modalities. It not only includes text and images, as well as videos and audio, but also provides synchronized text-image forgery samples and location tags, enabling it to be used for the performance evaluation of the MHFRT model.
[0329] II. Sample Generation
[0330] 1. Entity-level synchronized text-image forgery samples
[0331] Such samples are generated by synchronously replacing the entity face in image I o and the entity identity in the corresponding text T o . For example, replacing the face of A in a news image with the face of B, and at the same time replacing the name of A in the corresponding news text with the name of B. The specific process of generating such entity-level synchronized text-image forgery samples is as follows:
[0332] a1. Use a named entity recognition model to identify the entity name "PER1" from the original news title T o , and form an entity-level text forgery (TS) sample T s by replacing PER1 with PER2;
[0333] a2. For the original image I o corresponding to T o , randomly use two face replacement models, GHOST and SimSwap, to replace the face image of PER1 in I o with the face image of PER2 to form an entity-level image forgery (FS) sample I s ;
[0334] a3. Obtain the first M words of T s through text truncation or completion, and use an M-dimensional one-hot vector to mark the forged words in T s , and use the positioning box y box = {x 1 , y 1 , x 2 , y 2} to mark the image forgery area in I s ;
[0335] Among them, y i = 1 indicates that the i-th word in T s is forged, and y i = 0 indicates that the i-th word in T s is not forged.
[0336] 2. Attribute-level synchronized text-image forgery samples
[0337] By performing the same polarity deflection on the facial expressions in the image I o and the emotional descriptions in the corresponding text T o such samples are generated. For example, for the news text T o with a positive emotion, while forging the text emotion to be negative, the facial expressions in the corresponding image I o are synchronously edited to be negative. The specific process of generating such attribute-level text-image synchronous forgery samples is as follows:
[0338] b1. Use the Chinese-RoBERTa model to divide the news title T o into positive {o +}, negative {o -}, and neutral {o neu} according to the emotional polarity;
[0339] b2. Use the model GPT-4o to edit the text with {o +} emotion into text with {o -} emotion, and edit the text with {o -} and {o neu} emotions into text with {o +} emotion to form attribute-level text forgery (TA) samples T a , and use the vector y tok to mark the forged tokens in T a ;
[0340] b3. For the original image I o corresponding to T o , select the diffusion model Face-Adapter to edit the facial expressions with {o -} and {o neu} emotions into {o +} expressions;
[0341] b4. Select the GAN-based model HFGI to edit the facial expressions with {o +} emotion into {o -} expressions;
[0342] b5. By rendering the image editing area , form attribute-level image forgery (FA) samples I a , and use the vector y box to mark the image forgery area in I a .
[0343] 3. Audio-visual samples
[0344] Among them, the specific process of generating audio-visual samples is as follows:
[0345] All the captured videos V o are all clipped to a specified duration of 60 seconds and encoded into H.264 video streams using the open-source video encoding library Libx264;
[0346] All the audio A o is then extracted from the clipped videos and encoded according to the AAC audio encoding standard to generate the corresponding Mel spectrogram;
[0347] 4. Integration and perturbation
[0348] By integrating the TISF samples and the audio-visual samples, a multi-modal text-image synchronized forged media pool is formed. The specific formula is as follows:
[0349] P = {p m | p m = (I x , T x , A, V), x ∈ {o, s, a}}
[0350] where I x is the image sample;
[0351] T x is the text sample;
[0352] A is the audio sample;
[0353] V is the video sample;
[0354] x ∈ {o, s, a} is the forgery status of the modality, where o, s, and a represent unforged, entity-level forgery, and attribute-level forgery respectively;
[0355] And each text-image-audio-video pair p m provides 4 labels, namely the binary label y bin = {0, 1}, which is used to describe whether text-image synchronization forgery occurs in p m ;
[0356] The forgery type label is used to describe whether the j-th forgery type appears in p m , that is, entity-level image forgery FS, attribute-level image forgery FA, entity-level text forgery TS, and attribute-level text forgery TA;
[0357] y box is used to describe the image forgery area;
[0358] y tok is used to describe the text forgery token;
[0359] Meanwhile, to visually reflect the problem that the traces of text-image synchronization forgery may be masked by noise, 50% of the images in P were randomly perturbed, including operations such as JPEG compression and Gaussian blur.
[0360] III. Dataset Statistics
[0361] As Figure 2 shown, the statistical information of the DLSF dataset is presented;
[0362] From Figure 2 Figure a, it can be seen that the DLSF contains 2,200 text-image-audio-video pairs of samples, including 179 attribute-level TISF samples, i.e., (FA + TA), and 279 entity-level TISF samples, i.e., (FS + TS);
[0363] As Figure 2 shown in Figure b, the size distribution of the entity-level and attribute-level image forgery areas is presented. It can be seen that the size of the entity-level image forgery area is relatively concentrated, with as many as 81% of the samples having a forgery area of 17% - 22%. In contrast, the size of the attribute-level forgery area is relatively dispersed, with only 40% of the samples having a forgery area of 7% - 35%, and as many as 60% of the samples having a forgery area below 7%. This is because entity-level forgery focuses on face replacement, with a single forgery area, while attribute-level forgery focuses on expression forgery, involving complex and diverse forgery areas;
[0364] As Figure 2 shown in Figure c, the number distribution of entity-level and attribute-level text forgery tokens is presented. It can be seen that the number of entity-level text forgery tokens is relatively concentrated. Approximately 35% of the samples contain only two forgery tokens, and approximately 65% of the samples contain three forgery tokens. In contrast, the number of attribute-level text forgery tokens is relatively dispersed, with a certain proportion of samples containing 4 to 29 forgery tokens. This is because entity-level forgery focuses on identity replacement, with relatively single forgery tokens, while attribute-level forgery emphasizes emotional editing, requiring more tokens to intervene;
[0365] As Figure 2 shown in Figure d, it can be seen that all the text-images in the source pool O have emotional tendencies. Therefore, when constructing the DLSF dataset, 68 pairs of text-images with obvious positive emotions, 82 pairs with obvious neutral emotions, and 29 pairs with obvious negative emotions were selected for attribute-level TISF operations;
[0366] As Figure 2 shown in Figure e, by reverse-editing 82 pairs of neutral-emotion and 29 pairs of negative-emotion text-images into 111 pairs of positive-emotion text-images, and reverse-editing 68 pairs of positive-emotion text-images into negative-emotion text-images, a total of 179 attribute-level TISF samples were formed.
[0367] Step 2: Experimental Setup
[0368] I. Perform data preprocessing
[0369] Select the above DLSF dataset to evaluate the performance of the MHFRT model. At the same time, to meet the tensor dimension requirements of the TSE input in the video and audio coding section, three filters are used to segment the audio mel spectrogram in the DLSF dataset. Specifically:
[0370] Use a low - frequency filter to extract mel spectrogram data below 2000 Hz;
[0371] Use a band - pass filter to extract mel spectrogram data from 2000 Hz to 5000 Hz;
[0372] Use a high - frequency filter to extract spectrogram data above 5000 Hz, and stack the segmented spectrogram data into a three - dimensional array according to the three dimensions of frequency band, mel frequency, and time;
[0373] By analyzing the spectral characteristics of the audio signal in different frequency bands, fine - grained audio features are extracted. The specific results are shown in the appendix Figure 3 as follows.
[0374] II. Select a baseline model
[0375] Select the following 8 models commonly used in text, image, and text - image anomaly detection tasks as the evaluation baseline. Specifically:
[0376] BERT: Generate the embedding of the text to be detected through a BERT model and use it as hidden knowledge to input into the second BERT model. At the same time, input the text to be detected itself into the second BERT model. Finally, the second BERT model outputs the probability that the text is forged;
[0377] FAST - DetectGPT: The FAST - DetectGPT model reveals the differences in word - choice between large language models (LLMs) and humans in specific contexts through conditional probability curvature, and optimizes the zero - shot detector based on these differences to accurately distinguish between machine - generated content and human - created text;
[0378] ViT: The ViT model divides the image into fixed - size blocks and converts them into embedding vectors, adds position information encoding, captures the dependencies between blocks through the Transformer encoder, and uses the output of the classification token to predict the authenticity of the image;
[0379] MAT: The MAT model uses multiple spatial attention heads to focus on different local regions of the image, capture forgery traces, aggregate low-level texture and high-level semantic features through the attention map, enhance texture information to identify forgery details, and optimize parameters by combining region-independent loss and attention guidance strategy to improve the accuracy and robustness of image forgery detection;
[0380] TS: The TS model reveals unnatural changes in the image by capturing multi-scale high-frequency noise to identify forgery regions, uses residual-guided spatial attention to analyze subtle differences and anomalies in the image to focus on forgery traces; integrates information from different feature spaces through cross-modal attention to improve the accuracy of identifying forgery regions in the image;
[0381] CLIP: The CLIP model uses contrastive learning to detect image-text forgeries. During training, semantic consistency is ensured by increasing the similarity between matching images and texts, while reducing the similarity between non-matching images and texts to identify inconsistent image-text pairs. This method enables CLIP to effectively detect potential forgery traces;
[0382] CAFE: The CAFE model maps image and text features to a shared semantic space through cross-modal alignment. In this space, cross-modal ambiguity learning is used to evaluate the degree of ambiguity between images and texts. Based on these ambiguity degrees, the correlation between images and texts is captured through cross-modal fusion, and image-text forgeries are detected accordingly
[0383] HAMMER: The HAMMER model first obtains image-text encodings through forgery-aware contrastive learning and identifies forgery regions in the image through a BBox detector. Subsequently, cross-modal perceptual attention is used to fuse image-text features, and a binary classifier is used to judge the authenticity of image-text. A multi-label classifier is used to determine the forgery type, and a token anomaly detector is used to locate forged tokens in the text.
[0384] III. Parameter and Metric Settings
[0385] The parameter settings of the MHFRT model are as follows: The image encoder uses a 12-layer ViT-B / 16, the text encoder and the Token anomaly detector both select Bert-Chinese, and the output dimensions of the binary classifier, multi-label classifier, and BBox detector for image-text authenticity classification, image-text synchronous forgery type identification, and image forgery region detection are set to 2, 4, and 4 respectively;
[0386] The MHFRT model training uses the AdamW optimizer, the weight decay is set to 0.02, the initial learning rate is 1e-4, and it is gradually decayed to 1e-5 in the first 150 steps;
[0387] Meanwhile, to ensure the fairness of the experimental evaluation, the parameters of all baseline models were set according to the settings in the original literature;
[0388] In the experiment, the following 12 metrics were used to evaluate the model performance: ACC, AUC, and EER were used to evaluate the model's ability to identify the authenticity of images and texts;
[0389] MAP, CF1, and OF1 were used to evaluate the model's ability to identify the types of synchronized forgeries of images and texts;
[0390] IoUmean, IoU50, and IoU75 were used to evaluate the performance of the model in locating the forged regions of images;
[0391] Precision, Recall, and F1 were used to evaluate the performance of the model in locating the forged text tokens;
[0392] All models in the experiment were implemented based on the Pytorch deep learning framework and trained on an NVIDIA RTX3090 GPU. When training the model, the ratio of the training set to the test set was set to 7:3.
[0393] IV. Single-modal forgery of images / texts
[0394] For text forgery, BERT and FAST-DetectGPT were selected as baseline models to evaluate the performance of the MHFRT model in detecting and locating forged text tokens;
[0395] For image forgery, ViT, MAT, and TS were selected as baseline models to evaluate the performance of the MHFRT model in detecting and locating forged regions of images;
[0396] Since none of these five baseline models provided a forged location output function, for the fairness of the evaluation, the present invention added a token anomaly detector to BERT and FAST-DetectGPT to provide forged text token outputs, and added a BBox detector to ViT, MAT, and TS to provide forged region outputs of images;
[0397] The comparison results of the performance of the MHFRT model and the above five models in detecting and locating forged text tokens are shown in Table 2:
[0398] Table 2
[0399]
[0400] As can be seen from Table 2, the performance of the MHFRT model in identifying the authenticity of images / texts, locating the forged regions of images, and locating the forged text tokens is better than that of the five baseline models. This shows that by introducing audio and video as context information for images and texts, the ability of the model to detect and locate single-modal forgeries of images / texts can be enhanced.
[0401] V. Synchronized Image-Text Forgery
[0402] For synchronized image-text forgery, CLIP, CAFE, and HAMMER are selected as baseline models to evaluate the performance of the MHFRT model in detecting and locating synchronized image-text forgery;
[0403] Moreover, to ensure the fairness of the evaluation, the present invention makes necessary adjustments to CLIP and CAFE. Specifically: First, to address the issue that CLIP only provides the output of identifying the authenticity of images and texts, a cross-attention layer is added to it to fuse the processing results of images and texts. At the same time, a multi-label classifier, a BBox detector, and a token anomaly detector are added to provide the output of identifying the types of synchronized image-text forgery, locating the forged areas in the images, and locating the forged tokens in the texts;
[0404] Second, for the problem that CAFE only provides the output of identifying the authenticity of images and texts, a cross-attention layer is also added to it to fuse the image and text information. At the same time, a multi-label classifier and a BBox detector are added;
[0405] Meanwhile, since CAFE uses FastCNN instead of NLP's sequence tagging when processing texts and a token anomaly detector cannot be added to it, it does not participate in the comparison of the performance of locating forged text tokens;
[0406] The performance comparison results of the MHFRT model and the above three models in detecting and locating synchronized image-text forgery are shown in Table 3;
[0407] Table 3
[0408]
[0409]
[0410] As can be seen from Table 3, the MHFRT model outperforms the other three baseline models in the four tasks of identifying the authenticity of images and texts, identifying the types of synchronized image-text forgery, locating the forged areas in the images, and locating the forged tokens in the texts. This indicates that by introducing audio and video as context information for images and texts, the ability of the model to detect and locate synchronized image-text forgery can be enhanced;
[0411] At the same time, to avoid the deviation of a single metric and the problem of incomplete evaluation caused by the imbalance between positive and negative samples in the DLSF dataset, based on the F1 metric, HAMMER is used as the baseline to evaluate the performance of the MHFRT model in the multi-classification task of identifying the types of synchronized image-text forgery. The specific results are as Figure 4 ;
[0412] From Figure 4It can be seen that the classification performance of the MHFRT model is significantly better than that of HAMMER, which indicates that by introducing visual and audio as context information, the recognition ability of the text-image synchronization forgery type can be improved. At the same time, the recognition effect of the two models on FS+TS is better than that on FA+TA. The reason is that attribute synchronization forgery occurs in a smaller image area ( Figure 2 Figure b in Figure 2 and scattered text tokens (
[0413] VI. Influence of Different Visual and Audio Introduction Strategies
[0414] To evaluate the influence degree of different visual and audio introduction strategies on the text-image synchronization forgery detection and localization ability of the MHFRT model, the present invention constructs three comparison models, specifically:
[0415] Model M1: Remove the visual and audio input function in MHFRT;
[0416] Model M2: Remove the audio input function in MHFRT;
[0417] Model M3: Remove the video input function in MHFRT;
[0418] The performance results of the above four models in the four tasks of text-image authenticity recognition, text-image synchronization forgery type recognition, image forgery area localization, and text forgery token localization are specifically as follows in Table 4:
[0419] Table 4
[0420]
[0421]
[0422] It can be obtained from Table 4 that the performance of M1 is the lowest in all indicators, indicating that introducing visual and audio as text-image context information can significantly improve the text-image synchronization forgery detection and localization ability of the model; while the performance of M2 in image forgery area localization is better than that of M3, and M3 performs better than M2 in text forgery token localization, indicating that the video modality provides rich inter-image-block correlation knowledge for images and enhances the detection ability of abnormal image blocks, while the audio modality provides more sequential knowledge between tokens for text, which is conducive to identifying abnormal tokens. At the same time, the performance of MHFRT in 11 indicators is better than that of M2 and M3, indicating that introducing visual and audio simultaneously can provide more comprehensive text-image context information, thereby reducing the occurrence of false positives and false negatives.
[0423] VII. Influence of Different Training Optimization Objectives
[0424] To test the influence of different training optimization objectives on the performance of the MHFRT model, based on Six comparative optimization objectives were constructed, specifically:
[0425] Only retain
[0426] Remove from Remove
[0427] Remove from Remove
[0428] Remove from Remove
[0429] Remove from Remove
[0430] Remove from Remove
[0431] The performance of MHFRT in identifying the authenticity of images and texts, identifying the types of synchronized image-text forgeries, locating the forged areas of images, and locating the forged word elements of texts under different optimization objectives is shown in Table 5;
[0432] Table 5
[0433] Table 5
[0434]
[0435] It can be seen from Table 5 that when using or as the training optimization objective, the performance of MHFRT is lower than that when using This shows that in the process of multi-modal fusion, if image-text alignment and the alignment of images and texts with audio and video are not carried out, the model cannot fully utilize the semantic correlation between images and texts and the information of the semantic temporal consistency of the context, which is not conducive to discovering the traces of synchronized image-text forgeries;
[0436] Compared with using as the training optimization objective, when using that is, not paying attention to the performance of the task of identifying the types of synchronized image-text forgeries, the performance of MHFRT in the other three tasks will deteriorate. Similarly, when using that is, not paying attention to the performance of the task of locating the forged areas of images, the performance of MHFRT in the other three tasks will also be affected. When using When not focusing on the performance of the text forgery token localization task, the performance of MHFRT on the other three tasks will also deteriorate. In particular, when only focusing on the performance of the image-text authenticity recognition task while ignoring the other three tasks, not only does the performance of MHFRT on the other three tasks decline significantly, but the performance of the image-text authenticity recognition is also inhibited. The above analysis shows that using as the training optimization objective can comprehensively improve the performance of MHFRT on various tasks of image-text synchronous forgery detection and localization.
[0437] VIII. Visualization of Forgery Detection and Localization
[0438] As Figure 5 shown, it demonstrates the detection and localization of image-text synchronous forgery by the MHFRT model. Figure 5 In it, Figures a and b are the original image-text before entity-level and attribute-level image-text synchronous forgery respectively. Figures c and d are the image-text after entity-level and attribute-level image-text synchronous forgery and the detection and localization results given by the MHFRT model. From Figure c, it can be seen that the MHFRT model correctly detected that the forgery type is FS+TS, and accurately located the image forgery area (replacing Tian Yiqian's face with Deng Ziqi's face) and the text forgery token (changing "Tian Yiqian" to "Deng Ziqi"). And from Figure d, it can be seen that the MHFRT model correctly detected that the forgery type is FA+TA, and accurately located the image forgery area (the expression changes from sad to happy) and the text forgery token (the description of international relations changes from provocative to cooperative).
[0439] IX. Visualization of Forgery Attention
[0440] As Figure 6 shown, it demonstrates the attention of different models to image-text synchronous forgery. Among them, Figure 6 Figures a and e in it are the original image-text before entity-level and attribute-level image-text synchronous forgery respectively. Figures b and f are the image-text after entity-level and attribute-level image-text synchronous forgery respectively. Figures c and g are the attention of HAMMER to Figures b and f respectively. Figures d and h are the attention of MHFRT to Figures b and f respectively;
[0441] Among them, in Figure b, the face and name of the second person from the left are synchronously forged. In Figure f, the facial expression and text description of the person are synchronously forged;
[0442] Regarding entity-level image-text synchronous forgery, by comparing the attention to the first face from the left in Figures c and d, it can be seen that for the non-forged image area, the attention given by MHFRT is significantly lower than that of HAMMER;
[0443] Moreover, judging from the second face and text from the left in Figures c and d, for the entity-level image-text synchronization forgery area, HAMMER assigns less attention (the small red area on the face), while MHFRT assigns more attention (the large purple area on the face), where purple represents abnormality (this is because MHFRT treats images not appearing in the video as abnormal images);
[0444] For the attribute-level image-text synchronization forgery, by comparing Figures g and h, it can be seen that compared with HAMMER, MHFRT assigns more concentrated attention to the forgery area (the red area on the lips). This indicates that by introducing audio-visual as the image-text context information, MHFRT significantly improves the detection and localization capabilities of image-text synchronization forgery.
[0445] Experimental Summary
[0446] In view of the problem of detecting and localizing image-text synchronization forgery in multi-modal media, the present invention establishes a multi-modal hierarchical fusion reasoning MHFRT model and releases the first dataset DLSF specifically for evaluating the performance of such forgery detection and localization models;
[0447] Experiments on the DLSF dataset show that by introducing audio-visual as the image-text context information, MHFRT achieves high-precision detection and accurate localization of image-text synchronization forgery.
[0448] In summary
[0449] In the present invention, in the image-text encoding part, forgery-aware contrastive learning is used to align the overall semantic embeddings of images and texts, so as to better capture the semantic associations and potential inconsistent information between images and texts. Through fine-grained semantic interaction and deep feature fusion between the image-text and audio-visual modalities, the image features and text features are enhanced by using the image-text context information provided by the audio-visual, so as to facilitate a deeper revelation of image-text forgery traces. Through image-text synchronization forgery detection, accurate judgment of the authenticity of image-text pairs and effective identification of image-text synchronization forgery types are achieved. Through image-text forgery localization, high-precision identification of image forgery areas and text forgery tokens is realized, significantly improving the accuracy of multi-modal media image-text synchronization forgery detection and enhancing its interpretability, thereby achieving high-precision identification of forgery areas and forgery details, significantly improving the localization ability of forgery content, and significantly enhancing the accurate detection and localization performance of image-text synchronization forgery.
[0450] Only certain exemplary embodiments of the present invention have been described above by way of illustration. Without doubt, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of the claims of the present invention.
Claims
1. A method for detecting and locating forgery of multimodal media image and text synchronization, characterized by: The following steps are involved: Step 1: Create a multimodal encoding module and perform multimodal encoding; Step 2: Create a multimodal hierarchical fusion module to perform multimodal fusion; Step 3: Perform image and text synchronization forgery detection and forgery location; Step 4: Train the detection and location multimodal media image and text synchronization forgery model.
2. A method for detecting and locating multimodal media image and text synchronization forgery according to claim 1, characterized in that: In step 1, a multimodal encoding module is created to perform multimodal encoding, including the following steps: 1.
1. Encode the image and text; 1.
2. Encode the video and audio.
3. A method for detecting and locating multimodal media image and text synchronization forgery according to claim 2, characterized in that: In step 1.1, encode the image and text. The specific steps are as follows: 1.1.
1. For any image (I, T), use the image encoder ViT to encode the image I into E i (I) = {I cls ,I tok ; in, I cls is the global semantic embedding of I; I tok = {I tok1 ,...,I tokN } is the embedding of N image blocks of I; 1.1.
2. Use the text encoder BERT to encode the text T into E t (T) = {T cls ,T tok }; Among them, T cls is the global semantic embedding of T; T tok ={T tok1 ,...,T tokM } is the token embedding of M words in T; 1.1.
3. Based on InfoNCE loss, forgery-aware contrastive learning is used to align the overall semantic embedding of image I and text T. The specific formula is as follows: in, is the contrast loss from image to text; is the contrast loss from text to image; τ is the temperature hyperparameter; A negative sample set consisting of K texts that are inconsistent with the semantic embedding of I; A negative sample set consisting of K images that are inconsistent with the semantic embedding of T; is a negative sample T k Global semantic embedding of is a negative sample I k Global semantic embedding of Momentum encoder for text Global semantic encoding of T; Momentum encoder for text To T k The global semantic encoding of Momentum encoder for text Global semantic encoding of I; Momentum encoder for text to I k The global semantic encoding of h i is an image projection head; h t For text projection header; is the image momentum projection head, It is the image-text momentum projection head, which is used to map the image-text semantics into a low-dimensional space for similarity calculation; E p(I,T) It is the mathematical expectation of the joint probability distribution of image and text when the image is aligned to the text; E p(T,I) It is the mathematical expectation of the joint probability distribution of text and image when aligning text to image; 1.1.
4. Perform intra-modal embedding alignment, then the total image-text alignment loss of the entire image-text forgery-aware contrast learning is The calculation formula is as follows: in, is the image-to-image contrast loss; is the text-to-text contrastive loss.
4. A method for detecting and locating multimodal media image and text synchronization forgery according to claim 3, characterized in that: in step 1.2, the video and audio are encoded, and the specific steps are as follows: 1.2.
1. For any video-audio pair (V, A), divide the video into T segments, and take t key frames in each segment as video-level input; 1.2.
2. Use short-time Fourier transform to process T*t key frames, and use the obtained Mel spectrum as audio level input; 1.2.
3. Use transformer to process the audio-visual input to form the initial audio-visual embedding sequence, specifically: V 0 = [V cls , V 1|1 , V 2|1 ,..., V t|T and A 0 = [A cls , A1, A2,..., A T ; Among them, V cls It is the global semantics of the video; A cls It is the global semantics of audio; V i|j is the i-th key frame in the j-th video; A i is the audio clip corresponding to the i-th video; 1.2.
4. Update the initial video embedding sequence V using average pooling 0 , the specific formula is as follows: in, is element-by-element addition; 1.2.5 Temporal Encoder E Using Spatiotemporal Encoder TSE tm From V 0 and A 0 The time features of video and audio are extracted respectively, and the specific formula is as follows: Among them, m∈{V,A} is the video (V) and audio (A) modality label; For E tm The temporal features of the mode m extracted at the lth layer, and 1≤l≤L; and For E tm Input to layer 1; MSA and LN are multi-head attention and normalization layers respectively; SI (1≤SI≤T) is the video and audio timing segment number; 1.2.
6. Using the spatial encoder E in TSE spa From the time characteristics and Extract the spatiotemporal features of video and audio from the , the specific formula is as follows: in, Embedding of n patches of keyframes or Mel-spectrograms; POS is a learnable vector used for labeling The position of the n patches in the middle, the initial value is zero; It is a broadcast operation; For E spa The spatiotemporal features of the mode m extracted at the lth layer, and 1≤l≤L-1; For E spa Input to layer 1; FF is the feed-forward layer.
5. The method for detecting and locating multimodal media image and text synchronization forgery according to claim 1, characterized in that: In step 2, a multimodal hierarchical fusion module is created to perform multimodal fusion, including the following steps: 2.
1. Use a bidirectional cross-attention mechanism to fuse images and texts in pairs, perform fine-grained semantic interaction between images and texts, and make full use of the complementary information between images and texts. The specific formula is as follows: W Q , W K and W V It is a learnable parameter shared by images and texts, used to align entity or sentiment attribute embeddings in images and texts; E i (I)W Q is the attention score matrix of I; E t (T)W Q is the attention score matrix of T; BiCroAtt() is a bidirectional cross attention function; U I = {I cls ,I tok } is the embedding of I; U T ={T cls , T tok } is the embedding of T; I tok ={I1,...,I N } is the N image blocks in I; T tok ={T1,...,T M } is the embedding of M tokens in T; 2.
2. Use the cross-attention mechanism to fuse the video and audio in pairs, and use the complementary features between the video and audio. The specific formula is as follows: Among them, P V and P A They are respectively the spatiotemporal coding of video and audio; D is the encoding dimension; CroAtt() is the cross attention function; U V = {V cls ,V tok } is the fusion information of video and audio pairs; V cls Global semantic embedding for audio and video; V tok Embedding for local spatiotemporal features; 2.
3. Use cross-modal contrastive learning to compare the embedding similarity of audio-visual pairs and image-text pairs, and calculate the contrastive loss The specific formula is as follows: Among them, y I =y V and I ≠y V They are positive samples (real image-text pairs) and negative samples (synchronously forged image-text pairs); N is the sample size of each batch; sim() is the similarity function; α is the margin value, which is used to adjust the similarity of negative samples to prevent the model from over-relying on negative samples that are easy to separate; 2.
4. Cross-integrate pictures, texts, videos and audios.
6. A method for detecting and locating multimodal media image and text synchronization forgery according to claim 5, characterized in that: In step 2.4, the text, video and audio are cross-fused. The specific steps are as follows: 2.4.
1. Based on cross-attention, audio and video information is integrated into the image and text respectively, so that context information is included in both the image and the text. The specific formula is as follows: Among them, LN() is the layer normalization operation; Image embedding after integrating audio and video information into the image; Text embedding after integrating audio and video information into the text; 2.4.
2. Based on the cross-fusion information of images and audio-visual, and the cross-fusion information of text and audio-visual, self-attention and residual networks are used to obtain image enhancement features and text enhancement features. The specific formula is as follows: Among them, SAT() represents the self-attention mechanism; is the enhanced image embedding; For enhanced text embedding; and They are the overall semantic embeddings of the enhanced image and text respectively; and They are the enhanced image patches and text word embeddings respectively.
7. The method for detecting and locating multimodal media image and text synchronization forgery according to claim 1, characterized in that: In step 3, image-text synchronization forgery detection and forgery location are performed, including the following steps: 3.1、Image and text synchronization forgery detection; 3.
2. Forged positioning with simultaneous images and text.
8. The method for detecting and locating multimodal media image and text synchronization forgery according to claim 7, characterized in that: In step 3.1, the image-text synchronization forgery detection includes image-text authenticity recognition and image-text synchronization forgery type recognition, and the specific steps of the image-text synchronization forgery detection are: 3.1.
1. Establish the understanding of global semantics of images and texts. The specific formula is as follows: in, and They are respectively the enhanced global semantic embeddings of images and texts formed by multimodal hierarchical fusion; concat() is a concatenation operation; FC is the fully connected layer; cls I,T Global semantic embedding for image and text pairs; 3.1.
2. Use binary classifiers and multi-label classifiers to realize the recognition of true and false images and texts and the recognition of the type of simultaneous forgery of images and texts. The specific formula is as follows: Among them, C bin () and C mul () are binary classifiers and multi-label classifiers, respectively. Both are built based on multi-layer perceptron (MLP) and use Gaussian error linear unit (GELU) as the activation function; y bin and mul They are true and false labels for images and texts and synchronous forgery type labels; H() is the cross entropy; E (I,T)~P is the mathematical expectation of the entropy distribution; and are the optimization objectives for binary classifiers and multi-label classifiers, respectively.
9. The method for detecting and locating multimodal media image and text synchronization forgery according to claim 7, characterized in that: In step 3.2, the synchronous forgery of image and text is located, including the forged image area location and the forged text word location, and the specific steps of the synchronous forgery of image and text are as follows: 3.2.
1. Use BBox detector to locate the forged image area, and calculate the positioning loss by combining l1 norm and generalized intersection-union ratio The specific formula is as follows: in, Embedding for image patches; D i () is the BBox detector using MLP and GELU as architecture and activation function respectively; Sigmoid is the activation function; y box Forge region localization labels for normalized images; || || is the l1 norm; It is a generalized intersection and comparison; E IF The mathematical expectation of the deviation distribution for locating the forged regions in the image; 3.2.
2. Use token anomaly detector to locate forged words in text, and calculate the location loss by combining cross entropy and KL divergence The specific formula is as follows: in, Embedding for text word tokens; D t () is a token anomaly detector built based on BERT, used to predict The abnormal probability of each word token in ; y tok Forge token location tags for the text; For the momentum version It is a momentum version of the token anomaly detector; H() is the cross entropy, which is used here to measure D t The difference between the detection result and the true label of (); KL[||] is the Kullback-Leibler divergence, which is used to evaluate D t ()and The probability distribution difference between the test results; α∈(0,1) is a hyperparameter used to balance the positioning accuracy and stability of the model; E TF The mathematical expectation of the distribution of deviations from the location of anomalous words in a text.
10. The method for detecting and locating multimodal media image and text synchronization forgery according to claim 1, characterized in that: In step 4, the detection and location multimodal media image and text synchronization forgery model is trained, and the specific formula is as follows: in, The global optimization objective for the entire system training.
Citation Information
Cited By
Multi-mode data-oriented multi-agent deep forgery attack detection system
CN120431529A
Double-branch sensing CLIP evidence obtaining method for generalizable deep counterfeiting detection
CN120451587A
Image forgery detection method and device based on CLIP model
CN120997613A
Text and image bimodal fusion-based risk identification system
CN121960712A
A risk identification system based on text and image bimodal fusion
CN121960712B