A multimodal media tampering detection method, system, device, and medium based on multi-view comparative learning
By employing a multi-view comparative learning approach, utilizing noise-enhanced alignment learning, multi-label comparative learning, and prototype-guided comparative learning, the problems of scattered tampering distribution and noise interference confusion in multimodal media tampering detection are solved, achieving more efficient tampering detection and localization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HENGYANG NORMAL UNIV
- Filing Date
- 2025-07-02
- Publication Date
- 2026-05-26
Smart Images

Figure CN120747720B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multimedia analysis technology, and in particular relates to a multimodal media tampering detection method, system, device and medium based on multi-view comparative learning. Background Technology
[0002] With advancements in generative models, highly realistic fake content can be generated, leading to multimodal media manipulation (DGM). 4 The problem is becoming increasingly serious. Existing methods typically align the original image-text pairs through contrastive learning and use cross-encoders to aggregate information from the image and text modalities for detection and localization. However, these methods suffer from two main problems:
[0003] 1. Dispersed Manipulation Distribution (DMD): Because all types of tampering are treated as a single category, tampered samples are scattered in the feature space, making it difficult to effectively distinguish between different types of tampering.
[0004] 2. Noise Interference Confusion (NIC): Noise and subtle tampering in the data can cause the model to misclassify the original sample as a tampered sample, reducing the accuracy of classification. Summary of the Invention
[0005] The purpose of this invention is to provide a multimodal media tampering detection method, system, device, and medium based on multi-view comparative learning, so as to solve the problems existing in the prior art.
[0006] To achieve the above objectives, this invention provides a multimodal media tampering detection method based on multi-view contrastive learning, comprising:
[0007] Obtain a training dataset, which includes training image-training text pairs and corresponding tampering category labels;
[0008] Based on the visual-language model, a cross encoder is introduced, several multi-layer perceptron head structures are set up, and three contrastive learning methods are designed: noise enhancement, prototype-based, and multi-label tampering classification, to obtain an initial multi-view contrastive learning framework.
[0009] The initial multi-view contrast learning framework is trained based on the training dataset to obtain the trained multi-view contrast learning framework; the tampering detection task of the image-text pair data to be detected is performed based on the trained multi-view contrast learning framework.
[0010] Optionally, training the initial multi-view contrast learning framework based on the training dataset specifically includes:
[0011] The training images are input into the image encoder of the initial multi-view contrast learning framework to extract image embeddings, and the training text is input into the text encoder of the initial multi-view contrast learning framework to extract global and local embeddings. The outputs of the image encoder and the text encoder are input into the cross encoder for aggregation processing. Based on the aggregated data, a detection and localization task is performed, and the framework is trained according to the target loss function to obtain the trained multi-view contrast learning framework. The target loss function includes noise enhancement alignment learning loss, multi-label contrast learning loss, prototype guided contrast learning loss, binary classification loss, image tampering localization loss, multi-label classification loss, and text tampering localization loss.
[0012] Optionally, the training process based on the noise-enhanced alignment learning loss specifically includes:
[0013] Determine the hard-to-distinguish positive sample image corresponding to the image embedding, calculate the mean and standard deviation of the hard-to-distinguish positive sample image, generate a Gaussian noise vector based on the calculated mean and standard deviation and add it to the image embedding to obtain the image embedding with noise added.
[0014] The similarity distribution matching loss between the image embedding and the text embedding after adding noise is calculated based on cosine similarity. The similarity distribution matching loss is normalized according to the degree of noise interference to obtain the noise-enhanced alignment learning loss. The noise-enhanced alignment learning loss is used to guide the model to minimize the interference caused by noise and improve the alignment accuracy.
[0015] Optionally, the training process based on the multi-label contrastive learning loss specifically includes:
[0016] Based on the momentum contrast framework, two queues are used to maintain the embeddings input into the initial multi-view contrast learning framework.
[0017] Based on the InfoNCE loss function, the initial multi-view contrastive learning framework is subjected to contrastive learning of global and local views to obtain global view contrastive loss and local view contrastive loss.
[0018] By combining the global view contrast loss and the local view contrast loss, a multi-label contrast loss is obtained; the multi-label contrast loss is used to combine information from the global view and the local view to improve the detection and localization performance of the initial multi-view contrast learning framework.
[0019] Optionally, the process of training based on the prototype-guided contrastive learning loss specifically includes:
[0020] Initialize the prototype vector of each tampered category as the mean of the embeddings of the corresponding category samples;
[0021] A cross encoder is used to fuse image data, text data, and prototype embeddings to update the prototype vector;
[0022] The contrastive loss function is used to optimize the alignment between image data and text data and prototype embeddings, so that the embeddings of samples of the same class are close to the corresponding prototype vectors, while the embeddings of samples of different classes are far away from the prototype vectors of other classes. For other classes, the contrastive loss is calculated symmetrically by replacing the corresponding features.
[0023] A multimodal media tampering detection system based on multi-view contrastive learning includes:
[0024] The data acquisition module is used to acquire the training dataset, which includes training image-training text pairs and corresponding tampering category labels;
[0025] The model building and training module is used to introduce a cross-encoder on the basis of the vision-language model, set up several multi-layer perceptron head structures, and design three contrastive learning methods: noise enhancement, prototype-based, and multi-label tampering classification, to obtain an initial multi-view contrastive learning framework; the initial multi-view contrastive learning framework is trained based on the training dataset to obtain the trained multi-view contrastive learning framework.
[0026] The model application module is used to perform tamper detection tasks on the image-text pair data based on the trained multi-view contrastive learning framework.
[0027] An electronic device includes a memory and a processor, the memory storing a computer program, and the processor running the computer program to enable the electronic device to perform a multimodal media tampering detection method based on multi-view contrastive learning.
[0028] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the multimodal media tampering detection method based on multi-view contrastive learning.
[0029] The technical effects of this invention are as follows:
[0030] This invention utilizes multi-view comparative learning to not only enhance the differences between tampering categories but also reduce the distance within the same category, thereby improving detection performance. Compared to UFAFormer and CrUr, this invention achieves better results on several key metrics while maintaining low model complexity. Attached Figure Description
[0031] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0032] Figure 1 This is a schematic diagram of the multi-view contrastive learning framework (MPCL) model structure in an embodiment of the present invention;
[0033] Figure 2 This is a schematic diagram of the results of each training cycle in the MPCL model in this embodiment of the invention.
[0034] Figure 3 This is a flowchart illustrating the implementation of an embodiment of the present invention. Detailed Implementation
[0035] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.
[0036] It should be understood that the terminology used in this invention is merely for describing particular embodiments and is not intended to limit the invention. Furthermore, with respect to numerical ranges in this invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Every smaller range between any stated value or intermediate value within a stated range, and any other stated value or intermediate value within said range, is also included in this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.
[0037] Various modifications and variations can be made to the specific embodiments described in this specification without departing from the scope or spirit of the invention, as will be apparent to those skilled in the art. Other embodiments derived from this specification will also be obvious to those skilled in the art. This application specification and embodiments are merely exemplary.
[0038] The terms “include,” “including,” “have,” “contain,” etc., used in this article are all open-ended terms, meaning that they include but are not limited to.
[0039] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0040] like Figure 1 - Figure 3As shown, this embodiment provides a multimodal media tampering detection method based on multi-view contrastive learning, including: acquiring a training dataset, the training dataset including training image-training text pairs and corresponding tampering category labels; introducing a cross-encoder based on a visual-language model, setting several multilayer perceptron head structures, and designing three contrastive learning methods: noise enhancement, prototype-based, and multi-label tampering classification, to obtain an initial multi-view contrastive learning framework; training the initial multi-view contrastive learning framework based on the training dataset to obtain a trained multi-view contrastive learning framework; and performing a tampering detection task on the image-text pair data to be detected based on the trained multi-view contrastive learning framework.
[0041] This embodiment aims to provide an improved multi-view contrastive learning framework (MPCL) to effectively detect and locate media tampering in image and text bimodal media. By addressing the problems of scattered tampering distribution and noise interference in existing methods, it improves the accuracy and robustness of multi-label classification.
[0042] To achieve the above objectives, this embodiment provides a multi-view contrastive learning framework (MPCL), which includes the following three key components:
[0043] 1. Noise-Augmented Alignment Learning (NAAL):
[0044] By simulating noise in the original samples and optimizing the global similarity distribution, the robustness of the model to noise interference is improved.
[0045] A noise-enhanced representation is generated by embedding Gaussian noise based on the sample mean and variance into the original image.
[0046] The Similarity Distribution Matching (SDM) loss function is used to ensure the consistency of the global similarity distribution.
[0047] 2. Multi-Label Contrastive Learning (MLCL):
[0048] Clustering embeddings of different tampering types separately enhances the accuracy of multi-label classification.
[0049] By comparing and learning from global and local views, fine-grained tampering traces are captured, ensuring that features of different tampering categories are more distinct.
[0050] 3. Prototype-Guided Contrastive Learning (PCL):
[0051] Each tampered category is represented by a learnable prototype vector, attracting embeddings of the same category and repelling embeddings of different categories.
[0052] Prototype vectors interact with image and text embeddings through a cross-modal encoder, enhancing the discriminative power of features.
[0053] This embodiment uses the hierarchical multimodal tampering inference Transformer (HAMMER) as the baseline model. For example... Figure 1 As shown, the MPCL framework in this embodiment includes three encoders—an image encoder, a text encoder, and a cross encoder—and multiple multilayer perceptron (MLP) heads for classification and localization. Given an image-text pair (I... i ,T i In this embodiment, image I is first... i It is split into a series of M non-overlapping patches and fed into an image encoder to extract the image embedding v. cls ,v1,…,v M , where v cls As a [CLS] marker. Similarly, for input text T i In this embodiment, a text encoder is used to obtain global and local embeddings. cls ,t1,…,t N Where N is the sequence length. In this embodiment, the learnable query {p1, p2, ..., p5} is initialized as a prototype embedding, where 5 corresponds to the five categories (original and four tampered) in the dataset, and each prototype attempts to match the corresponding input image or text. To address the DMD and NIC problems, this embodiment further employs noise-enhanced alignment learning. Multi-label comparison learning Comparative learning with prototype-guided learning Furthermore, the cross-encoder aggregates image, prototype, and text embeddings for the prediction task, leveraging the last six layers of BERT combined with masked language modeling. Enhance cross-modal interaction. Following HAMMER, this embodiment uses the aggregate representation {m cls ,m1,…,m M Perform binary classification Image tampering location Multi-tag classification and text tampering location For simplicity, this embodiment refers to these losses collectively as The overall MPCL loss function formula is:
[0054]
[0055] In the formula, β, γ and η are set to 0.5, 0.1 and 0.1, respectively.
[0056] Noise-enhanced alignment learning (NAAL): In each batch, the most difficult-to-distinguish original samples are selected and a noise vector based on their mean and variance is added.
[0057] We apply similarity distribution matching (SDM) loss to optimize the global similarity distribution of image and text embeddings.
[0058] Noise augmentation representation construction: given the global embedding {v} in each batch cls ,t cls First, select an original image for embedding. The most difficult positive sample image associated with the embedding is then identified. In this embodiment, the mean μ and standard deviation σ of the most difficult positive sample are then calculated. Gaussian noise generated based on these parameters is subsequently added to the instance-normalized image embedding. This process can be represented as:
[0059]
[0060] in, This indicates noise enhancement. This represents the noise vector.
[0061] Noise-enhanced alignment: To maintain intra-class consistency in the presence of noise, this embodiment uses Similarity Distribution Matching (SDM) loss, which incorporates the cosine similarity distribution of the image-text pair embeddings into the Kullback-Leibler (KL) divergence to associate representations between different modalities. For a batch of image-text pairs, this embodiment creates a set of K pairs. Where (y) ij =1) means v cls,i and t cls,j This is the original embedding. Then, in this embodiment, the softmax function is used to calculate the probability of a matching pair:
[0062]
[0063] Where τ is the temperature hyperparameter, and sim() represents the cosine similarity between the two embeddings. The image-to-text SDM loss can be calculated using the following formula:
[0064]
[0065] in, This represents the true original probability distribution, where ∈ is a small constant added for numerical stability. The text-to-image SDM loss is calculated symmetrically, and the overall SDM loss is defined as:
[0066]
[0067] To further guide the model to minimize noise-induced interference, this embodiment employs a dynamic method to adjust the sensitivity of the loss function during the image-text alignment process with noise enhancement. This embodiment adjusts the sensitivity based on the degree of noise interference (denoted as σ). 2 )right Normalization is performed, with emphasis on penalizing high levels of interference. This process can be expressed as:
[0068]
[0069] Overall, noise-enhanced alignment learning is defined as:
[0070]
[0071] Multi-label contrastive learning (MLCL):
[0072] Use two queues to maintain the most recent image and text representations respectively, increase the batch size and include more samples from different categories.
[0073] The InfoNCE loss function is used to perform comparative learning of global and local views to ensure the embedding separation of different tampering categories.
[0074] Based on the noise-enhanced primal alignment learning used for binary classification, this embodiment focuses on more detailed multi-label classification challenges, taking a specific tampering category (e.g., face swapping) as an example. In this case, there is a significant semantic difference between the two modalities because the text accompanying the tampered image remains unchanged. The cross-encoder utilizes this difference to detect tampering during the fusion process. Therefore, the main objective of this embodiment is to amplify this difference and group similar tampering categories into unique clusters before fusion.
[0075] Following MoCo and HAMMER++, this embodiment utilizes two queues to maintain the M most recent image-text representations from a single-modal encoder. The introduction of queues provides a larger batch size and more samples from different categories. Furthermore, HAMMER++ introduces contrastive learning on both global and local views. This embodiment also leverages this loss to capture subtle tampering traces that may only affect a few local patches. Specifically, this embodiment employs the InfoNCE loss, which is formulated as follows:
[0076]
[0077] Where Q is the queue being maintained, and τ is the temperature parameter.
[0078] Unlike traditional cross-modal contrastive learning, this embodiment focuses on unimodal alignment, aiming to cluster the embeddings of four basic categories. The global view of CL can be represented as:
[0079]
[0080] A partial view in a CL is achieved using the following formula:
[0081]
[0082] Where M and N represent the number of image blocks and text tags, respectively.
[0083] The overall multi-label contrast loss is calculated as follows:
[0084]
[0085] Prototype-guided contrastive learning (PCL):
[0086] Initialize the prototype vector of each tampered category as the mean of the corresponding category sample embeddings.
[0087] A cross encoder is used to fuse the image, text, and prototype embeddings, and then the prototype vector is updated.
[0088] By comparing loss functions, the sample embeddings are aligned with the corresponding prototype vectors.
[0089] The core idea of prototype learning is to learn a set of representative prototypes, compressing the core features of each category in the data distribution. Compared to image-text contrastive learning, prototype learning provides a more intuitive way to align instance-level samples with their corresponding category prototypes.
[0090] Prototype initialization and update:
[0091] Given a batch of global embeddings This embodiment initializes the prototype P = {p1, p2, ..., p5} for each category by calculating the mean vector of the corresponding features. Specifically, for the original samples, this embodiment will... and They are considered to be in the same category. This process can be described as follows:
[0092]
[0093] Where, p c x represents the prototype of category c. cls B represents the embedding of the corresponding category c. c This represents the set of samples belonging to category c.
[0094] Following BLIP, this embodiment inputs queries, images, and text into a cross-encoder to obtain the updated prototype P' = {p'1, p'2, ..., p'5} and aggregate representation M = {m cls ,m1,…,m M}, where M represents the number of image patches. The built-in attention mechanism in the cross-encoder detects cross-modal differences by comparing image and text features with a predefined category-specific query. This process can be represented as:
[0095] {P',M}=Transformer(Concat[V,P],T,T)
[0096] This embodiment also introduces [a certain technique] during the fusion process. To help identify semantic differences.
[0097] To further smooth out changes in prototype features, this embodiment updates the prototype for each category using the features of the current mini-batch of samples. The update process can be represented by the following equation:
[0098] p′ i ←αp′ i +(1-α)x i
[0099] This embodiment uses single-modal alignment for optimization. Prototype-guided comparative learning can be described as follows:
[0100]
[0101] Losses for other categories are calculated symmetrically by replacing the corresponding features respectively. The overall goal of prototype-guided contrastive learning can be stated as:
[0102]
[0103] Table 1. In DGM 4 Performance comparison with state-of-the-art methods on datasets
[0104]
[0105] This embodiment compares the proposed MPCL method with recent state-of-the-art methods, as shown in Table 1. First, this embodiment compares it with the most influential works—HAMMER and HAMMER++. These methods further separate tampering features by aligning the original image-text pairs, leading to semantic inconsistencies when interacting with tampering features. As analyzed earlier in this embodiment, this approach inevitably results in the dispersion of tampering distributions and noise interference. The motivation of this embodiment is to address these problems. By leveraging multi-view contrastive learning, MPCL increases the differences between different tampering categories while reducing intra-category distances, thereby further amplifying semantic inconsistencies and achieving significant improvements across all metrics. Next, this embodiment evaluates its method against another method, UFAFormer, which is inspired by the success of frequency domain information in deep forgery detection, where tampering typically leaves detectable artifacts in the frequency domain. UFAFormer integrates frequency domain features, images, and text into a cross-encoder, aiming to capture spatial and frequency inconsistencies to improve detection performance. Notably, their multi-label classification metric shows only modest improvements. This limited improvement is primarily due to their methods merely supplementing information without effectively distinguishing between different types of tampering. Finally, this embodiment compares its method with CrUr, which captures semantic inconsistencies through bidirectional reconstruction of image and text features. While CrUr demonstrates competitiveness in text localization, particularly with precision, recall, and F1 score of 80.13%, 77.71%, and 78.90%, respectively, this comes at the cost of extremely high model complexity due to the stacking of multiple Transformer blocks. In contrast, the MPCL method of this embodiment outperforms CrUr on most key metrics by introducing a small number of prototype embeddings into the VLP, while maintaining lower model complexity. Therefore, MPCL shows a significant advantage in overall performance and efficiency.
[0106] Ablation studies:
[0107] Table 2. In DGM 4 Ablation study of the component proposed in this embodiment on the dataset
[0108]
[0109] In the 9th row of the methods in Table 2, the four icons from left to right represent noise-enhanced alignment learning, multi-label contrastive learning, masked language modeling, and prototype-guided contrastive learning, respectively.
[0110] This embodiment analyzes the effectiveness of the optimization objective in the MPCL framework. Here, this embodiment uses Image-Text Contrast Loss (ITC) as the baseline. Starting from No.1, this embodiment replaces ITC with Similarity Distribution Matching (SDM) in all subsequent experiments.
[0111] Baseline Comparison (ITC vs. SDM): As shown in Table 2, replacing ITC with SDM (No. 0 and No. 1) provided significant improvements across all metrics, with EER decreasing significantly from 17.66% to 15.57%. This indicates that SDM provides more effective cross-modal feature alignment, enabling a clearer distinction between original and tampered content, thus laying a more solid foundation for the integration of subsequent components.
[0112] Effectiveness of NAAL: The introduction of NAAL (No.2) further improved the performance of No.1, especially in binary classification metrics and text localization, with ACC increasing from 84.02% to 86.00% (+1.98%). This indicates that NAAL helps the model better resist noise interference by simulating noise and optimizing the global similarity distribution, thereby achieving more accurate classification and localization.
[0113] Effectiveness of MLCL: The introduction of multi-label contrastive learning (No. 3) significantly improved the performance of multi-label classification tasks. Specifically, mAP improved from 83.49% to 86.60% of the baseline SDM, CF1 from 76.38% to 80.47%, and OF1 from 76.44% to 80.81%. These improvements highlight the effectiveness of MLCL in distinguishing DGMs. 4 The effectiveness of MLCL across multiple tampering categories in the dataset is demonstrated. By clustering similar tampering types into different clusters, MLCL ensures that the features of different tampering categories are more explicit, reducing overlap and ambiguity in classification. The effectiveness of MLCL is further validated in No. 5, where combining MLCL with NAAL brings even greater improvements, further strengthening MLCL's role in enhancing feature discrimination. Furthermore, the improvements in CF1 and OF1 demonstrate MLCL's ability to maintain a balance between precision and recall, ensuring that the model not only accurately identifies multiple tampered labels but also captures more true positive examples from various categories.
[0114] Effectiveness of MLM: Masked Language Modeling (MLM) has proven effective in previous studies. Inspired by this, this embodiment introduces MLM into the fusion process within the MPCL framework, resulting in significant improvements across several metrics. By masking certain markers during cross-modal interactions, MLM forces the model to predict missing information, thereby deepening semantic understanding and improving alignment between image and text modalities.
[0115] Effectiveness of PCL: In the initial experiments, this embodiment introduced prototypes into the cross-encoder but did not provide a specific loss function for their learning (No. 6). This embodiment observed a slight decrease in performance across several metrics. Specifically, mAP decreased from 87.08% (No. 5) to 85.12%, and CF1 and OF1 decreased from 77.83% and 76.97% to 76.83% and 76.12%, respectively. This decrease indicates that while introducing prototypes into the cross-encoder provides additional semantic anchors, the lack of a specific optimization objective limits their effective contribution to the model learning process. However, when this embodiment introduced Prototype-Guided Contrastive Learning (PCL) and provided an explicit loss function for the optimization of these prototypes (No. 7), performance significantly improved, with mAP increasing to 88.35%, and CF1 and OF1 increasing to 84.68% and 84.66%, respectively. These results highlight the importance of setting clear objectives for prototype learning. By explicitly guiding prototype learning, PCL is able to better align features with corresponding tampering categories, resulting in more explicit and meaningful clusters. Furthermore, the improved image localization metric, with the IoU mean increasing from 75.53% to 77.52%, further validates that PCL not only improves the classification task but also enhances the accuracy of tamper localization in images. This indicates that PCL effectively strengthens the semantic representation of prototypes, making them more powerful anchors in cross-modal interactions, thereby improving the overall performance of the model.
[0116] t-SNE visualization: in Figure 2 In this embodiment, as the model progresses through each epoch (Epoch0 to EpochBest), the following key improvements are observed:
[0117] Prototype-guided clustering: In successive epochs, learnable prototype vectors effectively attract features of the same category, resulting in clusters of each modified category becoming more compact and unique. This embodiment clearly demonstrates the process of prototype-guided feature convergence as the clusters gradually tighten.
[0118] Multi-label contrastive learning: The MLCL component plays a crucial role in separating different tampering categories. As training progresses, clusters between different tampering categories become more apparent, and the overlap between categories is significantly reduced. This separation ensures that features of different tampering types do not interfere with each other, thus maintaining clear decision boundaries.
[0119] Noise-enhanced alignment learning: A significant effect of NAAL was observed when processing the original samples. Initially, the original features were mixed with the tampered features (as seen in early epochs). However, with the application of NAAL, the original features gradually aligned closer to their correct clusters, thus reducing confusion with the tampered features.
[0120] The MPCL method proposed in this example demonstrates a comprehensive performance improvement over traditional models such as HAMMER in the detection and localization of image-text pair tampering, especially in dealing with three types of high-difficulty challenges.
[0121] In terms of perceiving subtle alterations, traditional methods struggle to identify minute replacements hidden in complex backgrounds or image noise, while MPCL utilizes reconstruction and detail analysis techniques to accurately capture these hidden visual differences.
[0122] In terms of deep semantic understanding, when there is a semantic conflict between text and image, traditional methods are easily misled by surface concepts, while MPCL, through deep contextual understanding, can grasp the core semantic elements and accurately identify the logical inconsistencies in text replacement.
[0123] In terms of local change detection, traditional methods often suffer from missed detections when dealing with image attribute editing in small areas. However, MPCL, through precise local feature analysis, can reliably identify these subtle modifications, demonstrating its powerful capabilities in refined detection tasks.
[0124] Overall, these examples demonstrate that MPCL significantly outperforms HAMMER in detecting and locating tampered text. By leveraging multi-label contrastive learning and prototype-guided contrastive learning, MPCL improves its ability to distinguish between real and tampered content, reducing false positives and missed detections, especially in semantically complex texts.
[0125] Hyperparameter sensitivity analysis:
[0126] Table 3 in DGM 4 Impact of different hyperparameter settings on MPCL performance on the dataset
[0127]
[0128] The proposed MPCL framework uses multiple loss functions to optimize DGM. 4 Each subtask of the dataset. For loss functions inherited from ALBEF or HAMMER (such as L...). mlm and L DGM4 In this embodiment, the hyperparameter settings of these functions are directly adopted to ensure fair comparison. Therefore, this embodiment only performs sensitivity analysis on the weights of NAAL, MLC L, and PCL losses, which are unique loss functions in the method of this embodiment.
[0129] As shown in Table 3, the performance metrics of all subtasks show little variation under different hyperparameter settings. This indicates that the proposed MPCL framework is robust to changes in the weights of these loss functions. Whether in binary classification, multi-label classification, or localization tasks, MPCL maintains stable performance, further demonstrating its stability in handling multimodal manipulation detection tasks.
[0130] Single-domain performance comparison:
[0131] Table 4 DGM 4 Comparison of deepfake detection methods for datasets
[0132]
[0133] Table 5 shows the use of sequence labeling methods in DGM. 4 Comparison of fake text detection results on dataset
[0134]
[0135] This embodiment comprehensively compares the MPCL method with existing single-modal models. Through MPCL learning, the single-modal encoder in this embodiment has the ability to effectively cluster different tampering features. This enables the encoder to achieve high performance in DGM even with only fine-tuning using a simple prediction head. 4 The MPCL performs exceptionally well in the task. As shown in Tables 4 and 5, MPCL not only excels in binary classification tasks but also performs admirably in tamper localization tasks. Specifically, MPCL improves the semantic inconsistency between different tamper categories, enabling the encoder to better capture and distinguish these subtle differences.
[0136] These results demonstrate that MPCL has a stronger ability to accurately locate tampered content, further proving its broad applicability in multimodal data processing.
[0137] Cross-domain performance comparison:
[0138] Table 6 Comparison Results of Cross-Domain Face Spoofing Detection Generalization Ability
[0139]
[0140] The results in Table 6 highlight the robustness and effectiveness of the MPCL method, particularly in challenging scenarios where manipulation is unseen, with the model tested directly on different fake datasets without prior exposure to them. MPCL significantly outperforms other methods with a leading AUC of 66.80% in distinguishing between real and fake content under these conditions. Furthermore, MPCL exhibits the lowest EER (36.20%), demonstrating its superior ability to balance false positives and false negatives, a crucial factor in ensuring model reliability, especially when dealing with unseen manipulation. An accuracy of 65.67% further demonstrates MPCL's strong generalization ability.
[0141] MPCL's significant improvement stems from its multi-view contrastive learning mechanism, which effectively captures and clusters different forgery features, even when faced with unseen data. This mechanism enhances the model's ability to discern subtle semantic inconsistencies between forgery types, improving detection accuracy across different datasets. These results demonstrate that MPCL not only excels in standard deepfake detection tasks but also exhibits remarkable robustness and accuracy when applied to unseen manipulation, representing a major advancement in the field.
[0142] A multimodal media tampering detection system based on multi-view contrastive learning includes:
[0143] The data acquisition module is used to acquire the training dataset, which includes training image-training text pairs and corresponding tampering category labels;
[0144] The model building and training module is used to introduce a cross-encoder on the basis of the vision-language model, set up several multi-layer perceptron head structures, and design three contrastive learning methods: noise enhancement, prototype-based, and multi-label tampering classification, to obtain an initial multi-view contrastive learning framework; the initial multi-view contrastive learning framework is trained based on the training dataset to obtain the trained multi-view contrastive learning framework.
[0145] The model application module is used to perform tamper detection tasks on the image-text pair data based on the trained multi-view contrastive learning framework.
[0146] An electronic device includes a memory and a processor, the memory storing a computer program, and the processor running the computer program to enable the electronic device to perform a multimodal media tampering detection method based on multi-view contrastive learning.
[0147] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the multimodal media tampering detection method based on multi-view contrastive learning.
[0148] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A multimodal media tampering detection method based on multi-view comparative learning, characterized in that, include: Obtain a training dataset, which includes training image-training text pairs and corresponding tampering category labels; Based on the visual-language model, a cross encoder is introduced, several multi-layer perceptron head structures are set up, and three contrastive learning methods are designed: noise enhancement, prototype-based, and multi-label tampering classification, to obtain an initial multi-view contrastive learning framework. The initial multi-view contrast learning framework is trained based on the training dataset to obtain the trained multi-view contrast learning framework; the tampering detection task of the image-text pair data to be detected is performed based on the trained multi-view contrast learning framework. The training of the initial multi-view contrast learning framework based on the training dataset specifically includes: The training images are input into the image encoder of the initial multi-view contrast learning framework to extract image embeddings, and the training text is input into the text encoder of the initial multi-view contrast learning framework to extract global and local embeddings. The outputs of the image encoder and the text encoder are input into the cross encoder for aggregation processing. Based on the aggregated data, a detection and localization task is performed, and the framework is trained according to the target loss function to obtain the trained multi-view contrast learning framework. The target loss function includes noise enhancement alignment learning loss, multi-label contrast learning loss, prototype-guided contrast learning loss, binary classification loss, image tampering localization loss, multi-label classification loss, and text tampering localization loss. The training process based on the noise-enhanced alignment learning loss specifically includes: Determine the hard-to-distinguish positive sample image corresponding to the image embedding, calculate the mean and standard deviation of the hard-to-distinguish positive sample image, generate a Gaussian noise vector based on the calculated mean and standard deviation and add it to the image embedding to obtain the image embedding with noise added. The similarity distribution matching loss between the image embedding and the text embedding after adding noise is calculated based on cosine similarity. The similarity distribution matching loss is normalized according to the degree of noise interference to obtain the noise-enhanced alignment learning loss. The noise-enhanced alignment learning loss is used to guide the model to minimize the interference caused by noise and improve the alignment accuracy. The training process based on the multi-label contrastive learning loss specifically includes: Based on the momentum contrast framework, two queues are used to maintain the embeddings input into the initial multi-view contrast learning framework. Based on the InfoNCE loss function, the initial multi-view contrastive learning framework is subjected to contrastive learning of global and local views to obtain global view contrastive loss and local view contrastive loss. The multi-label contrast loss is obtained by combining the global view contrast loss and the local view contrast loss; the multi-label contrast loss is used to combine the information of the global view and the local view to improve the detection and localization performance of the initial multi-view contrast learning framework. The process of training based on the prototype-guided contrastive learning loss specifically includes: Initialize the prototype vector of each tampered category as the mean of the embeddings of the corresponding category samples; A cross encoder is used to fuse image data, text data, and prototype embeddings to update the prototype vector; The contrastive loss function is used to optimize the alignment between image data and text data and prototype embeddings, so that the embeddings of samples of the same class are close to the corresponding prototype vectors, while the embeddings of samples of different classes are far away from the prototype vectors of other classes. For other classes, the contrastive loss is calculated symmetrically by replacing the corresponding features.
2. A multimodal media tampering detection system based on multi-view comparative learning, employing the method described in claim 1, characterized in that, include: The data acquisition module is used to acquire the training dataset, which includes training image-training text pairs and corresponding tampering category labels; The model building and training module introduces a cross encoder on the basis of the vision-language model, sets up several multi-layer perceptron head structures, and designs three contrastive learning methods: noise enhancement, prototype-based, and multi-label tampering classification, to obtain an initial multi-view contrastive learning framework. The initial multi-view contrast learning framework is trained based on the training dataset to obtain the trained multi-view contrast learning framework. The model application module is used to perform tamper detection tasks on the image-text pair data based on the trained multi-view contrastive learning framework.
3. An electronic device, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program and the processor runs the computer program to enable the electronic device to perform a multimodal media tampering detection method based on multi-view contrast learning as described in claim 1.
4. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the multimodal media tampering detection method based on multi-view contrast learning as described in claim 1.