Multi-mode control detection and positioning method and system
By introducing large-scale models to generate descriptive and explanatory texts, combined with the cross-attention mechanism of manipulation guidance, the problem of insufficient semantic alignment in multimodal content manipulation detection and positioning is solved, and more efficient manipulation detection and positioning accuracy is achieved, and it is suitable for the field of multimodal content security.
Patent Information
- Application Number
- CN202510481491.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-09-05
AI Technical Summary
There are problems in the existing multimodal content manipulation detection and positioning techniques such as insufficient semantic alignment, insufficient attention to the manipulation area and insufficient utilization of auxiliary features, resulting in poor control area recognition effect.
By introducing large-scale models to generate descriptive and explanatory text, a large-model assisted alignment module based on comparison learning is designed, combined with the cross attention mechanism of manipulation guidance, the semantic alignment ability of multimodal input is strengthened, and the annotation information of the manipulation area is introduced during the training stage to adjust attention allocation.
The accuracy and reliability of multimodal manipulation detection and positioning have been significantly improved, especially in the ability to capture fine-grained manipulation areas, which improves the accuracy of manipulation detection and the perception ability of the model while maintaining the efficiency of the inference stage.
Smart Images

Figure CN120599211A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning and artificial intelligence technology, and in particular to a method and system for multimodal manipulation detection and positioning. Background Art
[0002] In recent years, with the rapid development of deep learning and artificial intelligence technologies, generative models have demonstrated powerful capabilities in the field of cross-modal content generation. In particular, large-scale generative models based on the Transformer architecture (such as GPT-4, Stable Diffusion, etc.) have been able to generate highly realistic content in modalities such as images and text. However, the advancement of these technologies has also brought serious information security issues, such as the forgery and manipulation of multimodal content. Therefore, how to efficiently detect and locate multimodal content manipulation (Detecting and Grounding Multi-Modal Media Manipulation, DGM4 for short) has become a key issue. The purpose of DGM4 is not only to detect the authenticity of multimodal content, but also to locate the manipulated content (such as image bounding boxes and text labels), so deeper reasoning about multimodal content manipulation is needed.
[0003] Current DGM4-related methods on the market primarily rely on a coarse-grained understanding of multimodal inputs. For example, the HAMMER model uses a shallow manipulation reasoning module to locate the manipulation area and designs a deep reasoning structure to identify the manipulated text and image regions. Furthermore, the VIKI model introduces a knowledge interaction mechanism to optimize the alignment of visual and linguistic features by minimizing geometric distance. However, these methods primarily rely on annotation information when processing manipulation content and do not adequately address the semantic alignment of multimodal inputs.
[0004] In summary, the existing technology has the following defects:
[0005] 1. Insufficient semantic alignment: Most current methods rely on annotation information and ignore the fine-grained semantic alignment between images and text. The lack of such alignment leads to poor recognition of manipulated areas.
[0006] 2. Insufficient attention to the control area: The attention mechanism of the existing model has difficulty in effectively allocating sufficient attention weight when processing the control area, thereby weakening the ability to perceive the control area.
[0007] 3. Insufficient utilization of manipulated sample features: When processing manipulated samples, there is a lack of full utilization of auxiliary features (such as descriptive text and explanatory text), which limits the detection and positioning performance of the model. Summary of the Invention
[0008] In order to overcome the above-mentioned defects in the existing technology, the present invention provides a method and system for multimodal manipulation detection and positioning, aiming to address the shortcomings of existing multimodal content manipulation detection and positioning technologies, and propose a solution that can efficiently perceive the fine-grained semantic alignment of multimodal content, so as to improve the accuracy and reliability of manipulation detection and positioning.
[0009] To achieve the above object, the present invention adopts the following technical solutions, including:
[0010] A method for multimodal manipulation detection and positioning, comprising the following steps:
[0011] S1, performing feature extraction on input multimodal data to obtain image features and text features respectively; the multimodal data includes images and text;
[0012] S2, generates descriptive text for the input image and calculates the alignment loss between the image and the descriptive text;
[0013] S3, based on the cross attention matrix, multimodal fusion of image features and text features is performed to obtain fused image features and fused text features; the manipulation guidance matrix is introduced to assign weights to the manipulation area, and the attention matrix loss after manipulation guidance is calculated;
[0014] S4 uses the fused multimodal features for detection and positioning, including: determining whether the multimodal data has been manipulated, identifying the type of manipulation, locating the manipulated area in the image, and locating the manipulated area in the text; based on the detection and positioning results, calculating the discrimination loss, classification loss, image positioning loss, and text positioning loss; introducing an indicator guidance matrix to emphasize the manipulated area in the image, and calculating the image discrimination loss after indicator guidance;
[0015] S5, calculates the total model loss based on the alignment loss between the image and the descriptive text, the attention matrix loss after manipulation guidance, the discrimination loss, the classification loss, the image localization loss, the text localization loss, and the image discrimination loss after instruction guidance;
[0016] S6, trains the model based on the total model loss, and the trained model is used to detect and locate the manipulation of multimodal data.
[0017] Preferably, in step S1, the image I is passed through the image encoder E v Perform block processing and extract visual features to obtain the image global classification feature i cls and image patch features Where N is the number of image blocks, Represents the features of the i-th image block; the text T is encoded by the text encoder E t Perform word segmentation and extract semantic features to obtain the global classification feature t of the text clsand text word features Where M is the number of words in the text, Represents the j-th word feature.
[0018] Preferably, in step S2, the image I is input into the multimodal large model to generate a descriptive text C, and the descriptive text C is passed through the text encoder E t Perform feature extraction to obtain the descriptive text global classification feature c cls , where H is the number of descriptive text words:
[0019] Alignment loss L between image I and descriptive text C I→C for:
[0020]
[0021] Where s(·,·) represents the feature similarity function; τ is the coefficient; and B is the set of small batch samples.
[0022] Preferably, in step S3, multimodal fusion of image features and text features is performed based on the cross attention matrix A:
[0023] F I =A*T w
[0024] F T =A*I w
[0025] Among them, F I is the fused image feature, F T is the fused text feature; image block feature Text word features
[0026] The calculation formula of the cross attention matrix A is:
[0027]
[0028] Among them, W t and W v is the feature mapping matrix; W t ∈R d×d‘ , W v ∈R d×d' ; d is the feature dimension, d' is the hidden dimension; d is the feature dimension; I w ∈R N×d , T w ∈R M×d ;A∈R N×M ;
[0029] Introduce the control guidance matrix G:
[0030]
[0031] Where G(i,j) is the value of the element in the i-th row and j-th column of the manipulation guidance matrix G, indicating whether the i-th image block or the j-th word is manipulated;
[0032] Attention matrix loss L after manipulation guidance MGCA for:
[0033]
[0034] Among them, A(i,j) is the element value of the i-th row and j-th column in the cross-attention matrix A, which represents the cross-attention between the i-th image block and the j-th word.
[0035] Preferably, in step S4, it is determined whether the multimodal data is manipulated and the discrimination loss L is calculated. bin for:
[0036]
[0037] y bin =Sigmoid(MLP(F T ))
[0038] Among them, y bin The predicted label for whether it is manipulated; is the real label of whether it is manipulated; F T is the fused text feature; MLP represents the mapping function; Sigmoid represents the activation function;
[0039] Identify the manipulation type and calculate the classification loss L mul for:
[0040]
[0041] y k =Sigmoid(MLP(F T ))
[0042] Among them, y k is the predicted probability of belonging to the k-th manipulation type, is the true probability of belonging to the kth manipulation type;
[0043] Locate the manipulated area in the image and calculate the image positioning loss L bbox for:
[0044] L bbox =|b pred -b true |+L IoU (bpred ,b true )
[0045] b pred =Sigmoid(MLP(F I ))
[0046] Among them, b pred is the predicted area of image manipulation, b true is the real area of image manipulation; F I is the fused image feature;
[0047] Locate the manipulation area in the text and calculate the text positioning loss L tok for:
[0048]
[0049] in, is the predicted manipulation probability of the jth word in the text, is the true manipulation probability of the jth word in the text.
[0050] Preferably, in step S4, an indication steering matrix P is introduced to emphasize the manipulated area in the image. The indication steering matrix P is:
[0051]
[0052] Among them, P i The element value of the i-th image block in the steering matrix P indicates whether the i-th image block is manipulated;
[0053] Calculate the image discrimination loss L after instruction guidance pmm for:
[0054]
[0055] Among them, F I is the fused image feature, Used to calculate the probability of each image block being manipulated, represents the fused i-th image block feature, and N is the number of image blocks.
[0056] The present invention also provides a multimodal manipulation detection and positioning system, characterized in that, applied to the multimodal manipulation detection and positioning method, the system includes: an image encoder, a text encoder, a large model auxiliary alignment module, a manipulation-guided attention module, a task head module, a manipulation area indicator, and a local modeling unit;
[0057] The image encoder is used to receive the input image, perform block processing and feature extraction on the image, and output image block features and image global classification features;
[0058] The text encoder is used to receive input text, perform word segmentation and feature extraction on the text, and output text word features and text global classification features;
[0059] The large model-assisted alignment module is used to receive input images and generate descriptive text, construct contrastive learning objectives, and enhance the alignment of images and descriptive text;
[0060] The manipulation-guided attention module is used to fuse image block features and text word features based on a cross-attention method, and to adjust the attention distribution and strengthen the attention weight of the manipulation area by receiving the annotation information of the manipulation area;
[0061] The task head module includes: a binary classifier for determining whether the multimodal data has been manipulated; a multi-classifier for identifying the type of manipulation; a bounding box detector for locating the manipulated area in the image based on the fused image features; and a text detector for locating the manipulated area in the text based on the fused text features.
[0062] The control area indicator is used to construct an indication guidance matrix based on the control area annotation information of the image;
[0063] The local modeling unit is used to enhance the training of bounding box detectors by assisting learning with an indicative guidance matrix.
[0064] The present invention also provides an electronic device, which includes a processor, a memory, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the method of multimodal manipulation detection and positioning is implemented.
[0065] The present invention also provides a computer program product, which includes a computer program / instruction, and when the computer program / instruction is executed by a processor, it implements the method of multimodal manipulation detection and positioning.
[0066] The present invention also provides a readable storage medium having a computer program stored thereon, and when the computer program is executed, the method for multimodal manipulation detection and positioning is implemented.
[0067] The advantages of the present invention are:
[0068] (1) The present invention aims to address the shortcomings of existing multimodal content manipulation detection and positioning technologies and propose a solution that can efficiently perceive the fine-grained semantic alignment of multimodal content, so as to improve the accuracy and reliability of manipulation detection and positioning.
[0069] (2) This paper introduces auxiliary text generated by large-scale models, using pre-trained multimodal large models and large language models to generate descriptive and explanatory text as additional semantic clues to improve the accuracy of image-text alignment. A large-scale model-assisted alignment based on contrastive learning is designed. By constructing a contrastive learning objective between the image and the auxiliary text, the semantic alignment capability of multimodal input is enhanced, especially for capturing fine-grained control areas.
[0070] (3) The present invention designs a cross-attention mechanism for manipulation guidance, which, combined with the annotation information in the training phase, significantly enhances the model's ability to allocate attention to the manipulation area, thereby improving the accuracy of manipulation detection and positioning, and further optimizing the model's perception of the manipulation area.
[0071] (4) The large model auxiliary alignment module (i.e., step S2) designed by the present invention is only introduced in the training phase and does not increase the computational burden of the inference phase, which ensures high efficiency in practical applications.
[0072] (5) The present invention has achieved significant performance improvement in multimodal manipulation detection and positioning tasks, which not only improves the accuracy of detection, but also enhances the model's ability to identify the manipulation area, and can better meet the actual requirements in the field of multimodal content security. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 The present invention provides a flowchart of a method and system for multi-modal manipulation detection and positioning.
[0074] Figure 2 Schematic diagram of the model architecture of the present invention. DETAILED DESCRIPTION
[0075] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0076] Depend on Figure 1 As shown, the multimodal manipulation detection and positioning system of the present invention includes the following main components: image encoder 1, text encoder 2, large model auxiliary alignment module 3, manipulation-guided attention module 4, task head module 5, and manipulation area indicator and local modeling unit 8. The structural composition, function, and connection relationship of each component are as follows:
[0077] The image encoder 1 serves as the input end of the system, is used to receive the input image, perform block processing and feature extraction on the image, and output image block features and image global classification features. The output of the image encoder is connected to the manipulation-guided attention module.
[0078] The text encoder 2 also serves as the input end of the system, which is used to receive input text, perform word segmentation and feature extraction on the text, and output text word (token) features and text global classification features. The output of the text encoder is connected to the manipulation-guided attention module.
[0079] The large model-assisted alignment module 3 is installed between the image encoder and the text encoder and includes an image description generation unit 6 and a text explanation generation unit 7. The image description generation unit 6 is used to receive image information and generate descriptive text, while the text explanation generation unit 7 is used to receive text information and generate explanatory text. Both serve as additional input information to enhance semantic alignment capabilities through comparative learning.
[0080] The manipulation-guided attention module 4 is connected to the image encoder 1, text encoder 2, and large-model assisted alignment module 3. Its core is a cross-attention matrix, which is used to fuse image block features and text word features. This module 4 is equipped with a manipulation-guided matrix. By receiving annotation information about the manipulation area, it adjusts the distribution of attention and strengthens the attention weight of the manipulation area. The fused multimodal features are divided into fused image features and fused text features, which are then input into the task head module.
[0081] The task head module 5 is set at the output end of the system and includes:
[0082] A binary classifier to determine whether multimodal input (image and text) has been manipulated;
[0083] Multiple classifiers to identify the type of manipulation;
[0084] A bounding box detector, connected to the image feature output of the manipulation-guided attention module 4, locates the manipulated area in the image based on the fused image features;
[0085] The text detector is connected to the text feature output of the manipulation-guided attention module 4 and marks the manipulation area in the text according to the fused text features.
[0086] In the control area indicator and the local modeling unit 8, the control area indicator constructs an indication guidance matrix based on the control area annotation information of the image; the local modeling unit performs binary classification on the processed image feature output end features, and uses the indication guidance matrix to assist learning to enhance the training of the bounding box detector.
[0087] These components are connected via data streams, completing image and text feature extraction, semantic alignment, attention optimization, and ultimately manipulation detection and localization. It's worth noting that the large-model assisted alignment module is only used during the training phase, so the inference phase doesn't consume significant computing power.
[0088] Depend on Figure 2 As shown, a multimodal manipulation detection and positioning method of the present invention includes the following steps:
[0089] S1, data preprocessing.
[0090] The input multimodal data (I, T) are fed into the image encoder E v and text encoder E t Perform feature extraction.
[0091] S11, image I passes through image encoder E v Perform block processing and extract visual features to obtain the image global classification feature i cls and image patch features Where N is the number of image blocks:
[0092]
[0093] S12, text T passes through the text encoder E t Perform word segmentation and extract semantic features to obtain the global classification feature t of the text cls and text word (token) features Where M is the number of text tokens (words):
[0094]
[0095] Among them, a text token refers to the basic unit in the text, which can be a word, a phrase or other language unit.
[0096] S2, large model assisted alignment.
[0097] S21, input the image I into the multimodal large model MLLM to generate descriptive text C, and pass the descriptive text C through the text encoder E t Perform word segmentation and extract semantic features to obtain the global classification feature c of the descriptive text cls and descriptive text words (tokens) features Where H is the number of descriptive text words:
[0098]
[0099] S22, input the text T into the language model LM to generate an explanatory text E, and pass the explanatory text E through the text encoder Et Perform word segmentation and extract semantic features to obtain the global classification features of the explanatory text e cls and explanatory text word (token) features Where Z is the number of explanatory text words:
[0100]
[0101] S23, construct a contrastive learning objective to enhance the alignment of image I and descriptive text C. The alignment loss of image I and descriptive text C is defined as:
[0102]
[0103] Where s(·,·) represents the feature similarity function; i cls is the global classification feature of image I; c cls is the global classification feature of the descriptive text C; τ is the temperature coefficient; B is a set of small batch samples, specifically a set composed of the global classification features of the descriptive text C of each sample (image) in the small batch sample.
[0104] S3, manipulation-guided feature fusion.
[0105] Cross-attention is used to perform multimodal fusion of image block features and text word features, and a manipulation guidance matrix is introduced to assign weights to the manipulation areas, as shown below:
[0106] S31, the calculation formula of the cross attention matrix A is:
[0107]
[0108] Among them, the image block features Text word features Represents the i-th image block feature; represents the jth word feature; W t and W v is the feature mapping matrix, W t ∈R d×d‘ , W v ∈R d×d' ; d is the feature dimension, d' is the hidden dimension. In this embodiment, d = 256, d' = 768; I w ∈R N×d , T w ∈R M×d ;A∈R N×M .
[0109] S32, feature fusion is performed based on the cross attention matrix A. The fused multimodal features are divided into image feature output and text feature output:
[0110] F I =A*T w
[0111] F T =A*I w
[0112] Among them, F I is the fused image feature, F T is the fused text feature.
[0113] S33, introduce the control guidance matrix G to give higher weight to the control area. The control guidance matrix G is defined as follows:
[0114]
[0115] Among them, G(i,j) is the element value of the i-th row and j-th column in the manipulation guidance matrix G, indicating whether the i-th image block or the j-th word is manipulated.
[0116] S33, the attention matrix loss after manipulation guidance is defined as:
[0117]
[0118] Among them, A(i,j) is the element value of the i-th row and j-th column in the cross-attention matrix A, which represents the cross-attention between the i-th image block and the j-th word.
[0119] S4, manipulation detection and positioning.
[0120] The fused multimodal features are input into the task head module to complete detection and positioning, as follows:
[0121] S41, the binary classifier determines whether the data is manipulated, and its loss is defined as:
[0122]
[0123] y bin =Sigmoid(MLP(F T ))
[0124] Among them, y bin The predicted label of whether it is manipulated is output by the binary classifier; is the real label of whether it is manipulated; F T is the fused text feature; MLP is the linear mapping function of the binary classifier; Sigmoid is the activation function.
[0125] In this embodiment, the fused text feature F T The number of query parameters is larger, and experiments have shown that the fused text features are more effective.
[0126] The indicator steering matrix P is introduced to emphasize the manipulated area in the image. The indicator steering matrix P is defined as follows:
[0127]
[0128] Among them, P i The element value of the i-th image block in the steering matrix P indicates whether the i-th image block is manipulated.
[0129] The loss after the indicated bootstrap is defined as:
[0130]
[0131] in, is a local modeling unit, and the probability of each image block being manipulated is calculated through a linear mapping function MLP of a binary classifier. N represents the sequence number of the image block, and F I is the fused image feature, Represents the fused i-th image block feature;
[0132] S42, multi-classifier identifies the manipulation type, and its loss is defined as:
[0133]
[0134] y k =Sigmoid(MLP(F T ))
[0135] Among them, y k is the predicted probability of the kth manipulation type output by the multi-classifier, is the true probability of belonging to the kth manipulation type; F T is the fused text feature; MLP is the linear mapping function of multiple classifiers, which specifically includes four manipulation types and adopts four classifiers in this embodiment; Sigmoid is the activation function.
[0136] S43, the bounding box detector locates the manipulated area in the image, using IoU regression loss:
[0137] L bbox =|b pred -b true |+L IoU (b pred ,b true )
[0138] b pred =Sigmoid(MLP(F I ))
[0139] Among them, b pred is the predicted region of the image manipulation output by the bounding box detector, b true is the real area of image manipulation; F I is the fused image feature; MLP is the linear mapping function of the bounding box detector, which outputs two coordinate positions.
[0140] S44, the text detector locates the manipulation area (word position) in the text, using cross entropy loss:
[0141]
[0142]
[0143] Among them, MLP is the linear mapping function of the text detector, which calculates the manipulation probability of each word; is the predicted manipulation probability of the jth word in the text output by the text detector, is the true manipulation probability of the jth word in the text.
[0144] S5, calculate the total model loss based on alignment loss, manipulation-guided attention matrix loss, discrimination loss, classification loss, image localization loss, text localization loss, and instruction-guided image discrimination loss.
[0145] The total loss of the model is:
[0146] L=α1L bin +α2L mul +α3L bbox +α4L tok +β1L MGCA +β2L I→C +β3L pmm
[0147] Among them, L is the total loss of the model, α1, α2, α3, α4, β1, β2, β3 are weight coefficients
[0148] S6, trains the model based on the total model loss, and the trained model is used to detect and locate the manipulation of multimodal data.
[0149] Through the above steps, the present invention achieves precise positioning of multimodal manipulation.
[0150] This embodiment significantly improves the detection accuracy and positioning capability of the control content through the carefully designed multi-modal control detection and positioning device and method. The specific effects are analyzed as follows:
[0151] The large-model assisted alignment module utilizes a multimodal large language model (VisCPM) and a language model (Mistral) to generate descriptive and explanatory text. By comparing and learning the semantic alignment of images and text, it effectively addresses the existing lack of attention to semantic alignment of manipulated content, significantly improving the model's ability to perceive multimodal manipulations. Specifically, the accuracy of authenticity discrimination exceeded 94%, and the accuracy of text manipulation localization increased by over 3% to 78.6%.
[0152] Manipulation-guided cross-attention mechanism: The introduced manipulation-guided matrix dynamically adjusts attention weights, allowing the model to focus more closely on unusual areas when processing manipulated samples, significantly improving manipulation detection accuracy. Specifically, the accuracy of manipulation recognition in image regions increased by over 1%, reaching 77%, and authenticity discrimination also saw slight improvements.
[0153] Control area indicator and local modeling unit: The local classification prior constructed based on the labeled information of the control area effectively optimizes the performance of the boundary detector, making the model more refined in locating the control area, avoiding the error accumulation caused by global feature loss, and making the model's regional control accuracy exceed 77%.
[0154] Compared with existing methods, the present invention has achieved significant performance improvement in multimodal manipulation detection tasks. Taking HAMMER and UFAFORMER as baselines, the present invention achieved a 1.29 improvement in positioning accuracy indicators, and achieved a nearly 2% mAP improvement in multi-classification task discrimination.
[0155] Large model-assisted alignment is only introduced in the training phase, and there is no need to load the large model in the inference phase, thus avoiding the waste of computing resources and ensuring high efficiency in practical applications.
[0156] Through the semantic alignment of multimodal features and refined attention to the manipulated areas, the present invention can achieve high-precision detection and positioning in complex scenarios (where face replacement and text manipulation coexist), and is suitable for the authenticity review tasks of current multimodal content.
[0157] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for multimodal manipulation detection and positioning, characterized in that: The following steps are involved: S1, performing feature extraction on input multimodal data to obtain image features and text features respectively; the multimodal data includes images and text; S2, generates descriptive text for the input image and calculates the alignment loss between the image and the descriptive text; S3, multimodal fusion of image features and text features based on the cross attention matrix to obtain fused image features and fused text features; Introduce the manipulation guidance matrix to assign weights to the manipulation area and calculate the attention matrix loss after manipulation guidance; S4, using the fused multimodal features for detection and positioning, including: determining whether the multimodal data has been manipulated, identifying the manipulation type, locating the manipulated area in the image, and locating the manipulated area in the text; Based on the detection and positioning results, the discrimination loss, classification loss, image positioning loss, and text positioning loss are calculated. The indicator guidance matrix is introduced to emphasize the manipulated area in the image, and the image discrimination loss after indicator guidance is calculated. S5, calculates the total model loss based on the alignment loss between the image and the descriptive text, the attention matrix loss after manipulation guidance, the discrimination loss, the classification loss, the image localization loss, the text localization loss, and the image discrimination loss after instruction guidance; S6, trains the model based on the total model loss, and the trained model is used to detect and locate the manipulation of multimodal data.
2. The method for multimodal manipulation detection and positioning according to claim 1, characterized in that: In step S1, the image I passes through the image encoder E v Perform block processing and extract visual features to obtain the image global classification feature i cls and image patch features Where N is the number of image blocks, Represents the i-th image block feature; The text T is passed through the text encoder E t Perform word segmentation and extract semantic features to obtain the global classification feature t of the text cls and text word features Where M is the number of words in the text, Represents the j-th word feature.
3. The method for multimodal manipulation detection and positioning according to claim 2, characterized in that: In step S2, the image I is input into the multimodal model to generate a descriptive text C, and the descriptive text C is passed through the text encoder E t Perform feature extraction to obtain the descriptive text global classification feature c cls , where H is the number of descriptive text words: Alignment loss L between image I and descriptive text C I→C for: Where s(·,·) represents the feature similarity function; τ is the coefficient; and B is the set of small batch samples.
4. The method for multimodal manipulation detection and positioning according to claim 2, characterized in that: In step S3, multimodal fusion of image features and text features is performed based on the cross attention matrix A: F I =A*T w F T =A*I w Among them, F I is the fused image feature, F T is the fused text feature; image block feature Text word features The calculation formula of the cross attention matrix A is: Among them, W t and W v is the feature mapping matrix; W t ∈R d×d‘ , W v ∈R d×d' ; d is the feature dimension, d' is the hidden dimension; d is the feature dimension; I w ∈R N×d , T w ∈R M×d ;A∈R N×M ; Introduce the control guidance matrix G: Where G(i,j) is the value of the element in the i-th row and j-th column of the manipulation guidance matrix G, indicating whether the i-th image block or the j-th word is manipulated; Attention matrix loss L after manipulation guidance MGCA for: Among them, A(i,j) is the element value of the i-th row and j-th column in the cross-attention matrix A, which represents the cross-attention between the i-th image block and the j-th word.
5. The method for multimodal manipulation detection and positioning according to claim 1, characterized in that: In step S4, determine whether the multimodal data is manipulated and calculate the discrimination loss L bin for: y bin =Sigmoid(MLP(F T )) Among them, y bin The predicted label for whether it is manipulated; is the real label of whether it is manipulated; F T is the fused text feature; MLP represents the mapping function; Sigmoid represents the activation function; Identify the manipulation type and calculate the classification loss L mul for: y k =Sigmoid(MLP(F T )) Among them, y k is the predicted probability of belonging to the k-th manipulation type, is the true probability of belonging to the kth manipulation type; Locate the manipulated area in the image and calculate the image positioning loss L bbox for: L bbox =|b pred -b true |+L IoU (b) pred ,b true ) b pred =Sigmoid(MLP(F I )) Among them, b pred is the predicted area of image manipulation, b true is the real area of image manipulation; F I is the fused image feature; Locate the manipulation area in the text and calculate the text positioning loss L tok for: in, is the predicted manipulation probability of the jth word in the text, is the true manipulation probability of the jth word in the text.
6. The method for multimodal manipulation detection and positioning according to claim 1, characterized in that: In step S4, the indicator steering matrix P is introduced to emphasize the manipulated area in the image. The indicator steering matrix P is: Among them, P i The element value of the i-th image block in the steering matrix P indicates whether the i-th image block is manipulated; Calculate the image discrimination loss L after instruction guidance pmm for: Among them, F I is the fused image feature, Used to calculate the probability of each image block being manipulated, represents the fused i-th image block feature, and N is the number of image blocks.
7. A multi-modal manipulation detection and positioning system, characterized in that: A method for multimodal manipulation detection and positioning as described in any one of claims 1 to 6 above, the system comprising: an image encoder, a text encoder, a large model auxiliary alignment module, a manipulation-guided attention module, a task head module, a manipulation area indicator, and a local modeling unit; The image encoder is used to receive the input image, perform block processing and feature extraction on the image, and output image block features and image global classification features; The text encoder is used to receive input text, perform word segmentation and feature extraction on the text, and output text word features and text global classification features; The large model-assisted alignment module is used to receive input images and generate descriptive text, construct contrastive learning objectives, and enhance the alignment of images and descriptive text; The manipulation-guided attention module is used to fuse image block features and text word features based on a cross-attention method, and to adjust the distribution of attention and strengthen the attention weight of the manipulation area by receiving the annotation information of the manipulation area; The task head module includes: a binary classifier for determining whether the multimodal data has been manipulated; a multi-classifier for identifying the type of manipulation; a bounding box detector for locating the manipulated area in the image based on the fused image features; and a text detector for locating the manipulated area in the text based on the fused text features. The control area indicator is used to construct an indication guidance matrix based on the control area annotation information of the image; The local modeling unit is used to enhance the training of bounding box detectors by assisting learning with an indicative guidance matrix.
8. An electronic device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements a multimodal manipulation detection and positioning method as described in any one of claims 1 to 6.
9. A computer program product, characterized in that It includes a computer program / instruction, which, when executed by a processor, implements a multimodal manipulation detection and positioning method as described in any one of claims 1 to 6.
10. A readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed, the method for multimodal manipulation detection and positioning according to any one of claims 1 to 6 is implemented.