Multi-modal tampering information detection method based on cooperation of large model and small model

Through the multimodal tamper information detection method that coordinates large models and small models, the fusion of prompt learning templates and cross-modal features is solved, and the problem of insufficient adaptation of subtle tampering and multi-scene in the existing technology is realized, and the precise tampering area positioning and classification of multimodal data is realized.

CN120354950APending Publication Date: 2025-07-22SHENZHEN KIM DAI INTELLIGENCE INNOVATION TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510776024.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing multimodal forgery media detection methods are difficult to identify subtle tampering and lack of multi-scene adaptation capabilities, especially when pixel-level modifications and text tampering adjustments, and lack a deep understanding of image-text emotional associations and domain knowledge.

Method used

The method of collaborating between large models and small models is adopted to design prompt learning templates for domain recognition and sentiment analysis, combined with small models to extract semantics and emotional features, and uses the contrast learning and cross-attention mechanism of manipulation perception to generate cross-modal interaction features. Finally, the decoder generates the tampered area mask and makes inference judgment.

Benefits of technology

It realizes accurate positioning and classification of tampered areas in multimodal data, improves the model's generalization ability of complex tampered scenes, and satisfies the pixel-level and word-level positioning and interpretability of tampered areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354950A_ABST
    Figure CN120354950A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal tampering information detection method based on cooperation of a large model and a small model, and the method comprises the following steps: designing a prompt learning template for a large language model, and carrying out the field recognition and emotion analysis of an input image and text through the large language model according to the prompt learning template, generating tampering field guide information and an emotion analysis result; extracting semantic features and emotional features of the image and the text by using a small model, generating cross-modal semantic interaction features and emotional interaction features through a contrast learning loss and cross attention mechanism of manipulation perception, and mining inconsistent clues of cross-modal semantics and emotions; inputting the generated cross-modal semantic interaction features and emotional interaction features into a decoder, and generating a tampering region mask of the image or the text; and inputting the tampering field guide information, the tampering region mask and the multi-modal interaction characteristics into a large language model to generate a tampering category, tampering region identification and a reasoning basis of a tampering identification process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-modal tampering information detection method based on the collaboration of a large model and a small model.

Background Art

[0002] With the rapid development of image editing technology and text generation models, the threshold for forging multimedia content has been significantly reduced. Although such technologies have promoted the progress of related industries, they have also given rise to malicious application scenarios. For example, generative text models can accurately rewrite the semantics or emotional tendencies of sentences, forming highly misleading graphic and text forgery content, which pose potential security and privacy threats to individuals and society, causing public concerns. Compared with single-modal forgery, multi-modal forged media are often more harmful due to stronger information complementarity and wider dissemination scenarios.

[0003] In response to the above problems, existing research has proposed a multimedia tampering information detection and localization method (DigitalGuardian Method, hereinafter referred to as DGM). Different from traditional false information detection that only provides binary classification results, DGM can not only give a judgment on whether an image or text has been tampered with, but also locate the tampered area of the image and identify the tampered characters in the text. Existing DGM methods include subtle artifact recognition, neural network to uncover forgery traces, frequency analysis, and learning the semantic similarity between images and texts, providing clues for locating tampered information. However, these methods face two major bottlenecks: 1. Pixel-level modifications or minor text tampering adjustments are difficult to identify through artifacts or frequency domain features; 2. Previous CLIP-based detectors usually lack effective input text prompts, which limits the adaptability of CLIP's multi-modal learning ability in detection tasks; 3. Existing models lack a deep understanding of the emotional association and domain knowledge between images and texts, and do not fully utilize the cross-modal association of multi-modal data, resulting in insufficient sensitivity to complex forgery content.

[0004] Therefore, the present invention is precisely produced based on the above deficiencies.

Summary of the Invention

[0005] The object of the present invention is to overcome the deficiencies of the prior art and provide a multi-modal tampering information detection method based on the collaboration of a large model and a small model to solve the problems of insufficient ability to detect subtle tampering and multi-scenario adaptation in the prior art.

[0006] The present invention is realized through the following technical solutions:

[0007] A multi-modal tampering information detection method based on the collaboration of a large model and a small model, characterized by including the following steps:

[0008] S1. The large language model generates multi-angle analysis and designs a prompt learning template for the large language model. The large language model performs domain recognition and sentiment analysis on the input image and text according to the prompt learning template, and generates tampering domain guidance information and sentiment analysis results;

[0009] S2, tampering clue extraction, using a small model to extract the semantic features and emotional features of the image and text, generating cross-modal semantic interaction features and emotional interaction features by manipulating the perceptual contrast learning loss and cross-attention mechanism, and mining cross-modal semantic and emotional inconsistency clues;

[0010] S3, generating a tampered region mask, inputting the generated cross-modal semantic interaction features and emotional interaction features into a decoder to generate a tampered region mask of the image or text;

[0011] S4, generation of tampering categories and reasoning judgments, inputting the tampering field guidance information, tampering area mask and multimodal interaction features into the large language model to generate tampering categories, identification of tampering areas and reasoning basis for the tampering identification process.

[0012] The multimodal tampered information detection method based on the collaboration of the large model and the small model as described above is characterized in that step S1 comprises:

[0013] S11, based on a preset prompt learning template, identifying the tampering field to which the tampering information of the image and text belongs, and determining the key parts that need to be focused on in the detection of the field;

[0014] S12. Based on the preset prompt learning template, the emotional tendency of the facial expression in the image and the emotional color of the text content are analyzed to generate the emotional analysis results.

[0015] The multimodal tampered information detection method based on the collaboration of large model and small model according to claim 1 is characterized in that step S2 comprises:

[0016] S21, multimodal feature extraction, using a small model to extract the visual semantic features of the image X v and the text semantic features X of the text t , and encode the sentiment analysis results to generate the visual sentiment feature E of the image v and the text sentiment feature E t ;

[0017] S22, contrastive learning, through contrastive learning loss function, close the semantic feature embedding and sentiment feature embedding of the untampered image-text pair, and use the tampered image-text pair as negative sample to push away its semantic and sentiment feature embedding;

[0018] S23. Multimodal feature fusion. Through a dual-branch attention mechanism, generate the image-text semantic interaction feature X v→t and the text-image semantic interaction feature X t→v and the image-text emotional interaction feature E v→t and the text-image emotional interaction feature E t→v .

[0019] The multimodal tampering information detection method based on the cooperation of a large model and a small model as described above is characterized in that: in the step S21, the small model includes a multimodal encoder for encoding image or text features and an image feature extraction network for encoding facial features, and the encoded feature X of the image v , the encoded feature X of the text t , the image emotional encoding E v , and the text emotional encoding E t are expressed as:

[0020] X v = {v cls , v pat}

[0021] X t = {t cls , t tok}

[0022] E v = {e vcls , e vpat}

[0023] E t = {e tcls , e tpat}

[0024] where e vcls represents the visual emotion category, e vpat represents the visual emotion feature, e tcls represents the text emotion category, and e tpat represents the text emotion feature.

[0025] The multimodal tampering information detection method based on the cooperation of a large model and a small model as described above is characterized in that: in the step S22, based on the InfoNCE loss, a contrastive loss from image to text is proposed, and the calculation formula is:

[0026]

[0027] where τ represents the temperature parameter, T - represents the negative text, T + represents the positive text, S() represents the similarity calculation formula, -E pDenote the contrastive learning loss as, and denote the number of samples as K;

[0028] The contrastive learning loss from text to image is calculated as follows:

[0029]

[0030] where I - denotes the negative sample image, and I + denotes the positive sample image;

[0031] The contrastive learning loss between text sentiment and image sentiment is calculated as follows:

[0032]

[0033] where E v + denotes the positive sample image sentiment, and E v - denotes the negative sample image sentiment;

[0034] The contrastive learning loss between image sentiment and text sentiment is calculated as follows:

[0035]

[0036] where E t + denotes the positive sample, and E t - denotes the negative sample.

[0037] The multimodal tampering information detection method based on the cooperation of a large model and a small model as described above is characterized in that: in the step S23, the dual-branch attention mechanism processes semantic features and emotional features respectively, and calculates the attention function through the normalized query Q, key K, and value V features. The calculation formula is:

[0038]

[0039] The calculation formula for the image-text interaction feature is:

[0040] X v→t = Attention(X t , X v , X v )

[0041] The calculation formula for the image sentiment and text sentiment interaction feature is:

[0042] E v→t = Attention(E t , E v , E v )

[0043] The calculation formula for the text-image interaction feature is as follows:

[0044] X t→v = Attention(X v , X t , X t )

[0045] The calculation formula for the text sentiment and image sentiment interaction feature is as follows:

[0046] E t→v = Attention(E v , E t , E t )

[0047] where d represents the feature dimension.

[0048] The multi-modal tampering information detection method based on the cooperation of the large model and the small model as described above is characterized in that: in step S3, the generated image-text semantic interaction feature X v→t , the text-image semantic interaction feature X t→v , the image-text sentiment interaction feature E v→t , and the text-image sentiment interaction feature E t→v are concatenated and input into the decoder, and the prediction result b of the tampering area is output. The calculation formula is: b = D v ([X v→t , E v→t )

[0049] The multi-modal tampering information detection method based on the cooperation of the large model and the small model as described above is characterized in that: the large language model is the GPT 4V large language model.

[0050] The multi-modal tampering information detection method based on the cooperation of the large model and the small model as described above is characterized in that: the multi-modal encoder is the ImageBind encoder, and the image feature extraction network is the Resnet50 network.

[0051] The multi-modal tampering information detection method based on the cooperation of the large model and the small model as described above is characterized in that: the decoder is the BBox decoder.

[0052] Compared with the prior art, the present invention has the following advantages:

[0053] 1. The domain recognition and sentiment analysis of the large model of the present invention provide a detection direction for the small model, and the fine-grained feature extraction of the small model makes up for the efficiency defect of the large model, forming a hierarchical detection framework. Moreover, through cross-modal feature fusion, contrast learning, and prompt guidance, the subtle inconsistencies between cross-modal semantics and sentiment are mined, enhancing the generalization ability of the model for complex tampering scenarios, meeting the pixel-level and word-level positioning of the tampered area, and realizing the accurate positioning of the tampered area in multi-modal data, the classification of tampering categories, and the interpretability of the reasoning process.

Description of the Drawings

[0054] Figure 1 It is a schematic diagram of the present invention.

Detailed Embodiments

[0055] The present invention will be further described below with reference to the drawings:

[0056] As Figure 1 shown, a multi-modal tampering information detection method based on the cooperation of a large model and a small model includes the following steps:

[0057] S1. The large language model generates multi-angle analysis. A prompt learning template is designed for the large language model. The large language model performs domain recognition and sentiment analysis on the input images and texts according to the prompt learning template, and generates tampering domain guidance information and sentiment analysis results.

[0058] S2. Tampering clue extraction. The small model is used to extract the semantic features and sentiment features of the images and texts, and through the contrast learning loss of manipulation perception and the cross-attention mechanism, cross-modal semantic interaction features and sentiment interaction features are generated, and the inconsistent clues between cross-modal semantics and sentiment are mined.

[0059] S3. Tampering area mask generation. The generated cross-modal semantic interaction features and sentiment interaction features are input into the decoder to generate a tampering area mask for the image or text.

[0060] S4. Generation of tampering category and reasoning judgment. The tampering domain guidance information, the tampering area mask, and the multi-modal interaction features are input into the large language model to generate the tampering category, the identification of the tampering area, and the reasoning basis for the tampering identification process.

[0061] Wherein step S1 includes: S11, based on the preset prompt learning template, identifying the tampering field to which the tampering information of the image and text belongs, and determining the key parts that need to be focused on in the detection of this field; S12, based on the preset prompt learning template, analyzing the emotional tendency of facial expressions in the image and the emotional color of the text content, and generating emotional analysis results. Specifically, the large language model is the GPT 4V large language model, and of course other large language models that support multimodal input can also be selected. In this step, GPT 4V is first used to identify the field involved in the tampering information according to the prompt learning template, and provide guidance on which aspects of the image or text information should be focused on in the tampering field. The template for prompt learning in step S11 is: "Pleasedetermine which domain the provided information belongs to and explain whichaspects should be focused on when assessing the manipulation of informationin that domain.", and obtain the domain information domain. Secondly, GPT 4V is used to analyze the sentiment of text and image information to provide guidance for detecting the tampering of facial attributes or text attribute categories, such as detecting whether a face with a positive expression is replaced with a face with a negative expression, or whether the emotional color of the text is tampered with. In step S12, the learning template is prompted as: "You are an expert in sentiment analysis. Please analyze which sentiment the images and text are belong to."

[0062] Step S2 includes: S21, multimodal feature extraction, using a small model to extract the visual semantic features X of the image v and the text semantic features X of the text t , and encode the sentiment analysis results to generate the visual sentiment feature E of the image v and the text sentiment feature E t ;

[0063] S22, contrastive learning, through contrastive learning loss function, close the semantic feature embedding and sentiment feature embedding of the untampered image-text pair, and use the tampered image-text pair as negative sample to push away its semantic and sentiment feature embedding;

[0064] S23, multimodal feature fusion, through the dual-branch attention mechanism, to generate image-text semantic interaction features X v→t , text-image semantic interaction feature Xt→v 、Image - text sentiment interaction feature E v→t and text - image sentiment interaction feature E t→v 。

[0065] The small model in step S21 includes a multi - modal encoder for encoding image or text features and an image feature extraction network for encoding facial features. Specifically, the multi - modal encoder is an ImageBind encoder, and the image feature extraction network is a Resnet50 network. The encoded feature X of the image v 、the encoded feature X of the text t 、the image sentiment encoding E v 、the text sentiment encoding E t are expressed as:

[0066] X v ={v cls ,v pat}

[0067] X t ={t cls ,t tok}

[0068] E v ={e vcls ,e vpat}

[0069] E t ={e tcls ,e tpat}

[0070] where e vcls represents the visual sentiment category, e vpat represents the visual sentiment feature, e tcls represents the text sentiment category, and e tpat represents the text sentiment feature.

[0071] To assist in the association of semantic features and emotions between the image and text modalities, the present invention designs a contrastive learning loss for semantic features and a contrastive learning loss for emotional features to align the semantic features and emotional features of the two modalities. To emphasize the semantic inconsistencies caused by these manipulations, in step S22, manipulation-aware contrastive learning is proposed for image and text embeddings. Different from ordinary cross-modal contrastive learning, which only pushes away the embeddings of unmatched pairs when pulling closer the embeddings of the original image-text pairs, the present invention, through a manipulation-aware contrastive learning loss function, takes the unmanipulated image-text pairs as positive samples during the training process, pulls closer the embedding distances of their visual semantic features and text semantic features, and takes the manipulated image-text pairs as negative samples, pushing away the embedding distances of their visual semantic features and text semantic features. At the same time, it aligns the image emotional features and text emotional features of the unmanipulated pairs and pushes away the emotional feature embeddings of the manipulated pairs, thereby further emphasizing the semantic inconsistencies they generate. Based on the InfoNCE loss, a contrastive loss from image to text is proposed, and the calculation formula is:

[0072]

[0073] where τ represents the temperature parameter, T - represents the negative text, T + represents the positive text, S() represents the similarity calculation formula, -E p represents the contrastive learning loss, and K represents the number of samples;

[0074] The contrastive learning loss from text to image has the calculation formula:

[0075]

[0076] where I - represents the negative sample image, I + represents the positive sample image;

[0077] The contrastive learning loss of text emotion - image emotion has the calculation formula:

[0078]

[0079] where E v + represents the positive sample image emotion, E v - represents the negative sample image emotion;

[0080] The contrastive learning loss of image emotion - text emotion has the calculation formula:

[0081]

[0082] where E t+ Indicates a positive sample, E t - Indicates a negative sample.

[0083] Since the replacement and the tampering of the attribute category will change the association between the image and the corresponding text (such as a person's name or sentiment), the present invention locates the manipulated image area by finding the local area that is inconsistent with the text embedding and the text sentiment, or locates the manipulated text area by finding the local area that is inconsistent with the image embedding and the image sentiment.

[0084] The lack of complementary information between modalities may hinder cross-modal semantic reasoning. To this end, the present invention designs a dual-branch cross-attention mechanism to guide the interaction between image and text features, image sentiment and text sentiment, so that semantic associations can be extracted. In step S23, the dual-branch attention mechanism processes semantic features and sentiment features respectively, and calculates the attention function through the normalized query Q, key K, and value V features. The calculation formula is:

[0085]

[0086] The calculation formula for the image-text interaction feature is:

[0087] X v→t = Attention(X t , X v , X v )

[0088] The calculation formula for the image sentiment and text sentiment interaction feature is:

[0089] E v→t = Attention(E t , E v , E v )

[0090] The calculation formula for the text-image interaction feature is:

[0091] X t→v = Attention(X v , X t , X t )

[0092] The calculation formula for the text sentiment and image sentiment interaction feature is:

[0093] E t→v = Attention(E v , E t , E t )

[0094] Where d represents the feature dimension.

[0095] In step S3, the generated image-text semantic interaction feature X v→t , text-image semantic interaction feature X t→v , image-text emotion interaction feature E v→t and text-image emotion interaction feature E t→v are concatenated and input into the decoder to output the prediction result b of the tampered area. The calculation formula is: b = D v ([X v→t , E v→t ). Specifically, the decoder is a BBox decoder (D v ), which is composed of two layers of MLP perceptrons.

[0096] In step S4, a prompt learning template also needs to be designed for the multimodal large language model to guide the multimodal large language model to complete the following tasks: 1. Analyze the tampering traces of images and texts: Through the Chain-of-Thought (hereinafter referred to as CoT) step, detect image tampering (such as face swapping / attribute tampering) and text tampering (such as keyword replacement) in sequence; 2. Generate a structured output: Require the model to output a JSON object containing labels (real / fake), tampering types, tampering areas, and tampering words to ensure the standardization of the results. The modified prompt learning template is:

[0097] ##Human: <semantictoken> <forgerytoken> <domain>You are a professional assistant specialized in detecting fake news.

[0098] ##Chain-of-Thought(CoT)Reasoning Process:

[0099] 1. First, analyze the image to determine if the face has been manipulated. If manipulated, record the type and classify the news as fake.

[0100] 2. Second, analyze the text for any manipulations. If manipulated, record the type and manipulated words, and classify the news as fake.

[0101] 3. If neither the image nor the text is manipulated, classify the news as real.

[0102] ##Rules:

[0103] Generate a JSON object with four properties: ‘Label’, ‘Image Manipulate’, ‘Text Manipulate’, and ‘Manipulated Words’.

[0104] ##Your Task:

[0105] Given a piece of Input Text and Image, your task is to analyze both the image and text to determine the authenticity of the news and provide a JSON-formatted result.

[0106] Input Text: {TEXT}

[0107] ##Your Response:

[0108] Let’s think step by step according to the above Chain-of-Thought,yourresponse:

[0109] The input of SemanticToken is X v→t ,E v→t ,the input of ForgeryToken is the predicted tampered area b, and the input of domain is the domain information.

[0110] Generally speaking, the present invention proposes a multi-angle tampered information detection method for the cooperation of large models and small models. This method first allows the LLM to generate analyses of text and images from the perspectives of sentiment and information domain. Then, the ImageBind encoder is used to encode the images, text, and the analyses produced by the LLM, and the cross-attention mechanism is used to learn the connections between these features, and a decoder is used to predict the mask of the tampered area. Finally, the predicted mask and the features processed by the encoding of the small model are input into the LLM to complete the identification and interpretation of the tampering type.< / domain> < / forgerytoken> < / semantictoken>

Claims

1. A multi-modal tampering information detection method based on the collaboration of large models and small models, characterized in that, Including the following steps: S1. The large language model generates multi - angle analysis, designs a prompt learning template for the large language model. The large language model performs domain recognition and sentiment analysis on the input images and texts according to the prompt learning template, and generates tampering domain guidance information and sentiment analysis results; S2. Tampering clue extraction, using a small model to extract the semantic features and sentiment features of the images and texts, generating cross - modal semantic interaction features and sentiment interaction features through contrastive learning loss of manipulation perception and cross - attention mechanism, and mining the inconsistencies of cross - modal semantics and sentiment; S3. Tampering region mask generation, inputting the generated cross - modal semantic interaction features and sentiment interaction features into a decoder to generate a tampering region mask for the image or text; S4. Generation of tampering category and inference judgment, inputting the tampering domain guidance information, tampering region mask and multi - modal interaction features into the large language model to generate the tampering category, the recognition of the tampering region and the inference basis for the tampering recognition process.

2. The multimodal tampering information detection method based on the collaboration of a large model and a small model according to claim 1, characterized in that The step S1 includes: S11. Based on a preset prompt learning template, identify the tampering domain to which the tampering information of the images and texts belongs, and determine the key parts that need to be concerned about in the detection of this domain; S12. Based on a preset prompt learning template, analyze the emotional tendency of the facial expressions in the image and the emotional color of the text content to generate a sentiment analysis result.

3. The multi-modal tampering information detection method based on the cooperation of a large model and a small model according to claim 1, wherein The step S2 includes: S21. Multi-modal feature extraction, using a small model to extract the visual semantic feature X of the image v and the text semantic feature X of the text t , and encoding the sentiment analysis result to generate the visual sentiment feature E of the image v and the text sentiment feature E of the text t ; S22. Contrastive learning, through a contrastive learning loss function, bringing closer the semantic feature embeddings and sentiment feature embeddings of non - tampered image - text pairs, and at the same time pushing away the semantic and sentiment feature embeddings of the tampered image - text pairs as negative samples; S23. Multi-modal feature fusion. Through a dual-branch attention mechanism, generate the image-text semantic interaction feature X v→t and the text-image semantic interaction feature X t→v and the image-text emotion interaction feature E v→t and the text-image emotion interaction feature E t→v .

4. The multimodal tampering information detection method based on the cooperation of a large model and a small model according to claim 3, wherein: In the step S21, the small model includes a multi-modal encoder for encoding image or text features and an image feature extraction network for encoding facial features, and the encoded feature X of the image v , the encoded feature X of the text t , the image emotion encoding E v , the text emotion encoding E t The expressions are as follows: X v = {v cls , v pat} X t = {t cls , t tok} E v = {e vcls , e vpat} E t = {e tcls , e tpat} where e vcls represents the visual emotion category, e vpat represents the visual emotion feature, e tcls represents the text emotion category, e tpat represents the text emotion feature.

5. The multimodal tampering information detection method based on the collaboration of a large model and a small model according to claim 3, wherein: In the step S22, according to the InfoNCE loss, the contrastive loss from image to text is proposed, and the calculation formula is: where τ represents the temperature parameter, T - represents the negative text, T + represents the positive text, S() represents the similarity calculation formula, -E p represents the contrastive learning loss, and K represents the number of samples; The contrastive learning loss from text to image, the calculation formula is: where \(I\) - represents a negative sample image, and \(I\) + represents a positive sample image; The contrastive learning loss of text sentiment - image sentiment, the calculation formula is: Among them, E v + represents the emotion of the positive sample image, E v - represents the emotion of the negative sample image; The contrastive learning loss of image sentiment - text sentiment, the calculation formula is: Among them, E t + represents the positive sample, and E t - represents the negative sample.

6. The multimodal tampering information detection method based on the cooperation of a large model and a small model according to claim 3, wherein: In the step S23, the dual - branch attention mechanism processes semantic features and sentiment features respectively, and calculates the attention function through normalized query Q, key K, and value V features. The calculation formula is: The calculation formula for image - text interaction features is: X v→t = Attention(X t , X v , X v ) The calculation formula for image sentiment and text sentiment interaction features is: E v→t = Attention(E t , E v , E v ) The calculation formula for text - image interaction features is: X t→v = Attention(X v , X t , X t ) The calculation formula for text sentiment and image sentiment interaction features is: E t→v = Attention(E v , E t , E t ) Where d represents the feature dimension.

7. The multimodal tampering information detection method based on the cooperation of large models and small models according to claim 3, wherein: In the step S3, the generated image-text semantic interaction feature X v→t , text-image semantic interaction feature X t→v , image-text emotion interaction feature E v→t and text-image emotion interaction feature E t-v are concatenated and input into the decoder to output the prediction result b of the tampered area. The calculation formula is: b = D v ([X v→t , E v→t ).

8. The multimodal tampering information detection method based on the collaboration of large models and small models according to claim 1, characterized in that: The large language model is the GPT 4V large language model.

9. The multimodal tampering information detection method based on the collaboration of a large model and a small model according to claim 4, characterized in that: The multi - modal encoder is the ImageBind encoder, and the image feature extraction network is the Resnet50 network.

10. The multimodal tampering information detection method based on the cooperation of a large model and a small model according to claim 7, characterized in that: The decoder is the BBox decoder.

Citation Information

Cited By

  • Image forgery detection method based on multi-modal large language model

    CN121236571A

  • Power generation equipment maintenance management method based on RCM and large and small model cooperation

    CN121504431A