VisionTransform image tampering positioning method and system based on multi-modal prompt guidance

The VisionTransformer method, guided by multimodal prompts, combines visual and textual features to solve the problem that complex tampering is difficult to identify with a single visual feature. It achieves high-precision localization of semantically consistent tampering and is suitable for real-time application scenarios such as social media and judicial evidence collection.

CN120976556APending Publication Date: 2025-11-18DALIAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511225510.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing image tampering localization technologies rely on single visual features, making it difficult to identify semantically consistent tampering in complex tampering scenarios. Furthermore, insufficient fusion of multimodal features results in weak model generalization ability and severe localization bias.

Method used

The VisionTransformer method, guided by multimodal prompts, generates text prompts through an LLaMA model. It combines VisionTransformer and a pre-trained language model to extract visual and textual features, uses cross-modal self-attention and cross-attention mechanisms for feature fusion, performs multi-scale processing through a spatial feature pyramid network, and finally outputs the tampered region location through a multilayer perceptron.

Benefits of technology

It significantly improves the ability to identify semantically consistent tampering such as deepfakes, enhances the accuracy and robustness of tampered area location, and adapts to the real-time image verification needs in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976556A_ABST
    Figure CN120976556A_ABST
Patent Text Reader

Abstract

The invention discloses a VisionTransform image tampering positioning method and a VisionTransform image tampering positioning system based on multi-modal prompt guidance, and relates to the cross technical field of computer vision, image tampering positioning and natural language processing. Generating a text prompt related to a to-be-detected image by using an LLaMA model, and respectively obtaining feature representations of the image and the text through a visual feature extractor and a text feature extractor; then, deep fusion and alignment of cross-modal features are realized through a multi-modal interaction prompt module; and finally, outputting an accurate tampering region positioning result in combination with the spatial feature pyramid network and the multi-layer sensor. According to the method, deep alignment of visual features and text semantics is realized through a cross-modal self-attention and cross-attention mechanism, and semantic association understanding of a model on a tampered region is remarkably improved; and meanwhile, in combination with a spatial feature pyramid network and a lightweight SegFormer decoder, the capturing capability of a multi-scale tampered region is effectively enhanced, and the performance is more excellent especially in a small tampering and large-region forging scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary fields of computer vision, image tampering localization, and natural language processing, specifically to a VisionTransformer image tampering localization method and system based on multimodal prompting guidance. Background Technology

[0002] In the digital information age, images, as a core carrier of information transmission, have been deeply integrated into key areas such as social media interaction, news dissemination and reporting, and judicial evidence preservation. However, with the rapid iteration of deep generative models (such as Generative Adversarial Networks (GANs)) and image editing tools (such as Photoshop and AI face-swapping software), image tampering technology has become increasingly sophisticated, with more covert methods and more realistic effects, posing a serious challenge to the social trust system and information security.

[0003] The problems caused by image tampering currently manifest in three main aspects: First, the spread of false information. Tampered images are used to create rumors and distort facts, causing public opinion chaos in a short period of time on social media, such as forging images of celebrities' remarks or tampering with photos of news events. Second, the credibility of evidence is damaged. If image evidence relied upon in fields such as the judiciary and security is tampered with, it may lead to misjudgments or missed judgments, affecting judicial fairness. Third, a crisis of social trust. Public trust in image information continues to decline, seriously undermining the foundation of the authenticity of information dissemination. Traditional tampering methods (such as copy-move, splicing, and removal) have gradually been upgraded. Deepfake technology combined with generative adversarial networks can generate more realistic content (such as AI face-swapping video frames and generating images of non-existent scenes), further increasing the difficulty of locating tampering.

[0004] Existing image tampering localization technologies have significant limitations: early methods based on statistical features and signal processing rely on manually designed features (such as pixel grayscale statistics and JPEG compression traces), which lack robustness in complex tampering scenarios (such as tampering under lighting adjustments and noise superposition); while mainstream deep learning methods have improved feature extraction capabilities through convolutional neural networks (CNNs), they mostly focus on visual features and neglect semantic information, making it difficult to deal with semantically consistent tampering types such as deepfakes; for example, fake faces generated by AI are semantically consistent with the original image background, making it difficult to distinguish between real and fake based solely on visual features. Single-modal analysis cannot cover diverse tampering methods, resulting in weak model generalization ability and frequent localization biases in practical applications (such as missing small tampered areas and misjudging normal areas as tampered areas).

[0005] Therefore, how to integrate multimodal information to overcome the bottleneck of single-modal technology and improve the accuracy and robustness of image tampering localization in complex scenarios has become a key issue that urgently needs to be addressed. Summary of the Invention

[0006] The purpose of this invention is to propose a VisionTransformer image tampering localization method and system based on multimodal prompts, which overcomes the shortcomings of existing technologies in image tampering localization, such as reliance on single visual features, weak semantic understanding ability, and insufficient fusion of multimodal features.

[0007] According to a first aspect of the present disclosure, a VisionTransformer image tampering localization method based on multimodal prompting guidance is provided, comprising the following steps:

[0008] The image to be detected is segmented into blocks to generate an image token; a text prompt related to the image content is generated using the LLaMA model and converted into a text token.

[0009] The image token is encoded using the VisionTransformer model to extract deep visual features of the image; the text token is encoded using a pre-trained language model to obtain semantic features of the text.

[0010] Deep visual features of images and semantic features of text are input into the multimodal interactive prompt module. Feature fusion is achieved through cross-modal self-attention and cross-attention mechanisms to generate multimodal joint features.

[0011] By utilizing a spatial feature pyramid network to perform multi-scale processing on multimodal joint features, the ability to capture tampered regions of different sizes is enhanced.

[0012] The processed features are classified by a multilayer perceptron, and the location mask of the tampered area is output to complete the image tampering location.

[0013] In one embodiment, the image to be detected is input into the LLaMA model, along with the guiding question "Please analyze the possible tampered areas and features in the image". The model generates a text prompt containing a description of the image content and a prediction of potential tampered areas through visual-language association reasoning. The output text sequence is then converted into a text token after word segmentation.

[0014] In one embodiment, the processing procedure of the multimodal interaction prompt module is as follows:

[0015] The deep visual features of the input image are first extracted by global average pooling, and then visual features of different subspaces are extracted by scale dot product method.

[0016] The semantic features of the input text are first projected using a 1×1 convolution to compress information redundancy, and then the text features of different subspaces are extracted using the scale dot product method.

[0017] In the formula, Represents the semantic features of the input text;

[0018] The processed visual features and text features are input into the cross-attention module, and deep fusion is achieved by combining residual connections to output multimodal joint features.

[0019] In one embodiment, visual features of different subspaces are extracted using the scale dot product method, the process of which is represented as follows:

[0020]

[0021] P SDP (Q,K,V)=Concat(Y1,…,Y8)W H

[0022] B j =P SDP (F u ,F u ,F u )

[0023] In the formula, Q represents the query vector, K represents the key vector, V represents the value vector, and W represents the value vector. i Q W i K W i V and W H All are trainable matrices, where d represents the dimension, φ(·) represents the softmax function, and P SDP Representing the SDP block, it is the module that performs scale dot product attention calculations; B j This represents the visual features obtained after processing by the SDP block.

[0024] Text features in different subspaces are extracted using the scale dot product method. The process is represented as follows:

[0025]

[0026] T SDP (Q,K,V)=Concat(X1,…,X8)W H

[0027] C n =T SDP (F l ,F l ,F l )

[0028] In the formula, T represents the semantic features of the input text, where m represents the modality identifier and n-1 represents the hierarchical information of the features; SDP T represents the text featuresSDP Block; C n This represents the text features obtained after processing by the SDP block.

[0029] In one embodiment, within the cross-attention module, visual features serve as the query vector, and text features serve as the key and value vectors. Through the attention weights between visual and text features, deep cross-modal information alignment and fusion are achieved. This process is described as follows:

[0030] F i n =P SDG (C n B j B j )+F i n-1

[0031]

[0032] In the formula, F i n This represents the fused features related to visual features after processing by the cross-attention module. This represents the fused features related to the text features after processing by the cross-attention module.

[0033] In one embodiment, the spatial feature pyramid network's processing procedure is as follows: multimodal joint features are downsampled using convolutional operations to generate low-resolution, high-semantic features; high-resolution, low-semantic features are upsampled using deconvolutional operations; the generated multi-scale features are processed using a lightweight SegFormer decoder, first unifying features of different scales to the same resolution, then concatenating and fusing them; finally, a tampering probability map is output through a linear layer; simultaneously, dilation and erosion operations are used to generate edge masks. Dilation expands edge regions to capture subtle tampering traces, while erosion removes noise to accurately define the contours of the tampered region.

[0034] In one embodiment, the loss function combines segmentation loss and edge detection loss, as follows:

[0035] L = L seg +λL edge

[0036] In the formula, L seg and L edge Both are binary cross-entropy loss functions, where λ is a hyperparameter used to balance the effects of segmentation loss and edge detection loss.

[0037] According to a second aspect of the present disclosure, a VisionTransformer image tampering localization system based on multimodal prompting guidance is provided, comprising:

[0038] The image segmentation and token generation module segments the image to be detected into blocks and generates an image token; it uses an LLaMA model to generate text prompts related to the image content and converts them into text tokens.

[0039] The single-modal feature encoding module encodes image tokens using the VisionTransformer model to extract deep visual features of the image; it also uses a pre-trained language model to encode text tokens to obtain text semantic features.

[0040] The multimodal feature fusion module inputs deep visual features of the image and semantic features of the text into the multimodal interactive prompt module, and achieves feature fusion through cross-modal self-attention and cross-attention mechanisms to generate multimodal joint features;

[0041] The multi-scale feature processing module utilizes a spatial feature pyramid network to perform multi-scale processing on multimodal joint features, thereby enhancing the ability to capture tampered regions of different sizes.

[0042] The tampered area localization output module classifies the processed features using a multilayer perceptron and outputs a localization mask for the tampered area, thus completing the image tampering localization.

[0043] According to a third aspect of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the memory. When the processor executes the program, it implements the VisionTransformer image tampering localization method based on multimodal prompting guidance.

[0044] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the VisionTransformer image tampering localization method based on multimodal prompting guidance.

[0045] Compared with existing technologies, the above-mentioned technical solutions adopted in this invention have the following advantages: 1. This invention proposes a multimodal prompt generation mechanism based on LLaMA, which can transform the visual content of the image to be detected into semantic text containing image content description and potential tampering region inference. This semantic text can provide high-level semantic guidance for subsequent tampering localization, effectively making up for the shortcomings of traditional single visual features that can only capture pixel-level details of images and cannot deeply understand image semantics. In particular, its adaptability to semantically consistent tampering scenarios such as deepfakes is significantly improved, breaking through the limitations of single visual modalities in complex tampering identification.

[0046] 2. By designing a multimodal interactive prompt module, a bidirectional interaction approach of self-attention and cross-attention is adopted. First, self-attention processing is applied to visual features and textual features separately to extract the core information of each modality. Then, cross-attention is used to achieve deep alignment and fusion of the two types of features. This fusion method allows visual features and textual semantic features to fully complement each other, significantly enhancing the model's ability to identify semantically consistent tampering (such as the natural splicing of AI-generated fake content with the original image background). This avoids localization bias caused by insufficient single-modal feature information and significantly improves the accuracy of tampering area localization.

[0047] 3. This invention combines the Spatial Feature Pyramid Network (SFPN) with a lightweight SegFormer decoder. SFPN enables multi-scale processing of multimodal joint features, enhancing the ability to capture tampered regions of varying sizes (from minute traces of tampering to large-scale forged areas). The lightweight SegFormer decoder effectively reduces model computational complexity while maintaining feature processing effectiveness. This combination significantly improves inference efficiency while maintaining high localization accuracy, making it suitable for real-time image verification scenarios that demand high processing speed, such as real-time social media review and on-site judicial evidence verification, thus expanding the practical application scope of the technology. Attached Figure Description

[0048] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an undue limitation of this application.

[0049] Figure 1 This is a framework diagram of the VisionTransformer image tampering localization method based on multimodal prompts;

[0050] Figure 2 Detailed design drawings for the multimodal interaction prompt module;

[0051] Figure 3 A comparison diagram showing the traditional method and the method proposed in this invention;

[0052] Figure 4 Image localization results from multiple datasets;

[0053] Figure 5 A diagram showing different variant versions of the model;

[0054] Figure 6 This is a robustness analysis diagram of an image subjected to noise attacks. Detailed Implementation

[0055] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.

[0056] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0057] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0058] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and systems according to various embodiments of this disclosure. It should be noted that each block in a flowchart or block diagram may represent a module, segment, or portion of code, which may include one or more executable instructions for implementing the logical functions specified in the various embodiments. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, may be implemented using a dedicated hardware-based system that performs the specified functions or operations, or using a combination of dedicated hardware and computer instructions.

[0059] Example 1:

[0060] like Figure 1 As shown, this embodiment provides a VisionTransformer image tampering localization method based on multimodal prompts, including the following steps:

[0061] S1. Divide the image to be detected into blocks to generate an image token; use the LLaMA model to generate text prompts related to the image content and convert them into text tokens;

[0062] Specifically, to verify the potential feasibility of using the LLaMA model to achieve multimodal image tamper localization (M-IML), the original image I... ori With guiding question T queThe input is given to the LLaMA model, with a guiding question set as "Please analyze the possible tampered regions and features in the image...". The LLaMA model can generate a detailed text description T based on the input image content. ans This textual description captures the deep semantic information contained within the visual content, providing semantic support for subsequent tampering detection. This process can be represented as:

[0063] T ans =LLaMA(I ori ,T que )

[0064] Through the above process, LLaMA is able to integrate visual information with language cues, and the resulting text description T ans It not only covers the inference of image tampering areas, but also provides high-level semantic support for subsequent multimodal interaction and tampering localization work.

[0065] S2. Encode the image token using the VisionTransformer model to extract deep visual features of the image; encode the text token using a pre-trained language model to obtain text semantic features;

[0066] Specifically, image tokens are generated from the image to be detected through 16×16 non-overlapping blocks and linear projection, and then input into the VisionTransformer (ViT) model for encoding. ViT captures global spatial relationships (such as the position of different regions and texture consistency) and local details (such as pixel transitions and edge features) between image tokens through the self-attention mechanism of multi-layer Transformer blocks. The final extracted "deep visual features of the image" contain not only pixel-level visual information, but also integrate the structural semantics of the image content, providing underlying feature support for subsequent identification of visual anomalies in tampered areas (such as splicing marks and unnatural edges).

[0067] The text tokens originate from text prompts (including image content descriptions and inferences about potential tampering regions) generated by the LLaMA model and are obtained through word segmentation. When encoding the text tokens using a pre-trained language model, the model captures the semantic relationships between text tokens through a self-attention mechanism (such as the logical correspondence between "tampered region" and "abnormal edge transition"), transforming natural language information into machine-recognizable "text semantic features." These features can extract high-level semantic guidance information, which can be used to focus on potential tampering regions in subsequent multimodal fusion, compensating for the inadequacy of single visual features in recognizing semantically consistent tampering (such as deepfakes).

[0068] S3. Input the deep visual features of the image and the semantic features of the text into the multimodal interactive prompt module, such as... Figure 2As shown, feature fusion is achieved through cross-modal self-attention and cross-attention mechanisms to generate multimodal joint features;

[0069] Specifically, the multimodal interactive prompting module enables targeted information exchange between visual and textual features; this process can be written as:

[0070]

[0071] For the visual feature branch, global average pooling is first performed on the deep visual features of the input image to extract the overall semantic features. This process can be written as:

[0072] F u =Avgpool(F i n-1 )

[0073] The process of extracting visual features from different subspaces using the scale dot product method can be written as follows:

[0074]

[0075] P SDP (Q,K,V)=Concat(Y1,…,Y8)W H

[0076] B j =P SDP (F u ,F u ,F u )

[0077] In the formula, Q represents the query vector, K represents the key vector, V represents the value vector, and W represents the value vector. i Q W i K W i V and W H All are trainable matrices, where d represents the dimension, φ(·) represents the softmax function, and P SDP Representing the SDP block, it is the module that performs scale dot product attention calculations; B j This represents the visual features obtained after processing by the SDP block.

[0078] For the text feature branch, the input text semantic features are first subjected to 1×1 convolution for dimensional projection to compress information redundancy, as shown below:

[0079]

[0080] Similarly, by extracting text features from different subspaces using the scale dot product method, this process can be written as:

[0081]

[0082] T SDP (Q,K,V)=Concat(X1,…,X8)W H

[0083] C n =T SDP (F l ,F l ,F l )

[0084] In the formula, T represents the semantic features of the input text, where m represents the modality identifier and n-1 represents the hierarchical information of the features; SDP T represents the text features SDP Block; C n This represents the text features obtained after processing by the SDP block.

[0085] In the multimodal interaction phase, visual and textual features are input into the cross-attention module for deep cross-modal information alignment and fusion. This process can be written as:

[0086] F i n =P SDG (C n B j B j )+F i n-1

[0087]

[0088] This module employs a dual-path input branch architecture, independently processing and deeply interacting with visual and textual features respectively, thereby fully exploring the potential synergistic effects between multimodal information. Specifically, visual and textual features first extract core information from their respective modalities through a self-attention mechanism, and then achieve information interaction between multimodalities using cross-attention. Simultaneously, residual connections are introduced to preserve and fuse the original features and the interacted features. The final output features possess both modal independence and cross-modal interaction information, effectively improving multimodal learning performance. Through this design, the multimodal interaction prompt module can more accurately capture and integrate key information from different modalities, providing richer feature representations for subsequent more precise tamper region localization.

[0089] After the above processing, the final output visual and text features are significantly improved in terms of semantic expression and spatial representation capabilities: they retain the specific advantages of each modality (such as the spatial details of visual features and the semantic guidance of text features) and achieve efficient feature complementarity; they can provide rich and robust feature support for subsequent tampering localization, and can effectively improve the model's generalization ability and localization accuracy in complex tampering scenarios (such as deepfakes and multi-region splicing), fully demonstrating the strong adaptability and excellent performance of the multimodal interactive prompt module.

[0090] S4. Utilize a spatial feature pyramid network to perform multi-scale processing on multimodal joint features, thereby enhancing the ability to capture tampered regions of different sizes.

[0091] Specifically, multi-scale feature maps are generated through convolution and deconvolution operations, achieving multi-scale representation of features while preserving the integrity of the model structure. A lightweight SegFormer decoder is used to process the multi-scale features, first unifying the resolution of features at different scales, then integrating multi-scale information through concatenation and fusion, and finally outputting a tampering probability map through a linear layer. This architecture effectively enhances the localization accuracy and robustness of tampered regions of different sizes while maintaining high computational efficiency, making it suitable for real-time application scenarios. Furthermore, edge masks are generated through dilation and erosion operations, specifically focusing on the boundaries of tampered regions: dilation expands the edge region, helping the model more accurately capture subtle tampering traces; erosion removes potential noise and precisely defines the contours of the tampered region. Through these methods, clearer and more accurate edge information can be obtained, further improving the accuracy of tampering localization.

[0092] S5. Classify the processed features using a multilayer perceptron and output the localization mask of the tampered area to complete the image tampering localization.

[0093] The loss function combines segmentation loss and edge detection loss to ensure that the model can accurately segment the boundaries of the target region while detecting tampered areas. The loss function is as follows:

[0094] L = L seg +λL edge

[0095] In the formula, L seg and L edge Both loss functions use binary cross-entropy, with λ being a hyperparameter used to balance the impact of segmentation and edge detection losses. By default, we set λ to 20, making the model focus more on the boundary regions of the tampered areas.

[0096] Given the rapid development of digital image tampering technology, the demand for image authenticity verification in fields such as social media and judicial evidence collection is becoming increasingly urgent. Addressing the shortcomings of existing methods in complex scenarios, such as insufficient accuracy in image tampering localization and weak semantic understanding capabilities, this invention proposes the aforementioned method. On the one hand, it solves the problems of traditional single-modal methods relying on visual features and having limited ability to identify semantically consistent tampering such as deepfakes; on the other hand, it overcomes the deficiencies of existing multimodal fusion strategies, such as insufficient feature interaction and insufficient localization robustness.

[0097] This invention is implemented within the PyTorch framework: an image is converted into its corresponding text representation using an LLaMA model, a step performed on an NVIDIA RTX 3080 GPU with 20GB of VRAM; the training phase involves training the model with 200 epochs of parameters, a step performed on an NVIDIA RTX 4090D GPU with 24GB of VRAM. Figure 3-6 As shown in Table 1-2, the image tampering localization effect of this example is better than the experimental results of other comparison schemes.

[0098] Table 1 Image Processing Localization Performance

[0099]

[0100] Table 2 Ablation Experiments of the Model

[0101]

[0102] Example 2:

[0103] This embodiment provides a VisionTransformer image tampering localization system based on multimodal prompting guidance, including:

[0104] The image segmentation and token generation module segments the image to be detected into blocks and generates an image token; it uses an LLaMA model to generate text prompts related to the image content and converts them into text tokens.

[0105] The single-modal feature encoding module encodes image tokens using the VisionTransformer model to extract deep visual features of the image; it also uses a pre-trained language model to encode text tokens to obtain text semantic features.

[0106] The multimodal feature fusion module inputs deep visual features of the image and semantic features of the text into the multimodal interactive prompt module, and achieves feature fusion through cross-modal self-attention and cross-attention mechanisms to generate multimodal joint features;

[0107] The multi-scale feature processing module utilizes a spatial feature pyramid network to perform multi-scale processing on multimodal joint features, thereby enhancing the ability to capture tampered regions of different sizes.

[0108] The tampered area localization output module classifies the processed features using a multilayer perceptron and outputs a localization mask for the tampered area, thus completing the image tampering localization.

[0109] Example 3:

[0110] An electronic device includes a memory, a processor, and a computer program stored in the memory and running thereon. When the processor executes the program, it implements the aforementioned VisionTransformer image tampering localization method based on multimodal prompting guidance, comprising:

[0111] The image to be detected is segmented into blocks to generate an image token; a text prompt related to the image content is generated using the LLaMA model and converted into a text token.

[0112] The image token is encoded using the VisionTransformer model to extract deep visual features of the image; the text token is encoded using a pre-trained language model to obtain semantic features of the text.

[0113] Deep visual features of images and semantic features of text are input into the multimodal interactive prompt module. Feature fusion is achieved through cross-modal self-attention and cross-attention mechanisms to generate multimodal joint features.

[0114] By utilizing a spatial feature pyramid network to perform multi-scale processing on multimodal joint features, the ability to capture tampered regions of different sizes is enhanced.

[0115] The processed features are classified by a multilayer perceptron, and the location mask of the tampered area is output to complete the image tampering location.

[0116] Example 4:

[0117] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned VisionTransformer image tampering localization method based on multimodal cue guidance, comprising:

[0118] The image to be detected is segmented into blocks to generate an image token; a text prompt related to the image content is generated using the LLaMA model and converted into a text token.

[0119] The image token is encoded using the VisionTransformer model to extract deep visual features of the image; the text token is encoded using a pre-trained language model to obtain semantic features of the text.

[0120] Deep visual features of images and semantic features of text are input into the multimodal interactive prompt module. Feature fusion is achieved through cross-modal self-attention and cross-attention mechanisms to generate multimodal joint features.

[0121] By utilizing a spatial feature pyramid network to perform multi-scale processing on multimodal joint features, the ability to capture tampered regions of different sizes is enhanced.

[0122] The processed features are classified by a multilayer perceptron, and the location mask of the tampered area is output to complete the image tampering location.

[0123] Those skilled in the art will understand that the modules or steps described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, which can then be stored in a storage device for execution by a computer device. Alternatively, they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. This disclosure is not limited to any particular combination of hardware and software.

[0124] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0125] While the specific embodiments of this disclosure have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of this disclosure are still within the scope of protection of this disclosure.

Claims

1. A VisionTransformer image tampering localization method based on multimodal prompting guidance, characterized in that, Includes the following steps: The image to be detected is segmented into blocks to generate an image token; a text prompt related to the image content is generated using the LLaMA model and converted into a text token. Image tokens are encoded using the VisionTransformer model to extract deep visual features from the image; A pre-trained language model is used to encode the text token to obtain the text semantic features; Deep visual features of images and semantic features of text are input into the multimodal interactive prompt module. Feature fusion is achieved through cross-modal self-attention and cross-attention mechanisms to generate multimodal joint features. By utilizing a spatial feature pyramid network to perform multi-scale processing on multimodal joint features, the ability to capture tampered regions of different sizes is enhanced. The processed features are classified by a multilayer perceptron, and the location mask of the tampered area is output to complete the image tampering location.

2. The VisionTransformer image tampering localization method based on multimodal prompting guidance according to claim 1, characterized in that, The image to be detected is input into the LLaMA model, along with the guiding question "Please analyze the possible tampered areas and features in the image". The model generates text prompts containing descriptions of the image content and inferences about potential tampered areas through visual-language association reasoning. The output text sequence is converted into a text token after word segmentation.

3. The VisionTransformer image tampering localization method based on multimodal prompting guidance according to claim 1, characterized in that, The processing procedure of the multimodal interaction prompt module is as follows: The deep visual features of the input image are first extracted by global average pooling, and then visual features of different subspaces are extracted by scale dot product method. The semantic features of the input text are first projected using a 1×1 convolution to compress information redundancy, and then the text features of different subspaces are extracted using the scale dot product method. The processed visual features and text features are input into the cross-attention module, and deep fusion is achieved by combining residual connections to output multimodal joint features.

4. The VisionTransformer image tampering localization method based on multimodal prompting guidance according to claim 3, characterized in that, Visual features of different subspaces are extracted using the scale dot product method. The process is represented as follows: P SDP (Q,K,V)=Concat(Y1,…,Y8)W H B j =P SDP (F u ,F u ,F u ) In the formula, Q represents the query vector, K represents the key vector, and V represents the value vector. and W H All are trainable matrices, where d represents the dimension, φ(·) represents the softmax function, and P SDP Representing the SDP block, it is the module that performs scale dot product attention calculations; B j This represents the visual features obtained after processing by the SDP block. Text features in different subspaces are extracted using the scale dot product method. The process is represented as follows: T SDP (Q,K,V)=Concat(X1,…,X8)W H C n =T SDP (F l ,F l ,F l ) In the formula, T represents the semantic features of the input text, where m represents the modality identifier and n-1 represents the hierarchical information of the features; SDP T represents the text features SDP Block; C n This represents the text features obtained after processing by the SDP block.

5. The VisionTransformer image tampering localization method based on multimodal prompting guidance according to claim 3, characterized in that, In the cross-attention module, visual features serve as the query vector, and text features serve as the key and value vectors. Through the attention weights between visual and text features, deep cross-modal information alignment and fusion are achieved. This process can be written as follows: In the formula, This represents the fused features related to visual features after processing by the cross-attention module. This represents the fused features related to the text features after processing by the cross-attention module.

6. The VisionTransformer image tampering localization method based on multimodal prompting guidance according to claim 1, characterized in that, The spatial feature pyramid network's processing steps are as follows: Multimodal joint features are downsampled using convolutional operations to generate low-resolution, high-semantic features; then, high-resolution, low-semantic features are upsampled using deconvolutional operations. A lightweight SegFormer decoder is used to process the generated multi-scale features, first unifying features of different scales to the same resolution, then concatenating and fusing them, and finally outputting a tampering probability map through a linear layer. Simultaneously, dilation and erosion operations are used to generate edge masks. Dilation expands edge regions to capture subtle tampering traces, while erosion removes noise to accurately define the contours of the tampered region.

7. The VisionTransformer image tampering localization method based on multimodal prompting guidance according to claim 1, characterized in that, The loss function is designed by combining segmentation loss and edge detection loss, as follows: L=L seg +λL edge In the formula, L seg and L edge Both are binary cross-entropy loss functions, where λ is a hyperparameter used to balance the effects of segmentation loss and edge detection loss.

8. A VisionTransformer image tampering localization system based on multimodal prompting guidance, characterized in that, include: The image segmentation and token generation module segments the image to be detected into blocks and generates an image token. Use the LLaMA model to generate text prompts related to image content, and convert them into text tokens; The single-modal feature encoding module encodes image tokens using the VisionTransformer model to extract deep visual features from the image; A pre-trained language model is used to encode the text token to obtain the text semantic features; The multimodal feature fusion module inputs deep visual features of the image and semantic features of the text into the multimodal interactive prompt module, and achieves feature fusion through cross-modal self-attention and cross-attention mechanisms to generate multimodal joint features; The multi-scale feature processing module utilizes a spatial feature pyramid network to perform multi-scale processing on multimodal joint features, thereby enhancing the ability to capture tampered regions of different sizes. The tampered area localization output module classifies the processed features using a multilayer perceptron and outputs a localization mask for the tampered area, thus completing the image tampering localization.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running thereon, characterized in that, When the processor executes the program, it implements the VisionTransformer image tampering localization method based on multimodal prompting guidance as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the VisionTransformer image tampering location method based on multimodal prompting guidance as described in any one of claims 1-7.

Citation Information

Cited By

  • Multi-modal geographic positioning method and system based on three-dimensional condition prompt learning

    CN122244168A

  • A Multimodal Geolocation Method and System Based on 3D Conditional Cueing Learning

    CN122244168B