A training-free image forgery localization method and system

CN122574606BActive Publication Date: 2026-09-22PEKING UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202611055089.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-09-22
Estimated Expiration
2046-07-16

AI Technical Summary

Technical Problem

本发明的目的是融合多模态大语言模型的跨模态推理能力与分割模型(如SAM)的细粒度分割能力,通过识别伪造线索生成像素级伪造定位结果,由此解决传统方法依赖训练数据导致泛化性差的技术问题

Benefits of technology

[0017]本发明可实现针对未知伪造类型的像素级精确定位;本发明采用多模态大语言模型与分割模型协同的免训练方式,无需依赖预先收集的大规模篡改数据集即可直接进行推理,克服了现有深度学习方法因训练数据受限导致对新型伪造手段泛化能力差的问题,可提高图像取证系统在面对未知攻击时的鲁棒性和适应性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574606B_ABST
    Figure CN122574606B_ABST
Patent Text Reader

Abstract

The application discloses a kind of training-free image forgery positioning method and system, belong to digital image processing field.The application does not need to rely on pre-collected labeled training data, by fusing the cross-modal inference ability of multi-modal large language model and the fine-grained positioning ability of segmentation model, realize pixel-level image forgery area identification and accurate positioning.The application first uses multi-modal large language model to carry out coarse-grained suspicious area exploration to the image to be detected, obtains candidate area;Then through the multi-round collaborative optimization of multi-modal large language model and segmentation model, the candidate area is refined in fine-grained iteration, and the forgery area is obtained.The application overcomes the technical problems that the existing training-based forgery positioning method has poor generalization ability for unobserved forgery types and relies on a large number of labeled data, can effectively deal with various image forgery scenarios, and improves the flexibility and adaptability of image forensics system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of digital image processing, and more specifically, relates to an image forgery localization method and system that utilizes a multimodal large language model and a segmentation model working in concert, without requiring training on specific forged data. Background Technology

[0002] Image forgery localization aims to achieve pixel-level localization of forged regions, requiring the detection of subtle differences between the real and forged content in images altered by different forgery methods. To this end, existing methods typically follow a training-then-inference paradigm, extracting forgery trace features in the spatial or frequency domains to handle images processed by different forgery techniques. Xiamen University, in its patent application "A Deep Learning-Based Method for Tampered Image Detection" (patent application number CN201910573995.3, patent publication number CN110349136A), constructed a multi-scale noise-constrained convolutional layer to extract high-frequency residuals, utilizing a two-stream network combined with a multi-task learning mechanism. Based on a large amount of labeled data, it optimized model parameters through backpropagation, thereby achieving classification and fine segmentation of tampered regions in images. However, the model's discriminative ability is limited to the range of forgery methods covered by the training set, exhibiting poor generalization ability to unknown forgery types. Furthermore, with the rapid iteration of forgery techniques, new forgery methods often emerge in unpredictable ways, resulting in a severe shortage of pre-collected training samples, causing existing methods to often lag behind in practical applications.

[0003] In recent years, multimodal large language models have demonstrated performance advantages in various computer vision tasks, including image retrieval, semantic segmentation, and multimodal understanding. However, directly applying multimodal large language models to forged image localization has limitations: on the one hand, existing multimodal large language models have not been exposed to relevant forgery techniques during the training phase, making it difficult to handle all the fine-grained differences between forged and real information, resulting in omissions and biases in forged region localization; on the other hand, images in real-world scenes often contain regions altered by various forgery methods, making it difficult for multimodal large language models to capture fine-grained forgery traces in different regions from a global perspective. In addition, while existing segmentation models (such as SAM) possess high-precision segmentation capabilities, they lack prior knowledge of forgery clues and cannot autonomously identify forged regions without precise cues.

[0004] In summary, there is a need for a localization method that can effectively integrate the semantic reasoning capabilities of multimodal large language models with the precise localization capabilities of segmentation models, and can handle various complex forgery scenarios without requiring training on specific data. This would solve the technical problems of coarse localization by large models in existing technologies and poor generalization caused by the reliance on training data in traditional methods. Summary of the Invention

[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a training-free image forgery localization method and system, enabling forgery region localization without a training process. The purpose of this invention is to integrate the cross-modal reasoning capabilities of multimodal large language models with the fine-grained segmentation capabilities of segmentation models (such as SAM), generating pixel-level forgery localization results by identifying forgery clues, thereby solving the technical problem of poor generalization caused by the reliance on training data in traditional methods.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A training-free image forgery localization method includes the following steps: A multimodal large language model is used to explore coarse-grained suspicious regions in the image to be detected, thereby obtaining candidate regions; Through multiple rounds of collaborative optimization using a multimodal large language model and a segmentation model, candidate regions are refined through fine-grained iterative processes to obtain forged regions.

[0007] Furthermore, the method of using a multimodal large language model to perform coarse-grained exploration of suspicious regions in the image to be detected includes: A multimodal large language model is used to analyze the image to be detected and generate a scene semantic description; Based on scene semantic description, forgery clues are extracted from three dimensions: semantic consistency, digital integrity, and physical environment consistency, and a candidate region list containing forgery clues is generated. An adaptive expansion mechanism is used to dynamically expand the bounding box of the candidate region to include the necessary contextual information.

[0008] Furthermore, the forgery clues are extracted from three dimensions: semantic consistency, digital integrity, and physical environment consistency. The following prompt words are used: The prompts for detecting semantic consistency include: defining the theme scene based on the semantic description of the image and the overall visual information of the image; examining each key object in the image based on the theme scene; and identifying objects that have logical conflicts with the theme scene. Tips for checking digital integrity include: examining the edges of all individual objects, identifying and establishing a real object as a reference, comparing other suspicious objects with the reference to look for typical traces of splicing forgery, and then checking whether the digital noise or film grain of the image is uniform, and whether the blockiness or color banding of the JPEG is consistent. Tips for checking the consistency of the physical environment include: checking whether all shadows are consistent in direction, softness, and depth; whether objects cast reasonable contact shadows where they touch the plane; and whether all objects and textures follow a consistent perspective and whether their relative sizes are logical.

[0009] Furthermore, the refinement of candidate regions through multi-round collaborative optimization of the multimodal large language model and the segmentation model includes: A multimodal large language model is used to perform binary classification and verification of candidate regions to filter out false positive regions; By iteratively decomposing the task using a multimodal large language model, generating masks using a segmentation model, and evaluating the quality of the masks using a multimodal large language model, the boundaries of the forged regions are gradually optimized. Integrate historical information from iterative cycles to generate a full-resolution binary mask and forensic report.

[0010] Furthermore, the step of using a multimodal large language model to perform binary classification and verification of candidate regions to filter out false positive regions includes: Each candidate region is prioritized based on its confidence score and the richness of associated clues. For each candidate region sorted by priority, a sub-region image is cropped from the original image based on its expanded bounding box; The sub-region image and its multi-dimensional evidence set are input into the multimodal large language model for binary classification. Candidate regions judged as normal regions by the multimodal large language model are then removed.

[0011] Furthermore, the iterative method of decomposing the task using a multimodal large language model, generating a mask using a segmentation model, and evaluating the mask quality using a multimodal large language model to gradually optimize the boundary of the forged region includes: The evidence collection task is decomposed based on the current forgery clues using a multimodal large language model, and then the task is executed and segmentation prompt words are generated; The segmentation model is used to segment sub-regions based on segmentation prompts to generate the current mask; The multimodal large language model judges whether the current segmentation effect meets the standard from three aspects: whether the original segmentation target is completely covered, whether it contains all the forged regions mentioned in the evidence chain, and whether it contains normal regions that are clearly unrelated to the forged evidence. If the segmentation effect does not meet the standard, the iteration is terminated; otherwise, the intersection-union ratio (IUU) of the current mask and the previous mask is calculated. If the IUU is greater than the set threshold or the maximum number of iterations is reached, the iteration is terminated; otherwise, the next round of optimization is entered.

[0012] Furthermore, the method of using a multimodal large language model to decompose the evidence collection task based on the current forgery clues includes thinking and task planning from four dimensions, transforming vague forgery suspicions into a list of executable, multi-angle verification tasks centered on specific objects; the four dimensions include: edge / boundary integrity, noise / texture consistency, lighting rationality, and semantic logic.

[0013] A training-free image forgery localization system, characterized in that it comprises: The coarse-grained suspicious region exploration module is used to explore the suspicious regions in the image to be detected using a multimodal large language model, and obtain candidate regions; The fine-grained iterative refinement module is used to refine candidate regions through multiple rounds of collaborative optimization of a multimodal large language model and a segmentation model, thereby obtaining forged regions.

[0014] The present invention also provides a computer device including a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the methods described above.

[0015] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a computer, implements the above-described method.

[0016] Compared with the prior art, the above-described technical solutions conceived in this invention can achieve the following beneficial effects.

[0017] This invention can achieve pixel-level precise localization for unknown forgery types; this invention adopts a training-free method that combines a multimodal large language model and a segmentation model, and can directly perform inference without relying on a pre-collected large-scale tampered dataset, overcoming the problem that existing deep learning methods have poor generalization ability to new forgery methods due to limited training data, and can improve the robustness and adaptability of image forensics systems when facing unknown attacks. This invention transforms complex forgery detection into a multi-dimensional evidence search task based on semantic, digital, and physical environment consistency through a coarse-grained suspicious region exploration stage using a multimodal large language model. This alleviates the difficulty of a single model in capturing high-level semantic contradictions or subtle physical anomalies, and ensures a high recall rate for potential forgery regions in images. The fine-grained iterative refining stage proposed in this invention utilizes a multi-expert collaborative optimization strategy. It can effectively filter out false alarm background noise introduced by the coarse-grained stage through multiple rounds of iterative discrimination, and can also use the zero-sample segmentation capability of the segmentation model to correct the boundaries of tampered regions with blurred edges or complex shapes, thereby achieving a leap from coarse candidate boxes to fine pixel-level masks. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the steps of a training-free image forgery localization method according to an embodiment of the present invention.

[0019] Figure 2 This is an overall framework diagram of a training-free image forgery localization method according to an embodiment of the present invention.

[0020] Figure 3 This is a flowchart of the main process of the fine-grained iterative refining stage in an embodiment of the present invention.

[0021] Figure 4 This is a block diagram of a training-free image forgery localization system according to an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0023] In one embodiment, the present invention provides a training-free image forgery localization method, such as... Figure 1 As shown, it includes the following steps: A multimodal large language model is used to explore coarse-grained suspicious regions in the image to be detected, thereby obtaining candidate regions; Through multiple rounds of collaborative optimization using a multimodal large language model and a segmentation model, candidate regions are refined through fine-grained iterative processes to obtain forged regions.

[0024] In one embodiment, the use of a multimodal large language model to perform coarse-grained suspicious region exploration of the image to be detected includes: A multimodal large language model is used to analyze the image to be detected and generate a scene semantic description; Based on scene semantic description, forgery clues are extracted from three dimensions: semantic consistency, digital integrity, and physical environment consistency, and a candidate region list containing forgery clues is generated. An adaptive expansion mechanism is used to dynamically expand the bounding box of the candidate region to include the necessary contextual information.

[0025] In one embodiment, the extraction of forgery clues from three dimensions—semantic consistency, digital integrity, and physical environment consistency—uses the following prompt words: The prompts for detecting semantic consistency include: defining the theme scene based on the semantic description of the image and the overall visual information of the image; examining each key object in the image based on the theme scene; and identifying objects that have logical conflicts with the theme scene. Tips for checking digital integrity include: examining the edges of all individual objects, identifying and establishing a real object as a reference, comparing other suspicious objects with the reference to look for typical traces of splicing forgery, and then checking whether the digital noise or film grain of the image is uniform, and whether the blockiness or color banding of the JPEG is consistent. Tips for checking the consistency of the physical environment include: checking whether all shadows are consistent in direction, softness, and depth; whether objects cast reasonable contact shadows where they touch the plane; and whether all objects and textures follow a consistent perspective and whether their relative sizes are logical.

[0026] In one embodiment, the fine-grained iterative refinement of candidate regions through multi-round collaborative optimization of a multimodal large language model and a segmentation model includes: A multimodal large language model is used to perform binary classification and verification of candidate regions to filter out false positive regions; By iteratively decomposing the task using a multimodal large language model, generating masks using a segmentation model, and evaluating the quality of the masks using a multimodal large language model, the boundaries of the forged regions are gradually optimized. Integrate historical information from iterative cycles to generate a full-resolution binary mask and forensic report.

[0027] In one embodiment, the step of using a multimodal large language model to perform binary classification and verification of candidate regions to filter out false positive regions includes: Each candidate region is prioritized based on its confidence score and the richness of associated clues. For each candidate region sorted by priority, a sub-region image is cropped from the original image based on its expanded bounding box; The sub-region image and its multi-dimensional evidence set are input into the multimodal large language model for binary classification. Candidate regions judged as normal regions by the multimodal large language model are then removed.

[0028] In one embodiment, the iterative approach of decomposing the task using a multimodal large language model, generating a mask using a segmentation model, and evaluating the mask quality using a multimodal large language model to progressively optimize the boundary of the forged region includes: The evidence collection task is decomposed based on the current forgery clues using a multimodal large language model, and then the task is executed and segmentation prompt words are generated; The segmentation model is used to segment sub-regions based on segmentation prompts to generate the current mask; The multimodal large language model judges whether the current segmentation effect meets the standard from three aspects: whether the original segmentation target is completely covered, whether it contains all the forged regions mentioned in the evidence chain, and whether it contains normal regions that are clearly unrelated to the forged evidence. If the segmentation effect does not meet the standard, the iteration is terminated; otherwise, the intersection-union ratio (IUU) of the current mask and the previous mask is calculated. If the IUU is greater than the set threshold or the maximum number of iterations is reached, the iteration is terminated; otherwise, the next round of optimization is entered.

[0029] In one embodiment, the use of a multimodal large language model to decompose the evidence collection task based on the current forgery clues includes thinking and task planning from four dimensions, transforming ambiguous forgery suspicions into an executable, multi-angle verification task list centered on specific objects; the four dimensions include: edge / boundary integrity, noise / texture consistency, lighting rationality, and semantic logic.

[0030] In one embodiment, the present invention provides a training-free image forgery localization method, comprising the following steps: S1. Coarse-grained Suspicious Region Exploration. The input target image is analyzed using Qwen2.5-VL-7B-Instruct to generate a global scene semantic description of the target image as prior knowledge. Specific prompts are designed to guide Qwen2.5-VL-7B-Instruct to check for forgery traces in the input image based on three dimensions: semantic consistency, digital integrity, and physical environment consistency, and to generate a list of candidate suspicious regions containing forgery clues.

[0031] S2. Fine-grained iterative refinement. A multi-expert hierarchical node collaboration and multi-round iterative optimization architecture is constructed to simulate the analytical logic of human forensic experts. Combining the cross-modal reasoning capabilities of Qwen2.5-VL-7B-Instruct and the precise segmentation capabilities of the segmentation model LangSAM, iterative judgment and segmentation optimization are performed on the expanded candidate regions to achieve pixel-level localization results.

[0032] In one embodiment, coarse-grained suspicious region exploration specifically includes the following steps: Generate scene semantic description: Use Qwen2.5-VL-7B-Instruct to generate a global scene semantic description of the target image, which serves as prior knowledge for subsequent forgery region detection; Multidimensional evidence search: Using Qwen2.5-VL-7B-Instruct as a forgery clue extractor, it is guided to perform detection from three dimensions: semantic consistency, digital integrity, and physical environment consistency, and outputs a list of candidate regions including coarse-grained bounding boxes, forgery clue descriptions, and confidence levels. Dynamic Region Expansion: Bounding boxes generated by Qwen2.5-VL-7B-Instruct often only cover the central object, easily missing forgery traces strongly related to the core region and revealing contradictory scene context information. To solve this problem, this invention introduces an adaptive expansion mechanism. The expansion ratio α of the dynamic region expansion is based on the candidate region. area ratio Confirmed, the formula is as follows: in, =area( ) / area(I), where area( ) represents a coarse-grained bounding box The area of ​​image I is represented by area(I).

[0033] In one embodiment, the prompt word used when Qwen2.5-VL-7B-Instruct is used as a forgery clue extractor is as follows: Clues for detecting semantic consistency in a scene: First, based on the image's semantic description and overall visual information, quickly define a core thematic scene, such as "an abandoned outdoor corner" or "a busy office desk." Using this thematic scene as a benchmark, examine each key object in the image, identifying those that logically conflict with the scene in terms of function, category, causality, or spatiotemporal context. For example, in a 19th-century Victorian-style room scene, the presence of a modern smartphone creates a clear conflict in terms of spatiotemporal context, making it a key point of conflict.

[0034] Tips for checking digital integrity: Starting from the top left corner of the image, systematically examine the edges of all individual objects from left to right and top to bottom to identify and establish a definite real object as a reference. Prioritize objects located on the focal plane, well-lit, and with the sharpest edges as references. Compare other suspicious objects with this reference, looking for typical signs of splicing forgery: abnormal pixel information such as sharp cuts, scratches, burrs, or jagged imperfections; abnormal smoothness or blurriness; inconsistent sharpness; color clipping; halos, etc. Then check if the digital noise or film grain of the image is uniform across the entire frame? Is the blockiness or color banding of the JPEG consistent? Tips for checking the consistency of the physical environment: Are all shadows consistent in direction, softness, and depth? Do objects cast reasonable contact shadows where they touch the plane? Do all objects and textures follow a unified perspective? Are their relative sizes logical? In one embodiment, fine-grained iterative refining specifically includes the following steps: Preliminary assessment: Based on the confidence and cue richness of candidate regions, the images of the cropped sub-regions are input into Qwen2.5-VL-7B-Instruct for binary classification verification, and regions judged as true are filtered out. Multi-round collaborative optimization: For unfiltered regions, Qwen2.5-VL-7B-Instruct is used to decompose the evidence collection task based on specific prompt words, LangSAM segmentation model is used to perform segmentation, and finally Qwen2.5-VL-7B-Instruct is used to evaluate the segmentation quality, and multiple rounds of iterative optimization are performed. Final judgment: Integrate the judgment history information of the entire evidence collection process (multi-round collaborative optimization) to generate the final result (full-resolution binary mask) and evidence collection report.

[0035] In one embodiment, the prompt words used when performing binary classification using Qwen2.5-VL-7B-Instruct are as follows: You are a top-tier digital image forensics expert. Your task is to quickly determine whether a suspicious area warrants further in-depth analysis. The criteria for judgment are: Are the clues truly visible in the image? Can the clues be explained by natural phenomena (such as lighting, perspective, or inherent object features) or normal image imperfections? If the clues cannot be reasonably explained, then "yes"; if they can, then "no". Briefly explain your reasoning in under 50 words, based on visible image features and without relying on external data.

[0036] In one embodiment, the prompt word set for guiding the Qwen2.5-VL-7B-Instruct decomposition task is as follows: "You" are a top-tier chief strategist in image forensics analysis. Your core competency is transforming vague forgery suspicions into a checklist of actionable, multi-faceted verification tasks centered around specific objects. "You" will approach task planning and analysis from the following four core forensic dimensions: 1) Edge / Boundary Integrity: The most direct trace of forgery, focusing on examining sharp cuts, scratches, burrs, and abnormal blurring left by splicing and copying operations; 2) Noise / Texture Consistency: Examining image content, noise / texture from content filler or foreign objects typically does not match the original image content; 3) Lighting and Shadow Rationality: Verifying whether the lighting and shadows of objects follow a unique physical law, focusing on checking whether the direction of the light source, shadows (presence, shape, and direction of shadows), highlights, and reflections are consistent with scene logic; 4) Semantic Logic: Checking whether the image content conforms to common sense and physical interaction logic, judging whether the story the image tells is reasonable.

[0037] In one embodiment, the multi-round collaborative optimization (i.e., the collaborative refinement process of intermediate nodes) includes: in the t-th iteration, Qwen2.5-VL-7B-Instruct decomposes specific forensic tasks based on the current forgery clues, Qwen2.5-VL-7B-Instruct executes the specific tasks and generates segmentation prompts. Then the segmentation model LangSAM is based on the prompt words. Segment the sub-regions to generate the current mask. The Qwen2.5-VL-7B-Instruct determines whether the current segmentation effect meets the standard based on the following three aspects: whether the original segmentation target is completely covered; whether it contains all the forged regions mentioned in the evidence chain; and whether it contains normal regions that are clearly unrelated to the forged evidence. If the conditions are met, the iteration terminates; otherwise, the current mask is calculated. Compared to the previous round of masking If the intersection-to-union ratio (IoU) is greater than the set threshold of 0.99 or the maximum number of iterations of 3 is reached, the iteration is terminated; otherwise, the next round of optimization is entered.

[0038] In one embodiment, this invention proposes a training-free image forgery localization method, comprising two parts: coarse-grained suspicious region exploration and fine-grained iterative refinement. In the coarse-grained suspicious region exploration stage, a multimodal large language model is first used to generate a scene semantic description. Customized prompt words guide the multimodal large language model as a forgery clue extractor to extract forgery clues from three dimensions: semantic consistency, digital integrity, and physical environment consistency, generating a high-recall candidate region list containing forgery clues. Subsequently, an adaptive expansion mechanism is introduced to dynamically expand the candidate region bounding boxes to include necessary contextual information. In the fine-grained iterative refinement stage, a multi-expert collaborative architecture is adopted, allowing Qwen2.5-VL-7B-Instruct to assume different expert roles and perform corresponding functions. First, Qwen2.5-VL-7B-Instruct is used to predict the authenticity of candidate regions to filter false positives. Then, through a cyclical iterative process of generating fine-grained prompt words using Qwen2.5-VL-7B-Instruct, generating masks using the segmentation model LangSAM, and evaluating the mask quality using Qwen2.5-VL-7B-Instruct, the boundaries of the forgery regions are gradually optimized. This invention introduces a multi-round iterative collaborative mechanism in the fine-grained iterative refinement stage. On the one hand, it solves the problem that multimodal large language models can only provide coarse bounding boxes or text descriptions and are difficult to achieve pixel-level localization directly. On the other hand, it utilizes the zero-shot segmentation capability of the segmentation model to respond to forgery clues and gradually eliminates background noise introduced in the coarse-grained suspicious region exploration stage through iterative feedback, thereby achieving accurate pixel-level forgery localization without any training data. In addition, this invention adopts a training-free inference paradigm instead of the traditional training-then-inference paradigm. On the one hand, it abandons the dependence on large-scale labeled and tampered datasets and avoids the overfitting problem caused by insufficient training data. On the other hand, it utilizes the general knowledge and reasoning ability of large models, enabling them to effectively cope with novel forgery attacks that have not appeared in the training set, improving the model's generalization ability and robustness in unknown forgery scenarios.

[0039] refer to Figure 2 This embodiment provides a training-free image forgery localization method, which includes the following steps: S1. Coarse-grained Suspicious Region Exploration. Coarse-grained suspicious region exploration aims to maximize the discovery of potential forgery regions in the image. This stage performs a comprehensive inspection of the entire image, identifying all potential regions of interest and providing a high-quality list of candidate regions for the fine-grained iterative refinement in the second stage. The coarse-grained suspicious region exploration stage can be broken down into the following three key steps: S11. Generate a scene semantic description using Qwen2.5-VL-7B-Instruct. This initial step is inspired by the human visual perception system. When humans observe an image, they don't immediately focus on pixel-level noise or edge details, but rather quickly scan to understand the overall scene context. After understanding the scene, the brain can quickly identify semantic anomalies using common sense. This step aims to simulate this efficient perception mechanism: for the image to be detected, Qwen2.5-VL-7B-Instruct first generates a global scene semantic description of the entire image, denoted as... This description will serve as prior knowledge to assist in subsequent semantic anomaly (semantic consistency) detection.

[0040] S12. Multidimensional Evidence Search. Using Qwen2.5-VL-7B-Instruct as a forgery clue extractor, prompts are set to guide it to systematically search for forgery clues from three different dimensions: Semantic consistency: Based on scene baseline analysis, check for illogical semantic contradictions in the image (e.g., an outdoor-specific object appearing in an indoor scene). Hints for detecting scene semantic consistency: First, based on the image's semantic description and overall visual information, quickly define a core thematic scene, such as "an abandoned outdoor corner" or "a busy office desk." Using this thematic scene as a benchmark, examine each key object in the image, identifying those that logically conflict with the scene in terms of function, category, causality, or spatiotemporal context. For example, in a 19th-century Victorian room scene, the presence of a modern smartphone creates a clear conflict in terms of spatiotemporal context, making it a key point of conflict.

[0041] Digital Integrity: Check for anomalous pixels at object edges (e.g., blurred edges, pixel jumps) or inconsistencies in noise patterns within the image (e.g., sudden increases or disappearances of noise in localized areas). Tips for checking digital integrity: Starting from the top left corner of the image, systematically examine the edges of all individual objects from left to right and top to bottom to identify and establish a definite real object as a reference. Prioritize objects located on the focal plane, well-lit, and with the sharpest edges as references. Compare other suspicious objects to this reference, looking for typical signs of splicing or forgery: sharp cuts, scratches, burrs, jagged imperfections, abnormal pixel information, abnormal smoothness or blurring, inconsistent sharpness, color bleeding, halos, etc. Then check if the digital noise or film grain of the image is uniform across the entire frame. Is the blockiness or color banding in the JPEG consistent? Physical environment consistency: Verify whether the lighting direction, shadow casting, and object proportions in the image conform to the physical laws of the real world (e.g., the direction of an object's shadow is contrary to the direction of the light source). Hints for checking physical environment consistency: Are all shadows consistent in direction, softness, and depth? Do objects cast reasonable contact shadows where they touch the plane? Do all objects and textures follow a unified perspective? Are their relative sizes logical? The output of this step is a list of candidate regions, denoted as... in, This represents a set of multiple candidate objects for suspicious regions. This represents the i-th candidate object for a suspicious region. This represents the total number of candidate regions. It is a textual description of the suspicious area. It is a coarse-grained bounding box. It is a list of the evidence that was found. {High, Medium, Low} represents the confidence score of the suspicious region given by Qwen2.5-VL-7B-Instruct.

[0042] S13. Dynamic Region Expansion: Bounding boxes generated by multimodal large language models often only cover the central object, easily missing forgery traces strongly related to the core region and revealing contradictory scene context information. To solve this problem, this invention introduces an adaptive expansion mechanism. The expansion ratio α of the dynamic region expansion is determined based on the candidate region. area ratio Confirmed, the formula is as follows: in, =area( ) / area(I), where area( ) represents a coarse-grained bounding box Let area(I) represent the area of ​​image I. Then the expanded bounding box... for: in, express Minimum x-coordinate of the top left corner express The smallest y-coordinate of the top left corner This represents the width of image I. This represents the height of image I.

[0043] Simultaneously, it is necessary to ensure that the coordinates of the expanded bounding box do not exceed the image boundary (if they do, the coordinates should be adjusted to the image edge coordinates). Finally, based on the expanded bounding box... Generate the corresponding coarse-grained mask .

[0044] S2. Fine-grained iterative refining. For example... Figure 3 As shown, the fine-grained iterative refinement part adopts a multi-expert hierarchical node collaboration and multi-round iterative optimization architecture to simulate the analysis logic of human forensic experts. At the same time, a collaborative mechanism between Qwen2.5-VL-7B-Instruct and the segmentation model LangSAM is introduced to balance semantic understanding ability and boundary localization accuracy. Figure 3 In this context, round 1 and round 2 refer to the leader-worker collaborative evidence verification and initial mask generation, as well as the evaluator's evaluation and mask optimization. The leader is the coordinating role responsible for task decomposition and result aggregation, the worker is the execution role focusing on a single evidence dimension and performing multi-dimensional verification, and the evaluator is the verification role verifying mask quality and generating optimization prompts. Figure 3 In this context, xN represents N rounds of iterative optimization.

[0045] The fine-grained iterative refinement may specifically include the following steps: the root node quickly filters the real regions in the candidate suspicious regions; the intermediate nodes optimize the boundary accuracy of the forged region mask through collaborative iteration of Qwen2.5-VL-7B-Instruct and LangSAM; and the leaf nodes output the final forged region mask and structured forensic report. S21. Preliminary judgment of the root node: The coarse-grained suspicious region exploration module obtains a candidate region list after dynamic region expansion. ,because False positive regions exist, necessitating rapid screening to identify normal candidate regions and improve overall efficiency. First, based on the confidence score and richness of associated clues, each candidate region is... Perform priority sorting, where Indicates that the image has passed through bounding box in The cropped sub-regions. Priority is calculated based on scores. accomplish: In the formula, Map confidence levels to numerical values ​​(high = 1, medium = 0.5, low = 0). The maximum number of clues found across all candidate regions.

[0046] Subsequently, for each candidate region sorted by priority According to its extended bounding box Sub-region images are cropped from the original image. This cropping operation effectively isolates irrelevant background, allowing the model to focus on local verification. Finally, the sub-region image and its multi-dimensional evidence set are input into Qwen2.5-VL-7B-Instruct, prompting Qwen2.5-VL-7B-Instruct to perform binary classification: determine whether the visual evidence (the three dimensions of evidence in S12) supports forged clues and cannot be explained by reasonable natural logic. If Qwen2.5-VL-7B-Instruct outputs a judgment result of "this sub-region is a normal region", then the corresponding candidate region is... Those that are screened out will not proceed to the next stage.

[0047] The prompts used when performing binary classification using Qwen2.5-VL-7B-Instruct are as follows: "You are a top-tier digital image forensics expert. Your task is to quickly determine whether a suspect area deserves further in-depth analysis. The criteria for judgment are: Are these clues truly visible in the image? Can the clues be explained by natural phenomena (such as lighting, perspective, or the object's inherent features) or normal image defects? If the clues cannot be reasonably explained, then 'yes'; if they can, then 'no'. Briefly explain your reasoning within 50 words, based on visible image features, without relying on external data."

[0048] S22. Multi-round collaborative optimization of intermediate nodes: After the initial screening by the root node, the high-suspicious regions that were not eliminated are... The intermediate nodes are then subjected to fine-grained verification. This invention refines the forged region through a collaborative optimization loop of Qwen2.5-VL-7B-Instruct and LangSAM, decoupling the task as follows: LangSAM generates precise boundaries based on prompts, while Qwen2.5-VL-7B-Instruct assumes different expert roles to handle semantic judgment and prompt optimization.

[0049] At the start of the iteration, Qwen2.5-VL-7B-Instruct receives candidate regions. With the initial set of clues This is then broken down into specific sub-tasks. The prompts used to guide Qwen2.5-VL-7B-Instruct in decomposing tasks are as follows: "You" are a top-tier chief strategist in image forensics analysis. Your core competency is to transform vague forgery suspicions into a list of actionable, multi-faceted verification tasks centered on specific objects. "You" will consider and plan tasks from the following four core forensic dimensions: 1) Edge / Boundary Integrity: The most direct trace of forgery, focusing on checking for sharp cuts, scratches, burrs, and abnormal blurring left by splicing and copying operations; 2) Noise / Texture Consistency: Checking image content, noise / texture from content filler or foreign objects usually does not match the original image content; 3) Lighting and Shadow Rationality: Verifying whether the lighting and shadow of objects follow a unique physical law, the core being checking whether the direction of the light source, shadows (presence, shape, and direction of shadows), highlights, and reflections are consistent with scene logic; 4) Semantic Logic: Checking whether the image content conforms to common sense and physical interaction logic, judging whether the story told by the image is reasonable. For example, if the clue set contains "object edge pixel anomalies", the specific verification instruction issued by Qwen2.5-VL-7B-Instruct is "zoom in on the edge of the target object and check if the edge has a scene depth inconsistency, excessively sharp clipping or jagged steps".

[0050] Subsequently, Qwen2.5-VL-7B-Instruct executes the previously assigned sub-tasks, focusing on a single forensic dimension (such as edge detection and illumination analysis), performing in-depth analysis of the sub-region, and generating segmentation prompts based on the analysis content. .

[0051] Then, the segmentation model LangSAM based on the prompt words Segment the sub-regions to generate the current mask. Qwen2.5-VL-7B-Instruct is evaluated from the following three aspects. Does the evidence completeness criterion need to be met? Is the original segmentation target completely covered? Does it include all forged regions mentioned in the evidence chain? Does it include normal regions that are clearly unrelated to the forged evidence? If the criteria are met, the iteration terminates; otherwise, the current mask is calculated. Compared to the previous round of masking If the intersection-to-union ratio (IoU) is greater than the set threshold of 0.99 or the maximum number of iterations of 3 is reached, the iteration is terminated; otherwise, the next round of optimization is entered.

[0052] S23. Leaf Node Final Judgment: As the final decision-making step, the leaf node integrates the iteration history of intermediate nodes to generate a clear forged region judgment result and standardized output. Its core responsibility is to ensure the accuracy, interpretability and usability of the final result.

[0053] First, for regions where intermediate nodes are identified as forged, leaf nodes are based on the expanded bounding box from the first stage. The refined sub-region mask Map back to the coordinate system of the original image I to generate a full-resolution binary mask. : in, This represents the coordinates of any pixel in the forged region. Represents the extended bounding box of the candidate region The minimum x-coordinate of the top left corner, Represents the extended bounding box of the candidate region The minimum y-coordinate of the top left corner. This represents the mask corresponding to the forged region within the sub-region.

[0054] Then, the leaf nodes summarize all clues, verification conclusions, and mask optimization records from the intermediate node iteration process, generating a structured report containing a complete chain of evidence, making the judgment results of the forged area interpretable.

[0055] In this invention, Qwen2.5-VL-7B-Instruct can be replaced with multimodal large language models such as Qwen3-VL-8B-Instruct and Llava-v1.5-7B-hf; LangSAM can be replaced with segmentation models such as X-SAM and Grounded-SAM.

[0056] The main innovative points of this invention include: 1) This invention provides a training-free method for locating image forgeries. It does not rely on a large-scale tampered dataset collected in advance. It integrates the cross-modal reasoning ability of a multimodal large language model with the fine-grained segmentation ability of a segmentation model. It can effectively deal with unknown forgery types and improve the generalization ability, robustness and adaptability of image forensics systems to new forgery methods.

[0057] 2) In the coarse-grained suspicious region exploration stage, complex forgery detection is transformed into a multi-dimensional evidence search task based on semantic consistency, digital integrity, and physical environment consistency. At the same time, an adaptive dynamic region expansion mechanism is introduced to determine the expansion ratio according to the area ratio of candidate regions, ensuring a high recall rate of potential forgery regions and avoiding the omission of relevant forgery traces and scene context information.

[0058] 3) In the fine-grained iterative refinement stage, a multi-expert hierarchical node collaboration and multi-round iterative optimization architecture is constructed to simulate the analysis logic of human forensic experts: the root node makes an initial judgment to filter false positive areas; the intermediate nodes optimize the boundaries of the fake areas through cyclical iteration of task decomposition, segmented execution, and quality assessment, and set the IoU threshold and the maximum number of iterations to control the iteration process; the leaf nodes integrate historical information to generate full-resolution binary masks and forensic reports, realizing the leap from coarse candidate boxes to fine pixel-level localization.

[0059] Another embodiment of the present invention provides a training-free image forgery localization system, such as... Figure 4 As shown, it includes: The coarse-grained suspicious region exploration module is used to explore the suspicious regions in the image to be detected using a multimodal large language model, and obtain candidate regions; The fine-grained iterative refinement module is used to refine candidate regions through multiple rounds of collaborative optimization of a multimodal large language model and a segmentation model, thereby obtaining forged regions.

[0060] The above division of modules is merely illustrative. In practical applications, the functions described above can be assigned to different functional modules as needed to complete all or part of the functions described in the aforementioned method. The specific working process of each module can be referred to the corresponding process in the aforementioned method embodiments, and will not be repeated here. Each of the above modules can be implemented entirely or partially through software, hardware, or a combination thereof.

[0061] Another embodiment of the present invention provides a computer device (computer, server, etc.) including a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing steps of the method of the present invention.

[0062] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disk) that stores a computer program, which, when executed by a computer, implements the steps of the method of the present invention.

[0063] Another embodiment of the present invention provides a computer program product, the computer program product including a computer program, which, when executed by a computer, implements the steps of the method of the present invention.

[0064] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A training-free image forgery localization method, characterized in that, Includes the following steps: A multimodal large language model is used to explore coarse-grained suspicious regions in the image to be detected, thereby obtaining candidate regions; Through multi-round collaborative optimization of a multimodal large language model and a segmentation model, candidate regions are refined through fine-grained iterative refinement to obtain forged regions; The method of using a multimodal large language model to explore coarse-grained suspicious regions in the image to be detected includes: A multimodal large language model is used to analyze the image to be detected and generate a scene semantic description; Based on scene semantic description, forgery clues are extracted from three dimensions: semantic consistency, digital integrity, and physical environment consistency, and a candidate region list containing forgery clues is generated. An adaptive expansion mechanism is used to dynamically expand the bounding box of the candidate region to include the necessary contextual information; The process of fine-grained iterative refinement of candidate regions through multi-round collaborative optimization of a multimodal large language model and a segmentation model includes: A multimodal large language model is used to perform binary classification and verification of candidate regions to filter out false positive regions; By iteratively decomposing the task using a multimodal large language model, generating masks using a segmentation model, and evaluating the quality of the masks using a multimodal large language model, the boundaries of the forged regions are gradually optimized. Integrate historical information from iterative cycles to generate a full-resolution binary mask and forensic report.

2. The method according to claim 1, characterized in that, The method extracts forgery clues from three dimensions: semantic consistency, digital integrity, and physical environment consistency. The following clue words are used: The prompts for detecting semantic consistency include: defining the theme scene based on the semantic description of the image and the overall visual information of the image; examining each key object in the image based on the theme scene; and identifying objects that have logical conflicts with the theme scene. Tips for checking digital integrity include: examining the edges of all individual objects, identifying and establishing a real object as a reference, comparing other suspicious objects with the reference to look for typical traces of splicing forgery, and then checking whether the digital noise or film grain of the image is uniform, and whether the blockiness or color banding of the JPEG is consistent. Tips for checking the consistency of the physical environment include: checking whether all shadows are consistent in direction, softness, and depth; whether objects cast reasonable contact shadows where they touch the plane; and whether all objects and textures follow a consistent perspective and whether their relative sizes are logical.

3. The method according to claim 1, characterized in that, The method of using a multimodal large language model to perform binary classification and verification of candidate regions to filter out false positive regions includes: Each candidate region is prioritized based on its confidence score and the richness of associated clues. For each candidate region sorted by priority, a sub-region image is cropped from the original image based on its expanded bounding box; The sub-region image and its multi-dimensional evidence set are input into the multimodal large language model for binary classification. Candidate regions judged as normal regions by the multimodal large language model are then removed.

4. The method according to claim 1, characterized in that, The iterative approach of decomposing the task using a multimodal large language model, generating a mask using a segmentation model, and evaluating the mask quality using a multimodal large language model progressively optimizes the boundaries of the forged region, including: The evidence collection task is decomposed based on the current forgery clues using a multimodal large language model, and then the task is executed and segmentation prompt words are generated; The segmentation model is used to segment sub-regions based on segmentation prompts to generate the current mask; The multimodal large language model judges whether the current segmentation effect meets the standard from three aspects: whether the original segmentation target is completely covered, whether it contains all the forged regions mentioned in the evidence chain, and whether it contains normal regions that are clearly unrelated to the forged evidence. If the segmentation effect does not meet the standard, the iteration is terminated; otherwise, the intersection-union ratio (IUU) of the current mask and the previous mask is calculated. If the IUU is greater than the set threshold or the maximum number of iterations is reached, the iteration is terminated; otherwise, the next round of optimization is entered.

5. The method according to claim 4, characterized in that, The method utilizes a multimodal large language model to decompose the evidence collection task based on the current forgery clues, including thinking and task planning from four dimensions, transforming vague forgery suspicions into an executable, multi-angle verification task list centered on specific objects. The four dimensions include: edge / boundary integrity, noise / texture consistency, lighting and shadow rationality, and semantic logic.

6. A training-free image forgery localization system, characterized in that, The system comprising performing the method of any one of claims 1 to 5, wherein the system includes: The coarse-grained suspicious region exploration module is used to explore the suspicious regions in the image to be detected using a multimodal large language model, and obtain candidate regions; The fine-grained iterative refinement module is used to refine candidate regions through multiple rounds of collaborative optimization of a multimodal large language model and a segmentation model, thereby obtaining forged regions.

7. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the method of any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer, implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Tampered image detection method based on deep learning

    CN110349136A

  • Image forgery detection method and system based on multi-modal large model

    CN121074611A

  • Counterfeit image detection method based on visual big language model

    CN121861465A