Forgery image detection method and device, medium and product

By dividing the image into local regions and using a method that fuses region dependency maps and style features, the problem of strong dependence on forged samples and insufficient cross-domain detection capability in existing technologies is solved, and efficient forged image detection is achieved under conditions of few samples.

CN120953779AActive Publication Date: 2025-11-14NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP

Patent Information

Application Number
CN202511493310.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-14
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Existing technologies rely heavily on a large number of forged samples in image forgery detection, lack sensitivity to key local regions, and have weak generalization ability in cross-domain detection tasks, making them difficult to adapt to forgery detection scenarios with few samples and complex and diverse scenarios.

Method used

A method combining region dependency maps and style features is adopted to divide the input image into multiple local regions. Style features are extracted through local and global network models, and then fused using an attention model. Finally, a classifier is used for forgery detection.

Benefits of technology

It improves the detection capability of forged images under few sample conditions, enhances the sensitivity to local details and the generalization performance of the model, and adapts to the needs of multi-source and diverse forgery detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953779A_ABST
    Figure CN120953779A_ABST
Patent Text Reader

Abstract

The invention relates to the field of forged image detection, and provides a forged image detection method and device, a medium and a product, and the method comprises the steps: dividing an input image into a plurality of local regions, and constructing a region dependence graph of the local regions; extracting local style features from the region dependence graph by using a local network model; extracting global style features of the input image by using the global network model; performing attention fusion on the local style features and the global style features by using an attention model to obtain a fusion vector; performing image forgery detection on the fusion vector by using a classifier; wherein the local network model, the global network model, the attention model and the classifier are optimized through training. According to the invention, high identification accuracy and robustness can be maintained in detection scenes of scarce samples, local detail forgery and cross-domain forgery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of forged image detection, and more specifically, to a forged image detection method, device, medium, and product. Background Technology

[0002] With the rapid development of deep learning technology, especially the widespread application of Generative Adversarial Networks (GANs) in image generation, the quality of generated fake images has reached a level that is indistinguishable from reality. This type of technology is widely used in visual tasks such as face replacement, expression transfer, image restoration, and style transfer. While bringing convenience, it also poses a serious challenge to the determination of the authenticity of image content. Currently, fake images have the potential to be misleading and harmful in news media, social media platforms, and financial security. Therefore, conducting research on high-precision and robust fake image detection is of great significance.

[0003] Currently, mainstream image forgery detection methods can be broadly categorized into the following three types: (1) An end-to-end detection method based on a classifier directly classifies real and fake images by constructing a deep neural network model; (2) Based on feature extraction and anomaly detection, a certain level of image features are extracted first, and then forgery anomaly detection is performed; (3) Based on frequency domain analysis, the abnormal patterns of frequency domain signals in the forged image are captured by means of Fourier transform and other methods.

[0004] Although the above methods have improved detection performance and automation to some extent, they still have the following significant shortcomings in practical applications: (1) Strong dependence on a large number of labeled fake samples: Current methods usually rely on large-scale fake image datasets for supervised training. Their performance drops significantly in scenarios with small samples or scarce labels, making them difficult to apply to real-world environments with insufficient data or emerging fake methods. (2) Ignoring the sensitivity of local image regions: Most detection methods focus on the overall statistical distribution and structural information of the image, often ignoring the stylistic anomalies of local regions (such as edges, textures, etc.) in forged images, resulting in insufficient ability to identify forged details; (3) Insufficient generalization ability and poor cross-domain adaptability: When faced with unknown forgery methods or different data distributions (such as different image styles, sources or forgery tools), the model shows poor transfer ability and is difficult to handle complex and diverse forgery detection scenarios.

[0005] In summary, existing technologies still face significant bottlenecks in addressing insufficient sample size, local forgery identification, and cross-domain detection. Therefore, there is an urgent need to propose a novel forgery image detection method that can effectively model local region features under limited sample conditions and possesses good generalization ability, in order to meet the current demands for multi-source, diverse, and high-precision forgery detection. Summary of the Invention

[0006] To address the problems of existing technologies, such as strong reliance on a large number of forged samples, insufficient sensitivity to key local areas, and weak generalization ability in cross-domain detection tasks, this invention provides a forged image detection method, device, medium, and product that is targeted at local areas, integrates style features, and has the ability to detect images with few samples, thereby improving the ability to detect forged images.

[0007] In a first aspect, the present invention provides a method for detecting forged images, comprising: The input image is divided into multiple local regions, and a region dependency map of each local region is constructed. Local style features are extracted from region dependency graphs using a local network model; Global style features of the input image are extracted using a global network model; An attention model is used to fuse local and global style features to obtain a fused vector. Image forgery detection is performed on the fused vector using a classifier; The local network model, global network model, attention model, and classifier are optimized through training.

[0008] In a preferred embodiment, dividing the input image into multiple local regions and constructing a region dependency map of the local regions includes: A local keypoint detector is applied to the input image to obtain the coordinate information of local keypoints, thereby dividing the input image into multiple semantically related local regions; Construct a region dependency graph based on the divided local regions; the region dependency graph is represented as follows: , where the node set Each node in the set corresponds to a local region, and the edge set Each edge in the diagram represents a dependency or semantic relationship between the corresponding local regions.

[0009] In a preferred embodiment, the local network model is a first UNet-GRU network, which includes a connected UNet encoder and a GRU network; wherein, a separate first UNet-GRU network is used to extract local style features for each region dependency graph.

[0010] In a preferred embodiment, the global network model is a second UNet-GRU network, which includes a connected UNet encoder and a GRU network.

[0011] In a preferred embodiment, the step of using an attention model to fuse local style features and global style features to obtain a fusion vector includes: Input all local region features into the LSTM network to obtain the local region dependency features; For each local region, the dependency features of neighboring local regions are weighted according to the weighted fusion dependency of the region dependency graph to obtain the weighted dependency features of the local region. The global style features are fused with each weighted dependency feature to obtain the region fusion features; Based on region fusion features, an image-level fusion vector is generated through attention pooling.

[0012] In a preferred embodiment, the weighted fusion in the region dependency graph is a weighted fusion dependency based on spatial dependency and style dependency; The spatial dependency is a spatial dependency based on weighted Euclidean distance; The style dependency is a style dependency based on feature vector similarity.

[0013] In a preferred embodiment, the classifier is an MLP classifier, and the fused vector is input into the MLP classifier; the MLP classifier outputs the image forgery probability score and the classification result.

[0014] In a second aspect, the present invention provides an electronic device, comprising: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor, and the at least one processor executes the instructions stored in the memory to perform the above-described method.

[0015] Thirdly, the present invention provides a computer-readable storage medium for storing instructions that, when executed, enable the above-described method to be implemented.

[0016] Fourthly, the present invention provides a computer program product that, when invoked by a computer, causes the computer to execute the above-described method.

[0017] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. Reduce dependence on the number of fake samples This invention introduces a region dependency graph and style modeling mechanism, which enables the model to capture key forgery features even when there are limited forgery samples, effectively alleviating the problem of decreased detection performance of traditional methods in scenarios with few samples.

[0018] 2. Enhance sensitivity to forgery of local details. This invention significantly improves the model's ability to detect subtle forgeries (such as eye reflections and mouth edges) and enhances the discriminability of local regions by dividing an image into multiple semantically related local regions, extracting style features of each local region independently, and fusing them with the dependency relationships of the region dependency graph.

[0019] 3. Improve the model's discriminative ability and generalization performance. This invention introduces a style-aware attention mechanism, enabling the model to adaptively focus on local regions with a high probability of forgery and integrate global style context information, thereby improving the accuracy and robustness of forgery detection, especially showing better generalization ability when dealing with different forgery sources or cross-domain forged images.

[0020] Therefore, this invention is not only applicable to scenarios with insufficient samples, but can also realize the detection of forged images in complex forgery environments, which has important theoretical value and broad practical application prospects. Attached Figure Description

[0021] Figure 1 A flowchart of a forgery image detection method based on a few-sample local region sensitive style provided in an embodiment of the present invention.

[0022] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0024] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0025] Example To address the shortcomings of existing technologies, such as strong reliance on a large number of forged samples, insufficient sensitivity to key local regions, and weak generalization ability in cross-domain detection tasks, this invention provides a forged image detection method that maintains high recognition accuracy and robustness in scenarios involving scarce samples, local detail forgery, and cross-domain forgery detection. This method achieves forged image detection through steps including region segmentation, modeling a region dependency graph, extracting local and global style features, attention fusion, and discriminative classification.

[0026] In view of this, such as Figure 1 As shown, this embodiment of the invention provides a method for detecting forged images based on few-sample local region sensitive style, including the following steps: S100: Divide the input image into multiple local regions and construct a region dependency map of the local regions; S200 uses a local network model to extract local style features from the region dependency graph; S300 utilizes a global network model to extract global style features from the input image; S400 uses an attention model to fuse local and global style features to obtain a fusion vector; S500 uses a classifier to detect image forgery on the fused vector; The local network model, global network model, attention model, and classifier are optimized through training.

[0027] The following details the specific implementation of the aforementioned image forgery detection method.

[0028] S100: Divide the input image into multiple local regions and construct region dependency maps for each local region. In this embodiment of the invention, a local keypoint detector is applied to the input image to obtain the coordinate information of local keypoints, thereby dividing the input image into semantically relevant categories. Local area For example, a face image can be divided into local areas such as the left eye, right eye, nose tip, corner of mouth, forehead, and chin.

[0029] Subsequently, a Region Dependency Graph (RDG) is constructed based on the divided local regions. This Region Dependency Graph is represented as follows: , where the node set Each node in the set corresponds to a local region, and the edge set Each edge in the diagram represents a dependency or semantic correlation between corresponding local regions, which guides subsequent feature fusion and attention allocation.

[0030] To facilitate subsequent feature fusion and attention allocation, the following three dependencies are calculated for each pair of local regions: (1) Spatial dependence based on weighted Euclidean distance:

[0031] in, Indicates a local area and local areas Spatial dependency between them For the first A local area, For the first A local area, Indicates a local area and local areas The weighted Euclidean distance between them It is an exponential function with the natural constant e as its base. The spatial scale factor controls the effect of the weighted Euclidean distance on the weight decay.

[0032] The weighted Euclidean distance It can be represented as:

[0033] in, For local areas The weight, For local areas The weights can be adjusted based on the size, shape, or other important factors of the local region; For local areas coordinates For local areas The coordinates.

[0034] (2) Style dependency based on feature vector similarity:

[0035]

[0036] in, For local areas and local areas Style dependency between them For local areas and local areas cosine similarity, For local areas eigenvectors, For local areas eigenvectors, To indicate the calculation of the norm, It is an exponential function with the natural constant e as its base. The style scaling factor adjusts the effect of cosine similarity on the weights.

[0037] Local area eigenvectors Deep convolutional neural networks can be used to extract local regions Extracting from the local region, the deep convolutional neural network can be selected from networks such as ResNet (Residual Network) and VGG (Visual Geometry Group). Specifically, it extracts from the local region. The input is fed into a trained deep convolutional neural network. The feature map output by the deep convolutional neural network undergoes pooling or dimensionality reduction operations to generate a fixed-size feature vector, which represents the local region. eigenvectors .

[0038] (3) Weighted fusion dependency based on spatial and style dependencies:

[0039] in, Indicates a local area and local areas Weighted fusion dependency between them These represent the weighting coefficients of spatial dependency, controlling the proportion of spatial information in the final dependency. This represents the weighting coefficients for style dependence, controlling the proportion of style information.

[0040] In this step, the region dependency graph serves as a static prior structure. This invention divides the input image into several local regions with semantic or structural consistency and constructs a region dependency graph between these regions. This graph describes the spatial or stylistic relationships between local regions, clearly defining the relationships between them and helping to allocate appropriate attention during feature fusion, avoiding interference from redundant information. Furthermore, this region dependency graph modeling mechanism is flexibly applicable to various image types, effectively expressing the internal structural features of the image, effectively focusing on local details, and enhancing sensitivity to forged regions, making it particularly suitable for few-sample forged image detection scenarios. In addition, this region dependency graph will be used in subsequent attention fusion to guide the information flow and saliency allocation between local regions.

[0041] S200 utilizes a local network model to extract local style features from a region dependency graph. In this embodiment of the invention, the local network model employs a first UNet-GRU network, which includes a connected UNet (Convolutional Networks for Biomedical Image Segmentation) encoder and a GRU (Gated Recurrent Unit) network; wherein, for each region dependency map, an independent first UNet-GRU network is used to extract local style features. Specifically: The region dependency map is input into the UNet encoder to extract the first style features of the local regions, such as color and texture, which are represented as follows:

[0042] in, Indicates a local area stylistic features, This indicates the operation of the UNet encoder.

[0043] The first style feature is input into the GRU network to model the spatial dependencies within a local region, resulting in the second style feature of that local region, represented as:

[0044] in, For GRU network operation, For local areas The second style feature is the local style feature to be extracted in this step.

[0045] In this step, each local region in this invention has its local style features (such as color, texture, contrast, etc.) extracted by the UNet encoder, and the spatial dependencies between local regions are modeled by the GRU network to uncover potential forgery traces. This can significantly enhance the model's sensitivity to changes in image details and improve its ability to identify local forgeries.

[0046] S300 utilizes a global network model to extract global style features from the input image; In this embodiment of the invention, the global network model is a second UNet-GRU network, which includes a connected UNet encoder and a GRU network. Specifically: The entire input image is fed into the second UNet-GRU network to extract global style features of the input image, providing contextual information to assist in local region determination, represented as:

[0047] in, For the input image Global style features.

[0048] The second UNet-GRU network provides overall image style information and contextual constraints, implicitly containing global dependencies between local regions. This provides an overall reference for local feature fusion, ensuring consistency between local and global styles. Therefore, this invention introduces overall image style encoding as a global style feature. By fusing this global style feature with various local style features, it effectively captures inconsistencies at the image style level, enhancing the model's ability to detect complex image attacks such as style transfer forgery and fusion forgery. The global style feature will participate in the subsequent fusion process along with local style features to further enhance the ability to determine style consistency.

[0049] S400 uses an attention model to fuse local and global style features to obtain a fusion vector; In this embodiment of the invention, the base network of the attention model LSTM-Attention uses an LSTM network (Long Short-Term Memory), and a weighted fusion dependency from the region dependency graph is introduced to perform attention-based weight adjustment. Specifically: S401, all local region features Inputting the data into an LSTM network models the contextual relationships between local regions, yielding the dependency features of these local regions, represented as follows:

[0050] in, For the first Local area Dependency characteristics, n For the number of features, For LSTM network operation.

[0051] S402, for each local area According to the weighted fusion dependency of the region dependency graph We weight the dependency features of neighboring local regions to obtain the local regions. Weighted dependency features:

[0052] in, For local areas The weighted dependency feature, For local areas Within the neighborhood of j Dependency features of local neighbor regions M For local areas The number of neighboring local regions within a given neighborhood is used to weight its neighborhood dependency features using structured attention, highlighting important regions and suppressing irrelevant information; thus, the weighted dependency features are... It not only includes its own information, but also incorporates the surrounding context, forming a more discriminative feature representation.

[0053] S403 will include global style features With each weighted dependency feature Integration, resulting in regional integration characteristics , represented as:

[0054] in, Represents a linear projection matrix that makes the global style features and Dimensions are consistent.

[0055] S404, based on regional fusion characteristics An image-level fusion vector is generated through attention pooling, as follows:

[0056] in, For the fusion vector, For activation function, Indicates regional integration characteristics Attention weights It is a learnable parameter vector.

[0057] In the feature fusion process, this invention uses an LSTM network to model the contextual relationships between regions, combines a region dependency graph to limit the distribution of attention, helps identify the salience of fake regions, and performs structured attention fusion. This can automatically identify the salience of fake regions, assign higher weights to key regions, thereby highlighting key regions, suppressing background redundancy interference, and improving the discrimination effect on local fake regions.

[0058] S500 uses a classifier to detect image forgery on the fused vector; In this embodiment of the invention, the classifier is an MLP (Multi-Layer Perceptron) classifier. The fused vector... Input an MLP classifier; the MLP classifier outputs an image forgery probability score and classification result.

[0059] The probability fraction is expressed as:

[0060] in, , representing the probability score that the input image is real. Or the probability score of the input image being fake. , Represents the weight parameters of the MLP classifier. This represents the bias vector.

[0061] The final classification result is expressed as follows:

[0062] Among them, the final classification results , This represents the result corresponding to the maximum probability score, i.e.: If probability fraction Greater than probability fraction The final classification result is This indicates that the input image is real; If probability fraction Less than probability fraction The final classification result is This indicates that the input image is fake.

[0063] S600 can be achieved by repeatedly executing steps S100 to S500 to train and optimize the local network model, global network model, attention model, and classifier. During the training phase, a weighted cross-entropy loss function can be used, combined with optimization techniques such as region regularization and fake sample equalization, to improve the model's stability and generalization ability under limited sample conditions, thereby adapting to the needs of multi-source, diverse, and high-precision fake image detection. The training data can include multiple types of fake images and real images to enhance the model's ability to recognize unknown fake styles.

[0064] Based on the same technical concept, embodiments of the present invention also provide an electronic device that can implement the forged image detection method flow provided in the above embodiments of the present invention. In one embodiment, the electronic device may be a server, a terminal device, or other electronic device. Figure 2 As shown, the electronic device may include: At least one processor and a memory connected to the at least one processor. In this embodiment of the invention, the specific connection medium between the processor and the memory is not limited. Figure 2 The example used is the connection between the processor and memory via a bus. The bus... Figure 2 The connections between other components are indicated by thick lines and are for illustrative purposes only, not as limiting information. Buses can be divided into address buses, data buses, control buses, etc., but for ease of representation, [the specific bus type is not shown here]. Figure 2 The processor is represented by a single thick line, but this does not imply that there is only one bus or one type of bus. Alternatively, a processor can also be called a controller; there are no restrictions on the name.

[0065] In this embodiment of the invention, the memory stores instructions that can be executed by at least one processor. By executing the instructions stored in the memory, at least one processor can execute a forged image detection method described above.

[0066] The processor is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory and calling data stored in memory, it can monitor the device's various functions and process data, thereby enabling overall monitoring of the device.

[0067] In an alternative design, the processor may include one or more processing units. The processor may integrate an application processor and a modem processor, wherein the application processor primarily handles the operating system, user interface, and applications, while the modem processor primarily handles wireless communication. It is understood that the modem processor may also not be integrated into the processor. In some embodiments, the processor and memory may be implemented on the same chip; in some embodiments, they may also be implemented separately on separate chips.

[0068] The processor can be a general-purpose processor, such as a CPU, digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the forged image detection method disclosed in the embodiments of this invention can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0069] Memory, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory can include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, etc. Memory is any other medium capable of carrying or storing desired program code in the form of instructions or data structures, and accessible by a computer, but is not limited thereto. In embodiments of the present invention, memory can also be a circuit or any other device capable of implementing storage functions, used to store program instructions and / or data.

[0070] By designing and programming the processor, the code corresponding to the forged image detection method described in the foregoing embodiments can be embedded into the chip, enabling the chip to execute the steps of the method described in the foregoing embodiments during operation. How to design and program the processor is a technique well-known to those skilled in the art and will not be elaborated upon here.

[0071] Based on the same inventive concept, embodiments of the present invention also provide a storage medium storing computer instructions that, when executed on a computer, cause the computer to perform a forged image detection method described above.

[0072] In some alternative embodiments, the present invention also provides that various aspects of a forged image detection method can also be implemented as a program product comprising program code that, when the program product is run on a device, causes the control device to perform the steps in a forged image detection method according to various exemplary embodiments of the present invention as described above.

[0073] It should be noted that although several units or sub-units of the apparatus have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the invention, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units. Furthermore, although the operation of the method of the invention is described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0074] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0075] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a server, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0076] Program code for performing the operations of this invention can be written using any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0077] In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0078] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0079] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0080] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for detecting forged images, characterized in that, include: The input image is divided into multiple local regions, and a region dependency map of each local region is constructed. Local style features are extracted from region dependency graphs using a local network model; Global style features of the input image are extracted using a global network model; An attention model is used to fuse local and global style features to obtain a fused vector. Image forgery detection is performed on the fused vector using a classifier; The local network model, global network model, attention model, and classifier are optimized through training.

2. The method for detecting forged images according to claim 1, characterized in that, The step of dividing the input image into multiple local regions and constructing a region dependency map of the local regions includes: A local keypoint detector is applied to the input image to obtain the coordinate information of local keypoints, thereby dividing the input image into multiple semantically related local regions; Construct a region dependency graph based on the divided local regions; the region dependency graph is represented as follows: , where the node set Each node in the set corresponds to a local region, and the edge set Each edge in the diagram represents a dependency or semantic relationship between the corresponding local regions.

3. The forged image detection method according to claim 1, characterized in that, The local network model is a first UNet-GRU network, which includes a connected UNet encoder and a GRU network; wherein, for each region dependency graph, an independent first UNet-GRU network is used to extract local style features.

4. The forged image detection method according to claim 1, characterized in that, The global network model is a second UNet-GRU network, which includes a connected UNet encoder and a GRU network.

5. The method for detecting forged images according to claim 1, characterized in that, The method of using an attention model to fuse local and global style features to obtain a fusion vector includes: Input all local region features into the LSTM network to obtain the local region dependency features; For each local region, the dependency features of neighboring local regions are weighted according to the weighted fusion dependency of the region dependency graph to obtain the weighted dependency features of the local region. The global style features are fused with each weighted dependency feature to obtain the region fusion features; Based on region fusion features, an image-level fusion vector is generated through attention pooling.

6. The method for detecting forged images according to claim 5, characterized in that, The weighted fusion in the region dependency graph is a weighted fusion dependency based on spatial dependency and style dependency; The spatial dependency is a spatial dependency based on weighted Euclidean distance; The style dependency is a style dependency based on feature vector similarity.

7. The method for detecting forged images according to claim 1, characterized in that, The classifier uses an MLP classifier, and the fused vector is input into the MLP classifier; the MLP classifier outputs the image forgery probability score and classification result.

8. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which executes the instructions stored in the memory to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store instructions that, when executed, cause the method as described in any one of claims 1-7 to be implemented.

10. A computer program product, characterized in that, When the computer program product is invoked by a computer, it causes the computer to perform the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Double-branch image restoration forgery detection method, system and device and storage medium

    CN113744153A

  • Counterfeit face detection method, device, equipment and medium

    CN120014717A

  • Deep counterfeit image detection method based on knowledge distillation and domain adversarial training

    CN120808126A

Cited By

  • Intelligent generated image detection method based on multi-granularity artifact feature fusion

    CN121147726A

  • Image forgery detection method based on multi-modal large language model

    CN121236571A