An Image Forgery Detection Method and System Based on Multi-Expert Model Decision Making

Through an image forgery detection method based on multi-expert model decision-making, combining the gating mechanism and multiple expert models, multi-dimensional detection and weighted fusion decisions are realized, which solves the problem of insufficient detection robustness and accuracy in the existing technology, and significantly improves detection accuracy and efficiency.

CN119762958BActive Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510274673.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The existing image forgery detection methods show low robustness and accuracy when facing complex and diverse forgery methods, especially in the case of unknown forgery technology, image compression and noise addition, the detection performance has significantly decreased.

Method used

An image forgery detection method based on multi-expert model decision is adopted. Through a gating mechanism, multiple expert models are combined, each model focuses on different features or forgery traces of the image, and multi-dimensional detection is realized, and the results are output comprehensively through a weighted fusion decision mechanism.

Benefits of technology

It improves the accuracy and efficiency of image forgery detection, can effectively identify complex forgery content, improves the detection ability of unknown forgery algorithms, and enhances the generalization ability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762958B_ABST
    Figure CN119762958B_ABST
Patent Text Reader

Abstract

The present invention discloses an image forgery detection method and system based on multi-expert model decision-making, belonging to the technical field of image forgery detection. Obtain normal images in multiple fields, generate forged images based on different forgery types and degrees of forgery and group them, and use the normal images and each group of forged images to pre-train a pre-screening model and a multi-expert model capable of predicting the forgery probability, where each expert model focuses on different features or forgery traces of the images; use the normal images and forged images to train an image forgery detection network including an image encoding module, a gating calculation module, a pre-screening model, and a multi-expert model; during training, the parameters of the pre-screening model and the multi-expert model are frozen; based on the trained image forgery detection network, identify whether the input image is forged. The present invention can achieve multi-dimensional detection of forged images, comprehensively output the results through a weighted fusion decision-making mechanism, make full use of the advantages of different models, and improve the detection accuracy and efficiency of forged images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image forgery detection, and particularly to an image forgery detection method and system based on multi-expert model decision-making. Background Art

[0002] With the development of generative artificial intelligence technology, the application of deepfake content in the fields of images and videos has become increasingly widespread, and the forgery means have become more complex. Traditional forgery content detection methods show low robustness and accuracy when facing increasingly complex and diverse forgery means.

[0003] Specifically, existing methods usually rely on a single detection model and are difficult to comprehensively identify different forgery features. Especially in complex scenarios, the model often tends to learn simple and obvious forgery features while ignoring the subtle forgery traces left by complex generative models. This limitation makes traditional methods prone to false positives or false negatives when detecting unknown or complex forgery content.

[0004] Existing multi-task learning methods usually adopt the ways of task sharing and task branching for forgery detection. By designing a shared underlying feature extraction module and setting up independent task branches for each sub-task to parallelly process different forgery detection tasks. The specific steps are as follows: First, all tasks share an underlying feature extraction network to extract the general features of images or videos; Second, each task processes specific forgery features through a dedicated branch network, such as the details of forged images, the localization of forged regions, and the identification of global forgery patterns, etc.; Third, by designing a fusion loss function, the losses of multiple tasks are weighted and fused to jointly optimize the parameters of the network, so that each task can be fully learned in the whole network; Finally, fusion and optimization. Although this method can significantly improve the detection performance, there may be information conflicts and interferences between tasks, and the optimization process is relatively complex. There are also the following problems:

[0005] 1) Insufficient generalization ability: The generalization ability of the detection model is insufficient when facing unknown forgery techniques, and it is difficult to cope with complex and changeable forgery means in practical applications.

[0006] 2) Insufficient robustness: The existing forgery detection technologies show a significant decrease in detection accuracy when facing common perturbations such as image compression and noise addition, and show poor robustness.

[0007] 3) Insufficient interpretability: Most traditional forgery detection methods rely on complex deep learning models and lack a clear explanation of the basis for model judgment. Summary of the Invention

[0008] In view of the above technical problems, the present invention proposes an image forgery detection method and system based on multi-expert model decision-making. By combining multiple expert models through a gating mechanism, each model focuses on different features or forgery traces of the image, realizing multi-dimensional detection, and comprehensively outputting the results through a weighted fusion decision-making mechanism, improving the detection accuracy and efficiency, and solving the problems of traditional methods in unknown algorithm detection, compression distortion, and forgery traceability.

[0009] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0010] In the first aspect, the present invention provides an image forgery detection method based on multi-expert model decision-making, including the following steps:

[0011] (1) Obtain normal images in multiple fields, generate forged images based on different forgery types and degrees of forgery, and group them. Use the normal images and each group of forged images to pre-train a pre-screening model and a multi-expert model capable of predicting the forgery probability. The multi-expert model includes a typical forgery trace detection model based on the EfficientNet network, a complex refinement trace detection model based on the ViT network, and a local forgery trace detection model based on the Xception network;

[0012] (2) Use the normal images and forged images to train an image forgery detection network including an image encoding module, a gating calculation module, a pre-screening model, and a multi-expert model; during training, the parameters of the pre-screening model and the multi-expert model are frozen;

[0013] (3) Based on the trained image forgery detection network, identify whether the input image is forged, including:

[0014] Use the image encoding module to extract image features; the pre-screening model takes the image features as input and identifies whether it is a forged image. If so, the judgment ends; if not, the gating calculation module takes the image features as input and generates a probability vector with the same dimension as the number of multi-expert models as the weight values of each expert model; screen the expert models with non-zero weight values and independently judge the forgery probability with the image features as input. Based on the non-zero weight values, fuse the forgery probability values output by each expert model, and identify the input image with a total forgery probability value higher than the threshold as a forged image.

[0015] Further, it also includes an image preprocessing process, and the preprocessing includes normalization, denoising, and enhancement.

[0016] Further, in step (1), the generation of forged images based on different forgery types and degrees of forgery includes:

[0017] The forgery types include texture forgery, boundary forgery, brightness forgery, facial feature forgery, expression forgery, illumination forgery, and synthetic splicing forgery;

[0018] The degree of forgery mentioned refers to the total degree of perturbation of a normal image under different types of forgery perturbations.

[0019] Furthermore, in step (1), the method for grouping forged images includes:

[0020] (1-1) Obtain the type of forgery and the degree of forgery corresponding to the forged image. If one or more of texture forgery, boundary forgery, brightness forgery, and synthetic stitching forgery are included in the type of forgery corresponding to the forged image, and the degree of forgery is higher than the threshold, then it is classified into the first group of forged images; otherwise, execute step (1-2);

[0021] (1-2) If one or more of texture forgery, boundary forgery, brightness forgery, and synthetic stitching forgery are included in the type of forgery corresponding to the forged image, and the degree of forgery is not higher than the threshold; or if the type of forgery is selected from any one of facial feature forgery, expression forgery, and lighting forgery, and the degree of forgery is higher than the threshold; then it is classified into the second group of forged images; otherwise, execute step (1-3);

[0022] (1-3) If the type of forgery corresponding to the forged image is selected from facial feature forgery, expression forgery, and lighting forgery, and the degree of forgery is not higher than the threshold; then it is classified into the third group of forged images;

[0023] (1-4) Select the images with local forged areas from the third group of forged images, and the selected images form the fourth group of forged images.

[0024] Furthermore, in step (1), the combination of the first group of forged images and several normal images is used as the training set for the pre-screening model, the combination of the second group of forged images and several normal images is used as the training set for the EfficientNet network, the combination of the third group of forged images and several normal images is used as the training set for the ViT network, and the combination of the fourth group of forged images and several normal images is used as the training set for the Xception network.

[0025] Furthermore, the pre-screening model uses a model based on the CNN network, and the gating calculation module uses a fully connected network.

[0026] Furthermore, it also includes the visualization step of the image forgery detection process, including:

[0027] Display the heat map corresponding to the input image;

[0028] And display the forgery probability, weight assignment, total forgery probability, and final recognition result independently judged by each expert model.

[0029] In a second aspect, the present invention provides an image forgery detection system based on multi-expert model decision-making for implementing the above-mentioned image forgery detection method based on multi-expert model decision-making.

[0030] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which when executed by a processor, implements the above-mentioned image forgery detection method based on multi-expert model decision-making.

[0031] In a fourth aspect, the present invention provides a computer electronic device, including a memory and a processor;

[0032] The memory is used for storing a computer program;

[0033] The processor is used for implementing the above-mentioned image forgery detection method based on multi-expert model decision-making when executing the computer program.

[0034] The beneficial effects of the present invention are as follows:

[0035] The present invention establishes a multi-expert fusion decision-making mechanism, dynamically selects an appropriate expert model for judgment based on the input image features. Each expert model focuses on different data features or forgery traces of the image. The multi-dimensional analysis based on the expert models can give full play to the advantages of each expert model, improve the recognition ability for different image forgery means, and is especially suitable for complex forgery scenarios. The present invention can not only identify known forgery methods but also cope with unknown forgery algorithms, effectively improving the generalization ability of the detection model and reducing the failure risk. In addition, a pre-screening model is used to preliminarily screen the input image to identify simple forged images, which not only meets the high real-time requirements for simple images but also meets the high-precision requirements for complex images, ensuring the accuracy and efficiency of image forgery detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a schematic framework diagram of the image forgery detection method based on multi-expert model decision-making shown in an embodiment of the present invention.

[0037] Figure 2 is a flowchart of the image forgery detection method based on multi-expert model decision-making shown in an embodiment of the present invention.

[0038] Figure 3 is a schematic diagram of the image forgery detection system based on multi-expert model decision-making shown in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The present invention will be further described and explained below in conjunction with the specific embodiments. The embodiments are only examples of the present disclosure and do not delimit the scope of limitation. The technical features of each embodiment of the present invention can be combined correspondingly without conflict.

[0040] The accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0041] The flowcharts shown in the accompanying drawings are only exemplary illustrations and do not necessarily include all steps. For example, some steps can be decomposed, while some steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0042] Figure 1 The following shows a schematic framework diagram of an image forgery detection method based on multi-expert model decision proposed by the present invention. A pre-screening model and a multi-expert model are introduced. The pre-screening model realizes the rapid identification of obvious forgeries. Each expert model focuses on different data features or forgery traces of the image. For the image features that the pre-screening model fails to identify as forgeries, the gating mechanism further calculates weights and sends them to each expert model with non-zero weights for independent detection. Finally, the detection results of each expert model are fused based on the weight values and output. At the same time, multi-dimensional output visualization results are provided to enhance the interpretability and transparency of the image forgery detection process.

[0043] Specifically, as Figure 2 shown, an image forgery detection method based on multi-expert model decision proposed by the present invention mainly includes the following steps:

[0044] S1, Obtain normal images in multiple fields, generate forged images based on different forgery types and degrees of forgery, and group them.

[0045] In this embodiment, the normal images are selected from public data sets such as ImageNet, COCO, FFHQ, etc., covering multi-field images such as natural scenes, faces, objects, etc. The normal images are pre-processed, such as normalization, denoising, and enhancement. For example: unify the image resolution (such as 256x256) and perform normalization processing on the images to remove low-quality image samples; remove unnecessary noise in the images by filtering, wavelet transform, or Fourier transform; perform image enhancement on the images, such as rotation, scaling, cropping, etc., to simulate the forgery content transformation that may occur in the real world, so as to improve the robustness of the model to different inputs and enable the system to maintain high accuracy and reliability in the face of various interferences and perturbations in the real environment.

[0046] For the pre-processed normal images, perturbations of different forgery types and degrees of forgery are performed to generate a large number of forged images.

[0047] In a specific implementation of the present invention, the forgery types include, but are not limited to, texture forgery, boundary forgery, brightness forgery, facial feature forgery, expression forgery, lighting forgery, and synthetic splicing forgery; the degree of forgery refers to the total perturbation degree of a normal image under the perturbation of different forgery types.

[0048] Among them, texture forgery refers to modifying local textures (such as adding noise, blurring), which may cause the textures in local areas to be inconsistent with the surroundings (such as being too smooth or repetitive). For example, the texture in the forged area may not match the real background. Boundary forgery refers to adding unnatural transitions at the image boundaries, which may cause unnatural jaggedness at the boundaries. Brightness forgery refers to adjusting local brightness, and the brightness in the forged area may be inconsistent with the surrounding areas. For example, the local brightness may be too high or too low, appearing unnatural. Facial feature forgery refers to forging by modifying facial features (such as eyes, nose, mouth), such as technologies like DeepFake, which may be difficult to distinguish with the naked eye. Expression forgery refers to modifying facial expressions (such as changing a smile to a serious expression), which may look natural, especially for minor modifications. For example, expression forgery generated by GAN may leave no obvious traces. Lighting forgery refers to adjusting the lighting of the picture. If processed properly, it will look very natural, especially when the forged lighting direction is consistent with the scene, and it is difficult for the naked eye to detect abnormalities. Synthetic splicing forgery refers to splicing different images together, which may result in inconsistencies in lighting, color, or perspective. For example, the lighting direction of the spliced person and the background may be inconsistent. These forgery types can be implemented using image processing tools or deep learning models. Different degrees of forged images can be generated by using one or integrating multiple forgery types. Here, the degree of forgery of multiple forgery types, that is, the total perturbation degree, refers to the mean value of the intensities of each forgery type. The definition of the perturbation degree of an image is usually based on the degree of difference between the image and the original image. Image quality assessment metrics (such as PSNR, SSIM) can be used to measure the difference between the forged image and the original image, so as to evaluate the degree of forgery of each forgery type.

[0049] For the generated forged images, a detailed data index table is established to record their corresponding forgery types and degrees of forgery, which is convenient for subsequent grouping and research and analysis of forged images in different groups, and provides support for training models focusing on different data features or forgery traces of images.

[0050] Taking the above forgery types as examples, texture forgery, boundary forgery, brightness forgery, and synthetic splicing forgery are easy to identify, while facial feature forgery, expression forgery, and lighting forgery are not easy to identify. Combining the degrees of forgery of each forgery type, a large number of the above-generated forged images are divided into four categories.

[0051] In a specific implementation of the present invention, the method for grouping forged images includes:

[0052] S1-1. Obtain the forgery type and forgery degree corresponding to the forged image. If one or more of texture forgery, boundary forgery, brightness forgery, and synthetic splicing forgery are included in the forgery type corresponding to the forged image, and the forgery degree is higher than the threshold, then classify it into the first forged image group; otherwise, execute S1-2;

[0053] S1-2. If one or more of texture forgery, boundary forgery, brightness forgery, and synthetic splicing forgery are included in the forgery type corresponding to the forged image, and the forgery degree is not higher than the threshold; or the forgery type is selected from any one of facial feature forgery, expression forgery, and illumination forgery, and the forgery degree is higher than the threshold; then classify it into the second forged image group; otherwise, execute S1-3;

[0054] S1-3. If the forgery type corresponding to the forged image is selected from facial feature forgery, expression forgery, and illumination forgery, and the forgery degree is not higher than the threshold; then classify it into the third forged image group;

[0055] S1-4. Select the images with local forged areas from the third forged image group, and the selected images form the fourth forged image group.

[0056] In practical applications, the data characteristics and forgery traces presented by images with different forgery types and degrees are significantly different. Simple image forgery may only involve basic image editing operations, with relatively single changes in data characteristics and obvious forgery traces; while complex image forgery may comprehensively use a variety of advanced technologies to deeply process the image, making the data characteristics highly complex and the forgery traces difficult to detect. If a single model is used to process images of all difficulties, the model is easily interfered by images of different difficulties during the learning process, making it difficult to accurately capture the key features of various images, resulting in a decline in the model's generalization ability and accuracy. By classifying forged images into four categories according to the forgery degree of the forgery type and using them to train four models respectively, the aim is to let each model focus on learning the unique data characteristics or forgery traces of images in a specific difficulty range. In this way, each model can be deeply optimized for images of a specific difficulty, avoiding interference between images of different difficulties, so as to more effectively extract and learn the key information of images corresponding to the difficulty, and improve the model's recognition ability and accuracy for images of different difficulties.

[0057] S2. Use normal images and forged images of each group to pre-train a pre-screening model and a multi-expert model that can predict the forgery probability.

[0058] The present invention designs and deploys multiple expert models, each of which is optimized for different types of forgery features or forgery algorithms, ensuring that the system can make comprehensive and accurate judgments when faced with various forged contents. The construction of expert models not only involves the selection of different model architectures (such as EfficientNet, ViT, and Xception), but also includes how to effectively configure and train these models in practical applications to ensure that they can work complementarily and synergistically.

[0059] S2-1, design of the pre-screening model.

[0060] The pre-screening model only needs to meet the purpose of initially screening simple forged images. For example, the MobileNet network based on CNN is a lightweight deep learning model, which can provide efficient image recognition and processing capabilities for resource-constrained environments; this model is trained using a combination of the first group of forged images and the same number of randomly sampled normal images as the training set, and can quickly extract high-level features and conduct preliminary classification of forged images and normal images.

[0061] Through simple preliminary judgment, the system can effectively avoid excessive processing of real data, thereby improving efficiency.

[0062] S2-2, design of the multi-expert model.

[0063] The multi-expert model should focus on different types of forgery features, and adopt a typical forgery trace detection model based on the EfficientNet network, a complex refined trace detection model based on the ViT network, and a local forgery trace detection model based on the Xception network. Multiple independent expert models detect specific forged contents, and each expert model outputs a confidence level.

[0064] In addition, each expert model adopts a hard sample replay strategy. By preferentially replaying and retraining the hard-to-detect samples in the model, it ensures that the system can effectively adapt to these new types of forgery features, thereby continuously improving the detection ability. This strategy enables the expert model to continuously adjust and improve its detection ability, maintain efficient and accurate recognition of forged contents, and ensure the long-term effectiveness and robustness of the system.

[0065] The following describes the three expert models in this embodiment.

[0066] (1) A typical forgery trace detection model based on the EfficientNet network.

[0067] The core objective of this model is to identify some common low-level features in forged images, such as unnatural boundaries introduced by common deepfake techniques, inconsistent brightness, fragmented textures, and inconsistencies in lighting, color, or perspective caused by synthetic stitching. Additionally, forgeries generated based on conventional deep learning models (such as GAN networks) generally involve forgeries of facial features, expression features, lighting features, etc. The model outputs detection results by analyzing the typical forgery features introduced by conventional forgery techniques in the image.

[0068] The typical forgery trace detection model in the present invention adopts the EfficientNet network, which proposes a compound scaling method to coordinately adjust the depth width and resolution to improve the model performance:

[0069]

[0070]

[0071]

[0072] where , , are scaling factors that satisfy , is the scaling level. By balancing the scaling factors, the EfficientNet network can significantly improve the representation ability of the network model while maintaining a low computational cost.

[0073] The typical forgery trace detection model takes lightweight and high efficiency as the core design, adopts a compound scaling strategy, and optimizes the computational efficiency. This model is trained using a combination of the second group of forged images and the same number of randomly sampled normal images as the training set. This expert model can maintain a high detection accuracy in large-scale forgery data detection while reducing the computational cost.

[0074] (2) Complex refinement trace detection model based on the ViT network.

[0075] The complex refinement data detection model in the present invention mainly detects images with more concealed forgery traces, generally including forgery of facial features, expression forgery, lighting forgery, etc. High-precision forgery models use technologies such as image enhancement, noise removal, and color correction to adjust lighting, color, texture, etc. There are no obvious stitching traces, and it is more natural and imperceptible visually. The texture, color, and lighting of the forged image are made more delicate.

[0076] In the present invention, Vision Transformer (ViT) is used to construct a complex refined data detection model, and the self-attention mechanism (Self-Attention) is used to model the global features of images. ViT is an image processing method based on the Transformer architecture. Specifically, the input image is divided into several small patches, each patch having a size of and the total size of the image being so that the image is divided into a total of small patches. Each patch is flattened and mapped to an embedding space of dimension D through a linear embedding:

[0077]

[0078] where is the th image patch, is the embedding matrix, and is the embedding vector of this patch. To preserve the spatial information in the image, ViT adds a positional encoding to the embedding vector to obtain the final input sequence:

[0079]

[0080] where is the embedding vector of the th image patch.

[0081] Finally, the similarity between each image patch is calculated through the self-attention mechanism and weighted and summed to obtain the global context information:

[0082]

[0083] where is the query vector, is the key vector, is the value vector, and is the dimension of the key vector. In this way, ViT can effectively capture long-range dependencies, focus more on capturing the global features of images or videos, and then detect complex refined forgery content.

[0084] This model is trained using a combination of the third group of forged images and the same number of randomly sampled normal images as the training set.

[0085] (3) Local forgery trace detection model based on the Xception network.

[0086] The local forgery trace detection model in the present invention focuses on the forgery traces in the local areas of the image, such as forgery in areas like the face, eyes, mouth, etc., aiming to identify those tiny forgery features that are difficult to detect, such as local facial feature forgery, expression forgery, lighting forgery, etc.

[0087] The local forgery trace detection model in the present invention uses the Xception network as the core backbone network.

[0088] Different from traditional convolutional neural networks, Xception adopts a depthwise separable convolution structure, which splits the standard convolution operation into two stages: one is depthwise convolution, which only performs convolution on each input channel; the other is pointwise convolution, which performs convolution across channels. Such a decomposition can significantly reduce the computational complexity and improve the learning ability for local features. Local forgery traces are usually small-area changes in the image, and depthwise separable convolution can effectively capture these changes. This design reduces the number of parameters, improves the computational efficiency, and at the same time retains the high sensitivity to local features, especially performing well when dealing with small-scale forgery features. The expert model is trained using a combination of the fourth group of forged images and the same number of randomly sampled normal images as the training set, and can provide high accuracy in local forgery detection.

[0089] S3. Train an image forgery detection network including an image encoding module, a gating calculation module, a pre-screening model, and a multi-expert model using normal images and forged images; during training, the parameters of the pre-screening model and the multi-expert model are frozen.

[0090] In a specific implementation of the present invention, the image encoding module can be implemented using VGG16 or ResNet50, which is used to extract image features with the image as the input. The gating calculation module uses a fully connected network, which takes the image features as the input, generates a probability vector with the same dimension as the number of multi-expert models as the weight values of each expert model, screens the expert models with non-zero weight values and independently judges the forgery probability with the image features as the input, fuses the forgery probability values output by each expert model based on the non-zero weight values, and identifies the input image with the total forgery probability value higher than the threshold as a forged image. Due to the diversity of forgery content, generally, the weight values of each expert model are non-zero results.

[0091] Assume the above-mentioned expert models 、 、 The probability values output by the forgery detection are respectively 、 、 ,and the corresponding weights obtained by gating calculation are 、 、 ,then the system obtains the final detection result through weighted fusion :

[0092]

[0093] The gating calculation module flexibly adjusts the weights of each expert model according to the input image features, ensuring that the appropriate expert model is selected according to the input image features to analyze the characteristics of the data, and automatically assigns the highest weight to the expert model that can most accurately process the current forged content. The present invention generates a final forged determination result by weighted fusion of the detection results of multiple expert models, balancing the contributions of different expert models and handling the problem of inconsistent judgments between models.

[0094] S4. Based on the trained image forgery detection network, identify whether the input image is forged.

[0095] The working process is as follows: The input image uses the image encoding module to extract image features, and then sends them to the pre-screening model. A simple detection algorithm is used to quickly screen out obvious forged data and identify whether it is a forged image. If so, the judgment ends; if not, the gating calculation module takes the image features as input and generates a probability vector with the same dimension as the number of expert models as the weight value of each expert model, and distributes the image features to the appropriate expert models through the gating mechanism. Each expert model in the expert model module extracts features and performs forgery detection on the input image features, independently judges the forgery probability, and fuses the forgery probability values output by each expert model based on non-zero weight values. An input image with a total forgery probability value higher than the threshold is identified as a forged image.

[0096] In a specific implementation of the present invention, a visualization output module can also be introduced into the overall process. The comprehensive obtained results are sent to the visualization output module to display the confidence level of the independent judgment of each expert model, the weight allocation, and the final fusion decision-making process, and provide visualization results such as forged area localization. In this embodiment, methods such as Grad-CAM are used to generate heatmaps to visualize the importance or significance of certain regions in the image. The forged regions of the image usually show abnormalities in the heatmap, for example: lower significance or abnormal significance distribution, and the heatmap distribution of the forged region may be inconsistent with other regions. By presenting the final detection results to the user through a visualization interface, outputs in the form of forged area localization, detection reports, etc. are provided, and at the same time, the interpretability of the decision-making chain is presented to help the user understand the decision-making process.

[0097] Each model in the present invention focuses on images of specific difficulty levels, capable of delving deeper into and learning the unique data features and forgery traces of corresponding images. Compared with a single model learning all images, the feature extraction is more accurate. Training models for different difficulty levels of images can effectively reduce interference during the model learning process, optimize the training effect of the models, and thus significantly improve the performance of the models in image recognition tasks of their respective difficulty levels. The good performance of each model on images of specific difficulty levels enables the entire system to have stronger generalization ability when facing images of different difficulties, and be able to more accurately identify and process various complex actual scenarios. The present invention has high flexibility and scalability. If new difficulty level images or forgery techniques emerge subsequently, the system's functions can be easily expanded by adding new model categories and training data to adapt to the ever-changing requirements.

[0098] Based on the same inventive concept, in this embodiment, an image forgery detection system based on multi-expert model decision-making is further provided, which is used to implement the above embodiment. The terms "module", "unit", etc. used hereinafter can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.

[0099] An image forgery detection system based on multi-expert model decision-making, as Figure 3 shown, includes:

[0100] An image sample generation module, which is used to obtain normal images in multiple fields, generate forged images based on different forgery types and degrees of forgery, and group them;

[0101] A primary training module, which is used to pre-train a pre-screening model and a multi-expert model capable of predicting the forgery probability by using normal images and each group of forged images. The multi-expert model includes a typical forgery trace detection model based on the EfficientNet network, a complex refinement trace detection model based on the ViT network, and a local forgery trace detection model based on the Xception network;

[0102] A secondary training module, which is used to train an image forgery detection network including an image encoding module, a gating calculation module, a pre-screening model, and a multi-expert model by using normal images and forged images; during training, the parameters of the pre-screening model and the multi-expert model are frozen;

[0103] An image forgery recognition module, which is used to recognize whether an input image is forged based on a trained image forgery detection network, includes: extracting image features using an image encoding module; the pre-screening model takes the image features as input to recognize whether it is a forged image. If so, the judgment ends; if not, the gating calculation module takes the image features as input to generate a probability vector with the same dimension as the number of multi-expert models as the weight values of each expert model; screening out the expert models with non-zero weight values and independently judging the forgery probability with the image features as input, fusing the forgery probability values output by each expert model based on the non-zero weight values, and recognizing the input image with a total forgery probability value higher than the threshold as a forged image.

[0104] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be elaborated here. The system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative work.

[0105] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The system embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.

[0106] In addition, it should be noted that the above-mentioned method for detecting image forgery based on multi-expert model decision can essentially be executed by a computer program. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the method provided in the above embodiment, which includes a memory and a processor;

[0107] The memory is used to store a computer program;

[0108] The processor is used to implement the method for detecting image forgery based on multi-expert model decision in the above embodiment when executing the computer program.

[0109] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0110] Therefore, based on the same inventive concept, in another preferred embodiment of the present invention, there is also provided a computer-readable storage medium corresponding to the method provided in the above embodiment. A computer program is stored on the storage medium. When the computer program is executed by a processor, it can implement the image forgery detection method based on multi-expert model decision-making in the above embodiment.

[0111] It can be understood that the above storage medium may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory. At the same time, the storage medium may also be various media such as a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc that can store program codes.

[0112] It can be understood that the above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0113] In addition, it should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system can refer to the corresponding process in the foregoing method embodiment, and will not be elaborated herein. In the various embodiments provided in the present application, the division of steps or modules in the system and method is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or steps can be combined or integrated together, and a module or step can also be split.

[0114] The embodiments described above are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.

Claims

1. An image forgery detection method based on multi-expert model decision making, characterized in that: The following steps are involved: (1) obtaining normal images in multiple fields, generating and grouping forged images based on different forgery types and forgery degrees, and using normal images and groups of forged images to pre-train a pre-screening model and a multi-expert model that can predict the probability of forgery. The multi-expert model includes a typical forgery trace detection model based on the EfficientNet network, a complex refinement trace detection model based on the ViT network, and a local forgery trace detection model based on the Xception network; (2) Using normal images and forged images, an image forgery detection network is trained, which includes an image encoding module, a gated calculation module, a pre-screening model, and a multi-expert model; During training, the parameters of the pre-screening model and the multi-expert model are frozen; (3) Based on the trained image forgery detection network, identify whether the input image is forged, including: Extract image features using image coding module; The pre-screening model takes image features as input to identify whether it is a forged image. If so, the judgment ends; if not, the gated calculation module takes image features as input to generate a probability vector with the same dimension as the number of multiple expert models as the weight value of each expert model; the expert model with non-zero weight value is screened and the forgery probability is independently judged with image features as input, and the forgery probability values ​​output by each expert model are fused based on the non-zero weight value, and the input image with a total forgery probability value higher than the threshold is identified as a forged image.

2. The image forgery detection method based on multi-expert model decision making according to claim 1 is characterized in that: Step (1) also includes an image preprocessing process, wherein the preprocessing includes standardization, denoising and enhancement.

3. The image forgery detection method based on multi-expert model decision making according to claim 1 is characterized in that: In step (1), generating a forged image based on different forgery types and forgery degrees includes: The types of forgery include texture forgery, boundary forgery, brightness forgery, facial feature forgery, expression forgery, lighting forgery and synthetic splicing forgery; The forgery degree refers to the total disturbance degree of a normal image under different forgery types of disturbances.

4. The image forgery detection method based on multi-expert model decision making according to claim 3 is characterized in that: In step (1), the method for grouping forged images includes: (1-1) Obtaining the forgery type and forgery degree corresponding to the forged image. If the forgery type corresponding to the forged image includes one or more of texture forgery, boundary forgery, brightness forgery, and synthetic splicing forgery, and the forgery degree is higher than a threshold, the forgery image is classified into the first forged image group; otherwise, executing step (1-2); (1-2) If the forgery type corresponding to the forged image includes one or more of texture forgery, boundary forgery, brightness forgery, and synthetic splicing forgery, and the degree of forgery is not higher than the threshold; or the forgery type is selected from any one of facial feature forgery, expression forgery, and lighting forgery, and the degree of forgery is higher than the threshold; then it is classified into the second forged image group; otherwise, step (1-3) is executed; (1-3) If the forgery type corresponding to the forged image is selected from facial feature forgery, expression forgery, and lighting forgery, and the degree of forgery is not higher than the threshold, then the image is classified into the third forged image group; (1-4) Selecting images whose forged regions are local regions from the third forged image group, and the selected images constitute a fourth forged image group.

5. The image forgery detection method based on multi-expert model decision making according to claim 3 is characterized in that: In step (1), the combination of the first forged image group and several normal images is used as the pre-screening model training set, the combination of the second forged image group and several normal images is used as the training set of the EfficientNet network, the combination of the third forged image group and several normal images is used as the training set of the ViT network, and the combination of the fourth forged image group and several normal images is used as the training set of the Xception network.

6. The image forgery detection method based on multi-expert model decision making according to claim 1 or 5, characterized in that: The pre-screening model adopts a model based on a CNN network, and the gating calculation module adopts a fully connected network.

7. The image forgery detection method based on multi-expert model decision making according to claim 1 is characterized in that: It also includes visualization steps of the image forgery detection process, including: Display the heat map corresponding to the input image; As well as, the forgery probability, weight distribution, total forgery probability and final recognition result independently judged by each expert model are displayed.

8. An image forgery detection system based on multi-expert model decision-making, used to implement the method described in claim 1; characterized in that: The system comprises: An image sample generation module, which is used to obtain normal images in multiple fields, generate forged images based on different forgery types and forgery degrees, and group them; A one-time training module, which is used to pre-train a pre-screening model and a multi-expert model capable of predicting the probability of forgery using normal images and each group of forged images, wherein the multi-expert model includes a typical forgery trace detection model based on an EfficientNet network, a complex refinement trace detection model based on a ViT network, and a local forgery trace detection model based on an Xception network; A secondary training module, which is used to train an image forgery detection network including an image encoding module, a gated calculation module, a pre-screening model and a multi-expert model using normal images and forged images; during training, the parameters of the pre-screening model and the multi-expert model are frozen; The image forgery identification module is used to identify whether the input image is forged based on the trained image forgery detection network, including: using the image encoding module to extract image features; the pre-screening model uses the image features as input to identify whether it is a forged image, and if so, the judgment ends; if not, the gated calculation module uses the image features as input to generate a probability vector with the same dimension as the number of multiple expert models as the weight value of each expert model; the expert model with non-zero weight value is screened and the forgery probability is independently judged using the image features as input, and the forgery probability values ​​output by each expert model are fused based on the non-zero weight value, and the input image with a total forgery probability value higher than the threshold is identified as a forged image.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the image forgery detection method based on multi-expert model decision-making according to any one of claims 1 to 7 is implemented.

10. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is configured to implement the image forgery detection method based on multi-expert model decision-making according to any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • General deep forgery detection method based on generative model

    CN117238015A

  • Deep counterfeit image detection method and device based on artifact domain adversarial learning

    CN118469968A