Image forgery detection and training method based on structural causality and related devices
The structural causal image authentication model decouples and concatenates features to enhance the accuracy of image authentication by focusing on altered regions, reducing interference from unaltered areas.
Patent Information
- Application Number
- CN202510324624.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-19
AI Technical Summary
When the existing image pseudo-identification technology deals with tampered areas, the untampered natural areas interferes with the identification of the forged areas, resulting in a decrease in the accuracy of pseudo-identification.
The image pseudo-recognition model based on structural causality is adopted, and the forged-related features and forged-independent features of the training image are decoupled and stitched through the backbone network, the forged bias elimination module (FBEM) and the classifier, and the forged-related features and forged-independent features are respectively trained to improve the accuracy of pseudo-recognition.
Effectively eliminate confusion and interference in natural areas, improve the accuracy of image false recognition, can accurately identify tampered areas, adapt to various tampered scenarios, and have high interpretability and robustness.
Smart Images

Figure CN119851067B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and more particularly, to an image forgery detection and training method based on structural causality and related devices. Background Art
[0002] With the rapid development of technology, it has become increasingly common to generate various highly realistic images and videos manually, which makes high-quality content creation more convenient. However, this also brings the risk of abuse of technology. For example, it may lead to the spread of false information. Therefore, it is crucial to develop image forgery detection technology.
[0003] In related technologies, when detecting the forgery of an image, it is usually based on the overall features of all regions included in the image to determine the authenticity of the image. However, the tampered region in the image is often only a small part, while most of the regions included in the image are actually real regions that have not been tampered with. For example, for a face image, the tampered region may only be the nose region or the mouth region, while most of the other regions of the face image have not actually been tampered with. At this time, if the forgery detection of the image is based on the overall features of the image, then most of the natural regions in the image that have not been tampered with will interfere with the recognition of the forged region in the image, which will reduce the accuracy of image forgery detection. Summary of the Invention
[0004] The present disclosure provides an image forgery detection and training method based on structural causality and related devices to at least solve the problem in the above related technologies that most of the natural regions in the image that have not been tampered with interfere with the recognition of the forged region, thereby reducing the accuracy of image forgery detection.
[0005] According to the first aspect of the embodiments of the present disclosure, a training method for an image forgery detection model based on structural causality is provided. The image forgery detection model includes a backbone network, a forgery bias elimination module (FBEM), and a classifier. The training method includes: obtaining training image samples, where the training image samples include two training images with different contents, and each training image corresponds to an image authenticity label and coordinates indicating the image tampering area; for each of the two training images, extracting deep features of the training image through the backbone network, and decoupling the deep features of the training image based on the coordinates corresponding to the training image through the FBEM to obtain a first feature and a second feature after decoupling, where the first feature is a forgery-related feature related to the image tampering area, and the second feature is a forgery-unrelated feature unrelated to the image tampering area; splicing any of the first features and the second features extracted from the two training images through the FBEM to obtain a spliced feature; inputting the spliced feature into the classifier to obtain a true / false prediction result of the image; calculating a loss based on the true / false prediction result of the image and the image authenticity label corresponding to the spliced feature, where the image authenticity label corresponding to the spliced feature is the image authenticity label of the training image corresponding to the first feature included in the spliced feature; adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss to train the image forgery detection model.
[0006] Optionally, the step of splicing any of the first features and the second features extracted from the two training images through the FBEM to obtain a spliced feature includes: splicing the first feature and the second feature from the same training image through the FBEM to obtain a first spliced feature; splicing the first features and the second features from different training images through the FBEM to obtain a second spliced feature; where the step of inputting the spliced feature into the classifier to obtain a true / false prediction result of the image; inputting the first spliced feature into the classifier to obtain a first true / false prediction result of the image; inputting the second spliced feature into the classifier to obtain a second true / false prediction result of the image; where the step of calculating the value of the loss function based on the true / false prediction result of the image and the image authenticity label corresponding to the spliced feature includes: calculating a first loss based on the first true / false prediction result and the image authenticity label corresponding to the first spliced feature; calculating a second loss based on the second true / false prediction result and the image authenticity label corresponding to the second spliced feature; the step of adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss includes: adjusting the parameters of the backbone network, the FBEM, and the classifier according to the first loss and the second loss.
[0007] Optionally, adjusting the parameters of the backbone network, the FBEM, and the classifier according to the first loss and the second loss includes: performing a weighted sum of the first loss and the second loss to obtain a weighted loss; adjusting the parameters of the backbone network, the FBEM, and the classifier according to the weighted loss.
[0008] Optionally, the training method further includes: calculating the conditional mutual information between the first feature and the second feature decoupled from the two training images, where the conditional mutual information is used to characterize the degree of decoupling between the first feature and the second feature; adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss includes: adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss and the conditional mutual information.
[0009] Optionally, adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss and the conditional mutual information includes: calculating the weighted sum result of the loss and the conditional mutual information; adjusting the parameters of the backbone network, the FBEM, and the classifier according to the weighted sum result.
[0010] Optionally, obtaining the training image samples includes: obtaining two original images with different contents; respectively performing tampering on the two original images to obtain the training image samples.
[0011] Optionally, the tampering method includes at least one of the following items: adding Gaussian noise, adding jitter, splicing, copy-pasting, erasing.
[0012] According to a second aspect of the embodiments of the present disclosure, there is provided an image forgery detection method based on structural causality. The image forgery detection method based on structural causality is implemented based on an image forgery detection model. The image forgery detection model includes a backbone network, a forgery bias elimination module (FBEM), and a classifier. The image forgery detection method based on structural causality includes: obtaining a target image to be subjected to image forgery detection; extracting target depth features of the target image through the backbone network, and decoupling the target depth features through the FBEM to obtain a decoupled first target feature and a second target feature, where the first target feature is a forgery-related feature related to the image tampering area, and the second target feature is a forgery-unrelated feature unrelated to the image tampering area; splicing the first target feature and the second target feature through the FBEM to obtain a target spliced feature; inputting the target spliced feature into the classifier to obtain a true / false prediction result of the target image.
[0013] According to a third aspect of the embodiments of the present disclosure, there is provided a training device for an image forgery detection model based on structural causality. The image forgery detection model includes a backbone network, a forgery bias elimination module (FBEM), and a classifier. The training device includes: an image sample acquisition module configured to acquire training image samples, where the training image samples include two training images with different contents, and each training image is corresponding to an image authenticity label and coordinate positions indicating the image tampering area; a decoupling module configured to, for each of the two training images, extract the depth features of the training image through the backbone network, and decouple the depth features of the training image based on the coordinate positions corresponding to the training image through the FBEM to obtain a first feature and a second feature after decoupling, where the first feature is a forgery-related feature related to the image tampering area, and the second feature is a forgery-unrelated feature unrelated to the image tampering area; a splicing module configured to splice any of the first features and the second features extracted from the two training images through the FBEM to obtain a spliced feature; a feature input module configured to input the spliced feature into the classifier to obtain a true / false prediction result of the image; a loss calculation module configured to calculate a loss based on the true / false prediction result of the image and the image authenticity label corresponding to the spliced feature, where the image authenticity label corresponding to the spliced feature is the image authenticity label of the training image corresponding to the first feature included in the spliced feature; a training module configured to train the image forgery detection model by adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss.
[0014] Optionally, the splicing module is configured to: splice the first feature and the second feature from the same training image through the FBEM to obtain a first spliced feature; splice the first feature and the second feature from different training images through the FBEM to obtain a second spliced feature; the feature input module is configured to: input the first spliced feature into the classifier to obtain a first true / false prediction result of the image; input the second spliced feature into the classifier to obtain a second true / false prediction result of the image; the loss calculation module is configured to: calculate a first loss based on the first true / false prediction result and the image authenticity label corresponding to the first spliced feature; calculate a second loss based on the second true / false prediction result and the image authenticity label corresponding to the second spliced feature; the training module is configured to: adjust the parameters of the backbone network, the FBEM, and the classifier according to the first loss and the second loss.
[0015] Optionally, the training module is configured to: perform a weighted sum of the first loss and the second loss to obtain a weighted loss; adjust the parameters of the backbone network, the FBEM, and the classifier according to the weighted loss.
[0016] Optionally, the training device further includes: a conditional mutual information calculation module configured to calculate the conditional mutual information between the first feature and the second feature decoupled from the two training images, where the conditional mutual information is used to characterize the decoupling degree between the first feature and the second feature; the training module is configured to: adjust the parameters of the backbone network, the FBEM, and the classifier according to the loss and the conditional mutual information.
[0017] Optionally, the training module is configured to: calculate the weighted sum result of the loss and the conditional mutual information; adjust the parameters of the backbone network, the FBEM, and the classifier according to the weighted sum result.
[0018] Optionally, the image sample acquisition module is configured to: acquire two original images with different contents; perform tampering on the two original images respectively to obtain the training image samples.
[0019] Optionally, the tampering method includes at least one of the following items: adding Gaussian noise, adding jitter, splicing, copy-pasting, erasing.
[0020] According to a fourth aspect of the embodiments of the present disclosure, there is provided an image forgery detection device based on structural causality. The image forgery detection device based on structural causality is implemented based on an image forgery detection model. The image forgery detection model includes a backbone network, a forgery bias elimination module (FBEM), and a classifier. The image forgery detection device based on structural causality includes: a target image acquisition module configured to acquire a target image to be subjected to image forgery detection; a feature decoupling module configured to extract a target depth feature of the target image through the backbone network and decouple the target depth feature through the FBEM to obtain a decoupled first target feature and a second target feature, where the first target feature is a forgery-related feature related to the image tampering area, and the second target feature is a forgery-unrelated feature unrelated to the image tampering area; a feature splicing module configured to splice the first target feature and the second target feature through the FBEM to obtain a target splicing feature; a genuine / fake result prediction module configured to input the target splicing feature into the classifier to obtain a genuine / fake prediction result of the target image.
[0021] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the instructions to implement the training method of the image forgery authentication model based on structural causality according to the present disclosure, or to implement the image forgery authentication method based on structural causality according to the present disclosure.
[0022] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the training method of the image forgery authentication model based on structural causality according to the present disclosure, or to execute the image forgery authentication method based on structural causality according to the present disclosure.
[0023] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, which when executed by a processor implements the training method of the image forgery authentication model based on structural causality according to the present disclosure, or implements the image forgery authentication method based on structural causality according to the present disclosure.
[0024] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0025] In the present disclosure, by decoupling the overall features of training images with different contents respectively, and splicing the forgery-related features and forgery-unrelated features decoupled out to train the image forgery authentication model, it can be ensured that the trained image forgery authentication model can accurately identify the tampered areas in the image, that is, it can effectively exclude the confusion interference brought by natural areas irrelevant to image forgery authentication, thereby improving the accuracy of image forgery authentication.
[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0028] Figure 1 is a schematic structural diagram showing an image forgery authentication model according to an exemplary embodiment of the present disclosure;
[0029] Figure 2 is a schematic logical diagram showing a structural causality model according to an exemplary embodiment of the present disclosure;
[0030] Figure 3It is a flowchart showing a method for training a structure-causal based image forgery detection model according to an exemplary embodiment of the present disclosure;
[0031] Figure 4 It is a flowchart showing a structure-causal based image forgery detection method according to an exemplary embodiment of the present disclosure;
[0032] Figure 5 It is a block diagram showing a structure-causal based image forgery detection model training apparatus according to an exemplary embodiment of the present disclosure;
[0033] Figure 6 It is a block diagram showing a structure-causal based image forgery detection apparatus according to an exemplary embodiment of the present disclosure;
[0034] Figure 7 It is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0035] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0037] It should be noted here that "at least one of several items" in the present disclosure all represents the inclusion of three parallel situations: "any one of the several items", "any combination of several items among them", and "the whole of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "executing at least one of step one and step two" means the following three parallel situations: (1) executing step one; (2) executing step two; (3) executing step one and step two.
[0038] With the rapid development of vision-language models, generating highly realistic visual media such as images and videos has become increasingly common, making high-quality content creation more convenient. However, this also brings the risk of technology abuse, which may lead to the spread of false information. Therefore, it is crucial to develop models for detecting visual media forgery.
[0039] Research in related technologies mainly focuses on identifying the authenticity of content and its forgery locations, and emphasizes improving the accuracy and generalization ability of detection. However, these methods often ignore the importance of explaining why certain results are judged to be true or false, which is crucial for ensuring reliable traceability and forensic analysis. In view of this, there is an urgent need to construct a forgery detection model that is both trustworthy and interpretable.
[0040] In addition, it should be noted that causal reasoning has received significant attention in fields such as natural language generation and information retrieval because it can model causal relationships and make interpretable and generalizable decisions by addressing confounding biases in the data. In the field of forgery detection, causal learning can clearly show the impact of confounding variables and help to untangle spurious correlations, thus becoming a reliable tool for improving visual media forgery detection.
[0041] To solve the above problems existing in related technologies, the image forgery detection and training method and related device provided by the present disclosure decouple the overall features of training images with different contents respectively, and splice the forgery-related features and forgery-unrelated features decoupled to train the image forgery detection model, which can ensure that the trained image forgery detection model can accurately identify the tampered areas in the image, that is, it can effectively exclude the confounding interference brought by natural areas irrelevant to image forgery detection, thereby improving the accuracy of image forgery detection.
[0042] It should be noted that the data and images involved in the present disclosure are all information fully authorized by all parties. Figure 1 FIG. is a schematic structural diagram of an image forgery detection model according to an exemplary embodiment of the present disclosure. The image forgery detection model can be a model obtained by improving on an existing Structural Causal Model (SCM). The existing structural causal model mainly uses a directed graph to model the causal relationship between variables, and this directed graph can generally be a Directed Acyclic Graph (DAG).
[0043] Figure 2 FIG. is a logical schematic diagram of a structural causal model according to an exemplary embodiment of the present disclosure. Refer to Figure 2, in a structural causal model, the causal effect from the input image X to the predicted label Y can be defined, and at the same time, the influence of domain knowledge K on the input image X and the predicted label Y can be considered, that is, the domain knowledge K can be introduced to affect the relationship between the input and the output. The "domain knowledge K" can refer to different tampering methods, that is, it can refer to various perturbations added to the image.
[0044] It should be noted that the structural causal model in the related art mainly can include a backbone network (BackboneEncoder) and a classifier. Specifically, the "backbone network" is mainly used to extract the deep features of the image; the "classifier" is mainly used for binary classification prediction of the image, that is, the "classifier" is mainly used to predict whether the image belongs to an untampered real image or a tampered fake image.
[0045] In the present disclosure, a "Forgery Bias Eliminating Module (FBEM)" is mainly added on the basis of the existing structural causal model. The "Forgery Bias Eliminating Module" is mainly used to decouple the deep features extracted from the image to obtain forgery-related features related to the image tampering area and forgery-unrelated features unrelated to the image tampering area. And the "Forgery Bias Eliminating Module" is also used to splice the decoupled forgery-related features and forgery-unrelated features, and then the spliced features can be obtained.
[0046] In this way, in the present disclosure, by setting the Forgery Bias Eliminating Module (FBEM), effective decomposition of the deep features can be achieved, that is, effective decoupling of the forgery-related features and the forgery-unrelated features can be ensured, and then the practicability and flexibility of the image forgery detection model in different tampering scenarios can be improved.
[0047] Figure 3 is a flowchart showing a training method of an image forgery detection model based on structural causality according to an exemplary embodiment of the present disclosure. As mentioned above, the image forgery detection model can include a backbone network, a Forgery Bias Eliminating Module (FBEM), and a classifier.
[0048] Referring to Figure 3 , in step 301, training image samples can be obtained, where the training image samples can include two training images with different contents, and each training image can correspond to an image authenticity label and the coordinate position for indicating the image tampering area.
[0049] Exemplarily, assuming that the training images are face images, each face image can correspond to an image authenticity label for indicating whether the face image has been tampered with, and the coordinate positions for indicating the specific tampered areas. Exemplarily, the tampered areas in the face image can be, but are not limited to, the nose area, the mouth area, the eye area, etc.
[0050] According to an exemplary embodiment of the present disclosure, two original images with different contents can be obtained. Then, the two original images can be tampered with respectively to obtain training image samples. Exemplarily, as described above, the original images can be face images. At this time, the face images of two persons can be tampered with respectively, and thus the training image samples for training the image forgery detection model can be obtained.
[0051] It should be noted that the tampering methods and tampered areas for the two original images with different contents can be the same or different. For example, the same tampering method can be used to tamper with the same area of the two original images; or, different tampering methods can be used to tamper with the same area of the two original images; or, different tampering methods can also be used to tamper with different areas of the two original images, etc. The present disclosure does not make specific limitations on this.
[0052] In this way, by tampering with the original images to train the image forgery detection model, it can be ensured that the trained image forgery detection model can effectively perform binary classification prediction on images, that is, it can be ensured that the trained image forgery detection model can effectively distinguish whether the image to be detected is a forged image that has been tampered with or a real image that has not been tampered with.
[0053] According to an exemplary embodiment of the present disclosure, the aforementioned tampering methods can include at least one of the following items: adding Gaussian noise, adding jitter, splicing, copy-pasting, erasing. It should be noted that the tampering methods provided by the present disclosure can be other methods in addition to the above methods. The present disclosure does not make specific limitations on this. The foregoing embodiments are only for exemplary illustration. In this way, by setting a variety of different image tampering methods, it can be ensured that the trained image forgery detection model can cope with various different tampering scenarios, and the generalization of the image forgery detection model is better.
[0054] In step 302, for each of the two training images, the depth feature Z of the training image can be extracted through the aforementioned backbone network, and the depth feature Z of the training image can be decoupled based on the corresponding coordinate position of the training image through the aforementioned FBEM to obtain the decoupled first feature and second feature. Specifically, two projection layers (Projection layer) included in the FBEM: φ and ψ can be used to decouple the depth feature Z into the first feature and the second feature. Moreover, the first feature can be a forgery-related feature related to the image tampering area , and the second feature can be a forgery-unrelated feature unrelated to the image tampering area .
[0055] Exemplarily, referring to Figure 1 , decoupling the depth feature Z of the i-th training image can obtain the forgery-related feature and the forgery-unrelated feature of the i-th training image; decoupling the depth feature Z of the j-th training image can obtain the forgery-related feature and the forgery-unrelated feature of the j-th training image.
[0056] In step 303, any first feature and second feature extracted from the two training images can be spliced through the FBEM to obtain a spliced feature. That is, the forgery-related feature and the forgery-unrelated feature extracted from the two training images can be spliced through the FBEM to obtain a spliced feature.
[0057] Exemplarily, referring to Figure 1 , the forgery-related feature and the forgery-unrelated feature extracted from the two training images can be spliced through the FBEM to obtain a spliced feature: CAT( , ); and, the forgery-related feature and the forgery-unrelated feature extracted from the two training images can also be spliced through the FBEM to obtain a spliced feature: CAT( , ), where "CAT()" represents splicing along the channel dimension, that is, it can represent connecting features along the depth direction. It should be noted that in the present disclosure, in addition to using the method of "CAT()" for feature splicing, other splicing methods can also be adopted, and the present disclosure does not make specific limitations on this. The aforementioned splicing method is only an exemplary illustration.
[0058] In step 304, the splicing feature can be input into the classifier as described above to obtain the authenticity prediction result of the image. That is, the splicing feature can be input into the classifier as described above to determine whether the corresponding image is a forged pseudo-image that has been tampered with or a real image that has not been tampered with.
[0059] In step 305, the loss can be calculated based on the authenticity prediction result of the image and the image authenticity label corresponding to the splicing feature. Among them, the image authenticity label corresponding to the splicing feature can be the image authenticity label of the training image corresponding to the first feature included in the splicing feature, that is, the image authenticity label corresponding to the splicing feature can be the image authenticity label of the training image from which the forgery-related feature included in the splicing feature comes.
[0060] Exemplarily, for the aforementioned splicing feature: , the image authenticity label corresponding to the splicing feature can be the image authenticity label of the i-th training image from which the forgery-related feature included in the splicing feature comes; or, for the splicing feature: , the image authenticity label corresponding to the splicing feature can be the image authenticity label of the i-th training image from which the forgery-related feature included in the splicing feature comes. , the image authenticity label corresponding to the splicing feature can be the image authenticity label of the i-th training image from which the forgery-related feature included in the splicing feature comes. comes.
[0061] In step 306, the parameters of the backbone network, FBEM, and classifier can be adjusted according to the loss to train the image forgery detection model.
[0062] According to an exemplary embodiment of the present disclosure, the first feature and the second feature from the same training image can be spliced through FBEM to obtain the first splicing feature. Exemplarily, as described above, the forgery-related feature and the forgery-unrelated feature both from the i-th training image can be spliced to obtain the first splicing feature: . And, the first feature and the second feature from different training images can also be spliced through FBEM to obtain the second splicing feature. Exemplarily, as described above, the forgery-related feature from the i-th training image and the forgery-unrelated feature from the j-th training image can be spliced through FBEM to obtain the second splicing feature: .
[0063] Then, the first splicing feature can be input into the classifier to obtain the first authenticity prediction result of the image. Exemplarily, the first splicing feature: The input classifier obtains the first authenticity prediction result for the i-th training image; and, the second splicing feature can also be input into the classifier to obtain the second authenticity prediction result of the image. Exemplarily, the second splicing feature: is input into the classifier to obtain the second authenticity prediction result of the i-th training image.
[0064] Next, based on the first authenticity prediction result and the image authenticity label corresponding to the first splicing feature, the first loss can be calculated. Exemplarily, based on the first authenticity prediction result and the first splicing feature: the corresponding image authenticity label, the first loss can be calculated, where the first loss can be the cross-entropy loss ; and, based on the second authenticity prediction result and the image authenticity label corresponding to the second splicing feature, the second loss can also be calculated. Exemplarily, based on the second authenticity prediction result and the second splicing feature: the corresponding image authenticity label, the second loss can be calculated, where the second loss can be the decoupling loss .
[0065] Then, the parameters of the backbone network, FBEM, and classifier can be adjusted according to the first loss and the second loss.
[0066] It should be noted that the above cross-entropy loss can be calculated by the following formula:
[0067]
[0068] where is the forgery-related feature decoupled from the i-th training image, is the forgery-unrelated feature decoupled from the i-th training image, is the image authenticity label corresponding to the i-th training image, BCELoss is the binary cross-entropy loss function (Binary Cross Entropy Loss), represents the "CAT()" concatenation operation, where, as mentioned above, "CAT()" represents concatenation along the channel dimension, that is, it can represent connecting features along the depth direction.
[0069] The above decoupling loss can be calculated by the following formula:
[0070]
[0071] where is the forgery-related feature decoupled from the i-th training image, is the forgery-unrelated feature decoupled from the j-th training image, is the image authenticity label corresponding to the i-th training image, and BCELoss is the Binary Cross Entropy Loss function. represents the "CAT()" concatenation operation. As mentioned above, "CAT()" represents concatenation along the channel dimension, that is, it can represent connecting features along the depth direction.
[0072] According to an exemplary embodiment of the present disclosure, the first loss and the second loss can be weighted and summed to obtain a weighted loss. Exemplarily, the aforementioned cross-entropy loss and the decoupling loss can be weighted and summed to obtain a weighted loss. Then, the parameters of the backbone network, FBEM, and classifier can be adjusted according to the weighted loss.
[0073] According to an exemplary embodiment of the present disclosure, the conditional mutual information between the first feature and the second feature decoupled from two training images can also be calculated , where the conditional mutual information can be used to characterize the degree of decoupling between the first feature and the second feature, that is, the conditional mutual information can be used to characterize the degree of decoupling between the forgery-related features and the forgery-unrelated features. And, the smaller the conditional mutual information , the higher the degree of decoupling between the features. Therefore, one of the objectives of the present disclosure is to minimize the conditional mutual information to ensure sufficient decoupling of both the forgery-related features and the forgery-unrelated features.
[0074] Exemplarily, the conditional mutual information can be calculated by the following formula:
[0075]
[0076] where N is the number of training images included in the training image samples, is the forgery-related feature of the i-th training image decoupled, is the forgery-unrelated feature of the i-th training image decoupled, is the forgery-unrelated feature of the j-th training image decoupled. When the image authenticity label corresponding to the i-th training image included in the training image samples is the same as the image authenticity label corresponding to the j-th training image, = 1; otherwise, that is, when the image authenticity label corresponding to the i-th training image included in the training image samples is different from the image authenticity label corresponding to the j-th training image, = 0.
[0077] Then, the parameters of the backbone network, FBEM, and classifier can be adjusted according to the aforementioned loss and conditional mutual information. Specifically, the weighted sum result L of the aforementioned loss and conditional mutual information can be calculated, and then the parameters of the backbone network, FBEM, and classifier can be adjusted according to this weighted sum result L. Exemplarily, L can be calculated by the following formula:
[0078] L =
[0079] where
[0080] is the weight corresponding to the decoupling loss , and is the weight corresponding to the conditional mutual information .
[0081] In this way, in the present disclosure, by introducing the forgery bias elimination loss function and the conditional mutual information regularization term, while ensuring the detection performance of the image forgery detection model, the computational cost can also be reduced, so that the image forgery detection method based on structural causality provided by the present disclosure can be seamlessly integrated into the existing system.
[0082] In addition, referring to Figure 1 , FBEM decouples a total of 4 features, namely: the forgery-related feature from the i-th training image, the forgery-unrelated feature from the i-th training image, the forgery-related feature from the j-th training image, and the forgery-unrelated feature from the j-th training image. When splicing these 4 features, two results can be obtained. The first result is the splicing result shown in Figure 1 : and ; the second splicing result can be: CAT( , ) and CAT( , ).
[0083] For the first splicing result, it can correspond to its own cross-entropy loss , decoupling loss , and conditional mutual information ; for the second splicing result, it can also correspond to its own cross-entropy loss , decoupling loss , and conditional mutual information . At this time, the cross-entropy loss corresponding to the first splicing result can be combined with the cross-entropy loss corresponding to the second splicing result Add them to obtain the total cross-entropy loss. Similarly, the decoupling loss corresponding to the first splicing result and the decoupling loss corresponding to the second splicing result can be added to obtain the total decoupling loss. The conditional mutual information corresponding to the first splicing result and the conditional mutual information corresponding to the second splicing result can be added to obtain the total conditional mutual information.
[0084] Next, the total cross-entropy loss, the total decoupling loss, and the total conditional mutual information can be weighted and summed, and then the parameters of the image forgery detection model can be adjusted based on the weighted sum result, so as to train the image forgery detection model.
[0085] In the present disclosure, a method for image forgery detection based on a structural causal model is proposed to explain the inference target of forgery detection from a causal perspective, which improves the accuracy of traditional forgery detection methods in estimating the required causal effects. Specifically, the present disclosure introduces a forgery bias elimination module (FBEM), which is mainly used to decouple forgery-related features and forgery-unrelated features in the feature representation. Further, the FBEM can decouple the features from the encoder through two projection layers φ and ψ, thereby avoiding additional complexity. And the FBEM is a lightweight feature decoupling plug-in, which is designed to be used in a plug-and-play manner, making it easy to integrate into existing systems, that is, it can be flexibly and seamlessly applied to various scenarios, thus allowing the corresponding functions to be realized without major modification of the current model.
[0086] On this basis, the present disclosure also proposes a forgery bias elimination loss function and a conditional mutual information regularization term to ensure that the image forgery detection model can effectively distinguish the aforementioned forgery-related features and forgery-unrelated features during the learning process, thereby enhancing the interpretability and robustness of the image forgery detection model, and at the same time ensuring a minimum computational cost.
[0087] Figure 4 is a flowchart showing a structural-causal-based image forgery detection method according to an exemplary embodiment of the present disclosure. The structural-causal-based image forgery detection method can be implemented based on an image forgery detection model trained by the training method according to the present disclosure, wherein the image forgery detection model can include a backbone network, a forgery bias elimination module (FBEM), and a classifier.
[0088] Refer to Figure 4, in step 401, the target image to be subjected to image forgery detection can be obtained. Exemplarily, the target image to be subjected to image forgery detection can be a person image, an animal image, a natural landscape image, etc. The present disclosure does not specifically limit the type of the target image to be subjected to image forgery detection.
[0089] In step 402, the target depth feature of the target image can be extracted through the backbone network, and the FBEM can be used to decouple the target depth feature to obtain the decoupled first target feature and second target feature. Specifically, two projection layers included in the FBEM: φ and ψ can be used to decouple the target depth feature into the first target feature and the second target feature. Moreover, the first target feature can be a forgery-related feature related to the image tampering area, and the second target feature can be a forgery-unrelated feature unrelated to the image tampering area.
[0090] In step 403, the FBEM can be used to splice the first target feature and the second target feature to obtain the target splicing feature. That is, the forgery-related feature and the forgery-unrelated feature in the target image can be spliced through the FBEM to obtain the target splicing feature.
[0091] In step 404, the target splicing feature can be input into the classifier to obtain the authenticity prediction result of the target image. That is, the target splicing feature can be input into the classifier to perform binary classification on the target image through the classifier, that is, to predict whether the target image is an untampered real image or a tampered forged image.
[0092] It should be noted that causal inference can effectively identify confounding factors in the data, and thus can help explain and reliably trace the forgery detection results. Therefore, the present disclosure proposes a forgery detection method based on a structural causal model (SCM). This method mainly uses causal relationships to capture the causal effect between the input image and the predicted label, thereby improving the accuracy and interpretability of forgery detection. Moreover, the forgery detection method provided by the present disclosure also has high credibility and traceability. Therefore, it can provide a more interpretable and reliable discrimination basis for forgery forensics, thereby coping with the increasingly serious problem of false information dissemination.
[0093] Figure 5 is a block diagram showing a training device 500 of an image forgery detection model based on structural causality according to an exemplary embodiment of the present disclosure. The image forgery detection model can include a backbone network, a forgery bias elimination module (FBEM), and a classifier.
[0094] Refer to Figure 5, the training device 500 of the structure-causality-based image forgery detection model may include an image sample acquisition module 501, a decoupling module 502, a splicing module 503, a feature input module 504, a loss calculation module 505, and a training module 506.
[0095] The image sample acquisition module 501 can acquire training image samples. Among them, the training image samples can include two training images with different contents, and each training image can correspond to an image authenticity label and the coordinate positions indicating the image tampering areas.
[0096] Exemplarily, assuming the training image is a face image, each face image can correspond to an image authenticity label indicating whether the face image has been tampered with, and the coordinate positions indicating the specific tampering areas. Exemplarily, the tampering areas in the face image can be, but are not limited to: the nose area, the mouth area, or the eye area, etc.
[0097] According to an exemplary embodiment of the present disclosure, the image sample acquisition module 501 can acquire two original images with different contents. Then, the image sample acquisition module 501 can perform tampering on the two original images respectively to obtain training image samples. Exemplarily, as described above, the original images can be face images. At this time, the face images of two people can be tampered with respectively, and then the training image samples for training the image forgery detection model can be obtained.
[0098] It should be noted that the tampering methods and tampering areas for the two original images with different contents can be the same or different. For example, the same tampering method can be used to tamper with the same area of the two original images; or, different tampering methods can be used to tamper with the same area of the two original images; or, different tampering methods can also be used to tamper with different areas of the two original images, etc. The present disclosure does not make specific limitations on this.
[0099] In this way, by tampering with the original images to train the image forgery detection model, it can be ensured that the trained image forgery detection model can effectively perform binary classification prediction on images, that is, it can be ensured that the trained image forgery detection model can effectively distinguish whether the image to be detected is a forged image that has been tampered with or an authentic image that has not been tampered with.
[0100] According to an exemplary embodiment of the present disclosure, the foregoing tampering methods may include at least one of the following items: adding Gaussian noise, adding jitter, splicing, copy-pasting, erasing. It should be noted that the tampering methods provided by the present disclosure may be other methods in addition to the above methods, and the present disclosure does not make specific limitations thereto. The foregoing embodiments are merely an exemplary illustration. In this way, by setting a variety of different image tampering methods, it can be ensured that the trained image forgery detection model can cope with various different tampering scenarios, and the generalization of the image forgery detection model is relatively good.
[0101] For each of the two training images, the decoupling module 502 may extract the depth feature Z of the training image through the foregoing backbone network, and may decouple the depth feature Z of the training image based on the coordinate position corresponding to the training image through the foregoing FBEM to obtain the decoupled first feature and second feature. Specifically, two projection layers (Projection layer) included in the FBEM: φ and ψ, may be used to decouple the depth feature Z into the first feature and the second feature. And, the first feature may be a forgery-related feature related to the image tampering area , and the second feature may be a forgery-unrelated feature unrelated to the image tampering area .
[0102] In this way, in the present disclosure, by setting the forgery bias elimination module (FBEM), effective decomposition of the depth feature can be achieved, that is, effective decoupling of the forgery-related feature and the forgery-unrelated feature can be ensured, and further, the practicability and flexibility of the image forgery detection model in different tampering scenarios can be improved.
[0103] The splicing module 503 may splice any first feature and second feature extracted from the two training images through the FBEM to obtain a spliced feature. That is, the forgery-related feature and the forgery-unrelated feature extracted from the two training images may be spliced through the FBEM to obtain a spliced feature.
[0104] The feature input module 504 may input the spliced feature into the classifier as described above to obtain the authenticity prediction result of the image. That is, the spliced feature may be input into the classifier as described above to determine whether the corresponding image is a forged image after tampering or a real image without tampering.
[0105] The loss calculation module 505 may calculate the loss based on the authenticity prediction result of the image and the image authenticity label corresponding to the spliced feature, where the image authenticity label corresponding to the spliced feature may be the image authenticity label of the training image corresponding to the first feature included in the spliced feature, that is, the image authenticity label corresponding to the spliced feature may be the image authenticity label of the training image from which the forgery-related feature included in the spliced feature comes.
[0106] The training module 506 can train the image forgery detection model by adjusting the parameters of the backbone network, FBEM, and classifier according to the loss.
[0107] According to an exemplary embodiment of the present disclosure, the splicing module 503 can splice the first feature and the second feature from the same training image through FBEM to obtain a first spliced feature. Moreover, the splicing module 503 can also splice the first feature and the second feature from different training images through FBEM to obtain a second spliced feature.
[0108] Then, the feature input module 504 can input the first spliced feature into the classifier to obtain a first authenticity prediction result of the image; moreover, the feature input module 504 can also input the second spliced feature into the classifier to obtain a second authenticity prediction result of the image.
[0109] Next, the loss calculation module 505 can calculate a first loss based on the first authenticity prediction result and the image authenticity label corresponding to the first spliced feature; moreover, the loss calculation module 505 can also calculate a second loss based on the second authenticity prediction result and the image authenticity label corresponding to the second spliced feature. Then, the training module 506 can adjust the parameters of the backbone network, FBEM, and classifier according to the first loss and the second loss.
[0110] According to an exemplary embodiment of the present disclosure, the training module 506 can perform weighted summation on the first loss and the second loss to obtain a weighted loss. Then, the training module 506 can adjust the parameters of the backbone network, FBEM, and classifier according to the weighted loss.
[0111] According to an exemplary embodiment of the present disclosure, the training device 500 for the image forgery detection model based on structural causality may further include a conditional mutual information calculation module. The conditional mutual information calculation module can calculate the conditional mutual information between the first feature and the second feature decoupled from two training images , where the conditional mutual information can be used to characterize the decoupling degree between the first feature and the second feature, that is, the conditional mutual information can be used to characterize the decoupling degree between the forgery-related features and the forgery-unrelated features. Moreover, the smaller the conditional mutual information , the higher the decoupling degree between the features. Therefore, one of the objectives of the present disclosure is to minimize the conditional mutual information to ensure sufficient decoupling between the forgery-related features and the forgery-unrelated features.
[0112] Then, the training module 506 can adjust according to the foregoing loss and conditional mutual information Adjust the parameters of the backbone network, FBEM, and classifier. Specifically, the training module 506 can calculate the weighted sum result of the aforementioned loss and conditional mutual information, and then adjust the parameters of the backbone network, FBEM, and classifier according to the weighted sum result.
[0113] Thus, in the present disclosure, by introducing the forgery bias elimination loss function and the conditional mutual information regularization term, while ensuring the detection performance of the image forgery detection model, the computational cost can also be reduced, enabling the structure-causality-based image forgery detection method provided by the present disclosure to be seamlessly integrated into existing systems.
[0114] Figure 6 FIG. is a block diagram showing a structure-causality-based image forgery detection apparatus 600 according to an exemplary embodiment of the present disclosure. The structure-causality-based image forgery detection apparatus 600 can be implemented based on an image forgery detection model trained according to the training method of the present disclosure, wherein the image forgery detection model can include a backbone network, a forgery bias elimination module (FBEM), and a classifier.
[0115] Referring to Figure 6 , the structure-causality-based image forgery detection apparatus 600 may include a target image acquisition module 601, a feature decoupling module 602, a feature splicing module 603, and a genuine / fake result prediction module 604.
[0116] The target image acquisition module 601 can acquire a target image to be subjected to image forgery detection. Exemplarily, the target image to be subjected to image forgery detection can be a human image, an animal image, a natural landscape image, etc., and the present disclosure does not specifically limit the type of the target image to be subjected to image forgery detection.
[0117] The feature decoupling module 602 can extract the target depth feature of the target image through the backbone network, and can decouple the target depth feature through the FBEM to obtain the decoupled first target feature and second target feature. Specifically, two projection layers included in the FBEM: φ and ψ can be used to decouple the target depth feature into the first target feature and the second target feature. Moreover, the first target feature can be a forgery-related feature related to the image tampering area, and the second target feature can be a forgery-unrelated feature unrelated to the image tampering area.
[0118] The feature splicing module 603 can splice the first target feature and the second target feature through the FBEM to obtain a target spliced feature. That is, the forgery-related feature and the forgery-unrelated feature in the target image can be spliced through the FBEM to obtain a target spliced feature.
[0119] The authenticity result prediction module 604 can input the target stitching feature into a classifier to obtain the authenticity prediction result of the target image. That is, the target stitching feature can be input into the classifier to perform binary classification on the target image through the classifier, that is, to predict whether the target image is an unaltered real image or a forged image that has been tampered with.
[0120] Figure 7 is a block diagram showing an electronic device 700 according to an exemplary embodiment of the present disclosure.
[0121] Referring to Figure 7 , the electronic device 700 includes at least one memory 701 and at least one processor 702. Instructions are stored in the at least one memory 701. When the instructions are executed by the at least one processor 702, a training method for an image forgery detection model based on structural causality or an image forgery detection method based on structural causality according to an exemplary embodiment of the present disclosure is executed.
[0122] As an example, the electronic device 700 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instructions. Here, the electronic device 700 does not have to be a single electronic device, and may also be a collection of devices or circuits that can individually or jointly execute the above instructions (or instruction sets). The electronic device 700 may also be part of an integrated control system or a system manager, or may be configured as a portable electronic device that can be interconnected with a local or remote (e.g., via wireless transmission) interface.
[0123] In the electronic device 700, the processor 702 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0124] The processor 702 can run instructions or code stored in the memory 701, where the memory 701 can also store data. The instructions and data can also be sent and received via a network interface device through a network, where the network interface device can adopt any known transmission protocol.
[0125] The memory 701 can be integrated with the processor 702. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor, etc. In addition, the memory 701 can include independent devices, such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The memory 701 and the processor 702 can be operatively coupled, or can communicate with each other, for example, through an I / O port, a network connection, etc., such that the processor 702 can read files stored in the memory.
[0126] In addition, the electronic device 700 may further include a video display (such as, a liquid crystal display) and a user interaction interface (such as, a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 700 may be connected to each other via a bus and / or a network.
[0127] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above-mentioned training method of the structure-causal based image forgery detection model or the structure-causal based image forgery detection method. Examples of the computer-readable storage medium herein include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or a computer such that the processor or the computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium may run in an environment deployed in computer devices such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0128] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including a computer program which, when executed by a processor, implements the training method of the structure-causal based image forgery detection model or the structure-causal based image forgery detection method according to the present disclosure.
[0129] Based on the structure-causal image forgery detection and training method and related devices of the present disclosure, by decoupling the overall features of training images with different contents respectively, and splicing the forgery-related features and forgery-unrelated features decoupled out to train the image forgery detection model, it can be ensured that the trained image forgery detection model can accurately identify the tampered areas in the image, that is, the confusion interference brought by natural areas irrelevant to image forgery detection can be effectively excluded, so as to improve the accuracy of image forgery detection.
[0130] According to an exemplary embodiment of the present disclosure, by tampering with the original image to train the image forgery detection model, it can be ensured that the trained image forgery detection model can effectively perform binary classification prediction on the image, that is, it can be ensured that the trained image forgery detection model can effectively distinguish whether the image to be detected is a forged image that has been tampered with or a real image that has not been tampered with.
[0131] According to an exemplary embodiment of the present disclosure, by setting a variety of different image tampering methods, it can be ensured that the trained image forgery detection model can cope with various different tampering scenarios, and the generalization of the image forgery detection model is better.
[0132] According to an exemplary embodiment of the present disclosure, by setting a forgery bias elimination module (FBEM), effective decomposition of deep features can be achieved, that is, effective decoupling of forgery-related features and forgery-unrelated features can be ensured, and further the practicability and flexibility of the image forgery detection model in different tampering scenarios can be improved.
[0133] According to an exemplary embodiment of the present disclosure, by introducing a forgery bias elimination loss function and a conditional mutual information regularization term, while ensuring the detection performance of the image forgery detection model, the calculation cost can also be reduced, so that the structure-causal image forgery detection method provided by the present disclosure can be seamlessly integrated into the existing system.
[0134] Those skilled in the art will readily conceive of other implementations of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0135] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A training method for an image forgery detection model based on structural causality, characterized in that, The image forgery detection model includes a backbone network, a forgery bias elimination module (FBEM), and a classifier. The training method includes: Obtain training image samples, where each training image sample includes two training images with different contents, and each training image is corresponding to an image authenticity label and the coordinate positions indicating the image tampering areas; For each of the two training images, extract the depth features of the training image through the backbone network, and decouple the depth features of the training image based on the corresponding coordinate positions through the FBEM to obtain a first feature and a second feature after decoupling, where the first feature is a forgery-related feature related to the image tampering area, and the second feature is a forgery-unrelated feature unrelated to the image tampering area; Through the FBEM, splice any of the first features and the second features extracted from the two training images to obtain a spliced feature; Input the spliced feature into the classifier to obtain the authenticity prediction result of the image; Calculate the loss based on the authenticity prediction result of the image and the image authenticity label corresponding to the spliced feature, where the image authenticity label corresponding to the spliced feature is the image authenticity label of the training image corresponding to the first feature included in the spliced feature; Train the image forgery detection model by adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss.
2. The training method according to claim 1, wherein The step of obtaining a spliced feature by splicing any of the first features and the second features extracted from the two training images through the FBEM includes: Splice the first feature and the second feature from the same training image through the FBEM to obtain a first spliced feature; Splice the first features and the second features from different training images through the FBEM to obtain a second spliced feature; Among them, input the spliced feature into the classifier to obtain the authenticity prediction result of the image; Input the first spliced feature into the classifier to obtain a first authenticity prediction result of the image; Input the second spliced feature into the classifier to obtain a second authenticity prediction result of the image; Among them, the step of calculating the value of the loss function based on the authenticity prediction result of the image and the image authenticity label corresponding to the spliced feature includes: Calculate a first loss based on the first authenticity prediction result and the image authenticity label corresponding to the first spliced feature; Calculate a second loss based on the second authenticity prediction result and the image authenticity label corresponding to the second spliced feature; The step of adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss includes: Adjust the parameters of the backbone network, the FBEM, and the classifier according to the first loss and the second loss.
3. The training method according to claim 2, wherein The step of adjusting the parameters of the backbone network, the FBEM, and the classifier according to the first loss and the second loss includes: Perform weighted summation on the first loss and the second loss to obtain a weighted loss; By adjusting the parameters of the backbone network, the FBEM, and the classifier according to the weighted loss.
4. The training method according to claim 1, wherein The training method further includes: Calculating the conditional mutual information between the first feature and the second feature decoupled from the two training images, where the conditional mutual information is used to characterize the degree of decoupling between the first feature and the second feature; The adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss includes: Adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss and the conditional mutual information.
5. The training method according to claim 4, characterized in that The adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss and the conditional mutual information includes: Calculating the weighted summation result of the loss and the conditional mutual information; Adjusting the parameters of the backbone network, the FBEM, and the classifier according to the weighted summation result.
6. The training method according to claim 1, wherein The obtaining the training image samples includes: Obtaining two original images with different contents; Tampering with the two original images respectively to obtain the training image samples.
7. The training method according to claim 6, wherein The way of tampering includes at least one of the following items: Adding Gaussian noise, adding jitter, splicing, copy-pasting, erasing.
8. An image forgery detection method based on structural causality, characterized in that, The image forgery detection method based on structural causality is implemented based on an image forgery detection model, and the image forgery detection model includes a backbone network, a forgery bias elimination module FBEM, and a classifier. The image forgery detection method based on structural causality includes: Obtaining a target image to be detected for image forgery; Extracting the target depth feature of the target image through the backbone network, and decoupling the target depth feature through the FBEM to obtain the decoupled first target feature and second target feature, where the first target feature is a forgery-related feature related to the image tampering area, and the second target feature is a forgery-unrelated feature unrelated to the image tampering area; Splicing the first target feature and the second target feature through the FBEM to obtain a target splicing feature; Inputting the target splicing feature into the classifier to obtain the authenticity prediction result of the target image.
9. A training device for an image forgery detection model based on structural causality, characterized in that, The image forgery detection model includes a backbone network, a forgery bias elimination module FBEM, and a classifier. The training device includes: An image sample acquisition module configured to acquire training image samples, where the training image samples include two training images with different contents, and each training image corresponds to an image authenticity label and the coordinate position for indicating the image tampering area; A decoupling module configured to, for each of the two training images, extract the depth feature of the training image through the backbone network, and decouple the depth feature of the training image based on the coordinate position corresponding to the training image through the FBEM to obtain the decoupled first feature and second feature, where the first feature is a forgery-related feature related to the image tampering area, and the second feature is a forgery-unrelated feature unrelated to the image tampering area; A splicing module, configured to splice any of the first features and the second features extracted from the two training images through the FBEM to obtain spliced features; A feature input module, configured to input the spliced features into the classifier to obtain a true / false prediction result of the image; A loss calculation module, configured to calculate a loss based on the true / false prediction result of the image and the true / false label of the image corresponding to the spliced features, where the true / false label of the image corresponding to the spliced features is the true / false label of the training image corresponding to the first feature included in the spliced features; A training module, configured to train the image forgery detection model by adjusting the parameters of the backbone network, the FBEM, and the classifier according to the loss.
10. An image forgery detection device based on structural causality, characterized in that, The structure-causality-based image forgery detection device is implemented based on an image forgery detection model, the image forgery detection model includes a backbone network, a forgery bias elimination module FBEM, and a classifier, and the structure-causality-based image forgery detection device includes: A target image acquisition module, configured to acquire a target image to be subjected to image forgery detection; A feature decoupling module, configured to extract target depth features of the target image through the backbone network and decouple the target depth features through the FBEM to obtain a decoupled first target feature and a second target feature, where the first target feature is a forgery-related feature related to the image tampering area, and the second target feature is a forgery-unrelated feature unrelated to the image tampering area; A feature splicing module, configured to splice the first target feature and the second target feature through the FBEM to obtain a target spliced feature; A true / false result prediction module, configured to input the target spliced feature into the classifier to obtain a true / false prediction result of the target image.
11. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the training method of the structure-causality-based image forgery detection model according to any one of claims 1 to 7, or to implement the structure-causality-based image forgery detection method according to claim 8.
12. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the training method of the structure-causality-based image forgery detection model according to any one of claims 1 to 7, or to execute the structure-causality-based image forgery detection method according to claim 8.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the training method of the structure-causality-based image forgery detection model according to any one of claims 1 to 7, or implements the structure-causality-based image forgery detection method according to claim 8.
Citation Information
Patent Citations
Face forgery detection method based on depth information decoupling
CN116704580A
Image authentic identification model training method, image authentic identification method, image authentic identification device and electronic equipment
CN117133039A