Three-branch image tampering trace detection neural network method based on cross-semantic feature fusion

By introducing a three-branch neural network method that integrates cross-semantic features in image tampering trace detection, the problem of poor detection performance and low-dimensional features being ignored in the prior art is solved, and higher detection accuracy and generalization are achieved.

CN120219929APending Publication Date: 2025-06-27王诗雨
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311823311.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing image tampering trace detection technology has problems such as inaccurate labeling, simple tampering process, poor detection performance and ignoring low-dimensional local features, resulting in poor detection capabilities.

Method used

A three-branch image tampering trace detection neural network method based on cross-semantic feature fusion is proposed. By tampering with striped semantic feature extraction network, camera-type semantic feature extraction network, and manipulating type semantic feature extraction network, fusing low-dimensional local features and high-dimensional global features, and optimizing the convolution module to improve detection accuracy.

Benefits of technology

It improves the accuracy and generalization of image tampering trace detection, enhances the ability to identify tampering traces, and effectively combats the spread of malicious tampering images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219929A_ABST
    Figure CN120219929A_ABST
Patent Text Reader

Abstract

The invention aims to provide a three-branch image tampering trace detection neural network method based on cross-semantic feature fusion, which is characterized in that a tampering strip semantic feature extraction network, a camera type semantic feature extraction network and a manipulation type semantic feature extraction network are used for extracting a tampering trace; respectively extracting tampering strip semantic features, camera type semantic features and manipulation type semantic features, and supervising and urging the model to fully extract low-dimensional local features of different areas in a tampering image; and meanwhile, a convolution module is optimized, low-dimensional features and high-dimensional global features in the tampered image are fused, and missing detection and false detection of tampering traces caused by insufficient features in an existing scheme are optimized. The invention innovatively provides a three-branch image tampering trace detection neural network method based on cross-semantic feature fusion, an existing scheme is optimized, the accuracy of image tampering trace detection is improved, the phenomenon of malicious tampering and spreading of images in society is attacked, and the method has great application value and theoretical prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of digital image forensics and information security for image tampering detection, and particularly to a three-branch neural network method for detecting image tampering traces based on cross-semantic feature fusion. Background Art

[0002] The increasingly advanced image processing technology has made the threshold of digital image editing lower and lower. With easily accessible image processing software, people can conveniently modify the content of images, and the tampered images are often very realistic, making it difficult to identify them with the naked eye. These tampered images have posed serious threats to personal privacy, social order, and national security. Therefore, detecting and locating the tampered areas in images has important practical significance.

[0003] However, there are two problems with existing image tampering trace detection technologies. On the one hand, the existing datasets for training have inaccurate annotations, and the tampering process is too simple and crude, and it can be judged whether it is tampered with by the naked eye, which is very different from the tampered images in real life. This results in unsatisfactory detection performance; on the other hand, existing deep learning-based methods only focus on high-dimensional global feature information and ignore more general low-dimensional local feature information. Even though attempts have been made to fuse the feature information of the two dimensions, there are still problems such as incomplete extraction of low-dimensional features and a large semantic span between low-dimensional features and high-dimensional features, resulting in poor detection ability of the model.

[0004] To solve the above problems, based on three semantic features: tampering stripe semantics, camera type semantics, and manipulation type semantics, the present invention proposes a three-branch neural network method for detecting image tampering traces based on cross-semantic feature fusion, which makes full use of low-dimensional local features to improve the generalization and accuracy of the algorithm. Summary of the Invention

[0005] The present invention proposes a three-branch neural network method for detecting image tampering traces based on cross-semantic feature fusion, which uses a tampering stripe semantic feature extraction network, a camera type semantic feature extraction network, and a manipulation type semantic feature extraction network to extract tampering stripe semantic features, camera type semantic features, and manipulation type semantic features respectively. Through the neural network, the low-dimensional local features are transformed into high-dimensional semantic features, enriching the feature information for discrimination in the overall network structure and bridging the semantic gap between low-dimensional features and high-dimensional features.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A three-branch neural network method for detecting image tampering traces based on cross-semantic feature fusion, comprising the following three steps:

[0008] Step 1: Produce a high-fidelity tampered image dataset;

[0009] Step 2: Design a three-branch image tampering trace detection neural network model based on cross-semantic feature fusion;

[0010] Step 3: Train a three-branch image tampering trace detection neural network model based on cross-semantic feature fusion;

[0011] The specific steps are as follows:

[0012] Step 1: Produce a high-fidelity tampered image dataset

[0013] (1) Use the Photoshop_python_api in Python to interact with Adobe Photoshop software and insert the image into the layer.

[0014] (2) Call the DoAction(·) method in the Win32com library in Python to execute the recorded operation actions related to Adobe Photoshop software, open Adobe Photoshop software and perform batch processing of image tampering according to the tampered area mask image.

[0015] (3) Finally, call the Composite(·) method and Save(·) method in the Psd_tools library in Python to perform format conversion on the image, and batch output the images processed by Adobe Photoshop to produce high-fidelity tampered image samples.

[0016] (4) Judge the pixels marked as positive examples in the semantic segmentation result. If there are heterogeneous pixels in its eight-neighborhood, mark this pixel as a tampered area edge pixel. Then perform the same judgment on the pixels in the original area to obtain the original area edge pixels. Combining the two, the ground truth map of the high-fidelity tampered image dataset is obtained.

[0017] (5) Produce 120,000 pairs of high-fidelity tampered images and their ground truth maps to obtain the high-fidelity tampered image dataset.

[0018] Step 2: Design a three-branch image tampering trace detection neural network model based on cross-semantic feature fusion

[0019] The present invention first designs a self-similarity convolution module and adds it to the existing U-Net network framework to become a semantic feature extraction network. Then, the image is input into the tampered strip semantic feature extraction network, the camera type semantic feature extraction network and the manipulation type semantic feature extraction network to obtain the corresponding tampered semantic feature maps. Finally, the tampered semantic feature maps obtained by the three semantic feature extraction networks are stacked and input into the cross-semantic feature fusion network to obtain the final image tampering trace detection result.

[0020] (1) Self-similar Convolution Module

[0021] The self-similar convolution module fuses global self-similarity convolution, sorting convolution, and ordinary convolution in a ratio of 1:2:2. The fused features pass through two consecutive ordinary convolution layers, then through a batch normalization layer, and finally undergo a non-linear transformation through an activation function.

[0022] The global self-similarity convolution extracts the intermediate feature layer of size H×W×C, scatters it according to the spatial distribution to obtain H×W feature vectors with a channel number of C, and then calculates the similarity between pairwise vectors. For any feature vector, the top k most similar vectors are taken on a global scale to form a feature matrix of size H×W×C×k, and convolution operations are performed through an ordinary convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1 to fuse the feature information of k feature blocks. The sorting convolution sorts the 3×3 pixels within a local range in descending order of similarity to the center point, clusters the features within the local range, and then performs convolution operations through an ordinary convolution with a convolution kernel size of 3×3 and a stride of 1 to obtain the self-similarity features on a local scale. The convolution kernel size of the ordinary convolution is 3×3 and the stride is 1.

[0023] (2) Semantic Feature Extraction Network

[0024] (2.1) Semantic Feature Extraction Network Framework

[0025] The Semantic Feature Extraction Network (SFEN) consists of 8 layers connected in series:

[0026] The first, second, third, and fourth layers are all composed of a self-similar convolution module, a convolution layer, and a pooling layer connected in series in sequence. The convolution kernel size in the convolution layer is 3×3 and the stride is 1. The kernel size in the pooling layer is 2×2. The input feature maps of the first, second, third, fourth, and fifth layers are 256×256×3, 128×128×64, 64×64×128, 32×32×256 respectively, and the output feature maps are 128×128×64, 64×64×128, 32×32×256, 16×16×512 respectively;

[0027] The fifth, sixth, seventh, and eighth layers are all composed of a self-similar convolution module, an upsampling layer, a convolution layer, and an activation function connected in series in sequence. The convolution kernel size in the convolution layer is 3×3, and the stride is 1. The input feature maps of the fifth, sixth, seventh, and eighth layers are 16×16×512, 32×32×256, 64×64×128, and 128×128×64 respectively, and the output feature maps are 32×32×256, 64×64×128, 128×128×64, and 256×256×1 respectively.

[0028] (2.2) Semantic Feature Extraction Network

[0029] The input image is I0, which passes through the tampered stripe semantic feature extraction network SFEN band , the camera type semantic feature extraction network SFEN cam and the manipulation type semantic feature extraction network SFEN man to obtain the respective intermediate feature maps and the tampered semantic feature maps. The specific formulas are as follows:

[0030]

[0031]

[0032]

[0033] Among them, f band1 , f band2 , f band3 are the feature maps of the seventh, sixth, and fifth layers of the tampered stripe semantic feature extraction network; is the prediction result of the tampered stripe semantic feature extraction network; f cam1 , f cam2 , f cam3 are the feature maps of the seventh, sixth, and fifth layers of the camera type semantic feature extraction network; is the prediction result of the camera type semantic feature extraction network; f man1 , f man2 , f man3 are the feature maps of the seventh, sixth, and fifth layers of the manipulation type semantic feature extraction network; is the prediction result of the manipulation type semantic feature extraction network.

[0034] (3) Cross-semantic Feature Fusion Network

[0035] (3.1) Cross-semantic Feature Fusion Network Framework

[0036] The cross-semantic feature fusion network (Cross-semantic Feature Fusion Network, CSFF) consists of 8 layers connected in series in sequence:

[0037] The first, second, third, and fourth layers are all composed of a self-similar convolution module, a convolutional layer, and a pooling layer connected in series in sequence. The convolutional kernel size in the convolutional layer is 3×3, and the stride is 1. The kernel size in the pooling layer is 2×2. The input feature maps of the first, second, third, fourth, and fifth layers are 256×256×3, 128×128×256, 64×64×512, 32×32×1024 respectively, and the output feature maps are 128×128×64, 64×64×128, 32×32×256, 16×16×512 respectively;

[0038] The fifth, sixth, seventh, and eighth layers are all composed of a self-similar convolution module, an upsampling layer, a convolutional layer, and an activation function connected in series in sequence. The convolutional kernel size in the convolutional layer is 3×3, and the stride is 1. The input feature maps of the fifth, sixth, seventh, and eighth layers are 16×16×512, 32×32×256, 64×64×128, 128×128×64 respectively, and the output feature maps are 32×32×256, 64×64×128, 128×128×64, 256×256×1 respectively.

[0039] (3.2) Cross-semantic Feature Fusion Network

[0040] Stack the intermediate feature maps of the tampered strip semantic feature extraction network, the camera type semantic feature extraction network, and the manipulation type semantic feature extraction network, as well as the tampered semantic feature map, and then pass through the cross-semantic feature fusion network CSFF to obtain the final image tampering trace detection result. The specific formula is as follows:

[0041]

[0042] I feat1 = concat(f band1 , f cam1 , f man1 )

[0043] I feat2 = concat(f band2 , f cam2 , fman2 )

[0044] I feat3 = concat(f band3 , f cam3 , f man3 )

[0045]

[0046] Among them, concat(·) represents stacking; f band1 , f band2, f band3 are the feature maps of the seventh, sixth, and fifth layers of the tampered stripe semantic feature extraction network; is the prediction result of the tampered stripe semantic feature extraction network; f cam1 , f cam2 , f cam3 are the feature maps of the seventh, sixth, and fifth layers of the camera type semantic feature extraction network; is the prediction result of the camera type semantic feature extraction network; f man1 , f man2 , f man3 are the feature maps of the seventh, sixth, and fifth layers of the manipulation type semantic feature extraction network; is the prediction result of the manipulation type semantic feature extraction network; I feat1 , I feat2 , I feat3 are the feature maps stacked with the output mappings of the first, second, and third layers of the cross-semantic feature fusion network respectively; is the prediction result of the cross-semantic feature fusion network.

[0047] Step 3: Train the three-branch image tampering trace detection neural network model based on cross-semantic feature fusion

[0048] (1) Training strategy

[0049] (1.1) Training of the tampered stripe semantic feature extraction network

[0050] This network branch mainly predicts "near-edge" pixels to alleviate the overall detection difficulty of the model. This branch uses the transition reinforcement loss function (L band ) to establish the constraint between the prediction result and the ground truth (GT) of the tampered stripe. The specific formula is as follows:

[0051]

[0052]

[0053]

[0054] where N represents the total number of pixels in the input tampered image; γ is the balance parameter; Mask is the spatial weight mask map; y band represents the tampered stripe map; is the prediction result of the tampered stripe semantic feature extraction network; log(·) represents the logarithmic function with base 2; ∑(·) represents the summation function; |·| represents getting the absolute value; and are the positive and negative sample weight parameters of the weighted cross-entropy loss function respectively. The specific formula is as follows:

[0055]

[0056]

[0057] Among them, N represents the total number of pixels in the input tampered image; N band represents the number of pixels of positive samples within the tampering strip ground truth (GT).

[0058] (1.2) Training of the camera type semantic feature extraction network

[0059] This network branch is trained on 40,000 camera type tampered images to learn the camera type semantic features in different regions of the images and predict the class information of different regions in the form of semantic segmentation. The cross-entropy loss function is used for the training of this network branch, and the specific formula is as follows:

[0060]

[0061] Among them, N represents the total number of pixel points in the input image; log(·) represents the logarithmic function with base 2; ∑(·) represents the summation function; y cam represents the ground truth (GT) of the semantic segmentation of the tampered image based on the camera type; is the prediction result of the camera type semantic feature extraction network.

[0062] (1.3) Training of the manipulation type semantic feature extraction network

[0063] This network branch is trained on 40,000 manipulation type tampered images to learn the manipulation type features in different regions of the images and predict the class information of different regions in the form of semantic segmentation. The cross-entropy loss function is used for the training of this network branch, and the specific formula is as follows:

[0064]

[0065] Among them, N represents the total number of pixel points in the input image; ∑(·) represents the summation function; log(·) represents the logarithmic function with base 2; y man represents the ground truth (GT) of the semantic segmentation of the tampered image based on the manipulation type; is the prediction result of the manipulation type semantic feature extraction network.

[0066] (1.4) Training of the cross-semantic feature fusion network

[0067] The cross-semantic feature fusion network stacks the tampered semantic feature maps output by three semantic feature extraction networks as inputs. At the same time, the intermediate feature maps of the three semantic feature maps are also stacked and input into the cross-semantic feature fusion network. A constraint is established between the prediction result of the cross-semantic feature fusion network and the tampered double-edge ground truth (GT). The loss function is the weighted cross-entropy loss function and the Dice Loss. The specific formula is as follows:

[0068]

[0069]

[0070]

[0071] where N represents the total number of pixels in the input tampered image; α is the balance parameter, defaulting to 0.6; y edges represents the tampered double-edge ground truth (GT); is the prediction result of the cross-semantic feature fusion network; log(·) represents the logarithmic function with base 2; ∑(·) represents the summation function; |·| represents obtaining the absolute value; and are the positive and negative sample weight parameters of the weighted cross-entropy loss function respectively, and their formulas are as follows:

[0072]

[0073]

[0074] where N edges represents the number of pixels of positive samples within the tampered edge ground truth (GT).

[0075] (2) Training parameter settings

[0076] The optimizer is SGD (Stochastic Gradient Descent), the initial learning rate is 0.007, the training image size is 256*256*3, the batch size is 128, and the training is carried out for 200 epochs. The high-fidelity tampered image dataset is normalized and then input into the model. The normalization method is to divide each pixel by 255.

[0077] Through the constraint of the loss function, the network parameters are updated using the deep learning framework Pytorch. When the output result converges, the training is stopped.

[0078] The beneficial effects of the present invention are as follows:

[0079] The present invention aims to provide a three-branch image forgery trace detection neural network method based on cross-semantic feature fusion. It uses a forgery stripe semantic feature extraction network, a camera type semantic feature extraction network, and a manipulation type semantic feature extraction network to extract forgery stripe semantic features, camera type semantic features, and manipulation type semantic features respectively, urging the model to fully extract the low-dimensional local features of different regions in the forged image. At the same time, the convolutional module is optimized to fuse the low-dimensional features and high-dimensional global features in the forged image, optimizing the situation of missed detection and false detection of forgery traces caused by insufficient features in the existing scheme. The present invention innovatively proposes a three-branch image forgery trace detection neural network method based on cross-semantic feature fusion, optimizing the existing scheme, improving the accuracy of image forgery trace detection, combating the phenomenon of maliciously forging and spreading images in society, and having great application value and theoretical prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] In order to enable those skilled in the art to more quickly and clearly understand the above and / or other objects, features, advantages and examples of the present application, some drawings are provided. It should be noted that the drawings forming part of the specification of the present application, the schematic embodiments and their descriptions are used to provide a further understanding of the present application and do not constitute an improper limitation of the present application.

[0081] Figure 1 is a sample of the forged image dataset for Step 1;

[0082] Figure 2 is the three-branch image forgery trace detection neural network model based on cross-semantic feature fusion for Step 2; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] The technical solution adopted by the present invention is as follows:

[0084] A three-branch image forgery trace detection neural network method based on cross-semantic feature fusion includes the following three steps:

[0085] Step 1: Produce a highly realistic forged image dataset;

[0086] Step 2: Design a three-branch image forgery trace detection neural network model based on cross-semantic feature fusion;

[0087] Step 3: Train the three-branch image forgery trace detection neural network model based on cross-semantic feature fusion;

[0088] The specific steps are as follows:

[0089] Step 1: Produce a highly realistic forged image dataset

[0090] (1) Interact with Adobe Photoshop software using the Photoshop_python_api in Python to insert an image into a layer.

[0091] (2) Call the DoAction(·) method in the Win32com library in Python to execute the recorded operation actions related to Adobe Photoshop software, open Adobe Photoshop software, and perform batch image tampering according to the tampered area mask image.

[0092] (3) Finally, call the Composite(·) method and Save(·) method in the Psd_tools library in Python to convert the image format, batch output the images processed by Adobe Photoshop, and produce highly realistic tampered image samples.

[0093] (4) Judge the pixels labeled as positive examples in the semantic segmentation result. If there are heterogeneous pixels in its eight-neighborhood, then label this pixel as a tampered area edge pixel. Then perform the same judgment on the pixels in the original area to obtain the original area edge pixels. Combining the two, the ground truth map of the highly realistic tampered image dataset is obtained.

[0094] (5) Produce 120,000 pairs of highly realistic tampered images and their ground truth maps to obtain a highly realistic tampered image dataset.

[0095] Step 2: Design a three-branch image tampering trace detection neural network model based on cross-semantic feature fusion

[0096] The present invention first designs a self-similarity convolution module and adds it to the existing U-Net network framework to become a semantic feature extraction network. Then, input the image into the tampered strip semantic feature extraction network, the camera type semantic feature extraction network, and the manipulation type semantic feature extraction network to obtain the corresponding tampered semantic feature maps. Finally, stack the tampered semantic feature maps obtained by the three semantic feature extraction networks and input them into the cross-semantic feature fusion network to obtain the final image tampering trace detection result.

[0097] (1) Self-similarity convolution module

[0098] The self-similarity convolution module fuses global self-similarity convolution, sorting convolution, and ordinary convolution in a ratio of 1:2:2. The fused features pass through two consecutive ordinary convolution layers, then through a batch normalization layer, and finally through an activation function for non-linear transformation.

[0099] The global self-similarity convolution extracts the intermediate feature layer of size H×W×C, scatters it according to the spatial distribution to obtain H×W feature vectors with the number of channels C, and then calculates the similarity between the vectors pairwise. For any feature vector, the top k most similar vectors are taken from the global scale to form a feature matrix of size H×W×C×k, and convolution operation is performed through a normal convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1 to fuse the feature information of k feature blocks. The sorting convolution sorts the 3×3 pixels within the local range in descending order according to the similarity with the center point, clusters the features within the local range, and then performs convolution operation through a normal convolution with a convolution kernel size of 3×3 and a stride of 1 to obtain the features of self-similarity at the local scale. The convolution kernel size of the normal convolution is 3×3 and the stride is 1.

[0100] (2) Semantic Feature Extraction Network

[0101] (2.1) Framework of Semantic Feature Extraction Network

[0102] The Semantic Feature Extraction Network (SFEN) consists of 8 layers connected in series:

[0103] The first, second, third, and fourth layers are all composed of a self-similarity convolution module, a convolution layer, and a pooling layer connected in series. The convolution kernel size in the convolution layer is 3×3 and the stride is 1. The kernel size in the pooling layer is 2×2. The input feature maps of the first, second, third, fourth, and fifth layers are 256×256×3, 128×128×64, 64×64×128, 32×32×256 respectively, and the output feature maps are 128×128×64, 64×64×128, 32×32×256, 16×16×512 respectively;

[0104] The fifth, sixth, seventh, and eighth layers are all composed of a self-similarity convolution module, an upsampling layer, a convolution layer, and an activation function connected in series. The convolution kernel size in the convolution layer is 3×3 and the stride is 1. The input feature maps of the fifth, sixth, seventh, and eighth layers are 16×16×512, 32×32×256, 64×64×128, 128×128×64 respectively, and the output feature maps are 32×32×256, 64×64×128, 128×128×64, 256×256×1 respectively.

[0105] (2.2) Semantic Feature Extraction Network

[0106] The input image is I0, passing through the tampered stripe semantic feature extraction network SFEN band and the camera type semantic feature extraction network SFEN camWith the manipulation type semantic feature extraction network SFEN man , the respective intermediate feature maps and tampering semantic feature maps are obtained, and the specific formulas are as follows:

[0107]

[0108]

[0109]

[0110] Among them, f band1 , f band2 , f band3 are the feature maps of the seventh, sixth, and fifth layers of the tampering stripe semantic feature extraction network; is the prediction result of the tampering stripe semantic feature extraction network; f cam1 , f cam2 , f cam3 are the feature maps of the seventh, sixth, and fifth layers of the camera type semantic feature extraction network; is the prediction result of the camera type semantic feature extraction network; f man1 , f man2 , f man3 are the feature maps of the seventh, sixth, and fifth layers of the manipulation type semantic feature extraction network; is the prediction result of the manipulation type semantic feature extraction network.

[0111] (3) Cross-semantic feature fusion network

[0112] (3.1) Cross-semantic feature fusion network framework

[0113] The cross-semantic feature fusion network (CSFF) consists of 8 layers connected in series in sequence:

[0114] The first, second, third, and fourth layers are each composed of a self-similar convolution module, a convolution layer, and a pooling layer connected in series in sequence. The convolution kernel size in the convolution layer is 3×3, the stride is 1, and the kernel size in the pooling layer is 2×2. The input feature maps of the first, second, third, fourth, and fifth layers are 256×256×3, 128×128×256, 64×64×512, 32×32×1024 respectively, and the output feature maps are 128×128×64, 64×64×128, 32×32×256, 16×16×512 respectively;

[0115] The fifth, sixth, seventh, and eighth layers are all composed of a self - similar convolution module, an up - sampling layer, a convolution layer, and an activation function connected in series in sequence. The convolution kernel size in the convolution layer is 3×3, and the stride is 1. The input feature maps of the fifth, sixth, seventh, and eighth layers are 16×16×512, 32×32×256, 64×64×128, and 128×128×64 respectively, and the output feature maps are 32×32×256, 64×64×128, 128×128×64, and 256×256×1 respectively.

[0116] (3.2) Cross - semantic Feature Fusion Network

[0117] Stack the intermediate feature maps of the tampered stripe semantic feature extraction network, the camera type semantic feature extraction network, and the manipulation type semantic feature extraction network, as well as the tampered semantic feature map, and then pass through the cross - semantic feature fusion network CSFF to obtain the final image tampering trace detection result. The specific formula is as follows:

[0118]

[0119] I feat1 = concat(f band1 , f cam1 , f man1 )

[0120] I feat2 = concat(f band2 , f cam2 , f man2 )

[0121] I feat3 = concat(f band3 , f cam3 , f man3 )

[0122]

[0123] Among them, concat(·) represents stacking; f band1 , f band2 , f band3 are the feature maps of the seventh, sixth, and fifth layers of the tampered stripe semantic feature extraction network; is the prediction result of the tampered stripe semantic feature extraction network; f cam1 , f cam2 , f cam3 are the feature maps of the seventh, sixth, and fifth layers of the camera type semantic feature extraction network; is the prediction result of the camera type semantic feature extraction network; f man1 , f man2 , f man3Are the feature maps of the seventh, sixth, and fifth layers of the manipulation type semantic feature extraction network; Is the prediction result of the manipulation type semantic feature extraction network; I feat1 、I feat2 、I feat3 Are the feature maps stacked with the output mappings of the first, second, and third layers of the cross-semantic feature fusion network respectively; Is the prediction result of the cross-semantic feature fusion network.

[0124] Step 3: Train a three-branch image forgery trace detection neural network model based on cross-semantic feature fusion

[0125] (1) Training strategy

[0126] (1.1) Training of the forgery stripe semantic feature extraction network

[0127] This network branch mainly predicts "near-edge" pixels to alleviate the overall detection difficulty of the model. This branch uses a transition reinforcement loss function (L band ) to establish the constraint between the prediction result and the forgery stripe ground truth (GT), and the specific formula is as follows:

[0128]

[0129]

[0130]

[0131] where N represents the total number of pixels in the input forged image; γ is a balance parameter; Mask is a spatial weight mask map; y band represents the forgery stripe map; is the prediction result of the forgery stripe semantic feature extraction network; log(·) represents the logarithmic function with base 2; ∑(·) represents the summation function; |·| represents getting the absolute value; and are the positive and negative sample weight parameters of the weighted cross-entropy loss function respectively, and the specific formula is as follows:

[0132]

[0133]

[0134] where N represents the total number of pixels in the input forged image; N band represents the number of positive sample pixels within the forgery stripe ground truth (GT).

[0135] (1.2) Training of the camera type semantic feature extraction network

[0136] This network branch is trained on 40,000 camera-type tampered images to learn the semantic features of different regions in the images and predict the class information of different regions in the form of semantic segmentation. The cross-entropy loss function is used for the training of this network branch, and the specific formula is as follows:

[0137]

[0138] where N represents the total number of pixel points in the input image; log(·) represents the logarithmic function with base 2; ∑(·) represents the summation function; y cam represents the semantic segmentation ground truth (GT) of the tampered image based on the camera type; is the prediction result of the camera-type semantic feature extraction network.

[0139] (1.3) Training of the manipulation-type semantic feature extraction network

[0140] This network branch is trained on 40,000 manipulation-type tampered images to learn the manipulation-type features of different regions in the images and predict the class information of different regions in the form of semantic segmentation. The cross-entropy loss function is used for the training of this network branch, and the specific formula is as follows:

[0141]

[0142] where N represents the total number of pixel points in the input image; ∑(·) represents the summation function; log(·) represents the logarithmic function with base 2; y man represents the semantic segmentation ground truth (GT) of the tampered image based on the manipulation type; is the prediction result of the manipulation-type semantic feature extraction network.

[0143] (1.4) Training of the cross-semantic feature fusion network

[0144] The cross-semantic feature fusion network stacks the tampered semantic feature maps output by three semantic feature extraction networks as inputs. At the same time, the intermediate feature maps of the three semantic feature maps are also stacked and input into the cross-semantic feature fusion network. A constraint is established between the prediction result of the cross-semantic feature fusion network and the tampered double-edge ground truth (GT), and its loss function is the weighted cross-entropy loss function and Dice Loss. The specific formula is as follows:

[0145]

[0146]

[0147]

[0148] Where N represents the total number of pixels in the input tampered image; α is a balance parameter, defaulting to 0.6; y edges represents the ground truth (GT) of the tampered double edges; is the prediction result of the cross-semantic feature fusion network; log(·) represents the logarithmic function with base 2; ∑(·) represents the summation function; |·| represents obtaining the absolute value; and are the positive and negative sample weight parameters of the weighted cross-entropy loss function respectively, and their formulas are as follows:

[0149]

[0150]

[0151] where N edges represents the number of pixels of positive samples within the ground truth (GT) of the tampered edges.

[0152] (2) Training parameter settings

[0153] The optimizer is SGD (Stochastic Gradient Descent), the initial learning rate is 0.007, the training image size is 256*256*3, the batch size is 128, and it is trained for 200 epochs. The high-fidelity tampered image dataset is normalized and then input into the model, and the normalization method is to divide each pixel by 255.

[0154] Through the constraint of the loss function, the network parameters are updated using the deep learning framework Pytorch, and the training is stopped when the output result converges.

[0155] The specific embodiments described herein are merely illustrative of the present invention. Those skilled in the art to which the present invention pertains can make various modifications, supplements, or use similar methods for substitution to the specific embodiments described herein, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

[0156] Although detailed descriptions of the present invention have been made and some specific embodiments have been cited, it is obvious that various changes or modifications can be made to those skilled in the art as long as they do not depart from the spirit and scope of the present invention.

[0157] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present application are included within the protection scope of the present application.

[0158] Matters not covered by this invention are all well-known technologies.

Claims

1. A three-branch image forgery trace detection neural network method based on cross-semantic feature fusion, comprising the following three steps: Step 1: Produce a high-fidelity forgery image dataset; Step 2: Design a three-branch image forgery trace detection neural network model based on cross-semantic feature fusion; Step 3: Train a three-branch image forgery trace detection neural network model based on cross-semantic feature fusion.

2. The three-branch image forgery trace detection neural network method based on cross-semantic feature fusion according to claim 1, wherein in Step 1 of producing the high-fidelity forgery dataset, the following steps are included: (1) Use Photoshop_python_api in Python to interact with Adobe Photoshop software and insert the image into the layer. (2) Call the DoAction(·) method in the Win32com library in Python to execute the recorded operation actions related to Adobe Photoshop software, open Adobe Photoshop software and perform batch processing of image forgery according to the forgery area mask image. (3) Finally, call the Composite(·) method and Save(·) method in the Psd_tools library in Python to perform format conversion on the image, and batch output the images processed by Adobe Photoshop to produce high-fidelity forgery image samples. (4) Judge the pixels labeled as positive examples in the semantic segmentation result. If there are heterogeneous pixels in its eight-neighborhood, then label this pixel as a forgery area edge pixel. Then perform the same judgment on the pixels in the original area to obtain the original area edge pixels. Combining the two, the ground truth map of the high-fidelity forgery image dataset is obtained. (5) Produce 120,000 pairs of high-fidelity forgery images and their ground truth maps to obtain a high-fidelity forgery image dataset.

3. The three-branch image forgery trace detection neural network method based on cross-semantic feature fusion according to claim 1, wherein in Step 2 of designing a three-branch image forgery trace detection neural network model based on cross-semantic feature fusion, the following steps are included: The present invention first designs a self-similarity convolution module and adds it to the existing U-Net network framework to become a semantic feature extraction network. Then, the image is input into the forgery stripe semantic feature extraction network, the camera type semantic feature extraction network and the manipulation type semantic feature extraction network to obtain the corresponding forgery semantic feature maps. Finally, the forgery semantic feature maps obtained by the three semantic feature extraction networks are stacked and input into the cross-semantic feature fusion network to obtain the final image forgery trace detection result. (1) Self-similarity convolution module The self-similarity convolution module fuses global self-similarity convolution, sorting convolution and ordinary convolution in a ratio of 1:2:

2. The fused features pass through two consecutive ordinary convolution layers, then through a batch normalization layer, and finally through an activation function for non-linear transformation. The global self-similarity convolution extracts the intermediate feature layer of size \(H\times W\times C\), scatters it according to the spatial distribution to obtain \(H\times W\) feature vectors with the number of channels \(C\), and then calculates the similarity between the vectors pairwise. For any feature vector, the top \(k\) closest vectors are taken on the global scale to form a feature matrix of size \(H\times W\times C\times k\), and convolution operation is performed through a normal convolution with a convolution kernel size of \(3\times3\), a stride of \(1\), and a padding of \(1\) to fuse the feature information of \(k\) feature blocks. The sorting convolution sorts the \(3\times3\) pixels within the local range from large to small according to the similarity with the center point, clusters the features within the local range, and then performs convolution operation through a normal convolution with a convolution kernel size of \(3\times3\) and a stride of \(1\) to obtain the features of self-similarity on the local scale. The convolution kernel size of the normal convolution is \(3\times3\) and the stride is \(1\). (2) Semantic Feature Extraction Network (2.1) Semantic Feature Extraction Network Framework The Semantic Feature Extraction Network (SFEN) consists of 8 layers connected in series in turn: The first, second, third, and fourth layers are all composed of a self-similarity convolution module, a convolution layer, and a pooling layer connected in series in turn. The convolution kernel size in the convolution layer is \(3\times3\) and the stride is \(1\). The kernel size in the pooling layer is \(2\times2\). The input feature maps of the first, second, third, fourth, and fifth layers are \(256\times256\times3\), \(128\times128\times64\), \(64\times64\times128\), \(32\times32\times256\) respectively, and the output feature maps are \(128\times128\times64\), \(64\times64\times128\), \(32\times32\times256\), \(16\times16\times512\) respectively; The fifth, sixth, seventh, and eighth layers are all composed of a self-similarity convolution module, an upsampling layer, a convolution layer, and an activation function connected in series in turn. The convolution kernel size in the convolution layer is \(3\times3\) and the stride is \(1\). The input feature maps of the fifth, sixth, seventh, and eighth layers are \(16\times16\times512\), \(32\times32\times256\), \(64\times64\times128\), \(128\times128\times64\) respectively, and the output feature maps are \(32\times32\times256\), \(64\times64\times128\), \(128\times128\times64\), \(256\times256\times1\) respectively. (2.2) Semantic Feature Extraction Network The input image is I0, which passes through the tampered strip semantic feature extraction network SFEN band , the camera type semantic feature extraction network SFEN cam and the manipulation type semantic feature extraction network SFEN man , obtaining the respective intermediate feature maps and the tampered semantic feature maps. The specific formulas are as follows: Among them, f band1 , f band2 , f band3 are the feature maps of the seventh, sixth, and fifth layers of the tampered stripe semantic feature extraction network; is the prediction result of the tampered stripe semantic feature extraction network; f cam1 , f cam2 , f cam3 are the feature maps of the seventh, sixth, and fifth layers of the camera type semantic feature extraction network; is the prediction result of the camera type semantic feature extraction network; f man1 , f man2 , f man3 are the feature maps of the seventh, sixth, and fifth layers of the manipulation type semantic feature extraction network; is the prediction result of the manipulation type semantic feature extraction network. (3) Cross-semantic Feature Fusion Network (3.1) Cross-semantic Feature Fusion Network Framework The Cross-semantic Feature Fusion Network (CSFF) consists of 8 layers connected in series in turn: The first, second, third, and fourth layers are all composed of a self - similar convolution module, a convolution layer, and a pooling layer connected in series in sequence. In the convolution layer, the convolution kernel size is 3×3 and the stride is 1. In the pooling layer, the kernel size is 2×2. The input feature maps of the first, second, third, fourth, and fifth layers are 256×256×3, 128×128×256, 64×64×512, 32×32×1024 respectively, and the output feature maps are 128×128×64, 64×64×128, 32×32×256, 16×16×512 respectively; The fifth, sixth, seventh, and eighth layers are all composed of a self - similar convolution module, an up - sampling layer, a convolution layer, and an activation function connected in series in sequence. In the convolution layer, the convolution kernel size is 3×3 and the stride is 1. The input feature maps of the fifth, sixth, seventh, and eighth layers are 16×16×512, 32×32×256, 64×64×128, 128×128×64 respectively, and the output feature maps are 32×32×256, 64×64×128, 128×128×64, 256×256×1 respectively. (3.2) Cross - semantic Feature Fusion Network Stack the intermediate feature maps of the tampering stripe semantic feature extraction network, the camera type semantic feature extraction network, and the manipulation type semantic feature extraction network, as well as the tampering semantic feature map, and then pass through the cross - semantic feature fusion network CSFF to obtain the final image tampering trace detection result. The specific formula is as follows: I feat1 = concat(f band1 , f cam1 , f man1 ) I feat2 = concat(f band2 , f cam2 , f man2 ) I feat3 = concat(f band3 , f cam3 , f man3 ) Among them, concat(·) represents stacking; f band1 、f band2 、f band3 are the feature maps of the seventh, sixth, and fifth layers of the semantic feature extraction network for tampered strips; is the prediction result of the semantic feature extraction network for tampered strips; f cam1 、f cam2 、f cam3 are the feature maps of the seventh, sixth, and fifth layers of the semantic feature extraction network for camera types; is the prediction result of the semantic feature extraction network for camera types; f man1 、f man2 、f man3 are the feature maps of the seventh, sixth, and fifth layers of the semantic feature extraction network for manipulation types; is the prediction result of the semantic feature extraction network for manipulation types; I feat1 、I feat2 、I feat3 are the feature maps stacked with the output mappings of the first, second, and third layers of the cross-semantic feature fusion network respectively; is the prediction result of the cross-semantic feature fusion network.

4. According to the method for a three - branch image tampering trace detection neural network based on cross - semantic feature fusion as claimed in claim 1, in step three, training the three - branch image tampering trace detection neural network model based on cross - semantic feature fusion includes the following steps: (1) Training Strategy (1.1) Training of the Tampering Stripe Semantic Feature Extraction Network This network branch mainly predicts "near-edge" pixels to ease the overall detection difficulty of the model. This branch uses a transition-enhanced loss function (L band ) to establish the constraint between the prediction result and the ground truth (GT) of the tampered strip. The specific formula is as follows: where N represents the total number of pixels in the input tampered image; γ is the balance parameter; Mask is the spatial weight mask map; y band represents the tampered strip map; is the prediction result of the tampered strip semantic feature extraction network; log(·) represents the logarithmic function with base 2; ∑(·) represents the summation function; |·| represents getting the absolute value; and are the positive and negative sample weight parameters of the weighted cross-entropy loss function respectively, and the specific formula is as follows: Where N represents the total number of pixels in the input tampered image; N band represents the number of pixels of positive samples within the tampering stripe ground truth (GT). (1.2) Training of the Camera Type Semantic Feature Extraction Network This network branch is trained on 40,000 camera - type tampered images to learn the camera - type semantic features in different regions of the images and predict the class information of different regions in the form of semantic segmentation. The cross - entropy loss function is used for the training of this network branch. The specific formula is as follows: where N represents the total number of pixel points in the input image; log(·) represents the logarithmic function with base 2; ∑(·) represents the summation function; y cam represents the semantic segmentation ground truth (GT) of the tampered image based on the camera type; is the prediction result of the camera type semantic feature extraction network. (1.3) Training of the Manipulation Type Semantic Feature Extraction Network This network branch is trained on 40,000 manipulation - type tampered images to learn the manipulation - type features in different regions of the images and predict the class information of different regions in the form of semantic segmentation. The cross - entropy loss function is used for the training of this network branch. The specific formula is as follows: where N represents the total number of pixel points in the input image; ∑(·) represents the summation function; log(·) represents the logarithmic function with base 2; y man represents the semantic segmentation ground truth (GT) of the tampered image based on the manipulation type; is the prediction result of the manipulation type semantic feature extraction network. (1.4) Training of the Cross - semantic Feature Fusion Network The cross - semantic feature fusion network stacks the tampering semantic feature maps output by the three semantic feature extraction networks as inputs. At the same time, the intermediate feature maps of the three semantic feature maps are also stacked and input into the cross - semantic feature fusion network. A constraint is established between the prediction result of the cross - semantic feature fusion network and the tampering double - edge ground truth (GT). Its loss function is the weighted cross - entropy loss function and Dice Loss. The specific formula is as follows: Where N represents the total number of pixels in the input tampered image; α is a balance parameter, defaulting to 0.6; y edges represents the ground truth (GT) of the tampered double edges; is the prediction result of the cross-semantic feature fusion network; log(·) represents the logarithmic function with base 2; ∑(·) represents the summation function; |·| represents obtaining the absolute value; and are the positive and negative sample weight parameters of the weighted cross-entropy loss function respectively, and its formula is as follows: Where N edges represents the number of pixels of positive samples within the tampering edge ground truth (GT). (2) Training Parameter Settings The optimizer is SGD (Stochastic Gradient Descent), the initial learning rate is 0.007, the training image size is 256*256*3, the batch size is 128, and the training is carried out for 200 epochs. The high-fidelity tampered image dataset is normalized before being input into the model, and the normalization method is to divide each pixel by 255. Through the constraint of the loss function, the network parameters are updated using the deep learning framework Pytorch, and the training is stopped when the output results converge.