Image forgery detection method, device, server and storage medium
By performing multi-level residual quantization and domain feature extraction on the image, and forged detection of dynamic fusion features, the problems of low computational efficiency and insufficient generalization capabilities of convolutional neural networks in multi-background and multi-distribution domain image detection are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411524515.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-10-30
AI Technical Summary
The forgery detection method based on convolutional neural networks is inefficient in computing efficiency, and the relationship between the characteristics of different network structure layers is not understood, resulting in insufficient generalization capabilities of the model and low output accuracy.
By extracting the input images, performing multi-level residual quantization, capturing the quantized particle size distribution features, performing domain feature extraction, dynamically fusing shared features and unique features, and classifying them to output fake detection results.
It improves the robustness and accuracy of model output in complex scenarios, effectively deal with the diverse background and cross-domain feature distribution in the image, and understands the relationship between features through feature fusion, without deepening the model network structure and improving the generalization ability of the model.
Smart Images

Figure CN119048846B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an image forgery detection method, apparatus, server and storage medium. Background Art
[0002] Image forgery detection is a key research topic in the fields of computer vision and image processing. There are various types of forged images, involving multiple forms such as tampering, splicing, ghosting, face swapping, etc. These forgery techniques can be used to create misleading information, thus having a profound impact on fields such as news, law, medicine, and social media. In image forgery detection, multi-background and cross-domain scenarios refer to the situation where the background information and feature distributions in the image are complex and diverse. This makes it difficult for traditional handcrafted feature-based detection methods to handle. For example, in the forgery detection task when dealing with multi-background images, it is necessary to distinguish the subtle differences between the complex background and foreground objects to prevent background interference in detection. When the forgery detection task deals with cross-domain scenario images, different regions in the image may adopt different editing techniques (such as splicing, tampering, deep forgery, etc.), and the feature distribution differences between these regions make the detection task more complex. In related technologies, the forgery detection method based on convolutional neural network has made remarkable progress in overcoming the detection tasks under multi-background and cross-domain scenarios by automatically learning image features. Compared with traditional methods, the forgery detection method based on convolutional neural network can handle more complex scenarios and forgery forms. However, when facing images with multi-background and multi-distribution domains, the model often requires a deeper network structure, which brings problems of computational efficiency. Moreover, when extracting features from different layers of the network structure, due to the failure to understand the relationship between features, the generalization ability of the model is insufficient, and the output accuracy of the model is low when facing new scenarios. Summary of the Invention
[0003] The present invention provides an image forgery detection method, apparatus, server and storage medium to solve the defects of the forgery detection method based on convolutional neural network, such as low computational efficiency, failure to understand the relationship between features of different network structure layers, resulting in insufficient generalization ability of the model, and low output accuracy of the model when facing new scenarios.
[0004] The present invention provides an image forgery detection method, including:
[0005] Performing feature extraction on the input image to obtain initial image features;
[0006] Performing multi-level residual quantization on the initial image features to capture the quantization granularity distribution features in the input image and obtain a quantization granularity distribution feature image;
[0007] Extract domain features from the quantized granularity distribution feature image to obtain shared features and unique features. The shared features are the features included in at least two domains in the quantized granularity distribution feature image, and the unique features are the features included in each domain in the quantized granularity distribution feature image;
[0008] Dynamically fuse the shared features and the unique features, classify the features after the dynamic fusion, and output a forgery detection result according to the classification result.
[0009] According to the image forgery detection method provided by the present invention, the multi-level residual quantization of the initial image features to capture the quantized granularity distribution features in the input image includes:
[0010] Input the input image into a convolutional neural network to obtain an initial feature representation, and create a multi-layer residual codebook, where the residual codebook includes multiple center vectors;
[0011] Calculate the distance between each feature vector in the current layer and all the center vectors in the residual codebook, and regard each feature vector as a sample of the center vector closest to it;
[0012] After all the samples in each center vector are assigned, update the position of the center vector according to the average value of all the samples belonging to the same center vector;
[0013] Calculate the residual value between each feature vector and the closest center vector, and use the residual value as the input for the next layer of iterative calculation until the convergence condition is met. Take the center vector in the residual codebook as the quantized granularity distribution feature of the input image.
[0014] According to the image forgery detection method provided by the present invention, the convergence condition includes:
[0015] The update amplitude of each center vector in each residual codebook is less than a preset threshold;
[0016] Or,
[0017] The number of times of the iterative calculation reaches a preset maximum number of times.
[0018] According to the image forgery detection method provided by the present invention, the multi-level residual quantization of the initial image features to capture the quantized granularity distribution features in the input image further includes:
[0019] Obtain a sensitivity index according to the difference between the residual value at a certain position in the current layer and the residual value at the same position in the next layer;
[0020] During the residual codebook update process, the sensitivity index is used to distinguish the sensitivity of the residual values of each layer when calculating the residual value between each feature vector and the nearest center vector.
[0021] According to the image forgery detection method provided by the present invention, the cross-domain feature extraction of the quantization granularity distribution features to obtain shared features includes:
[0022] Input the quantization granularity distribution features into a deep neural network, where the deep neural network is a multi-layer perceptron, and the multi-layer perceptron includes multiple hidden layers, activation functions, and an output layer;
[0023] The multiple hidden layers are used to fuse the quantization granularity distribution features of different levels through connection or weighted summation to generate a comprehensive feature representation;
[0024] The activation function is used to capture the non-linear relationship between the quantization granularity distribution features according to the comprehensive feature representation to obtain the final feature representation;
[0025] The output layer is used to compress the final feature representation to between 0 and 1 to obtain shared features.
[0026] According to the image forgery detection method provided by the present invention, the calculation method of the shared features includes:
[0027] Fuse the quantization granularity distribution features with the weight matrix of the deep neural network, and accumulate the fusion result with the bias term;
[0028] Perform an activation function calculation on the result of accumulating the fusion result and the bias term to obtain shared features.
[0029] According to the image forgery detection method provided by the present invention, the cross-domain feature extraction of the quantization granularity distribution features to obtain unique features includes:
[0030] Input the quantization granularity distribution features into a convolutional layer to extract local features;
[0031] Input the local features into a feature extraction network to generate unique features.
[0032] According to the image forgery detection method provided by the present invention, the calculation method of the unique features includes:
[0033] Obtain multiple basis functions in the feature extraction network and the weights corresponding to each basis function;
[0034] Take the result of multiplying and accumulating multiple basis functions with the corresponding weights as unique features.
[0035] According to the image forgery detection method provided by the present invention, the dynamic fusion of the shared features and the unique features includes:
[0036] Concatenate the shared features and the unique features to obtain a comprehensive feature;
[0037] Input the comprehensive feature into a fully connected layer or a self-attention mechanism, use the gating weights to weight the features of each layer in the fully connected layer or the self-attention mechanism, and perform weighted summation on the weighted features of all levels to obtain the dynamically fused features;
[0038] Among them, the gating weights are obtained according to the network learning parameters obtained after training based on the fully connected layer or the self-attention mechanism.
[0039] According to the image forgery detection method provided by the present invention, the calculation method of the gating weights includes:
[0040] Perform feature fusion on the comprehensive feature and the gating weight learning parameters in the fully connected layer or the self-attention mechanism, and accumulate the fusion result with the gating bias term learning parameters;
[0041] Obtain the gating weights according to the accumulation result.
[0042] According to the image forgery detection method provided by the present invention, the classification of the dynamically fused features and the output of the forgery detection result according to the classification result include:
[0043] Perform prediction on the dynamically fused features through an activation function : , where is the Sigmoid activation function, which is used to output the forgery probability value, and the final prediction result represents the probability that the image is forged;
[0044] According to the final prediction result and the threshold perform output result classification:
[0045]
[0046] Among them, is the output result, being 0 indicates that the image is determined to be a real image, being 1 indicates that the image is determined to be a forged image;
[0047] The threshold is set according to different scenarios and different requirements, specifically including:
[0048] When performing legal evidence image forgery detection, set the threshold to a first threshold, where the first threshold is used to increase the probability of outputting a forged image;
[0049] When performing news content image forgery detection, set the threshold to a second threshold, where the second threshold is used to balance the probabilities of outputting a forged image and a genuine image;
[0050] When performing social media image forgery detection, set the threshold to a third threshold, where the third threshold is used to decrease the probability of outputting a forged image;
[0051] Among them, the first threshold is less than the second threshold, and the second threshold is less than the third threshold.
[0052] According to the image forgery detection method provided by the present invention, the convolutional neural network is obtained after training. The method for obtaining the training data set of the convolutional neural network includes:
[0053] Obtain multiple original images, preprocess the multiple original images to obtain an initial data set. The preprocessing includes cropping the original images to a unified size;
[0054] Augment the images in the initial data set through random rotation, flipping, and adding noise operations to obtain an augmented data set;
[0055] Normalize the pixel values of the images in the augmented data set to obtain a training data set.
[0056] The present invention also provides an image forgery detection device, including:
[0057] A first extraction module, configured to extract features from an input image to obtain initial image features;
[0058] A quantization module, configured to perform multi-level residual quantization on the initial image features, capture the quantization granularity distribution features in the input image, and obtain a quantization granularity distribution feature image;
[0059] A second extraction module, configured to extract domain features from the quantization granularity distribution feature image to obtain shared features and unique features. The shared features are the features included in at least two domains in the quantization granularity distribution feature image, and the unique features are the features included in each domain in the quantization granularity distribution feature image;
[0060] A fusion module, configured to dynamically fuse the shared features and the unique features, classify the features after the dynamic fusion, and output a forgery detection result according to the classification result.
[0061] The present invention also provides a server, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the image forgery detection method described in any one of the above is implemented.
[0062] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the image forgery detection method described in any one of the above is implemented.
[0063] The image forgery detection method, device, server, and storage medium provided by the present invention obtain initial image features by performing feature extraction on an input image; perform multi-level residual quantization on the initial image features to capture the quantization granularity distribution features in the input image and obtain a quantization granularity distribution feature image; perform domain feature extraction on the quantization granularity distribution feature image to obtain shared features and unique features, where the shared features are the features included in at least two domains in the quantization granularity distribution feature image, and the unique features are the features included in each domain in the quantization granularity distribution feature image; perform dynamic fusion on the shared features and the unique features, classify the features after the dynamic fusion, and output a forgery detection result according to the classification result. The present invention adopts a design idea combining hierarchical feature extraction, step-by-step residual quantization, and dynamic fusion, improves the robustness and accuracy of the model output in complex scenarios, can effectively handle diverse backgrounds and cross-domain feature distributions in images, understands the relationship between features through feature fusion, does not require deepening the model network structure, improves the model generalization ability, and provides a more reliable solution for the forgery detection task. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0065] Figure 1 is a schematic flowchart of the image forgery detection method provided by the embodiment of the present invention;
[0066] Figure 2 is a schematic flowchart of the vector codebook generation method provided by the embodiment of the present invention;
[0067] Figure 3 is a schematic diagram of the step-by-step residual coding network structure provided by the embodiment of the present invention;
[0068] Figure 4It is a schematic diagram of the shared feature extraction network structure provided by an embodiment of the present invention;
[0069] Figure 5 It is a schematic diagram of the functional structure of the image forgery detection device provided by an embodiment of the present invention;
[0070] Figure 6 It is a schematic diagram of the functional structure of the electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0071] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0072] Figure 1 It is a flowchart of the image forgery detection method provided by an embodiment of the present invention. As Figure 1 shown, the image forgery detection method includes:
[0073] Step 101: Extract features from the input image to obtain initial image features;
[0074] Step 102: Perform multi-level residual quantization on the initial image features to capture the quantization granularity distribution features in the input image, and obtain a quantization granularity distribution feature image;
[0075] Step 103: Extract domain features from the quantization granularity distribution feature image to obtain shared features and unique features. The shared features are the features included in at least two domains in the quantization granularity distribution feature image, and the unique features are the features included in each domain in the quantization granularity distribution feature image;
[0076] In an embodiment of the present invention, the quantization granularity distribution feature image is divided into multiple spatial domains according to the feature distribution and feature semantics. The shared features obtained after extracting the domain features from the quantization granularity distribution feature image are the features included in at least two spatial domains in the quantization granularity distribution feature image, and the unique features are the features included in each spatial domain in the quantization granularity distribution feature image. For example, the quantization granularity distribution feature image is divided into 3 spatial domains according to the feature distribution and feature semantics. The features included in each spatial domain itself are unique features, and the features included in two or three spatial domains are shared features.
[0077] Step 104: Dynamically fuse the shared features and the unique features, classify the features after dynamic fusion, and output a forgery detection result according to the classification result.
[0078] Traditional image forgery detection methods and forgery detection methods based on convolutional neural networks have made significant progress in overcoming detection tasks in multi-background and cross-domain scenarios by automatically learning image features. Compared with traditional methods, forgery detection methods based on convolutional neural networks can handle more complex scenarios and forgery forms. However, when faced with images with multi-background and multi-distribution domains, the model often requires a deeper network structure, which brings problems of computational efficiency. Moreover, when extracting features from different layers of the network structure, due to the failure to understand the relationship between features, the generalization ability of the model is insufficient, and the output accuracy of the model is low when faced with new scenarios.
[0079] The image forgery detection method provided by the embodiments of the present invention extracts initial image features by performing feature extraction on the input image; performs multi-level residual quantization on the initial image features to capture the quantization granularity distribution features in the input image and obtain a quantization granularity distribution feature image; performs domain feature extraction on the quantization granularity distribution feature image to obtain shared features and unique features, where the shared features are the features included in at least two domains in the quantization granularity distribution feature image, and the unique features are the features included in each domain in the quantization granularity distribution feature image. The embodiments of the present invention adopt a design idea combining hierarchical feature extraction, step-by-step residual quantization, and dynamic fusion to improve the robustness and accuracy of the model output in complex scenarios, can effectively handle diverse backgrounds and cross-domain feature distributions in images, understand the relationship between features through feature fusion, do not require deepening the model network structure, improve the generalization ability of the model, and provide a more reliable solution for the forgery detection task.
[0080] Based on any of the above embodiments, the performing multi-level residual quantization on the initial image features to capture the quantization granularity distribution features in the input image includes:
[0081] Step 201: Input the input image into a convolutional neural network to obtain an initial feature representation, and create a multi-layer residual codebook, where the residual codebook includes multiple center vectors;
[0082] Step 202: Calculate the distance between each feature vector in the current layer and all center vectors in the residual codebook, and regard each feature vector as a sample of the center vector closest to it;
[0083] Step 203: After all samples in each center vector are assigned, update the position of the center vector according to the average value of all samples belonging to the same center vector;
[0084] Step 204: Calculate the residual value between each feature vector and the nearest central vector, and use the residual value as the input for the next-layer iterative calculation until the convergence condition is met. Then, use the central vector in the residual codebook as the quantization granularity distribution feature of the input image.
[0085] In the embodiments of the present invention, the convergence condition includes:
[0086] The update amplitude of each central vector in each residual codebook is less than a preset threshold; or, the number of times of the iterative calculation reaches a preset maximum number.
[0087] In the embodiments of the present invention, the hierarchical residual coding is a technique for processing and encoding image features layer by layer. Through the quantization operation of the residuals, it can sensitively capture the subtle changes in the image, thus laying a solid foundation for high-precision forgery detection. The specific steps of the hierarchical residual coding quantization include:
[0088] Initial residual setting: The input image , first undergoes preliminary processing by a convolutional neural network to obtain an initial feature representation . The specific expression is:
[0089]
[0090] Where represents the mapping function of the convolutional neural network. In this initial stage, the features extracted by the network usually cover the basic elements of the image, such as basic information like edges and textures.
[0091] Hierarchical coding and residual update: First, it is necessary to design the codebook of the layer. In the initial-layer codebook, record the distance between each vector sample and the central vector of the codebook. The residual codebook records the distance between each vector sample and the residual of the codebook vector sample. The purpose is to achieve high-precision codebook recording for the insufficient accuracy of the original codebook vector. For example, initialize the vector codebook . First, determine the size of the vector codebook, that is, the number of central vectors in the codebook. Initialize the vector codebook of the layer , where one vector codebook , and each vector is a -dimensional vector.
[0092] In the embodiments of the present invention, the initialization method can be randomly selecting training samples or using an initialization algorithm such as the K-means++ method. The present application does not limit the specific initialization method.
[0093] After initialization is complete, samples need to be assigned to the nearest centroid vector. For each sample (the output of the layer), calculate its distance (usually Euclidean distance) from all codebook vectors .
[0094] According to the formula , assign the sample to the nearest centroid vector , that is is assigned the value of the sample .
[0095] Once all samples have been assigned to the nearest centroid vectors, update each centroid vector to the mean of the samples belonging to that centroid vector: , where is the set of all samples assigned to the centroid vector , and is the number of samples in this set.
[0096] In each layer , update the residuals in real time. The update calculation method is: . Repeat the above steps, find the nearest encoding in the next layer's codebook for the recorded residuals and update the residual codebook vectors until the change in the centroid vectors is less than a certain preset threshold or the maximum number of iterations is reached. Finally, the codebook is the set of stable centroid vectors. After multiple iterations, the codebook
[0097] reaches convergence, that is, the update amplitude of each centroid vector of each vector codebook is small enough, or other stopping conditions are met (such as reaching the maximum number of iterations). At this time, the centroid vectors in the codebook are the best partition of the input data feature space. As
[0098] shown, according to the input feature vector Figure 2 , the nearest vector encoding can be found on the first-layer codebook. At the same time, according to the residual between it and the vector encoding, the vector encoding closest to the residual can be found in the next layer's codebook, and so on. Therefore, the sum of the input feature vector on the codebook , where , represents the vector closest to each level of residual in the codebook of each layer .
[0099] Through such multi-layer iterative calculations, it can be ensured that the residuals are fully and finely quantified, so that extremely subtle feature differences in the image can be accurately captured. The construction method of the residual codebook is the same as that of the initial vector codebook.
[0100] In a specific example, as Figure 3 shown, the quantization times is set to 8, and the size of a codebook in vector quantization is set to 64 x 256, that is , and the number of layers of the codebook is related to the quantization precision required. In this example, it is set to count = 6 layers.
[0101] In some embodiments of the present invention, the multi-level residual quantization of the initial image features to capture the quantization granularity distribution features in the input image further includes:
[0102] Obtaining a sensitivity index according to the difference between the residual value at a certain position in the current layer and the residual value at the same position in the next layer;
[0103] During the update process of the residual codebook, the sensitivity index is used to distinguish the sensitivity of the residual values of each layer when calculating the residual value between each feature vector and the nearest central vector.
[0104] In the embodiments of the present invention, progressive residual coding is a technique for processing and quantifying features layer by layer. By gradually reducing and updating the residuals, it captures subtle differences in the image. It has important applications in image forgery detection, which can improve the accuracy of feature extraction and the robustness of detection. Through a multi-stage coding process, progressive residual coding can effectively handle complex backgrounds and forgery forms, thereby improving the detection effect.
[0105] In the embodiments of the present invention, after careful processing of multi-level residual quantization, the finally obtained quantization residuals , show extremely high sensitivity to tiny forgery traces in the image. In order to quantify this sensitivity between each level of residuals, a sensitivity index is introduced, and the expression is: where represents the set of pixels in the image.
[0106] During the quantization process, the sensitivity index is introduced in the residual codebook update step to adjust the quantization weights, that is , the detection ability of the model is significantly enhanced by introducing sensitivity indicators. Moreover, a higher quantization accuracy is achieved for the residual components with higher sensitivity, thereby improving the detection accuracy. The residual components with lower sensitivity can moderately reduce the quantization accuracy and reduce the computational and storage overheads. In multi-background and cross-domain scenarios, quantization based on sensitivity indicators can dynamically adapt to the complexity of different regions and improve the robustness of the model.
[0107] Based on any of the above embodiments, the cross-domain feature extraction of the quantization granularity distribution features to obtain shared features includes:
[0108] Step 301, input the quantization granularity distribution features into a deep neural network, where the deep neural network is a multi-layer perceptron, and the multi-layer perceptron includes multiple hidden layers, activation functions, and an output layer;
[0109] Step 302, the multiple hidden layers are used to fuse the quantization granularity distribution features at different levels through connection or weighted summation to generate a comprehensive feature representation;
[0110] Step 303, the activation function is used to capture the non-linear relationship between the quantization granularity distribution features according to the comprehensive feature representation to obtain the final feature representation;
[0111] Step 304, the output layer is used to compress the final feature representation between 0 and 1 to obtain shared features.
[0112] In the embodiment of the present invention, the shared features have the following mathematical expression:
[0113]
[0114] where, is the quantization granularity distribution feature, is the activation function, is the weight matrix, is the bias term, is a multi-layer perceptron or a convolutional network.
[0115] In the embodiment of the present invention, as Figure 4 shown, the extraction process of the shared features is realized by means of the non-linear transformation function . In this process, the deep neural network (DNN) plays a key role and is usually used to obtain more complex and abstract features. The mathematical expression is: Here, represents the quantized feature, is the activation function (such as ReLU), is the weight matrix, is the bias term. And It can be a powerful multi-layer perceptron (MLP) or an efficient convolutional network, whose main role is to capture complex non-linear relationships in high-dimensional space. In this embodiment, it is a 5-layer MLP network, including 3 hidden layers (128, 64, and 32 units each), ReLU activation, and finally using Sigmoid output.
[0116] Based on any of the above embodiments, the cross-domain feature extraction of the quantization granularity distribution feature to obtain unique features includes:
[0117] Step 401: Input the quantization granularity distribution feature into the convolutional layer to extract local features;
[0118] Step 402: Input the local features into the feature extraction network to generate unique features.
[0119] In the embodiment of the present invention, the unique features have the following mathematical expression:
[0120]
[0121] where, is the quantization granularity distribution feature, represents the feature extraction function, is a parameter set for a specific distribution type, is the weight, is the basis function, and K is the number of basis functions.
[0122] In the embodiment of the present invention, the extraction of unique features closely depends on predefined distribution types, such as key factors like color and texture. By using convolutional layers or other specific network architectures such as the Visual Geometry Group network and Residual network, etc., to generate unique features, the specific mathematical expression is: where represents the unique feature extraction function, is a parameter set for a specific distribution type, is the weight, is the basis function.
[0123] The embodiment of the present invention accurately maps image features into a clear distribution space, thus creating favorable conditions for subsequent in-depth processing and accurate analysis.
[0124] After successfully completing the residual quantization operation, the next crucial step is to accurately extract cross-domain features from the quantized features. The core goal of this link is to generate a more comprehensive and rich feature representation, providing strong support for subsequent forgery detection tasks. The cross-domain feature extraction network structure provided by the embodiment of the present invention includes:
[0125] Input: Quantized features , whose dimensions are precisely defined as .
[0126] Shared feature extraction: A deep neural network (such as a multi-layer perceptron (MLP) with 5 hidden layers, where the number of neurons in each layer is 128, 64, and 32 respectively) is used to generate complex and abstract shared features .
[0127] Specific feature extraction: Through an efficient convolutional layer (such as a 3x3 convolutional kernel, stride of 1, and padding of 1) or a specific feature extraction network (such as a partial layer of VGG16 used in this example), specific features are accurately generated .
[0128] Output: Cross-domain features and , providing key inputs for subsequent feature fusion and weight calculation
[0129] In the embodiments of the present invention, the cross-domain feature extraction strategy aims to comprehensively capture the rich features of images, covering shared and specific cross-domain information. These carefully extracted features will provide richer and more comprehensive input information for subsequent forgery detection tasks, significantly enhancing the model's accurate recognition ability of forgery traces
[0130] Based on any of the above embodiments, the dynamic fusion of the shared features and the specific features includes:
[0131] Concatenating the shared features and the specific features to obtain a comprehensive feature
[0132] Inputting the comprehensive feature into a fully connected layer or a self-attention mechanism, weighting the features of each layer in the fully connected layer or the self-attention mechanism using gating weights, and performing weighted summation on all levels of weighted features to obtain the dynamically fused features
[0133] Wherein, the gating weights are obtained according to the network learning parameters obtained after training the fully connected layer or the self-attention mechanism
[0134] In the embodiments of the present invention, the calculation expression of the gating weights is:
[0135]
[0136] Wherein, is the comprehensive feature, and are the learning parameters of the gating weights, is the activation function
[0137] In the embodiment of the present invention, first, an advanced feature fusion technology is adopted to fuse the shared features and the unique features to generate comprehensive features. The mathematical expression is: Wherein, represents the feature concatenation operation, and the dimension of the fused features is precisely defined as (the dimensions of the shared and unique features are the same). Then, a gating mechanism is used to calculate the weight . The comprehensive feature is input into a highly flexible fully connected layer or self-attention mechanism, and the gating weight is obtained through network learning. The mathematical expression is: Here, and represent the learning parameters of the gating weight, and the function cleverly restricts the calculation result strictly within the range of [0, 1], thus providing a solid guarantee for the accurate calculation and effective application of the weight.
[0138] Make full use of the calculated gating weight to accurately weight the features of each layer to generate the final model output. The mathematical expression is: Here, represents the feature function of the th layer, and is the weighted feature, which provides the key output for the final prediction of the model. The weighted features of all levels are weighted and summed to obtain the final feature vector Wherein, is the processing function for the features of the th layer.
[0139] In the embodiment of the present invention, the network structure of dynamic fusion includes:
[0140] Input: The fused feature , whose dimension is precisely defined as .
[0141] Gating calculation network: A fully connected layer (with 128 neurons) or a self-attention mechanism (with 8 heads) is used to accurately calculate the gating weight . The output dimension is closely related to the specific number of tasks, that is dimensions.
[0142] Output: The weighted feature , which provides the key basis for the final prediction of the model.
[0143] In the embodiments of the present invention, the application of the gating mechanism cleverly introduces cross-domain information into the model by dynamically adjusting the feature weights, thereby significantly enhancing the robustness and accuracy of the model. The precise calculation of the gating weights can effectively integrate features at different levels, enabling the model to exhibit excellent adaptability when dealing with complex and diverse tasks.
[0144] During the training process, an optimization algorithm such as the efficient Adaptive Moment Estimation (Adam) algorithm is adopted. The learning rate is set to 0.001, the exponential decay rate of the first moment estimation is 0.9, and the exponential decay rate of the second moment estimation is 0.999 to precisely optimize the learning of the gating weights and ensure that the performance of the model reaches the optimal.
[0145] The embodiments of the present invention are applicable to processing images under different resolutions and different lighting conditions. The dynamic weight mechanism is adjusted dynamically according to the feature distribution in a specific scenario.
[0146] Based on any of the above embodiments, the output of the forgery detection result according to the dynamically fused features includes:
[0147] Predicting through an activation function for the dynamically fused features : , where is the Sigmoid activation function, used to output the forgery probability value, and the final prediction result represents the probability that the image is forged;
[0148] According to the final prediction result and the threshold to classify the output result:
[0149]
[0150] where, is the output result, being 0 indicates that the image is determined to be a real image, being 1 indicates that the image is determined to be a forged image;
[0151] In the embodiments of the present invention, a threshold is set to determine the classification result of the model output. The threshold usually takes values between 0 and 1. For example, a commonly used threshold is 0.5. If the output probability of the model is greater than or equal to , it is determined to be a forged image; otherwise, it is determined to be a real image. The selection of the threshold can be adjusted through cross-validation or on the validation set during the training phase to optimize the performance of the model.
[0152] In the embodiments of the present invention, the threshold Set according to different scenarios and different requirements, specifically including:
[0153] When performing legal evidence image forgery detection, set the threshold to the first threshold, and the first threshold is used to increase the probability of the output being a forged image;
[0154] In court, digital images are often used as evidence. The authenticity of these images directly affects the outcome of the case. Therefore, it is crucial to ensure that the images have not been tampered with in any way. In this case, it is very important to avoid false negatives (i.e., wrongly determining a forged image as genuine), because this may lead to miscarriages of justice. Even if it means that some genuine images may be misjudged as forged (false positives), it is necessary to minimize undetected forgeries. Therefore, the threshold can be set relatively low, such as 0.4. The purpose of doing this is to increase the sensitivity to potential forged images. Even if some genuine images are marked as suspicious, they can be confirmed through further forensic analysis.
[0155] When performing news content image forgery detection, set the threshold to the second threshold, and the second threshold is used to balance the probabilities of the output being a forged image and the output being a genuine image;
[0156] The content published by news agencies must maintain a high level of authenticity and credibility. If the pictures or videos published are found to be forged, it will not only damage the reputation of the media but also may trigger a public trust crisis. News content hopes to minimize false alarms for genuine content while ensuring the authenticity of the content. Therefore, it is necessary to find a balance point that can effectively identify forged content and minimize misjudgments of genuine content. Therefore, the threshold can be set to an intermediate value, such as 0.5. Such a setting aims to balance sensitivity and specificity, being able to capture most forged content without overly affecting the normal publication of genuine content. If the model performs well, this default value is usually a reasonable choice.
[0157] When performing social media image forgery detection, set the threshold to the third threshold, and the third threshold is used to decrease the probability of the output being a forged image;
[0158] Users upload a large amount of image and video content on social media platforms. The platforms need to automatically detect and remove posts containing forged content to maintain a healthy community environment. Social media platforms need to consider the user experience and avoid user dissatisfaction caused by a large number of false positives. At the same time, it is also necessary to prevent the widespread dissemination of harmful forged content. Therefore, it is necessary to control the proportion of false positives while ensuring high accuracy. Therefore, the threshold can be set relatively high, such as 0.6. The purpose of this is to reduce the situation of false positives and ensure that only when the model is very certain that an image is forged will it be marked. This helps to maintain user trust and effectively remove obvious forged content.
[0159] In image forgery detection, the selection of the threshold is crucial for balancing false positives and false negatives. Different application scenarios may need to adjust the threshold according to their specific requirements. Legal evidence authentication: The threshold is relatively low (such as 0.4) to ensure as few false negatives as possible. News media content review: The threshold is moderate (such as 0.5) to balance sensitivity and specificity. Social media platforms: The threshold is relatively high (such as 0.6) to reduce false positives and protect the user experience. By adjusting the threshold, the performance of the model can be optimized according to the needs of different application scenarios, so as to better serve the specific goals in practical applications.
[0160] Based on any of the above embodiments, the convolutional neural network is obtained after training. During the model training process, the L2 norm is used to calculate the loss function. The method for obtaining the training data set of the convolutional neural network includes:
[0161] Obtain multiple original images, and preprocess the multiple original images to obtain an initial data set. The preprocessing includes cropping the original images to a unified size;
[0162] In the embodiments of the present invention, the images are cropped to a unified size, such as 256x256 pixels, through image cropping and scaling to ensure that the images input into the model have the same size. At the same time, for images with too high a resolution, appropriate scaling is performed to reduce the amount of calculation and memory occupancy.
[0163] Augment the images in the initial data set through random rotation, flipping, and adding noise operations to obtain an augmented data set;
[0164] In the embodiments of the present invention, the data set is augmented through operations such as random rotation, flipping, and adding noise to increase the diversity of the data, thereby improving the generalization ability of the model. For example, in a specific instance of the present invention, the images are horizontally flipped with a probability of 50%, and randomly rotated with a probability of 30% (the rotation angle is between -30 degrees and 30 degrees).
[0165] Normalize the pixel values of the images in the augmented dataset to obtain the training dataset.
[0166] Before performing the image forgery detection task, data preprocessing is a crucial step. First, a large number of image datasets need to be collected, including real images and various types of forged images, such as splicing, copy-pasting, image enhancement, etc. For the collected images, the following preprocessing operations are performed:
[0167] In the embodiment of the present invention, the pixel values of the images are normalized to the range of [0, 1] to facilitate the training and convergence of the model. The calculation formula is: , where is the original pixel value, and are the minimum and maximum values of the image pixel values respectively.
[0168] The embodiment of the present invention also includes predicting the confidence of image forgery. For the image forgery task with only image-level annotations, it is assumed that only image-level labels are available during the training phase. Given an image, if this image does not contain forgery events, this image is defined as a normal image, and the label T = 0; otherwise, if the image contains at least one forgery event, then this image is labeled as forged, and the label T = 1.
[0169] The image forgery detection method provided by the embodiment of the present invention extracts features from the input image to obtain an initial image feature representation. The extracted image features are input into the residual encoding module for multi-level residual quantization to obtain a hierarchical multi-distribution representation of the quantization granularity. Then, different domain-level representations of the image are obtained: using two representation extraction methods, shared and specific, to further refine the hierarchical representation of the image. The processed image features are input into the hierarchical multi-distribution model, changing the calculation method and the adjustment method of the dynamic weight to adapt to the task requirements. Finally, it is judged whether there are forgery traces in the input image according to the model output result. It overcomes the problem of difficult forgery detection tasks in multi-background and cross-domain scenarios, and provides a more reliable solution for the forgery detection task.
[0170] Next, the image forgery detection device provided by the present invention will be described. The image forgery detection device described below can be mutually corresponded and referred to with the image forgery detection method described above.
[0171] Figure 5 is a schematic functional structure diagram of the image forgery detection device provided by the embodiment of the present invention. As Figure 5 shown, the image forgery detection device includes:
[0172] The first extraction module 501 is used to extract features from the input image to obtain the initial image features;
[0173] The quantization module 502 is used to perform multi-level residual quantization on the initial image features, capture the quantization granularity distribution features in the input image, and obtain a quantization granularity distribution feature image;
[0174] The second extraction module 503 is used to extract domain features from the quantization granularity distribution feature image to obtain shared features and unique features. The shared features are the features included in at least two domains in the quantization granularity distribution feature image, and the unique features are the features included in each domain in the quantization granularity distribution feature image;
[0175] The fusion module 504 is used to dynamically fuse the shared features and the unique features, classify the features after the dynamic fusion, and output a forgery detection result according to the classification result.
[0176] The image forgery detection device provided by the present invention obtains initial image features by performing feature extraction on the input image; performs multi-level residual quantization on the initial image features, captures the quantization granularity distribution features in the input image, and obtains a quantization granularity distribution feature image; extracts domain features from the quantization granularity distribution feature image to obtain shared features and unique features. The shared features are the features included in at least two domains in the quantization granularity distribution feature image, and the unique features are the features included in each domain in the quantization granularity distribution feature image; dynamically fuses the shared features and the unique features, classifies the features after the dynamic fusion, and outputs a forgery detection result according to the classification result. By adopting the design idea of combining hierarchical feature extraction, step-by-step residual quantization and dynamic fusion, the robustness and accuracy of the model output in complex scenarios are improved, and it can effectively cope with the diverse backgrounds and cross-domain feature distributions in the image. By understanding the relationship between features through feature fusion, it is not necessary to deepen the model network structure, and the model generalization ability is improved, providing a more reliable solution for the forgery detection task.
[0177] In the embodiment of the present invention, the first extraction module 501 is further used for: the input image , first undergoes preliminary convolutional neural network processing to obtain an initial feature representation . The specific expression is:
[0178] Where represents the mapping function of the convolutional neural network. At this initial stage, the features extracted by the network usually cover the basic elements of the image, such as basic information such as edges and textures.
[0179] In the embodiment of the present invention, the quantization module 502 is further used for:
[0180] Input the input image into a convolutional neural network to obtain an initial feature representation. For each layer of the convolutional neural network, create a residual codebook, which includes multiple center vectors; calculate the distances between each feature vector in the current layer and all the center vectors in the residual codebook, and assign each feature vector as a sample of the center vector it is closest to; after all samples in each center vector are assigned, update the position of the center vector according to the average value of all samples belonging to the same center vector; calculate the residual value between each feature vector and the closest center vector, and use the residual value as the input for the next layer's iterative calculation until the convergence condition is met, and use the center vectors in the residual codebook as the quantization granularity distribution features of the input image.
[0181] In an embodiment of the present invention, the convolutional neural network is obtained after training, and the method for obtaining the training data set of the convolutional neural network includes:
[0182] Obtain multiple original images, preprocess the multiple original images to obtain an initial data set, and the preprocessing includes cropping the original images to a unified size;
[0183] Augment the images in the initial data set through random rotation, flipping, and adding noise operations to obtain an augmented data set;
[0184] Normalize the pixel values of the images in the augmented data set to obtain a training data set.
[0185] In an embodiment of the present invention, progressive residual coding is a technique for processing and encoding image features layer by layer. Through quantization operations on residuals, it can sensitively capture subtle changes in images, thus laying a solid foundation for high-precision forgery detection. The specific steps of progressive residual coding quantization include:
[0186] First, it is necessary to design the codebook of the layer. In the initial layer codebook, record the distances between each vector sample and the center vector of the codebook. The residual codebook records the distances between each vector sample and the residuals of the codebook vector samples. The purpose is to achieve high-precision codebook recording for the inaccurate original codebook vectors. For example, initialize the vector codebook , first, determine the size of the vector codebook, that is, the number of center vectors in the codebook. Initialize the vector codebook of the layer , where one vector codebook , and each vector is a -dimensional vector.
[0187] In the embodiments of the present invention, the initialization method may be to randomly select training samples or use a certain initialization algorithm such as the K-means++ method. The present application does not limit the specific initialization method.
[0188] After the initialization is completed, samples need to be assigned to the nearest center vector. For each sample (the output of the th layer), calculate its distance from all codebook vectors (usually the Euclidean distance).
[0189] According to the formula , assign the sample to the nearest center vector , that is, is assigned the value of the sample .
[0190] Once all samples are assigned to the nearest center vector, update each center vector to the mean of the samples belonging to this center vector: , where is the set of all samples assigned to the center vector , and is the number of samples in this set.
[0191] In each layer , the residual is updated in real time, and the update calculation method is: , repeat the above steps, find the nearest code in the next-layer codebook for the recorded residual and update the residual codebook vector until the change in the center vector is less than a certain preset threshold or the maximum number of iterations is reached. Finally, the codebook is the set of stable center vectors. After multiple iterations, the codebook
[0192] reaches convergence, that is, the update amplitude of each center vector of each vector codebook is small enough, or other stopping conditions are met (such as reaching the maximum number of iterations). At this time, the center vectors in the codebook are the best partition of the input data feature space. In the embodiments of the present invention, the quantization module 502 is further configured to:
[0193] Obtain a sensitivity index according to the difference between the residual value at a certain position in the current layer and the residual value at the same position in the next layer;
[0194] During the update process of the residual codebook, use the sensitivity index to distinguish the sensitivity of the residual values of each layer when calculating the residual value between each feature vector and the nearest center vector.
[0195]
[0196] In the embodiment of the present invention, the second extraction module 503 is further configured to:
[0197] Input the quantization granularity distribution feature into a deep neural network, where the deep neural network is a multi-layer perceptron, and the multi-layer perceptron includes a plurality of hidden layers, an activation function, and an output layer;
[0198] The plurality of hidden layers are used to fuse the quantization granularity distribution features at different levels through connection or weighted summation to generate a comprehensive feature representation;
[0199] The activation function is used to capture the non-linear relationship between the quantization granularity distribution features according to the comprehensive feature representation to obtain the final feature representation;
[0200] The output layer is used to compress the final feature representation to between 0 and 1 to obtain a shared feature.
[0201] In the embodiment of the present invention, the shared feature has the following mathematical expression:
[0202]
[0203] where is the quantization granularity distribution feature, is the activation function, is the weight matrix, is the bias term, is a multi-layer perceptron or a convolutional network.
[0204] In the embodiment of the present invention, the second extraction module 503 is further configured to
[0205] Input the quantization granularity distribution feature into a convolutional layer to extract local features;
[0206] Input the local features into a feature extraction network to generate unique features.
[0207] In the embodiment of the present invention, the unique feature has the following mathematical expression:
[0208]
[0209] where is the quantization granularity distribution feature, represents a feature extraction function, is a parameter set for a specific distribution type, is the weight, is the basis function, and K is the number of basis functions.
[0210] In an embodiment of the present invention, the fusion module 504 is further configured to:
[0211] Concatenate the shared feature and the unique feature to obtain a comprehensive feature;
[0212] Input the comprehensive feature into a fully connected layer or a self-attention mechanism, weight the features of each layer in the fully connected layer or the self-attention mechanism using gating weights, and perform weighted summation on all levels of weighted features to obtain a dynamically fused feature;
[0213] Wherein, the gating weights are obtained according to network learning parameters obtained after training based on the fully connected layer or the self-attention mechanism.
[0214] In an embodiment of the present invention, the gating weights The calculation expression is:
[0215]
[0216] Wherein, Is the comprehensive feature, And Are the learning parameters of the gating weights, Is the activation function.
[0217] In an embodiment of the present invention, outputting a forgery detection result based on the dynamically fused feature includes:
[0218] Predict the dynamically fused feature through an activation function For prediction: , where Is the Sigmoid activation function, used to output a forgery probability value, and the final prediction result Represents the probability that the image is forged;
[0219] According to the final prediction result And a threshold Perform output result classification:
[0220]
[0221] Wherein, Is the output result, Being 0 indicates that the image is determined to be a real image, Being 1 indicates that the image is determined to be a forged image;
[0222] The threshold Is set according to different scenarios and different requirements, specifically including:
[0223] When performing forgery detection on legal evidence images, set the threshold to a first threshold, where the first threshold is used to increase the probability of outputting forged images;
[0224] When performing forgery detection on news content images, set the threshold to a second threshold, where the second threshold is used to balance the probabilities of outputting forged images and real images;
[0225] When performing forgery detection on social media images, set the threshold to a third threshold, where the third threshold is used to decrease the probability of outputting forged images;
[0226] Among them, the first threshold is less than the second threshold, and the second threshold is less than the third threshold.
[0227] In the embodiments of the present invention, the threshold is set according to different scenarios and different requirements, specifically including:
[0228] When performing forgery detection on legal evidence images, set the threshold to a first threshold, where the first threshold is used to increase the probability of outputting forged images;
[0229] In court, digital images are often used as evidence. The authenticity of these images directly affects the judgment result of the case. Therefore, it is crucial to ensure that the images have not been tampered with in any way. In this case, it is very important to avoid false negatives (i.e., wrongly determining a forged image as real), because this may lead to miscarriage of justice. Even if it means that some real images may be misjudged as forged (false positives), the situation of undetected forgeries needs to be minimized. Therefore, the threshold can be set relatively low, such as 0.4. The purpose of doing this is to increase the sensitivity to potential forged images. Even if some real images are marked as suspicious, they can be confirmed through further forensic analysis.
[0230] When performing forgery detection on news content images, set the threshold to a second threshold, where the second threshold is used to balance the probabilities of outputting forged images and real images;
[0231] The content released by news agencies must maintain a high level of authenticity and credibility. If the pictures or videos released are found to be forged, it will not only damage the reputation of the media but may also trigger a public trust crisis. While ensuring the authenticity of the content, news content also hopes to minimize false reports of real content. Therefore, a balance needs to be found that can effectively identify forged content while minimizing misjudgments of real content. Thus, the threshold can be set to an intermediate value, such as 0.5. Such a setting aims to balance sensitivity and specificity, being able to capture most forged content without overly affecting the normal release of real content. If the model performs well, this default value is usually a reasonable choice.
[0232] When performing social media image forgery detection, set the threshold to a third threshold, and the third threshold is used to reduce the probability of the output being a forged image;
[0233] Users upload a large amount of picture and video content on social media platforms. The platforms need to automatically detect and remove those posts containing forged content to maintain a healthy community environment. Social media platforms need to consider the user experience and avoid user dissatisfaction caused by a large number of false positives. At the same time, it is also necessary to prevent the widespread dissemination of harmful forged content. Therefore, while ensuring high accuracy, it is necessary to control the proportion of false positives. Thus, the threshold can be set relatively high, such as 0.6. The purpose of this is to reduce the situation of false positives and ensure that only when the model is very certain that an image is forged will it be marked. This helps to maintain user trust and can also effectively remove obvious forged content.
[0234] In image forgery detection, the choice of threshold is crucial for balancing false positives and false negatives. Different application scenarios may need to adjust the threshold according to their specific requirements. Legal evidence authentication: The threshold is relatively low (such as 0.4) to ensure as few false negatives as possible. News media content review: The threshold is moderate (such as 0.5) to balance sensitivity and specificity. Social media platforms: The threshold is relatively high (such as 0.6) to reduce false positives and protect the user experience. By adjusting the threshold, the performance of the model can be optimized according to the needs of different application scenarios, so as to better serve the specific goals in actual applications.
[0235] Figure 6 An example of the entity structure diagram of a server is shown as Figure 6As shown in the figure, the server may include: a processor 610, a communications interface 620, an image forgery detection device (memory) 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the image forgery detection device 630 communicate with each other through the communication bus 640. The image forgery detection device 630 includes computer programs, an operating system, and acquired data. The processor 610 can call the logical instructions in the image forgery detection device 630 to execute the image forgery detection method, which includes: extracting features from the input image to obtain initial image features; performing multi-level residual quantization on the initial image features to capture the quantization granularity distribution features in the input image; performing cross-domain feature extraction on the quantization granularity distribution features to obtain shared features and unique features; dynamically fusing the shared features and the unique features, and outputting a forgery detection result based on the features after dynamic fusion.
[0236] In addition, when the logical instructions in the above-mentioned image forgery detection device 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only image forgery detection devices (ROM, Read-Only Memory), random access image forgery detection devices (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0237] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the image forgery detection method provided by the above-mentioned various methods. The method includes: extracting features from the input image to obtain initial image features; performing multi-level residual quantization on the initial image features to capture the quantization granularity distribution features in the input image; performing cross-domain feature extraction on the quantization granularity distribution features to obtain shared features and unique features; dynamically fusing the shared features and the unique features, and outputting a forgery detection result based on the features after dynamic fusion.
[0238] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0239] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0240] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0241] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting image forgery, characterized in that: include: Extract features from the input image to obtain initial image features; Performing multi-level residual quantization on the initial image features to capture the quantized particle size distribution features in the input image and obtain a quantized particle size distribution feature image; Performing domain feature extraction on the quantitative particle size distribution feature image to obtain shared features and unique features, wherein the shared features are features included in at least two domains in the quantitative particle size distribution feature image, and the unique features are features included in each domain in the quantitative particle size distribution feature image; The shared features and the unique features are dynamically fused, the dynamically fused features are classified, and a forgery detection result is output according to the classification result.
2. The image forgery detection method according to claim 1, characterized in that: The performing multi-level residual quantization on the initial image features to capture the quantized particle size distribution features in the input image includes: Inputting the input image into a convolutional neural network to obtain an initial feature representation, and creating a multi-layer residual codebook, wherein the residual codebook includes multiple center vectors; Calculate the distance between each feature vector of the current layer and all the center vectors in the residual codebook, and take each feature vector as a sample of the nearest center vector; After all samples in each center vector are assigned, the position of the center vector is updated according to the average value of all samples belonging to the same center vector; The residual value between each feature vector and the nearest center vector is calculated, and the residual value is used as the input of the next layer of iterative calculation until the convergence condition is met, and the center vector in the residual codebook is used as the quantized granularity distribution feature of the input image.
3. The image forgery detection method according to claim 2, characterized in that: The convergence conditions include: The update amplitude of each center vector in each residual codebook is less than a preset threshold; or, The number of iterative calculations reaches a preset maximum number.
4. The image forgery detection method according to claim 2, characterized in that: The performing multi-level residual quantization on the initial image features to capture the quantized particle size distribution features in the input image also includes: The sensitivity index is obtained based on the difference between the residual value at a certain position in the current layer and the residual value at the same position in the next layer; In the residual codebook updating process, the sensitivity index is used to distinguish the sensitivity of each layer of residual values when calculating the residual value between each feature vector and the nearest center vector.
5. The image forgery detection method according to claim 1, characterized in that: The extracting domain features from the quantitative particle size distribution feature image to obtain shared features includes: Inputting the quantized particle size distribution feature into a deep neural network, wherein the deep neural network is a multi-layer perceptron, and the multi-layer perceptron includes multiple hidden layers, an activation function, and an output layer; The multiple hidden layers are used to fuse the quantitative granularity distribution features at different levels by connection or weighted summation to generate a comprehensive feature representation; The activation function is used to capture the nonlinear relationship between the quantitative particle size distribution features according to the comprehensive feature representation to obtain a final feature representation; The output layer is used to compress the final feature representation to between 0 and 1 to obtain shared features.
6. The image forgery detection method according to claim 5, characterized in that: The method for calculating the shared features includes: Performing feature fusion on the quantized particle size distribution feature and the weight matrix of the deep neural network, and accumulating the fusion result and the bias term; An activation function is performed on the result of accumulating the fusion result and the bias term to obtain a shared feature.
7. The image forgery detection method according to claim 1, characterized in that: The extracting domain features from the quantitative particle size distribution feature image to obtain unique features includes: Inputting the quantized particle size distribution features into a convolutional layer to extract local features; The local features are input into a feature extraction network to generate unique features.
8. The image forgery detection method according to claim 7, characterized in that: The calculation method of the unique features includes: Obtaining multiple basis functions in the feature extraction network and the weight corresponding to each basis function; The results of multiplying multiple basis functions with corresponding weights and accumulating them are used as unique features.
9. The image forgery detection method according to claim 1, characterized in that: The dynamically fusing the shared features and the unique features includes: The shared features and the unique features are combined to obtain comprehensive features; Input the comprehensive features into a fully connected layer or a self-attention mechanism module, weight the features of each layer in the fully connected layer or the self-attention mechanism module using a gated weight, and perform weighted summation of the weighted features of all layers to obtain dynamically fused features; The gating weight is obtained according to the network learning parameters obtained after training the fully connected layer or the self-attention mechanism module.
10. The image forgery detection method according to claim 9, characterized in that: The method for calculating the gating weight includes: Performing feature fusion on the comprehensive features and the gated weight learning parameters in the fully connected layer or the self-attention mechanism module, and accumulating the fusion result with the gated bias item learning parameters; According to the accumulated results, the gating weight is obtained.
11. The image forgery detection method according to claim 10, characterized in that: The step of classifying the dynamically fused features and outputting a forgery detection result according to the classification result includes: The dynamic fusion features are activated by the activation function Make predictions: ,in It is the Sigmoid activation function, which is used to output the forgery probability value and the final prediction result. represents the probability that the image is forged; According to the final prediction results and threshold Classify the output results: ; in, To output the result, A value of 0 indicates that the image is judged to be a real image. A value of 1 indicates that the image is judged to be a forged image; The threshold Set according to different scenarios and different needs, including: When performing forgery detection of legal evidence images, the threshold is set to a first threshold, and the first threshold is used to increase the probability that the output is a forged image; When performing news content image forgery detection, the threshold is set to a second threshold, and the second threshold is used to balance the probability of outputting a forged image and outputting a real image; When performing social media image forgery detection, setting the threshold to a third threshold, wherein the third threshold is used to reduce the probability that the output is a forged image; The first threshold is smaller than the second threshold, and the second threshold is smaller than the third threshold.
12. The image forgery detection method according to claim 2, characterized in that: The convolutional neural network is obtained after training, and the method for obtaining the training data set of the convolutional neural network includes: Acquire multiple original images, and preprocess the multiple original images to obtain an initial data set, wherein the preprocessing includes cropping the original images into a uniform size; The images in the initial data set are expanded by randomly rotating, flipping, and adding noise to obtain an expanded data set; The pixel values of the images in the expanded data set are normalized to obtain a training data set.
13. An image forgery detection device, characterized in that: include: The first extraction module is used to extract features from the input image to obtain initial image features; A quantization module, used for performing multi-level residual quantization on the initial image features, capturing the quantized particle size distribution features in the input image, and obtaining a quantized particle size distribution feature image; A second extraction module is used to perform domain feature extraction on the quantitative particle size distribution feature image to obtain shared features and unique features, wherein the shared features are features included in at least two domains in the quantitative particle size distribution feature image, and the unique features are features included in each domain in the quantitative particle size distribution feature image; The fusion module is used to dynamically fuse the shared features and the unique features, classify the dynamically fused features, and output a forgery detection result according to the classification result.
14. A server comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the image forgery detection method according to any one of claims 1 to 12 is implemented.
15. A non-transitory readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image forgery detection method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Deep forgery detection method and corresponding device
CN115311525A
Face forgery detection system and method based on convolutional neural network
CN116824708A