A method for identifying deepfake compressed face images based on deep learning
Through dynamic adaptive discrete cosine transformation and multi-dimensional entropy state load mutual feed mechanism, the frequency domain and airspace characteristics are extracted and interacted with each other, and the problem of low accuracy of face forgery images after compression is solved, and high-accurate face forgery identification is achieved.
Patent Information
- Application Number
- CN202411558517.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-11-04
AI Technical Summary
When detecting compressed face forged images, the detection accuracy of the prior art is low, making it difficult to cope with the problems of blurred and obfuscated feature addition caused by compression processing.
The frequency domain mode is extracted by dynamic adaptive discrete cosine transform, combined with the multi-dimensional entropy state load mutual feed mechanism, to realize the two-way interaction of high-dimensional global and local information, and deeply explore the nonlinear fake trace characteristics.
It improves the identification accuracy of compressed images, effectively retains the fake features that are blurred during compression, and clearly distinguishes confusing features, significantly improving the identification performance.
Smart Images

Figure CN119541058B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition in artificial intelligence, and particularly relates to a method for identifying deepfake compressed face images based on deep learning. Background Art
[0002] Deepfake refers to the use of deep learning and computer vision technologies to synthesize realistic fake face images and videos. Identifying deepfake face images is the process of recognizing and detecting fake face images generated using artificial intelligence technologies.
[0003] Currently, for the dual-branch deepfake face image identification network that combines the spatial domain and the frequency domain, traditional discrete cosine transforms are used to extract frequency domain modal information with fixed spectra, and a combination of convolutional neural networks and Transformers is used to extract local and global features, so as to identify fake face images. The main problems with this approach are as follows: (1) The detection accuracy for compressed face fake images is very low. Images and videos uploaded on social media are often compressed, and compression will blur many fake features in the face image and add new confounding features to disrupt the fake identification results. Current detection methods have difficulty dealing with compression. (2) Traditional discrete cosine transforms use fixed frequency decompositions and tend to favor low-frequency spectral information when extracting frequency domain modes, but high-frequency spectral information often retains more fake traces. The method of traditional discrete cosine transforms for extracting frequency domain modes contains fewer high-frequency fake features. (3) When using convolutional neural networks and Transformers to extract local and global features, convolutional neural networks focus on extracting local features, and Transformers focus on extracting global features, and there is little interaction between global and local features. This makes the identification network have weak analytical ability for non-linear boundaries and difficult to identify non-linear fake traces. Summary of the Invention
[0004] To solve the above problems, the present invention proposes a method for identifying deepfake compressed face images based on deep learning. It adopts a dual-branch network architecture that combines the spatial domain and the frequency domain, designs a method of dynamic adaptive discrete cosine transform to flexibly extract frequency domain modes containing dynamic spectral information that are more suitable for deepfake compressed face images, and at the same time designs a multi-dimensional entropy state load mutual feedback mechanism to allow full two-way interaction and feedback between high-dimensional global and local information, and deeply excavate non-linear fake trace features. During the training of the deep learning model, it learns to effectively extract, fuse, and identify dual-modal fake traces in the spatial and frequency domains of the image. The present invention can accurately identify compressed face images fabricated by deepfake means.
[0005] The technical solution of the present invention is as follows:
[0006] A method for identifying deepfake compressed face images based on deep learning, comprising the following steps:
[0007] Step 1: Perform frequency-domain modal extraction on the deepfake compressed face image using dynamic adaptive discrete cosine transform;
[0008] Step 2: Extract high-dimensional frequency-domain payload features using a multi-head self-attention mechanism and a multi-dimensional entropy state payload mutual feedback mechanism;
[0009] Step 3: Extract spatial domain features based on a pre-trained deep residual network;
[0010] Step 4: Perform feature fusion on the high-dimensional frequency-domain payload features and spatial domain features based on an attention mechanism;
[0011] Step 5: Use a classifier containing a fully connected layer to perform discriminant classification on the fused features.
[0012] Further, the specific process of Step 1 is as follows:
[0013] Step 1.1: Project the deepfake compressed face image X ∈ R N×N into the universal domain space through unified preprocessing, denoted as the face image matrix X N , and apply a learnable dynamic discrete cosine transform matrix M N to X D , and M D is continuously updated through the feedback of the loss function during the training process. The calculation formula for generating matrix elements is:
[0014]
[0015] where N represents the size of the dynamic discrete cosine transform matrix; represents the value of the element in the i-th row and j-th column of the dynamic discrete cosine transform matrix at the t-th round of training; represents the update amount according to the loss function during the (t - 1)-th round of training;
[0016] Step 1.2: Use M D to perform discrete cosine transform on X N to obtain the high-dimensional frequency-domain response result X freq , and the calculation formula is:
[0017]
[0018] where represents the high-dimensional frequency-domain response result at the t-th round of training; represents the dynamic discrete cosine transform matrix at the t-th round of training; T represents transpose;
[0019] Step 1.3. Update the adaptive multi-frequency filter F A ∈ {F h , F m , F l , F a}, F h , F m , F l , F a correspond to the high-frequency adaptive filter, intermediate-frequency adaptive filter, low-frequency adaptive filter, and full-frequency adaptive filter parts respectively. The calculation formula is:
[0020]
[0021] where is the adaptive multi-frequency filter in the t-th round of training; represent the high-frequency adaptive filter, intermediate-frequency adaptive filter, low-frequency adaptive filter, and full-frequency adaptive filter in the t-th round of training respectively; σ(·) represents the Sigmoid activation function; ε represents the adaptive parameter; represent the update amount according to the loss function in the (t - 1)-th round of training; B h , B m , B l , B a correspond to the high-frequency basic filter, intermediate-frequency basic filter, low-frequency basic filter, and full-frequency basic filter respectively. The calculation formula is:
[0022]
[0023] where represents the matrix row number of the basic filter; represents the matrix column number of the basic filter;
[0024] Step 1.4. Decompose and fuse X A using F freq to obtain the final frequency-domain modal image Y freq . The calculation formula is:
[0025]
[0026] where represents the frequency-domain modal image in the t-th round of training; F conv (·) represents the convolution mapping operation of the frequency-domain modal feature channel; represents the concatenation operation; * represents matrix multiplication;
[0027] Step 1.5. During the training process, feedback through the loss function to M D and F APerform dynamic adaptive update adjustment, M D and F A The update rule calculation formula is:
[0028]
[0029] Wherein, represents The update amount of the (t - 1)-th round of training according to the loss function, represents the value of the element in the i-th row and j-th column of the dynamic discrete cosine transform matrix during the (t - 1)-th round of training; represents The update amount of the (t - 1)-th round of training according to the loss function, represents the adaptive multi-frequency filter during the (t - 1)-th round of training; η represents the learning rate; L(t - 1) represents the loss function of the (t - 1)-th round;
[0030] Step 1.6, Use the extracted frequency-domain modal image Y freq Perform shallow frequency-domain modal forgery feature extraction using a 5×5 convolutional neural network, and then segment the shallow frequency-domain modal forgery features into individual payloads. The calculation formulas for the number of payloads Q and the payload dimension D are:
[0031]
[0032] Wherein, Ω h , Ω w , Ω c respectively represent the longitudinal dimension, transverse dimension, and number of feature channels of the payload in the entropy state space; represents the boundary control factor of the perception window.
[0033] Furthermore, the specific process of step 2 is:
[0034] Step 2.1, Capture the long-term context dependence relationship of the entropy state space of the payload through the multi-head self-attention mechanism. After passing through the multi-head self-attention mechanism, Q payloads form Q high-dimensional payloads. All the high-dimensional payloads constitute the high-dimensional payload set S = {s 1 , s 2 , …, s Q}, and s Q is the Q-th high-dimensional payload;
[0035] Step 2.2, Use the multi-dimensional entropy state payload mutual feedback mechanism to perform multi-dimensional fine-grained feature interaction and feedback attention calculation.
[0036] Furthermore, the specific process of step 2.2 is:
[0037] Step 2.2.1. The entropy states of each high-dimensional load in the high-dimensional load set are asymmetric non-normal distributions. The entropy state size of each high-dimensional load is measured through the standard deviation statistical functional, and the high-dimensional loads are globally hierarchically sorted according to the entropy state size. The calculation formula is:
[0038]
[0039] where R(S) represents the sorted high-dimensional load set; E[·] represents the entropy state expectation value; F Rank (·) represents the sorting algorithm;
[0040] Step 2.2.2. The sorted high-dimensional load set R(S) is decomposed into two subsets, which are respectively mapped to different entropy dimension regions. The two subsets are the low-entropy state high-dimensional load subset R(S) a and the high-entropy state high-dimensional load subset R(S) b , where a and b represent different entropy dimensions; the high-entropy state loads are included in R(S) b ;
[0041] Step 2.2.3. Adjust the size and offset of the sensing window to non-uniformly and adaptively reconstruct R(S) b . When reconstructing, the boundary control factor of the sensing window is set to twice the original;
[0042] The high-entropy state high-dimensional load subset after reconstruction is R(S) b ′. The calculation formulas for the number Q′ of high-dimensional loads and the dimension D′ of high-dimensional loads of R(S) b ′ are:
[0043]
[0044] Step 2.2.4. Perform two-way cross-dimensional interaction on R(S) a and R(S) b ′. The specific calculation formula is:
[0045]
[0046] where cross_atten a and cross_atten b respectively represent the high-dimensional load features after two-way cross-dimensional interaction of R(S) a and R(S) b ′; g cross_attention (·) represents the cross-attention fusion method;
[0047] Step 2.2.5. For cross_atten bPerform shape correction on cross_atten b Restore it to the same shape as cross_atten a After becoming the same shape, perform cascading on cross_atten a and cross_atten b in the entropy state dimension to obtain the final high-dimensional load feature in the frequency domain.
[0048] Furthermore, in step 3, pre-train the deep residual network ResNet50 on the ImageNet dataset. ResNet50 extracts features through stacking residual blocks, and the calculation formula for each residual block is:
[0049]
[0050] where represents the residual convolution operation performed by the k-th residual block; represents the spatial domain feature extracted in the k-th residual block; represents the spatial domain feature extracted in the (k - 1)-th residual block; represents the total number of residual blocks.
[0051] Furthermore, the specific process of step 4 is as follows:
[0052] Step 4.1: First, cascade the spatial domain feature and the high-dimensional load feature in the frequency domain, and then calculate the attention weight W atten of the cascaded feature. The calculation formula is:
[0053]
[0054] where Features spa represents the spatial domain feature; Features freq represents the high-dimensional load feature in the frequency domain; Conv(·) represents a 1×1 convolution operation; Φ(·) represents the RELU activation function;
[0055] Step 4.2: Multiply the attention weight of the cascaded feature element-wise with the cascaded feature to obtain the fused feature.
[0056] Furthermore, the specific process of step 5 is as follows:
[0057] Step 5.1: Use a fully connected layer to calculate the unnormalized score of the fused feature; the fully connected layer projects the dimension of the fused feature to the dimension of the category to be classified. The projection formula of the fully connected layer is as follows:
[0058]
[0059] where Indicates the classification category, Indicates the forgery category, Indicates the genuine category; Indicates the category Normalized score; Features fus Indicates the fused features; Indicates the category Weight matrix of the corresponding fully connected layer; Indicates the category Bias of the corresponding fully connected layer;
[0060] Step 5.2: Use the Softmax activation function to convert the unnormalized score into the corresponding probability, and select the category with the highest probability as the output category determined by the model. The probability calculation formula of the Softmax activation function is as follows:
[0061]
[0062] Among them, P represents the probabilities of different categories; Are the normalized scores of the forgery category and the genuine category, respectively.
[0063] The beneficial technical effects brought by the present invention are as follows.
[0064] (1) Achieve efficient and comprehensive frequency-domain modal capture. The dynamic adaptive discrete cosine transform can adaptively adjust the weights extracted from different frequency domains according to the characteristics of the input compressed face image, dynamically enhance the sensitivity to the forgery area, and thus more accurately capture the frequency-domain modes related to deep forgery.
[0065] (2) Improve the ability to capture high-dimensional local and global features in the frequency domain. The multi-dimensional entropy state load mutual feedback mechanism uses standard deviation statistical functional ranking to preferentially process high-entropy state high-dimensional loads containing highly critical information, adaptively reconstructs high-entropy state high-dimensional loads through a perception window, and conducts complex information circulation through two-way cross-dimensional interaction, thereby enhancing the sensitivity to subtle forgery traces in the frequency domain.
[0066] (3) Improve the discrimination accuracy of compressed images. The present invention uses the dynamic adaptive discrete cosine transform to extract the frequency-domain modes of face images, effectively retains the forgery features blurred during the compression process, and through the multi-dimensional entropy state load mutual feedback mechanism, can clearly distinguish the confused features, effectively extract the subtle forgery features of the frequency-domain modes, and greatly improve the discrimination accuracy of compressed images. Description of the Drawings
[0067] Figure 1 Is a process diagram of the method for identifying deep forgery compressed face images based on deep learning of the present invention.
[0068] Figure 2 It is a process diagram for changing the scale of high-dimensional payloads using a sensing window in the multi-dimensional entropy state payload mutual feedback mechanism of the present invention.
[0069] Figure 3 It is an ROC curve diagram of the present invention on the C23 version of the FaceForensics++ dataset.
[0070] Figure 4 It is an ROC curve diagram of the present invention on the C40 version of the FaceForensics++ dataset. Detailed implementation manners
[0071] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners:
[0072] As Figure 1 shown, a method for identifying deepfake compressed face images based on deep learning includes the following steps:
[0073] Step 1: Perform frequency-domain modal extraction on the deepfake compressed face image using dynamic adaptive discrete cosine transform; for an input deepfake compressed face image, the present invention first uses dynamic adaptive discrete cosine transform to extract the frequency-domain modal image. The specific process is as follows:
[0074] Step 1.1: Project the input deepfake compressed face image X ∈ R N×N onto the general domain space through unified preprocessing, denoted as the face image matrix X N , and apply a learnable dynamic discrete cosine transform matrix M N to X D . M D is continuously updated through the feedback of the loss function during the training process. The calculation formula for generating matrix elements is:
[0075]
[0076] where t represents the number of training epochs; i represents the row subscript of the matrix; j represents the column subscript of the matrix; N represents the size of the dynamic discrete cosine transform matrix; represents the value of the element in the i-th row and j-th column of the dynamic discrete cosine transform matrix at the t-th training epoch; represents the update amount according to the loss function at the (t - 1)-th training epoch.
[0077] Step 1.2: Use M D to perform discrete cosine transform on x N to obtain the high-dimensional frequency-domain response result X freq . The calculation formula is:
[0078]
[0079] Among them, represents the high-dimensional frequency-domain response result of the t-th round of training; represents the dynamic discrete cosine transform matrix of the t-th round of training; T represents transpose.
[0080] Step 1.3. Update the adaptive multi-frequency filter F A ∈{F h , F m , F l , F a}, where F h , F m , F l , F a correspond to the high-frequency adaptive filter, intermediate-frequency adaptive filter, low-frequency adaptive filter, and full-frequency adaptive filter parts respectively, and the calculation formula is:
[0081]
[0082] Among them, is the adaptive multi-frequency filter of the t-th round of training; represent the high-frequency adaptive filter, intermediate-frequency adaptive filter, low-frequency adaptive filter, and full-frequency adaptive filter of the t-th round of training respectively; σ(·) represents the Sigmoid activation function; ε represents the adaptive parameter; represent respectively the update amount of the (t - 1)-th round of training according to the loss function; B h , B m , B l , B a represent the high-frequency basic filter, intermediate-frequency basic filter, low-frequency basic filter, and full-frequency basic filter respectively. The basic filter is used to decompose the high-dimensional frequency-domain response result into high-frequency band, intermediate-frequency band, low-frequency band, and full-frequency band, corresponding to the first 1 / 16, 1 / 16 to 1 / 8, the last 1 / 8, and all of the entire spectrum respectively. The calculation formulas of each basic filter are:
[0083]
[0084] Among them, represents the matrix row number of the basic filter; represents the matrix column number of the basic filter;
[0085] Step 1.4. Use F A to decompose and fuse X freq to obtain the final frequency-domain modal image Y freq , and the calculation formula is:
[0086]
[0087] Among them, represents the frequency-domain modal image in the t-th round of training; F conv (·) represents the convolutional mapping operation of the frequency-domain modal feature channels; represents the concatenation operation; * represents matrix multiplication.
[0088] Step 1.5, M D and F A in the dynamic adaptive discrete cosine transform method are learnable and are dynamically adaptively updated and adjusted through the feedback of the loss function during training to better capture the frequency-domain data. M D and F A The update rule calculation formula is:
[0089]
[0090] Among them, represents the update amount in the (t - 1)-th round of training according to the loss function, represents the value of the element in the i-th row and j-th column of the dynamic discrete cosine transform matrix in the (t - 1)-th round of training; represents the update amount in the (t - 1)-th round of training according to the loss function, represents the adaptive multi-frequency filter in the (t - 1)-th round of training; η represents the learning rate; L(t - 1) represents the loss function in the (t - 1)-th round.
[0091] Step 1.6, Use a 5×5 convolutional neural network to extract shallow frequency-domain modal forgery features from the extracted frequency-domain modal image Y freq , and then segment the shallow frequency-domain modal forgery features into individual payloads. The calculation formulas for the number Q of payloads and the payload dimension D are:
[0092]
[0093] Among them, Ω h , Ω w , Ω c represent the longitudinal dimension, transverse dimension, and number of feature channels of the payload in the entropy state space respectively; represents the boundary control factor of the perception window, The size of which can be determined according to specific situations such as the resolution of the image and hardware conditions.
[0094] Step 2, Adopt the multi-head self-attention mechanism and the multi-dimensional entropy state payload mutual feedback mechanism to extract frequency-domain high-dimensional payload features; The present invention realizes the deep extraction of global and local frequency-domain high-dimensional payload features. The specific process is as follows:
[0095] Step 2.1. The multi-head self-attention mechanism first captures the long-term context dependencies of the entropy state space of the payloads. After passing through the multi-head self-attention mechanism, Q payloads form Q high-dimensional payloads. All the high-dimensional payloads constitute the high-dimensional payload set S = {s 1 , s 2 , …, s Q}, where s Q is the Qth high-dimensional payload; the multi-head self-attention mechanism ignores the local detail forgery features, and the ignored local detail forgery features may often be important features affecting the identification result.
[0096] Step 2.2. To further capture the local detail forgery features, a multi-dimensional entropy state payload mutual feedback mechanism is then used for multi-dimensional fine-grained feature interaction and feedback attention calculation. The specific implementation steps of the multi-dimensional entropy state payload mutual feedback mechanism are as follows:
[0097] Step 2.2.1. The entropy states of each high-dimensional payload in the high-dimensional payload set are asymmetric non-normal distributions. Through the standard deviation statistical functional, the entropy state size of each high-dimensional payload is measured, and the high-dimensional payloads are globally hierarchically sorted according to the entropy state size. High entropy state means high criticality, ensuring that high-entropy state high-dimensional payloads are preferentially processed in the multi-scale expression of information, so as to achieve deeper correlation extraction in the high-dimensional feature space. The calculation formula for this process is:
[0098]
[0099] where R(S) represents the sorted high-dimensional payload set; E[·] represents the entropy state expectation value; F Rank (·) represents the sorting algorithm.
[0100] Step 2.2.2. The sorted high-dimensional payload set R(S) is decomposed into two subsets, which are respectively mapped to different entropy dimension regions. The two subsets are the low-entropy state high-dimensional payload subset R(S) a , and the high-entropy state high-dimensional payload subset R(S) b , where a and b represent different entropy dimensions; the high-entropy state payloads are included in R(S) b .
[0101] Step 2.2.3. Perform multi-scale entropy state reconstruction based on a deformable perception window. R(S) b is adaptively reconstructed through a deformable perception window to achieve dynamic adjustment of the high-dimensional payload scale and distribution. This reconstruction is achieved by adjusting the size of the perception window (when reconstructing, the boundary control factor of the perception window is set to twice the original, i.e., ) and the offset of the perception window to non-uniformly and adaptively reconstruct R(S) b . Figure 2As an example of the reconstruction process, Figure 2 it is assumed that there are initially 9 high-dimensional loads. These 9 high-dimensional loads form a 3×3 matrix form and are numbered sequentially from 1 to 9. The dashed box represents the sensing window. The size of the sensing window is assumed to be 2×2, and the offset of the sensing window is assumed to be 1. The sensing window moves according to the offset each time. Each time the sensing window will enclose 2×2, that is, 4 high-dimensional loads. These 4 high-dimensional loads will re-form a new larger-scale brand-new load. It can be seen from the process in the figure that the sensing window first selects the 1st, 2nd, 4th, and 5th high-dimensional loads to form the new first high-dimensional load. The sensing window moves, selects the 2nd, 3rd, 4th, and 5th high-dimensional loads to form the new second high-dimensional load. The sensing window continues to move, selects the 4th, 5th, 7th, and 8th high-dimensional loads to form the new third high-dimensional load. The sensing window moves again, selects the 5th, 6th, 8th, and 9th high-dimensional loads to form the new fourth high-dimensional load. The movement of the sensing window ends, and the change in the scale of the high-dimensional load is completed. Finally, these four high-dimensional loads are combined into a brand-new high-dimensional load with a larger scale. This brand-new high-dimensional load is in the form of a 4×4 matrix.
[0102] The high-entropy state high-dimensional load subset after reconstruction is R(S) b ′. The number and dimension of the high-dimensional loads have both changed, and load overlaps have formed. These overlaps achieve repeated interactions of local information, thereby improving the capture of high-order micro-features and their complex correlations, especially enhancing the analytical ability for non-linear boundaries. The calculation formulas for the number Q′ of high-dimensional loads and the dimension D′ of high-dimensional loads of R(S) b ′ are as follows:
[0103]
[0104] Step 2.2.4. In order to establish complex interconnections between the global and local high-dimensional information, perform two-way cross-dimensional interactions on R(S) a and R(S) b ′. The specific calculation formula is as follows:
[0105]
[0106] Among them, cross_atten a and cross_atten b respectively represent the high-dimensional load features after two-way cross-dimensional interaction of R(S) a and R(S) b ′; g cross_attention (·) represents the cross-attention fusion method.
[0107] Step 2.2.5. For cross_atten a and cross_atten bPerform cascading, that is, perform cross-scale high-dimensional load feature reconstruction because Cross_atten a and cross_atten b are high-dimensional load features of different scales. Therefore, it is necessary to correct the shape of cross_atten b and restore Cross_atten b to the same shape as Cross_atten a . After becoming the same shape, cascading these two high-dimensional load features in the entropy state dimension will obtain the final high-dimensional load feature in the frequency domain.
[0108] Step 3: Extract spatial domain features based on the pre-trained deep residual network (abbreviated as ResNet50 network); adopt a dual-branch parallel method. While extracting high-dimensional load features in the frequency domain, use the pre-trained ResNet50 network to extract spatial domain features from the input compressed face image.
[0109] While performing frequency domain mode extraction and high-dimensional load feature extraction in the frequency domain, the input face image also enters the spatial domain branch. On the spatial domain branch, use the deep residual network ResNet50 pre-trained on the InmageNet dataset to extract spatial domain features. ResNet50 is a residual network that performs feature extraction by stacking residual blocks. The calculation formula for each residual block is:
[0110]
[0111] where represents the residual convolution operation performed by the k-th residual block; represents the spatial domain features extracted in the k-th residual block; represents the spatial domain features extracted in the (k - 1)-th residual block; represents the total number of residual blocks.
[0112] Step 4: Perform feature fusion on the high-dimensional load features in the frequency domain and the spatial domain features based on the attention mechanism; based on the splicing of the spatial domain features and the high-dimensional load features in the frequency domain, the present invention uses the attention mechanism to calculate the weights of the spliced features, and enhances important features and suppresses unimportant features through the attention mechanism. The specific process is as follows:
[0113] Step 4.1: First, cascade the spatial domain features and the high-dimensional load features in the frequency domain, and then calculate the attention weights of the cascaded features. The attention weights can learn which forged features are more important, so as to enhance these forged features. The calculation formula for the attention weight W atten is:
[0114]
[0115] Among them, Features spa represents the spatial domain feature; Features freq represents the high-dimensional load feature in the frequency domain; Conv(·) represents a 1×1 convolution operation; Φ(·) represents the RELU activation function.
[0116] Step 4.2: Multiply the attention weights of the cascaded features element-wise with the cascaded features to obtain the fused features. In this way, the important forged features in the concatenated features will be enhanced, and the unimportant forged features will be suppressed, achieving a better feature fusion effect.
[0117] Step 5: Use a classifier containing a fully connected layer to perform discriminative classification on the fused features. In the present invention, the fused features are fed into the classifier to generate a discriminative result, which is either real or forged.
[0118] The present invention constructs a binary classifier for mapping the fused features to the target categories, that is, obtaining the results of the two categories of forged or real. This binary classifier uses a fully connected layer and a Softmax activation function for category judgment; the specific process is as follows:
[0119] Step 5.1: Use the fully connected layer to calculate the unnormalized scores of the fused features. The fully connected layer projects the dimension of the fused features to the category dimension to be classified. In the present invention, the category dimension to be classified is 2, and the projection formula of the fully connected layer is as follows:
[0120]
[0121] Among them, represents the classification category, represents the forged category, represents the real category; represents the normalized score of category ; Features fus represents the fused features; represents the weight matrix of the fully connected layer corresponding to category ; represents the bias of the fully connected layer corresponding to category .
[0122] Step 5.2: Based on the unnormalized scores, judge the category to obtain the classification result of the model of the present invention. Use the Softmax activation function to convert the unnormalized scores into corresponding probabilities, and select the category with the larger probability as the output category determined by the model of the present invention. The probability calculation formula of the Softmax activation function is as follows:
[0123]
[0124] Among them, P represents the probabilities of different categories; They are the normalized scores for the forged category and the real category respectively.
[0125] The deepfake-compressed face image discrimination method based on deep learning provided by the present invention realizes high-accuracy discrimination of compressed face images by designing a dynamic adaptive discrete cosine transform method and a multi-dimensional entropy state payload mutual feedback mechanism. Using F3-Net, which was the first to propose using the traditional discrete cosine transform method for frequency-domain modal extraction, as a baseline, the experimental results were compared. The evaluation metrics include: accuracy (ACC), receiver operating characteristic curve (ROC curve), and area under the ROC curve (AUC).
[0126] ACC is used to measure the proportion of samples correctly classified by the model on the test set and is the most intuitive evaluation metric. The specific calculation formula is:
[0127]
[0128] where ACC is the accuracy value.
[0129] The ROC curve (also known as the sensitivity curve) has the true positive rate (TPR) as the vertical axis and the false positive rate (FPR) as the horizontal axis, and can intuitively evaluate the classification performance of the compressed face image discrimination classifier. The specific calculation formula is:
[0130]
[0131] where TPR is the true positive rate value; FPR is the false positive rate value; TP represents true positives, that is, the number of samples with the actual label being the positive class and the prediction result also being the positive class; FP represents false positives, that is, the number of samples with the true label being the negative class and being wrongly predicted as the positive class; TN represents true negatives, that is, the number of samples with the actual label being the negative class and the prediction result also being the negative class; FN represents false negatives, that is, the number of samples with the true label being the positive class and being wrongly predicted as the negative class.
[0132] AUC represents the area under the ROC curve and uses a numerical value to evaluate the classification ability.
[0133] Performance evaluation was carried out on the FaceForensics++ dataset. This dataset contains 1000 real videos and 4000 forged videos. The real videos are from the original videos on YouTube, and the forged videos are generated based on the original videos using 4 deepfake methods and contain two compressed versions: the C23 version and the C40 version. Among them, C23 is the video dataset after high-quality compression, and C40 is the video dataset after low-quality compression. Face images were collected from this dataset through a data balancing method for performance evaluation. The comparison results of the accuracy and area under the sensitivity curve between the method of the present invention and F3-Net are shown in Table 1 below:
[0134] Comparison Results of C23 Version and C40 Version in Different Methods
[0135]
[0136]
[0137] It can be seen from the comparison results that the higher the compression ratio, the lower the image quality, and the better the performance improvement of the present invention. On the C23 high-quality compression version, the ACC increased by 0.64% and the AUC increased by 1.35%; on the C40 low-quality compression version, the ACC increased by 7.75% and the AUC increased by 6.51%, showing very good discrimination performance for compressed face images.
[0138] The ROC curves of the present invention on the C23 and C40 versions of the FaceForensics++ dataset are respectively as Figure 3 、 Figure 4 shown, and through Figure 3 、 Figure 4 it can be intuitively seen that the present invention has good classification performance.
[0139] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A deep learning-based method for identifying deep fake compressed face images, characterized in that: The steps include: Step 1: Use dynamic adaptive discrete cosine transform to extract frequency domain modality of deep fake compressed face image; Step 2: Use multi-head self-attention mechanism and multi-dimensional entropy state load mutual feedback mechanism to extract frequency domain high-dimensional load features; The specific process is: Step 2.1: Capture the long-term dependencies of the entropy space context of the payload through the multi-head self-attention mechanism. After the Q payloads pass through the multi-head self-attention mechanism, they form Q high-dimensional payloads. All high-dimensional payloads constitute a high-dimensional payload set S = {s1, s2, ..., s Q },s Q is the Qth high-dimensional load; Step 2.2: Use the multi-dimensional entropy load mutual feedback mechanism to perform multi-dimensional fine-grained feature interaction and feedback attention calculation; The specific process is: Step 2.2.1: The entropy state of each high-dimensional load in the high-dimensional load set is an asymmetric non-normal distribution. The entropy state of each high-dimensional load is measured by the standard deviation statistical functional, and the high-dimensional loads are globally sorted according to the entropy state size. The calculation formula is: Where R(S) represents the sorted high-dimensional load set; E[·] represents the expected value of the entropy state; F Rank (·) indicates a sorting algorithm; Step 2.2.2: Decompose the sorted high-dimensional load set R(S) into two subsets, which are mapped to different entropy dimension regions. The two subsets are low-entropy high-dimensional load subsets R(S) a and high entropy state high dimensional load subset R(S) b , where a and b represent different entropy dimensions; the high entropy state load is included in R(S) b middle; Step 2.2.3: Adjust the size and offset of the perception window To adaptively reconstruct R(S) b ,During reconstruction, the boundary control factor of the perception window is set to twice the original; The reconstructed high-entropy high-dimensional load subset is R(S) b ′,R(S) b The calculation formulas for the number of high-dimensional loads Q′ and the dimension of high-dimensional loads D′ are: Step 2.2.4: R(S) a and R(S) b ′ performs two-way cross-dimensional interaction; the specific calculation formula is: Among them, cross_atten a 、cross_atten b Respectively represent R(S) a 、R(S) b ′High-dimensional load characteristics after bidirectional cross-dimensional interaction; g cross_attention (·) indicates the cross-attention fusion method; Step 2.2.5: cross_atten b Perform shape correction and set cross_atten b Revert to cross_atten a Same shape, after changing to the same shape, cross_atten a and cross_atten b The final frequency domain high-dimensional load characteristics are obtained by cascading in the entropy state dimension; Step 3: Extract spatial features based on the pre-trained deep residual network; Step 4: Based on the attention mechanism, feature fusion is performed on the frequency domain high-dimensional load features and the spatial domain features; Step 5: Use a classifier containing a fully connected layer to identify and classify the fused features.
2. According to the deep learning-based deep forged compressed face image identification method of claim 1, it is characterized in that: The specific process of step 1 is as follows: Step 1.1: Deep fake compressed face image X∈R N×N After unified preprocessing, it is projected into the universal domain space and recorded as the face image matrix X N , and for X N Apply a learnable dynamic discrete cosine transform matrix M D , M D The loss function is continuously updated through feedback during training, and the calculation formula for generating matrix elements is: Where N represents the size of the dynamic discrete cosine transform matrix; Represents the value of the element in the i-th row and j-th column of the dynamic discrete cosine transform matrix during the t-th round of training; express The update amount for the t-1th round of training according to the loss function; Step 1.2: Use M D X N Perform discrete cosine transform to obtain high-dimensional frequency domain response result X freq , the calculation formula is: in, Represents the high-dimensional frequency domain response result of the tth round of training; represents the dynamic discrete cosine transform matrix of the tth round of training; T represents transpose; Step 1.3: Update the adaptive multi-frequency filter F A ∈{F h , F m , F l , F a }, F h 、F m 、F l 、F a They correspond to the high-frequency adaptive filter, the intermediate-frequency adaptive filter, the low-frequency adaptive filter, and the full-frequency adaptive filter respectively, and the calculation formula is: in, is the adaptive multi-frequency filter for training the tth round; They represent the high-frequency adaptive filter, medium-frequency adaptive filter, low-frequency adaptive filter, and full-frequency adaptive filter of the t-th round of training respectively; σ(·) represents the Sigmoid activation function; ε represents the adaptive parameter; Respectively The update amount of the training round t-1 according to the loss function; B h , B m , B l , B a They represent high-frequency basic filter, medium-frequency basic filter, low-frequency basic filter, and full-frequency basic filter respectively, and the calculation formula is: in, Represents the matrix row number of the basic filter; Represents the matrix column number of the basic filter; Step 1.4: Use F A X freq Decompose and fuse to obtain the final frequency domain modal image Y freq , the calculation formula is: in, represents the frequency domain modal image of the tth round of training; F conv (·) represents the convolution mapping operation of the frequency domain modal feature channel; Indicates cascade operation; * indicates matrix multiplication; Step 1.5: During the training process, the loss function is used to feedback M D and F A Perform dynamic adaptive update adjustment, M D and F A The update rule calculation formula is: in, express The update amount for the t-1th round of training according to the loss function, Represents the value of the element in the i-th row and j-th column of the dynamic discrete cosine transform matrix during the t-1th round of training; express The update amount for the t-1th round of training according to the loss function, represents the adaptive multi-frequency filter of the t-1th round of training; η represents the learning rate; L(t-1) represents the loss function of the t-1th round; Step 1.6: Extract the frequency domain modal image Y freq A 5×5 convolutional neural network is used to extract shallow frequency domain modal forgery features, and then the shallow frequency domain modal forgery features are segmented into individual loads. The calculation formulas for the number of loads Q and the load dimension D are: Among them, Ω h ,Ω w ,Ω c They represent the vertical dimension, horizontal dimension, and number of characteristic channels of the load in the entropy space respectively; Represents the boundary control factor of the perception window.
3. According to claim 2, the deep learning-based deep forged compressed face image identification method is characterized in that: In step 3, the deep residual network ResNet50 is pre-trained on the InmageNet dataset. ResNet50 extracts features by stacking residual blocks. The calculation formula for each residual block is: in, represents the residual convolution operation performed by the kth residual block; represents the spatial features extracted from the kth residual block; represents the spatial features extracted from the k-1th residual block; Represents the total number of residual blocks.
4. According to claim 3, the deep learning-based deep forged compressed face image identification method is characterized in that: The specific process of step 4 is as follows: Step 4.1: First, concatenate the spatial domain features and the frequency domain high-dimensional load features, and then calculate the attention weight W of the concatenated features. atten , the calculation formula is: Among them, Features spa Represents spatial features; Features freq represents the high-dimensional load feature in the frequency domain; Conv(·) represents a 1×1 convolution operation; Φ(·) represents the RELU activation function; Step 4.2: Multiply the attention weight of the cascade feature by the cascade feature element by element to obtain the fused feature.
5. According to claim 4, the deep learning-based deep forged compressed face image identification method is characterized in that: The specific process of step 5 is as follows: Step 5.
1. Use the fully connected layer to calculate the unnormalized score of the fused feature. The fully connected layer projects the dimension of the fused feature to the dimension of the category to be classified. The formula for the fully connected layer projection is as follows: in, Indicates the classification category, Indicates the forged category, represents the real category; Indicates category Normalized score of Features fus Indicates fusion features; Indicates category The corresponding weight matrix of the fully connected layer; Indicates category The corresponding fully connected layer bias; Step 5.2: Use the Softmax activation function to convert the unnormalized scores into corresponding probabilities, and select the category with the largest probability as the output category determined by the model. The probability calculation formula of the Softmax activation function is as follows: Among them, P represents the probability of different categories; are the normalized scores of the fake and real categories respectively.
Citation Information
Patent Citations
Deep counterfeit video identification method and system based on double features of spatial domain and frequency domain
CN113935365A
Deep counterfeit video identification method based on frequency domain information and multi-task learning
CN115187891A