Counterfeit face detection method and device, electronic equipment and storage medium
Through the method of dual-branch feature extraction and multi-dimensional attention fusion, the shortcomings of deep fake detection algorithm in accuracy and generalization are solved, and efficient recognition and anti-interference capabilities of fake facial images are achieved.
Patent Information
- Application Number
- CN202510826494.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-16
AI Technical Summary
Existing deepfake detection algorithms have deficiencies in accuracy and generalization, making it difficult to effectively identify forged facial images.
A dual-branch feature extraction method is adopted, combining the RGB branch and the difference branch. Through the multi-scale edge enhancement and detail perception interaction modules, the multi-dimensional attention fusion module is used to perform feature fusion to improve the detection accuracy.
It significantly improves the accuracy and robustness of deep fake detection, can effectively identify fake facial images, and has good anti-interference capabilities.
Smart Images

Figure CN120656225A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep fake detection technology, and in particular to a method, device, electronic device and storage medium for detecting fake faces. Background Art
[0002] As the performance of deep generative models continues to improve, they can generate realistic facial images. Forged content can not only be used to generate fake news, but also to create pornographic videos, and can also be used for political attacks and financial fraud, posing a huge threat to different fields. Therefore, how to detect forged facial images is a technical problem that urgently needs to be solved to maintain social security and public order.
[0003] Existing deep fake detection algorithms still have shortcomings in accuracy and generalization, and their performance needs to be further improved. Summary of the Invention
[0004] In order to address the deficiencies of the prior art, the present invention provides a forged face detection method, device, electronic device, and storage medium; On the one hand, a forged face detection method is provided, including: Obtain a facial image to be detected, input the facial image to be detected into the trained deepfake detection model, and obtain a detection result of whether the facial image to be detected is a real image or a forged image; Among them, the trained deep fake detection model is used to: Extract channel difference image from the face image to be detected; Using the RGB branch, feature extraction is performed on the face image to be detected to obtain the unenhanced RGB features of the first stage. The unenhanced RGB features of the first stage are subjected to multi-scale edge enhancement processing to obtain the enhanced RGB features of the first stage. The enhanced RGB features of the first stage are processed at different stages to obtain the RGB features of the second, third and fourth stages. Using the difference branch, feature extraction is performed on the channel difference image to obtain the unenhanced difference features of the first stage. The unenhanced difference features of the first stage are subjected to multi-scale edge enhancement processing to obtain the enhanced difference features of the first stage. The enhanced difference features of the first stage are processed at different stages to obtain the difference features of the second, third and fourth stages. The RGB features of the fourth stage and the difference features of the fourth stage are fused with multi-dimensional attention to obtain fused features; the fused features are classified to obtain the true or false prediction probability of the face image to be detected.
[0005] On the other hand, a fake face detection system is provided, including: a processing module configured to: obtain a facial image to be detected, input the facial image to be detected into a trained deepfake detection model, and obtain a detection result of whether the facial image to be detected is a real image or a forged image; Among them, the trained deep fake detection model is used to: extract a channel difference image from the face image to be detected; use the RGB branch to perform feature extraction on the face image to be detected to obtain the unenhanced RGB features of the first stage, perform multi-scale edge enhancement processing on the unenhanced RGB features of the first stage to obtain the enhanced RGB features of the first stage, perform different stages of processing on the enhanced RGB features of the first stage to obtain the RGB features of the second, third and fourth stages; use the difference branch to perform feature extraction on the channel difference image to obtain the unenhanced difference features of the first stage, perform multi-scale edge enhancement processing on the unenhanced difference features of the first stage to obtain the enhanced difference features of the first stage, perform different stages of processing on the enhanced difference features of the first stage to obtain the difference features of the second, third and fourth stages; perform multi-dimensional attention fusion on the RGB features of the fourth stage and the difference features of the fourth stage to obtain fused features; classify the fused features to obtain the true or false prediction probability of the face image to be detected.
[0006] In another aspect, an electronic device is provided, comprising: a memory for non-transitory storage of computer-readable instructions; and a processor for executing said computer-readable instructions, When the computer-readable instructions are executed by the processor, the method described in the first aspect is executed.
[0007] On the other hand, a storage medium is provided, which non-temporarily stores computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the method described in the first aspect is executed.
[0008] On the other hand, a computer program product is provided, comprising a computer program, wherein the computer program is configured to implement the method described in the first aspect when running on one or more processors.
[0009] The above technical solution has the following advantages or beneficial effects: Based on the interaction and fusion of dual-branch features, this method has higher detection accuracy and better robustness than existing deep learning-based detection methods. Using this method, the authenticity of facial images can be accurately determined.
[0010] First, to compensate for the incomplete forgery cues caused by a single RGB domain, this paper introduces a difference branch as a supplement. Second, to highlight subtle forgery cues and improve detection performance, a multi-scale edge enhancement module based on the Sobel operator is designed. A detail-aware interaction module is also designed to enable mutual learning between the two-branch features and further enhance detailed edge information. Finally, a multi-dimensional attention fusion module is used to highlight important features to obtain more discriminative fused features. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0012] Figure 1 Schematic diagram of the internal modules of the forged face detection model of Example 1; Figure 2 Schematic diagram of the internal modules of the VSS block of Example 1; Figure 3 This is a schematic diagram of the internal modules of the scanning block of Example 1; Figure 4 This is a schematic diagram of the internal structure of the multi-scale edge enhancement module in Example 1; Figure 5 This is a schematic diagram of the internal structure of the detail perception interaction module of Example 1; Figure 6 This is a schematic diagram of the internal structure of the multi-dimensional attention fusion module in Example 1; Figure 7 This is a comparison chart of the robustness under color saturation attack in Example 1; Figure 8 This is a comparison chart of the robustness under color contrast attack in Example 1; Figure 9 This is a comparison chart of the robustness under block distortion attack in Example 1; Figure 10 This is a comparison chart of the robustness under Gaussian noise attack of Example 1; Figure 11 This is a comparison chart of the robustness under Gaussian blur attack in Example 1; Figure 12 This is a comparison chart of the robustness under JPEG compression attack in Example 1; Figure 13 This is a comparison diagram of the robustness under rotation attack of Example 1; Figure 14 This is a comparison chart of the robustness under affine transformation attack in Example 1. DETAILED DESCRIPTION
[0013] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0014] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention. The terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0015] In the embodiments of the present invention, "and / or" simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, in the description of the present invention, "plurality" refers to two or more than two.
[0016] In addition, to facilitate a clear description of the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, the words "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or execution order, and the words "first" and "second" do not necessarily mean that they are different.
[0017] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0018] All data in this embodiment is obtained in compliance with laws and regulations and based on the consent of the user, and is used legally.
[0019] Example 1 This embodiment provides a method for detecting fake faces; Fake face detection methods, including: Obtain a facial image to be detected, input the facial image to be detected into the trained deepfake detection model, and obtain a detection result of whether the facial image to be detected is a real image or a forged image; Among them, the trained deep fake detection model is used to: Extract channel difference image from the face image to be detected; Using the RGB branch, feature extraction is performed on the face image to be detected to obtain the unenhanced RGB features of the first stage. The unenhanced RGB features of the first stage are subjected to multi-scale edge enhancement processing to obtain the enhanced RGB features of the first stage. The enhanced RGB features of the first stage are processed at different stages to obtain the RGB features of the second, third and fourth stages. Using the difference branch, feature extraction is performed on the channel difference image to obtain the unenhanced difference features of the first stage. The unenhanced difference features of the first stage are subjected to multi-scale edge enhancement processing to obtain the enhanced difference features of the first stage. The enhanced difference features of the first stage are processed at different stages to obtain the difference features of the second, third and fourth stages. The RGB features of the fourth stage and the difference features of the fourth stage are fused with multi-dimensional attention to obtain fused features; the fused features are classified to obtain the true or false prediction probability of the face image to be detected.
[0020] The beneficial effect of the above technical solution is that it can improve the accuracy and robustness of deep fake detection.
[0021] Furthermore, the method further includes: The enhanced RGB features of the first stage, the RGB features of the second and third stages, and the difference features of the corresponding stages are respectively subjected to detail perception interaction processing to obtain the RGB features after interaction in the first, second and third stages and the difference features after interaction; the RGB features after interaction in each stage are used as the input value for the next stage processing in the RGB branch; the difference features after interaction in each stage are used as the input value for the next stage processing in the difference branch.
[0022] It should be understood that the above solution is to perform detail perception interaction processing on the RGB features enhanced in the first stage and the difference features enhanced in the first stage to obtain the RGB features after the first stage interaction and the difference features after the first stage interaction; Perform detail perception interaction processing on the RGB features of the second stage and the difference features of the second stage to obtain the RGB features after the second stage interaction and the difference features after the second stage interaction; The RGB features of the third stage and the difference features of the third stage are subjected to detail perception interaction processing to obtain the RGB features after the third stage interaction and the difference features after the third stage interaction.
[0023] Furthermore, the training process of the trained deep fake detection model includes: Constructing a training set, wherein the training set is face images with known authenticity labels; The training set is input into the deep fake detection model to train the model. When the loss function value of the model no longer decreases or the number of iterations exceeds the set number, the training is stopped to obtain the trained deep fake detection model.
[0024] Furthermore, if Figure 1 As shown, the trained deep fake detection model includes: RGB branch and difference branch; The RGB branch includes: a first embedding module, a first VSS module, a first multi-scale edge enhancement module, a first adder, a first downsampling layer, a second VSS module, a second adder, a second downsampling layer, a third VSS module, a third adder, a third downsampling layer, and a fourth VSS module connected in sequence; The difference branch includes: a second embedding module, a fifth VSS module, a second multi-scale edge enhancement module, a fourth adder, a fourth downsampling layer, a sixth VSS module, a fifth adder, a fifth downsampling layer, a seventh VSS module, a sixth adder, a sixth downsampling layer, and an eighth VSS module connected in sequence; The output end of the fourth VSS module and the output end of the eighth VSS module are both connected to the input end of the multidimensional attention fusion module, the output end of the multidimensional attention fusion module is connected to the input end of the classifier, and the output end of the classifier outputs the true and false discrimination result; The output end of the first multi-scale edge enhancement module is connected to the first input end of the first detail perception interaction module, the output end of the second multi-scale edge enhancement module is connected to the second input end of the first detail perception interaction module, the first output end of the first detail perception interaction module is connected to the input end of the first adder; and the second output end of the first detail perception interaction module is connected to the input end of the fourth adder. The output end of the second VSS module is connected to the first input end of the second detail perception interaction module, the output end of the sixth VSS module is connected to the second input end of the second detail perception interaction module, the first output end of the second detail perception interaction module is connected to the input end of the second adder; the second output end of the second detail perception interaction module is connected to the input end of the fifth adder; Among them, the output end of the third VSS module is connected to the first input end of the third detail perception interaction module, the output end of the seventh VSS module is connected to the second input end of the third detail perception interaction module, the first output end of the third detail perception interaction module is connected to the input end of the third adder; the second output end of the third detail perception interaction module is connected to the input end of the sixth adder.
[0025] Furthermore, the internal structures of the first embedding module and the second embedding module are consistent. The first embedding module decomposes the image into 16 square image blocks of the same size and then flattens each image block into a single vector embedding as the input of the subsequent model.
[0026] Furthermore, if Figure 2 As shown, the internal structures of the first VSS module, the second VSS module, the third VSS module, the fourth VSS module, the fifth VSS module, the sixth VSS module, the seventh VSS module and the eighth VSS module are consistent. The first VSS module includes: A first-layer normalization module, a scanning module, a seventh adder, a second-layer normalization module, a feedforward neural network, and an eighth adder connected in sequence; The input end of the seventh adder is also connected to the input end of the first layer normalization module; The input end of the eighth adder is also connected to the input end of the second-layer normalization module.
[0027] Furthermore, if Figure 3 As shown, the scanning module includes: The first linear layer, depth-wise separable convolution layer DWConv, activation function layer SiLU, SS2D (2D Selective Scan Module) unit, third layer normalization module and second linear layer are connected in sequence.
[0028] Furthermore, if Figure 4 As shown, the internal structure of the first multi-scale edge enhancement module is consistent with that of the second multi-scale edge enhancement module. The first multi-scale edge enhancement module includes: The first branch, the second branch and the third branch are arranged in parallel; The first branch includes: a first edge operator calculation layer, a first batch of normalization layers, a first activation function layer ReLU, and a first channel attention layer connected in sequence; wherein the input end of the first edge operator calculation layer is the input end of the first multi-scale edge enhancement module; The second branch includes: a first global maximum pooling layer, a second edge operator calculation layer, a second batch normalization layer, a second activation function layer ReLU, a first bilinear interpolation layer, and a second channel attention layer connected in sequence; the input end of the first global maximum pooling layer is connected to the input end of the first multi-scale edge enhancement module; The third branch includes: a second global maximum pooling layer, a third edge operator calculation layer, a third batch normalization layer, a third activation function layer ReLU, a second bilinear interpolation layer and a third channel attention layer connected in sequence; the input end of the second global maximum pooling layer is connected to the output end of the first global maximum pooling layer; The output end of the first channel attention layer, the output end of the second channel attention layer, and the output end of the third channel attention layer are all connected to the input end of the ninth adder; The output end of the ninth adder is connected to the input end of the first convolutional layer, the output end of the first convolutional layer is connected to the input end of the fourth batch normalization layer, the output end of the fourth batch normalization layer is connected to the input end of the fourth activation function layer ReLU, the output end of the fourth activation function layer ReLU is connected to the input end of the tenth adder, and the input end of the tenth adder is also connected to the input end of the first multi-scale edge enhancement module; the output end of the tenth adder is the output end of the first multi-scale edge enhancement module.
[0029] Furthermore, if Figure 4 As shown, the internal structures of the first channel attention layer, the second channel attention layer, and the third channel attention layer are consistent; the first channel attention layer includes: The third global maximum pooling layer, the eleventh adder, the second convolutional layer, the fifth activation function layer ReLU, the third convolutional layer, the first activation function layer Sigmoid and the first multiplier are connected in sequence; Among them, the input end of the third global maximum pooling layer is the input end of the first channel attention layer; the output end of the first multiplier is the output end of the first channel attention layer; the input end of the first channel attention layer is also connected to the input end of the first multiplier; the input end of the first channel attention layer is also connected to the input end of the average pooling layer, and the output end of the average pooling layer is also connected to the input end of the eleventh adder.
[0030] Furthermore, if Figure 5 As shown, the internal structures of the first detail perception interaction module, the second detail perception interaction module, and the third detail perception interaction module are consistent. The first detail perception interaction module includes: a first input end of the first detail perception interaction module and a second input end of the first detail perception interaction module; the first input end of the first detail perception interaction module is connected to the input end of the fourth edge operator calculation layer, the output end of the fourth edge operator calculation layer is connected to the input end of the fifth batch normalization layer, and the output end of the fifth batch normalization layer is connected to the input end of the sixth activation function layer ReLU; The first input end of the first detail perception interaction module is connected to the input end of the fourth convolutional layer; the output end of the fourth convolutional layer is connected to the input end of the first flattening layer, and the output end of the first flattening layer is connected to the input end of the first Reshape layer; the output end of the first Reshape layer is connected to the input end of the second multiplier, and the output end of the second multiplier is connected to the input end of the third multiplier, and the output end of the third multiplier is the first output end of the first detail perception interaction module; The output end of the sixth activation function layer ReLU is connected to the input end of the fifth convolutional layer and the input end of the sixth convolutional layer respectively; the output end of the fifth convolutional layer is connected to the input end of the second flattening layer; the output end of the sixth convolutional layer is connected to the input end of the third flattening layer; the output end of the third flattening layer is connected to the input end of the fifth multiplier; the output end of the second flattening layer is connected to the input end of the fifth multiplier; The second input end of the first detail perception interaction module is connected to the input end of the fifth edge operator calculation layer, the output end of the fifth edge operator calculation layer is connected to the input end of the sixth batch normalization layer, and the output end of the sixth batch normalization layer is connected to the input end of the seventh activation function layer ReLU; The second input end of the first detail perception interaction module is connected to the input end of the ninth convolutional layer; the output end of the ninth convolutional layer is connected to the input end of the sixth flattening layer, and the output end of the sixth flattening layer is connected to the input end of the second Reshape layer; the output end of the second Reshape layer is connected to the input end of the fourth multiplier, and the output end of the fourth multiplier is connected to the input end of the fifth multiplier, and the output end of the fifth multiplier is the second output end of the first detail perception interaction module; The output end of the seventh activation function layer ReLU is connected to the input end of the seventh convolutional layer and the input end of the eighth convolutional layer respectively; the output end of the seventh convolutional layer is connected to the input end of the fourth flattening layer; the output end of the eighth convolutional layer is connected to the input end of the fifth flattening layer; the output end of the fourth flattening layer is connected to the input end of the second multiplier; and the output end of the fifth flattening layer is connected to the input end of the third multiplier.
[0031] Furthermore, the internal calculation processes of the first, second, third, fourth, and fifth edge operator calculation layers are consistent. The calculation process of the first edge operator calculation layer is consistent with the calculation process of the edge detection operator Sobel, specifically including: Two 3×3 convolution kernels are used to calculate the gradients of the feature map in the horizontal and vertical directions respectively. The horizontal convolution kernel is used to detect longitudinal edges, and the vertical convolution kernel is used to detect horizontal edges. Finally, the complete edge features are obtained by integrating the information in the two directions through square roots.
[0032] Furthermore, the internal working principles of the first, second, third, fourth, fifth and sixth flattening layers are consistent, and the first flattening layer is used to convert a multi-dimensional vector into a one-dimensional vector.
[0033] Furthermore, the first Reshape layer and the second Reshape layer are both used to convert a tensor into a new shape without changing the number of elements in the tensor, for subsequent matrix multiplication with the attention map.
[0034] Furthermore, the RGB branch is used to extract features from the face image to be detected to obtain unenhanced RGB features of the first stage, multi-scale edge enhancement processing is performed on the unenhanced RGB features of the first stage to obtain enhanced RGB features of the first stage, and different stages of processing are performed on the enhanced RGB features of the first stage to obtain RGB features of the second, third and fourth stages, including: The face image is input into the RGB branch and divided into non-overlapping image blocks through the first embedding module; Input non-overlapping image blocks into the first VSS module to obtain the unenhanced RGB features of the first stage, which are also the first-scale RGB features. Input the first-scale RGB features into the first multi-scale edge enhancement module, and obtain the second-scale RGB features and the third-scale RGB features through two-level global maximum pooling of the first multi-scale edge enhancement module. The edges of the RGB features of each scale are calculated using the Sobel operator, and the channel attention layer is used to distribute different weights to different channels to obtain edge-enhanced RGB features of three scales. The edge-enhanced RGB features of the three scales are added and residually connected with the first-scale RGB features. The first multi-scale edge enhancement module outputs the enhanced RGB features of the first stage; The enhanced RGB features of the first stage are input into the first downsampling layer and the second VSS module to extract the RGB features of the second stage; the RGB features of the second stage are input into the second downsampling layer and the third VSS module to extract the RGB features of the third stage; the RGB features of the third stage are input into the third downsampling layer and the fourth VSS module to extract the RGB features of the fourth stage.
[0035] Furthermore, the method adopts a difference branch to perform feature extraction on the channel difference image to obtain unenhanced difference features of the first stage, performs multi-scale edge enhancement processing on the unenhanced difference features of the first stage to obtain enhanced difference features of the first stage, and performs different stages of processing on the enhanced difference features of the first stage to obtain difference features of the second, third and fourth stages, including: The cropped face image is passed through the difference image extraction module to perform channel difference. That is, for the RGB three-channel image, the R channel component is subtracted from the G channel component to obtain the first difference, the G channel component is subtracted from the B channel component to obtain the second difference, and the B channel component is subtracted from the R channel component to obtain the third difference. The first, second and third differences are then channel-connected to obtain a channel difference image. Input the channel difference image into the difference branch, and divide the channel difference image into non-overlapping image blocks after passing through the second embedding module; Input the non-overlapping channel difference image blocks into the fifth VSS module to obtain the unenhanced difference features of the first stage, which are also the first-scale difference features. Input the first-scale difference features into the second multi-scale edge enhancement module. After the two-level global maximum pooling operation of the second multi-scale edge enhancement module, the second-scale difference features and the third-scale difference features are obtained. The edge of each scale difference feature is calculated by the Sobel operator, and the channel attention layer is used to distribute different weights to different channels to obtain edge-enhanced difference features of three scales. The edge-enhanced difference features of the three scales are added and residually connected with the first-scale difference feature, and finally the enhanced difference features of the first stage are output; The enhanced difference features of the first stage are input into the fourth downsampling layer and the sixth VSS module to extract the difference features of the second stage; the difference features of the second stage are input into the fifth downsampling layer and the seventh VSS module to extract the difference features of the third stage; the difference features of the third stage are input into the sixth downsampling layer and the eighth VSS module to extract the difference features of the fourth stage.
[0036] Furthermore, if Figure 5 As shown in FIG, the enhanced RGB features of the first stage, the RGB features of the second and third stages, and the difference features of the corresponding stages are respectively subjected to detail perception interaction processing to obtain the RGB features after interaction of the first, second and third stages and the difference features after interaction; the detail perception interaction processing process of the three stages is consistent, wherein the detail perception interaction processing process of the first stage includes: The first-stage enhanced RGB features are input, and the RGB edge features are calculated by the Sobel operator. The first-stage enhanced RGB features are linearly mapped to obtain the RGB query features. The RGB edge features are linearly mapped to obtain the RGB key features and RGB value features. The same operation is used to process the difference features enhanced in the first stage to obtain difference key features, difference query features and difference value features; Then perform cross attention calculation: The first attention map is generated using the RGB query feature and the difference key feature. The first attention map is then element-wise multiplied with the difference value feature to perform feature interaction and obtain the RGB feature after the first stage of interaction. The second attention map is generated using the difference query feature and RGB key feature. The second attention map is then element-wise multiplied with the RGB value feature to perform feature interaction and obtain the difference feature after the first stage interaction.
[0037] Furthermore, the method of using the RGB features after interaction in each stage as the input value for the next stage processing in the RGB branch; and using the difference features after interaction in each stage as the input value for the next stage processing in the difference branch specifically includes: The RGB features after the first stage of interaction are used as the input value of the first adder; The difference features after the first stage of interaction are used as the input value of the fourth adder; The RGB features after the second stage of interaction are used as the input value of the second adder; The difference features after the second stage of interaction are used as the input value of the fifth adder; The RGB features after the third stage of interaction are used as the input value of the third adder; The difference features after the third stage of interaction are used as the input value of the sixth adder; The fourth VSS module outputs the RGB features of the fourth stage; The eighth VSS module outputs the difference characteristics of the fourth stage.
[0038] Furthermore, the RGB features of the fourth stage and the difference features of the fourth stage are fused with multi-dimensional attention to obtain fused features, which specifically include: Connect the RGB features of the fourth stage and the difference features of the fourth stage along the channel dimension to obtain the connected features; For the connection features, global maximum pooling and global average pooling operations are performed along the channel, width, and height directions respectively. The two pooling results in each direction are connected to calculate the attention map to obtain the attention map in the channel direction, the attention map in the width direction, and the attention map in the height direction; Then multiply the connection feature, the channel direction attention map, the width direction attention map and the height direction attention map to obtain the enhanced feature; The enhanced features are split into enhanced RGB features and enhanced difference features along the channel dimension, and the enhanced RGB features and enhanced difference features are flattened channel by channel. The cosine similarity of each channel vector is calculated, and the attention vector is calculated based on the obtained similarity. The attention vector is multiplied with the enhanced RGB features and enhanced difference features to highlight the important channels. Finally, the original information is retained through the residual connection to obtain the RGB features that retain the original information and the difference features that retain the original information; The RGB features that retain the original information and the difference features that retain the original information are added together to output the fused features.
[0039] For example, the RGB branch uses VMamba as the backbone network, and the input size of the RGB branch is For a face image, the first VSS module of the RGB branch outputs the unenhanced RGB features of the first stage , the second VSS module outputs the RGB features of the second stage , the third VSS module outputs the RGB features of the third stage , the fourth VSS module outputs the RGB features of the fourth stage .
[0040] For example, the size is The RGB image is input into the difference image extraction module to obtain the size of The channel difference image, the difference branch uses VMamba as the backbone network, and the input size of the difference branch is The channel difference image, the fifth VSS module of the difference branch outputs the unenhanced difference feature of the first stage , the sixth VSS module outputs the difference characteristics of the second stage , the seventh VSS module outputs the difference characteristics of the third stage , the eighth VSS module outputs the difference characteristics of the fourth stage .
[0041] Exemplarily, the first multi-scale edge enhancement module is used to enhance edge information in features and highlight traces of detailed forgery; For the RGB branch, the input of the first multi-scale edge enhancement module is , the output is the edge enhanced feature , For the difference branch, the input of the second multi-scale edge enhancement module is , the output is the edge enhanced feature .
[0042] For the RGB branch, in the first multi-scale edge enhancement module, two-level global maximum pooling is used to extract the first-scale RGB features. The second scale RGB features and the third scale RGB features Then the first, second and third scale RGB features are edge enhanced based on the Sobel operator, which is expressed as:
[0043]
[0044]
[0045] in, BN represents batch normalization; BI represents bilinear interpolation, which is used to restore the feature to its original size; in, CA Represents channel attention, which is used to give different weights to different channels to highlight important information, expressed as:
[0046] in, GAP represents global average pooling; GMP represents global maximum pooling; Represents the Sigmoid function.
[0047] The edge enhancement features of the three scales are weighted summed and combined with Perform residual connection and finally output the enhanced RGB features of the first stage , expressed as:
[0048] in, 、 、 are learnable weight parameters.
[0049] For example, the internal working principles of the first detail perception interaction module, the second detail perception interaction module and the third detail perception interaction module are consistent. The first detail perception interaction module uses cross attention to achieve edge-aware dual-branch feature interaction, thereby promoting representation learning and extracting more subtle forgery clues. The input of the first detail perception interaction module is the RGB features enhanced in the first stage. Difference features from the first stage enhancement The output of the first detail perception interaction module is the RGB feature after information interaction and the difference characteristics after information interaction .
[0050] Furthermore, the first detail perception interaction module is configured to: First, use the Sobel operator, batch normalization, and ReLU to extract RGB features Edge information of RGB features ; First, use Sobel operator, batch normalization, and ReLU to extract difference features The edge information of the difference feature ; Then, 1×1 convolution is used to transform the RGB features Mapping to RGB query features ; Using two 1×1 convolutions Mapping to RGB key features and RGB value features , and perform flattening and Reshape operations. In the mapping process, the channel dimension reduction factor is introduced r , to reduce computational complexity while retaining sufficient information.
[0051] Process different features through the same operation and its edge information , get the difference query features , difference key features and difference value features , and perform flattening and reshape operations.
[0052] For the RGB branch, use RGB query features and differential key features Multiplying the attention map , expressed as:
[0053] The difference value feature Multiply the elements with the attention map to obtain the RGB detail perception feature, and resize the RGB detail perception feature to obtain the final interaction feature , expressed as:
[0054] in, are trainable weight parameters.
[0055] For difference branches, use the difference query feature and RGB key features Multiplying the attention map , expressed as:
[0056] The RGB value features Multiply the elements with the attention map to obtain the difference detail perception feature, and resize the difference detail perception feature to obtain the final interaction feature , expressed as:
[0057] in, are trainable weight parameters.
[0058] Furthermore, the RGB features of the fourth stage and the difference features of the fourth stage are subjected to multi-dimensional attention fusion to obtain fusion features, which specifically includes: The input of the multi-dimensional attention fusion module is the RGB features of the fourth stage The difference between the fourth stage and ; The output of the multi-dimensional attention fusion module is the fusion feature ; For the input RGB features and differential characteristics , connect along the channel to get the connection feature ; Then, the connection features Perform global average pooling and global maximum pooling in the channel dimension, and use the Permute function to connect the features. Transpose, perform global average pooling and global maximum pooling in the height and width dimensions respectively, and finally connect along the pooling dimension to obtain the pooling feature , expressed as:
[0059] For convenience, we use Represents the pooled features along the channel, height, and width; Indicates feature connection along the channel; and Represents the dimension transposition function, which is used to transform The dimension from Adjust to and to perform pooling operations.
[0060] Using pooling features, calculate multi-dimensional attention map , expressed as:
[0061] Use the Permute function to transpose the high and wide dimension attention maps to the original dimensions, and then use the multidimensional attention map to connect the features Enhance to obtain multi-dimensional enhanced features , expressed as:
[0062] in, and express and The transpose of and The dimension of and Restore to and .
[0063] For multi-dimensional enhanced features , divided into two parts according to the connection order: RGB branch and difference branch and .
[0064] Then, the dual-branch features are flattened into vectors channel by channel to obtain the features and .calculate and The cosine similarity of the corresponding channel obtains a similarity vector, and then calculates the cosine attention vector , expressed as:
[0065] in, Indicates cosine similarity calculation; k indicates the number of channels.
[0066] Using cosine attention vector to analyze dual-branch features and weighted, and with and Perform residual connection to obtain channel weighted features and , expressed as:
[0067]
[0068] final, and Add and output multi-dimensional attention fusion features , expressed as:
[0069] The classifier includes: a layer normalization module, a global average pooling layer and a linear layer connected in sequence.
[0070] The input of the classifier is the fusion feature The output is the predicted probability of a face image being real or fake. The six modules are integrated into an overall model for training and testing. The overall model is trained and tested using four standard datasets: FaceForensics++ (FF++), Celeb-DF-v2 (CDF), WildDeepfake (WDF), and DFDCp. The FF++ dataset includes two versions: FF++c23 and FF++c40. The model is trained using cross-entropy loss, expressed as:
[0071] in, is the true label of the input face image; is the output predicted probability of the model.
[0072] Other training parameter settings are shown in Table 1. Save the trained network parameters.
[0073] Table 1 Model training parameter settings
[0074] At this point, all network models of the method of the present invention have been built and trained, and the weight files have been saved.
[0075] When performing forgery detection on the input face image, Figure 1 The complete process shown uses the trained weight parameters to detect forgery on the input test face image.
[0076] After testing, the algorithm has good accuracy and robustness, and can accurately classify the test face images into true and false ones.
[0077] To verify the effectiveness of the present invention, five datasets, FF++ (HQ), FF++ (LQ), CDF, WDF, and DFDCp, were used to test the detection performance of the present invention. Furthermore, eight post-processing attacks were applied to the test data: saturation transformation, contrast transformation, blocking distortion, Gaussian noise, Gaussian blur, JPEG compression, rotation, and affine transformation to verify the robustness of the present invention.
[0078] Table 2 The performance of the present invention and other eleven forgery detection methods on the FF++ dataset and Compare
[0079] The present invention is compared with other deep learning-based algorithms (a backbone network: Xception; ten deep fake detection models: MesoNet, Multi-task, Face X-ray, F3-Net, SPSL, HFI-Net, GocNet, Zhang et al., ID3, and HIFE). To ensure the effectiveness of the experimental results, the experimental data of the eleven algorithms are divided into the same dataset and the network training parameters are set the same. Table 2 shows the accuracy of the eleven deep fake detection methods on the FF++ (HQ) and FF++ (LQ) datasets. and AUC values Comparison,Table 3 shows the accuracy of nine deep fake detection methods on CDF, WDF and DFDCp datasets. and AUC values Comparison. As can be seen from Table 2 and Table 3, the present invention has higher accuracy. The best results are achieved on the FF++c23, FF++c40, WDF and DFDCp datasets, and the second best results are achieved on the CDF dataset. The best results were achieved on the FF++c40, CDF, WDF, and DFDCp datasets, and the second best results were achieved on the FF++c23 dataset. Figure 5 The robustness comparison of the present invention against six post-processing attacks is shown, and the evaluation indicators are It can be seen that the present invention has good robustness to saturation change, contrast change, block distortion, Gaussian noise and affine transformation.
[0080] Table 3 The performance of the present invention and other nine forgery detection methods on three datasets and Compare
[0081] The present invention includes six parts: an RGB branch, a difference branch, a multi-scale edge enhancement module, a detail perception interaction module, a multi-dimensional attention fusion module and a classifier. The steps are: using the difference image extraction module to perform difference calculations on adjacent channels to calculate the channel difference image. Then, in order to highlight the edge information, a multi-scale edge enhancement module is designed using the Sobel operator to obtain subtle forgery clues. Next, in order to achieve mutual learning of dual-branch features to obtain a more complete feature representation and further highlight the detail edge information, a detail perception interaction module is designed by fusing the Sobel operator and cross attention. Finally, in order to more effectively highlight important information and better achieve feature fusion, a multi-dimensional attention fusion module is designed to perform attention weighting on features from multiple dimensions, and highlight the significant channels of the dual-branch features based on cosine similarity, ultimately obtaining a more complete fusion feature.
[0082] Figure 7 This is a comparison chart of the robustness under color saturation attack in Example 1; Figure 8 This is a comparison chart of the robustness under color contrast attack in Example 1; Figure 9 This is a comparison chart of the robustness under block distortion attack in Example 1; Figure 10 This is a comparison chart of the robustness under Gaussian noise attack of Example 1; Figure 11 This is a comparison chart of the robustness under Gaussian blur attack in Example 1; Figure 12 This is a comparison chart of the robustness under JPEG compression attack in Example 1; Figure 13 This is a comparison diagram of the robustness under rotation attack of Example 1; Figure 14 This is a comparison chart of the robustness under affine transformation attack in Example 1.
[0083] Example 2 This embodiment provides a forged face detection system, including: a processing module configured to: obtain a facial image to be detected, input the facial image to be detected into a trained deepfake detection model, and obtain a detection result of whether the facial image to be detected is a real image or a forged image; Among them, the trained deep fake detection model is used to: extract a channel difference image from the face image to be detected; use the RGB branch to perform feature extraction on the face image to be detected to obtain the unenhanced RGB features of the first stage, perform multi-scale edge enhancement processing on the unenhanced RGB features of the first stage to obtain the enhanced RGB features of the first stage, perform different stages of processing on the enhanced RGB features of the first stage to obtain the RGB features of the second, third and fourth stages; use the difference branch to perform feature extraction on the channel difference image to obtain the unenhanced difference features of the first stage, perform multi-scale edge enhancement processing on the unenhanced difference features of the first stage to obtain the enhanced difference features of the first stage, perform different stages of processing on the enhanced difference features of the first stage to obtain the difference features of the second, third and fourth stages; perform multi-dimensional attention fusion on the RGB features of the fourth stage and the difference features of the fourth stage to obtain fused features; classify the fused features to obtain the true or false prediction probability of the face image to be detected.
[0084] It should be noted that the above-mentioned processing modules correspond to the steps in Example 1. The examples and application scenarios implemented by the above-mentioned modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0085] The descriptions of the various embodiments in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0086] The proposed system can be implemented in other ways. For example, the system embodiment described above is merely illustrative. For example, the above module division is only a logical function division. In actual implementation, other division methods may be used. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not implemented.
[0087] Example 3 This embodiment also provides an electronic device, comprising: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in the above embodiment one.
[0088] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0089] The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0090] During implementation, each step of the above method may be completed by an integrated logic circuit of hardware in a processor or by instructions in the form of software.
[0091] The method in Example 1 can be directly implemented as being executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software module can be located in a storage medium well-established in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not given here.
[0092] Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with this embodiment can be implemented using electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0093] Example 4 This embodiment further provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first embodiment is performed.
[0094] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A fake face detection method, characterized in that: include: Obtain a facial image to be detected, input the facial image to be detected into the trained deepfake detection model, and obtain a detection result of whether the facial image to be detected is a real image or a forged image; Among them, the trained deep fake detection model is used to: Extract channel difference image from the face image to be detected; Using the RGB branch, feature extraction is performed on the face image to be detected to obtain the unenhanced RGB features of the first stage. The unenhanced RGB features of the first stage are subjected to multi-scale edge enhancement processing to obtain the enhanced RGB features of the first stage. The enhanced RGB features of the first stage are processed at different stages to obtain the RGB features of the second, third and fourth stages. Using the difference branch, feature extraction is performed on the channel difference image to obtain the unenhanced difference features of the first stage. The unenhanced difference features of the first stage are subjected to multi-scale edge enhancement processing to obtain the enhanced difference features of the first stage. The enhanced difference features of the first stage are processed at different stages to obtain the difference features of the second, third and fourth stages. The RGB features of the fourth stage and the difference features of the fourth stage are fused with multi-dimensional attention to obtain fused features; the fused features are classified to obtain the true or false prediction probability of the face image to be detected.
2. The forged face detection method according to claim 1, wherein: The method further comprises: The enhanced RGB features of the first stage, the RGB features of the second and third stages, and the difference features of the corresponding stages are respectively subjected to detail perception interaction processing to obtain the RGB features after interaction in the first, second and third stages and the difference features after interaction; the RGB features after interaction in each stage are used as the input value for the next stage processing in the RGB branch; the difference features after interaction in each stage are used as the input value for the next stage processing in the difference branch.
3. The forged face detection method according to claim 1, wherein: The trained deepfake detection model includes: RGB branch and difference branch; The RGB branch includes: a first embedding module, a first VSS module, a first multi-scale edge enhancement module, a first adder, a first downsampling layer, a second VSS module, a second adder, a second downsampling layer, a third VSS module, a third adder, a third downsampling layer, and a fourth VSS module connected in sequence; The difference branch includes: a second embedding module, a fifth VSS module, a second multi-scale edge enhancement module, a fourth adder, a fourth downsampling layer, a sixth VSS module, a fifth adder, a fifth downsampling layer, a seventh VSS module, a sixth adder, a sixth downsampling layer, and an eighth VSS module connected in sequence; The output end of the fourth VSS module and the output end of the eighth VSS module are both connected to the input end of the multidimensional attention fusion module, the output end of the multidimensional attention fusion module is connected to the input end of the classifier, and the output end of the classifier outputs the true and false discrimination result; The output end of the first multi-scale edge enhancement module is connected to the first input end of the first detail perception interaction module, the output end of the second multi-scale edge enhancement module is connected to the second input end of the first detail perception interaction module, the first output end of the first detail perception interaction module is connected to the input end of the first adder; and the second output end of the first detail perception interaction module is connected to the input end of the fourth adder. The output end of the second VSS module is connected to the first input end of the second detail perception interaction module, the output end of the sixth VSS module is connected to the second input end of the second detail perception interaction module, the first output end of the second detail perception interaction module is connected to the input end of the second adder; the second output end of the second detail perception interaction module is connected to the input end of the fifth adder; Among them, the output end of the third VSS module is connected to the first input end of the third detail perception interaction module, the output end of the seventh VSS module is connected to the second input end of the third detail perception interaction module, the first output end of the third detail perception interaction module is connected to the input end of the third adder; the second output end of the third detail perception interaction module is connected to the input end of the sixth adder.
4. The forged face detection method according to claim 3, wherein: The internal structure of the first multi-scale edge enhancement module is consistent with that of the second multi-scale edge enhancement module. The first multi-scale edge enhancement module includes: The first branch, the second branch and the third branch are arranged in parallel; The first branch includes: a first edge operator calculation layer, a first batch of normalization layers, a first activation function layer ReLU, and a first channel attention layer connected in sequence; wherein the input end of the first edge operator calculation layer is the input end of the first multi-scale edge enhancement module; The second branch includes: a first global maximum pooling layer, a second edge operator calculation layer, a second batch normalization layer, a second activation function layer ReLU, a first bilinear interpolation layer, and a second channel attention layer connected in sequence; the input end of the first global maximum pooling layer is connected to the input end of the first multi-scale edge enhancement module; The third branch includes: a second global maximum pooling layer, a third edge operator calculation layer, a third batch normalization layer, a third activation function layer ReLU, a second bilinear interpolation layer and a third channel attention layer connected in sequence; the input end of the second global maximum pooling layer is connected to the output end of the first global maximum pooling layer; The output end of the first channel attention layer, the output end of the second channel attention layer, and the output end of the third channel attention layer are all connected to the input end of the ninth adder; The output end of the ninth adder is connected to the input end of the first convolutional layer, the output end of the first convolutional layer is connected to the input end of the fourth batch normalization layer, the output end of the fourth batch normalization layer is connected to the input end of the fourth activation function layer ReLU, the output end of the fourth activation function layer ReLU is connected to the input end of the tenth adder, and the input end of the tenth adder is also connected to the input end of the first multi-scale edge enhancement module; the output end of the tenth adder is the output end of the first multi-scale edge enhancement module.
5. The forged face detection method according to claim 3, wherein: The internal structures of the first detail perception interaction module, the second detail perception interaction module, and the third detail perception interaction module are consistent. The first detail perception interaction module includes: a first input end of the first detail perception interaction module and a second input end of the first detail perception interaction module; the first input end of the first detail perception interaction module is connected to the input end of the fourth edge operator calculation layer, the output end of the fourth edge operator calculation layer is connected to the input end of the fifth batch normalization layer, and the output end of the fifth batch normalization layer is connected to the input end of the sixth activation function layer ReLU; The first input end of the first detail perception interaction module is connected to the input end of the fourth convolutional layer; the output end of the fourth convolutional layer is connected to the input end of the first flattening layer, and the output end of the first flattening layer is connected to the input end of the first Reshape layer; the output end of the first Reshape layer is connected to the input end of the second multiplier, and the output end of the second multiplier is connected to the input end of the third multiplier, and the output end of the third multiplier is the first output end of the first detail perception interaction module; The output end of the sixth activation function layer ReLU is connected to the input end of the fifth convolutional layer and the input end of the sixth convolutional layer respectively; the output end of the fifth convolutional layer is connected to the input end of the second flattening layer; the output end of the sixth convolutional layer is connected to the input end of the third flattening layer; the output end of the third flattening layer is connected to the input end of the fifth multiplier; the output end of the second flattening layer is connected to the input end of the fifth multiplier; The second input end of the first detail perception interaction module is connected to the input end of the fifth edge operator calculation layer, the output end of the fifth edge operator calculation layer is connected to the input end of the sixth batch normalization layer, and the output end of the sixth batch normalization layer is connected to the input end of the seventh activation function layer ReLU; The second input end of the first detail perception interaction module is connected to the input end of the ninth convolutional layer; the output end of the ninth convolutional layer is connected to the input end of the sixth flattening layer, and the output end of the sixth flattening layer is connected to the input end of the second Reshape layer; the output end of the second Reshape layer is connected to the input end of the fourth multiplier, and the output end of the fourth multiplier is connected to the input end of the fifth multiplier, and the output end of the fifth multiplier is the second output end of the first detail perception interaction module; The output end of the seventh activation function layer ReLU is connected to the input end of the seventh convolutional layer and the input end of the eighth convolutional layer respectively; the output end of the seventh convolutional layer is connected to the input end of the fourth flattening layer; the output end of the eighth convolutional layer is connected to the input end of the fifth flattening layer; the output end of the fourth flattening layer is connected to the input end of the second multiplier; and the output end of the fifth flattening layer is connected to the input end of the third multiplier.
6. The forged face detection method according to claim 1, wherein: The method adopts the RGB branch to extract features of the face image to be detected to obtain unenhanced RGB features of the first stage, performs multi-scale edge enhancement processing on the unenhanced RGB features of the first stage to obtain enhanced RGB features of the first stage, and performs different stages of processing on the enhanced RGB features of the first stage to obtain RGB features of the second, third and fourth stages, including: The face image is input into the RGB branch and divided into non-overlapping image blocks through the first embedding module; Input non-overlapping image blocks into the first VSS module to obtain the unenhanced RGB features of the first stage, which are also the first-scale RGB features. Input the first-scale RGB features into the first multi-scale edge enhancement module, and obtain the second-scale RGB features and the third-scale RGB features through two-level global maximum pooling of the first multi-scale edge enhancement module. The edges of the RGB features of each scale are calculated using the Sobel operator, and the channel attention layer is used to distribute different weights to different channels to obtain edge-enhanced RGB features of three scales. The edge-enhanced RGB features of the three scales are added and residually connected with the first-scale RGB features. The first multi-scale edge enhancement module outputs the enhanced RGB features of the first stage; The enhanced RGB features of the first stage are input into the first downsampling layer and the second VSS module to extract the RGB features of the second stage; the RGB features of the second stage are input into the second downsampling layer and the third VSS module to extract the RGB features of the third stage; the RGB features of the third stage are input into the third downsampling layer and the fourth VSS module to extract the RGB features of the fourth stage; or, The method adopts the difference branch to extract features from the channel difference image to obtain unenhanced difference features of the first stage, performs multi-scale edge enhancement processing on the unenhanced difference features of the first stage to obtain enhanced difference features of the first stage, and performs different stages of processing on the enhanced difference features of the first stage to obtain difference features of the second, third and fourth stages, including: The cropped face image is passed through the difference image extraction module to perform channel difference. That is, for the RGB three-channel image, the R channel component is subtracted from the G channel component to obtain the first difference, the G channel component is subtracted from the B channel component to obtain the second difference, and the B channel component is subtracted from the R channel component to obtain the third difference. The first, second and third differences are then channel-connected to obtain a channel difference image. Input the channel difference image into the difference branch, and divide the channel difference image into non-overlapping image blocks after passing through the second embedding module; Input the non-overlapping channel difference image blocks into the fifth VSS module to obtain the unenhanced difference features of the first stage, which are also the first-scale difference features. Input the first-scale difference features into the second multi-scale edge enhancement module. After the two-level global maximum pooling operation of the second multi-scale edge enhancement module, the second-scale difference features and the third-scale difference features are obtained. The edge of each scale difference feature is calculated by the Sobel operator, and the channel attention layer is used to distribute different weights to different channels to obtain edge-enhanced difference features of three scales. The edge-enhanced difference features of the three scales are added and residually connected with the first-scale difference feature, and finally the enhanced difference features of the first stage are output; The enhanced difference features of the first stage are input into the fourth downsampling layer and the sixth VSS module to extract the difference features of the second stage; the difference features of the second stage are input into the fifth downsampling layer and the seventh VSS module to extract the difference features of the third stage; the difference features of the third stage are input into the sixth downsampling layer and the eighth VSS module to extract the difference features of the fourth stage; or, The enhanced RGB features of the first stage, the RGB features of the second and third stages, and the difference features of the corresponding stages are respectively subjected to detail perception interaction processing to obtain the RGB features and difference features after interaction of the first, second and third stages. The detail perception interaction processing process of the three stages is consistent, and the detail perception interaction processing process of the first stage includes: The first-stage enhanced RGB features are input, and the RGB edge features are calculated by the Sobel operator. The first-stage enhanced RGB features are linearly mapped to obtain the RGB query features. The RGB edge features are linearly mapped to obtain the RGB key features and RGB value features. The same operation is used to process the difference features enhanced in the first stage to obtain difference key features, difference query features and difference value features; Then perform cross attention calculation: The first attention map is generated using the RGB query feature and the difference key feature. The first attention map is then element-wise multiplied with the difference value feature to perform feature interaction and obtain the RGB feature after the first stage of interaction. The second attention map is generated using the difference query feature and RGB key feature. The second attention map is then element-wise multiplied with the RGB value feature to perform feature interaction and obtain the difference feature after the first stage interaction.
7. The forged face detection method according to claim 1, wherein: The RGB features of the fourth stage and the difference features of the fourth stage are fused with multi-dimensional attention to obtain fused features, including: Connect the RGB features of the fourth stage and the difference features of the fourth stage along the channel dimension to obtain the connected features; For the connection features, global maximum pooling and global average pooling operations are performed along the channel, width, and height directions respectively. The two pooling results in each direction are connected to calculate the attention map to obtain the attention map in the channel direction, the attention map in the width direction, and the attention map in the height direction; Then multiply the connection feature, the channel direction attention map, the width direction attention map and the height direction attention map to obtain the enhanced feature; The enhanced features are split into enhanced RGB features and enhanced difference features along the channel dimension, and the enhanced RGB features and enhanced difference features are flattened channel by channel. The cosine similarity of each channel vector is calculated, and the attention vector is calculated based on the obtained similarity. The attention vector is multiplied with the enhanced RGB features and enhanced difference features to highlight the important channels. Finally, the original information is retained through the residual connection to obtain the RGB features that retain the original information and the difference features that retain the original information; The RGB features that retain the original information and the difference features that retain the original information are added together to output the fused features.
8. A fake face detection system, characterized by: include: a processing module configured to: obtain a facial image to be detected, input the facial image to be detected into a trained deepfake detection model, and obtain a detection result of whether the facial image to be detected is a real image or a forged image; Among them, the trained deep fake detection model is used to: extract a channel difference image from the face image to be detected; use the RGB branch to perform feature extraction on the face image to be detected to obtain the unenhanced RGB features of the first stage, perform multi-scale edge enhancement processing on the unenhanced RGB features of the first stage to obtain the enhanced RGB features of the first stage, perform different stages of processing on the enhanced RGB features of the first stage to obtain the RGB features of the second, third and fourth stages; use the difference branch to perform feature extraction on the channel difference image to obtain the unenhanced difference features of the first stage, perform multi-scale edge enhancement processing on the unenhanced difference features of the first stage to obtain the enhanced difference features of the first stage, perform different stages of processing on the enhanced difference features of the first stage to obtain the difference features of the second, third and fourth stages; perform multi-dimensional attention fusion on the RGB features of the fourth stage and the difference features of the fourth stage to obtain fused features; classify the fused features to obtain the true or false prediction probability of the face image to be detected.
9. An electronic device, comprising: a memory for non-transitory storage of computer-readable instructions; as well as a processor for executing said computer-readable instructions, When the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is executed.
10. A storage medium, characterized in that: Non-transitory storage of computer-readable instructions, wherein when the non-transitory computer-readable instructions are executed by a computer, the method according to any one of claims 1 to 7 is performed.