Face depth spoofing detection device and method based on multiple gram texture

By using a face forgery detection device based on deep neural networks and Gram matrices, and leveraging global semantics and multi-gram texture feature extraction modules, the problem of detecting subtle texture features in deep forged face images that is difficult to detect in existing technologies is solved, and higher accuracy forged image detection is achieved.

CN115909459BActive Publication Date: 2025-12-05BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211509797.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2025-12-05
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively detect subtle texture features in deepfake face images, resulting in insufficient accuracy in forgery detection.

Method used

A face forgery detection device based on deep neural networks and Gram matrices is adopted. Through a global semantic feature extraction module, a multi-gram texture feature extraction module, and a feature fusion module, spatial local texture and global style texture features of the image are extracted and enhanced respectively, and feature fusion is performed to achieve high-precision detection.

Benefits of technology

It improves the accuracy of face forgery image detection, enhances the network's expressive power, and can effectively distinguish the texture differences between the forgery area and the background area, thus achieving more efficient forgery image discrimination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115909459B_ABST
    Figure CN115909459B_ABST
Patent Text Reader

Abstract

The application discloses a face deep fake detection device and method based on multiple gram texture, which comprises a global semantic feature extraction module, a multiple gram texture feature extraction module, and a feature fusion module.The global semantic feature extraction module extracts a face image to be detected and shallow image features and global semantic features from the face image.The multiple gram texture feature extraction module constructs the shallow image features extracted by the global semantic feature extraction module into shallow image features and strengthens the texture characteristics of the input face image.The multiple gram texture feature extraction module comprises a spatial dimension local gram texture feature module, a channel dimension global gram texture construction module, and a feature contrast strengthening module.The feature fusion module obtains the global semantic features extracted by the global semantic feature extraction module and the final gram texture features constructed by the multiple gram texture feature extraction module, splices and fuses the global semantic features and the final gram texture features to obtain fused features, and realizes discrimination on fake images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, and in particular relates to a face deep forgery detection device and method based on deep neural networks and Gram matrices. Background Technology

[0002] Deepfake technology for faces includes face replacement, expression modification, and attribute editing. With the support of deep learning technology and the gradual improvement of synthesis technology, the difficulty of obtaining such forgeries has been greatly reduced. More and more fake images have appeared on the Internet, and the proliferation of these false information has had a great negative impact on the politics and security of society.

[0003] Currently, several image forgery detection methods exist for detecting and recognizing deepfake faces. Generally, these methods focus on higher-level semantic information within the entire image, distinguishing between forged and genuine images in a high-level semantic feature space. However, often the forgery traces and clues of deepfake faces exist locally and manifest as subtle textures, and current technologies lack forgery detection methods that can deeply process these characteristics. Summary of the Invention

[0004] To address the problems and shortcomings of the existing technologies, this invention proposes a face forgery detection device and method based on deep neural networks and Gram matrices. First, a shallow neural network is used to model the low-level features of an image and output its corresponding features. Then, a multi-gram texture feature extraction module and a high-level semantic extraction module are used to extract the features. Finally, a feature fusion module is used to fuse the features, thereby achieving higher accuracy in detecting forged faces.

[0005] A face deepfake detection device based on multi-gram texture according to an embodiment of the present invention includes: a global semantic feature extraction module, a multi-gram texture feature extraction module, and a feature fusion module, wherein...

[0006] The global semantic feature extraction module is used to extract the face image to be tested for effective input into the device, and to extract shallow image features and global semantic features from the face image;

[0007] The multi-gram texture feature extraction module includes a spatial dimension local gram texture feature module, a channel dimension global gram texture feature module, and a feature contrast enhancement module. The multi-gram texture feature extraction module constructs gram texture features from the shallow image features extracted by the global semantic feature extraction module and enhances the texture characteristics of the input face image.

[0008] The spatial dimension local gram texture feature module describes the spatial dimension feature correlation of the face image by calculating the gram matrix between pixels of the shallow image features, thereby realizing the modeling of spatial local texture. It also uses a spatial attention mechanism to calculate the mask of the shallow image features, distinguishes the tampered area and the background area in the input shallow image features, and extracts the spatial gram texture features of the tampered area and the background area of ​​the image respectively.

[0009] The channel-dimensional global gram texture feature module models the overall style texture of the image by calculating the gram matrix between different channels of the input face image. It also uses a spatial attention mechanism to calculate the mask of shallow image features, distinguishes the tampered area and the background area in the input shallow image features, and extracts the global gram style texture features of the tampered area and the background area of ​​the image respectively.

[0010] The feature contrast enhancement module performs feature residual calculation on the spatial gram texture features of the image tampering region and the background region of the shallow image features to obtain spatial gram texture contrast enhancement features. It also performs feature residual calculation on the channel gram style texture features of the image tampering region and the background region to obtain global channel gram texture contrast enhancement features. Finally, it splices the gram texture contrast enhancement features of the spatial dimension and the channel dimension to obtain the fused multi-gram texture features.

[0011] The feature fusion module obtains the global semantic features extracted by the global semantic feature extraction module and the fused multigram texture features constructed by the multigram texture feature extraction module. The global semantic features and the fused multigram texture features are then spliced ​​and fused to obtain fused features, thereby enabling the identification of forged images.

[0012] In an optional implementation, both the global semantic feature extraction module and the multigram texture feature extraction module employ convolutional neural networks.

[0013] In an optional implementation, the spatial dimension local gram texture feature module can divide the input shallow image features into square feature blocks of 4, 5, 7, 9, etc., and the channel dimension global gram texture feature module can compress the input shallow image features into a number of channels selected from 32 to 96.

[0014] In an optional implementation, the face image extracted by the global semantic feature extraction module is 320×320 pixels in size.

[0015] According to an embodiment of the present invention, a method for detecting deepfake faces based on multi-gram textures is provided, comprising the following steps:

[0016] S1: Construct a face image dataset that includes face images synthesized using deepfake technology. This face image dataset has image-level annotations, in which 1 represents a fake image and 0 represents a real image.

[0017] S2: Scale the face image obtained in step S1 and input it into the global semantic feature extraction module to obtain the global semantic features of the face image. and shallow image features Where i=1,2, and global semantic features The output dimension is Global semantic features used to describe the input face image, shallow image features and The dimension is ,in, , These represent the length and width of the shallow image features, respectively. , This constitutes the spatial dimension of shallow image features. Let i be the channel dimension of the shallow image features, where i = 1, 2;

[0018] S3: The shallow image features obtained in step S2... The data is sequentially fed into the channel-dimensional global gram texture feature module and processed, specifically including the following steps:

[0019] S3-1: The channel-dimensional global Gram texture feature module utilizes spatial attention to calculate the mask. 1. Processing the received shallow image features ,in Dimensions , , These are shallow image features. Length and width, The range of all elements in 1 is [0,1], thus we obtain 1 and ,in Representing shallow image features Shallow image features corresponding to the background region, features Representing shallow image features The shallow image features corresponding to the tampered area in the middle, where i=1,2, and * denotes matrix bitwise multiplication;

[0020] S3-2: Obtain the features input from step S3-1 ,right The channel dimensions are compressed and the features of different channels are unfolded to obtain the image embedding features of the tampered region. Its size is ,in Representation of features Compressed channel dimensions, Representation of features The expanded spatial dimensions ;

[0021] S3-3: Obtain the features input in step S3-2 Calculate the channel dimension Gram matrix:

[0022]

[0023] The channel dimension Gram matrix The global Gram texture features of the image tampering region are obtained after convolutional and pooling layers. ,

[0024] S3-4: Obtain the features input from step S3-1 , for features The channel dimensions are compressed and the features of different channels are expanded to obtain the image embedding features of the background region. Its size is ,in Representation of features Compressed channel dimensions, Representation of features Spatial dimensions after feature expansion ;

[0025] S3-5: Obtain the features input in step S3-4 Calculate the Gram matrix for the channel dimension:

[0026]

[0027] The channel dimension Gram matrix The global Gram texture features of the image background region are obtained through convolutional layers and pooling layers. ;

[0028] S3-6: Obtain and The difference between them is calculated to obtain the global channel Gram texture contrast enhancement feature. ;

[0029] S4: Obtain the shallow image layer features from the input in step S2. The spatial dimension local gram texture feature module is input sequentially and processed, specifically including the following steps:

[0030] S4-1: The spatial dimension local gram texture feature module utilizes spatial attention to calculate the mask. Processing received shallow image features This leads to the acquisition of shallow image features. Shallow image features corresponding to the mid-background region 2. Shallow image features Shallow image features corresponding to the tampered area ;

[0031] S4-2: Obtain features from step S4-1 , for features The image is divided into blocks along its spatial dimension and then expanded to obtain the embedding features of each block. The embedding features of this block The size is ,in =S*S , ,in , These are shallow image features. Length and width, S The size of the block indicates the size of the segment. N Indicates the number of blocks. Represents the image embedding features of the blocks;

[0032] S4-3: Obtain the input features from step S4-2 Calculate the Gram matrix in terms of spatial dimensions:

[0033]

[0034] The spatial dimensional gram matrix is ​​passed through convolutional and pooling layers to obtain the spatial gram texture features of the image tampering region. ;

[0035] S4-4: Features obtained from step S4-1 ,right The image is divided into blocks along its spatial dimension and then expanded to obtain the embedding features of each block. The embedding features of this block The size is ,in =S*S , ,in , These are shallow image features. Length and width, S The size of the block indicates the size of the segment. N Indicates the number of blocks. Represents the image embedding features of the blocks;

[0036] S4-5: Obtain the input features from step S4-4 Calculate the Gram matrix in terms of spatial dimensions:

[0037]

[0038] The spatial gram texture features of the image background region are obtained by passing the channel-dimensional gram matrix through convolutional and pooling layers. ;

[0039] S4-6: Calculation and The difference between them yields the final spatial gram texture contrast enhancement feature. ;

[0040] S5: The feature contrast enhancement module receives features from steps S4 and S3. and For i=1,2, a concatenation operation is performed to obtain multigram texture features. i=1,2 represent the multigram texture features obtained after inputting different shallow image features into the multigram texture feature extraction module. These multigram texture features are further fused through a concatenation operation. , , This represents the multi-gram texture features after fusion;

[0041] S6: The feature fusion module receives the global semantic features obtained in step S2. The fused multigram texture features obtained in step S5 And through splicing operations Achieve feature fusion, This means that the final image features F are obtained and then fed into a classifier to detect forgery of the input image, thus obtaining the final detection result.

[0042] S7: Repeat steps S2 to S6 until the loss function converges, complete the training and save all parameters in the global semantic feature extraction module in step S2, the channel dimension global gram texture feature module in step S3, and the spatial dimension local gram texture feature module in step S4.

[0043] S8: Load all parameters saved in step S7, set the validation set and test set images as input data, execute steps S2 to S6 on all input data, and calculate relevant indicators based on the detection results of step S6.

[0044] In an optional implementation, the global semantic feature extraction module in step S2 is a convolutional neural network with a 36-layer depthwise separable convolutional network as the backbone network, and the channel-dimensional global gram texture feature module in step S3 and the spatial-dimensional local gram texture feature module in step S4 are each convolutional neural networks with two layers of grouped convolutions.

[0045] In an optional implementation, the forgery detection process in step S6 is obtained using the binary entropy loss function:

[0046]

[0047] in Represents the binary entropy loss function. This is the result of a batch output from step S6. The image-level labels for the batch of face images obtained in step S1. , , The image number is the number of the face image in this batch. Indicates the first One prediction result, Indicates the first Image tags.

[0048] According to the embodiments of the present invention, at least the following beneficial effects are achieved: the multi-gram texture feature extraction module can effectively extract the texture features of the input image; the multi-gram texture feature extraction module not only follows the original definition of the gram matrix and realizes the extraction of global texture style, but also realizes the extraction of local texture features through the spatial dimension local gram texture feature module, thereby enhancing the expressive power of the network; the feature contrast enhancement module can effectively extract the texture differences between the fake and non-fake regions, thereby enhancing the texture features and achieving more effective discrimination. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the embodiments will be briefly introduced below. The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings. The accompanying drawings are schematic and should not be construed as limiting the present invention in any way. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a schematic diagram of a face depth forgery detection device based on multiple gram textures according to an embodiment of the present invention.

[0051] Figure 2This is a flowchart of the training process of a face depth forgery detection method based on multi-gram texture according to an embodiment of the present invention;

[0052] Figure 3 This is a schematic diagram of the structure and process of the multi-gram texture feature extraction module in a face depth forgery detection device based on multi-gram texture according to an embodiment of the present invention. Detailed Implementation

[0053] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.

[0054] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0055] refer to Figure 1 , Figure 3 This paper provides a detailed description of a face depth forgery detection device based on multi-gram texture according to an embodiment of the present invention. Figure 1 This is a schematic diagram of a face depth forgery detection device based on multi-gram texture according to an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure and process of the multi-gram texture feature extraction module in a face depth forgery detection device based on multi-gram texture according to an embodiment of the present invention.

[0056] According to an embodiment of the present invention, a face deepfake detection device based on multi-gram texture is provided, comprising: a global semantic feature extraction module, a multi-gram texture feature extraction module, and a feature fusion module, wherein the multi-gram texture feature extraction module includes a spatial dimension local gram texture feature module, a channel dimension global gram texture feature module, and a feature contrast enhancement module.

[0057] The global semantic feature extraction module is used to extract the face image to be tested for effective input into the device, and to extract shallow image features and global semantic features from the face image. This global semantic feature extraction module can obtain a face image and extract shallow image features using a shallow CNN (Convolutional Neural Network), thus realizing the extraction of shallow image features from the face image. The spatial dimensions of the shallow image features extracted by the shallow CNN in the convolutional neural network are C×H×W, representing the number of channels, height, and width of the image feature, respectively, where C represents the channel dimension, and H and W represent the spatial dimensions.

[0058] The multi-gram texture feature extraction module includes a spatial dimension local gram texture feature module, a channel dimension global gram texture feature module, and a feature contrast enhancement module. The multi-gram texture feature extraction module constructs gram texture features from the shallow image features extracted by the global semantic feature extraction module and enhances the texture characteristics of the input face image.

[0059] The spatial dimension local gram texture feature module describes the spatial dimension feature correlation of the face image by calculating the gram matrix between pixels of the shallow image features, thereby achieving modeling of spatial local texture. This spatial dimension local gram texture feature module can have a two-layer grouped convolutional neural network.

[0060] The channel-dimensional global gramm texture feature module models the overall style texture of an image by calculating the gramm matrix between different channels of the input face image. This channel-dimensional global gramm texture feature module can have a two-layer grouped convolutional neural network.

[0061] The feature contrast enhancement module (also known as the texture contrast enhancement module) performs feature residual calculation on the spatial gram texture features of the image tampering region and the background region of the shallow image features to obtain spatial gram texture contrast enhancement features. It also performs feature residual calculation on the channel gram style texture features of the image tampering region and the background region to obtain global channel gram texture contrast enhancement features. Finally, it splices the gram texture contrast enhancement features of the spatial dimension and the channel dimension to obtain the fused multi-gram texture features.

[0062] The feature fusion module obtains the global semantic features extracted by the global semantic feature extraction module and the final gram texture features constructed by the multi-gram texture feature extraction module. The global semantic features and the final gram texture features are then spliced ​​and fused to obtain fused features, thereby enabling the identification of forged images.

[0063] In the face deepfake detection apparatus based on multi-gram texture provided according to an embodiment of the present invention, a multi-gram texture feature extraction module can effectively extract the texture features of the input face image. Furthermore, by providing a multi-gram texture feature extraction module, not only is the original definition of the gram matrix followed, achieving global texture style extraction, but a spatial dimension local gram texture feature module is also provided to extract local texture features, enhancing the network's expressive power. Additionally, by providing a feature contrast enhancement module, the texture differences between fake and non-fake regions can be effectively extracted, thereby enhancing texture features and achieving more effective discrimination.

[0064] In an optional implementation, both the global semantic feature extraction module and the multigram texture feature extraction module employ convolutional neural networks (CNNs).

[0065] In an optional implementation, the spatial dimension local Gram texture feature module can divide the input shallow image features into squared feature blocks of any number of values, such as 4, 5, 7, or 9. To shorten training time and reduce computational load, it can be divided into 16 feature blocks. The channel dimension global Gram texture feature module can compress the input shallow image features into a number of channels selected from 32 to 96. To shorten training time and reduce computational load, the shallow image features can be compressed into 32 channels.

[0066] In an optional implementation, the face image extracted by the global semantic feature extraction module can be a square with a size between 256×256 and 512×512. To ensure extraction quality and reduce computation, a size of 320×320 pixels can be selected.

[0067] refer to Figure 2 , Figure 3 This paper provides a detailed description of a face depth forgery detection method based on multi-gram texture according to an embodiment of the present invention. Figure 2 This is a flowchart of the training process of a face depth forgery detection method based on multi-gram texture according to an embodiment of the present invention. Figure 3 This is a schematic diagram illustrating the structure and process of the multi-gram texture feature extraction module in a face depth forgery detection device based on multi-gram texture according to an embodiment of the present invention. This face depth forgery detection method based on multi-gram texture can be executed by the face depth forgery detection method based on multi-gram texture provided above according to an embodiment of the present invention.

[0068] According to an embodiment of the present invention, a method for detecting deep face forgery based on multi-gram texture is provided, specifically including the following steps S1 to S8.

[0069] S1: Construct a face image dataset consisting of face images synthesized using deepfake technology, with image-level annotations. These annotations can be 1 to represent a fake image and 0 to represent a real image.

[0070] S2: The global semantic feature extraction module obtains the face image from step S1, scales it, and extracts the global semantic features of the face image. and shallow image features , where i=1,2, and global semantic features The output dimension is Shallow image features The dimension is ,in, , These represent the length and width of the shallow image features, respectively. , This constitutes the spatial dimension of shallow image features. Let be the channel dimension of the shallow image features, where i = 1, 2. This global semantic feature... This module describes the global semantic features of the input face image. It obtains the face image and extracts shallow image features using a shallow CNN (Convolutional Neural Network), thus extracting shallow image features from the face image. The spatial dimensions of the shallow image features extracted by the shallow CNN are C×H×W, representing the number of channels, height, and width of the image features, respectively. Here, C represents the channel dimension, and H and W represent the spatial dimensions.

[0071] S3: The shallow image features obtained in step S2... The data is sequentially fed into the channel-dimensional global gramm texture feature module of the multi-gramm texture feature extraction module and processed, specifically including the following steps S3-1 to S3-6. This multi-gramm texture feature extraction module includes a spatial-dimensional local gramm texture feature module, a channel-dimensional global gramm texture feature module, and a feature contrast enhancement module. The channel-dimensional global gramm texture feature module may include convolutional layers and pooling layers.

[0072] S3-1: The channel-dimensional global Gram texture feature module utilizes spatial attention to calculate the mask. Processing received shallow image features ,in Dimensions , , These are shallow image features. Given the length and width of the mask, and the fact that all elements in the mask have a value range of [0,1], we can obtain... and ,in Representing shallow image features Shallow image features corresponding to the background region, features Representing shallow image features The shallow image features corresponding to the tampered region, where i = 1, 2, and * denotes matrix bitwise multiplication. See also Figure 3 The above uses spatial attention to calculate the mask. It can be provided by spatial attention module 2, which can be included in the channel-dimensional global gram texture feature module, or it can be provided separately.

[0073] S3-2: Obtain the features input from step S3-1 ,right The channel dimensions are compressed and the features of different channels are unfolded to obtain the image embedding features of the tampered region. Its size is ,in Representation of features Compressed channel dimensions, Representation of features The expanded spatial dimensions Optionally, features The compressed channel dimension C can be set to 32.

[0074] S3-3: Obtain the features input in step S3-2 Calculate the channel dimension Gram matrix:

[0075]

[0076] The channel dimension Gram matrix The global gram texture features of the image tampering region are obtained through convolutional and pooling layers of the channel-dimensional global gram texture feature module. In the above formula, T represents the transpose of the eigenvector.

[0077] S3-4: Obtain the features input from step S3-1 , for features The channel dimensions are compressed and the features of different channels are expanded to obtain the image embedding features of the background region. Its size is ,in Representation of features Compressed channel dimensions, Representation of features Spatial dimensions after feature expansion Optionally, features The compressed channel dimension C can be set to 32.

[0078] S3-5: Obtain the features input from step S3-4 Calculate the Gram matrix for the channel dimension:

[0079]

[0080] The channel dimension Gram matrix The global gram texture features of the image background region are obtained through convolutional and pooling layers of the channel-dimensional global gram texture feature module. In the above formula, T represents the transpose of the eigenvector.

[0081] S3-6: Obtain and The differences between them are then calculated to obtain the final global channel gram texture features. .

[0082] Through the above steps, the multigram texture feature extraction module can effectively extract the texture features of the input face image.

[0083] S4: Obtain the shallow image layer features from the input in step S2. The spatial dimension local gram texture feature module of the multi-gram texture feature extraction module is sequentially input and processed, specifically including the following steps S4-1 to S4-6. This spatial dimension local gram texture feature module may include convolutional layers and pooling layers.

[0084] S4-1: The spatial dimension local gram texture feature module utilizes spatial attention to calculate the mask. Processing received shallow image features This leads to the acquisition of shallow image features. Shallow image features corresponding to the mid-background region and shallow image features Shallow image features corresponding to the tampered area See also Figure 3 The above uses spatial attention to calculate the mask. This can be provided by the spatial attention module 1. The spatial attention module 1 can be included in the spatial dimension local gram texture feature module, or it can be provided separately.

[0085] S4-2: Obtain features from step S4-1 and , for features The image is divided into blocks along its spatial dimension and then expanded to obtain the embedding features of each block. ,feature The size is ,in N=S*S , ,in , These are shallow image features. Length and width, S The size of the block indicates the size of the segment. Represents the image embedding features of the blocks. Optionally, for The size S of the spatial dimension block can be set to 4 for better results.

[0086] S4-3: Obtain the features input from step S4-2 Calculate the Gram matrix for the channel dimension:

[0087]

[0088] The channel-dimensional gram matrix is ​​passed through convolutional and pooling layers to obtain the global gram texture features. In the above formula, T represents the transpose of the eigenvector.

[0089] S4-4: Features obtained from step S4-1 ,right The image is divided into blocks along its spatial dimension and then expanded to obtain the embedding features of each block. The embedding features of this block The size is ,in N=S*S , ,in S The size of the block indicates the size of the segment. Represents the image embedding features of the blocks. Optionally, for The size S of the spatial dimension block can be set to 4.

[0090] S4-5: Obtain the embedding features of the blocks input from step S4-4. Calculate the Gram matrix for the channel dimension:

[0091]

[0092] The channel-dimensional gram matrix is ​​passed through convolutional and pooling layers to obtain the global gram texture features. In the above formula, T represents the transpose of the eigenvector.

[0093] S4-6: Calculation and The difference between them yields the final global space gram texture feature. .

[0094] S5: The feature contrast enhancement module receives the features obtained from steps S4 and S3. and The features are obtained by concatenating i=1,2 and then performing a splicing operation to obtain multigram texture features. i=1,2 represent the multigram texture features obtained after inputting different shallow image features into the multigram texture feature extraction module. These multigram texture features are then fused through a concatenation operation. , , This represents the multigram texture features after fusion.

[0095] In the above steps, by providing a multi-gram texture feature extraction module, the texture features of the input face image can be effectively extracted. Furthermore, in these steps, the multi-gram texture feature extraction module not only follows the original definition of the Gram matrix and achieves global texture style extraction, but also provides a spatial dimension local Gram texture feature module to achieve local texture feature extraction, enhancing the network's expressive power. Additionally, in the above steps, by providing a feature contrast enhancement module, the texture differences between forged and non-forged regions can be effectively extracted, thereby enhancing texture features and achieving more effective discrimination.

[0096] S6: Apply the global semantic features obtained from step S2 The fused multigram texture features obtained in step S5 The input is given to the feature fusion module, where a splicing operation is performed. The final image features are obtained by feature fusion. This indicates that the final image features The input image is fed into a classifier to detect forgery and obtain the final detection result. This forgery detection can be achieved by feeding the final image features into different classifiers. In an optional implementation, a fully connected layer of a deep neural network can be used as a classifier to perform this forgery detection processing.

[0097] S7: Repeat steps S2-S6 until the loss function converges, complete the training and save all parameters in the global semantic feature extraction module in step S2, the channel-dimensional global gramm texture feature module in step S3, and the spatial-dimensional local gramm texture feature module in step S4.

[0098] S8: Load all parameters saved in step S7, set the validation set and test set images as input data, execute steps S2-S6 on all input data, and calculate relevant indicators based on the detection results of step S6.

[0099] In one optional implementation, the global semantic feature extraction module in step S2 may be a convolutional neural network with a 36-layer depthwise separable convolutional network as the backbone. Alternatively, the channel-dimensional global gram texture feature module in step S3 and the spatial-dimensional local gram texture feature module in step S4 may each be a convolutional neural network with two layers of grouped convolutions.

[0100] In an optional implementation, the forgery detection process in step S6 is performed using a loss function:

[0101]

[0102] in, This is the result of a batch output from step S6. The image-level labels for the batch of face images obtained in step S1. , , The image number is the number of the face image in this batch. Indicates the first One prediction result, Indicates the first Image tags, This represents the binary entropy loss function.

[0103] Optionally, the number of iterations in step S6 can be set to... =10 rounds, to achieve better results and reduce computational load. In other implementations, the number of rounds can be increased or decreased as needed.

[0104] Optionally, the size of the convolutional kernel in steps S3-3 and S3-5 can be set to... =3, to achieve better results and reduce computational cost. In other implementations, the size of the convolution kernel can be increased or decreased as needed.

[0105] Optionally, the size of the convolutional kernel in steps S4-3 and S4-5 can be set to... =3, to achieve better results and reduce computational cost. In other implementations, the size of the convolution kernel can be increased or decreased as needed.

[0106] The settings in the above-described method provided by the embodiments of the present invention have the advantages of short training time and small number of parameters.

[0107] The following details the verification process of the face depth forgery detection method based on multi-gram texture provided by the present invention. To verify the effectiveness and practicality of the invention, FaceForensic++ (face forensics++ dataset) was used as the training dataset (5000 videos). The model was trained according to steps S1-S7 above, using Adam (adaptive momentum estimation optimization algorithm) as the optimizer, with a learning rate set to 0.0005. 72% of the training data was used to train the model, and 14% was used as the validation model. The model was trained for 10 iterations, with the learning rate decreasing to half every 5 iterations. Finally, the model with the best evaluation metric on the validation set was saved as the final result.

[0108] The model was evaluated using the FaceForensic++ dataset, which contains 140 real videos and 560 fake videos. The trained model was evaluated according to step S8 above and compared with the real labels. Under the LQ (low quality images) setting of the FaceForensic++ dataset, the accuracy index of this method reached 91.2% and the AUC (area of ​​return) reached 93.8%. Under the HQ (high quality images) setting, the accuracy index of this method reached 96.4% and the AUC reached 98.89%. These values ​​indicate that the method provided by the embodiment of the present invention yields high-level results, demonstrating the efficiency and feasibility of the present invention.

[0109] The robustness of the model was evaluated using CelebDF (Celebrity Deepfake Dataset). The evaluation dataset contains 590 real videos and 5963 fake videos. The model trained on the FaceForensics++ dataset was evaluated according to step S8 above. The AUC reached 86.78% in setting v1 (version 1) and 77.66% in setting v2 (version 2). These values ​​indicate that the method provided by the embodiment of the present invention yields high-level results, demonstrating the efficiency and feasibility of the present invention.

[0110] It should be understood that the foregoing only illustrates some embodiments, and changes, modifications, additions, and / or variations can be made without departing from the scope and spirit of the disclosed embodiments. These embodiments are illustrative and not restrictive. Furthermore, the described embodiments relate to those currently considered most practical and preferred, and should be understood as not being limited to the disclosed embodiments, but rather intended to cover different modifications and equivalent arrangements included within the spirit and scope of those embodiments. Moreover, the various embodiments described above can be used in conjunction with other embodiments; for example, an aspect of one embodiment can be combined with an aspect of another embodiment to achieve yet another embodiment. Additionally, individual features or components of any given component can constitute another embodiment.

[0111] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.

[0112] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A face deepfake detection device based on multi-gram texture, characterized in that, The device comprises: a global semantic feature extraction module which extracts a human face image to be tested and extracts shallow image features and global semantic features from the human face image; a multi-gram texture feature extraction module which constructs the shallow image features extracted by the global semantic feature extraction module into shallow image features and strengthens the texture characteristics of the input human face image, the multi-gram texture feature extraction module comprising a spatial dimension local gram texture feature module, a channel dimension global gram texture feature module and a feature contrast strengthening module, the spatial dimension local gram texture feature module describes the feature correlation of the spatial dimension of the human face image by calculating the gram matrix between the pixels of the shallow image features, and models the spatial local texture, the channel dimension global gram texture feature module models the style texture of the whole image by calculating the gram matrix between different channels of the input human face image, the feature contrast strengthening module performs feature residual calculation on the spatial gram texture features of the image tampering region and the background region of the shallow image features to obtain spatial gram texture contrast strengthening features, performs feature residual calculation on the channel gram style texture features of the image tampering region and the background region to obtain global channel gram texture contrast strengthening features, and splices the gram texture contrast strengthening features of the spatial dimension and the channel dimension to obtain fused multi-gram texture features; a feature fusion module which obtains the global semantic features extracted by the global semantic feature extraction module and the final gram texture features constructed by the multi-gram texture feature extraction module, splices and fuses the global semantic features and the final gram texture features to obtain fused features, and realizes the discrimination of the fake image.

2. The human face deep fake detection device based on multi-gram texture according to claim 1, wherein: the global semantic feature extraction module and the multi-gram texture feature extraction module both adopt a convolutional neural network.

3. The human face deep fake detection device based on multi-gram texture according to claim 1, wherein: the spatial dimension local gram texture feature module divides the input shallow image features into feature blocks selected from the square of a number selected from 4, 5, 7 and 9, the channel dimension global gram texture feature module compresses the input shallow image features into a number of channels selected from 32-96.

4. The human face deep fake detection device based on multi-gram texture according to claim 1, wherein: the global semantic feature extraction module extracts human face images with any square size between 256x256 and 512x512.

5. A face deepfake detection method based on multi-gram texture, characterized in that, The method comprises the following steps: S1: constructing a human face image dataset comprising human face images synthesized by using deep fake technology, the human face image dataset having image-level labels, wherein 1 represents a fake image and 0 represents a real image; S2: a global semantic feature extraction module obtains the face image from step S1 and scales it, and obtains the global semantic feature of the face image and shallow image features , wherein i = 1, 2, wherein the global semantic feature has an output dimension of , used to describe the global semantic feature of the input face image, the shallow image feature has a dimension of , wherein , respectively the length and width of the shallow image feature, and constitute the spatial dimension of the shallow image feature, is the channel dimension of the shallow image feature; S3: The shallow image features obtained in step S2... The data is sequentially fed into the channel-dimensional global gram texture feature module and processed, specifically including the following steps: S3-1: The channel dimension global gram texture feature module calculates a mask using spatial attention Processing the received shallow image features wherein The dimension is , , The length and width of the shallow image features respectively, all elements in the mask have a value range of [0, 1], and then and wherein represents the shallow image features corresponding to the background region in the shallow image features , and the features represent the shallow image features corresponding to the tampered region in the shallow image features , wherein i = 1, 2, and wherein * represents matrix bitwise multiplication, S3-2: Obtain the features input from step S3-1 , the channel dimension of the tampered region is compressed and the features of different channels are unfolded to obtain image embedding features of the tampered region , the size of which is , wherein represents the features , the compressed channel dimension, represents the features , the unfolded space dimension, ,​ S3-3: obtaining the features inputted in step S3-2 , computing the channel dimension gram matrix The channel-wise Gram matrix The global Gram texture features of the image tampered region are obtained through the convolutional layer and the pooling layer , S3-4: Obtain the feature input from step S3-1 , the channel dimension of the feature is compressed and the features of different channels are unfolded to obtain the image embedding feature of the background region , the size of which is , wherein represents the channel dimension of the feature after compression, represents the spatial dimension of the feature after feature unfolding, , S3-5: obtaining the features input in step S3-4 gramian matrix of the channel dimension The channel-wise Gram matrix The global Gram texture features of the image background region are obtained through the convolutional layer and the pooling layer , S3-6: obtaining and and calculating the difference between them to obtain the final global channel gram texture feature ; S4: Obtain the shallow image layer features input from step S2 The spatial dimension local Gromov texture feature module is sequentially input and processed, specifically including the following steps S4-1: The spatial dimension local gram texture feature module calculates a mask using spatial attention processing the received shallow image features , and obtaining the shallow image features corresponding to the background region in the middle image , and the shallow image features corresponding to the tampered region in the middle image , S4-2: obtain features from step S4-1 and , the spatial dimension of the features is partitioned and unfolded to obtain partitioned embedding features , the size of the partitioned embedding features is , where N=S*S , , where , are the length and width of the shallow image features , respectively, S denotes the size of the partitioning, denotes the image embedding features of the partition, S4-3: obtain the features input from step S4-2 , compute the Gram matrix of the channel dimension: The channel dimension gram matrix is subjected to a convolution layer and a pooling layer to obtain a global gram texture feature ; S4-4: The feature obtained according to step S4-1 is split into blocks and unfolded to obtain a block-wise embedding feature , the spatial dimension of the block-wise embedding feature is , , the size of the block-wise embedding feature is , , wherein N=S*S , , wherein S denotes the size of the block size, denotes the image embedding feature of the block. S4-5: obtaining the features input from step S4-4 , computing the Gram matrix of the channel dimension: The channel dimension gram matrix is subjected to a convolution layer and a pooling layer to obtain a global gram texture feature ; S4-6: Compute and the difference between the two results in the final global spatial gram texture feature ; S5: The feature contrast enhancement module receives the features obtained by step S4 and step S3 and , i = 1, 2, and a plurality of Gram texture features are obtained through a splicing operation , respectively represent the plurality of Gram texture features obtained by inputting different shallow image features into the plurality of Gram texture feature extraction modules, the plurality of Gram texture features are fused through a splicing operation to obtain fused plurality of Gram texture features , ; S6: The feature fusion module receives the global semantic features obtained in step S2 and the fused multi-gram texture features obtained in step S5 and performs a stitching operation to realize feature fusion and obtain final image features The final image features F are then sent to a classifier to realize forgery discrimination of the input image and obtain a final detection result. S7: Repeat steps S2 to S6 until the loss function converges, complete the training and save all parameters in the global semantic feature extraction module in step S2, the channel dimension global gram texture feature module in step S3 and the spatial dimension local gram texture feature module in step S4; S8: Load all parameters saved in step S7, set the validation set and test set images as input data, execute steps S2 to S6 on all input data, and calculate the relevant indicators according to the detection results of step S6.

6. The face deep fake detection method based on multiple gram textures according to claim 5, wherein: The global semantic feature extraction module in step S2 is a convolutional neural network with a 36-layer depth separable convolution network as a backbone network.

7. The face deep fake detection method based on multiple gram textures according to claim 5, wherein: The channel dimension global gram texture feature module in step S3 and the spatial dimension local gram texture feature module in step S4 are each a convolutional neural network with two layers of grouped convolution.

8. The face deep fake detection method based on multiple gram textures according to claim 5, wherein: The process of fake discrimination in step S6 is obtained by a binary entropy loss function: wherein, is the result of one batch of step S6, is the image-level label of the batch of face images obtained in step S1, , , is the number of the face image in the batch, denotes the th prediction result, denotes the label of the th image.

Citation Information

Patent Citations

  • Deep forgery detection method and device, storage medium and computer equipment

    CN114463805A

  • System and Method for Extremely Efficient Image and Pattern Recognition and Artificial Intelligence Platform

    US20200184278A1