Face counterfeit image detection method and device
The training neural network model extracts facial features and artifact features, and combines the visual converter for partition analysis, which solves the problem of low detection efficiency and accuracy of facial forged images in the prior art, and achieves more efficient forged image recognition.
Patent Information
- Application Number
- CN202411883411.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-07-08
AI Technical Summary
The existing facial forgery image detection methods rely on expert experience design models, making it difficult to dig out the detailed characteristics of facial tampering, resulting in low detection efficiency and accuracy.
A facial forgery image detection method is adopted to extract facial features, artifact features and partitioned features through the trained neural network model, feature extraction is used using the Encoder and Decoder modules of the U-Net model, partition analysis is performed by combining the visual converter (ViT), and detection results are output through the softmax classification layer.
The accuracy and efficiency of facial forgery image detection are improved, and the forgery image can be accurately and efficiently identified.
Smart Images

Figure CN120279604A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and device for detecting forged facial images. Background Art
[0002] For media messages, it is usually to change the facial information involved therein to forge the identity information of the person. With the development of artificial intelligence technology, the generated effect of forged facial images is becoming more and more realistic, resulting in the difficulty of distinguishing the authenticity of media messages.
[0003] Existing methods for detecting forged facial images rely on expert experience to design models. These models are difficult to mine the detailed features of facial tampering, and the detection efficiency and accuracy of forged facial images are relatively low. Summary of the Invention
[0004] The present invention provides a method and device for detecting forged facial images to solve the problem of relatively low detection efficiency and accuracy of forged facial images in the prior art.
[0005] The present invention provides a method for detecting forged facial images, including the following steps:
[0006] Obtain a facial image to be detected;
[0007] Input the facial image to be detected into a trained neural network model to obtain a detection result of the facial image to be detected output by the neural network model;
[0008] The neural network model includes a first convolutional module, a second convolutional module, a visual transformation module, and a classification layer;
[0009] The first convolutional module is used to extract the facial features of the facial image to be detected;
[0010] The second convolutional module is used to extract the artifact features of the facial image to be detected;
[0011] The visual transformation module is used to extract the partition features of the facial image to be detected;
[0012] The classification layer is used to output a detection result of the facial image to be detected.
[0013] According to a method for detecting forged facial images provided by the present invention, before the step of inputting the facial image to be detected into a trained neural network model to obtain a detection result of the facial image to be detected output by the neural network model, it further includes:
[0014] Obtain real facial images and forged facial images;
[0015] Generate a high-frequency artifact image based on the real face image and the forged face image;
[0016] Use the real face image and the forged face image as sample data, and use the high-frequency artifact image as the sample label data corresponding to the sample data to train the neural network model.
[0017] According to a method for detecting forged face images provided by the present invention, the generating a high-frequency artifact image based on the real face image and the forged face image includes:
[0018] Perform high-pass filtering on the forged face image to obtain a high-frequency image; and calculate the absolute value of the difference between the real face image and the forged face image to obtain an artifact image;
[0019] Generate a high-frequency artifact image based on the overlapping part of the high-frequency image and the artifact image.
[0020] According to a method for detecting forged face images provided by the present invention, the inputting the face image to be tested into the trained neural network model to obtain the detection result of the face image to be tested output by the neural network model includes:
[0021] Input the face image to be tested into the first convolution module to obtain the face features of the face image to be tested output by the first convolution module;
[0022] Input the face features into the second convolution module to obtain the artifact features of the face image to be tested output by the second convolution module;
[0023] Based on the face features and the artifact features, obtain the partition features of the face image to be tested through the visual transformation module;
[0024] Input the partition features into the classification layer to obtain the detection result of the face image to be tested output by the classification layer.
[0025] According to a method for detecting forged face images provided by the present invention, the obtaining the partition features of the face image to be tested through the visual transformation module based on the face features and the artifact features includes:
[0026] Determine the weight of the artifact features based on the artifact features;
[0027] Perform linear weighting on the feature vectors in the face features based on the weight of the artifact features to obtain weighted features;
[0028] According to the facial region where each weighted feature is located, the weighted feature is partitioned by the visual conversion module to obtain the partitioned feature output by the visual conversion module.
[0029] According to a method for detecting forged facial images provided by the present invention, the partitioned features include eye features, facial center features, and facial contour features.
[0030] The present invention also provides a device for detecting forged facial images, including the following modules:
[0031] An acquisition module, configured to acquire a facial image to be detected;
[0032] A detection module, configured to input the facial image to be detected into a trained neural network model to obtain a detection result of the facial image to be detected output by the neural network model;
[0033] The neural network model includes a first convolutional module, a second convolutional module, a visual conversion module, and a classification layer;
[0034] The first convolutional module is used to extract facial features of the facial image to be detected;
[0035] The second convolutional module is used to extract artifact features of the facial image to be detected;
[0036] The visual conversion module is used to extract partitioned features of the facial image to be detected;
[0037] The classification layer is used to output a detection result of the facial image to be detected.
[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, the method for detecting forged facial images as described in any one of the above is implemented.
[0039] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for detecting forged facial images as described in any one of the above is implemented.
[0040] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for detecting forged facial images as described in any one of the above is implemented.
[0041] A method for detecting forged facial images provided by the present invention detects a facial image to be tested through a trained neural network model. Among them, the neural network model includes a first convolutional module, a second convolutional module, a visual transformation module, and a classification layer. When detecting the facial image to be tested, first, the facial features of the facial image to be tested are extracted through the first convolutional module, the artifact features of the facial image to be tested are extracted through the second convolutional module, and the partition features of the facial image to be tested are extracted through the visual transformation module, so that the attention area of the neural network model changes from the entire face to detailed artifacts, thereby promoting the neural network model to extract effective tampering features and improving the detection accuracy of the neural network model. Then, the detection result of the facial image to be tested is output through the classification layer, realizing accurate and efficient detection of forged facial images. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 FIG. 1 is a schematic flowchart of a method for detecting forged facial images provided by the present invention.
[0044] Figure 2 FIG. 2 is a schematic flowchart of a method for detecting forged facial images provided by the present invention.
[0045] Figure 3 FIG. 3 is a schematic diagram of the generation process of high-frequency artifact images in a method for detecting forged facial images provided by the present invention.
[0046] Figure 4 FIG. 4 is a schematic diagram of the triple partitioning of features by ViT in a method for detecting forged facial images provided by the present invention.
[0047] Figure 5 FIG. 5 is a schematic diagram of the feature partitioning result in a method for detecting forged facial images provided by the present invention.
[0048] Figure 6 FIG. 6 is a scatter plot generated during the process of the neural network model detecting the facial image to be tested in a method for detecting forged facial images provided by the present invention.
[0049] Figure 7 FIG. 7 is a schematic structural diagram of a device for detecting forged facial images provided by the present invention.
[0050] Figure 8 FIG. 8 is a schematic structural diagram of an electronic device provided by the present invention. Detailed implementation manners
[0051] Tampering with video media usually involves altering facial information to forge a person's identity information. To filter out valid and reliable media information, the present invention studies a method for forensic authentication of face manipulation to ensure the reliability of media messages.
[0052] In the early stage, face anti-spoofing was performed by collecting statistical data such as facial texture and illumination to perform regression classification. As the generation effect of forged images becomes more and more realistic, the corresponding detection methods have been gradually improved. To enhance the ability of the model to extract forged features and the detection accuracy of forged images, researchers, based on deep learning models, regard the true / false binary classification problem as a fine-grained task and adopt multi-task techniques to enrich feature mining.
[0053] In fact, it is difficult for forged images to be exactly the same as real images. In the frequency domain, the common differences between the images generated after upsampling and natural images will be revealed. Therefore, the extraction of frequency domain features will help to discover artifact clues. Existing frequency-based face forged image detection methods include: the two-stream cooperation framework F 3 -Net that combines frequency with Convolutional Neural Networks (CNNs), Frequency-aware Discriminative Feature Learning Supervised by Single-Center Loss for Face Forgery Detection (FDFL) that constructs frequency clues in fine-grained CNNs, and a detection model that combines frequency with multi-task CNNs.
[0054] In the prior art, face forged image detection methods rely on expert experience to design models, resulting in the problem of low detection accuracy. Although conventional deep learning solutions can improve traditional methods, it is still difficult to discover the details of face tampering. And existing frequency methods, although they can enhance the anti-spoofing performance of deep learning methods by leveraging the correlation between frequency and tampering artifacts, the utilization of frequency features is relatively rough, so it is difficult to exert the advantages of frequency in the face tampering detection task.
[0055] To solve the above problems, the present invention designs a multi-task learning network that combines frequency, a vision transformer that combines feature correlation analysis, and a contrastive loss for feature space optimization. The three cooperate with each other to jointly construct a method for detecting face forged images.
[0056] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0057] The following will describe Figures 1 to 8 a method and device for detecting forged facial images of the present invention.
[0058] Figure 1 is one of the flow diagrams of a method for detecting forged facial images provided by the present invention. As Figure 1 shown, the method includes step 101 and step 102.
[0059] Step 101: Obtain a facial image to be detected.
[0060] Specifically, the facial image to be detected can be a face image intercepted from a video, a face photo intercepted from a media message, or any face image downloaded from other databases, etc. The facial image to be detected can be either a real, untreated facial image, a completely forged facial image, or a facial image with local forgery on the basis of a real image. Before being detected, the authenticity of the facial image to be detected cannot be determined.
[0061] Step 102: Input the facial image to be detected into a trained neural network model to obtain a detection result of the facial image to be detected output by the neural network model; the neural network model includes a first convolutional module, a second convolutional module, a visual transformation module, and a classification layer;
[0062] The first convolutional module is used to extract facial features of the facial image to be detected;
[0063] The second convolutional module is used to extract artifact features of the facial image to be detected;
[0064] The visual transformation module is used to extract partition features of the facial image to be detected;
[0065] The classification layer is used to output a detection result of the facial image to be detected.
[0066] Specifically, use a trained neural network model to detect the facial image to be detected, and obtain a detection result output by the neural network model. The detection result is either real or forged.
[0067] Figure 2 is the second flow diagram of a method for detecting forged facial images provided by the present invention. AsFigure 2 as shown
[0068] The neural network model includes a first convolutional module, a second convolutional module, a vision transformation module, and a classification layer. Among them, the first convolutional module is the Encoder module in the U-Net model. The Encoder module can extract the facial features of the facial image to be tested input into the neural network model through downsampling; the second convolutional module is the Decoder module in the U-Net model. The Decoder module can extract the artifact features of the facial image to be tested input into the neural network model through upsampling; the vision transformation module is the Vision Transformer (ViT), which is a network that uses the self-attention mechanism to capture the relationships between elements in a sequence. In the embodiments of the present invention, ViT is used to extract the partition features of the facial image to be tested input into the neural network model; the classification layer is the output layer of the neural network model for binary classification and can be constructed based on functions for classification such as the sigmoid function or the softmax function. In the embodiments of the present invention, taking the softmax function as an example, the classification layer of the neural network model is constructed, and the classification layer can output the detection result of the facial image to be tested.
[0069] In the embodiments of the present invention, first, the backbone part of the U-Net model (i.e., the Encoder module and the Decoder module) is used for feature extraction, and then ViT is combined to partition the extracted features to obtain partition features, so that the attention area of the neural network model changes from the entire face to the detailed artifacts in different areas. Finally, the classification layer is constructed through the softmax function to output the detection result of the facial image to be tested, improving the detection accuracy and detection efficiency of the neural network model.
[0070] A method for detecting forged facial images provided by the present invention detects a facial image to be tested through a trained neural network model. Among them, the neural network model includes a first convolutional module, a second convolutional module, a vision transformation module, and a classification layer. When detecting the facial image to be tested, the first convolutional module extracts the facial features of the facial image to be tested, the second convolutional module extracts the artifact features of the facial image to be tested, and the vision transformation module extracts the partition features of the facial image to be tested, and the classification layer outputs the detection result of the facial image to be tested, so that the attention area of the neural network model changes from the entire face to the detailed artifacts, thereby promoting the neural network model to extract effective tampering features, improving the detection accuracy of the neural network model, and realizing accurate and efficient detection of forged facial images.
[0071] Optionally, before inputting the facial image to be tested into the trained neural network model to obtain the detection result of the facial image to be tested output by the neural network model, the method further includes:
[0072] Obtain a real face image and a forged face image;
[0073] Generate a high-frequency artifact image based on the real face image and the forged face image;
[0074] Use the real face image and the forged face image as sample data, and use the high-frequency artifact image as the sample label data corresponding to the sample data to train the neural network model.
[0075] Specifically, before using the neural network model to detect a face image to be measured, the neural network model needs to be trained so that the neural network model can learn the relevant features in the forged face image, so as to accurately discriminate the forged face image.
[0076] In the embodiment of the present invention, first, a large number of real face images and corresponding forged faces (i.e., Figure 2 the manipulated faces in) images are used as the original data set. Among them, the real face images and the forged face images correspond one by one to form a true and false image pair. The number of real face images and the corresponding forged face images can be 100, 500, 1000, etc., and the specific number is not limited; then, based on the original data set, a high-frequency artifact image is generated; finally, using the real face image and the forged face image as sample data, and using the high-frequency artifact image as the sample label data corresponding to the sample data, a supervised learning training is performed on the neural network model, so that the model fully learns the features in the high-frequency artifact image. Thus, the trained neural network model can accurately and quickly detect an unknown face image to be measured.
[0077] Optionally, the generating a high-frequency artifact image based on the real face image and the forged face image includes:
[0078] Perform high-pass filtering on the forged face image to obtain a high-frequency image; and calculate the absolute value of the difference between the real face image and the forged face image to obtain an artifact image;
[0079] Generate a high-frequency artifact image based on the overlapping part of the high-frequency image and the artifact image.
[0080] Specifically, further processing based on the real face image and the forged face image can generate a high-frequency artifact image. Figure 3 is a schematic diagram of the high-frequency artifact image generation process in a face forgery image detection method provided by the present invention, as Figure 3 shown. The high-frequency artifact image generation process includes:
[0081] (1) The real face (i.e., Figure 3The GT graph) images and the corresponding forged faces (i.e., Figure 3 The manipulated faces) images are matched one by one to form genuine and forged image pairs, and high-pass filtering is performed on the forged face images among them to generate high-frequency images.
[0082] The process of performing high-pass filtering on the forged face images is as follows: After discrete Fourier transform, high-frequency signal filtering is performed, and then the filtered image is restored through inverse Fourier transform to obtain a high-frequency image. The corresponding frequency transformation formula is:
[0083] M Fre = IDFT{H (u,v) ·DFT{I input}}
[0084] Among them, M Fre represents the high-frequency image, IDFT{·} represents the inverse (or reverse) Fourier transform, H (u,v) represents filtering, u represents the horizontal component, v represents the vertical component, and DFT{I input} represents performing discrete Fourier transform on the input image I input .
[0085] (2) Calculate the absolute value of the difference between the real face image and the forged face image to obtain the artifact image, that is, calculate the absolute difference map (i.e., the artifact image) between the input real face image and the corresponding forged face image. The calculation formula is:
[0086] M AD = |I input - I GT |
[0087] Among them, M AD represents the artifact image, I input represents the input forged face image, and I GT represents the real face image corresponding to the forged face image.
[0088] (3) Based on the overlapping part of the high-frequency image and the artifact image, generate a high-frequency artifact image and use the high-frequency artifact image as the model label. That is, calculate the boundary sets set() of the values in M Fre and M AD respectively, and then store them in two dictionaries dicts(). Then, the intersection of the two dictionaries is the overlapping part of the high-frequency image and the artifact image. According to the overlapping part, it can be converted into a high-frequency artifact image. The corresponding conversion formula is:
[0089] M label = Bin{dict{set(M Fre )}∩dict{set(M AD )}}
[0090] Among them, M label represents the high-frequency artifact image, Bin{·∩·} represents calculating the intersection, and dict{set(M Fre )} represents the dictionary for storing the boundary set set(M Fre ) of the values in the high-frequency image M Fre ), and dict{set(M AD )} represents the dictionary for storing the boundary set set(M AD ) of the values in the artifact image M AD ).
[0091] In some embodiments, before further processing the real face image and the forged face image, it is also necessary to preprocess the real face image and the forged face image according to the actual requirements of the network model. For example, perform grayscale transformation on the input image, or align and crop the overlapping target image regions using the detected key points to meet the processing requirements of the network model for the image.
[0092] In the embodiment of the present invention, by performing high-pass filtering on the forged face image, a high-frequency image is obtained, and the absolute value of the difference between the real face image and the forged face image is calculated to obtain an artifact image. Then, based on the overlapping part of the high-frequency image and the artifact image, a high-frequency artifact image is generated, so as to incorporate frequency cues into the absolute difference label, thereby narrowing the positioning range of the artifact and further exerting the frequency characteristics. From Figure 3 it can be seen that the high-frequency artifact image generated in the embodiment of the present invention includes both facial edge artifacts and high-frequency artifacts. Compared with the traditional method, the sample label data made from the high-frequency artifact image solves the problems of inaccurate positioning of the artifact label and inconsistent multi-task segmentation labels for face forgery, making the sample label data include more comprehensive and detailed face tampering features, efficiently capturing tampering features for the authenticity detection task, thereby improving the training effect of the neural network model, enabling the neural network model to learn more comprehensive and detailed features through training, and improving the detection accuracy of the neural network model for the face image to be tested in subsequent use.
[0093] Optionally, the step of inputting the face image to be tested into the trained neural network model to obtain the detection result of the face image to be tested output by the neural network model includes:
[0094] Inputting the face image to be tested into the first convolutional module to obtain the facial features of the face image to be tested output by the first convolutional module;
[0095] Inputting the facial features into the second convolutional module to obtain the artifact features of the face image to be tested output by the second convolutional module;
[0096] Based on the facial features and the artifact features, through the visual transformation module, obtain the partition features of the facial image to be measured;
[0097] Input the partition features into the classification layer to obtain the detection result of the facial image to be measured output by the classification layer.
[0098] Specifically, the neural network model includes a first convolutional module, a second convolutional module, a visual transformation module, and a classification layer. Among them, the first convolutional module is the Encoder module in the U-Net model. The Encoder module can extract the facial features of the facial image to be measured input into the neural network model through downsampling; the second convolutional module is the Decoder module in the U-Net model. The Decoder module can extract the artifact features of the facial image to be measured input into the neural network model through upsampling; the visual transformation module is the Vision Transformer (ViT), which is a network that uses the self-attention mechanism to capture the relationships between elements in a sequence. In the embodiments of the present invention, ViT is used to extract the partition features of the facial image to be measured input into the neural network model; the classification layer is the output layer of the neural network model for binary classification, which can be constructed based on functions for classification such as the sigmoid function or the softmax function. In the embodiments of the present invention, taking the softmax function as an example, the classification layer of the neural network model is constructed, and the classification layer can output the detection result of the facial image to be measured.
[0099] The specific process of inputting the facial image to be measured into the trained neural network model to obtain the detection result of the facial image to be measured output by the neural network model is as follows:
[0100] Input the facial image to be measured into the Encoder module. After the Encoder module performs downsampling on the facial image to be measured, output the facial features of the facial image to be measured;
[0101] Input the facial features output by the Encoder module into the Decoder module. After the Decoder module performs upsampling on the facial image to be measured, output the artifact features of the facial image to be measured;
[0102] Input both the facial features output by the Encoder module and the artifact features output by the Decoder module into ViT (see also Figure 2 ), and through ViT, divide the features according to the facial regions where the features are located (such as eyes and eyebrows, the center of the face, and the facial contour, etc.) to obtain the partition features output by ViT;
[0103] The partition features output by the ViT are input into the softmax classification layer. The softmax classification layer judges the authenticity of the face image to be tested based on the partition features and outputs any one of the detection results of a real image or a forged image.
[0104] In an embodiment of the present invention, the backbone part of the U-Net model (i.e., the Encoder module and the Decoder module) is combined with the ViT and the softmax classification layer to construct a neural network model. The face features are extracted through the Encoder module, the artifact features are extracted through the Decoder module, and then the extracted features are partitioned by the ViT to obtain partition features, so that the attention area of the neural network model changes from the entire face to the detailed artifacts in different areas. Finally, through the softmax classification layer, the detection result of the face image to be tested is output, improving the detection accuracy and detection efficiency of the neural network model.
[0105] Optionally, obtaining the partition features of the face image to be tested based on the face features and the artifact features through the vision transformation module includes:
[0106] Determine the weight of the artifact features based on the artifact features;
[0107] Linearly weight the feature vectors in the face features based on the weight of the artifact features to obtain weighted features;
[0108] According to the face area where each weighted feature is located, partition the weighted features through the vision transformation module to obtain the partition features output by the vision transformation module.
[0109] Specifically, the vision transformation module is the Vision Transformer (ViT). The ViT can further analyze the face features and the artifact features, perform correlation analysis on the features from the whole to the local, and perform feature partitioning, thereby enhancing the parsing performance of the neural network model for the features.
[0110] In the embodiments of the present invention, facial features and artifact features are jointly used as the input of the ViT. Through the self-attention mechanism in the ViT, the weights of the artifact features are determined, and then linearly weighted in combination with the facial features to increase the weights of the artifact features, obtaining weighted features, so that the model pays more attention to the artifact features in the facial features. Then, the weighted features are input into the multi-head attention module in the ViT, and contrastive loss analysis is performed according to the facial region where each weighted feature is located, that is, the training batch data is divided into positive and negative samples according to the supervised sample label data, and a contrastive learning paradigm is constructed based on any positive sample. According to the contrastive result, the weighted features are partitioned to obtain partitioned features, so as to provide a basis for the classification output of the subsequent classification layer. The expression of the contrastive loss function is as follows:
[0111]
[0112] Among them, L SC is the contrastive loss function, q represents the anchor sample feature, k + represents the positive sample feature, τ represents the temperature coefficient, which is used to control the distinguishability of the model for negative samples, and k i represents the negative sample feature.
[0113] In the embodiments of the present invention, the ViT is combined to partition the facial features and artifact features, further refine and analyze the extracted features, and optimize the distribution of the features through the contrastive loss, so that the features belonging to the same region range are divided into the same region, and the features belonging to different regions are divided into different regions, improving the intra-class compactness and inter-class distinguishability of the features, thereby improving the detection accuracy of the neural network model for the facial images to be tested.
[0114] Optionally, the partitioned features include eye features, facial center features, and facial contour features.
[0115] Specifically, in the embodiments of the present invention, triple partitioning is performed through the ViT (in other cases, multiple partitioning can also be performed according to feature categories, and the present invention takes triple partitioning as an example for illustration), that is, the partitioned features include eye features, facial center features, and facial contour features.
[0116] Figure 4 is a schematic diagram of the triple partitioning of features by the ViT in a facial forgery image detection method provided by the present invention, as Figure 4As shown in the figure. The three-partition vision transformer in ViT groups patches with highly correlated encoded features while retaining the initial position embedding information, thus dividing them into a triple partition that matches facial features. Among them, a single partition of the triple partition consists of 8 patches with initial position embeddings. Each partition is fed into the multi-head attention module in ViT, parsed separately and then combined, enabling the model to consider the correlation between partitions and patches during the detection process, thereby improving the classification accuracy of the neural network model in subsequent classification tasks.
[0117] Figure 5 It is a schematic diagram of the feature partition result in a facial forgery image detection method provided by the present invention. As Figure 5 shown, after partitioning the input features through ViT, the partitioned features obtained include eye features (such as eyes and eyebrows), facial center features, and facial contour features.
[0118] In the embodiment of the present invention, partitioning the input features through ViT can obtain partitioned features including eye features, facial center features, and facial contour features, enabling the neural network model to further analyze more detailed features, thereby improving the detection accuracy of the subsequent facial image to be detected.
[0119] Based on the above embodiment, the present invention also experimentally verified the performance of the above neural network model, that is, during the process of the neural network model detecting the facial image to be detected, a scatter plot was used to qualitatively analyze the neural network model.
[0120] Figure 6 It is a scatter plot generated during the process of the neural network model detecting the facial image to be detected in a facial forgery image detection method provided by the present invention. As Figure 6 shown. From Figure 6 the scatter plot presenting the sample feature distribution, it can be seen that the model trained using high-frequency artifact images as sample label data can widen the distribution distance between genuine and fake image samples (such as the + frequency module in the figure); while after triple partitioning of the extracted features by combining ViT, the optimization of the sample distribution is more obvious (such as the + vision transformer in the figure); and when using high-frequency artifact images as sample label data for training and combining ViT to triple partition the extracted features, the distinguishability can be further increased to a visible level (such as the + frequency and vision transformer in the figure).
[0121] Based on the above experiments, in the embodiments of the present invention, the backbone part of the U-Net model (i.e., the Encoder part and the Decoder part) is used for feature extraction, and the extracted features are triple-partitioned in combination with ViT to obtain partitioned features, so that the attention area of the neural network model focuses from the entire face to the detailed feature blocks, thereby completing more accurate positioning. Finally, a classification layer is constructed through the softmax function to output the detection result of the face image to be detected, improving the detection accuracy and detection efficiency of the neural network model.
[0122] A face forgery image detection device provided by the present invention will be described below. The face forgery image detection device described below can be correspondingly referred to the face forgery image detection method described above.
[0123] Based on any of the above embodiments, Figure 7 is a schematic structural diagram of a face forgery image detection device provided by the present invention, as Figure 7 shown. Embodiments of the present invention provide a face forgery image detection device, including an acquisition module 701 and a detection module 702, where:
[0124] The acquisition module 701 is used to acquire a face image to be detected; the detection module 702 is used to input the face image to be detected into a trained neural network model to obtain the detection result of the face image to be detected output by the neural network model;
[0125] The neural network model includes a first convolutional module, a second convolutional module, a vision transformation module, and a classification layer;
[0126] The first convolutional module is used to extract the face features of the face image to be detected;
[0127] The second convolutional module is used to extract the artifact features of the face image to be detected;
[0128] The vision transformation module is used to extract the partitioned features of the face image to be detected;
[0129] The classification layer is used to output the detection result of the face image to be detected.
[0130] A facial forgery image detection device provided by the present invention detects a to-be-detected facial image through a trained neural network model. Among them, the neural network model includes a first convolutional module, a second convolutional module, a visual transformation module, and a classification layer. When detecting the to-be-detected facial image, the first convolutional module extracts the facial features of the to-be-detected facial image, the second convolutional module extracts the artifact features of the to-be-detected facial image, and the visual transformation module extracts the partition features of the to-be-detected facial image. And the classification layer outputs the detection result of the to-be-detected facial image, so that the attention area of the neural network model changes from the entire face to the detailed artifacts, thereby promoting the neural network model to extract effective tampering features, improving the detection accuracy of the neural network model, and achieving accurate and efficient detection of facial forgery images.
[0131] Figure 8 An exemplary schematic physical structure diagram of an electronic device is shown as Figure 8 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the facial forgery image detection method, and the method includes:
[0132] Obtain a to-be-detected facial image;
[0133] Input the to-be-detected facial image into the trained neural network model to obtain the detection result of the to-be-detected facial image output by the neural network model;
[0134] The neural network model includes a first convolutional module, a second convolutional module, a visual transformation module, and a classification layer;
[0135] The first convolutional module is used to extract the facial features of the to-be-detected facial image;
[0136] The second convolutional module is used to extract the artifact features of the to-be-detected facial image;
[0137] The visual transformation module is used to extract the partition features of the to-be-detected facial image;
[0138] The classification layer is used to output the detection result of the to-be-detected facial image.
[0139] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0140] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the face forgery image detection method provided by the above-mentioned various methods. The method includes:
[0141] Obtain a face image to be detected;
[0142] Input the face image to be detected into a trained neural network model to obtain the detection result of the face image to be detected output by the neural network model;
[0143] The neural network model includes a first convolutional module, a second convolutional module, a visual transformation module, and a classification layer;
[0144] The first convolutional module is used to extract the face features of the face image to be detected;
[0145] The second convolutional module is used to extract the artifact features of the face image to be detected;
[0146] The visual transformation module is used to extract the partition features of the face image to be detected;
[0147] The classification layer is used to output the detection result of the face image to be detected.
[0148] On yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the face forgery image detection method provided by the above-mentioned various methods. The method includes:
[0149] Obtain a face image to be detected;
[0150] Input the facial image to be measured into a trained neural network model to obtain the detection result of the facial image to be measured output by the neural network model;
[0151] The neural network model includes a first convolutional module, a second convolutional module, a vision transformation module, and a classification layer;
[0152] The first convolutional module is used to extract the facial features of the facial image to be measured;
[0153] The second convolutional module is used to extract the artifact features of the facial image to be measured;
[0154] The vision transformation module is used to extract the partition features of the facial image to be measured;
[0155] The classification layer is used to output the detection result of the facial image to be measured.
[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0157] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0158] It should be noted that in this text, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes such element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0159] "Determining B based on A" in the embodiments of the present application means that the factor A should be considered when determining B. It is not limited to "determining B only based on A", and should also include: "determining B based on A and C", "determining B based on A, C, and E", "determining C based on A and further determining B based on C", etc. Additionally, it may also include using A as a condition for determining B. For example, "when A meets the first condition, use the first method to determine B"; for another example, "when A meets the second condition, determine B"; for yet another example, "when A meets the third condition, determine B based on the first parameter", etc. Of course, it may also be using A as a condition for the factor of determining B. For example, "when A meets the first condition, use the first method to determine C and further determine B based on C", etc.
[0160] The term "plurality" in the present invention means two or more, and other quantifiers are similar thereto.
[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting forged facial images, characterized in that, Including: Obtain a face image to be measured; Input the face image to be measured into a trained neural network model, and obtain a detection result of the face image to be measured output by the neural network model; The neural network model includes a first convolution module, a second convolution module, a visual transformation module, and a classification layer; The first convolution module is used to extract facial features of the face image to be measured; The second convolution module is used to extract artifact features of the face image to be measured; The visual transformation module is used to extract partition features of the face image to be measured; The classification layer is used to output a detection result of the face image to be measured.
2. The facial forgery image detection method according to claim 1, wherein Before the step of inputting the face image to be measured into a trained neural network model and obtaining a detection result of the face image to be measured output by the neural network model, it further includes: Obtain a real face image and a forged face image; Generate a high-frequency artifact image based on the real face image and the forged face image; Use the real face image and the forged face image as sample data, and use the high-frequency artifact image as sample label data corresponding to the sample data to train the neural network model.
3. The facial forgery image detection method according to claim 2, characterized in that, The step of generating a high-frequency artifact image based on the real face image and the forged face image includes: Perform high-pass filtering on the forged face image to obtain a high-frequency image; and calculate the absolute value of the difference between the real face image and the forged face image to obtain an artifact image; Generate a high-frequency artifact image based on the overlapping part of the high-frequency image and the artifact image.
4. The facial forgery image detection method according to claim 1, characterized in that, The step of inputting the face image to be measured into a trained neural network model and obtaining a detection result of the face image to be measured output by the neural network model includes: Input the face image to be measured into the first convolution module, and obtain facial features of the face image to be measured output by the first convolution module; Input the facial features into the second convolution module, and obtain artifact features of the face image to be measured output by the second convolution module; Based on the facial features and the artifact features, obtain partition features of the face image to be measured through the visual transformation module; Input the partition features into the classification layer, and obtain a detection result of the face image to be measured output by the classification layer.
5. The facial forgery image detection method according to claim 4, wherein, The step of obtaining partition features of the face image to be measured through the visual transformation module based on the facial features and the artifact features includes: Determine the weight of the artifact features based on the artifact features; Perform linear weighting on the feature vectors in the facial features based on the weight of the artifact features to obtain weighted features; According to the facial region where each weighted feature is located, partition the weighted features through the visual transformation module to obtain partition features output by the visual transformation module.
6. The face forgery image detection method according to claim 5, wherein, The partition features include eye features, facial center features, and facial contour features.
7. A facial forgery image detection device, characterized in that, Including: An acquisition module, configured to obtain a face image to be measured; A detection module, configured to input the face image to be measured into a trained neural network model, and obtain a detection result of the face image to be measured output by the neural network model; The neural network model includes a first convolutional module, a second convolutional module, a visual transformation module, and a classification layer; The first convolutional module is used to extract the facial features of the to-be-detected facial image; The second convolutional module is used to extract the artifact features of the to-be-detected facial image; The visual transformation module is used to extract the partition features of the to-be-detected facial image; The classification layer is used to output the detection result of the to-be-detected facial image.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the facial forgery image detection method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the facial forgery image detection method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the facial forgery image detection method according to any one of claims 1 to 6.