Counterfeit face identification method and system

By building a multi-feature fusion framework, using lightweight CNN and multiple deep learning models for feature extraction and fusion, the problems of insufficient generalization capabilities and susceptibility to attacks in the existing technology are solved, and fake face detection with high accuracy, robustness and generalization capabilities are achieved.

CN120220256AActive Publication Date: 2025-06-27EAST CHINA JIAOTONG UNIVERSITY
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202510677496.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-27
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

The existing forged face detection methods are insufficient in generalization when dealing with new forged methods, cross-data set migration and low-resolution images, are vulnerable to adversarial attacks, and are sensitive to image compression, and lacks multi-dimensional feature fusion.

Method used

By constructing a multi-feature fusion framework, obtain the face images and data sets to be identified, the features are extracted and clustered using lightweight CNN, and then the data set is input into multiple deep learning models for joint training of feature extractors. The features are mapped to the shared latent feature space using a variational autoencoder, fuse the features extracted by different models, and distributed through KL divergence constraints.

Benefits of technology

The accuracy, robustness and generalization capabilities of forged face detection are improved, and the resistance to different image quality and new forgery methods is enhanced, and the tight and robust fusion of multi-dimensional features is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220256A_ABST
    Figure CN120220256A_ABST
Patent Text Reader

Abstract

The invention provides a fake face identification method and system. The method comprises the following steps: acquiring a face image to be identified and a real face and fake face image data set which are paired, and clustering and dividing the data set; inputting the divided data set into multiple models to carry out joint training of a feature extractor; fixing the trained feature extractor, and mapping the extracted features to a shared potential feature space by using a variational auto-encoder to realize deep collaborative fusion of different model features; fixing the trained feature extractor and variational auto-encoder, obtaining the features of the input image, and sending the features to a classifier for training; and carrying out counterfeit identification on the face image to be detected by using the trained counterfeit face identification network. Wherein the generalization of the model is effectively improved through clustering division of a data set, and different model features are fused through a variational auto-encoder, so that feature expression is tighter and more robust, and complementarity between different models is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of forged face identification, and particularly relates to a forged face identification method and system. Background Art

[0002] With the rapid development of deep learning technology, image generation technology based on artificial intelligence has been widely applied. Especially driven by generative models such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and Diffusion Models, the synthesis of highly realistic face images has become a reality. This technology has played an important role in fields such as virtual character creation, film and television special effects, and augmented reality (AR), but at the same time has brought serious security risks, such as identity forgery, false information dissemination, and malicious fraud. Therefore, an efficient, accurate, and robust identification method for forged face images has become one of the key technologies in the field of artificial intelligence security.

[0003] Currently, forged face detection mainly includes two major categories of methods: spatial domain analysis-based and frequency domain analysis-based. Spatial domain-based methods rely on deep neural networks to extract local features, such as texture, edge information, and unnatural facial details, and are identified through a classification model. Representative methods include XceptionNet, ResNet, EfficientNet based on CNN, as well as ViT and Swin Transformer based on the Transformer architecture. Although these methods perform well on standard datasets, their generalization ability still has limitations when dealing with new forgery methods, cross-dataset migration, and low-resolution images. Frequency domain-based methods use the abnormal features of GAN-generated images in the frequency domain distribution for detection. For example, Fourier transform analysis can reveal high-frequency abnormalities in forged images, while wavelet transform can extract features of forged images at different frequencies through multi-scale analysis. Although these methods make up for the deficiencies of spatial domain methods to a certain extent, their detection performance will still decline when faced with image compression, degradation, or noise pollution. Existing technologies generally have problems such as insufficient generalization ability, susceptibility to adversarial attacks, sensitivity to image compression, and lack of multi-dimensional feature fusion. Summary of the Invention

[0004] Based on this, embodiments of the present invention provide a forged face identification method and system, aiming to improve the accuracy, robustness, and generalization ability of detection by constructing a multi-feature fusion framework.

[0005] The first aspect of the embodiments of the present invention provides a forged face identification method, and the method includes: Obtain the face image to be authenticated, as well as a dataset of paired real and forged face images. Use a lightweight CNN to extract features, perform clustering via K-nearest neighbors in the feature space, and divide the data into multiple target datasets; Input the target datasets into a deep learning model for joint training of the feature extractor. The deep learning model includes at least Inception, ResNet50, ViT, and EfficientNet; Determine the trained feature extractor, and use a variational autoencoder to map the extracted features to a shared latent feature space. Fuse the features extracted by different models, and constrain the distribution via KL divergence. Then, reconstruct the original features from the latent space through the decoder to ensure that key information is not lost during the encoding process; Use the trained feature extractor and variational autoencoder to obtain features, and input the features into a classifier for fine-tuning to complete the training of the forged face authentication network; Input the face image to be authenticated into the trained forged face authentication network and output the authentication result.

[0006] Furthermore, in the step of inputting the target datasets into the deep learning model for joint training of the feature extractor, randomly input the target datasets into the deep learning model; after feature extraction, perform contrastive learning. Then, each deep learning model inputs the features into the classification head to output results, calculates the loss according to the loss function, and performs iterative training so that each deep learning model can serve as a unique feature extractor, and use the corresponding test set for performance evaluation.

[0007] Furthermore, in the step of calculating the loss according to the loss function and performing iterative training, the calculation formula of the total loss function is: ; where L is the total loss function, L NCE is the contrastive learning loss function, and L CE is the cross-entropy loss function; The calculation formula of the contrastive learning loss function is: ; where represents the features extracted from the current positive class sample, represents the features obtained by extracting positive class samples by different deep learning models, represents the features extracted by different deep learning models, T represents the temperature coefficient, which is used to control the steepness of the distribution, represents and similarity; The calculation formula of the cross-entropy loss function is: ; Among them, N represents the number of samples, M represents the number of models, represents the weight of the j th deep learning model, represents the correct label of the i th sample, represents the prediction result made by the i th sample for the j th deep learning model.

[0008] Furthermore, in the step of using the corresponding test set for performance evaluation, three metrics, Precision, Recall, and F1-score, are used for evaluation. Specifically, ; Among them, TP represents the number of correctly predicted samples, FP represents the number of incorrectly predicted samples, Precision represents the precision rate, Recall represents the recall rate, and F1-score represents the harmonic mean of Precision and Recall.

[0009] Furthermore, in the step of determining the trained feature extractor, mapping the extracted features to a shared latent feature space using a variational autoencoder, fusing the features extracted by different models, and constraining the distribution through KL divergence, and then reconstructing the original features from the latent space through a decoder, fix the trained feature extractor, extract the features of the dataset, and simply concatenate them to obtain the feature , is the number of input channels, H and W are the height and width of the input respectively. Input the feature into the variational autoencoder, and map it to the shared latent space through convolution to obtain the feature and the feature . Specifically, Input the feature into the first convolutional block, and through a non-linear activation function and batch normalization, obtain . The calculation formula is: ; Among them, represents a 4×4 convolution operation with a stride of 2 and a padding of 1, represents a batch normalization operation, represents a non-linear activation function; Input into the second convolutional block, and through a non-linear activation function and batch normalization, obtain . The calculation formula is: ; Among them, Represents a 4×4 convolution operation with a stride of 2 and a padding of 1; Input into two convolutional layers respectively, and obtain features and feature respectively. The calculation formula is: ; ; Among them, and both represent 3×3 convolution operations with a stride of 1 and a padding of 1; Subsequently, the KL divergence loss is used to constrain the feature distributions of feature and feature ; Then, the reparameterization trick is used to obtain the latent representation z, that is, the fused feature z. Then, the decoder is used to reconstruct the features of z. The calculation formula of the latent representation z is: ; Among them, is sampled as a random variable, and sampling is performed through the reparameterization trick to ensure that the gradient of the latent representation z can be transmitted, so that feature and feature can still be optimized through backpropagation.

[0010] Furthermore, in the step of constraining the feature distributions of feature and feature by the KL divergence loss, KL divergence is used for regularization, and the calculation formula is: ; Among them, D KL represents the KL divergence loss, d represents the dimension of the latent space, σ i and μ i respectively represent the features of the i-th sample.

[0011] Furthermore, in the step of reconstructing the features of z by the decoder, the latent representation z is input into the first transposed convolution block, and through the non-linear activation function and batch normalization, R1 is obtained. The calculation formula is: ; Among them, represents a 4×4 transposed convolution operation with a stride of 2 and a padding of 1; Then, R1 is input into the second transposed convolution block, and through the non-linear activation function and batch normalization, R2 is obtained. The calculation formula is: ; Among them, represents a 4×4 transposed convolution operation with a stride of 2 and a padding of 1; Finally, R2 is input into the last transposed convolution layer to obtain the reconstructed feature , and the calculation formula is: ; Among them, represents a 4×4 transposed convolution operation with a stride of 2 and a padding of 1.

[0012] Furthermore, the goal of the decoder is to generate the reconstructed feature from the latent representation z, which is as close as possible to the features of the original input. The mean square error is used to measure the quality of the decoder's reconstructed data, and the calculation formula is: ; Among them, represents the reconstruction loss, N represents the total number of all features, X represents the original feature, represents the reconstructed feature.

[0013] Furthermore, in the step of inputting the feature into the classifier for fine-tuning, the feature after the variational autoencoder is fused, that is, the latent representation z is gradually input into the convolutional block to obtain z1, and the calculation formula is: ; Among them, represents a 3×3 transposed convolution operation with a stride of 1 and a padding of 1; Then z1 is subjected to global average pooling, and input into the fully connected layer for dimensionality reduction to obtain z1, and the calculation formula is: ; Among them, represents the global average pooling operation, represents the dimensionality reduction operation of the fully connected layer; Finally, after normalization and the Sigmoid activation function, the predicted probability is obtained, and the calculation formula is: ; Among them, Sigmoid represents the activation function that maps the output to between [0,1]; The loss function used by the classifier is the cross-entropy loss, and the calculation formula is: ; Among them, represents the cross-entropy loss, N represents the total number of all samples, represents the true label, Indicates the classification result of the classifier.

[0014] The second aspect of the embodiments of the present invention provides a forged face authentication system for implementing the forged face authentication method described in the first aspect. The system includes: An acquisition module, configured to acquire a face image to be authenticated, as well as a dataset of paired real faces and forged face images, extract features using a lightweight CNN, perform clustering in the feature space through K-nearest neighbors, and divide the data into multiple target datasets; A training module, configured to input the target dataset into a deep learning model for joint training of a feature extractor. The deep learning model at least includes Inception, ResNet50, ViT, and EfficientNet; A fusion module, configured to determine a trained feature extractor, map the extracted features to a shared latent feature space using a variational autoencoder, fuse the features extracted by different models, and constrain the distribution through KL divergence, and then reconstruct the original features from the latent space through a decoder to ensure that key information is not lost during the encoding process; A fine-tuning module, configured to obtain features using the trained feature extractor and variational autoencoder, and input the features into a classifier for fine-tuning to complete the training of the forged face authentication network; An input module, configured to input the face image to be authenticated into the trained forged face authentication network and output an authentication result.

[0015] A forged face authentication method and system provided in the embodiments of the present invention, by acquiring a face image to be authenticated and a dataset of paired real faces and forged face images and clustering and dividing the dataset; inputting the divided dataset into multiple models for joint training of a feature extractor; fixing the trained feature extractor, and using a variational autoencoder to map the extracted features to a shared latent feature space to achieve deep collaborative fusion of features of different models; fixing the trained feature extractor and variational autoencoder, obtaining the features of the input image and feeding them into a classifier for training; using the trained forged face authentication network to perform forged authentication on the face image to be detected. Among them, the generalization of the model is effectively improved by clustering and dividing the dataset, and the features of different models are fused through a variational autoencoder, making the feature expression more compact and robust, thereby promoting the complementarity between different models. Description of the Drawings

[0016] Figure 1 Is a flowchart for implementing a forged face authentication method provided in Embodiment 1 of the present invention; Figure 2 Is a schematic diagram of the process of dividing the dataset and extracting features by multiple models; Figure 3Schematic diagram of the classifier training process based on multi-model feature extraction and variational autoencoder feature fusion; Figure 4 Schematic diagram of the structure for feature fusion based on variational autoencoder; Figure 5 Block diagram of a forged face identification system provided in the second embodiment of the present invention; Figure 6 Block diagram of an electronic device provided in the third embodiment of the present invention. Detailed implementation manners

[0017] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0018] It should be noted that when an element is referred to as being "fixed on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are only for the purpose of illustration.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0020] Embodiment 1 According to an embodiment of the present invention, an embodiment of a forged face identification method is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0021] In the first embodiment, a forged face identification method is provided, which can be used in an electronic device, such as a computer. Please refer to Figure 1 , Figure 1 shows an implementation flowchart of a forged face identification method provided in the first embodiment of the present invention, which specifically includes steps S01 to S05.

[0022] Step S01: Obtain the face image to be authenticated, as well as a dataset of paired real and forged face images. Then, use a lightweight CNN to extract features and perform clustering in the feature space through K-nearest neighbors to divide the data into multiple target datasets.

[0023] Specifically, use the publicly available datasets DeepFake and FaceForensics++. Use the lightweight MobileNet to perform basic feature extraction on the datasets, and then use the K-nearest neighbor clustering method to cluster the extracted face features into N categories. Select one category of data as the validation set, and the remaining N - 1 categories of data as the training set to ensure the generalization ability of the model under different data distributions. Subsequently, follow the same strategy, that is, in each division, select one category of data as the validation set and the rest as the training set, and finally form N independent training-validation dataset combinations. Uniformly crop the face image to be authenticated and the dataset size to 224×224 to ensure input consistency.

[0024] Step S02: Input the target dataset into a deep learning model for joint training of the feature extractor. The deep learning model includes at least Inception, ResNet50, ViT, and EfficientNet.

[0025] Please refer to Figure 2 and Figure 3 , Figure 2 which are the schematic diagrams of the process of dataset division and feature extraction by multiple models. Figure 3 Figure [X] is the schematic diagram of the classifier training process based on multi-model feature extraction and variational autoencoder feature fusion. In the embodiments of the present invention, the divided dataset is randomly input into deep learning models such as Inception, ResNet50, ViT, and EfficientNet. After feature extraction, contrastive learning is performed. In order to prevent different models from learning similar features and to prompt them to extract more diverse and complementary information, this embodiment uses the InfoNCE (Information Noise-Contrastive Estimation) loss function to optimize the model, and its formula is: ; where, represents the features extracted from the current positive sample, represents the features obtained by different deep learning models extracting positive samples, represents the features obtained by different deep learning models, T represents the temperature coefficient used to control the steepness of the distribution, represents and Similarity. This formula emphasizes that the model should learn important features relevant to the target task, enabling samples of the same category (such as real / fake) to remain relatively consistent in the latent space. At the same time, by imposing regularization constraints, it avoids the features extracted by different models from being too similar, thereby enhancing the specificity of the model and enabling each model to have its own feature representation.

[0026] Then, the features of each model are input into the classification head to output the results, and the cross-entropy loss function is used. Its formula is: ; where N represents the number of samples, M represents the number of models, represents the weight of the j th deep learning model, represents the correct label of the i th sample, represents the i th sample and the j th prediction result made by the

[0027] The calculation formula of the total loss function is: ; where L is the total loss function, L NCE is the contrastive learning loss function, and L CE is the cross-entropy loss function. By optimizing the loss, the feature redundancy of different models is reduced, allowing each model to learn information from different perspectives.

[0028] In addition, the performance of the feature extractor classification results is evaluated using the corresponding test set. Among them, three indicators, Precision, Recall, and F1-score, are used for evaluation. Specifically, ; where TP represents the number of correctly predicted samples, and FP represents the number of incorrectly predicted samples. Precision represents the precision rate, which reflects the proportion of samples that truly belong to the positive class among all samples predicted as the positive class (AI-generated); Recall represents the recall rate, which reflects the proportion of all truly positive class samples that are correctly predicted as the positive class; F1-score represents the harmonic mean of Precision and Recall.

[0029] Step S03: Determine the trained feature extractor, and use the variational autoencoder to map the extracted features to a shared latent feature space, fuse the features extracted by different models, and constrain the distribution through the KL divergence. Then, reconstruct the original features from the latent space through the decoder to ensure that key information is not lost during the encoding process.

[0030] Please refer toFigure 4 is a schematic diagram of the structure for feature fusion based on a variational autoencoder. After step S02, the feature extractor has been trained. Then we fix the feature extractor network, and after simply concatenating the extracted features, we get the feature , is the number of input channels, H and W are the height and width of the input respectively. The feature is input into the variational autoencoder and mapped to the shared latent space through convolution to obtain the feature and the feature . Specifically, the feature is input into the first convolutional block, and through a non-linear activation function and batch normalization, we get . The calculation formula is: ; where represents a 4×4 convolution operation with a stride of 2 and a padding of 1, represents a batch normalization operation, represents a non-linear activation function; The is input into the second convolutional block, and through a non-linear activation function and batch normalization, we get . The calculation formula is: ; where represents a 4×4 convolution operation with a stride of 2 and a padding of 1; The are respectively input into two convolutional layers to respectively obtain the features and the feature . The calculation formula is: ; ; where and both represent 3×3 convolution operations with a stride of 1 and a padding of 1; Subsequently, the KL divergence loss is used to constrain the feature distributions of the features and the feature ; Then the reparameterization trick is used to obtain the latent representation z, that is, the fused feature z. Then the decoder is used to reconstruct the features of z. The calculation formula for the latent representation z is: ; where is sampled as a random variable, and sampling is performed through the reparameterization trick to ensure that the gradient of the latent representation z can be propagated, so that the features and the feature Can still be optimized by backpropagation.

[0031] Constraining features through KL divergence loss and features In the step of the feature distribution, KL divergence is used for regularization, and the calculation formula is: ; Among them, D KL represents the KL divergence loss, d represents the dimension of the latent space, σ i and μ i respectively represent the features of the i-th sample. It can be understood that the features and features are respectively obtained by the variational autoencoder. By constraining the feature distribution in the latent space with KL divergence, it is made to approach the standard normal distribution, improving the smoothness of the latent space, thereby enhancing the robustness of the encoder to noise and perturbations.

[0032] In the step of using the decoder to reconstruct features from z, the latent representation z is input into the first transposed convolutional block, and through the non-linear activation function and batch normalization, R1 is obtained, and the calculation formula is: ; Among them, represents a 4×4 transposed convolutional operation, with a stride of 2 and a padding of 1; Then R1 is input into the second transposed convolutional block, and through the non-linear activation function and batch normalization, R2 is obtained, and the calculation formula is: ; Among them, represents a 4×4 transposed convolutional operation, with a stride of 2 and a padding of 1; Finally, R2 is input into the last transposed convolutional layer to obtain the reconstructed feature , and the calculation formula is: ; Among them, represents a 4×4 transposed convolutional operation, with a stride of 2 and a padding of 1.

[0033] At the same time, in order to ensure that key information is not lost during the encoding process, the reconstructed feature generated by the decoder through the latent representation z is made as close as possible to the features of the original input. The mean square error is used to measure the quality of the decoder's reconstructed data, and the calculation formula is: ; Among them, represents the reconstruction loss, N represents the total number of all features, XRepresents the original feature, Represents the reconstructed feature. Backpropagation is performed by calculating the mean squared error while iterating the encoder and decoder to ensure that key information is not lost during the encoding process.

[0034] In step S04, the trained feature extractor and variational autoencoder are used to obtain features, and the features are input into the classifier for fine-tuning to complete the training of the forged face discrimination network.

[0035] Fix the parameters of the feature extraction network and variational autoencoder, and input the extracted and fused features into the final classifier for final training. It should be noted that the features fused by the variational autoencoder, that is, the latent representation z, are gradually input into the convolutional block to obtain z1. The calculation formula is: ; Among them, Represents a 3×3 transposed convolutional operation with a stride of 1 and a padding of 1; To reduce the computational amount and retain key information at the same time, global average pooling is performed on z1, and it is input into the fully connected layer for dimensionality reduction to obtain z1. The calculation formula is: ; Among them, Represents the global average pooling operation, Represents the dimensionality reduction operation of the fully connected layer; Finally, after normalization and the Sigmoid activation function, the predicted probability is obtained. The calculation formula is: ; Among them, Sigmoid represents the activation function that maps the output to between [0,1]; The loss function used by the classifier is the cross-entropy loss. The calculation formula is: ; Among them, Represents the cross-entropy loss, N represents the number of all samples, Represents the true label, Represents the classification result of the classifier.

[0036] To evaluate the model performance, all datasets are randomly input into the trained model to obtain the corresponding predicted labels. The AUC-ROC is used as the evaluation metric. The calculation formula of AUC is as follows: ; Among them, TP represents the number of correctly predicted samples, FP represents the number of wrongly predicted samples. The value of AUC represents the area under the ROC curve, ranging from 0 to 1. When the value of AUC is closer to 1, it indicates better performance of the classifier.

[0037] Step S05: Input the face image to be identified into the trained forged face identification network, and output the identification result.

[0038] Specifically, the training of the forged face identification network is completed, and then the face image to be identified is input into the trained multi-model feature fusion network to output the probability (between 0 and 1) of the image to be identified. The category of the image to be identified is a real image (0) or a forged face image (1). The closer the probability is to 1, the greater the possibility of it being a forged face image.

[0039] In summary, for the forged face identification method in the above embodiments of the present invention, the method obtains the face image to be identified and a paired dataset of real faces and forged face images and clusters and partitions the dataset; inputs the partitioned dataset into multiple models for joint training of the feature extractor; fixes the trained feature extractor, and uses the variational autoencoder to map the extracted features to a shared latent feature space to achieve deep collaborative fusion of different model features; fixes the trained feature extractor and variational autoencoder, obtains the features of the input image and sends them to the classifier for training; uses the trained forged face identification network to perform forged identification on the face image to be detected. Among them, the generalization of the model is effectively improved by clustering and partitioning the dataset, and the feature expressions are made more compact and robust by fusing different model features through the variational autoencoder, thereby promoting the complementarity between different models.

[0040] Embodiment 2 Please refer to Figure 5 , Figure 5 which is a structural block diagram of a forged face identification system provided in Embodiment 2 of the present invention. The forged face identification system 200 is used to implement the above embodiments and preferred implementation manners, and those already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0041] Specifically, the forged face identification system 200 includes: an acquisition module 21, a training module 22, a fusion module 23, a fine-tuning module 24, and an input module 25, where: The acquisition module 21 is used to acquire the face image to be identified, as well as a dataset of paired real faces and forged face images, extract features using a lightweight CNN, perform clustering through K-nearest neighbors in the feature space, and divide the data into multiple target datasets; The training module 22 is used to input the target data set into a deep learning model for joint training of the feature extractor. The deep learning model at least includes Inception, ResNet50, ViT, and EfficientNet. Specifically, the target data set is randomly input into the deep learning model. After feature extraction, contrast learning is performed. Then, each deep learning model inputs the features into the classification head to output results, calculates the loss according to the loss function, and performs iterative training so that each deep learning model can be used as a unique feature extractor and is evaluated using the corresponding test set. In the step of calculating the loss according to the loss function and performing iterative training, the calculation formula of the total loss function is: ; where L is the total loss function, L NCE is the contrast learning loss function, and L CE is the cross-entropy loss function; The calculation formula of the contrast learning loss function is: ; where represents the features extracted from the current positive class sample, represents the features obtained by extracting positive class samples by different deep learning models, represents the features extracted by different deep learning models, and T represents the temperature coefficient used to control the steepness of the distribution, represents and similarity; The calculation formula of the cross-entropy loss function is: ; where N represents the number of samples, M represents the number of models, represents the weight of the j th deep learning model, represents the correct label of the i th sample, represents the prediction result of the i th sample by the j th deep learning model. In addition, in the step of evaluating the performance using the corresponding test set, three indicators, Precision, Recall, and F1-score, are used for evaluation. Specifically, ; where TP represents the number of correctly predicted samples, FP represents the number of incorrectly predicted samples, Precision represents the precision rate, Recall represents the recall rate, and F1-score represents the harmonic mean of Precision and Recall; The fusion module 23 is used to determine the trained feature extractor, map the extracted features to a shared latent feature space using a variational autoencoder, fuse the features extracted by different models, and constrain the distribution through the KL divergence. Then, the original features are reconstructed from the latent space through the decoder to ensure that key information is not lost during the encoding process. Specifically, the trained feature extractor is fixed, the features of the dataset are extracted, and the features are obtained after simple concatenation. , is the number of input channels, H and W are the height and width of the input respectively. The features are input into the variational autoencoder and mapped to the shared latent space through convolution to obtain the features and the features . Specifically, The features are input into the first convolutional block, and through the non-linear activation function and batch normalization, is obtained. The calculation formula is: ; Among them, represents a 4×4 convolution operation with a stride of 2 and a padding of 1, represents the batch normalization operation, represents the non-linear activation function; The are input into the second convolutional block, and through the non-linear activation function and batch normalization, is obtained. The calculation formula is: ; Among them, represents a 4×4 convolution operation with a stride of 2 and a padding of 1; The are respectively input into two convolutional layers, and the features and the features are respectively obtained. The calculation formula is: ; ; Among them, and both represent 3×3 convolution operations with a stride of 1 and a padding of 1; Subsequently, the feature distributions of the features and the features are constrained through the KL divergence loss; Then, the reparameterization trick is used to obtain the latent representation z, that is, the fused feature z. Then, the decoder is used to reconstruct the features of z. The calculation formula of the latent representation z is: ; Among them, Sampled as a random variable, sampling is performed through the reparameterization trick to ensure that the gradient of the latent representation z can be passed, so that the features and features can still be optimized through backpropagation. In the step of constraining the feature distributions of features and features by KL divergence loss, KL divergence is used for regularization, and the calculation formula is: ; where D KL represents the KL divergence loss, d represents the dimension of the latent space, σ i and μ i respectively represent the features of the i-th sample. In the step of reconstructing features from z using the decoder, the latent representation z is input into the first transposed convolutional block, and through a non-linear activation function and batch normalization, R1 is obtained, and the calculation formula is: ; where represents a 4×4 transposed convolutional operation with a stride of 2 and a padding of 1; Then R1 is input into the second transposed convolutional block, and through a non-linear activation function and batch normalization, R2 is obtained, and the calculation formula is: ; where represents a 4×4 transposed convolutional operation with a stride of 2 and a padding of 1; Finally, R2 is input into the last transposed convolutional layer to obtain the reconstructed feature , and the calculation formula is: ; where represents a 4×4 transposed convolutional operation with a stride of 2 and a padding of 1. The goal of the decoder is to make the reconstructed feature generated from the latent representation z as close as possible to the features of the original input. The mean squared error is used to measure the quality of the data reconstructed by the decoder, and the calculation formula is: ; where represents the reconstruction loss, N represents the total number of all features, X represents the original feature, represents the reconstructed feature; The fine-tuning module 24 is used to obtain features by using the trained feature extractor and variational autoencoder, and input the features into the classifier for fine-tuning to complete the training of the forged face discrimination network. Specifically, the features after the fusion of the variational autoencoder, that is, the latent representation z, are gradually input into the convolutional block to obtain z1. The calculation formula is: ; Among them, represents a 3×3 transposed convolutional operation with a stride of 1 and a padding of 1; Then, global average pooling is performed on z1, and it is input into the fully connected layer for dimensionality reduction to obtain z1. The calculation formula is: ; Among them, represents the global average pooling operation, represents the dimensionality reduction operation of the fully connected layer; Finally, after normalization and the Sigmoid activation function, the predicted probability is obtained. The calculation formula is: ; Among them, Sigmoid represents the activation function that maps the output to the range [0, 1]; The loss function used by the classifier is the cross-entropy loss. The calculation formula is: ; Among them, represents the cross-entropy loss, N represents the number of all samples, represents the true label, represents the classification result of the classifier; The input module 25 is used to input the to-be-discriminated face image into the trained forged face discrimination network and output the discrimination result.

[0042] Embodiment 3 On the other hand, the present invention also proposes an electronic device. Please refer to Figure 6 , which shows the electronic device in Embodiment 3 of the present invention, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor. When the processor 10 executes the computer program 30, the above-mentioned forged face discrimination method is implemented.

[0043] Among them, in some embodiments, the processor 10 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips, and is used to run the program code stored in the memory 20 or process data, such as executing an access restriction program, etc.

[0044] Among them, the memory 20 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 20 can be an internal storage unit of the electronic device in some embodiments, such as the hard disk of the electronic device. The memory 20 can also be an external storage device of the electronic device in other embodiments, such as a plug-in hard disk equipped on the electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc. Further, the memory 20 can also include both the internal storage unit and the external storage device of the electronic device. The memory 20 can be used not only to store application software and various types of data of the electronic device, but also to temporarily store the data that has been output or will be output.

[0045] It should be noted that Figure 6 The structure shown does not constitute a limitation on the electronic device. In other embodiments, the electronic device may include fewer or more components than shown in the figure, or combine certain components, or have a different component arrangement.

[0046] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the forgery face authentication method as described above is implemented.

[0047] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0048] More specific examples (a non-exhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer disk cartridges (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0049] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.

[0050] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0051] The above embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

Claims

1. A method for identifying forged faces, characterized in that, The method includes: Obtain the face image to be authenticated, as well as a dataset of paired real and forged face images, and use a lightweight CNN to extract features. Perform clustering in the feature space through K-nearest neighbors to divide the data into multiple target datasets; Input the target datasets into a deep learning model for joint training of the feature extractor. The deep learning model includes at least Inception, ResNet50, ViT, and EfficientNet; Determine the trained feature extractor, and use a variational autoencoder to map the extracted features to a shared latent feature space, fuse the features extracted by different models, and constrain the distribution through KL divergence. Then reconstruct the original features from the latent space through the decoder to ensure that key information is not lost during the encoding process; Use the trained feature extractor and variational autoencoder to obtain features, and input the features into a classifier for fine-tuning to complete the training of the forged face authentication network; Input the face image to be authenticated into the trained forged face authentication network and output the authentication result.

2. The forgery face identification method according to claim 1, characterized in that In the step of inputting the target datasets into the deep learning model for joint training of the feature extractor, randomly input the target datasets into the deep learning model; after feature extraction, perform contrastive learning, and then each deep learning model inputs the features into the classification head to output results. Calculate the loss according to the loss function and perform iterative training so that each deep learning model can serve as a unique feature extractor, and use the corresponding test set for performance evaluation.

3. The forgery face identification method according to claim 2, wherein In the step of calculating the loss according to the loss function and performing iterative training, the calculation formula of the total loss function is: ; Among them, L is the total loss function, and L NCE is the contrastive learning loss function, and L CE is the cross-entropy loss function; The calculation formula of the contrastive learning loss function is: ; Among them, represents the features extracted from the current positive class samples, represents the features obtained by different deep learning models from positive class samples, represents the features obtained by different deep learning models. T represents the temperature coefficient, which is used to control the steepness of the distribution, represents and similarity; The calculation formula of the cross-entropy loss function is: ; Among them, N represents the number of samples, and M represents the number of models. represents the weight of the j th deep learning model, represents the correct label of the i th sample, represents the prediction result of the i th sample made by the j th deep learning model.

4. The forgery face authentication method according to claim 3, characterized in that, In the step of using the corresponding test set for performance evaluation, three metrics, Precision, Recall, and F1-score, are used for evaluation. Specifically, ; where TP represents the number of correctly predicted samples, FP represents the number of incorrectly predicted samples, Precision represents the precision rate, Recall represents the recall rate, and F1-score represents the harmonic mean of Precision and Recall.

5. The forgery face identification method according to claim 4, wherein In the step of determining the trained feature extractor, using the variational autoencoder to map the extracted features to the shared latent feature space, fusing the features extracted by different models, constraining the distribution through the KL divergence, and then reconstructing the original features from the latent space through the decoder, fix the trained feature extractor, extract the features of the dataset, and simply concatenate them to obtain the features , is the number of input channels, H and W are the height and width of the input respectively. Input the features into the variational autoencoder, and map them to the shared latent space through convolution to obtain the features and the features , specifically, Input the feature into the first convolutional block, and through the non-linear activation function and batch normalization, obtain , and the calculation formula is: ; Among them, represents a 4×4 convolution operation with a stride of 2 and a padding of 1, represents a batch normalization operation, represents a non-linear activation function; Input into the second convolutional block, and through a non-linear activation function and batch normalization, obtain , and the calculation formula is: ; Among them, represents a 4×4 convolution operation with a stride of 2 and a padding of 1; Input into two convolutional layers respectively to obtain features and feature , and the calculation formula is as follows: ; ; Among them, and both represent a 3×3 convolution operation with a stride of 1 and a padding of 1; Subsequently, the KL divergence loss is used to constrain the features and the features feature distribution; Then use the reparameterization trick to obtain the latent representation z, that is, the fused feature z, and then use the decoder to reconstruct the features of z. The calculation formula of the latent representation z is: ; Among them, is sampled as a random variable, and sampling is performed through the reparameterization trick to ensure that the gradient of the latent representation z can be passed, so that the feature and the feature can still be optimized through backpropagation.

6. The forgery face authentication method according to claim 5, wherein The step of constraining features through KL divergence loss and features In the step of the feature distribution, KL divergence is used for regularization, and the calculation formula is as follows: ; Among them, D KL represents the KL divergence loss, d represents the dimension of the latent space, σ i and μ i represent the features of the i-th sample respectively.

7. The forgery face authentication method according to claim 6, wherein In the step of using the decoder to reconstruct the features of z, input the latent representation z into the first transposed convolutional block, and through a non-linear activation function and batch normalization, obtain R1. The calculation formula is: ; Among them, represents a 4×4 transposed convolution operation with a stride of 2 and a padding of 1; Then input R1 into the second transposed convolutional block, and through a non-linear activation function and batch normalization, obtain R2. The calculation formula is: ; Among them, represents a 4×4 transposed convolution operation with a stride of 2 and a padding of 1; Finally, input R2 into the last transposed convolutional layer to obtain the reconstructed features , and the calculation formula is: ; Among them, represents a 4×4 transposed convolution operation with a stride of 2 and a padding of 1.

8. The forgery face identification method according to claim 7, wherein The goal of the decoder is to generate reconstructed features from the latent representation z , which are as close as possible to the features of the original input. The mean squared error is used to measure the quality of the decoder's reconstructed data, and the calculation formula is as follows: ; wherein, represents the reconstruction loss, N represents the total number of all features, X represents the original feature, represents the reconstructed feature.

9. The forgery face identification method according to claim 8, wherein In the step of inputting the features into the classifier for fine-tuning, gradually input the fused features of the variational autoencoder, that is, the latent representation z, into the convolutional block to obtain z1. The calculation formula is: ; Among them, represents a 3×3 transposed convolution operation with a stride of 1 and a padding of 1; Then perform global average pooling on z1 and input it into the fully connected layer for dimensionality reduction to obtain z1. The calculation formula is: ; Among them, represents a global average pooling operation, represents a dimensionality reduction operation of a fully connected layer; Finally, after normalization and passing through the Sigmoid activation function, the predicted probability is obtained. , and the calculation formula is: ; where Sigmoid represents the activation function that maps the output to between [0,1]; The loss function used by the classifier is the cross-entropy loss, and the calculation formula is as follows: ; Among them, represents the cross-entropy loss, N represents the number of all samples, represents the true label, represents the classification result of the classifier.

10. A forged face authentication system, characterized in that, To implement the forged face identification method according to any one of claims 1-9, the system includes: An acquisition module, configured to acquire the face image to be identified, as well as a dataset of paired real face and forged face images, extract features using a lightweight CNN, perform clustering through K-nearest neighbors in the feature space, and divide the data into multiple target datasets; A training module, configured to input the target dataset into a deep learning model for joint training of the feature extractor, and the deep learning model includes at least Inception, ResNet50, ViT, and EfficientNet; A fusion module, configured to determine the trained feature extractor, map the extracted features to a shared latent feature space using a variational autoencoder, fuse the features extracted by different models, and constrain the distribution through KL divergence, and then reconstruct the original features from the latent space through a decoder to ensure that key information is not lost during the encoding process; A fine-tuning module, configured to obtain features using the trained feature extractor and variational autoencoder, and input the features into a classifier for fine-tuning to complete the training of the forged face identification network; An input module, configured to input the face image to be identified into the trained forged face identification network and output an identification result.

Citation Information

Patent Citations

  • Method for realizing generative false face image identification based on deep convolutional neural network

    CN111597983A

  • Forged image recognition model training method and forged image recognition method

    CN112686331A

  • Cooperative multi-feature clustering unsupervised pedestrian re-identification method and system

    CN116092122A

  • Image clustering method based on composite subspace

    CN116758320A

  • Image anomaly detection method fusing AVAE and SE modules

    CN117036786A