A highly robust deepfake face detection method

The Deep Image Prior method eliminates the adversarial perturbation in the deep fake face detection model, and improves the robustness of the model through ensemble learning, solving the problem of poor detection robustness in the prior art, and achieving efficient detection under adversarial attacks.

CN115588226BActive Publication Date: 2025-05-09CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211354009.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-01
Publication Date
2025-05-09
Estimated Expiration
2042-11-01

AI Technical Summary

Technical Problem

The existing deep-fake face detection model has poor detection robustness in the face of adversarial attacks and is overly dependent on the training dataset, making it difficult to effectively remove adversarial perturbations in real environments.

Method used

The convolutional neural network is trained through the Deep Image Prior method to obtain high impedance of image noise, eliminate hostile perturbations in perturbed images, and integrate the fake face classifier with the reconstructed image classifier through ensemble learning to improve the robustness of the model.

Benefits of technology

The robustness of the fake face detection model under adversarial attacks is improved, the adversarial ability of the model is enhanced, and the stability of detection accuracy is maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115588226B_ABST
    Figure CN115588226B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of artificial intelligence security and relates to a highly robust deep fake face detection method; the present invention mainly includes four parts: firstly, an original data set is obtained and screened to obtain a training sample; a perturbation attack is performed on a fake face detector, thereby interfering with the classification accuracy of the fake face detector and obtaining a perturbation sample; a convolutional neural network is used to eliminate adversarial perturbations in the perturbation samples to obtain a reconstructed image classifier; the reconstructed image classifier and the fake face detector subjected to the perturbation attack are integrated to finally obtain a deep fake face detection model; the present invention improves the robustness of the model and simultaneously improves the model detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence security and relates to a highly robust deep fake face detection method. Background Art

[0002] With the advent of deep learning and the availability of large datasets, fake face detection technology has achieved impressive results, and the most advanced fake face detection technology has been widely used in many fields. However, there are still groups of attackers targeting fake face detection technology, who spend time and energy to manipulate faces and try various methods to try to fool fake face detection detectors.

[0003] Traditional deep fake detection technology is generally based on image-level forensics, based on traditional signal processing, relying on specific tampering evidence, and using the frequency domain features and statistical features of the image for distinction. This method is limited by the quality of the image. Once the quality of the forged face image improves, the corresponding recognition effect will decrease. Although a large number of studies have focused on defending against synthetic adversarial attacks, these methods cannot better adapt to the current diverse forms of attack technology, and cannot make correct judgments when facing different hostile perturbation attacks, resulting in poor model detection robustness. In order to solve the problem of detection accuracy, many researchers have made great research efforts on deep fake detection and have achieved quite good results, but there are still some challenges:

[0004] 1. The fake face detection model lacks adversarial resistance. Most of the existing fake face detection models use deep neural network technology, but the neural network itself has adversarial sample attacks and lacks adversarial resistance, so it is easily affected by adversarial attacks, which makes the face detection model unable to correctly predict whether the adversarial face image is the same as the original image.

[0005] 2. Over-reliance on training data sets. Actual environmental disturbance attacks are unpredictable, and training models adapted to a large number of training sets are not reliable in removing disturbance attacks.

[0006] 3. Dynamic balance between robustness and accuracy. The blind pursuit of improving the robustness of model detection reduces the accuracy of model detection. How to dynamically balance the relationship between the two is obviously another problem that needs to be faced in the current research on improving robustness. Summary of the invention

[0007] To solve the above problems, the present invention provides a highly robust deep fake face detection method, which is characterized by comprising the following steps:

[0008] S1. Obtain a deep fake face image dataset and preprocess it to obtain a training image set;

[0009] S2. Take the training image set as the input of the fake face classifier, and use FGSM and CW2 to simultaneously perform adversarial attack training on the fake face classifier to obtain the perturbed image set;

[0010] S3. Using the Deep Image Prior method, a convolutional neural network is used to learn the perturbed image, obtain high resistance to image noise, and eliminate hostile perturbations in the perturbed image;

[0011] S4. Based on the high impedance of image noise found in S3, all the perturbed images in the perturbed image set are reconstructed and trained by the convolutional neural network in S3 to obtain a reconstructed image set;

[0012] S5. Improve the convolutional neural network, train the improved convolutional neural network with the reconstructed image set to obtain a reconstructed image classifier, and use a binary cross entropy loss function to calculate the classification loss;

[0013] S6. Integrate the fake face classifier with the trained reconstructed image classifier to obtain a deep fake face detection model, perform ensemble training, and calculate the loss using the classification ensemble loss function;

[0014] S7. Input the image to be detected into the deep fake face detection model trained in S6 to obtain the detection result.

[0015] Furthermore, the process of obtaining the training image set in step S1 includes:

[0016] S11. Download a deep fake face image set from a public dataset, or forge a deep fake face image set using fake face generation technology;

[0017] S12. Use a fake face classifier to detect all images in the deep fake face image set, and collect the images whose detection results are fake to form a training image set.

[0018] Furthermore, step S3 uses the Deep Image Prior method to obtain high noise impedance of the image. The purpose is to iterate the convolutional neural network on a single disturbed image to obtain prior information before the initialized convolutional neural network learns the specific generator network structure parameters, thereby completing the restoration of the disturbed image. The objective function constructed for this purpose is expressed as:

[0019]

[0020] Among them, x* represents the final target image, x′ represents the perturbation image, represents the generated image of the convolutional neural network, is a task-dependent data item, representing the perturbation image x′ and the generated image Minimize cross entropy between ; represents the regularization term that captures the prior information of the generated image;

[0021] Further Interpreted as the perturbation image x′ and the generated image The domain-dependent distance loss or domain-dependent similarity loss between the two domains is used, and the surjective function is introduced. The improved objective function is:

[0022]

[0023] Furthermore, when eliminating the hostile disturbance in the perturbed image in step S3, the mean square error is calculated pixel by pixel as the similarity measure, and the objective function is further optimized, which is expressed as:

[0024] min{MSE(y(χ,z),x′)}

[0025] Among them, MSE() represents mean square error, y() represents the mapping model, which is used to generate images to calculate similarity metrics, χ represents an adjustable parameter, and z represents a randomized vector seed.

[0026] Furthermore, an improved ResNet-50 network is used in the reconstructed image classifier for image classification. The improved ResNet-50 network deletes all BN layers based on the existing ResNet-50 network structure. The classification loss of the reconstructed image classifier is calculated using a binary cross entropy loss function, which is expressed as:

[0027]

[0028] in, represents the averaging operation, x′ represents the perturbed image, represents the reconstructed image, and D() represents the reconstructed image classifier.

[0029] Furthermore, the classification ensemble loss function is expressed as:

[0030]

[0031] Among them, α, β, and γ represent adjustable parameters, represents the loss of the reconstructed image classifier, represents the loss of the fake face classifier.

[0032] Furthermore, the loss of the fake face classifier is expressed as:

[0033]

[0034] in, Represents the target image output by the reconstructed image classifier Minimize the cross entropy loss between the perturbed image x′ and represents the regularized L2 reconstruction loss, ω∈R C×H×W For the encoding tensor, C, H, W represent the image height, width, and number of channels, and μ represents the network hyperparameters.

[0035] Beneficial effects of the present invention:

[0036] Based on the DIP framework, the termination condition of the iterative process is improved, the process of learning images is reasonably terminated, and the target natural images are screened out. The present invention is suitable for situations where there is no data set training model, that is, the purpose of removing adversarial perturbations is achieved in a real environment, so as to improve the robustness of the forged face detector detection model, and by integrating the classification loss of the classifier and the detector, the relationship between correctness and robustness is balanced, so that the detection robustness of the forged face detection model is improved while the original detection accuracy is maintained without reduction.

[0037] Mainstream defense schemes against adversarial fake face images tend to overly adapt to interference in a large number of data sets, but fail to make correct decisions on adversarial attacks that are invisible in real environments. Recent studies have found that adding adversarial noise to deep fake face images can affect the detection accuracy of deep fake face detectors. To solve this problem, the present invention proposes an effective and efficient solution suitable for real environments. By improving DIP, we propose a method that can remove adversarial perturbation noise in deep fake face images, that is, to eliminate adversarial perturbations by iteratively optimizing the generated convolutional neural network in an unsupervised manner. In order to improve the quality of reconstructed images, the ResNet network is improved, and unnecessary BN layers in the residual block are removed by introducing residual dense blocks (RRDB). Effective residual dense blocks are used. While expanding the model size, deeper networks with channel attention are explored, and image quality evaluation is performed by PSNR to improve the perceptual quality. In order to improve the robustness of the model without reducing the detection accuracy of the model, ensemble learning is introduced to integrate several classifiers to enhance the final fake face detection classification network, thereby improving the overall prediction robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a flow chart of the deep fake face detection method of the present invention;

[0039] Figure 2 This is a model diagram of the deep fake face detection system of the present invention;

[0040] Figure 3 A schematic diagram of eliminating adversarial attacks according to the present invention;

[0041] Figure 4 The network structure diagram of the reconstructed image classifier of the present invention;

[0042] Figure 5 Schematic diagram of DIP learning according to an embodiment of the present invention. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] The present invention provides a highly robust deep fake face detection method, which can improve the robustness of fake face detection while resisting perturbation attacks, and improve the accuracy of fake face detection. The present invention mainly includes four parts: first, the original data set is obtained and screened to obtain training samples; a perturbation attack is performed on the fake face detector, thereby interfering with the classification accuracy of the fake face detector and obtaining perturbation samples; a convolutional neural network is used to eliminate the adversarial perturbation in the perturbation sample to obtain a reconstructed image classifier; the reconstructed image classifier and the fake face detector after the perturbation attack are integrated to finally obtain a deep fake face detection model.

[0045] In one embodiment, if Figure 1 As shown, a highly robust deep fake face detection method includes the following steps:

[0046] S1. Obtain a deep fake face image dataset and preprocess it to obtain a training image set;

[0047] Specifically, the process of obtaining a training image set includes:

[0048] S11. Download a deep fake face image set from a public dataset, or forge a deep fake face image set using fake face generation technology;

[0049] S12. Use an existing deep fake detector to detect all images in the deep fake face image set, and collect the images whose detection results are forged to form a training image set.

[0050] S2. Take the training image set as the input of the fake face classifier, and use FGSM and CW2 to simultaneously perform adversarial attack training on the fake face classifier to obtain the perturbed image set;

[0051] S3. Using the Deep Image Prior method, a convolutional neural network is used to learn the perturbed image, obtain high resistance to image noise, and eliminate hostile perturbations in the perturbed image;

[0052] S4. Based on the high impedance of image noise found in S3, all the perturbed images in the perturbed image set are reconstructed and trained by the convolutional neural network in S3 to obtain a reconstructed image set;

[0053] S5. Improve the convolutional neural network, train the improved convolutional neural network with the reconstructed image set to obtain a reconstructed image classifier, and use a binary cross entropy loss function to calculate the classification loss;

[0054] S6. Integrate the fake face classifier with the trained reconstructed image classifier to obtain a deep fake face detection model, perform ensemble training, and calculate the loss using the classification ensemble loss function;

[0055] S7. Input the image to be detected into the deep fake face detection model trained in S6 to obtain the detection result.

[0056] In one embodiment, the system architecture of the deep fake face detection method proposed by the present invention is as follows: Figure 2 As shown, the training process includes:

[0057] STEP 1. Download 10K fake face images generated by the thispersondoesntexists website as a deep fake face image set; input all images in the deep fake face image set into the existing fake face classifier, the fake face classifier will judge the image and output the detection result of whether the image is a fake image or a real image, and collect all the images with fake detection results to form a training image set.

[0058] STEP 2. Then perform a perturbation attack on the fake face classifier used above to interfere with its judgment, that is, add adversarial perturbations to all images in the training image set. After the perturbed training images are input into the fake face classifier, its detection results become true, thus obtaining the perturbation-interfered fake face classifier and the perturbation image set.

[0059] STEP 3. The DIP framework proposed by the Deep Image Prior method starts from the convolutional neural network itself and starts to learn a single noisy image (also referred to as a perturbed image in this embodiment) from scratch to obtain high noise resistance of the image. In layman's terms, the high noise resistance of the image means that in the process of learning the noisy image, the convolutional neural network will first learn the features of the non-noise image and then learn the noise in the noisy image. Based on the high noise resistance of the image, the convolutional neural network is used to remove the adversarial interference in the perturbed image to obtain a reconstructed image, such as Figure 3 shown.

[0060] STEP4. Introduce the improved ResNet-50 network into the convolutional neural network and build a reconstructed image classifier, such as Figure 4As shown in the figure. The improved ResNet-50 network is based on the existing ResNet-50 network structure, but all BN layers are removed. Originally, the BN layer uses the mean and variance in a batch of samples to normalize features during the training of the ResNet-50 network, and uses the mean and variance of the entire training dataset during the test. However, when the statistical difference between the training dataset and the test dataset is large, the BN layer will introduce unpleasant noise. In the image deblurring task, removing the BN layer has been shown to improve performance and reduce computational complexity.

[0061] STEP5. Most fake face detection will fail to generalize due to the huge differences in the generated images. This is due to the imbalance of the data set or the different network architectures, loss functions and image preprocessing methods, which result in a model that has good recognition effects on some data sets but fails to achieve the expected results on other data sets. In order to train a model with good robustness and accuracy, the present invention adopts an ensemble learning method to enhance the decision-making effect by integrating several weak classifiers. Specifically, the perturbation interference fake face classifier is integrated with the reconstructed image classifier to obtain a deep fake face detection model. The overall Figure 2 shown.

[0062] Preferably, although the accuracy of the current fake face detection technology continues to improve, it is still susceptible to adversarial samples, resulting in poor robustness of the fake face detection model, causing the fake face detection model to be misclassified. This is because most studies focus on evaluating the effectiveness of their methods on a limited number of known Deepfake fake face generation networks or simple data sets. However, due to the extremely fast development speed of the generation network, the detection technology that can be applied in the previous generation network may be deleted or destroyed. In real scenarios, Deepfake will suffer from adversarial noise attacks of various disturbances that are not easily detected, which has also become the biggest obstacle to developing a highly robust Deepfake detection model. Therefore, this embodiment considers implementing a hostile perturbation attack in STEP2.

[0063] In adversarial perturbation attacks, white-box attacks assume that the adversary has full access to the attacked model, including the model architecture and parameters. Black-box attacks assume that the adversary has limited or almost no information about the attacked model. Black-box attacks may involve varying degrees of access to the attacked model, such as access to prediction probabilities, prediction categories, and even training data. In this implementation, Fast Gradient Sign Method (FGSM) and Carlini and Wagner L2attack (CW2) norm attacks are used based on white-box and black-box environments to create a forged face classifier, which greatly reduces the accuracy of the forged face classifier.

[0064] Specifically, FGSM is an algorithm for generating adversarial samples based on gradients, which belongs to the non-targeted attack in adversarial attacks (i.e., the adversarial samples are not required to be specified by the model prediction, as long as they are different from the original sample predictions). This embodiment uses FGSM for interference to generate a perturbed image that causes the forged face classifier to misjudge, expressed as:

[0065] x′=x+εsign(▽ x J(x,y,θ))(1)

[0066] Where x′ represents the perturbed image, x represents the training image, y represents the true category of the training image x, θ represents the weight parameter of the fake face classifier, J() represents the loss function, ▽ x It represents the partial derivative of the training image x, sign() represents the sign function, and ε represents a hyperparameter used to control the perturbation size of each pixel. By retaining the minimum ε, the size of the disturbance can be limited.

[0067] Specifically, CW2 is a slow but stronger attack that only considers the interference of the attack on the image, that is, the goal of the adversarial attack is to manipulate the image itself, and the goal of the adversarial attack is to misclassify the perturbed image. In the process of generating adversarial samples by CW2, there are two attack goals. The first goal is to minimize the L2 norm of the perturbed image x′ and the training image x, which is expressed as:

[0068]

[0069] The second goal is to try to make the perturbation cause misclassification, which is expressed as follows:

[0070]

[0071] y C =min(max{|C s (x)-C s (x′ i )|},-κ) (4)

[0072] Among them, y(x′) represents the output that effectively leads to misclassification, C s (x) represents the classification probability of the fake face classifier for the training image x, C s (x i ′) represents the perturbed image x generated by the fake face classifier for the i-th time i ′, the classification probability, x i ′ represents the perturbed image generated for the i-th time during the CW2 adversarial perturbation process for the training image x, and κ represents a parameter that defines a threshold, which is adjusted so that the logically incorrect prediction class exceeds the true target class. Further:

[0073]

[0074] In order to ensure that the perturbed image x′ can be within the interval [0,1], formulas (2), (3), (4), and (5) are integrated to obtain:

[0075]

[0076] in, represents the disturbance factor of the CW2 hostile interference process, tanh()(-1≤tanh()≤1) means that the generated perturbation image always falls within the interval [0,1], λ represents the intensity parameter controlling the two targets, It represents how to balance two objectives and transform the nonlinear problem into a linear problem.

[0077] Preferably, in real scenarios, the images we get are often images after disturbances are added by hostile attacks, so we do not have enough data sets to learn the difference between normal deep fake face images and images after disturbances are added. In STEP3, in order to eliminate the impact of adversarial attacks on fake face classifiers, this embodiment defines the elimination of adversarial attacks as an image denoising problem, that is, an image reconstruction problem. Based on the theory proposed by Deep Image Prior, the input perturbed image is learned through the learning ability of the convolutional neural network, that is, the main purpose is to repeatedly iterate on a single perturbed image to obtain prior information before the initialized convolutional neural network learns the specific generator network structure parameters, thereby completing the restoration of the perturbed image; the objective function constructed for this purpose is expressed as:

[0078]

[0079] Where x* represents multiple generated images reshaped based on the neural network, and the final destination image is obtained by minimizing the domain-dependent distance or dissimilarity between the perturbed image and the generated image, and x′ represents the perturbed image. represents the generated image of the convolutional neural network, is a task-dependent data item, representing the perturbation image x′ and the generated image Minimize cross entropy between ; Represents the captured image Regularization term of prior information;

[0080] Specifically, starting from the real scene, using the high resistance to image noise found in DIP, the convolutional neural network can naturally eliminate the adversarial perturbation from the image containing the perturbation attack in an unsupervised manner, and the image restoration can be completed without a training set. Reconstructing the original deep fake image and restoring the training image x from the perturbed image x′ is usually an uncertain problem, so regularization is crucial. By combining the key ideas of DIP, the formula (7) Interpreted as the perturbation image x′ and the generated image The domain-dependent distance loss or domain-dependent similarity loss between the two domains is used, and the surjective function is introduced. The improved objective function is:

[0081]

[0082] In general, prior knowledge favors natural images rather than damaged images. Good reconstructed images can be successfully classified in the optimization trajectory, so we have to consider the optimization term to obtain better reconstructed images. The image restoration framework described above is used to remove adversarial interference from adversarial samples (perturbed images). The mean square error (MSE) is calculated pixel by pixel on the image as a similarity measure. In order to be able to represent the result of eliminating adversarial interference, continue to optimize formula (8):

[0083] min{MSE(y(χ,z),x′)}(9)

[0084] Among them, MSE() represents mean square error, y() represents the mapping model, which is applied to image generation to calculate similarity metrics, χ represents an adjustable parameter, and z represents a randomized vector seed to replace g(χ).

[0085] Specifically, the convolutional neural network itself plays a priori role. Through the convolutional layer, the convolutional neural network obtains the internal structure and self-similarity of the image without adversarial noise from the perturbed image. This embodiment proposes an active defense strategy, that is, actively removing the noise from the image containing adversarial noise. In the process of the convolutional neural network learning the noisy image from 0 to 1, it is assumed that the undisturbed image and the image containing adversarial noise have different behaviors in the entire iterative optimization process, and the image containing adversarial noise in the entire iterative process, the denoised image usually appears at a certain position in the optimization trajectory of the iterative curve inclination angle, such as Figure 5 As shown; therefore, it is only necessary to send the generated image at this time to the existing classifier in the intermediate iteration; then the classifier performs screening to obtain the appropriate image, and uses it as the result of the final denoised natural image.

[0086] Specifically, the structure of the convolutional neural network is as follows Figure 3As shown in the figure, it uses a full convolution structure of downsampling and upsampling with a nonlinear activation function, where downsampling is achieved based on convolution adjustment stride, and averaging and maximum pooling are performed based on lanzeos interpolation, and upsampling uses nearest neighbor upsampling to enlarge and restore the original size of the image.

[0087] Preferably, the reconstructed image classifier established in STEP4 is as follows: Figure 4 As shown in Figure 1, similar to popular adversarial detection classifiers, the decision boundary between real images and adversarial images should be learned. However, the reconstructed image classifier proposed in this paper is not trained using adversarial images pre-computed with known attacks, but distinguishes between real images and adversarial images in an active supervised learning manner. The proposed method does not require a large number of pre-computed adversarial image datasets for training, and uses a binary convolutional neural network to distinguish between the perturbed image x′ and the reconstructed image And trained with binary cross entropy loss:

[0088]

[0089] in, represents the averaging operation, and D() represents the ResNet classifier (that is, the reconstructed image classifier).

[0090] Specifically, the improved ResNet-50 network is used to calculate the inclination angle in the DIP image denoising process, and the classifier is trained for multiple periods on the training dataset. Simply put, according to the DIP method, a convolutional neural network is used to learn a noisy image (perturbed image) from scratch. The convolutional neural network will generate an image at different periods (different positions) of the learning process, such as Figure 5 As shown, a clean and noise-free image must appear at one or several positions. What this embodiment needs to do is to select an appropriate position from these positions. Figure 5 There is an obvious turning point (inclination) in the loss graph, and this turning point is taken as the appropriate position. Given a target neural network, the loss of the perturbed image and the image generated by the target neural network at a certain position is calculated, and the loss function is defined as the L1 norm of their difference:

[0091]

[0092] in, Refers to the image generated by the target neural network at a certain position, y l () represents the true value y of the model at layer l l (x′) and predicted value The value of .

[0093] Specifically, the results of removing adversarial attacks are measured using the true loss of the classifier at different epochs:

[0094]

[0095] in, represents the loss of measuring the effect of removing adversarial attacks, L x′ Represents the reconstructed images generated at different times for x′ Taking the average operation, D() represents the classification of the generated reconstructed image by the ResNet classifier. Represents the reconstructed images produced at different time periods j.

[0096] Preferably, in STEP5, using an ensemble learning classifier instead of a single classifier can improve the generalization ability of the model, and the gain obtained by the combination is more affected by the content selection presented to the combined classifier rather than by the actual selected combined classifier method. Specifically, the perturbation interference forged face classifier D and the reconstructed image classifier C are integrated to obtain a deep fake face detection model, and the training image x is used as the input of the deep fake face detection model to minimize the detection error of the perturbation interference forged face classifier D and the reconstructed image classifier C, while improving the detection robustness of D without reducing the accuracy, and the image quality of the reconstructed image is evaluated to obtain a reconstructed image with clean peak detection.

[0097] Specifically, in the process of integrated model learning, the loss function is mainly used to balance the relationship between the fake face classifier D and the reconstructed image classifier C, which is expressed as the cross entropy loss between the image judged as fake by the fake face classifier D and the image obtained by DIP reconstruction (reconstructed image classifier C) when the fake face classifier D divides the data set. H(·,·) represents the cross entropy loss. The first part is to minimize the loss between the target image output by the reconstructed image classifier and the perturbation image x′. The second part is to minimize the cross entropy loss of the generated image during the DIP learning perturbation image. The third part is to minimize the cross entropy of the target image screened by the reconstructed image classifier after being detected by the fake face classifier D. The total loss is the weighted sum of all individual losses, as follows:

[0098]

[0099] in, Represents the target image output by the reconstructed image classifier Minimize the cross entropy loss between the perturbed image x′ and represents the regularized L2 reconstruction loss, ω∈R C×H×WFor the encoding tensor, C, H, W represent the image height, width, and number of channels, and μ represents the network hyperparameters.

[0100] In the present invention, unless otherwise clearly stipulated and limited, the terms such as "installation", "setting", "connection", "fixation" and "rotation" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral one; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium; it can be the internal connection of two elements or the interaction relationship between two elements. Unless otherwise clearly defined, ordinary technicians in this field can understand the specific meanings of the above terms in the present invention according to the specific circumstances.

[0101] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A highly robust deep fake face detection method, characterized in that: The following steps are involved: S1. Obtain a deep fake face image dataset and preprocess it to obtain a training image set; S2. Take the training image set as the input of the fake face classifier, and use FGSM and CW2 to simultaneously perform adversarial attack training on the fake face classifier to obtain the perturbed image set; S3. Using the Deep Image Prior method, a convolutional neural network is used to learn the perturbed image, obtain high resistance to image noise, and eliminate hostile perturbations in the perturbed image; S4. Based on the high impedance of image noise found in S3, all the perturbed images in the perturbed image set are reconstructed and trained by the convolutional neural network in S3 to obtain a reconstructed image set; S5. Improve the convolutional neural network, train the improved convolutional neural network with the reconstructed image set to obtain a reconstructed image classifier, and use a binary cross entropy loss function to calculate the classification loss; The improved ResNet-50 network is used in the reconstructed image classifier for image classification. The improved ResNet-50 network is based on the existing ResNet-50 network structure with all BN layers deleted. The binary cross entropy loss function is used to calculate the classification loss of the reconstructed image classifier, which is expressed as: in, represents the averaging operation, x′ represents the perturbed image, represents the reconstructed image, D() represents the reconstructed image classifier, D(x) represents the classification loss of ResNet-50 relative to the average classifier for x in image classification, and D(x′) represents the classification loss of ResNet-50 relative to the average classifier for x′ in image classification; S6. Integrate the fake face classifier with the trained reconstructed image classifier to obtain a deep fake face detection model, perform ensemble training, and calculate the loss using the classification ensemble loss function; The classification ensemble loss function is expressed as: Among them, α, β, and γ represent adjustable parameters, represents the loss of the reconstructed image classifier, represents the loss of fake face classifier; H(.,.) represents the cross entropy loss; The loss of the fake face classifier is expressed as: in, Represents the target image output by the reconstructed image classifier Minimize the cross entropy loss between the perturbed image x′ and represents the regularized L2 reconstruction loss, ω∈R C×H×W is the encoding tensor, C, H, W represent the image height, width and number of channels, and μ represents the network hyperparameters; S7. Input the image to be detected into the deep fake face detection model trained in S6 to obtain the detection result.

2. A highly robust deep fake face detection method according to claim 1, characterized in that: The process of obtaining the training image set in step S1 includes: S11. Download a deep fake face image set from a public dataset, or forge a deep fake face image set using fake face generation technology; S12. Use a fake face classifier to detect all images in the deep fake face image set, and collect the images whose detection results are fake to form a training image set.

3. A highly robust deep fake face detection method according to claim 1, characterized in that: Step S3 uses the Deep Image Prior method to obtain the high impedance of image noise. The purpose is to use the convolutional neural network to repeatedly iterate on a single disturbed image to obtain prior information before the initialized convolutional neural network learns the specific generator network structure parameters, thereby completing the restoration of the disturbed image. The objective function constructed for this purpose is expressed as: Among them, x* represents the final target image, x′ represents the perturbation image, represents the generated image of the convolutional neural network, is a task-dependent data item, representing the perturbation image x′ and the generated image Minimize cross entropy between ; Represents the captured image Regularization term of prior information; Further Interpreted as the perturbation image x′ and the generated image The domain-related distance loss or domain-related similarity loss between them is used, and the surjective function g is introduced: The improved objective function is:

4. A highly robust deep fake face detection method according to claim 3, characterized in that: When eliminating the hostile disturbance in the perturbed image in step S3, the mean square error is calculated pixel by pixel as the similarity measure, and the objective function is further optimized, which is expressed as: min{MSE(y(χ,z),x′)} Among them, MSE() represents mean square error, y() represents the mapping model, which is used to generate images to calculate similarity metrics, χ represents an adjustable parameter, and z represents a randomized vector seed.

Citation Information

Patent Citations

  • Forgery image detection method and system based on posterior probability, and storage medium

    CN114332536A

  • Graphics processing unit for detecting cheating using neural networks

    CN114588636A