High-robustness deep forgery detection method and device based on submerged space countermeasure purification

By introducing Gaussian noise into the latent space for diffusion and denoising, an adversarial purifier is built, which solves the robustness of the deep forgery detection model under unknown attacks, and realizes efficient detection in an adversarial attack environment.

CN120339814APending Publication Date: 2025-07-18WUHAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510467745.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing deep forgery detection models have significantly reduced their defense effects when facing unknown attack strategies, making it difficult to maintain robustness in the anti-attack environment.

Method used

By introducing Gaussian noise into the latent space for diffusion and denoising processes, an anti-purifier is built to eliminate anti-noise interference, enhance the robustness of image reconstruction, and purify the image before detection.

Benefits of technology

Maintaining strong defense capabilities under unknown attack strategies improves the robustness and reliability of the deep forgery detection model without modifying the existing detector structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339814A_ABST
    Figure CN120339814A_ABST
Patent Text Reader

Abstract

The invention discloses a high-robustness deep forgery detection method and device based on submerged space adversarial purification, and the method comprises the steps: constructing a robust auto-encoder training frame fusing adversarial attack and a dynamic disturbance mechanism, so as to enhance the robustness of the mapping of a submerged space and an image space; therefore, the original real semantics of the image reconstructed by the auto-encoder can be accurately kept. For different types of adversarial attacks, Gaussian distribution noise is introduced into a submerged space for diffusion, and a submerged vector is gradually restored by using a denoising solver based on a U-Net architecture, so that the purified vector is highly consistent with the original submerged vector, and the adversarial noise in the deep counterfeit content is effectively removed. In addition, the confrontation purifier provided by the invention has a plug-and-play characteristic, does not need to re-train an existing deep counterfeit detector, and shows good generalization ability at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence information security, and more specifically, to a highly robust deep fake detection method and device based on latent space adversarial purification. Background Art

[0002] With the breakthrough of deep learning technology and the rapid improvement of computing power, the field of artificial intelligence-generated content has witnessed explosive development, greatly expanding the creativity and application boundaries of content generation. As the core technology in this field, generative adversarial networks (GANs) have made breakthroughs in aspects such as image synthesis, editing, super-resolution reconstruction, and video generation, and their academic value and application potential have reached a broad consensus in the industrial community. However, the universality of technology has brought about a dual effect: on the one hand, the popularization of open-source frameworks has lowered the threshold for using deep fake tools, and the continuous optimization of algorithms has made the generated content increasingly approaching reality, thus promoting the innovation of emerging industries such as virtual reality; on the other hand, technology diffusion has also brought new security challenges. Malicious users can use forged audio and video to carry out precise fraud, manipulate public opinion, and even commit identity theft through high-precision deep fakes, threatening social security and the trust system.

[0003] The core dilemma faced by the current deep fake detection system lies in the widespread existence of adversarial vulnerability. Attackers only need to inject tiny perturbations into the original input to render the detection model ineffective, and this threat has spread from the digital space to the physical world. For example, autonomous driving systems may cause traffic accidents due to misjudging road signs, forged medical images may interfere with clinical diagnosis, and highly realistic political deep fake content may seriously damage social trust. It is worth noting that the vulnerability of existing detection models in deep fake scenarios is particularly prominent, that is, malicious attackers can embed adversarial noise specifically designed for detectors in forged images generated by GANs, rendering the vast majority of current detection methods ineffective. Different from the detectable traces caused by the degradation of the quality of forged content during the dissemination process, this adversarial noise has a high degree of concealment and pertinence, making its destructiveness significantly enhanced.

[0004] In order to improve the robustness of deep fake detection models in adversarial attack environments, the current mainstream methods mainly rely on adversarial training to enhance the defense capabilities of the models. However, this method still has significant limitations in practical applications. Since the defense side usually cannot obtain the detailed parameters of the attack algorithm, when the model faces unknown attack strategies, its defense effect often drops significantly, resulting in impaired detection performance. Summary of the Invention

[0005] In view of the problems existing in the prior art, the present invention starts from the perspective of data preprocessing and proposes a plug-and-play adversarial purification method. This method is based on the diffusion and denoising processes in the latent space. By first mapping the input samples to the latent space and disrupting the distribution of adversarial noise by adding Gaussian noise, a perturbation-free sample close to the original clean image is reconstructed after denoising, thereby reducing the interference of adversarial noise on the detection model. Since this method does not rely on a specific attack strategy or detection model, but eliminates the impact of adversarial attacks through a general denoising mechanism, it can still maintain strong defense capabilities when facing unknown attacks. In addition, this method can be deployed as an independent module before the deepfake detector. Without modifying the detector structure, it can significantly improve its robustness in the adversarial attack environment, thereby enhancing the reliable recognition ability of deepfake content.

[0006] To achieve the above object, the first aspect of the present invention provides a highly robust deepfake detection method based on latent space adversarial purification, including:

[0007] Collect original image data, preprocess the collected image data, and divide it into training data and test data;

[0008] Use the training data to train an adversarial purifier based on a diffusion model. Among them, the adversarial purifier based on the diffusion model includes an autoencoder and a latent space adversarial purification model. The autoencoder includes an encoder and a decoder. The encoder is used to extract and compress the features of the image data. The latent space adversarial purification model includes a diffusion stage and a denoising stage. In the diffusion stage, Gaussian noise is gradually added to the latent vector according to a preset noise step size and intensity distribution to construct a noise sequence; in the denoising stage, the noise in the noise sequence is gradually removed to restore the latent vector to a purified sample close to the original image. The decoder is used to reconstruct the image features and map them back to the image space;

[0009] Input the test data into the trained adversarial purifier based on the diffusion model to obtain the adversarially purified image;

[0010] Input the adversarially purified image into a pre-constructed deepfake detection model to obtain a detection result.

[0011] In one implementation, preprocessing the collected image data includes:

[0012] Adjust the size and normalize the collected image.

[0013] In one implementation, the method further includes:

[0014] Introduce random rotation and noise addition techniques to enhance the training data.

[0015] In one implementation, when training an adversarial purifier based on a diffusion model using training data, the PGD (Projected Gradient Descent) adversarial attack algorithm is used to perturb the original image to generate adversarial samples. During the gradient calculation process, a transfer attack strategy is introduced, and the VGG-19 is used as a surrogate model to obtain gradient direction information.

[0016] In one implementation, training an adversarial purifier based on a diffusion model using training data includes:

[0017] Input the original image and the adversarial sample into the latent space, and perturb the latent vector z of the original image in the direction of the latent vector z of the adversarial sample in the latent space to obtain the perturbed latent vector z adv ; r ;

[0018] Input the latent vector z of the adversarial sample adv into the latent space adversarial purification model, and obtain the adversarially purified latent vector z through the operations of the diffusion stage and the denoising stage p ;

[0019] Obtain the reconstructed image I through the decoder D for the perturbed latent vector z r ; In the training process, the optimal model parameters are obtained by minimizing the following loss function value: rec , where:

[0020]

[0021] where z p represents the adversarially purified latent vector, is used to measure the difference between the latent vector z of the original image and the adversarially purified latent vector z p , and the mean squared error metric is adopted. q φ (z|x) represents the posterior distribution parameterized by the encoder, p(z) represents the prior distribution of the latent vector of the original image, which is used to constrain the distribution structure of the encoder output. The D KL divergence is used to measure the difference between the posterior distribution q φ (z|x) generated by the encoder and the prior distribution p(z), and is used as a regularization term in the training process to constrain the model to converge to the optimal value.

[0022] In one implementation, inputting the test data into the trained adversarial purifier based on the diffusion model to obtain the adversarially purified image includes:

[0023] During the diffusion stage, the adversarial sample z adv undergoes a t-step forward diffusion process to obtain a sample containing standard Gaussian noise and the original adversarial noise , which is formally expressed as follows:

[0024]

[0025] Among them, represents the latent vector of the adversarial example after t steps of the diffusion process, represents the cumulative weight coefficient of t steps in the diffusion process, ∈ is standard Gaussian noise, and σ′ represents the perturbation;

[0026] In the denoising stage, the denoising operation of each step is modeled through a parameterized conditional probability distribution p θ (z t-1 |z t ), and the specific form is:

[0027]

[0028] Among them, z t represents the latent vector obtained after t steps of diffusion, z t-1 represents the latent vector at t - 1 steps in the reverse diffusion process, and is predicted from z t through a denoising solver based on the U - Net and attention mechanism. The mean μ θ (z t , t) and the variance are predicted by the denoising solver.

[0029] In one implementation, the pre - constructed deepfake detection model is a model based on the ResNet structure, including a convolutional neural network and an output layer. The adversarially purified image is input into the pre - constructed deepfake detection model to obtain a detection result, including:

[0030] Extract discriminative features from the adversarially purified image through the convolutional neural network;

[0031] Output a trust score based on the discriminative features through the output layer.

[0032] Based on the same inventive concept, the second aspect of the present invention provides a highly robust deepfake detection method and device based on latent space adversarial purification, including:

[0033] A data acquisition and pre - processing module, configured to collect original image data, pre - process the collected image data, and divide it into training data and test data;

[0034] A training module for training an adversarial purifier based on a diffusion model using training data. The adversarial purifier based on the diffusion model includes an autoencoder and a latent space adversarial purification model. The autoencoder includes an encoder and a decoder. The encoder is used to extract and compress features of image data. The latent space adversarial purification model includes a diffusion stage and a denoising stage. In the diffusion stage, Gaussian noise is gradually added to the latent vector according to a preset noise step size and intensity distribution to construct a noise sequence. In the denoising stage, the noise in the noise sequence is gradually removed to restore the latent vector to a purified sample close to the original image. The decoder is used to reconstruct image features and map them back to the image space.

[0035] An adversarial purification module for inputting test data into the trained adversarial purifier based on the diffusion model to obtain an adversarially purified image.

[0036] A detection module for inputting the adversarially purified image into a pre-constructed deepfake detection model to obtain a detection result.

[0037] Based on the same inventive concept, the third aspect of the present invention provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, it implements the high-robust deepfake detection method based on latent space adversarial purification described in the first aspect.

[0038] Based on the same inventive concept, the fourth aspect of the present invention provides a computer device including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the high-robust deepfake detection method based on latent space adversarial purification described in the first aspect.

[0039] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:

[0040] The present invention provides a high-robust deepfake detection method based on latent space adversarial purification, constructs an adversarial purifier based on a diffusion model, and constructs a robust autoencoder training framework integrating adversarial attacks and dynamic perturbation mechanisms to enhance the robustness of the mapping between the latent space and the image space, thereby ensuring that the image reconstructed by the autoencoder can accurately retain the original true semantics. For different types of adversarial attacks, the present invention introduces Gaussian distributed noise for diffusion in the latent space and uses a denoising solver to gradually restore the latent vector, making the purified vector highly consistent with the original latent vector, thereby effectively removing the adversarial noise in the deepfake content. In addition, the adversarial purifier proposed by the present invention has the plug-and-play characteristic, does not require retraining of the existing deepfake detector, and at the same time exhibits good generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0042] Figure 1 This is a flowchart of the high-robust deepfake detection method based on latent space adversarial purification according to an embodiment of the present invention;

[0043] Figure 2 This is an overall framework diagram of the high-robust deepfake detection method based on latent space adversarial purification according to an embodiment of the present invention;

[0044] Figure 3 This is a specific application diagram of the high-robust deepfake detection method based on latent space adversarial purification according to an embodiment of the present invention;

[0045] Figure 4 This is a module diagram of the high-robust deepfake detection device based on latent space adversarial purification according to an embodiment of the present invention. Specific Embodiments

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0047] Embodiment 1

[0048] This embodiment discloses a high-robust deepfake detection method based on latent space adversarial purification. Please refer to Figure 1 , including:

[0049] S1: Collect the original image data, preprocess the collected image data, and divide it into training data and test data.

[0050] In the specific implementation process, collect suspected forged pictures from the social network as the original image data. The preprocessing includes operations such as size adjustment and normalization. For example, crop pictures with different resolutions to a unified size and normalize pictures in different formats to scale them to a unified dimension.

[0051] S2: Use the training data to train the adversarial purifier based on the diffusion model. The adversarial purifier based on the diffusion model includes an autoencoder and a latent space adversarial purification model. The autoencoder includes an encoder and a decoder. The encoder is used to extract and compress the features of the image data. The latent space adversarial purification model includes a diffusion stage and a denoising stage. In the diffusion stage, according to the preset noise step size and intensity distribution, Gaussian noise is gradually added to the latent vector to construct a noise sequence. In the denoising stage, the noise in the noise sequence is gradually removed to restore the latent vector to a purified sample close to the original image. The decoder is used to reconstruct the image features and map them back to the image space.

[0052] Specifically, using the training data to train the adversarial purifier based on the diffusion model is to train the autoencoder in the adversarial purifier based on the diffusion model.

[0053] S3: Input the test data into the trained adversarial purifier based on the diffusion model to obtain the adversarially purified image.

[0054] S4: Input the adversarially purified image into the pre-constructed deepfake detection model to obtain the detection result.

[0055] Specifically, the pre-constructed deepfake detection model is an image classifier and can be obtained based on the existing convolutional network structure. Please refer to Figure 2 , which is the overall framework diagram of the high-robust deepfake detection method based on latent space adversarial purification in the embodiments of the present invention.

[0056] In one implementation, preprocess the collected image data, including:

[0057] Adjust the size and perform normalization processing on the collected image.

[0058] In the specific implementation process, when adjusting the size of the image, use the nearest neighbor interpolation method to adjust it to a unified size of 256×256. Next, perform a cropping operation to extract the corresponding region from the image according to the set target size. During the cropping process, the pixels outside the target size range will be removed, and the region that meets the target size will be retained. The final size of the image will be standardized to 224×224 to meet the subsequent operation requirements.

[0059] In addition, to ensure that the numerical distribution of the input data is more uniform, improve the training stability, and help the deep learning model converge better, in this implementation, the image data is transformed into a distribution with a mean μ of 0 and a standard deviation σ of 1. For each channel, use the formula for normalization processing to map the pixel values from the range [0, 255] to [0, 1].

[0060] To improve the generalization ability of subsequent model training, this embodiment further includes:

[0061] Introduce random rotation and noise addition techniques to enhance the training data.

[0062] In the specific implementation process, multiple data augmentation methods are used to expand the training data. This link introduces preprocessing techniques such as random rotation and noise addition to build a multi-dimensional feature enhancement mechanism, thereby improving the generalization ability of subsequent model training.

[0063] In step S2, when training the adversarial purifier based on the diffusion model using the training data, the PGD adversarial attack algorithm is used to perturb the original image to generate adversarial samples. In the gradient calculation process, a transfer attack strategy is introduced, and the VGG-19 is used as a surrogate model to obtain gradient direction information.

[0064] Specifically, to improve the anti-interference ability of the pre-trained autoencoder in the generation task, this study introduces an adversarial attack strategy for optimization. First, the PGD adversarial attack algorithm is used to perturb the preprocessed image x to generate an adversarial sample x adv . To enhance the transferability of the attack and overcome the problem of invisible gradients in black-box attacks, a transfer attack strategy is introduced in the gradient calculation process, and the VGG-19 is used as a surrogate model to obtain gradient direction information. The generated x adv is used to train the autoencoder, which consists of an encoder and a decoder. The encoder maps the image to a low-dimensional latent space, and the decoder maps the latent vector back to the image space to reconstruct the image. Subsequently, the original image x and the corresponding adversarial sample x adv are used as a pair of inputs to the encoder and projected into the latent space respectively to obtain the corresponding latent vectors z and z adv . Further, in the latent space, a perturbation is applied to z along the direction of z adv , and the perturbation intensity is dynamically adjusted by a random step size σ. Finally, the perturbed latent vector z r =z + σ is input into the decoder D and mapped back to the image space to improve the adversarial robustness of the model.

[0065] In the training process of step S3, the image data to be detected is input into the encoder and compressed to a low-dimensional space to reduce the computational amount and obtain the output latent vector. Subsequently, the latent vector is then subjected to a diffusion process. In the forward diffusion stage, a latent vector with random noise is generated, and Gaussian noise is gradually injected into the latent vector to generate a sequence of noisy latent vectors {z1, z2,..., z t}. To retain the original semantic information, the noise addition process sets the time step t so that it does not completely approximate the pure noise distribution; subsequently, in the reverse purification stage, the noisy latent vector z tAs the initial input, a denoising solver based on the U-Net architecture is used to gradually predict z t-1 , z t-2 ,..., z0. Among them, at each step, the latent vector z i adjusts the noise component through the latent vector z i-1 at the previous moment to gradually approximate the original clean latent representation z0. Finally, through the inverse diffusion process, the denoised latent vector z0 that eliminates noise interference and retains real semantic features is reconstructed, and is mapped back to the image space through the decoder E to generate a high-fidelity denoised image.

[0066] In the detection process of step S4, the adversarially denoised image is input into a pre-constructed deepfake detection model, whose dimension is the same as that of the original image. Through feature extraction and fully connected layer processing, a confidence level is output to determine the authenticity of the image.

[0067] In one implementation, the pre-training process of the autoencoder in step S2 is as follows:

[0068] Step 2.1, both the original image x and the adversarial sample x adv are input into the latent space. Since the latent vector z adv gradually approaches z, therefore, a step size is randomly selected in the direction of z along z adv during the training process to make the image obtained by passing z = z + s through the decoder as similar as possible to the original image. r

[0069] Step 2.2, the obtained vector z r is input into the decoder D to obtain the reconstructed image I rec . The training process of this model obtains the optimal model parameters by minimizing the following loss function value:

[0070]

[0071] Among them, z p represents the adversarially denoised latent vector, is used to measure the difference between the original latent vector z and the predicted or reconstructed latent vector z p , and the mean squared error metric is adopted. q φ (z|x) represents the posterior distribution parameterized by the encoder, indicating the distribution of the latent vector z under the condition of the given input image x, and is the key loss component for training the autoencoder. p(z) represents the prior distribution of the latent variable z, which is used to constrain the distribution structure of the encoder output. D KL divergence is used to measure the difference between the posterior distribution q φ (z|x) generated by the encoder and the prior distribution p(z), and is used as a regularization term during the training process to constrain the model to converge to the optimal value.

[0072] In one implementation, the latent space adversarial purification process in step S3 is described in detail as follows:

[0073] Step S3.1, adversarial sample z adv After t steps of the forward diffusion process, a sample containing standard Gaussian noise and the original adversarial noise is obtained Formally expressed as follows:

[0074]

[0075] Among them, represents the latent vector of the adversarial sample after t steps of the diffusion process. In this process, Gaussian noise is gradually injected into the original latent vector to generate a latent representation containing random noise and the original adversarial perturbation. represents the cumulative weight coefficient of t steps in the diffusion process, which is composed of the product of the noise injection coefficients of each step (the ratio of signal retention to noise injection for the original signal at each step), and is used to regulate the signal retention and noise injection ratio in the forward diffusion process. ∈ is the standard Gaussian noise. As the number of steps t increases, the value of gradually decreases, while gradually increases. During the adversarial attack process, in order to ensure the concealment of the adversarial sample, the perturbation σ′ is very small relative to the image x. Therefore, by selecting an appropriate number of steps t, the of the Gaussian noise is large enough compared to the perturbation noise .

[0076] Step S3.2, the goal of the reverse diffusion process is to gradually denoise from the noise state z t and finally recover the original data z0. Its core is to model the denoising operation of each step through the parameterized conditional probability distribution p θ (z t-1 |z t ), and the specific form is:

[0077]

[0078] Among them, z t represents the latent vector obtained after t steps of diffusion, in which Gaussian noise has been gradually injected. z t-1 represents the latent vector at the (t - 1)th moment in the reverse diffusion process, which is predicted by the denoising solver from z t , and a cleaner latent representation closer to the original data is gradually recovered. The mean μ θ (z t ,t) and variance Prediction is based on the neural network U-Net and the Attention-based structure. The entire generation process can be decomposed into a product form of Markov chains:

[0079]

[0080] According to the derivation of the forward diffusion process and variational inference, the mean ω θ (z t , t) is parameterized as:

[0081]

[0082] where p θ (z 0:t ) represents the joint distribution from the initial latent vector to the t-th step, which is defined by the parameter θ to describe the probability model of the latent vector sequence in the inverse diffusion process. p(z t ) represents the prior distribution of the latent vector at the t-th step to describe the noise state of the diffusion model. α t represents the noise injection coefficient at the t-th step.

[0083] In one implementation, step S4 specifically includes:

[0084] Step S4.1, extracting discriminative features from the adversarially purified image through a convolutional neural network;

[0085] Step S4.2, outputting a trust score based on the discriminative features through an output layer.

[0086] Specifically, for the task of authenticating the authenticity of the purified image, an image classification model is designed as a classifier to extract discriminative features through its deep residual network structure and output the final trust score. This model jumps the input features to the output layer through skip connections, and finally outputs through an average pooling layer, a fully connected layer, and a softmax activation function. By calculating the difference between the predicted class and the true class y, the cross-entropy loss function is selected to calculate and update the parameters of the classifier iteratively:

[0087]

[0088] y i represents the predicted class, and y represents the true class.

[0089] The technical solution of the present invention will be further specifically described below through specific embodiments in conjunction with the accompanying drawings.

[0090] The method of this embodiment is to solve the problem that the existing deep fake detection model is insufficient to cope with the escape detection of forged images in an adversarial context. For the convenience of description, now taking Figure 2The model algorithm in it illustrates the specific implementation process of the deepfake detection method.

[0091] Taking the data to be detected obtained from social media image data as the input, first, feature extraction and compression are performed through an encoder. Among them, the encoder consists of structures such as convolutional layers and pooling layers. Through convolutional operations and downsampling operations, high-dimensional image data is mapped to a low-dimensional latent space vector. Subsequently, the latent vector is input into the latent space adversarial purification model, which includes two stages: diffusion and denoising. In the diffusion stage, Gaussian noise is gradually added to the latent vector according to the preset noise step size and intensity distribution to construct a noise sequence. In the denoising stage, the denoising model based on the conditional U-Net architecture uses its powerful feature extraction ability and attention mechanism to gradually remove the noise in the noise sequence, so that the latent vector is restored to a purified sample close to the original image. Subsequently, the purified latent vector is input into the decoder, and image features are reconstructed through deconvolution and upsampling operations and mapped back to the image space. Finally, the processed image is input into the deepfake detection model to complete the forgery detection task.

[0092] Figure 3 It is a specific application diagram of the high-robust deepfake detection method based on latent space adversarial purification in the present invention, and its detailed description process is as follows:

[0093] Step 1, taking the images to be detected suspected of forgery in the social network as the input, preprocessing the input to obtain continuous vectors suitable for model processing. These vectors are input into the deployed adversarial purifier based on the diffusion model, which specifically includes three key components: an encoder, a latent space adversarial purification model, and a decoder. Subsequently, the purified latent vector is input into the decoder to decode the latent vector into an image, obtaining a picture with adversarial noise removed, which helps the deepfake detector to identify.

[0094] Step 2, inputting the purified picture into the deepfake detector based on the ResNet structure, extracting features through the convolutional neural network, obtaining the confidence of the category through the fully connected layer and the softmax function. If the score of the confidence is higher than the threshold (for example, 0.5), it is determined as an abnormal sample, and if it is lower than the threshold, it is determined as a normal sample.

[0095] Embodiment 2

[0096] Based on the same inventive concept, this embodiment discloses a high-robust deepfake detection method and device based on latent space adversarial purification. Please refer to Figure 4 , including:

[0097] The data acquisition and preprocessing module 401 is used to collect the original image data, preprocess the collected image data, and divide it into training data and test data;

[0098] A training module 402 is used to train an adversarial purifier based on a diffusion model using training data. The adversarial purifier based on the diffusion model includes an autoencoder and a latent space adversarial purification model. The autoencoder includes an encoder and a decoder. The encoder is used to extract and compress features of image data. The latent space adversarial purification model includes a diffusion stage and a denoising stage. In the diffusion stage, Gaussian noise is gradually added to the latent vector according to a preset noise step size and intensity distribution to construct a noise sequence. In the denoising stage, the noise in the noise sequence is gradually removed to restore the latent vector to a purified sample close to the original image. The decoder is used to reconstruct image features and map them back to the image space.

[0099] An adversarial purification module 403 is used to input test data into the trained adversarial purifier based on the diffusion model to obtain an adversarially purified image.

[0100] A detection module 404 is used to input the adversarially purified image into a pre-constructed deepfake detection model to obtain a detection result.

[0101] Since the device introduced in the second embodiment of the present invention is the device used to implement the high-robust deepfake detection method based on latent space adversarial purification in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device used in the method of the first embodiment of the present invention falls within the scope of protection of the present invention.

[0102] Embodiment Three

[0103] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method described in Embodiment One.

[0104] Since the computer-readable storage medium introduced in the third embodiment of the present invention is the computer-readable storage medium used to implement the high-robust deepfake detection method based on latent space adversarial purification in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the computer-readable storage medium, so it will not be elaborated here. Any computer-readable storage medium used in the method of the first embodiment of the present invention falls within the scope of protection of the present invention.

[0105] Embodiment Four

[0106] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method described in Embodiment One.

[0107] Since the computer device introduced in the fourth embodiment of the present invention is the computer device adopted in the high-robust deepfake detection method based on latent space adversarial purification in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the computer device, so it will not be elaborated here. Any computer device adopted in the method of the first embodiment of the present invention falls within the scope of protection of the present invention.

[0108] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0109] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0110] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments and all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and variations.

Claims

1. A highly robust deepfake detection method based on latent space adversarial purification, characterized in that Including: Collecting original image data, preprocessing the collected image data, and dividing it into training data and test data; Training an adversarial purifier based on a diffusion model using the training data. The adversarial purifier based on the diffusion model includes an autoencoder and a latent space adversarial purification model. The autoencoder includes an encoder and a decoder. The encoder is used to extract and compress features from the image data. The latent space adversarial purification model includes a diffusion stage and a denoising stage. In the diffusion stage, Gaussian noise is gradually added to the latent vector according to a preset noise step size and intensity distribution to construct a noise sequence. In the denoising stage, the noise in the noise sequence is gradually removed to restore the latent vector to a purified sample close to the original image. The decoder is used to reconstruct the image features and map them back to the image space; Inputting the test data into the trained adversarial purifier based on the diffusion model to obtain an adversarially purified image; Inputting the adversarially purified image into a pre-constructed deepfake detection model to obtain a detection result.

2. The high-robust deepfake detection method based on latent space adversarial purification according to claim 1, wherein, Preprocessing the collected image data, including: Adjusting the size of the collected image and performing normalization processing.

3. The high-robust deepfake detection method based on latent space adversarial purification according to claim 1, wherein The method further includes: Introducing random rotation and noise addition techniques to enhance the training data.

4. The high-robust deepfake detection method based on latent space adversarial purification according to claim 1, wherein When training the adversarial purifier based on the diffusion model using the training data, the PGD adversarial attack algorithm is used to perturb the original image to generate adversarial samples. In the gradient calculation process, a transfer attack strategy is introduced, and the VGG-19 is used as a proxy model to obtain gradient direction information.

5. The high-robust deepfake detection method based on latent space adversarial purification according to claim 4, wherein Training the adversarial purifier based on the diffusion model using the training data, including: Input the original image and the adversarial sample into the latent space, and apply a perturbation to the latent vector z of the original image in the direction of the latent vector z of the adversarial sample in the latent space to obtain the perturbed latent vector z adv ; r ; Input the latent vector z of the adversarial example adv into the latent space adversarial purification model, and obtain the adversarially purified latent vector z through the operations of the diffusion stage and the denoising stage p ; The perturbed latent vector z r is passed through the decoder D to obtain the reconstructed image I rec , and in the training process, the following loss function value is minimized to obtain the optimal model parameters: Among them, z p represents the latent vector after adversarial purification, which is used to measure the difference between the latent vector z of the original image and the latent vector z p after adversarial purification. The mean squared error metric is adopted. q φ q(z|x) represents the posterior distribution parameterized by the encoder, and p(z) represents the prior distribution of the latent vector of the original image, which is used to constrain the distribution structure of the encoder output. D KL divergence is used to measure the difference between the posterior distribution q φ (z|x) and the prior distribution p(z), and is used as a regularization term during training to constrain the model to converge to the optimal value.

6. The high-robust deepfake detection method based on latent space adversarial purification according to claim 1, wherein, Inputting the test data into the trained adversarial purifier based on the diffusion model to obtain an adversarially purified image, including: Diffusion stage adversarial sample z adv After t steps of the forward diffusion process, a sample containing standard Gaussian noise and the original adversarial noise is obtained The formal representation is as follows: Among them, represents the latent vector of the adversarial sample after t steps of the diffusion process, represents the cumulative weight coefficient at the t-th step in the diffusion process, ∈ is standard Gaussian noise, and σ′ represents the perturbation; In the denoising stage, the denoising operation at each step is modeled by a parameterized conditional probability distribution p θ (z t-1 |z t ), and the specific form is as follows: Among them, z t represents the latent vector obtained after t steps of diffusion, and z t-1 represents the latent vector at step t-1 in the reverse diffusion process, which is predicted from z t by a denoising solver based on the U-Net and attention mechanism. The mean μ θ (z t , t) and the variance are predicted by the denoising solver.

7. The high-robust deepfake detection method based on latent space adversarial purification according to claim 1, wherein The pre-constructed deepfake detection model is a model based on the ResNet structure, including a convolutional neural network and an output layer. Inputting the adversarially purified image into the pre-constructed deepfake detection model to obtain a detection result, including: Extracting discriminative features from the adversarially purified image through the convolutional neural network; Outputting a confidence score based on the discriminative features through the output layer.

8. A highly robust deepfake detection method and device based on latent space adversarial purification, characterized in that, Including: A data collection and preprocessing module for collecting original image data, preprocessing the collected image data, and dividing it into training data and test data; A training module for training an adversarial purifier based on a diffusion model using the training data. The adversarial purifier based on the diffusion model includes an autoencoder and a latent space adversarial purification model. The autoencoder includes an encoder and a decoder. The encoder is used to extract and compress features from the image data. The latent space adversarial purification model includes a diffusion stage and a denoising stage. In the diffusion stage, Gaussian noise is gradually added to the latent vector according to a preset noise step size and intensity distribution to construct a noise sequence. In the denoising stage, the noise in the noise sequence is gradually removed to restore the latent vector to a purified sample close to the original image. The decoder is used to reconstruct the image features and map them back to the image space; An adversarial purification module for inputting the test data into the trained adversarial purifier based on the diffusion model to obtain an adversarially purified image; The detection module is used to input the adversarially purified image into a pre-constructed deepfake detection model to obtain a detection result.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the high-robust deepfake detection method based on latent space adversarial purification as described in any one of claims 1 to 7.

10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the high-robust deepfake detection method based on latent space adversarial purification as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Three-dimensional Gaussian splash appearance robustness enhancement method and system based on generative potential optimization

    CN121190365A

  • Face image depth forgery detection method and system

    CN121600582A