Face deep counterfeiting adaptive detection method with privacy efficiency

By a method of acquiring image frequency information in depth forgery detection, adaptive analysis and reconstruction of images, the problem of large overhead of user sensitive information leakage and calculation overhead in the prior art is solved, and efficient privacy protection and low overhead depth forgery detection are achieved.

CN119992620APending Publication Date: 2025-05-13CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510056358.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing deep forgery detection technology is prone to leaking user sensitive information during transmission and auditing, resulting in identity information security threats. The existing privacy protection methods are expensive to calculate and communicate, making it difficult to promote in large-scale application scenarios.

Method used

By acquiring image frequency information, using feature importance evaluation to perform adaptive analysis of frequency domain channels, setting weight thresholds, retaining key channels, reconstructing images, and inputting protected images to detect models for real and fake face classification.

Benefits of technology

It realizes deep forgery detection without exposing the original image, reduces computing and communication overhead, and improves the privacy protection efficiency and practicality of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992620A_ABST
    Figure CN119992620A_ABST
Patent Text Reader

Abstract

The invention relates to a face deep counterfeiting adaptive detection method with privacy efficiency, and belongs to the field of computer vision and image security. The method comprises three main stages: a frequency domain channel analysis stage, a reconstruction protection stage and a forgery detection stage. And a frequency domain channel analysis stage: acquiring image frequency information, and analyzing the frequency domain channel by using feature importance evaluation to obtain the contribution degree of each channel to forgery detection. And a reconstruction protection stage: setting a weight threshold, retaining a key channel, and reconstructing an image. And a forgery detection stage: inputting the protected image into the detection model to realize real and forgery face classification. According to the method, the channel selection process is optimized through feature importance evaluation, identity related information is hidden on the original image, and the detection model still has the forgery detection capability when coping with the protected image. The method is especially suitable for a deep counterfeiting detection method based on a frequency domain, the detection accuracy can be effectively improved, and meanwhile the privacy protection requirement of a user can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for adaptively detecting deep fake human faces with privacy performance, and belongs to the fields of computer vision and image security. Background Art

[0002] The rapid development of deep fake technology has made it possible to generate highly realistic fake facial images. This technology uses deep neural networks to generate fake facial images that are highly similar to real images, greatly improving the efficiency and realism of virtual character or film and television special effects production, but it has also caused many security risks at the social level. Criminals use this technology to spread false information, defraud and other illegal activities, which poses a serious threat to personal privacy and social information security. Therefore, in order to cope with the security threats brought by highly realistic deep fake facial images, related fields have successively carried out research on deep fake detection technology.

[0003] Deep fake detection usually requires the collection and analysis of a large number of users' facial images, but in this process, all original images will be directly exposed to data reviewers and detection models. The resulting problem is that user sensitive information is easily leaked during transmission and review, and identity information is also facing potential security threats. Users do not want to leak fake images involving themselves to others for adverse dissemination. For this reason, it is particularly urgent and important to design an effective image identity protection mechanism that can effectively protect the privacy of images while performing fake detection, thereby avoiding direct exposure of original facial images. This can not only protect the user's personal privacy, but also enhance the credibility of deep fake detection models and their applications. At the privacy protection level, existing technologies have proposed a variety of solutions that combine cryptography or secure multi-party computing to protect images from being leaked during the detection process. For example, the paper "Privacy-preserving DeepFake face image detection" builds a secure communication protocol based on the additive secret sharing framework, and completes the collaborative calculation of the pre-trained model through non-colluding dual servers, realizing forged face detection without exposing the input original image; the paper "PP-DFD:APrivacy-Preserving Deepfake Detection" adopts the homomorphic encryption method to map the image processing and detection process into the encryption domain, and ensures that the image privacy is protected without affecting the detection accuracy. This scheme can obtain similar detection accuracy under both plaintext and encrypted conditions, and performs forgery detection while maintaining the privacy of the image. Although the above methods achieve privacy protection for deep fake detection through technical means such as secure multi-party computing, homomorphic encryption, and secret sharing, there are still some shortcomings. The operations of homomorphic encryption, secret sharing, and secure multi-party computing often require a large number of complex operations in the encryption domain, resulting in a significant increase in the computational overhead and execution time of the detection process. In addition, it is necessary to frequently exchange intermediate ciphertexts or keys between servers or between clients and servers, resulting in a large amount of communication data and high implementation and deployment costs. When the network size or the number of participants increases, the secure communication protocols required by these methods will become increasingly complex, which is not conducive to promotion in large-scale application scenarios and has low practicality. In order to further improve the privacy protection efficiency of deep fake detection, a method for directly protecting the original image is urgently needed to reduce the risk of privacy leakage that may occur during subsequent transmission and detection, and minimize the computational burden and communication overhead. Summary of the invention

[0004] In order to further improve the privacy protection efficiency of deep fake detection and enhance its practicality, the present invention proposes a face deep fake adaptive detection method with privacy effectiveness, which includes the following steps:

[0005] S1: Obtain image frequency information, use feature importance evaluation to adaptively analyze frequency domain channels, and obtain the contribution of each channel to forgery detection;

[0006] S2: Set weight threshold, retain key channels, and reconstruct the image;

[0007] S3: Input the protected image into the detection model to classify real and fake faces.

[0008] Furthermore, the frequency channel feature importance evaluation analysis specifically includes the following steps:

[0009] S1.1: The image to be detected is converted to the frequency domain through fast Fourier transform (FFT) to obtain the frequency representation of the image, and its mathematical expression is as follows:

[0010]

[0011] In the formula, X represents the input image to be detected, is the fast Fourier transform operation, F(X) is the frequency domain image, Represents the complex domain, C, H, and W are the number of channels, height, and width of the frequency domain image respectively;

[0012] S1.2: Divide the frequency characteristics into multiple channels, and its mathematical expression is as follows:

[0013]

[0014] Where r represents the radius distance of the frequency amplitude; r i is the frequency radius threshold used to divide the channel and determine the boundary of each channel; N represents the number of divided channels; channel C i Contains the frequency features within the i-th frequency channel. In this way, the frequency domain image can be divided into different channels such as low frequency, medium frequency, and high frequency, which is suitable for subsequent dynamic weighted selection;

[0015] S1.3: For each divided channel C i , initialize the weight parameter w i ;

[0016] S1.4: Gumbel Softmax is used to learn the weight parameters of each channel. During the training process, the loss function of the forgery detection task is fed back to dynamically evaluate the contribution of each channel to the forgery detection task. The mathematical expressions involved are as follows:

[0017]

[0018] In the formula, p i is the weight probability of the i-th channel, g iis the Gumbel noise term, τ is the temperature parameter, which controls the randomness and sparsity of the selection;

[0019]

[0020] Where, f(·) represents the forgery detection model; Represents the predicted probability of the mth sample;

[0021]

[0022] In the formula, Represents the loss function value, and the cross entropy loss needs to be minimized during the training process; y m Represents the true label value of the mth sample (1 represents a real image, 0 represents a fake image).

[0023] Furthermore, the steps of learning the weight parameters of each channel using Gumbel Softmax include:

[0024] Add Gumbel noise: First add Gumbel noise g to the logits i , making it random, its mathematical expression is as follows;

[0025]

[0026] Where U is a random variable sampled from uniform distribution U(0,1);

[0027] Apply Softmax: Then apply the Softmax function, and control the smoothness of the output through the temperature parameter τ. When τ approaches infinity, the output is close to a uniform distribution, and when τ approaches 0, the output is close to a one-hot vector, but it is still continuous, so that gradient calculation can be performed.

[0028] Furthermore, the reconstruction protection method specifically comprises the following steps:

[0029] S2.1: Set the channel weight threshold T, retain the model with the highest training accuracy in the analysis phase, obtain and sort the channel weight parameters of the model for the image in the frequency domain, and divide according to the threshold T;

[0030] S2.2: Sort the weight parameters from high to low, set the channels beyond the T range to zero, and remove the frequency information that contributes little to the detection task. The mathematical expression is as follows:

[0031]

[0032] In the formula, F selected (X) represents the frequency domain image retained after screening, is the indicator function, when the weight p of the i-th channel i When it belongs to the first T channels with the highest weight, the value is 1, otherwise it is 0. Its mathematical expression is as follows:

[0033]

[0034] S2.3: Restore the filtered frequency map to the spatial domain through inverse Fourier transform (IFFT) to obtain the reconstructed image Its mathematical expression is as follows:

[0035]

[0036] In the formula, represents inverse Fourier transform;

[0037] S2.4: Replace the brightness channel of the reconstructed image with the brightness channel of the original image through YCrCb color space mapping to ensure that the image is visually consistent with the original image and obtain the final protected image

[0038] Furthermore, the method for adaptively detecting deep fake faces with privacy protection specifically comprises the following steps:

[0039] S3.1: Build a pre-trained forgery detection model f(·);

[0040] S3.2: Protected Image Input into the forgery detection model to extract features of the input image;

[0041] S3.3: Use a binary classifier to classify the extracted features and output the predicted probability P of the image authenticity.

[0042] Furthermore, the forgery detection model is a deep neural network model, including but not limited to Xception, ResNet, EfficientNet and other network structures suitable for image binary classification tasks.

[0043] Compared with the prior art, the present invention has the following advantages:

[0044] (1) Source privacy protection: The present invention directly removes sensitive facial features from the original image to eliminate the risk of leakage in subsequent transmission links.

[0045] (2) Low computational and communication overhead: No tedious encryption and decryption and multi-party collaborative operations are required, which greatly reduces the computational and communication volume of the detection process.

[0046] (3) High scalability and flexibility: No need to customize hardware or make large-scale changes to existing detection models, and it can be easily integrated into a variety of detection pipelines.

[0047] (4) Taking into account detection accuracy: Only the frequency domain channels that are most critical for counterfeit detection are screened, which can both hide identity information and retain the model detection performance.

[0048] (5) The method is simple and easy to deploy: The process is intuitive and easy to implement, making it easy to quickly implement and promote in various projects. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a flow chart of a method for adaptively detecting deep fake faces with privacy performance according to the present invention. DETAILED DESCRIPTION

[0050] The following will be combined with the accompanying drawings in the embodiments of the present invention to describe the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention based on the content of this specification. The various details of this specification can be appropriately modified or changed according to different viewpoints and needs without violating the spirit of the present invention. It should be pointed out that the illustrations shown in the following embodiments are for illustration purposes only to illustrate the basic concept of the present invention. If there is no conflict, these embodiments and the features therein can be used in combination with each other.

[0051] See also Figure 1 The present invention proposes a privacy-effective adaptive face deep forgery detection method, which specifically includes the following steps:

[0052] S1: Obtain image frequency information, use feature importance evaluation to adaptively analyze frequency domain channels, and obtain the contribution of each channel to forgery detection;

[0053] S2: Set weight threshold, retain key channels, and reconstruct the image;

[0054] S3: Input the protected image into the detection model to classify real and fake faces.

[0055] Furthermore, the frequency channel feature importance evaluation analysis specifically includes the following steps:

[0056] S1.1: The image to be detected is converted to the frequency domain through fast Fourier transform (FFT) to obtain the frequency representation of the image, and its mathematical expression is as follows:

[0057]

[0058] In the formula, X represents the input image to be detected, is the fast Fourier transform operation, F(X) is the frequency domain image, Represents the complex domain, C, H, and W are the number of channels, height, and width of the frequency domain image respectively;

[0059] S1.2: Divide the frequency characteristics into multiple channels, and its mathematical expression is as follows:

[0060]

[0061] Where r represents the radius distance of the frequency amplitude; r i is the frequency radius threshold used to divide the channel and determine the boundary of each channel; N represents the number of divided channels; channel C i Contains the frequency features within the i-th frequency channel. In this way, the frequency domain image can be divided into different channels such as low frequency, medium frequency, and high frequency, which is suitable for subsequent dynamic weighted selection;

[0062] S1.3: For each divided channel C i , initialize the weight parameter w i ;

[0063] S1.4: Gumbel Softmax is used to learn the weight parameters of each channel. During the training process, the loss function of the forgery detection task is fed back to dynamically evaluate the contribution of each channel to the forgery detection task. The mathematical expressions involved are as follows:

[0064]

[0065] In the formula, p i is the weight probability of the i-th channel, g i is the Gumbel noise term, τ is the temperature parameter, which controls the randomness and sparsity of the selection;

[0066]

[0067] Where, f(·) represents the forgery detection model; Represents the predicted probability of the mth sample;

[0068]

[0069] In the formula, Represents the loss function value, and the cross entropy loss needs to be minimized during the training process; y m Represents the true label value of the mth sample (1 represents a real image, 0 represents a fake image).

[0070] Furthermore, the steps of learning the weight parameters of each channel using Gumbel Softmax include:

[0071] Add Gumbel noise: First add Gumbel noise g to the logits i , making it random, its mathematical expression is as follows;

[0072]

[0073] Where U is a random variable sampled from uniform distribution U(0,1);

[0074] Apply Softmax: Then apply the Softmax function, and control the smoothness of the output through the temperature parameter τ. When τ approaches infinity, the output is close to a uniform distribution, and when τ approaches 0, the output is close to a one-hot vector, but it is still continuous, so that gradient calculation can be performed.

[0075] Furthermore, the reconstruction protection method specifically comprises the following steps:

[0076] S2.1: Set the channel weight threshold T, retain the model with the highest training accuracy in the analysis phase, obtain and sort the channel weight parameters of the model for the image in the frequency domain, and divide according to the threshold T;

[0077] S2.2: Sort the weight parameters from high to low, set the channels beyond the T range to zero, and remove the frequency information that contributes little to the detection task. The mathematical expression is as follows:

[0078]

[0079] In the formula, F selected (X) represents the frequency domain image retained after screening, is the indicator function, when the weight p of the i-th channel i When it belongs to the first T channels with the highest weight, the value is 1, otherwise it is 0. Its mathematical expression is as follows:

[0080]

[0081] S2.3: Restore the filtered frequency map to the spatial domain through inverse Fourier transform (IFFT) to obtain the reconstructed image Its mathematical expression is as follows:

[0082]

[0083] In the formula, represents inverse Fourier transform;

[0084] S2.4: Replace the brightness channel of the reconstructed image with the brightness channel of the original image through YCrCb color space mapping to ensure that the image is visually consistent with the original image and obtain the final protected image

[0085] Furthermore, the method for adaptively detecting deep fake faces with privacy protection specifically comprises the following steps:

[0086] S3.1: Build a pre-trained forgery detection model f(·);

[0087] S3.2: Protected Image Input into the forgery detection model to extract features of the input image;

[0088] S3.3: Use a binary classifier to classify the extracted features and output the predicted probability P of the image authenticity.

[0089] Furthermore, the forgery detection model is a deep neural network model, including but not limited to Xception, ResNet, EfficientNet and other network structures suitable for image binary classification tasks.

[0090] Although the embodiments of the present invention have been shown and described, it is possible for those skilled in the art to make various changes, modifications, substitutions or deformations to these embodiments without violating the principles and spirit of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A privacy-effective adaptive face deepfake detection method, characterized in that: The method mainly includes the following steps: S1: Obtain image frequency information, use feature importance evaluation to adaptively analyze frequency domain channels, and obtain the contribution of each channel to forgery detection; S2: Set weight threshold, retain key channels, and reconstruct the image; S3: Input the protected image into the detection model to classify real and fake faces.

2. The method for self-adaptive face deepfake detection with privacy protection according to claim 1, characterized in that: Step S1 specifically includes the following steps: S1.1: The image to be detected is converted to the frequency domain through fast Fourier transform (FFT) to obtain the frequency representation of the image, and its mathematical expression is as follows: In the formula, X represents the input image to be detected, is the fast Fourier transform operation, F(X) is the frequency domain image, Represents the complex domain, C, H, and W are the number of channels, height, and width of the frequency domain image respectively; S1.2: Divide the frequency characteristics into multiple channels, and its mathematical expression is as follows: Where r represents the radius distance of the frequency amplitude; i is the frequency radius threshold used to divide the channel and determine the boundary of each channel; N represents the number of divided channels; channel C i Contains the frequency features within the i-th frequency channel. In this way, the frequency domain image can be divided into different channels such as low frequency, medium frequency, and high frequency, which is suitable for subsequent dynamic weighted selection; S1.3: For each divided channel C i , initialize the weight parameter w i ; S1.4: Gumbel Softmax is used to learn the weight parameters of each channel. During the training process, the loss function of the forgery detection task is fed back to dynamically evaluate the contribution of each channel to the forgery detection task. The mathematical expressions involved are as follows: In the formula, p i is the weight probability of the i-th channel, g i is the Gumbel noise term, τ is the temperature parameter, which controls the randomness and sparsity of the selection; Where f(·) represents the forgery detection model; Represents the predicted probability of the mth sample; In the formula, Represents the loss function value, and the cross entropy loss needs to be minimized during the training process; y m Represents the true label value of the mth sample (1 represents a real image, 0 represents a fake image).

3. The method of using Gumbel Softmax to learn the weight parameters of each channel according to claim 2, characterized in that: Gumbel-Softmax is a method for implementing differentiable discrete selection in neural networks. It makes the discrete selection process differentiable by introducing a combination of Gumbel noise and Softmax function, allowing gradients to backpropagate through the selection process. This enables the model to automatically select or weight different frequency channels during training. The steps of Gumbel Softmax are as follows: 1) Add Gumbel noise: First add Gumbel noise g to the logits i , making it random, its mathematical expression is as follows; Where U is a random variable sampled from uniform distribution U(0,1); 2) Apply Softmax: Then apply the Softmax function, and control the smoothness of the output through the temperature parameter τ. When τ approaches infinity, the output is close to a uniform distribution, and when τ approaches 0, the output is close to a one-hot vector, but it is still continuous, so that gradient calculation can be performed.

4. The method for self-adaptive face deep fake detection with privacy protection according to claim 1, characterized in that: Step S2 specifically includes the following steps: S2.1: Set the channel weight threshold T, retain the model with the highest training accuracy in the analysis phase, obtain and sort the channel weight parameters of the model for the image in the frequency domain, and divide according to the threshold T; S2.2: Sort the weight parameters from high to low, set the channels beyond the T range to zero, and remove the frequency information that contributes little to the detection task. The mathematical expression is as follows: In the formula, F selected (X) represents the frequency domain image retained after screening, Indicates that when the weight p of the i-th channel i When it belongs to the first T channels with the highest weight, the value is 1, otherwise it is 0. Its mathematical expression is as follows: S2.3: Restore the filtered frequency map to the spatial domain through inverse Fourier transform (IFFT) to obtain the reconstructed image Its mathematical expression is as follows: In the formula, represents inverse Fourier transform; S2.4: Replace the brightness channel of the reconstructed image with the brightness channel of the original image through YCrCb color space mapping to ensure that the image is visually consistent with the original image and obtain the final protected image 5. The method for self-adaptive face deepfake detection with privacy protection according to claim 1, characterized in that: Step S3 specifically includes the following steps: S3.1: Build a pre-trained forgery detection model f(·); S3.2: Protected Image Input into the forgery detection model to extract features of the input image; S3.3: Use a binary classifier to classify the extracted features and output the predicted probability P of the image authenticity.

6. The forgery detection model according to claim 5, characterized in that: The forgery detection model is a deep neural network model, including but not limited to Xception, ResNet, EfficientNet and other network structures suitable for image binary classification tasks. The forgery detection model that focuses on classification based on frequency domain features has lower detection accuracy than other detection models.