Non-supervision face forgery detection method and system based on frequency domain supplementation

By adopting an unsupervised detection method based on frequency domain supplementation in face forgery detection, using the frequency domain enhanced variational autoencoder and unsupervised detection network, the fused multi-scale features and calculate anomaly score map is solved, and the problem of difficulty in capturing image details and global information in the existing technology is solved, and high-precision and robust forgery detection is achieved.

CN120183052AActive Publication Date: 2025-06-20GOLDEN TIMES CULTURE COMM

Patent Information

Application Number
CN202510653184.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

When faced with diverse and complex forgery technologies, existing face forgery detection technology is difficult to fully capture the details and global information of the image, resulting in low recognition accuracy of forgery traces.

Method used

Unsupervised face forgery detection method based on frequency domain supplementation is adopted, multi-level encoding features are extracted through ResNet reference network and frequency domain enhancement variational autoencoder, frequency domain enhancement processing is performed, feature reconstruction is carried out in combination with dynamic spectrum loss mechanism, and global and local pooling is carried out through unsupervised detection network to generate fused multi-scale features, and finally calculate the abnormal score graph for forgery detection.

Benefits of technology

It effectively improves the accuracy and robustness of face forgery detection, and can identify complex forgery images under unsupervised conditions, with wide application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183052A_ABST
    Figure CN120183052A_ABST
Patent Text Reader

Abstract

The invention relates to an unsupervised face forgery detection method and system based on frequency domain supplementation, and belongs to the technical field of computer vision. The detection method comprises the following steps: inputting a face image to be detected into a reference network based on ResNet, and extracting multi-level coding features through an encoder of a frequency domain enhanced variational auto-encoder; performing frequency domain enhancement processing on the multi-level coding features, generating high-frequency enhancement features, inputting the high-frequency enhancement features to a decoder of the frequency domain enhancement variational auto-encoder, performing feature reconstruction in combination with a dynamic spectrum loss mechanism, and generating reconstructed frequency domain supplementary features; global pooling and local pooling are carried out on the multi-level coding features, global representation and local representation are generated and connected along a channel axis, and fused multi-scale features are generated; calculating an abnormal score graph based on the fused multi-scale features and the reconstructed frequency domain supplementary features; and determining an image authenticity result of the to-be-detected face image according to the pixel-level score of the abnormal score graph. According to the invention, the accuracy of face counterfeiting detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to an unsupervised face forgery detection method and system based on frequency domain supplementation. Background Art

[0002] With the continuous development of deepfake technology, the quality of forged images and videos has been continuously improved, resulting in increasing challenges for face forgery detection technology. Face forgery detection technology aims to identify and detect whether a face image has been tampered with or forged, and has important application values in the fields of digital image processing, social security, and legal evidence. In recent years, with the rise of technologies such as generative adversarial networks (GANs) and deepfake videos, the authenticity of forged images has approached or even exceeded that of real images, which poses a serious threat to privacy protection, trust mechanisms, and social security. Therefore, accurate and efficient forgery detection technology is crucial for maintaining social order and protecting personal privacy.

[0003] Common face forgery detection methods can generally be divided into three categories: detection methods based on image patch features, detection methods based on deep learning, and detection methods based on frequency domain feature enhancement. Among them, the detection methods based on image patch features mainly extract local features from the image, analyze and compare the detailed parts in the image to detect forgery traces. However, due to the continuous upgrading of forgery technology, traditional image patch-based feature extraction methods often perform poorly in terms of robustness and adaptability, and are difficult to cope with complex forgery methods. Detection methods based on deep learning train a deep neural network (DNN) model to automatically learn the complex features of the image, which can improve the forgery detection performance to a certain extent. However, such methods usually require a large amount of labeled data for training and have a large computational overhead, making them unsuitable for real-time applications. In recent years, detection methods based on frequency domain feature enhancement have gradually received attention. By enhancing the high-frequency information of the image and suppressing the noise in the forged image, the accuracy and generalization ability of detection can be effectively improved.

[0004] Currently, most common detection methods are based on the extraction and comparison of local features. However, when faced with diverse and complex forgery technologies, they often cannot fully capture the details and global information in the image, resulting in a low recognition accuracy of forgery traces. In particular, the subtle local forgery traces in forged images are often ignored in traditional methods, or there are great difficulties in fusing local features and global features, resulting in the inability to achieve effective synchronous modeling. Moreover, due to the diversity of forged images and the lack of labeled data, existing models are prone to training biases. Especially when faced with images generated by different forgery technologies, it is often difficult to maintain stable detection performance, reducing the accuracy of forgery detection. Summary of the Invention

[0005] To improve the accuracy of face forgery detection, this application provides an unsupervised face forgery detection method and system based on frequency domain supplementation.

[0006] In the first aspect, this application provides an unsupervised face forgery detection method based on frequency domain supplementation, adopting the following technical solutions: An unsupervised face forgery detection method based on frequency domain supplementation, the detection method includes: Input the face image to be detected into a benchmark network based on ResNet, and extract multi-level encoded features through the encoder of the frequency domain enhanced variational autoencoder; Perform frequency domain enhancement processing on the multi-level encoded features to generate high-frequency enhanced features; Input the high-frequency enhanced features into the decoder of the frequency domain enhanced variational autoencoder, and perform feature reconstruction in combination with the dynamic spectrum loss mechanism to generate the reconstructed frequency domain supplemented features; Perform global pooling and local pooling on the multi-level encoded features through an unsupervised detection network to generate a global representation and a local representation, and connect the global representation and the local representation along the channel axis to generate a fused multi-scale feature; Calculate an anomaly score map based on the fused multi-scale feature and the reconstructed frequency domain supplemented feature; Determine the authenticity result of the face image to be detected according to the pixel-level scores of the anomaly score map.

[0007] By adopting the above technical solutions, the frequency domain enhanced variational autoencoder, the unsupervised detection network and the frequency domain supplemented features are introduced, and the global and local features of the image are comprehensively utilized, which can effectively detect forgery traces in face images under unsupervised conditions. By enhancing the image details through frequency domain enhancement and using multi-scale feature fusion to enhance the recognition ability of forged images, and finally making pixel-level judgments through the anomaly score map, the high accuracy and robustness of the detection are ensured. This technical solution can effectively identify complex forged images and has broad application prospects.

[0008] Optionally, the step of performing frequency domain enhancement processing on the multi-level encoded features to generate high-frequency enhanced features includes: Perform discrete cosine transform on each layer of encoded features of the multi-level encoded features to convert them into frequency domain features; Perform low-pass filtering weighting on the low-frequency components of the frequency domain features, and perform adaptive weight assignment on the high-frequency components of the frequency domain features; Convert the weighted frequency domain features to the spatial domain through inverse discrete cosine transform to obtain high-frequency enhanced features.

[0009] By adopting the above technical solution, the image features are transformed from the spatial domain to the frequency domain, and then through low-pass filtering and adaptive weight assignment, the high-frequency details are emphasized, thereby enhancing the expressiveness of forgery traces. Finally, through the inverse discrete cosine transform (IDCT), the enhanced frequency-domain features are transformed back to the spatial domain to generate high-frequency enhanced features. This process not only improves the detail information of the image, especially in the case where the forgery traces are relatively weak, but also makes the subtle changes in the forged image more significant through high-frequency enhancement, providing a more reliable basis for subsequent forgery detection.

[0010] Optionally, the steps of constructing the frequency-domain enhanced variational autoencoder include: Input the input feature x into the encoder E of the pre-constructed variational autoencoder for convolutional sampling, and map it to the latent space z; Quantize the features in the latent space z, and generate a quantized latent embedding through formula (1) : ; (1) where the embedding space C = {c1, c2,..., c∣C∣} is a preset discrete vector set, z ij is the j-th feature of the i-th layer in the latent space z, and c k represents the nearest vector in the embedding space; Input the quantized latent embedding into the decoder of the variational autoencoder G for layer-by-layer reconstruction to generate a preliminary reconstructed image , and the calculation process is: ; ... (2) ; ; (3) In the above formula (2), represents the i-th layer of the encoder, represents the feature before activation of the i-th layer of the encoder, represents the frequency supplement feature of the i-th layer, is used to extract the frequency-rich features learned from the feature before activation of the encoder; in the above formula (3), represents the corresponding function of the i-th layer of the decoder, represents the output feature of the decoder; Perform frequency-domain processing on the feature before activation of the encoder and the output feature of the decoder to generate a frequency supplement feature ; Perform frequency-domain processing on the frequency supplement feature Output features of the decoder Perform a discrete Fourier transform to calculate the focus frequency loss FFL and the dynamic spectrum loss SL. The specific calculation formulas are as follows: ; (4) ; (5) ; (6) ; (7) In the above formula (4), e and i represent the Euler number and the imaginary unit respectively, M×N is the spatial resolution of the feature map, f(x,y) is the value at (x,y) of each feature map, and F(u,v) is the corresponding value at the (u,v) coordinate on the spectrum; in the above formula (5), , representing the weight applied to each frequency; , representing the main error function based on the frequency domain; in the above formula (6), represents the Gaussian kernel function, μ is the mean, σ i is the variance, * represents the convolution operation, and are the weighted features; Combine the focus frequency loss FFL, the dynamic spectrum loss SL, the pixel-level reconstruction loss, and the perceptual loss to calculate the total reconstruction loss L rec , and the specific calculation formula is as follows: ; (8) In the above formula (8), α and β are hyperparameters, is the norm pixel-level reconstruction loss, is the perceptual loss; Based on the total reconstruction loss L rec Optimize the model parameters to obtain a trained frequency-domain enhanced variational autoencoder.

[0011] By adopting the above technical solutions, optimization mechanisms such as frequency-domain processing and focus frequency loss (FFL), dynamic spectrum loss (SL), etc. are introduced, effectively improving the reconstruction quality of the variational autoencoder (VAE). By performing frequency-domain enhancement on the features of the encoder and decoder, not only the traditional pixel-level reconstruction error is considered, but also the error constraint at the frequency level is introduced, enhancing the model's learning ability for the frequency-domain features of images. The combination of quantization of the latent embedding and frequency-domain features ensures the retention of image details and reconstruction accuracy. At the same time, the use of focus frequency loss and dynamic spectrum loss effectively reduces the reconstruction error in the low-frequency and high-frequency parts. Finally, by comprehensively optimizing the total reconstruction loss, higher-quality and more realistic image reconstruction results can be obtained, improving the model's expressiveness and reconstruction effect, especially in terms of high-frequency details and visual perception.

[0012] Optionally, the global pooling and local pooling are performed on the multi-level encoded features through an unsupervised detection network to generate a global representation and a local representation, and the global representation and the local representation are concatenated along the channel axis. The calculation formula for generating the fused multi-scale features includes: Performing global average pooling on the input feature x to generate a global representation is: ; (9) In the above formula (9), Pooling represents the global pooling operation; Dividing the input feature x into q×q regions, and performing local average pooling on each region to generate a local representation is: ; (10) In the above formula (10), Pooling represents the local pooling operation; Concatenating the global representation with the local representation along the channel axis, and inputting it into the convolutional network ψ (⋅), to obtain the fused multi-scale feature is: ; (11) In the above formula (11), is the fused multi-scale feature, ψ (⋅) is the convolutional network, representing the concatenation of the global representation and the local representation along the channel axis.

[0013] By adopting the above technical solution, global pooling and local pooling are performed on the input features based on an unsupervised detection network, and the global representation and the local representation are combined to generate fused multi-scale features. Global average pooling retains the overall information of the input features, while local average pooling captures the local detailed features. After concatenating the global representation and the local representation along the channel axis, the global and local information can be effectively combined to fully express the multi-scale features. This method improves the network's perception ability of different scale information, thereby enhancing the network's feature extraction and discrimination ability. Especially when dealing with tasks with complex details and multi-level structures, the model's expressiveness and detection accuracy can be significantly improved.

[0014] Optionally, the steps for calculating the anomaly score map based on the fused multi-scale features and the reconstructed frequency domain supplementary features include: Calculating the mean square error and cosine similarity between the fused multi-scale feature and the reconstructed frequency domain supplementary feature to obtain the pixel similarity is: ; (12) In the above formula (12), h and w represent the coordinate indices of the spatial features, is the indicator function, and are the mean squared error and cosine similarity respectively, and are hyperparameters, k represents the threshold for determining the similarity metric MSE; According to the pixel similarity the reconstructed feature is: ; (13) According to the input feature x and the reconstructed feature, the mean squared error loss is calculated as: ; (14) In the above formula (14), C, H, and W are the number of channels, height, and width of the feature map respectively, represents the Euclidean distance; According to the reconstructed feature the anomaly score map d is obtained, and the specific formula includes: ; (15) In the above formula (15), h and w represent the positions of each pixel.

[0015] By adopting the above technical solution, based on the fusion of multi-scale features and the reconstructed frequency-domain supplementary features, the pixel similarity is calculated and the anomaly score map is generated, thereby effectively detecting the anomaly region. First, by calculating the mean squared error and cosine similarity between the fused feature and the frequency-domain supplementary feature, the similarity between pixels can be accurately measured, and then the reconstructed feature can be obtained. This process can capture the differences between the input feature and the reconstructed feature, and use these differences to calculate the mean squared error loss, thereby generating the anomaly score map for each pixel position. This technical solution can accurately identify the anomaly region in the image. By combining multi-scale features and frequency-domain information, the anomaly detection is made more refined and robust, especially suitable for processing images with complex structures and noises.

[0016] Optionally, the steps of determining the authenticity result of the face image to be detected according to the pixel-level scores of the anomaly score map include: Interpolate the anomaly score map to the original size of the input face image to be detected; Perform global max pooling on the interpolated anomaly score map, and extract the maximum anomaly score in the anomaly score map as the final anomaly score; Based on the comparison between the preset threshold and the final anomaly score, determine the authenticity result of the face image to be detected.

[0017] By adopting the above technical solution, the abnormal score map is aligned with the input image, enabling the abnormal scores to be accurately compared at the pixel level. Through global max pooling, the most prominent forgery traces in the image are extracted, ensuring the accuracy of forgery detection. Finally, by comparing with a preset threshold, the authenticity result of the image is determined. This technical solution can accurately and effectively detect the forged parts in the image by combining the abnormal score map, pooling operation, and threshold judgment, with high detection accuracy and good robustness.

[0018] Optionally, the detection method further includes: Combining the total reconstruction loss L rec with the mean square error loss to obtain a total objective function as : ; (16) In the above formula (16), and are preset hyperparameters that respectively control the influence of the total reconstruction loss and the mean square error loss on the total objective function.

[0019] By adopting the above technical solution, the weighted combination of the total reconstruction loss and the mean square error loss enables the method to take into account both global and local features in forgery detection, and can more accurately capture the differences between forged images and real images. Through this combination, the detection method can make fine distinctions when dealing with different types of forged images, ensuring more accurate detection results.

[0020] In a second aspect, the present application provides an unsupervised face forgery detection system based on frequency domain supplementation, adopting the following technical solution: An unsupervised face forgery detection system based on frequency domain supplementation, the detection system includes: A feature extraction module, configured to input the face image to be detected into a benchmark network based on ResNet, and extract multi-level encoded features through the encoder of the frequency domain enhanced variational autoencoder; A frequency domain enhancement module, configured to perform frequency domain enhancement processing on the multi-level encoded features to generate high-frequency enhanced features; A reconstruction module, configured to input the high-frequency enhanced features into the decoder of the frequency domain enhanced variational autoencoder, and perform feature reconstruction in combination with the dynamic spectrum loss mechanism to generate reconstructed frequency domain supplemented features; A multi-scale feature fusion module, configured to perform global pooling and local pooling on the multi-level encoded features through an unsupervised detection network to generate a global representation and a local representation, and connect the global representation and the local representation along the channel axis to generate fused multi-scale features; An abnormal score map calculation module, configured to calculate an abnormal score map based on the fused multi-scale features and the reconstructed frequency-domain supplementary features; An authenticity verification module, configured to determine the authenticity result of the face image to be detected according to the pixel-level scores of the abnormal score map.

[0021] In a third aspect, the present application provides a computer device, adopting the following technical solution: A computer device includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method as described in the first aspect.

[0022] In a fourth aspect, the present application provides a computer-readable storage medium, adopting the following technical solution: A computer-readable storage medium stores a computer program that can be loaded and executed by a processor to implement any one of the methods in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 FIG. is a first flowchart of an unsupervised face forgery detection method according to an embodiment of the present application.

[0024] Figure 2 FIG. is a schematic diagram of the overall network model structure of an unsupervised face forgery detection method according to an embodiment of the present application.

[0025] Figure 3 FIG. is a second flowchart of an unsupervised face forgery detection method according to an embodiment of the present application.

[0026] Figure 4 FIG. is a third flowchart of an unsupervised face forgery detection method according to an embodiment of the present application.

[0027] Figure 5 FIG. is a schematic diagram of a frequency-domain supplementary module according to an embodiment of the present application.

[0028] Figure 6 FIG. is a schematic diagram of a dynamic frequency-domain loss according to an embodiment of the present application.

[0029] Figure 7 FIG. is a fourth flowchart of an unsupervised face forgery detection method according to an embodiment of the present application.

[0030] Figure 8 FIG. is a fifth flowchart of an unsupervised face forgery detection method according to an embodiment of the present application.

[0031] Figure 9 FIG. is a sixth flowchart of an unsupervised face forgery detection method according to an embodiment of the present application. Detailed implementation manners

[0032] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the following further describes the present application in detail with reference to the accompanying Figures 1 - 9 drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0033] An embodiment of the present application discloses an unsupervised face forgery detection method based on frequency domain supplementation.

[0034] Referring to Figure 1 and Figure 2 , an unsupervised face forgery detection method based on frequency domain supplementation, the detection method includes: Step S101: Input the face image to be detected into a benchmark network based on ResNet, and extract multi-level encoded features through the encoder of the frequency domain enhanced variational autoencoder; Among them, ResNet (Residual Network) is a deep convolutional neural network structure, which can effectively alleviate the problem of gradient disappearance in deep networks by introducing residual connections, and can retain information in deeper networks, so as to extract high-level abstract features. The input face image to be detected can extract global and local feature information through the ResNet network.

[0035] In some embodiments, the variational autoencoder (VAE) is a generative model that compresses image data into a distribution in the latent space through an encoder and then restores the original data through a decoder. The frequency domain enhanced variational autoencoder (FA-VAE) is used to transform image features into the frequency domain, which can capture more delicate frequency information, which is important for forgery detection, especially for the high-frequency information changes commonly found in face forgery.

[0036] Exemplarily, assuming that the input image is a standard face image, after being extracted by the ResNet network, the features of the image will be encoded into a low-dimensional latent space, which includes various levels of features of the face, such as abstract representations of eyes, nose, mouth, etc. The multi-level features extracted by ResNet can effectively capture the global and local information of the face image, while the frequency domain enhanced variational autoencoder enhances the high-frequency details of the image by extracting multi-level encoded features (F1, F2, F3), which can improve the detection ability of forgery traces.

[0037] Step S102: Perform frequency domain enhancement processing on the multi-level encoded features to generate high-frequency enhanced features; Among them, frequency-domain enhancement extracts high-frequency information by performing frequency-domain transformation on image features (such as discrete Fourier transform, discrete cosine transform, etc.). Since forged traces in images are usually manifested in details (such as small changes and abnormal textures in the image), frequency-domain enhancement can extract these details.

[0038] Exemplarily, after performing frequency-domain enhancement processing on the input image, high-frequency details (such as textures, noise, etc.) in the image will be emphasized, which helps to highlight forged traces and makes these details more prominent for subsequent forgery detection.

[0039] Step S103: Input the high-frequency enhanced features into the decoder of the frequency-domain enhanced variational autoencoder, and perform feature reconstruction in combination with the dynamic spectrum loss mechanism to generate the reconstructed frequency-domain supplementary features. Among them, the decoder of the frequency-domain enhanced variational autoencoder restores the detailed information of the image by converting the high-frequency enhanced features back to the image space (or latent space). Combining with the dynamic spectrum loss mechanism enables the decoder to focus on optimizing the high-frequency components of the image during the reconstruction process, thereby more accurately reconstructing the features of the forged image.

[0040] Specifically, the decoder converts the frequency-domain enhanced features into a reconstructed version of the image, where details (such as forged traces) are clearer. At the same time, the reconstruction process is optimized through the spectrum loss mechanism, making the difference in details between the reconstructed image and the real image more significant.

[0041] Step S104: Perform global pooling and local pooling on the multi-level encoded features through an unsupervised detection network to generate a global representation and a local representation, and connect the global representation and the local representation along the channel axis to generate fused multi-scale features. Among them, the global pooling operation extracts the overall information of the image, while the local pooling extracts the local features of the image. After combining these two operations, the face image can be analyzed from both global and local perspectives, enhancing the comprehensiveness of detection. For example, global pooling will obtain the feature representation of the entire image (such as the general contour of the face), while local pooling can capture detailed information (such as the specific structure of the eyes and mouth).

[0042] Furthermore, connect the global representation and the local representation along the channel axis to form a comprehensive feature map. These fused multi-scale features can capture the detailed information of the image from multiple angles. After being processed by a convolutional network, more rich features are extracted for forgery detection.

[0043] Exemplarily, assume that the global representation contains the basic contour of the face, and the local representation captures the detailed features such as eyes and nose. After further processing by the convolutional network, a multi-scale feature map that fuses global and local information can be obtained.

[0044] Step S105: Calculate an anomaly score map based on the fused multi-scale features and the reconstructed frequency-domain supplementary features; Among them, the fused multi-scale features and the frequency-domain supplementary features are combined, and by calculating the difference between the two, an anomaly score map is generated. The score map can highlight the abnormal areas in the image and indicate the possible forged parts.

[0045] Exemplarily, by calculating the difference (such as the reconstruction error) between the two features, a score map is generated, where the high-score areas may represent the forged parts of the image.

[0046] Step S106: Determine the authenticity result of the face image to be detected according to the pixel-level scores of the anomaly score map.

[0047] Among them, according to the pixel-level scores in the anomaly score map, it is determined which areas may have forgery traces. By setting a threshold to judge whether the image is forged. Usually, if the scores of some areas in the anomaly score map exceed the predetermined threshold, then that area may be a forgery trace, and finally the whole image is judged to be forged.

[0048] In the above embodiments, a frequency-domain enhanced variational autoencoder, an unsupervised detection network, and frequency-domain supplementary features are introduced. By comprehensively using the global and local features of the image, the forgery traces in the face image can be effectively detected under unsupervised conditions. Through frequency-domain enhancement to improve image details, using multi-scale feature fusion to enhance the recognition ability of forged images, and finally making pixel-level judgments through the anomaly score map, the high-precision and robustness of the detection are ensured. This technical solution can effectively identify complex forged images and has a wide range of application prospects.

[0049] Refer to Figure 2 As shown, an image processing model based on the encoder-decoder architecture of the embodiments of the present application is shown. The input image is first subjected to feature extraction through the encoder (E). Among them, the image features are extracted through multiple encoding modules (such as E1, E2, E3), and feature representations a1, a2, and a3 are generated. These features are processed by DSL (feature learning module) to generate more refined features c1, c2, and c3. Then, the image is reconstructed through the decoder (G) to generate an image output. The model also includes local and global information processing, using a convolutional neural network (CNN) to perform further feature learning and visual generation on the input image, and finally outputting an image. This model encodes, processes, and reconstructs the input image through the encoder-decoder architecture and deep learning strategy, aiming to complete the accurate restoration or forgery detection of the image.

[0050] Refer to Figure 3 , as an implementation manner of step S102, the steps of performing frequency-domain enhancement processing on the multi-level encoded features to generate high-frequency enhanced features include: Step S201: Perform discrete cosine transform on each layer of the multi-level encoded features to convert them into frequency domain features; Among them, the multi-level encoded features of the image are extracted through a deep neural network (such as ResNet). The encoded features at different levels represent different levels of image information - from low-level edge, corner, etc. to high-level object recognition information (such as features of the nose, eyes, etc. of a face). Performing discrete cosine transform (DCT) on these features respectively can effectively convert the image from the spatial domain to the frequency domain, and then analyze the performance of each level of image features in the frequency domain.

[0051] Specifically, the discrete cosine transform DCT is a commonly used frequency domain conversion method. By converting the spatial domain (pixel values in the original image) to the frequency domain, it helps to extract the frequency components of the image. The frequency domain representation can reveal the high-frequency and low-frequency information of the image. Among them, the high-frequency part usually represents the details of the image (such as texture, noise, etc.), and the low-frequency part represents the overall structure of the image (such as contour, smooth part of the region).

[0052] Exemplarily, assume that the input image is a face image with forgery traces. After being processed by the ResNet network, the obtained encoded features contain information at different levels, including details, structures, etc. Performing DCT conversion on the encoded features of each layer can obtain the corresponding frequency domain features, which is convenient for subsequent frequency domain analysis.

[0053] Step S202: Perform low-pass filtering weighting on the low-frequency components of the frequency domain features and perform adaptive weight allocation on the high-frequency components of the frequency domain features; Among them, the low-pass filter allows low-frequency components to pass through and suppresses high-frequency components. In image processing, the low-frequency part usually represents the smooth area or large-scale structure of the image, such as the contour of a person and the facial contour. By performing low-pass filtering weighting, the expressiveness of these low-frequency components can be enhanced, making the overall structure of the image more prominent.

[0054] For the high-frequency part in the frequency domain, the role of high-frequency components is strengthened through adaptive weight allocation. Adaptive weight allocation dynamically adjusts according to the features of different frequencies, usually determining the weights of certain high-frequency components according to the characteristics of the forged image and its specific forgery traces. For example, forgery traces usually appear as small changes in the image, and these changes are mostly concentrated in the high-frequency region. Therefore, assigning higher weights to the high-frequency part helps to highlight these forged features.

[0055] Exemplarily, for the low-frequency region in the face image (such as the face contour), its smoothness effect is enhanced through low-pass filtering weighting; while for the high-frequency region in the image (such as details like eyes and skin texture), its performance in the frequency domain is further enhanced through adaptive weight allocation, making the forgery traces more prominent and facilitating subsequent detection.

[0056] It can be understood that low-pass filtering weighting enhances the overall structure of the image, while adaptive weight allocation helps detect minute forgery traces in forged images by boosting the weights of high-frequency features, especially detailed information. Through this method, the visibility of forgery traces can be significantly improved in the frequency domain.

[0057] Step S203: Convert the weighted frequency-domain features to the spatial domain through inverse discrete cosine transform to obtain high-frequency enhanced features.

[0058] Among them, the inverse discrete cosine transform (IDCT) is the inverse operation of DCT and is used to convert frequency-domain information back to the spatial domain. By performing IDCT on the weighted frequency-domain features to restore them to the spatial domain, a new image feature is obtained. This process enables the detailed information enhanced in the frequency domain to be reflected in the spatial domain of the image, facilitating further image analysis and processing.

[0059] Specifically, through IDCT conversion, the enhanced features originally existing in the frequency domain (especially the detailed information in the high-frequency part) will return to the image in a more obvious form. At this time, the detailed parts in the image (such as forgery traces) are more prominently presented, making subsequent forgery detection more sensitive and accurate.

[0060] Exemplarily, after the frequency-domain features are weighted by low-pass filtering and adaptively weighted and then converted by IDCT, the generated image will more prominently display its detailed information. For example, the originally imperceptible forgery traces will become more obvious in the image, facilitating further processing and judgment by the detection network.

[0061] In the above embodiments, the image features are converted from the spatial domain to the frequency domain, and then through low-pass filtering and adaptive weight allocation, the high-frequency details are emphasized, thereby enhancing the expressiveness of forgery traces. Finally, through the inverse discrete cosine transform (IDCT), the enhanced frequency-domain features are converted back to the spatial domain to generate high-frequency enhanced features. This process not only improves the detailed information of the image, especially when the forgery traces are relatively weak, but also makes the minute changes in the forged image more significant through high-frequency enhancement, providing a more reliable basis for subsequent forgery detection.

[0062] Refer to Figure 4 , as an embodiment of the frequency-domain enhanced variational autoencoder, the steps of constructing the frequency-domain enhanced variational autoencoder include: Step S301: Input the input feature x into the encoder E of the pre-constructed variational autoencoder for convolutional sampling, and map it to the latent space z; where the input feature x is a multi-level encoded feature (F1, F2, F3); Step S302: Quantize the features in the latent space z, and generate a quantized latent embedding through formula (1) : ; (1) where the embedding space C = {c1, c2, …, c∣C∣} is a preset discrete vector set, z ij is the j-th feature of the i-th layer in the latent space z, and c k represents the nearest vector in the embedding space; Step S303: Input the quantized latent embedding into the decoder of the variational autoencoder G for layer-by-layer reconstruction to generate a preliminary reconstructed image , and the calculation process is: ; ... (2) ; ; (3) In the above formula (2), represents the i-th layer of the encoder, represents the feature before activation of the i-th layer of the encoder, represents the frequency complementary feature of the i-th layer, which is used to extract the rich frequency features learned from the feature before activation of the encoder; in the above formula (3), represents the corresponding function of the i-th layer of the decoder, represents the output feature of the decoder; Step S304: Perform frequency domain processing on the feature before activation of the encoder and the output feature of the decoder to generate a frequency complementary feature ; Step S305: Perform discrete Fourier transform on the frequency complementary feature and the output feature of the decoder, and calculate the focus frequency loss FFL and the dynamic spectrum loss SL. The specific calculation formulas are: ; (4) ; (5) ; (6) ; (7) In the above formula (4), e and i represent the Euler number and the imaginary unit respectively, M×N is the spatial resolution of the feature map, f(x,y) is the value at (x,y) of each feature map, and F(u,v) is the corresponding value at the (u,v) coordinate on the frequency spectrum; in the above formula (5), , representing the weight applied to each frequency; , representing the main error function based on the frequency domain; in the above formula (6), represents the Gaussian kernel function, μ is the mean, σ i is the variance, * represents the convolution operation, and are the weighted features; Referring to Figure 5 , the frequency domain enhanced variational autoencoder of the embodiment of the present application is based on a frequency domain supplement module, Figure 5 As shown, it is a schematic diagram of the frequency domain supplement module constructed in the embodiment of the present application. This module receives the input signal and undergoes a series of operations (including normalization (Norm), ReLU activation, Dropout, and convolution (Conv)). These operations are repeated N times to extract the features in the signal. Finally, through the processing of these steps, the output signal is supplemented and enhanced and synthesized with other signals. The purpose of this frequency domain supplement module is to enhance the feature expression of the input signal through frequency domain feature processing, thereby improving the performance of the overall model in image or signal processing.

[0063] Step S306, combine the focus frequency loss FFL, dynamic spectrum loss SL, pixel-level reconstruction loss, and perceptual loss to calculate the total reconstruction loss L rec , and the specific calculation formula is: ; (8) In the above formula (8), α and β are hyperparameters, is the norm pixel-level reconstruction loss, is the perceptual loss; Step S307, based on the total reconstruction loss L rec Optimize the model parameters to obtain the trained frequency domain enhanced variational autoencoder.

[0064] Referring to Figure 6 , as shown is the schematic diagram of the dynamic frequency domain loss constructed in the embodiment of the present application. After the input features are subjected to discrete Fourier transform, they are converted to the frequency domain to generate the focus frequency loss (FFL). First, b i in the figure and the output c iThe spectrum F(u,v) is calculated through the discrete Fourier transform respectively, and then the loss value is calculated using the difference in characteristic frequencies. The focus frequency loss (FFL) is calculated through a formula, and a low-pass filter is applied according to the frequency domain characteristics to adjust the penalty weights of the low-frequency and high-frequency parts, ensuring the accurate restoration of the overall image layout and main features. Then, frequency domain weighting processing is performed through convolution operation and Gaussian kernel (K_i) to obtain the weighted features and , and then the frequency loss SL( and ) is calculated. Finally, the total reconstruction loss is used to measure the image reconstruction quality. This loss function combines the focus frequency loss, feature map loss, and pixel-level and perceptual losses to ensure the quality and visual effect of the reconstructed image

[0065] In the above embodiments, optimization mechanisms such as frequency domain processing, focus frequency loss (FFL), and dynamic spectrum loss (SL) are introduced, effectively improving the reconstruction quality of the variational autoencoder (VAE). Through frequency domain enhancement on the features of the encoder and decoder, not only the traditional pixel-level reconstruction error is considered, but also the error constraint at the frequency level is introduced, enhancing the model's learning ability for image frequency domain features. The combination of quantized latent embedding and frequency domain features ensures the retention of image details and reconstruction accuracy. At the same time, the use of focus frequency loss and dynamic spectrum loss effectively reduces the reconstruction error in the low-frequency and high-frequency parts. Finally, through comprehensive optimization of the total reconstruction loss, higher-quality and more realistic image reconstruction results can be obtained, improving the model's expressiveness and reconstruction effect, especially in terms of high-frequency details and visual perception

[0066] Referring to Figure 7 , as an embodiment of step S104, global pooling and local pooling are performed on the multi-level encoded features through an unsupervised detection network to generate a global representation and a local representation, and the global representation and the local representation are concatenated along the channel axis. The calculation formula for generating the fused multi-scale features includes: Step S401, perform global average pooling on the input feature x to generate a global representation as: ; (9) In the above formula (9), Pooling represents the global pooling operation Step S402, divide the input feature x into q×q regions, and perform local average pooling on each region to generate a local representation as: ; (10) In the above formula (10), Pooling represents the local pooling operation Step S403, the global representation Concatenated with the local representation In series along the channel axis and input into the convolutional network ψ (⋅), to obtain the fused multi-scale features It is: ; (11) In the above formula (11), is the fused multi-scale feature, ψ (⋅) is the convolutional network, represents the concatenation of the global representation and the local representation along the channel axis.

[0067] In the above embodiment, based on the unsupervised detection network, global pooling and local pooling are performed on the input features, and the global representation and the local representation are combined to generate the fused multi-scale features. Global average pooling retains the overall information of the input features, while local average pooling captures the local detailed features. After connecting the global representation and the local representation along the channel axis, the global and local information can be effectively combined to fully express the multi-scale features. This method improves the network's perception ability of different scale information, thereby enhancing the network's feature extraction and discrimination ability. Especially when dealing with tasks with complex details and multi-level structures, the model's expressiveness and detection accuracy can be significantly improved.

[0068] Referring to Figure 8 , as an embodiment of step S105, the steps of calculating the anomaly score map based on the fused multi-scale features and the reconstructed frequency-domain supplementary features include: Step S501, calculate the mean square error and cosine similarity between the fused multi-scale features and the reconstructed frequency-domain supplementary features to obtain the pixel similarity It is: ; (12) In the above formula (12), h and w represent the coordinate indices of the spatial features, is the indicator function, and are the mean square error and cosine similarity respectively, and are hyperparameters, and k represents the threshold for determining the similarity metric MSE; Step S502, obtain the reconstructed feature according to the pixel similarity It is: ; (13) Step S503, calculate the mean square error loss according to the input feature x and the reconstructed feature as: ; (14) In the above formula (14), C, H, and W are the number of channels, height, and width of the feature map, respectively. represents the Euclidean distance; Step S504, according to the reconstructed feature obtain the anomaly score map d, and the specific formula includes: ; (15) In the above formula (15), h and w represent the positions of each pixel.

[0069] Among them, the score of each pixel position (h, w) represents the degree of anomaly at that position. The larger the value, the more significant the difference from the real face feature, and the higher the forgery possibility.

[0070] In the above embodiment, based on the fused multi-scale features and the reconstructed frequency-domain supplementary features, the pixel similarity is calculated and the anomaly score map is generated, thereby effectively detecting the anomaly region. First, by calculating the mean square error and cosine similarity between the fused features and the frequency-domain supplementary features, the similarity between pixels can be accurately measured, and then the reconstructed features can be obtained. This process can capture the differences between the input features and the reconstructed features, and use these differences to calculate the mean square error loss, thereby generating an anomaly score map for each pixel position. This technical solution can accurately identify the anomaly region in the image. By combining multi-scale features and frequency-domain information, the anomaly detection becomes more refined and robust, especially suitable for processing images with complex structures and noises.

[0071] Referring to Figure 9 , as an embodiment of step S106, the steps of determining the authenticity result of the face image to be detected according to the pixel-level scores of the anomaly score map include: Step S601, interpolate the anomaly score map to the original size of the input face image to be detected; Among them, the size of the anomaly score map is usually obtained by extracting features from the input image. These feature maps may become smaller due to processes such as pooling and convolution operations. Therefore, in order to align the anomaly score map with the input image, interpolation processing is required. This interpolation process can restore the anomaly score map to the size of the original image through image upsampling or interpolation algorithms (such as bilinear interpolation or nearest neighbor interpolation), so that the anomaly score of each pixel point corresponds to the corresponding image pixel one by one, facilitating subsequent processing.

[0072] Exemplarily, assume that the size of the input image is 224×224, and the size of the anomaly score map is 56×56. The anomaly score map is upsampled from 56×56 to 224×224 through interpolation, so that the anomaly score at each position can correspond to the corresponding position in the original image.

[0073] Step S602: Perform global max pooling on the interpolated anomaly score map, and extract the maximum anomaly score in the anomaly score map as the final anomaly score; Among them, the global pooling operation extracts the most significant anomaly region in the image by performing pooling on the anomaly score map. Max pooling is an operation that selects the maximum value in the pooling window. Global max pooling finds the maximum anomaly score in the entire anomaly score map, representing the degree of anomaly in the most abnormal and most likely forged region of the image.

[0074] Exemplarily, assume that in the interpolated anomaly score map, the score of a certain region is the highest (e.g., 0.8), and this region is the most severely abnormal part. Through global max pooling, this maximum value is extracted as the final anomaly score.

[0075] Step S603: Compare the preset threshold with the final anomaly score to determine the authenticity result of the face image to be detected.

[0076] Among them, the preset threshold is used to judge the forgery degree of the image. If the final anomaly score exceeds this threshold, it can be considered that there are signs of forgery in the image. Otherwise, the image is considered a genuine image. The selection of the threshold usually depends on the score distribution of genuine images and forged images during the training process. By comparing the final anomaly score with this threshold, a judgment can be made on whether the image is a forged image.

[0077] Exemplarily, assume that the preset threshold is 0.7. When the final anomaly score is 0.8, it indicates that the abnormal part in the image is significant, and the system determines that the image is forged; if the final anomaly score is 0.5, the image is judged to be genuine.

[0078] In the above embodiment, the anomaly score map is aligned with the input image, so that the anomaly scores can be accurately compared at the pixel level. Through global max pooling, the most significant forgery traces in the image are extracted, ensuring the accuracy of forgery detection. Finally, by comparing with the preset threshold, the authenticity result of the image is determined. This technical solution can accurately and effectively detect the forged part in the image by combining the anomaly score map, pooling operation, and threshold judgment, and has high detection accuracy and good robustness.

[0079] As a further embodiment of the detection method, it further includes: Combining the total reconstruction loss L rec with the mean square error loss to obtain the total objective function as : ; (16) In the above formula (16), and They are preset hyperparameters that respectively control the influence of the total reconstruction loss and the mean square error loss on the total objective function.

[0080] Among them, by weighted combination of the total reconstruction loss and the mean square error loss, it guides the model to optimize both the reconstructed image quality and the feature distribution alignment simultaneously, ensuring the similarity between the reconstructed image output by the decoder and the input source image, and at the same time constraining the distribution consistency between the real face features and the fused features.

[0081] In the above embodiments, the weighted combination of the total reconstruction loss and the mean square error loss enables the method to take into account both global and local features in forgery detection, and can more accurately capture the differences between forged images and real images. Through this combination, the detection method can make fine distinctions when dealing with different types of forged images, ensuring more accurate detection results.

[0082] The technical solution of this application proposes a multi-feature fusion face forgery detection method (FDS-Net) based on frequency domain supplementation and unsupervised learning. Its core lies in combining the strategies of frequency domain enhancement and local-global feature fusion, significantly improving the detection accuracy and robustness of forged images. First, a variational autoencoder (VAE) based on frequency domain enhancement is adopted. By converting the face image into the frequency domain and enhancing the high-frequency spectral features, it can effectively magnify the forgery traces of forged images and suppress the noise in real images, making the distinction between forged and real images more obvious, thus improving the detection performance. This technology effectively addresses the deficiencies of traditional methods in forged image detection, especially for the processing of high-frequency details.

[0083] Secondly, the unsupervised network adopted in the embodiments of this application combines local detail features and global semantic information, and enhances the model's detection ability for forged images through multi-scale perception. Without manual labels, this network automatically extracts the internal structure and forgery clues of images through self-supervised learning, can adapt to different types of forgery methods, and maintains strong robustness and generalization ability. This innovative design avoids information loss in traditional methods, improves the detection accuracy, and has high adaptability. Experimental results show that FDS-Net performs excellently on multiple public datasets (such as Celeb-DF and DFDC). It not only exceeds many existing methods in terms of forgery detection accuracy and generalization ability, but also demonstrates strong cross-dataset adaptability. The innovation of this technical solution lies in combining frequency domain enhancement and unsupervised learning to achieve efficient and accurate detection of forged images in the absence of a large amount of labeled data, and has strong practical application value. Especially in scenarios that require large-scale automated detection, it can effectively improve the efficiency and accuracy of face forgery detection.

[0084] The embodiments of this application also disclose an unsupervised face forgery detection system based on frequency domain supplementation.

[0085] An unsupervised face forgery detection system based on frequency domain supplementation, comprising: A feature extraction module, configured to input a face image to be detected into a benchmark network based on ResNet, and extract multi-level encoded features through an encoder of a frequency domain enhanced variational autoencoder; A frequency domain enhancement module, configured to perform frequency domain enhancement processing on the multi-level encoded features to generate high-frequency enhanced features; A reconstruction module, configured to input the high-frequency enhanced features into a decoder of the frequency domain enhanced variational autoencoder, and perform feature reconstruction in combination with a dynamic spectrum loss mechanism to generate reconstructed frequency domain supplementary features; A multi-scale feature fusion module, configured to perform global pooling and local pooling on the multi-level encoded features through an unsupervised detection network to generate a global representation and a local representation, and connect the global representation and the local representation along the channel axis to generate fused multi-scale features; An anomaly score map calculation module, configured to calculate an anomaly score map based on the fused multi-scale features and the reconstructed frequency domain supplementary features; An authenticity verification module, configured to determine the authenticity result of the face image to be detected according to the pixel-level scores of the anomaly score map.

[0086] The unsupervised face forgery detection system based on frequency domain supplementation in the embodiments of the present application can implement any one of the above forgery detection methods, and the specific working processes of each module in the forgery detection system can refer to the corresponding processes in the above method embodiments.

[0087] In several embodiments provided by the present application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0088] The embodiments of the present application also disclose a computer device.

[0089] The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements an unsupervised face forgery detection method based on frequency domain supplementation as described above.

[0090] The embodiments of the present application also disclose a computer-readable storage medium.

[0091] The computer-readable storage medium stores a computer program that can be loaded and executed by a processor to implement any one of the unsupervised face forgery detection methods based on frequency domain supplementation as described above.

[0092] Among them, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or component; the program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0093] It should be noted that in the above embodiments, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0094] The above are all preferred embodiments of the present application. The protection scope of the present application is not limited thereby. Any feature disclosed in this specification (including the abstract and drawings), unless specifically described, can be replaced by other equivalent or similar-purpose alternative features. That is, unless specifically described, each feature is only an example of a series of equivalent or similar features.

Claims

1. An unsupervised face forgery detection method based on frequency domain complementation, characterized in that: The detection method comprises: The face image to be detected is input into the benchmark network based on ResNet, and the encoder of the frequency domain enhanced variational autoencoder is used to extract multi-level coding features; Performing frequency domain enhancement processing on the multi-level coding features to generate high-frequency enhancement features; Inputting the high-frequency enhancement feature into the decoder of the frequency-domain enhancement variational autoencoder, performing feature reconstruction in combination with a dynamic spectrum loss mechanism, and generating a reconstructed frequency-domain supplementary feature; Performing global pooling and local pooling on the multi-level coding features through an unsupervised detection network to generate a global representation and a local representation, and connecting the global representation and the local representation along a channel axis to generate a fused multi-scale feature; Calculating an abnormal score map based on the fused multi-scale features and the reconstructed frequency domain supplementary features; The image authenticity result of the face image to be detected is determined according to the pixel level score of the abnormal score map.

2. The unsupervised face forgery detection method based on frequency domain complementation according to claim 1 is characterized in that: The step of performing frequency domain enhancement processing on the multi-level coding features to generate high-frequency enhancement features includes: Performing discrete cosine transform on each layer of coding features of the multi-level coding features to convert them into frequency domain features; Performing low-pass filtering weighting on the low-frequency components of the frequency domain features, and performing adaptive weight allocation on the high-frequency components of the frequency domain features; The weighted frequency domain features are converted to the spatial domain through inverse discrete cosine transform to obtain high-frequency enhanced features.

3. The unsupervised face forgery detection method based on frequency domain complementation according to claim 1, characterized in that: The steps of constructing the frequency domain enhanced variational autoencoder include: Input the input feature x into the encoder E of the pre-built variational autoencoder for convolution sampling and mapping to the latent space z; The features in the latent space z are quantized and the quantized latent embedding is generated by formula (1) : ;(1) Among them, the embedding space C={c1,c2,…,c|C|} is a preset discrete vector set, z ij is the jth feature of the i-th layer in the latent space z, c k represents the nearest vector in the embedding space; Embedding the quantized potential Input to the decoder of the variational autoencoder G and reconstruct layer by layer to generate a preliminary reconstructed image , the calculation process is: ; ...(2) ; ;(3) In the above formula (2), represents the i-th layer of the encoder, represents the features before the activation of the i-th layer of the encoder, represents the frequency complementary feature of the i-th layer, Used to extract features before activation from the encoder The rich frequency features learned in ; In the above formula (3), represents the corresponding function of the decoder layer i, represents the output features of the decoder; The pre-activation features of the encoder And the output characteristics of the decoder Perform frequency domain processing to generate frequency supplementary features ; Supplementary characteristics of the frequency And the output characteristics of the decoder Perform discrete Fourier transform to calculate the focus frequency loss FFL and dynamic spectrum loss SL. The specific calculation formula is: ;(4) ;(5) ;(6) ;(7) In the above formula (4), e and i represent the Euler number and imaginary unit respectively, M×N is the spatial resolution of the feature map, f(x,y) is the value at (x,y) of each feature map, and F(u,v) is the corresponding value at the (u,v) coordinate on the spectrum; in the above formula (5), , represents the weight applied to each frequency; , represents the main error function based on the frequency domain; in the above formula (6), represents the Gaussian kernel function, μ is the mean, σ i is the variance, * indicates the convolution operation, and is the weighted feature; Combining the focus frequency loss FFL, dynamic spectrum loss SL, pixel-level reconstruction loss and perceptual loss, the total reconstruction loss L is calculated rec , the specific calculation formula is: ;(8) In the above formula (8), α and β are hyperparameters. for norm pixel-level reconstruction loss, It is a perceived loss; Based on the total reconstruction loss L rec Optimize the model parameters to obtain the trained frequency domain enhanced variational autoencoder.

4. The unsupervised face forgery detection method based on frequency domain complementation according to claim 3 is characterized in that: The multi-level coding features are globally pooled and locally pooled through an unsupervised detection network to generate a global representation and a local representation, and the global representation and the local representation are connected along the channel axis to generate a calculation formula for fusion multi-scale features, including: Perform global average pooling on the input feature x to generate a global representation for: ;(9) In the above formula (9), Pooling represents the global pooling operation; Divide the input feature x into q×q regions, perform local average pooling on each region, and generate local representation for: ;(10) In the above formula (10), Pooling represents the local pooling operation; The global representation With local representation Concatenate along the channel axis and input into the convolutional network ψ (⋅), and obtain the fused multi-scale features for: ;(11) In the above formula (11), To fuse multi-scale features, ψ (⋅) is a convolutional network, represents the concatenation of the global representation and the local representation along the channel axis.

5. The unsupervised face forgery detection method based on frequency domain complementation according to claim 4 is characterized in that: The step of calculating an abnormal score map based on the fused multi-scale features and the reconstructed frequency domain supplementary features comprises: Calculate the fused multi-scale features The pixel similarity is obtained by comparing the mean square error and cosine similarity of the reconstructed frequency domain complementary features. for: ;(12) In the above formula (12), h and w represent the coordinate index of the spatial feature, is the indicator function, and are mean square error and cosine similarity respectively, and is a hyperparameter, k represents the threshold that determines the similarity metric MSE; According to the pixel similarity Get the reconstruction features for: ;(13) According to the input feature x and the reconstructed feature The mean square error loss is calculated as: ;(14) In the above formula (14), C, H, and W are the number of channels, height, and width of the feature map, respectively. represents the Euclidean distance; According to the reconstruction features Get the abnormal score graph d, the specific formula includes: ;(15) In the above formula (15), h and w represent the position of each pixel.

6. The unsupervised face forgery detection method based on frequency domain complementation according to claim 5, characterized in that: The step of determining the authenticity of the face image to be detected according to the pixel-level scores of the abnormal score map comprises: Interpolating the anomaly score map to the original size of the input image to be detected; Performing global maximum pooling on the interpolated anomaly score map, and extracting the maximum anomaly score in the anomaly score map as the final anomaly score; The authenticity result of the face image to be detected is determined based on a comparison between a preset threshold and the final anomaly score.

7. The unsupervised face forgery detection method based on frequency domain complementation according to any one of claims 1 to 6, characterized in that: The detection method further comprises: The total reconstruction loss L rec And mean square error loss Perform weighted combination and get the total objective function as : ;(16) In the above formula (16), and To preset hyperparameters, we control the impact of total reconstruction loss and mean square error loss on the overall objective function.

8. An unsupervised face forgery detection system based on frequency domain complementation, characterized in that: The detection system comprises: The feature extraction module is used to input the face image to be detected into the ResNet-based baseline network and extract multi-level coding features through the encoder of the frequency domain enhanced variational autoencoder; A frequency domain enhancement module, used to perform frequency domain enhancement processing on the multi-level coding features to generate high-frequency enhancement features; A reconstruction module, used to input the high-frequency enhancement features into the decoder of the frequency-domain enhancement variational autoencoder, perform feature reconstruction in combination with a dynamic spectrum loss mechanism, and generate reconstructed frequency-domain supplementary features; A multi-scale feature fusion module, used to perform global pooling and local pooling on the multi-level coding features through an unsupervised detection network to generate a global representation and a local representation, and connect the global representation and the local representation along a channel axis to generate a fused multi-scale feature; An abnormal score map calculation module, used to calculate the abnormal score map based on the fused multi-scale features and the reconstructed frequency domain supplementary features; The authenticity verification module is used to determine the image authenticity result of the face image to be detected according to the pixel level score of the abnormal score map.

9. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.

10. A computer-readable storage medium, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Deep fake face detection method based on frequency learning

    CN114898437A

  • Forgery image detection method, electronic equipment and computer readable storage medium

    CN115984178A

  • Test stage training face forgery detection method and system containing uncertainty guidance

    CN117275068A

  • Multi-scale feature fusion depth forgery detection method based on reconstruction learning

    CN119068318A

  • Medical image segmentation method based on global and local feature reconstruction network

    WO2023151141A1

Cited By

  • Multi-core feature representation learning system and method for enhancing remote sensing image based on content retrieval

    CN120853029A

  • Microservice system anomaly detection method, device and equipment based on time-frequency feature contrast enhancement, and storage medium

    CN121144144A

  • Face image anti-counterfeiting detection method and device based on frequency spectrum reconstruction

    CN121354195A

  • Face image anti-counterfeiting detection method and device based on spectrum reconstruction

    CN121354195B

  • A method and system for image authenticity determination based on frequency domain perturbation response differences

    CN122574605A