Face image detection method and system based on frequency anomaly injection attention

CN122551414BActive Publication Date: 2026-09-18TIBET UNIVERSITY FOR NATIONALITIES
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202611058132.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-09-18
Estimated Expiration
2046-07-16

AI Technical Summary

Technical Problem

[0006]本发明的目的在于针对现有技术的不足,提供一种基于频率异常注入注意力的人脸图像检测方法及系统,以解决现有技术中检测模型极易陷入捷径学习、双流网络计算冗余高以及相位谱噪声干扰导致特征难以解耦的技术问题,在保证检测精度的前提下,显著提升跨域泛化能力与实时推理效率

Benefits of technology

本发明通过对待检测人脸图像进行预处理提取包含幅度谱和相位谱的频域特征,基于幅度谱生成动态滤波掩码并对相位谱进行定向滤波提纯,能够精准筛选出与深度伪造相关的纯净频率异常特征;再将双通道频率异常特征与空间基础特征拼接融合后生成单通道空间注意力掩码,并以该空间注意力掩码作为信息瓶颈,通过残差调制机制对图像基础特征进行加权干预与特征解耦,可主动打破模型对域内低维捷径特征的依赖,强制模型学习跨域通用的伪造本质痕迹,无需搭建复杂的双流网络架构即可实现特征的有效解耦,在保障检测精度的同时显著提升跨域泛化能力与推理效率,最终输出稳定可靠的人脸图像真假检测结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551414B_ABST
    Figure CN122551414B_ABST
Patent Text Reader

Abstract

The application discloses a face image detection method and system based on frequency anomaly injection attention, and relates to the technical field of artificial intelligence security and multimedia forensics. The method comprises the following steps: performing pretreatment extraction on the face image to obtain a frequency domain feature; generating a dynamic filter mask based on the amplitude spectrum, filtering and purifying the phase spectrum by using the dynamic filter mask to obtain a dual-channel frequency anomaly feature; generating a single-channel spatial attention mask through a convolution residual network; taking the spatial attention mask as an information bottleneck, performing weighted intervention and feature decoupling on the basic feature of the face image to be detected through a residual modulation mechanism to obtain an intervened and decoupled feature; and inputting the intervened and decoupled feature into a classifier for classification to output a final true or false detection result. The application effectively improves the generalization ability and reasoning efficiency of deep fake face detection in cross-domain scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence security and multimedia forensics, and in particular to a face image detection method and system based on frequency anomaly injection attention. Background Technology

[0002] With the rapid evolution of generative artificial intelligence technology, deepfake technology has diverged into two categories: "generative synthesis" (such as full-face synthesis generated by diffusion models) and "manipulation attacks" (such as face swapping and facial expression replay). The widespread dissemination of deepfake content poses a serious threat to personal privacy; therefore, high-precision and robust deepfake face detection technology has become a key research focus in the field of artificial intelligence security. Although existing deepfake detection technologies have made some progress, they still face the following three major technical bottlenecks: First, feature entanglement leads to poor cross-domain generalization. Existing single-stream detection models (such as Xception and EfficientNet) typically learn all types of forgery features in a "one-size-fits-all" manner. However, physical forgeries (such as splicing boundaries generated by face swapping and high-frequency texture anomalies) and semantic forgeries (such as logical inconsistencies caused by facial expression replay and failure of facial feature coordination) belong to completely orthogonal feature dimensions. Hybrid learning causes the model to easily overfit to a specific feature when facing unknown attacks, resulting in a significant performance drop on cross-domain datasets.

[0003] Second, the input distribution of multimodal pre-trained models is not well-matched. With the increasing trend of introducing large-scale pre-trained models like CLIP to assist in detection, effectively combining traditional CNNs pre-trained on ImageNet (such as EfficientNet) with Vision Transformers pre-trained on large-scale image-text pairs (such as CLIP) has become a challenge. The input data distributions (mean, variance) of the two differ significantly, and directly concatenating the inputs leads to difficulties in feature space alignment, affecting model convergence.

[0004] Third, there is redundancy and inefficiency in computational resources. Existing state-of-the-art methods (such as large Transformer-based models) perform deep inference on all input samples indiscriminately. In reality, approximately 24%-74% of fake samples in real-world scenarios contain obvious statistical anomalies that can be identified without using large models. Existing methods lack a tiered processing mechanism, resulting in a huge waste of computational power and making it difficult to meet the needs of real-time video stream detection.

[0005] In summary, existing detection models are prone to getting stuck in shortcut learning, have high computational redundancy in two-stream networks, and suffer from phase spectrum noise interference that makes feature decoupling difficult. There is an urgent need for a deep fake face detection method that can both decouple features to improve robustness, solve the input distribution adaptation problem, and have efficient inference capabilities. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of existing technologies by providing a face image detection method and system based on frequency anomaly injection attention. This addresses the technical problems in existing technologies, such as the detection model being prone to getting stuck in shortcut learning, high computational redundancy in dual-stream networks, and difficulty in decoupling features due to phase spectrum noise interference. While ensuring detection accuracy, this invention significantly improves cross-domain generalization ability and real-time inference efficiency.

[0007] To achieve the above objectives, the present invention provides the following solution: A face image detection method based on frequency anomaly injection attention includes the following steps: S1: Acquire the face image to be detected, preprocess the face image to extract frequency domain features, the frequency domain features include amplitude spectrum and phase spectrum; S2: Generate a dynamic filter mask based on the amplitude spectrum, and use the dynamic filter mask to filter and purify the phase spectrum to obtain the frequency anomaly characteristics of the dual channels; S3: The frequency anomaly features of the dual channels are spliced ​​and fused with the basic features of the face image to be detected, and a single-channel spatial attention mask is generated through a convolutional residual network; S4: Using the spatial attention mask as an information bottleneck, the basic features of the face image to be detected are weighted and decoupled through a residual modulation mechanism to obtain the features after intervention and decoupling. S5: Input the decoupled features after the intervention into the classifier for classification, and output the final true / false detection result.

[0008] Furthermore, the preprocessing and extraction of the face image specifically includes: Convert the face image to be detected into a single-channel grayscale image; The single-channel grayscale image is transformed from the spatial domain to the complex frequency domain using a two-dimensional fast discrete Fourier transform. By performing polar coordinate decomposition on the complex spectrum, the amplitude spectrum and phase spectrum in real form are obtained.

[0009] Furthermore, the generation of the dynamic filter mask based on the amplitude spectrum specifically involves: Construct multiple concentric ring-shaped binary masks; The amplitude spectrum is divided into multiple sub-bands with different frequency step sizes. The average frequency energy feature in each sub-band is calculated. The average frequency energy feature is then used to generate a band weight vector through adaptive learning. The frequency band weight vector is remapped onto the corresponding binary mask to generate a smooth two-dimensional dynamic filter mask.

[0010] Furthermore, the specific steps for filtering and purifying the phase spectrum using the dynamic filter mask are as follows: First, the phase spectrum is averaged and aggregated along the channel dimension to obtain a single-channel phase spectrum; The sine and cosine components of the single-channel phase spectrum are extracted and spliced ​​together along the channel dimension to form the basic phase features of the dual channels; The basic phase characteristics of the dual channels are spatially multiplied element-wise with the two-dimensional dynamic filter mask to obtain the frequency anomaly characteristics of the dual channels.

[0011] Furthermore, the number of concentric ring-shaped binary masks is 16.

[0012] Furthermore, the frequency anomaly features of the dual channels are concatenated and fused with the basic features of the face image to be detected. A single-channel continuous spatial attention mask is then generated through a convolutional residual network. Specifically: The frequency anomaly features of the dual channels are concatenated with the three-channel RGB spatial basic features of the face image to be detected in the channel dimension to form a five-dimensional space-frequency fusion tensor. The five-dimensional space-frequency fusion tensor is sequentially input into a convolutional residual network containing an initial convolutional layer, multiple residual blocks, and a final convolutional layer. Initial convolution operation, multi-layer residual convolution operation, and final convolution operation are performed sequentially to extract deep space-frequency fusion features. Perform a Sigmoid activation normalization operation on the deep spatial-frequency fusion features to generate a single-channel spatial attention mask.

[0013] Furthermore, the residual modulation mechanism is used to perform weighted intervention and feature decoupling on the basic features of the face image to be detected, specifically as follows: The basic features of the face image to be detected are enhanced based on the residual modulation formula, which is: in, The basic features of the face image to be detected. This represents element-wise multiplication for spatial broadcasting along the channel dimension. It is an identity tensor of all 1s. The spatial attention mask, This is the core hyperparameter for controlling the probe injection intensity.

[0014] Furthermore, the decoupled features from the intervention are input into a classifier for classification, specifically as follows: The decoupled features after intervention are input into a visual backbone network to extract deep features. The visual backbone network is a convolutional neural network or a visual Transformer image encoder. After performing global average pooling on deep features, the data is input into a fully connected classification layer, which outputs the probability of forgery of the face image. The forgery probability is compared with a preset threshold, and the final result of the authenticity detection is output.

[0015] The present invention also provides a face image detection system based on frequency anomaly injection attention, the system comprising a frequency domain transformation module, a dynamic frequency filtering module, a spatial attention generation module, a feature intervention module, and a classification decision module; The frequency domain transformation module is used to extract the amplitude spectrum and phase spectrum of the face image; The dynamic frequency filtering module is used to generate a dynamic filtering mask and refine it to obtain frequency anomaly features. The spatial attention generation module is used to generate a spatial attention mask; The feature intervention module is used to achieve feature weighting intervention and decoupling through residual modulation; The classification decision module is used to output the true / false detection results.

[0016] The present invention discloses the following technical effects: This invention extracts frequency domain features, including amplitude and phase spectra, from the face image to be detected through preprocessing. A dynamic filtering mask is generated based on the amplitude spectrum, and the phase spectrum is purified by directional filtering. This accurately identifies clean frequency anomaly features related to deepfakes. The dual-channel frequency anomaly features are then concatenated and fused with the spatial basic features to generate a single-channel spatial attention mask. This spatial attention mask serves as an information bottleneck, and a residual modulation mechanism is used to weight and decouple the image's basic features. This proactively breaks the model's dependence on low-dimensional shortcut features within the domain, forcing the model to learn cross-domain, universal traces of forgery. Effective feature decoupling can be achieved without building a complex two-stream network architecture. While ensuring detection accuracy, this significantly improves cross-domain generalization ability and inference efficiency, ultimately outputting stable and reliable face image authenticity detection results. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the overall process architecture of the method of the present invention; Figure 3 The diagram shows the internal structure and calculation flowchart of the Frequency Anomaly Injection Attention (FAIA) module in this invention. Figure 4 This is a comparison diagram of singular value spectral lines in the depth feature space in an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] like Figures 1-4 As shown, this embodiment provides a face image detection method based on Frequency Anomaly Injection Attention (FAIA). In this embodiment, the method is implemented in the PyTorch deep learning framework, and the hardware environment uses an NVIDIA RTX 5070 GPU.

[0022] Step S1: Image preprocessing and frequency domain transformation Obtain the RGB format face image to be detected and resize it uniformly. , represented as input tensor To extract structured frequency features and reduce computational redundancy, the three-channel RGB image is first converted into a single-channel grayscale image. .

[0023] Subsequently, the grayscale image is transformed from the spatial domain to the complex frequency domain using a two-dimensional fast discrete Fourier transform (2D-FFT): For the acquired complex spectrum Polar coordinate decomposition is performed to separate the amplitude spectrum and phase spectrum in real form. Specifically, the amplitude spectrum... Reflects the brightness and contrast distribution, phase spectrum of the image. This implicitly encodes the object's edges and structural contours, calculated using the following formula: Step S2: Dynamic frequency filtering and phase purification (FAIA core) Because the original phase spectrum contains a large amount of natural high-frequency noise unrelated to forgery, direct use will lead to model performance degradation. This embodiment constructs a dynamic frequency filter "guided by the amplitude spectrum", which specifically includes the following sub-steps: S2.1 Bandwidth Energy Assessment: Construct 16 concentric ring-shaped binary masks to represent the amplitude spectrum. The frequency band is uniformly divided into 16 sub-bands from low to high frequency. The average frequency energy within each sub-band is calculated to form a 16-dimensional energy feature vector.

[0024] S2.2 Adaptive Weight Generation: The 16-dimensional energy feature vector is input into a lightweight multilayer perceptron (MLP) evaluator. This MLP adaptively learns and outputs a 16-dimensional frequency band weight vector. The learned weight vectors are remapped onto the corresponding two-dimensional concentric ring mask to generate a mask of size [size missing]. Smooth two-dimensional dynamic filter mask .

[0025] S2.3 Phase Precision Purification: Considering the periodicity of the phase, to reduce redundancy, the separated phase spectra are first averaged and aggregated along the channel dimension to obtain the single-channel phase spectrum. Extract their sinusoidal components respectively. Sum and cosine components These are then concatenated along the channel dimension to form a dual-channel basic phase feature. Finally, the dynamic filtering mask is used... Perform spatial element-wise multiplication filtering on it: in, This indicates splicing along the channel dimension. This represents element-wise multiplication. This yields a dual-channel pure frequency anomaly characteristic that completely filters out background natural noise. .

[0026] Step S3: Generate spatial attention mask The dual-channel frequency anomaly characteristics The three-channel RGB spatial features of the face image to be detected are directly concatenated along the channel dimension to form a five-dimensional space-frequency fusion tensor. This tensor contains interwoven spatial and frequency anomaly information.

[0027] Fusion tensors The input is fed into a feature transformation micro-network. In this embodiment, the micro-network sequentially includes: one initial convolutional layer (for channel dimensionality reduction and feature alignment), eight residual blocks (for deep feature extraction), and one final convolutional layer. After the final convolutional layer, the output features are normalized using a sigmoid activation function, and the collapsed output features range from [value range missing]. Single-channel continuous spatial attention mask between .

[0028] Step S4: Injecting information into the visual backbone network This invention uses the FAIA module as a plug-and-play probe, without disrupting the weight structure of the original visual backbone network (such as Xception or CLIP-Large image encoder). Enhanced feature image. Calculated using the residual modulation mechanism: Among them, input As a fundamental feature of space, This represents element-wise multiplication for spatial broadcasting along the channel dimension. It is an identity tensor of all 1s. This is the core hyperparameter for controlling the probe injection intensity.

[0029] In this computational mechanism, the identity tensor guarantees the identity mapping flow of the fundamental spatial information, while This step forces the allocation of extremely high activation gains to regions with abnormal frequencies (areas with a high incidence of forgery artifacts). This step constructs a strict information bottleneck in the shallow layers of the network, actively depriving the backbone network of the path to overfit low-dimensional "shortcut features" such as specific backgrounds, lighting conditions, or compression artifacts, and forcing the network to achieve feature decoupling in subsequent deep computations.

[0030] Step S5: Network Training and Classification Decision Features enhanced with frequency attention The input is fed into a dynamic normalization layer, and then enters the visual backbone network for forward propagation. The high-dimensional feature map output by the backbone network is flattened by a global average pooling layer and then fed into a fully connected layer for multi-dimensional mapping, finally outputting the predicted forgery probability value of the sample. .

[0031] During the network training phase, the model uses the standard binary cross-entropy loss function for end-to-end optimization. The loss function formula is as follows: in, The input samples are assigned their true physical labels (0 for real images, 1 for fake images). During the inference and decision-making phase, the predicted probabilities are used... The system determines the relationship between the value of the data and a set threshold (usually 0.5), and outputs the final result for determining whether the data is true or false.

[0032] Experimental Results Verification and Analysis To verify the effectiveness of the Frequency Anomaly Injected Attention (FAIA) detection method and system proposed in this invention, this embodiment conducted extensive zero-shot tests on the authoritative public dataset FaceForensics++ (FF++) and the extremely rigorous cross-domain dataset Celeb-DF-v2 (CDF) in the field of deep pseudo-detection.

[0033] 1. Experimental Environment and Parameter Settings Hardware environment: All experiments were conducted on a single commercial-grade NVIDIA GPU (such as the RTX 5070 series) to verify the industrial deployment feasibility of the present invention.

[0034] Software environment: The deep learning framework used is PyTorch, and OpenCV is used for image grayscale conversion and preprocessing.

[0035] Evaluation metrics: The area under the receiver operating characteristic curve (AUC), equal error rate (EER), and frames per second (FPS) were used as the main evaluation metrics.

[0036] 2. Cross-domain generalization ability test results To verify the model's generalization ability when faced with unknown generation algorithms, this embodiment trains the model on the FF++ dataset and tests it directly on the unknown domain CDF. Experimental results show: Performance breakthrough: When the FAIA module of this invention is injected into the CLIP-Large visual backbone network, the frame-level cross-domain AUC on CDF soars to 79.44%, and the EER drops significantly to 28.10%.

[0037] Comparative analysis: Compared to the unprotected CLIP-Large baseline model (AUC 74.38%), this invention achieves an improvement of over 5 percentage points. Furthermore, it surpasses the latest mainstream state-of-the-art methods (such as CFM's 78.39% and SAFE's 78.60%). This fully demonstrates that the FAIA module, acting as an information bottleneck, successfully deprives the network of its dependence on intra-domain shortcut features, forcing it to learn high-frequency forgery traces applicable across domains.

[0038] 3 Attention Calibration and Feature Decoupling Analysis This embodiment verifies the probe's internal mechanism through heatmap visualization and singular value decomposition (SVD): Attention calibration: Unprotected CNN baselines are highly susceptible to local overfitting, while ViT baselines exhibit severe attentional dissipation (focusing on background noise). After injecting the FAIA mask, the network attention is successfully pulled back and precisely aligned to the facial core fake regions that produce semantic inconsistencies.

[0039] Feature manifold unrolling: SVD singular value spectrum analysis of the extracted deep features shows that the singular values ​​of the baseline model exhibit a cliff-like decay (features fall into low-rank degradation); while the singular value distribution of the model of this invention shows a significant "long-tail effect", proving that the method achieves a high degree of feature decoupling and effectively expands the representation dimension of the feature manifold.

[0040] 4. Multidimensional robustness under extreme image degradation In the real world, videos often suffer from severe image quality compression and degradation during transmission. This example tests the model's resilience to degradation without any targeted data augmentation. Facing destructive super-resolution ( During the attack, the baseline model's AUC collapsed catastrophically (dropping to about 62%), while the system equipped with FAIA successfully salvaged the detection performance, maintaining it at a high level of close to 70%.

[0041] In brightness jitter and Gaussian blur tests, this system also demonstrated extremely high immunity. This proves that the structural anomaly features extracted by the dynamic frequency filtering mechanism have extremely strong anti-interference capabilities.

[0042] 5. Computational Efficiency and Deployment Feasibility Analysis Unlike existing dual-stream frequency domain sensing networks (which typically require adding large frequency branches), the FAIA of this invention is an extremely lightweight, plug-and-play probe: Extremely low parameter overhead: Compared to the single-stream baseline model, introducing the FAIA module only increases the number of parameters by about 0.66M, and the model size remains almost unchanged.

[0043] Extremely fast inference speed: Despite precise frequency domain filtering at full resolution, thanks to the high parallelism of the GPU, the system of this invention still maintains a real-time inference throughput of up to 125 FPS. This speed far exceeds the 112 FPS of complex dual-stream architectures (such as SRM networks), perfectly meeting the needs of industrial-grade real-time video stream deepfake detection.

[0044] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0045] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A face image detection method based on frequency anomaly injection attention, characterized in that, Includes the following steps: S1: Acquire the face image to be detected, preprocess the face image to extract frequency domain features, the frequency domain features include amplitude spectrum and phase spectrum; S2: Generate a dynamic filter mask based on the amplitude spectrum, and use the dynamic filter mask to filter and purify the phase spectrum to obtain the frequency anomaly characteristics of the dual channels; S3: The frequency anomaly features of the dual channels are spliced ​​and fused with the basic features of the face image to be detected, and a single-channel spatial attention mask is generated through a convolutional residual network; S4: Using the spatial attention mask as an information bottleneck, the basic features of the face image to be detected are weighted and decoupled through a residual modulation mechanism to obtain the features after intervention and decoupling. S5: Input the decoupled features after the intervention into the classifier for classification, and output the final true / false detection result; The frequency anomaly features of the dual-channel image are concatenated and fused with the basic features of the face image to be detected. A single-channel continuous spatial attention mask is then generated through a convolutional residual network. Specifically: The frequency anomaly features of the dual channels are concatenated with the three-channel RGB spatial basic features of the face image to be detected in the channel dimension to form a five-dimensional space-frequency fusion tensor. The five-dimensional space-frequency fusion tensor is sequentially input into a convolutional residual network containing an initial convolutional layer, multiple residual blocks, and a final convolutional layer. Initial convolution operation, multi-layer residual convolution operation, and final convolution operation are performed sequentially to extract deep space-frequency fusion features. Perform a Sigmoid activation normalization operation on the deep space-frequency fusion features to generate a single-channel spatial attention mask; The residual modulation mechanism is used to perform weighted intervention and feature decoupling on the basic features of the face image to be detected. Specifically, this involves: The basic features of the face image to be detected are enhanced based on the residual modulation formula, which is: in, The basic features of the face image to be detected. This represents element-wise multiplication for spatial broadcasting along the channel dimension. It is an identity tensor of all 1s. The spatial attention mask, This is the core hyperparameter for controlling the probe injection intensity.

2. The face image detection method based on frequency anomaly injection attention as described in claim 1, characterized in that, The preprocessing and extraction of the face image specifically involves: Convert the face image to be detected into a single-channel grayscale image; The single-channel grayscale image is transformed from the spatial domain to the complex frequency domain using a two-dimensional fast discrete Fourier transform. By performing polar coordinate decomposition on the complex spectrum, the amplitude spectrum and phase spectrum in real form are obtained.

3. The face image detection method based on frequency anomaly injection attention according to claim 1, characterized in that, The dynamic filter mask is generated based on the amplitude spectrum as follows: Construct multiple concentric ring-shaped binary masks; The amplitude spectrum is divided into multiple sub-bands with different frequency step sizes. The average frequency energy feature in each sub-band is calculated. The average frequency energy feature is then used to generate a band weight vector through adaptive learning. The frequency band weight vector is remapped onto the corresponding binary mask to generate a smooth two-dimensional dynamic filter mask.

4. The face image detection method based on frequency anomaly injection attention according to claim 1, characterized in that, The specific steps for filtering and purifying the phase spectrum using the dynamic filter mask are as follows: First, the phase spectrum is averaged and aggregated along the channel dimension to obtain a single-channel phase spectrum; The sine and cosine components of the single-channel phase spectrum are extracted and spliced ​​together along the channel dimension to form the basic phase features of the dual channels; The basic phase characteristics of the dual channels are spatially multiplied element-wise with the dynamic filter mask to obtain the frequency anomaly characteristics of the dual channels.

5. The face image detection method based on frequency anomaly injection attention according to claim 3, characterized in that, The number of concentric ring-shaped binary masks is 16.

6. The face image detection method based on frequency anomaly injection attention according to claim 1, characterized in that, The specific package for inputting the decoupled features from the intervention into the classifier for classification is as follows: The decoupled features after intervention are input into a visual backbone network to extract deep features. The visual backbone network is a convolutional neural network or a visual Transformer image encoder. After performing global average pooling on deep features, the data is input into a fully connected classification layer, which outputs the probability of forgery of the face image. The forgery probability is compared with a preset threshold, and the final result of the authenticity detection is output.

7. A face image detection system based on frequency anomaly injection attention, characterized in that, The system is used to perform the detection method as described in any one of claims 1 to 6, and the system includes a frequency domain transformation module, a dynamic frequency filtering module, a spatial attention generation module, a feature intervention module, and a classification decision module; The frequency domain transformation module is used to extract the amplitude spectrum and phase spectrum of the face image; The dynamic frequency filtering module is used to generate a dynamic filtering mask and refine it to obtain frequency anomaly features. The spatial attention generation module is used to generate a spatial attention mask; The feature intervention module is used to achieve feature weighting intervention and decoupling through residual modulation; The classification decision module is used to output the true / false detection results.

Citation Information

Patent Citations

  • Deep counterfeit image processing method and system based on frequency enhanced self-attention

    CN118115481A

  • Face image anti-counterfeiting detection method and device based on frequency spectrum reconstruction

    CN121354195A

  • Face living body detection method and system for multi-modal feature fusion

    CN122116492A