Makeup face fraud detection method and system based on frequency spectrum optimization and multi-level attention interaction

Through spectrum optimization and multi-level attention interaction, the problem of insufficient generalization performance in makeup face detection is solved, effective detection of makeup faces is achieved, and detection accuracy and robustness are improved.

CN120356267APending Publication Date: 2025-07-22SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510324389.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When faced with different degrees of makeup faces, the existing facial fraud detection technology has insufficient generalization performance, making it difficult to effectively distinguish between real faces and makeup faces, resulting in a decrease in detection accuracy.

Method used

Using spectrum optimization and multi-level attention interaction methods, makeup augmented samples are generated through makeup transfer algorithms, and the frequency domain characteristics of makeup area images are decomposed as amplitude components and phase components. Learnable tokens are introduced for optimization, and high-frequency texture features are extracted using adaptive high-pass filters, combined with multi-level attention modules for dynamic weighting. Finally, the makeup degree judgment threshold is obtained through comparison learning, and the binary classification prediction results are corrected.

Benefits of technology

It improves the generalization performance of the face fraud detection model for makeup and faces, improves the accuracy and robustness of detection, and reduces misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356267A_ABST
    Figure CN120356267A_ABST
Patent Text Reader

Abstract

The invention discloses a makeup face fraud detection method and system based on frequency spectrum optimization and multi-level attention interaction. The method comprises the following steps: extracting face images in each training domain data set; generating makeup augmentation samples with different makeup degrees through a makeup migration algorithm; detecting a face mark point of the makeup augmentation sample and extracting a key makeup area; decoupling the amplitude component and the phase component of the frequency domain feature of the makeup area image, and optimizing the frequency spectrum component; dynamically adjusting the area weight to obtain an overall high-frequency texture feature value; obtaining a makeup degree threshold value through comparative learning; according to the dichotomy loss of the dichotomy prediction branch and the comparison loss of the makeup degree judgment branch, carrying out weighted summation to obtain a total loss function, and training to obtain a makeup degree judgment threshold value; and based on the makeup degree judgment threshold, correcting the dichotomy prediction branch result, and obtaining a final dichotomy prediction probability. According to the invention, the generalization performance of face fraud detection in a makeup face application scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face fraud detection, and in particular to a makeup face fraud detection method and system based on spectrum optimization and multi-level attention interaction. Background Art

[0002] Face fraud refers to the act of deceiving the face recognition system by forging, simulating or using non-real facial information, achieving illegal access, impersonation or bypassing identity authentication. The main attack methods include photo printing, video playback and face masks. Most of the existing face fraud detection technologies have good in-library detection effects, but poor cross-library detection performance. The main reason is that the data inside and outside the library are often collected under different conditions, such as different shooting equipment, ambient lighting and presentation equipment, resulting in domain shift between the data inside and outside the library. Therefore, when the diversity of training data is insufficient, the model is prone to overfitting during the in-library learning process, resulting in poor generalization performance.

[0003] At present, makeup mainly includes foundation, eye shadow, lipstick and blush, and the degree of makeup varies with the makeup technique and color of the cosmetics. For special extreme heavy makeup (such as role-playing, etc.), since it can significantly modify the facial features, it is likely to be used by attackers to tamper with facial information. In order to improve the security of the system, this type of makeup should be regarded as a new type of attack; for the light makeup commonly seen in real life, although there are changes in facial texture and color of some areas like heavy makeup, this degree of makeup does not change the identity information of the face. If it is misjudged as fraud by the system, it will bring many inconveniences to the user. Therefore, the light makeup face should still be judged as a real face. It can be seen that different degrees of makeup faces will interfere with the system judgment to a certain extent. However, no face fraud detection technology has been specifically studied for the above problems. The existing methods are not robust and generalizable to different degrees of makeup faces, and it is difficult to meet the detection model's detection accuracy requirements in makeup face application scenarios. Summary of the invention

[0004] To overcome the defects and deficiencies of the existing technologies, the present invention provides a makeup face fraud detection method and system based on spectrum optimization and multi-level attention interaction. The present invention constructs a binary classification prediction branch and a makeup degree judgment branch, performs makeup augmentation on some input images based on a makeup transfer algorithm as input samples for the makeup degree judgment branch, and improves the generalization of the overall model for makeup faces through a feature extraction and optimization module. The method decomposes the image features of the makeup area into amplitude components and phase components in the frequency domain by means of spectrum optimization, and introduces learnable tokens to optimize each spectrum component respectively; designs an adaptive high-pass filter to extract the high-frequency texture information in the optimized image features; effectively integrates multi-level attention and dynamically weights the high-frequency texture features through the capture and integration of local channel attention and global spatial attention; and obtains a makeup degree judgment threshold through contrastive learning to measure the difference in high-frequency texture features between the input samples and the real samples so as to correct the binary classification prediction results.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] The present invention provides a makeup face fraud detection method based on spectrum optimization and multi-level attention interaction, including the following steps:

[0007] Divide the dataset, frame the videos in each dataset and extract the face images, and set the corresponding authenticity labels and domain labels;

[0008] Generate makeup augmentation samples with different makeup degrees through a makeup transfer algorithm, and set makeup degree labels;

[0009] Input the training domain data and the makeup augmentation samples into the binary classification prediction branch to output a binary classification loss;

[0010] Input the makeup augmentation samples into the makeup degree judgment branch, detect the face landmark points of the makeup augmentation samples, and extract the image of each makeup area;

[0011] Decouple the amplitude component and the phase component of the frequency domain features of the makeup area image, and perform optimization processing on the spectrum components;

[0012] Extract the high-frequency texture features of each makeup area based on an adaptive high-pass filter;

[0013] Dynamically weight the high-frequency texture features of each makeup area to obtain the overall high-frequency texture feature value of the makeup augmentation sample;

[0014] Construct positive and negative sample pairs according to the makeup degree labels, obtain a makeup degree threshold through contrastive learning, and output a contrastive loss;

[0015] The total loss function is obtained by weighted summation of the binary classification loss and the contrastive loss, and training is performed based on the total loss function to obtain the judgment threshold for the makeup degree.

[0016] Based on the judgment threshold for the makeup degree, correct the result of the binary classification prediction branch and output the binary classification prediction probability.

[0017] As a preferred technical solution, decouple the amplitude component and the phase component of the frequency domain features of the makeup area image, and optimize the spectral components, specifically including:

[0018] Decompose the frequency domain features of the makeup area image into an amplitude component and a phase component, respectively introduce learnable tokens, calculate the similarity matrix between the two spectral components and the corresponding tokens, use it as the attention weight to weight the tokens, and then add them to the spectral components, and output the optimized spectral components through a multi-layer perceptron. After optimizing the spectral components, the frequency domain features are obtained.

[0019] As a preferred technical solution, decompose the frequency domain features of the makeup area image into an amplitude component and a phase component, which is expressed as:

[0020]

[0021] where M and N are the width and height of the input area image, u and v are the frequency indices in the frequency domain, f(x,y) is the pixel value of the input image in the spatial domain, F(u,v) is the corresponding complex number in the frequency domain, representing the frequency domain features, Re(F(u,v)) is the real part of F(u,v), Im(F(u,v)) is the imaginary part of F(u,v), atan2 is the two-variable arctangent function, |F(u,v)| and θ(u,v) respectively represent the amplitude component and the phase component of the input image in the frequency domain;

[0022] Introduce a learnable token to the amplitude component and the phase component respectively, denoted as T α and T ρ , calculate the similarity matrix between each spectral component and its corresponding token through dot product, and perform a normalization operation on the similarity matrix of the amplitude component, which is expressed as:

[0023]

[0024] where d represents the number of image feature channels, Norm represents the normalization operation, A α and A ρ respectively represent the similarity matrices between the amplitude component and the phase component and their corresponding tokens;

[0025] Use Softmax to respectively process the two similarity matrices A α and Aρ Convert it into a probability distribution between 0 and 1, which is used as the attention weight and acts on T α and T ρ Obtain the weighted token, denoted as and Expressed as:

[0026]

[0027] After adding the amplitude component α and the phase component ρ to respectively, the optimized spectral component is obtained through the output of a multi-layer perceptron, denoted as and Expressed as:

[0028]

[0029] Combine the two optimized spectral components and to obtain the optimized frequency-domain feature, denoted as The formula is expressed as:

[0030]

[0031] where j represents the imaginary unit.

[0032] As a preferred technical solution, the high-frequency texture features of each makeup area are extracted based on a high-pass filter, specifically including:

[0033] Multiply each frequency point by an alternating positive and negative matrix (-1) in the frequency domain u+v to obtain the frequency-domain feature after spectral centering, denoted as Expressed as:

[0034]

[0035] The high-pass filter performs high-pass filtering on the frequency-domain feature after spectral centering Calculate the amplitude of the filtered frequency-domain feature and normalize it to obtain the high-frequency texture feature value of the regional image, specifically expressed as:

[0036]

[0037] where represents the high-frequency feature obtained after filtering, is the real part of is the imaginary part of represents the amplitude component of, α HFis the normalized high-frequency amplitude, that is, the high-frequency texture feature value of the regional image.

[0038] As a preferred technical solution, the cut-off frequency radius of the high-pass filter is:

[0039]

[0040] where k is the cut-off frequency radius coefficient, and M and N are the width and height of the input regional image;

[0041] Calculate the distance from each position in the high-pass filter to its center, compare it with the cut-off frequency radius, set the positions greater than the cut-off frequency radius to 1 to retain the high-frequency components, and otherwise set them to 0 to filter out the low-frequency components, which is expressed as:

[0042]

[0043] where D(u, v) is the distance from the point (u, v) in the frequency domain to the center point of the spectrum and H(u, v) is the binary form frequency response function of the high-pass filter.

[0044] As a preferred technical solution, dynamically weight the high-frequency texture features of each makeup area to obtain the high-frequency texture feature value of the overall makeup augmented sample, specifically including:

[0045] Perform Patch Embedding on the input images of each makeup area, convert the images into information units that can be understood and processed by the model, denoted as x, and send them into the local branch and the global branch respectively;

[0046] The local branch is used to capture channel attention and enhance cross-channel interaction, which is expressed as:

[0047]

[0048] where W 1×1 represents a 1×1 convolution, which is used to adjust the channel dimension, CS represents channel randomization, and x represents the output after channel randomization;

[0049] After preliminary weighting through the channel attention mechanism, adjust the weight of the feature intensity of each channel to obtain the local channel attention, which is expressed as:

[0050]

[0051] where GlobalAvgPool is the global average pooling layer, FC1 and FC2 are two fully connected layers, and F local is the output local channel attention;

[0052] The global branch is used to capture spatial attention and perform information aggregation, expressed as:

[0053] σ(x) = LN(AvgPool(GELU(Linear(x))))

[0054] Wherein, Linear represents a linear layer, GELU represents the GELU activation function adopted, AvgPool represents an average pooling layer, LN represents layer normalization, and σ(x) represents the features output by the global branch;

[0055] Perform a linear transformation on σ(x) to obtain Q, K, and V, and introduce a learnable Query Embedding, denoted as QE, into Q. The global attention mechanism is expressed as:

[0056] Q = σ(x)W Q +QE, K = σ(x)W K , V = σ(x)W V

[0057]

[0058] F global = A × V

[0059] Wherein, W Q 、W K and W V are all learnable weight matrices, d k is the feature dimension after weight mapping, A is the attention weight after Softmax normalization, and F global represents the global spatial attention;

[0060] Use the global spatial attention calculated by the Sigmoid activation function as a gate to act on the local channel attention to optimize the channel attention result, and output the channel and spatial multi-level attention F all , expressed as:

[0061] F all = Sigmoid(F global ) × F local

[0062] Based on the channel and spatial multi-level attention F all Dynamically weight each makeup area image, and output the high-frequency texture feature value of the overall makeup augmented sample.

[0063] As a preferred technical solution, in the local branch, the operations of channel randomness specifically include:

[0064] The features after convolution are divided into groups along the channel dimension. Each group uses depthwise separable convolution (DWConv) for per-channel convolution and pointwise convolution. Per-channel convolution provides an independent processing method for each channel, and pointwise convolution combines independent channel features through random selection;

[0065] In the global branch, before feature pooling, a Linear layer and the GELU activation function are used for mapping and activation, and after feature pooling, layer normalization is used to normalize the output.

[0066] As a preferred technical solution, positive and negative sample pairs are constructed according to the makeup degree label, and the makeup degree threshold is obtained through contrastive learning, and the contrastive loss is output, specifically including:

[0067]

[0068] Among them, B is the number of samples in a batch during model training, u i and v i respectively represent the high-frequency texture features of two samples whose distances are to be calculated in the batch, D(u, v) represents the distance value obtained through normalized Euclidean distance calculation, briefly denoted as D, V[u i ,v i represents the variance between u i and v i , y i is the sample pair label, m is the minimum distance of the negative sample pair, and L con represents the contrastive loss.

[0069] As a preferred technical solution, training is performed based on the total loss function, and the makeup degree judgment threshold is obtained through training, specifically including:

[0070] Statistical average distances of positive and negative sample pairs, and calculate their mean as the makeup degree threshold T, expressed as:

[0071]

[0072] Among them, D p represents the average distance of positive sample pairs, D n represents the average distance of negative sample pairs;

[0073] Based on the makeup degree judgment threshold, correct the results of the binary classification prediction branch and output the binary classification prediction probability, specifically including:

[0074] When the difference in high-frequency texture features between the sample and the true sample is less than T, it is determined that the makeup degree of the sample belongs to the normal range;

[0075] When the difference in high-frequency texture features between the sample and the real sample is greater than T, it is determined that the makeup degree of the sample is relatively thick, and the prediction result is corrected, which is expressed as:

[0076] P adj =(1 - β)×P bin +β×I(Δα test )

[0077] Among them, P bin and P adj respectively represent the predicted values of the binary classification branches before and after correction. β is the adjustment coefficient, and I(Δα test ) is the correction direction, which is expressed as:

[0078]

[0079] Among them, Δα test represents the difference in high-frequency texture features between the sample and the real sample.

[0080] The present invention also provides a makeup face fraud detection system based on spectrum optimization and multi-level attention interaction, including: a data set preprocessing module, a makeup augmentation module, a binary classification loss output module, a makeup area extraction module, a spectrum decomposition optimization module, a high-frequency texture feature extraction module, a multi-level attention interaction module, a makeup degree threshold update module, a total loss function construction module, a model training module, and a prediction result output module;

[0081] The dataset preprocessing module is used to divide the dataset, frame the videos in each dataset and extract face images, and set corresponding authenticity labels and domain labels; the makeup augmentation module is used to generate makeup augmentation samples with different makeup degrees through the makeup transfer algorithm and set makeup degree labels; the binary classification loss output module is used to input the training domain data and makeup augmentation samples into the binary classification prediction branch and output the binary classification loss; the makeup area extraction module is used to input the makeup augmentation samples into the makeup degree judgment branch, detect the face landmark points of the makeup augmentation samples, and extract the images of each makeup area; the spectral decomposition optimization module is used to decouple the amplitude component and phase component of the frequency domain features of the makeup area images and perform optimization processing on the spectral components; the high-frequency texture feature extraction module is used to extract the high-frequency texture features of each makeup area based on a high-pass filter; the multi-level attention interaction module is used to dynamically weight the high-frequency texture features of each makeup area to obtain the high-frequency texture feature value of the overall makeup augmentation sample; the makeup degree threshold update module is used to construct positive and negative sample pairs according to the makeup degree labels, obtain the makeup degree threshold through contrastive learning, and output the contrastive loss; the total loss function construction module is used to obtain the total loss function by weighted summation of the binary classification loss and the contrastive loss; the model training module is used to train based on the total loss function to obtain the makeup degree judgment threshold; the prediction result output module is used to correct the result of the binary classification prediction branch based on the makeup degree judgment threshold and output the binary classification prediction probability.

[0082] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0083] (1) The present invention performs makeup augmentation on some input images based on the makeup transfer algorithm and uses them as input samples for the makeup degree judgment branch to solve the problem that the number and diversity of makeup samples in the existing dataset are insufficient for model training, and improves the generalization of the overall model to makeup faces through subsequent feature extraction and optimization modules.

[0084] (2) Through the method of spectrum decomposition optimization, the present invention decomposes the input makeup area image features into amplitude components and phase components in the frequency domain, that is, decouples the texture features and structural features in the image. Both real faces and corresponding makeup faces have depth information and consistent facial feature position information, so their structural features are similar; while different levels of makeup will introduce different information such as colors and brightness, and there will be significant differences in their texture features from real faces. Therefore, spectrum decomposition is beneficial to analyzing the essential differences between real faces and makeup faces with different levels; on this basis, learnable tokens are introduced to optimize the two spectrum components respectively, and the spectrum components are optimized to obtain frequency domain features. In addition, the present invention proposes an adaptive high-frequency texture feature extraction module, which designs a matching high-pass filter for makeup area images of different sizes to calculate the high-frequency texture features of each area.

[0085] (3) The present invention proposes a multi-level attention interaction module for extracting and aggregating local channel information and global spatial information. Among them, global spatial attention acts as a gate on local channel attention to guide the model to enhance or suppress local features, obtaining channel-space multi-level attention, realizing the organic integration of global information and local details. Combining spatio-frequency domain fusion, the channel-space multi-level attention is used to dynamically weight the high-frequency texture features of each makeup area, thereby obtaining the overall high-frequency texture feature value.

[0086] (4) The present invention constructs a makeup degree threshold update module, which trains a makeup degree threshold through contrastive learning to measure the difference in high-frequency texture features between the input sample and the real sample, and corrects the prediction value of the binary classification branch according to the comparison result. The model learns the influence of the makeup intensity on the fraud detection binary classification result, enabling the makeup degree judgment branch proposed by the present invention to more effectively assist the binary classification branch, thereby improving the generalization of the overall model to makeup faces. Brief Description of the Drawings

[0087] Figure 1 It is a schematic flow chart of the makeup face fraud detection method based on spectrum optimization and multi-level attention interaction of the present invention;

[0088] Figure 2 It is a schematic overall architecture diagram of the makeup face fraud detection system based on spectrum optimization and multi-level attention interaction of the present invention;

[0089] Figure 3 It is a schematic structural diagram of the spectrum decomposition optimization module of the present invention;

[0090] Figure 4 It is a schematic structural diagram of the adaptive high-frequency texture feature extraction module of the present invention;

[0091] Figure 5 It is a structural schematic diagram of the multi-level attention interaction module of the present invention. DETAILED DESCRIPTION

[0092] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0093] Example 1

[0094] This embodiment uses the face fraud detection datasets Oulu-NPU, CASIA-FASD, Replay-Attack and MSU-MFSD for training and testing, wherein some makeup samples in the public dataset MT-Dataset in the field of makeup detection migration and high-definition makeup face images collected from the Internet are used as reference makeup samples, and some real faces without makeup in Oulu-NPU, CASIA-FASD, Replay-Attack and MSU-MFSD are augmented with makeup through a typical makeup migration algorithm to generate samples with different degrees of makeup.

[0095] Among them, the Oulu-NPU dataset contains 55 subjects, collected in three scenarios, with a total of 990 real videos and 3960 fraud attack videos, the attack methods are paper printing attack and replay attack, and are shot with 6 types of mobile phones; the CASIA-FASD dataset contains 300 real videos and 300 fraud attack videos, collected from 50 subjects, shot with a high-resolution Sony NEX-5 camera and a low-quality USB camera, the attack methods include distorted photos, cut photos and video replay, including a variety of postures and expression changes; the Replay-Attack dataset contains 50 subjects, with a total of 200 real videos and 1000 fraud attack videos, the attack methods are A4 paper printing attack, iPhone replay attack and iPad replay attack, shot with Mac front camera; the MSU-MFSD dataset is collected from 35 subjects, including 70 real videos and 210 fraud videos, shot with MacBook Air The images were taken with the built-in camera of a 13-inch laptop and the front camera of a Google Nexus 5 Android phone; MT-Dataset is a dataset in the field of makeup transfer, with a total of 2719 makeup face images and 1114 non-makeup face images, which were collected by web crawlers. In addition, due to the lack of a public makeup face library, this embodiment uses two makeup transfer methods, BeautyGAN (Instance-level Facial Makeup Transfer with Deep Generative Adversarial Network) and EleGANt (Exquisite and Locally Editable GAN for Makeup Transfer), and uses some makeup samples in MT-Dataset as reference makeup samples to perform makeup transfer on some real faces without makeup in the Oulu-NPU, CASIA-FASD, Replay-Attack, and MSU-MFSD datasets to obtain samples with different degrees of makeup. This example is carried out on a Linux system and is implemented based on the deep learning framework Pytorch. The graphics card used in the experiment is GTX1080Ti and the CUDA version is 11.4.

[0096] like Figure 1 As shown, this embodiment provides a makeup face fraud detection method based on spectrum optimization and multi-level attention interaction, comprising the following steps:

[0097] S1: Divide the data set into a training set, a validation set, and a test set, divide the videos in each data set into frames, crop the face area in each frame image to obtain the face image, and set the corresponding authenticity label and domain label for it;

[0098] In this embodiment, the data set is divided into a training set, a validation set and a test set. Specifically, the Oulu-NPU dataset is divided into training set, validation set and test set in the ratio of 4:3:4, the CASIA-FASD dataset is divided into training set, validation set and test set in the ratio of 2:1:2, the Replay-Attack dataset is divided into training set, validation set and test set in the ratio of 4:3:4, and the MSU-MFSD dataset is divided into training set, validation set and test set in the ratio of 3:2:2. The VideoCapture class in the OpenCV open source software library is used to decode and frame the videos of each dataset. The Dlib library face detector get_frontal_face_detector is used to perform face recognition on each frame of video image, obtain the four coordinate values of the face area box, crop the video frame image according to the face area box, and adjust the image size to 256×256. The cropped face image set of each video is saved as a separate folder to prevent different videos from interfering with each other. The authenticity label of the real face sample is set to 1, and the authenticity label of the attack face sample is set to 0.

[0099] S2: Select some real face images without makeup in the training domain data, and use the makeup face images with the face area cropped out from the makeup transfer dataset MT-Dataset as reference samples. Generate makeup augmented samples with different degrees of makeup through the makeup transfer algorithm, and set corresponding makeup degree labels;

[0100] In this embodiment, for some real face images without makeup in the training set of the Oulu-NPU dataset, the makeup face images in the MT-Dataset are used as reference makeup samples, and the BeautyGAN and EleGANt methods are used to perform makeup migration respectively, and make-up face images of different degrees are generated and resized to 256×256, where the makeup degree label of the light makeup sample is 0, and the makeup degree label of the heavy makeup sample is 1. The makeup face image set corresponding to each video face image set is saved as an independent folder, and each makeup face image set is added to the training set of the Oulu-NPU dataset. For the real face images without makeup in the validation set and test set of the Oulu-NPU dataset, the high-definition makeup face images collected from the Internet are used as reference makeup samples, and the same method is used to perform makeup migration and resize, and similarly added to the validation set and test set of the Oulu-NPU dataset; the same operation is performed on the CASIA-FASD, Replay-Attack, and MSU-MFSD datasets, where the high-definition makeup face images referenced by the validation sets and test sets of different datasets are different;

[0101] In this embodiment, the training set, validation set, and test set respectively include the real faces in each data set and the light makeup and heavy makeup faces obtained by augmenting some of the real faces through makeup transfer algorithms such as BeautyGAN and EleGANt;

[0102] In this embodiment, during model training, the training domain data and the corresponding makeup-augmented samples are input into a binary classification prediction branch based on various face fraud detection algorithms, and the makeup-augmented samples are input into the makeup degree judgment branch proposed by the present invention. The same is true for model testing.

[0103] S3: Use the Dlib library to detect facial landmarks of makeup augmentation samples, extract key makeup areas such as eyes (including left and right eyes) and lips, and expand them accordingly with the set coefficients to obtain makeup areas including the eye and lip areas;

[0104] In this embodiment, the image paths of the makeup samples in the training set are traversed and read. The Dlib library face detector get_frontal_face_detector is used to perform face recognition on the images, and face landmark points are extracted. In this embodiment, the number of face landmark points selected is 68. The minimum bounding rectangle of the left eye is extracted according to the left eye landmark point set (36 - 41). Keeping the center position of the bounding rectangle frame unchanged, the minimum bounding rectangle frame of the left eye is enlarged with a certain expansion coefficient, thereby obtaining the makeup area image of the left eye (including the periorbital area). The same method is applied to the right eye (landmark point set is 42 - 47) and the lips (landmark point set is 48 - 59) to obtain the makeup area images of the right eye (including the periorbital area) and the lips (including the perioral area). Among them, the expansion coefficient of the minimum bounding rectangle frame of both eyes is 1.8, and the expansion coefficient of the minimum bounding rectangle frame of the lips is 1.3.

[0105] S4: Construct a spectral decomposition optimization module to decompose the frequency domain features of the makeup area image into amplitude components and phase components. Learnable tokens are respectively introduced, and the similarity matrices between the two spectral components and the corresponding tokens are calculated. After weighting the tokens with the similarity matrices as attention weights and adding them to the spectral components, the optimized spectral components are output through a multi-layer perceptron to perform optimization processing on the spectral components.

[0106] In this embodiment, the spectral decomposition optimization module is used to decouple the texture information and structural information of the makeup area image features, promote the model to focus on learning the texture information more relevant to the fraud detection task, and enhance the capture of key texture information in the makeup area. As Figure 3 shown, the specific steps of this module are as follows:

[0107] First, perform Fourier transform on each key makeup area image, and decompose the image features into amplitude components and phase components in the frequency domain, which are respectively expressed as:

[0108]

[0109] θ(u, V) = atan2(Im(F(u, v)), Re(F(u, v)))

[0110] Among them, M and N are the width and height of the input region image, u and v are the frequency indices in the frequency domain, f(x, y) is the pixel value of the input image in the spatial domain, F(u, v) is the complex number in the corresponding frequency domain, representing the frequency domain features, Re(F(u, v)) is the real part of F(u, v), Im(F(u, v)) is the imaginary part of F(u, v), atan2 is the two-variable arctangent function, |F(u, v)| and θ(u, v) respectively represent the amplitude component (abbreviated as α) and phase component (abbreviated as ρ) of the input image in the frequency domain. The amplitude component α characterizes the energy distribution of the image and reflects the texture information; the phase component ρ characterizes the spatial layout of the image and reflects the structural information.

[0111] Introduce a learnable token to α and ρ respectively, denoted as T α and T ρ , that is, a set of feature vectors initialized to random values and continuously updated during model training, used to optimize the above spectral components and enhance the model's capture of key texture information in the makeup area. Specifically, calculate the similarity matrix between each spectral component and its corresponding token through dot product, and train and optimize it so that the model gradually learns the makeup texture features related to the task. In addition, this module only normalizes the similarity matrix of the amplitude component, aiming to weaken the deviation in the amplitude component of image features between different domains to alleviate domain shift. The formula is expressed as:

[0112]

[0113] Among them, d represents the number of image feature channels, Norm represents the normalization operation, A α and A ρ respectively represent the similarity matrices between the amplitude component and the phase component and their corresponding tokens.

[0114] Use Softmax to convert the two similarity matrices A α and A ρ into probability distributions between 0 and 1, as attention weights and act on T α and T ρ to obtain the weighted tokens, denoted as and Expressed as:

[0115]

[0116] After adding the spectral components α and ρ to respectively, and passing through a Multilayer Perceptron (MLP), the optimized spectral components are output, denoted as and Specifically expressed as:

[0117]

[0118] The original input features may contain redundant information or be insufficient to characterize the key features. Through parameter learning, the MLP automatically adjusts the weights and relationships of the input features to obtain a feature representation that is more suitable for the current task. For the two optimized spectral components and are combined to obtain the optimized frequency-domain features, denoted as The formula is expressed as:

[0119]

[0120] S5: Construct an adaptive high-frequency texture feature extraction module to perform Fourier transform and spectral centering on the optimized features, and design a matching high-pass filter according to the size of the input makeup area image to extract the high-frequency texture features of each makeup area;

[0121] In this embodiment, the adaptive high-frequency texture feature extraction module is used to extract the high-frequency texture features of each makeup area image. As Figure 4 shown, the specific steps of this module are as follows:

[0122] First, to make the spectrum symmetric for subsequent filtering, the zero-frequency component in the optimized frequency-domain features is moved to the center position through spectral centering, and at the same time, the high-frequency components are rearranged so that they are distributed in the edge area. Specifically, by multiplying each frequency point by an alternating positive and negative matrix (-1) in the frequency domain u+v to obtain the frequency-domain features after spectral centering, denoted as The formula is expressed as:

[0123]

[0124] Design a matching high-pass filter for the input makeup area image, denoted as H(u, v), whose cut-off frequency radius (denoted as D0) depends on the size of the area image. Considering that the frequency distribution of the image is closely related to its size, the minimum dimension of the image (i.e., the smaller value of the width or height) is used to calculate the cut-off frequency, aiming to ensure that the design of the filter frequency response can adapt to the area images with a large aspect ratio difference in this embodiment and avoid the overinfluence of the longer side of the area image on the filtering effect. The cut-off frequency radius D0 is expressed as:

[0125]

[0126] where k is the cut-off frequency radius coefficient, and a smaller value should be selected to retain as much high-frequency information as possible. In this embodiment, k = 0.1 is selected.

[0127] Calculate the distance from each position in the filter to its center, and compare it with the cut-off frequency radius. The positions greater than the cut-off frequency radius are set to 1 to retain the high-frequency components, and vice versa to 0 to filter out the low-frequency components, which is specifically expressed as:

[0128]

[0129] where D(u, v) is the distance from the point (u, v) in the frequency domain to the center point of the spectrum and H(u, v) is the binary form frequency response function of the high-pass filter.

[0130] Apply the above high-pass filter to the frequency domain features after spectrum centering for high-pass filtering, calculate the amplitude of the filtered frequency domain features and normalize it to obtain the high-frequency texture feature values of the regional image, which is expressed by the formula:

[0131]

[0132] where represents the high-frequency features obtained after filtering, is the real part of is the imaginary part of represents the amplitude component of, α HF is the normalized high-frequency amplitude, that is, the high-frequency texture feature value of the above regional image.

[0133] S6: Construct a multi-level attention interaction module, where the local branch is used to capture channel attention and enhance channel interaction, and the global branch is used to capture spatial attention and perform gating processing on the local channel attention, realizing the fusion of channel and spatial attention, and dynamically adjusting the weights of each makeup area to obtain the high-frequency texture feature values of the overall makeup augmented sample;

[0134] In this embodiment, the multi-level attention interaction module is used to capture and integrate channel and spatial attention, and is used to dynamically weight the high-frequency texture features of each region to obtain the high-frequency texture feature values of the overall makeup augmented sample. As Figure 5 shown, this module includes two branches of local channel attention and global spatial attention, specifically including:

[0135] First, perform Patch Embedding on each input makeup area image to convert the image into information units that can be understood and processed by the model, denoted as x, and send them into the local branch and global branch of this module respectively.

[0136] The local branch is used to capture channel attention and enhance cross-channel interaction, which consists of operations such as 1×1 convolution, channel randomization, and channel weighting, and is specifically expressed as:

[0137]

[0138] Among them, W 1×1 represents 1×1 convolution, which is used to adjust the channel dimension for subsequent channel grouping; CS represents channel randomization, represents the output after channel randomization.

[0139] In channel randomization, first, the convolutional features are divided into groups along the channel dimension, and each group uses depthwise separable convolution DWConv for per-channel convolution and pointwise convolution. Among them, per-channel convolution provides an independent processing method for each channel, while pointwise convolution combines these independent channel features through random selection. The above method further breaks the fixed association between channels by adding a randomization mechanism.

[0140] Finally, it is preliminarily weighted through the channel attention mechanism to adjust the weight of the feature intensity of each channel, enhance important channels, and suppress unimportant channels to obtain local channel attention. This process consists of a global average pooling layer, two fully connected layers, and a Sigmoid activation function, and the formula is expressed as follows:

[0141]

[0142] Among them, GlobalAvgPool is the global average pooling layer, FC1 and FC2 are the two fully connected layers respectively, and F local is the output local channel attention.

[0143] The global branch is used to capture spatial attention and perform information aggregation. Specifically, first, the input features are activated and pooled. This process consists of a linear layer, an activation layer, an average pooling layer, and layer normalization operations, and is expressed as:

[0144] σ(x) = LN(AvgPool(GELU(Linear(x))))

[0145] Among them, Linear represents the linear layer, GELU represents the GELU activation function (Gaussian Error Linear Unit) used, AvgPool represents the average pooling layer, LN represents layer normalization (Layer Norm), and σ(x) represents the feature obtained after the input feature x passes through the above activation and pooling operations.

[0146] In this embodiment, since average pooling may lose some information in the image features, before feature pooling, a Linear layer and a GELU activation function are used for mapping and activation to compress and extract useful information in advance. After pooling, layer normalization is used to normalize the output, so that the variance of the features is consistent, avoiding the problem of unstable training caused by inconsistent feature distributions.

[0147] On this basis, a linear transformation is performed on σ(x) to obtain Query (denoted as Q), Key (denoted as K), and Value (denoted as V) for calculating the attention output. In addition, a learnable Query Embedding (denoted as QE) is introduced into Q, and QE is broadcast to each position of Q to enhance its adaptability and expressive ability, making the attention mechanism more sensitive to the current specific task. The global attention mechanism can be expressed as:

[0148] Q = σ(x)W Q +QE, K = σ(x)W K , V = σ(x)W V

[0149]

[0150] F global = A × V

[0151] Among them, W Q , W K and W V are all learnable weight matrices, d k is the feature dimension after weight mapping, A is the attention weight after Softmax normalization, and the global spatial attention F global is obtained by weighting V with A.

[0152] Finally, the channel and spatial attentions are organically fused. The global spatial attention is calculated by the Sigmoid activation function and used as a gate to act on the local channel attention to guide the enhancement or suppression of local features, further optimizing the channel attention result from a global level, and outputting the channel and spatial multi-level attention (denoted as F all ), which is specifically expressed as:

[0153] F all = Sigmoid(F globa1 ) × F local

[0154] F all is used to dynamically weight each makeup area image, and thus the overall high-frequency texture feature value of the makeup sample image is output, which is expressed as follows:

[0155] α face = F all_leye α HF_leye + F all_reye α HF_reye + F all_lip α HF_lip

[0156] Among them, F all_leye , F all_reye and F all_lip are respectively the channel and spatial multi-level attentions of the left eye, right eye and lips in the makeup sample during model training, and α HG_leye , α HF_reye and α HF_lip are respectively the high-frequency texture feature values of the left eye, right eye and lips in the makeup sample, and α face is the overall high-frequency texture feature value of the makeup sample obtained after attention weighting in the corresponding area.

[0157] S7: Construct a makeup degree threshold update module, construct positive and negative sample pairs according to the makeup degree labels of the samples, obtain the makeup degree threshold through contrastive learning during the training process, correct the binary classification prediction result accordingly, and output the contrastive loss L con ;

[0158] In this embodiment, the makeup degree threshold is used to measure the high-frequency texture feature difference between the input sample and the real sample. This module is the key module connecting the binary classification prediction branch and the makeup degree judgment branch proposed by the present invention, and specifically includes:

[0159] First, construct positive and negative sample pairs according to the makeup degree labels between the samples. Specifically, the features of the light makeup sample and the real face should be relatively close, and the features of the heavy makeup sample and other heavy makeup samples should also be similar. Therefore, the light makeup and the real face, and the heavy makeup and the heavy makeup are regarded as positive sample pairs. Similarly, the heavy makeup and the real face, and the heavy makeup and the light makeup are regarded as negative sample pairs. On this basis, the distance between similar samples is minimized and the distance between different categories is maximized through the contrastive loss.

[0160] In this embodiment, the standardized Euclidean distance is used to enhance the contrast effect of the high-frequency texture features. The standardization process can measure the distance between samples more fairly, make the distance measurement between positive and negative sample pairs more distinguishable, thus helping the model to better distinguish between heavy makeup and light makeup samples and providing a stable basis for contrastive learning. Specifically expressed as:

[0161]

[0162] Among them, B is the number of samples in a batch during model training, u i and v irespectively represent the high-frequency texture features of two samples for which distances are to be calculated in a batch, D(u, v) represents the distance value obtained through the calculation of the standardized Euclidean distance, briefly denoted as D, y i is the label of the sample pair (the label of the same-class sample pair y i = 1, the label of the different-class sample pair y i = 0), m is the minimum distance of the negative sample pairs, in this embodiment, m = 1.0, L con represents the contrast loss, V[u i , v i represents the variance between u i and v i , respectively represented by x1 and x2 for u i and v i , specifically expressed as:

[0163]

[0164] wherein, x i is the high-frequency texture feature of each of the above samples, i is its index subscript, x1 and x2 respectively represent the high-frequency texture features of the two samples in the [u i , v i sample pair, is the mean of x1 and x2.

[0165] During the training process, the relationship between sample pairs is optimized under the constraint of the above contrast loss. The model adjusts the feature distances between sample pairs during each iteration, making the distances of positive sample pairs as close as possible and the distances of negative sample pairs as large as possible.

[0166] S8: Obtain the total loss function (denoted as L all ) by weighted summing the binary classification loss of the binary classification prediction branch and the contrast loss of the makeup degree judgment branch. The specific formula is:

[0167] L all = ω1 × L bin + ω2 × L con

[0168] wherein, L bin , L con are respectively the binary classification loss of the binary classification prediction branch and the contrast loss of the makeup degree judgment branch, ω1, ω2 are the weights of the corresponding losses, and in this embodiment, both ω1 and ω2 are taken as 1.

[0169] S9: Train the overall model based on the total loss function L all , iteratively update and save each network parameter in the model, etc.;

[0170] In this embodiment, the Adam optimizer is used as the model training optimizer. The initial learning rate is 0.0001, the exponential decay rate of the first moment estimation is 0.9, and the exponential decay rate of the second moment estimation is 0.999. The network weights are updated to minimize the total loss function, and the network model and the best weights are saved after training is completed.

[0171] S10: During model testing, the test set samples are input into the trained network for feature extraction and prediction. The two-class prediction branch result is corrected using the makeup degree judgment threshold obtained from model training, and the final two-class prediction probability is obtained and the prediction result is output.

[0172] In this embodiment, after the model training is completed, the average distances of the positive sample pairs and the negative sample pairs are statistically calculated and denoted as D p and D n , and their mean value is used as the makeup degree threshold, denoted as T, which is specifically expressed as:

[0173]

[0174] T is dynamically learned through the above process and is used to measure the high-frequency texture feature difference between the input sample and the real sample during the test phase, and thus judge the makeup degree of the sample. The purpose of judging the makeup degree is not to semantically distinguish between light makeup and heavy makeup, but to judge whether the makeup degree of the current sample will interfere with the two-classification result of fraud detection. Specifically, when the high-frequency texture feature difference (denoted as Δα test ) between the sample and the real sample is less than T, it means that the makeup degree of the sample belongs to the normal range and should not be recognized as an attack by the model; while when Δα test is greater than T, it means that the makeup degree of the sample is relatively heavy and may cause an attack, that is, it may affect the two-classification. At this time, the prediction result of the face fraud detection branch needs to be corrected. Combining the setting of the sample authenticity label, the corrected two-class prediction value should be closer to 0, and vice versa, which is specifically expressed as:

[0175] P adj =(1 - β)×P bin +β×I(Δα test )

[0176] where P bin and P adj respectively represent the prediction values of the two-classification branches before and after correction, β is the adjustment coefficient that controls the intensity of correction, and its value range is (0, 1). I(Δα test ) is the correction direction, which is defined as:

[0177]

[0178] In this embodiment, the benchmark metrics adopted include False Positive Rate (FPR), False Negative Rate (FNR), Half Total Error Rate (HTER), and the Area Under the Curve (AUC) of the Receiver Operating Characteristic Curve (ROC).

[0179] The False Positive Rate (FPR) refers to the ratio of the number of fraud attack faces misjudged as real faces to the number of faces labeled as fraud attack faces. The formula is as follows:

[0180]

[0181] The False Negative Rate (FNR) refers to the ratio of the number of real faces misjudged as fraud attack faces to the number of faces labeled as real faces. The formula is as follows:

[0182]

[0183] The Half Total Error Rate (HTER) refers to the average of the False Positive Rate and the False Negative Rate. The smaller the HTER value, the better the detection effect of the model. The formula is as follows:

[0184]

[0185] The ROC curve is obtained by calculating the False Positive Rate (FPR) and the True Positive Rate (TPR) at different thresholds, and plotting a curve with FPR as the abscissa and TPR as the ordinate. The AUC is the area under the ROC curve. The larger the AUC value, the better the detection effect of the model.

[0186] To illustrate the insufficient generalization of existing face fraud detection models in the makeup scenario, the Oulu-NPU (hereinafter referred to as O), CASIA (hereinafter referred to as C), Replay-Attack (hereinafter referred to as I), and MSU-MFSD (hereinafter referred to as M) datasets, as well as the makeup face datasets generated from the O, C, I, and M datasets by the makeup transfer method, are used to conduct training and cross-database testing in two modes: without makeup samples and with makeup samples. The specific method is to train on three databases and test on another database. For example, OCI_M means that the model trained on the training set of non-makeup samples in the O, C, and I databases is tested on the test set of the M dataset without makeup samples, and the same applies to OCM_I, ICM_O, and OMI_C; OCI_M(mk) means that the model trained on the training set of makeup samples in the O, C, and I databases is tested on the test set of the M dataset with makeup samples, and the same applies to OCM_I(mk), ICM_O(mk), and OMI_C(mk). The cross-database test results of various typical face fraud detection algorithms in the two modes of without makeup samples and with makeup samples are shown in Tables 1 and 2:

[0187] Table 1 Cross-database test results without makeup samples

[0188]

[0189] Table 2 Cross-database test results with makeup samples

[0190]

[0191]

[0192] Comparing the results in Tables 1 and 2, it can be seen that the generalization performance of various typical face fraud detection algorithms decreases to varying degrees when dealing with makeup faces. Taking SAFAS as an example, the performance of this algorithm is stable under the test protocol without makeup faces, but it significantly declines when tested on makeup faces. This algorithm separates the features of different domains through supervised contrastive learning and optimizes the true and false classification hyperplanes of each domain to align and converge them into a global true and false classification hyperplane as the final trained classifier. However, due to the excessive domain shift in the makeup scenario, the domain shift simulated during model training is difficult to cover, resulting in poor classification performance of the model when dealing with makeup faces.

[0193] To prove the effectiveness and universality of the makeup degree judgment module (hereinafter referred to as MkDeg) proposed in the present invention, the present invention is combined with various typical face fraud detection algorithms. The former serves as the makeup degree judgment branch to assist the binary classification prediction branch based on the latter. Using the same cross-database protocol as in Table 1 and Table 2, several typical algorithms before and after adding MkDeg are trained with makeup samples and cross-database tested. The comparison results are shown in Table 3:

[0194] Table 3 Comparison results of cross-database tests for makeup samples

[0195]

[0196]

[0197] It can be seen from the experimental results in Table 3 that the MkDeg module proposed by the present method can significantly improve the generalization performance of each model when facing makeup faces in the unknown domain. Taking SAFAS as an example, the AUC of this algorithm increases by 4.5% and the HTER decreases by 36.8% after combining with MkDeg, indicating that MkDeg can effectively guide the model to use the makeup degree threshold to assist the binary classification branch. The cross-database test comparison results of several other typical algorithms combined with MkDeg all strongly prove that the present invention can improve the generalization of the face fraud detection model in the face of makeup faces to a certain extent.

[0198] In this embodiment, through the method of spectrum decomposition optimization, the image features of the makeup area are decomposed into amplitude components and phase components in the frequency domain. Learnable tokens are introduced into the above two spectrum components respectively to enhance the model's capture of key texture information in the makeup area; an adaptive high-pass filter is used to extract the high-frequency texture information in the optimized image features of the makeup area; a multi-level attention interaction module is constructed to extract and integrate local channel attention and global spatial attention, and the high-frequency texture features of each makeup area are dynamically weighted by the aggregated channel-space multi-level attention to obtain the overall high-frequency texture feature value of the makeup sample; the makeup degree threshold is obtained through contrastive learning training and is used to correct the prediction results of the binary classification branch during testing. The experimental results prove that the present invention can improve the generalization performance of face fraud detection in the application scenario of makeup faces, and has certain effectiveness and universality.

[0199] Example 2

[0200] Such as Figure 2As shown in the figure, this embodiment provides a makeup face fraud detection system based on spectrum optimization and multi-level attention interaction, which is used to implement the makeup face fraud detection method based on spectrum optimization and multi-level attention interaction in Embodiment 1 above. The system includes: a dataset preprocessing module, a makeup augmentation module, a binary classification loss output module, a makeup area extraction module, a spectrum decomposition optimization module, a high-frequency texture feature extraction module, a multi-level attention interaction module, a makeup degree threshold update module, a total loss function construction module, a model training module, and a prediction result output module;

[0201] In this embodiment, the dataset preprocessing module is used to divide the dataset, frame the videos in each dataset and extract face images, and set corresponding true / false labels and domain labels; the makeup augmentation module is used to generate makeup augmentation samples with different makeup degrees through the makeup transfer algorithm and set makeup degree labels; the binary classification loss output module is used to input the training domain data and makeup augmentation samples into the binary classification prediction branch and output the binary classification loss; the makeup area extraction module is used to input the makeup augmentation samples into the makeup degree judgment branch, detect the face landmark points of the makeup augmentation samples, and extract the images of each makeup area; the spectrum decomposition optimization module decouples the amplitude component and phase component of the frequency domain features of the makeup area images and performs optimization processing on the spectrum components; the high-frequency texture feature extraction module is used to extract the high-frequency texture features of each makeup area based on a high-pass filter; the multi-level attention interaction module is used to dynamically weight the high-frequency texture features of each makeup area to obtain the overall high-frequency texture feature value of the makeup augmentation sample; the makeup degree threshold update module is used to construct positive and negative sample pairs according to the makeup degree labels, obtain the makeup degree threshold through contrastive learning, and output the contrastive loss; the total loss function construction module is used to obtain the total loss function by weighted summation of the binary classification loss and the contrastive loss; the model training module is used to train based on the total loss function to obtain the makeup degree judgment threshold; the prediction result output module is used to correct the result of the binary classification prediction branch based on the makeup degree judgment threshold and output the binary classification prediction probability.

[0202] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for detecting makeup face fraud based on spectrum optimization and multi-level attention interaction, characterized in that It includes the following steps: Divide the dataset, frame the videos in each dataset and extract face images, and set corresponding authenticity labels and domain labels; Generate makeup augmentation samples with different makeup degrees through the makeup transfer algorithm, and set makeup degree labels; Input the training domain data and makeup augmentation samples into the binary classification prediction branch, and output the binary classification loss; Input the makeup augmentation samples into the makeup degree judgment branch, detect the face landmark points of the makeup augmentation samples, and extract the images of each makeup area; Decouple the amplitude component and phase component of the frequency domain features of the makeup area images, and optimize the frequency spectrum components; Extract the high-frequency texture features of each makeup area based on an adaptive high-pass filter; Dynamically weight the high-frequency texture features of each makeup area to obtain the overall high-frequency texture feature value of the makeup augmentation sample; Construct positive and negative sample pairs according to the makeup degree labels, obtain the makeup degree threshold through contrastive learning, and output the contrastive loss; Sum the binary classification loss and the contrastive loss with weights to obtain the total loss function, and train based on the total loss function to obtain the makeup degree judgment threshold; Correct the result of the binary classification prediction branch based on the makeup degree judgment threshold, and output the binary classification prediction probability.

2. The method for detecting makeup face fraud based on spectrum optimization and multi-level attention interaction according to claim 1, wherein Decouple the amplitude component and phase component of the frequency domain features of the makeup area images, and optimize the frequency spectrum components, specifically including: Decompose the frequency domain features of the makeup area images into amplitude components and phase components, introduce learnable tokens respectively, calculate the similarity matrix between the two frequency spectrum components and the corresponding tokens, use it as the attention weight to weight the tokens and then add them to the frequency spectrum components, and output the optimized frequency spectrum components through a multi-layer perceptron to optimize the frequency spectrum components.

3. The method for detecting makeup face fraud based on spectrum optimization and multi-level attention interaction according to claim 2, wherein Decompose the frequency domain features of the makeup area images into amplitude components and phase components, expressed as: where M and N are the width and height of the input area image, u, v are the frequency indices in the frequency domain, f(x, y) is the pixel value of the input image in the spatial domain, F(u, v) is the corresponding complex number in the frequency domain, representing the frequency domain features, Re(F(u, v)) is the real part of F(u, v), Im(F(u, v)) is the imaginary part of F(u, v), atan2 is the two-variable arctangent function, |F(u, v)| and θ(u, v) respectively represent the amplitude component and phase component of the input image in the frequency domain; Introduce a learnable token, denoted as T, to the amplitude component and the phase component respectively α and T ρ , calculate the similarity matrix between each spectral component and its corresponding token through dot product, and perform a normalization operation on the similarity matrix of the amplitude component, expressed as: where d represents the number of channels of image features, Norm represents the normalization operation, A α and A ρ respectively represent the similarity matrices between the amplitude component and the phase component and the corresponding tokens; Use Softmax to convert two similarity matrices A α and A ρ into probability distributions between 0 and 1, which are used as attention weights and applied to T α and T ρ to obtain the weighted tokens, denoted as and represented as: After adding the amplitude component α and the phase component ρ to respectively, the optimized spectral component is obtained through the output of a multi-layer perceptron, denoted as and which is expressed as: For two optimized spectral components and are combined to obtain the optimized frequency-domain features, denoted as The formula is expressed as: where j represents the imaginary unit.

4. The method for detecting makeup face fraud based on spectrum optimization and multi-level attention interaction according to claim 1, characterized in that, Extract the high-frequency texture features of each makeup area based on the high-pass filter, specifically including: Multiply each frequency point by an alternating positive and negative matrix (-1) in the frequency domain u+v Obtain the frequency-domain features after spectral centering, denoted as Expressed as: Frequency domain features after spectral centering by a high-pass filter Perform high-pass filtering, calculate the amplitude of the filtered frequency domain features and normalize them to obtain the high-frequency texture feature values of the regional image, which are specifically expressed as: Among them, represents the high-frequency features obtained after filtering, is the real part of, is the imaginary part of, represents the amplitude component of, α HF is the normalized high-frequency amplitude, that is, the high-frequency texture feature value of the regional image.

5. The method for detecting makeup face fraud based on spectrum optimization and multi-level attention interaction according to claim 1, characterized in that, The cut-off frequency radius of the high-pass filter is: where k is the cut-off frequency radius coefficient, and M and N are the width and height of the input area image; Calculate the distance from each position in the high-pass filter to its center, compare it with the cut-off frequency radius, set the positions greater than the cut-off frequency radius to 1 to retain the high-frequency components, otherwise set them to 0 to filter out the low-frequency components, expressed as: where D(u, v) is the distance from the point (u, v) in the frequency domain to the center point of the spectrum and H(u, v) is the binary form frequency response function of the high-pass filter.

6. The method for detecting makeup face fraud based on spectrum optimization and multi-level attention interaction according to claim 1, wherein Dynamically weight the high-frequency texture features of each makeup area to obtain the overall high-frequency texture feature value of the makeup augmentation sample, specifically including: Patch Embedding is performed on each input makeup area image to convert the image into an information unit that can be understood and processed by the model, denoted as x, and is fed into the local branch and the global branch respectively; The local branch is used to capture channel attention and enhance cross-channel interaction, expressed as: Among them, W 1×1 represents a 1×1 convolution for adjusting the channel dimension, CS represents channel randomization, represents the output after channel randomization; After preliminary weighting through the channel attention mechanism, the feature intensity of each channel is weighted and adjusted to obtain the local channel attention, expressed as: Among them, GlobalAvgPool is the global average pooling layer, FC1 and FC2 are two fully connected layers, and F local is the output local channel attention; The global branch is used to capture spatial attention and perform information aggregation, expressed as: σ(x) = LN(AvgPool(GELU(Linear(x)))) Where, Linear represents the linear layer, GELU represents the GELU activation function adopted, AvgPool represents the average pooling layer, LN represents layer normalization, and σ(x) represents the features output by the global branch; Linear transformation is performed on σ(x) to obtain Q, K, and V, and a learnable Query Embedding, denoted as QE, is introduced into Q. The global attention mechanism is expressed as: Q = σ(x)W Q + QE, K = σ(x)W K , V = σ(x)W V F global = A × V Among them, W Q , W K and W V are all learnable weight matrices, d k is the feature dimension after weight mapping, A is the attention weight after Softmax normalization, and F global represents the global spatial attention; The global spatial attention is calculated through the Sigmoid activation function and used as a gate to act on the local channel attention, optimizing the result of the channel attention and outputting the multi-level attention F of channels and space all , which is expressed as: F all = Sigmoid(F global ) × F local Channel and Spatial Multi-level Attention F all Dynamically weight the images of each makeup area and output the high-frequency texture feature values of the overall makeup augmented sample.

7. The makeup face fraud detection method based on spectrum optimization and multi-level attention interaction according to claim 6, characterized in that In the local branch, the channel random operations specifically include: The convolution features are divided into groups along the channel dimension. Each group uses depthwise separable convolution DWConv for per-channel convolution and pointwise convolution. Per-channel convolution provides an independent processing method for each channel, and pointwise convolution combines independent channel features by random selection; In the global branch, a Linear linear layer and a GELU activation function are used for mapping and activation before feature pooling, and layer normalization operation is used to normalize the output after feature pooling.

8. The method for detecting makeup face fraud based on spectrum optimization and multi-level attention interaction according to claim 1, characterized in that Positive and negative sample pairs are constructed according to the makeup degree label, and the makeup degree threshold is obtained through contrastive learning, and the contrastive loss is output, specifically including: Among them, B is the number of samples in a batch during model training, u i and v i respectively represent the high-frequency texture features of two samples for which the distance is to be calculated in the batch. D(u, v) represents the distance value obtained through standardized Euclidean distance calculation, abbreviated as D. V[u i , v i represents the variance between u i and v i . y i is the label of the sample pair, m is the minimum distance of the negative sample pair, and L con represents the contrastive loss.

9. The method for detecting makeup face fraud based on spectrum optimization and multi-level attention interaction according to claim 1, characterized in that Training is performed based on the total loss function to obtain the makeup degree judgment threshold, specifically including: The average distance between positive sample pairs and negative sample pairs is statistically calculated, and its mean value is used as the makeup degree threshold T, expressed as: Among them, D p represents the average distance of positive sample pairs, and D n represents the average distance of negative sample pairs; Based on the makeup degree judgment threshold, the result of the binary classification prediction branch is corrected, and the binary classification prediction probability is output, specifically including: When the difference in high-frequency texture features between the sample and the real sample is less than T, it is determined that the makeup degree of the sample belongs to the normal range; When the difference in high-frequency texture features between the sample and the real sample is greater than T, it is determined that the makeup degree of the sample is relatively thick, and the prediction result is corrected, expressed as: P adj = (1 - β) × P bim + β × I(Δα test ) where P bin and P adj represent the predicted values of the binary classification branches before and after correction respectively, β is the adjustment coefficient, and I(Δα test ) is the correction direction, expressed as: Among them, Δα test represents the high-frequency texture feature difference between the sample and the true sample.

10. A makeup face fraud detection system based on spectrum optimization and multi-level attention interaction, characterized in that, Including: Dataset preprocessing module, makeup augmentation module, binary classification loss output module, makeup area extraction module, spectrum decomposition optimization module, high-frequency texture feature extraction module, multi-level attention interaction module, makeup degree threshold update module, total loss function construction module, model training module, prediction result output module; The dataset preprocessing module is used to divide the dataset, frame the videos in each dataset and extract the face images, and set the corresponding authenticity labels and domain labels; The makeup augmentation module is used to generate makeup augmentation samples with different makeup degrees through the makeup transfer algorithm and set the makeup degree labels; The binary classification loss output module is used to input the training domain data and the makeup augmented samples into the binary classification prediction branch and output the binary classification loss; The makeup area extraction module is used to input the makeup augmented samples into the makeup degree judgment branch, detect the facial landmark points of the makeup augmented samples, and extract the images of each makeup area; The spectral decomposition optimization module is used to decouple the amplitude component and the phase component of the frequency domain features of the makeup area image and perform optimization processing on the spectral components; The high-frequency texture feature extraction module is used to extract the high-frequency texture features of each makeup area based on a high-pass filter; The multi-level attention interaction module is used to dynamically weight the high-frequency texture features of each makeup area to obtain the high-frequency texture feature value of the overall makeup augmented sample; The makeup degree threshold update module is used to construct positive and negative sample pairs according to the makeup degree labels, obtain the makeup degree threshold through contrastive learning, and output the contrastive loss; The total loss function construction module is used to obtain the total loss function by weighted summation of the binary classification loss and the contrastive loss; The model training module is used to train based on the total loss function to obtain the makeup degree judgment threshold; The prediction result output module is used to correct the result of the binary classification prediction branch based on the makeup degree judgment threshold and output the binary classification prediction probability.

Citation Information

Cited By

  • Face image anti-counterfeiting detection method and device based on frequency spectrum reconstruction

    CN121354195A