Deep counterfeit multi-label sorting and positioning method based on multi-domain feature fusion
By using multi-domain feature fusion and graph attention network methods in deep forgery detection, the problem of difficulty in dealing with sequence depth forgery operation in the prior art is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510148811.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-23
AI Technical Summary
Existing deep forgery detection methods are difficult to effectively handle multi-step editing in sequence depth forgery operations, resulting in insufficient detection accuracy and robustness, especially when facing data constructed by unknown forgery methods.
The deep fake multi-label sorting and positioning method based on multi-domain feature fusion is adopted. The spatial domain and frequency domain features are extracted respectively through DINOv2 and DCT transformations, and the cross attention mechanism is used to fusion, and the probability prediction of multi-label content forgery and the sorting of forgery parts is carried out in combination with the graph attention network.
It improves the accuracy and robustness of deep forgery detection, can more effectively identify and locate modified facial components and their tampering order in sequence depth forgery operations, and enhances the ability to respond to complex deep forgery technologies.
Smart Images

Figure CN120032234A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep fakes, and in particular to a deep fake multi-label sorting and positioning method based on multi-domain feature fusion. Background Art
[0002] With the rapid development of deep learning technology, face forgery technology has become increasingly accessible, requiring only a small amount of professional skills and equipment to generate highly realistic forged faces. These technologies, such as variational autoencoders and generative adversarial networks, have spawned numerous free applications and open source projects, making it extremely easy to create forged images or videos. However, this convenience also brings risks, as face forgery technology can be maliciously abused in many fields such as politics and pornography, causing serious trust crises and social problems. In order to ensure the safety and responsibility of multimedia environments, it is now urgent to study more robust face forgery detection and forensics methods to identify and prevent these potential threats.
[0003] At present, most of the work on visual deepfake detection focuses on image and video content, which can be roughly divided into three categories: detection methods based on specific artifacts, detection methods based on deep learning technology, and other types of visual deepfake detection methods. In the early days, detection methods based on specific artifacts mainly relied on using internal statistical features of images or manually designed features to reveal the differences between real and fake videos, or by looking for mixed boundary information for detection. However, with the continuous advancement of facial synthesis technology, fake faces have become more and more realistic, which makes it extremely challenging to mine subtle fake clues, making it difficult for these methods to reliably detect face fakes. In addition, image processing technology may weaken artifacts in fake images. These artifacts often rely on subtle image patterns and are easily affected by post-processing steps such as JPEG compression, limiting the generalization ability of fake detection algorithms. With the rapid development of deep learning, many works have gradually revolved around the improvement of CNN or RNN networks. Excellent network structure optimization methods can more effectively capture and extract subtle features in images or videos, thereby improving the accuracy and robustness of tasks.
[0004] Deepfake technology can not only replace the target person's face with the face of a specific individual, but also modify the target person's expression, action, accessories, hair color, etc. while maintaining the target person's identity. However, the current mainstream deepfake detection methods mainly detect forged images of single facial operations. With the popularity of facial editing applications, users can easily perform multi-step editing, such as adjusting expressions, changing hairstyles, etc. These multi-step operations may affect each other, resulting in an increase in the diversity of forged images, forming an emerging trend - Sequential Deepfake Operation (Seq-Deepfake). This complex operation mode poses challenges to traditional deepfake detection methods, making it difficult for them to accurately predict the overall results. Therefore, existing deepfake detection methods need to adapt to and cope with this continuous multi-step operation mode of face tampering, and study more refined deepfake localization and sorting tasks. By identifying the tampered facial components and the order of their tampering, the challenges posed by sequential deepfake operations can be better addressed. Summary of the invention
[0005] The main technical problem that the present invention solves is as follows: For deep fake detection tasks, a deep fake multi-label sorting and positioning method based on multi-domain feature fusion is proposed to solve the problem that existing deep fake detection tasks lack more refined deep fake positioning and sorting tasks to deal with serial deep fake operations. At the same time, by combining the fusion of spatial domain and frequency domain features, a more discriminative and universal feature representation is discovered, thereby improving the accuracy and robustness of deep fake detection and effectively responding to the increasingly complex technical challenges of deep fakes.
[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0007] A deep fake multi-label sorting and positioning method based on multi-domain feature fusion, comprising:
[0008] Step S1: preprocessing the sequence deep fake dataset;
[0009] Step S2: Input the face RGB image into the spatial feature extractor based on DINOv2 to extract spatial domain features. DINOv2 consists of DINO and iBOT. DINO adopts the student-teacher model training method and uses the cross entropy between the probability distributions generated by category tokens based on image cropping to learn to extract global image features. iBOT randomly blocks part of the image area and improves the local information extraction ability of the network by minimizing the difference in fusion features in the occluded area of the image.
[0010] Step S3: Decompose the face RGB image using discrete cosine transform to obtain several frequency domain coefficients and spectra, and flatten and reshape the obtained spectrum so that components with the same frequency are grouped into one channel to form a new input, and then finally convert the original color input of the face RGB image into a feature representation in the frequency domain through an encoder;
[0011] Step S4: Using the spatial domain features obtained in step S2 and the frequency domain features obtained in step S3, they are respectively connected and fused through the cross-attention module to form a new multi-domain representation of deep fake face images with spatial domain and frequency domain, and then input into the graph attention network to capture the complex relationship between different features, including: the graph attention network includes a graph convolutional network, an attention mechanism, attention pooling and a feedforward neural network, wherein the graph convolutional network captures the correlation between face labels through a reweighted correlation matrix, and the attention mechanism further enhances the relationship between image regions and labels and the correlation between labels. Subsequently, the attention pooling technology is used to aggregate the feature map, and the features of each region are represented as a single feature vector. Finally, the feature vector is input into the feedforward neural network and the sigmoid function is used to realize the multi-label content forgery probability prediction and forged part sorting, ultimately providing a more fine-grained forged content detection and analysis result.
[0012] According to the technical solution of the present invention, a deep fake spatial domain feature extractor is designed to extract global and local information from the image, covering the overall appearance features, movements, regional textures and other features of the character; a frequency domain feature extractor is designed to obtain frequency components by performing discrete cosine transform on the image, thereby revealing possible coding traces or unusual frequency domain distributions in the image; at the same time, the features of the spatial domain and frequency domain are spliced and fused through a cross-attention mechanism to obtain more universal discriminative features to cope with the influence of the disappearance of discriminative artifacts caused by compression operations; then the correlation and importance between the fused feature sequences are learned through a graph attention network, and the multi-label content forgery probability prediction and forged part sorting are realized through a feedforward neural network and a sigmoid function, ultimately providing a more fine-grained forged content detection and analysis result.
[0013] The beneficial effects of the present invention are:
[0014] The present invention provides a deep fake multi-label sorting and positioning method based on multi-domain feature fusion. In the field of deep fake detection, there are problems such as artifact blurring caused by compression operation of fake videos, and insufficient generalization ability of the model when facing data constructed by unknown fake methods. The multi-domain feature fusion method is used to capture image domain and frequency domain information, and the complementarity between the two domains is used to improve the anti-interference ability of the model. In addition, traditional binary classification methods are difficult to effectively process data that has undergone multiple steps of fake operations, and have limitations in sequence detection tasks. By introducing a graph attention network to learn face relationship features for multi-label sequence data, the similarity of regions in the image is analyzed to identify the specific facial components that have been modified and determine the fake order, so as to achieve more accurate deep fake analysis results. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic diagram of the deep fake multi-label sorting and positioning method based on multi-domain feature fusion in the present invention;
[0016] Figure 2 It is a schematic diagram of multi-domain feature fusion in the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0018] Figure 1 Schematic diagram of the deep fake multi-label sorting and positioning method based on multi-domain feature fusion in the present invention. Figure 2 This is a schematic diagram of multi-domain feature fusion in the present invention. Figure 1 and Figure 2 This paper describes a deep fake multi-label ranking and localization method based on multi-domain feature fusion. Figure 1 As shown, the method comprises the following steps:
[0019] Step S1: Supplement the sequence deep fake dataset by preprocessing the face deep fake dataset. The preprocessing may include: obtaining the face deep fake dataset, extracting video frames from the video to be detected, and performing operations such as face cropping, normalization, image horizontal flipping, and denoising, and then supplementing it to the sequence deep fake dataset to obtain richer deep fake data. The sequence deep fake dataset is a public dataset established in advance.
[0020] In one embodiment, step S1 may specifically include:
[0021] Step S11: For the video to be detected, firstly extract the video frame through OpenCV, and then perform preprocessing operations such as face cropping, normalization, image horizontal flipping, and denoising on it;
[0022] Step S12: Add the processed image data to the sequential deepfake dataset to effectively enrich the diversity of deepfake data and improve the robustness and generalization ability of the deepfake detection model. ,in is an image, Is not longer than An ordered subset of Contains possibly modified instances.
[0023] Step S2: Input the face RGB image into the spatial feature extractor based on DINOv2 to extract spatial domain features. DINOv2 consists of DINO and iBOT. DINO adopts the student-teacher model training method and uses the cross entropy between the probability distributions generated by the category tokens based on image cropping to learn to extract the global features of the image. iBOT randomly blocks part of the image area and improves the local information extraction ability of the network by minimizing the difference in the fused features of the occluded area of the image.
[0024] In one embodiment, step S2 specifically includes:
[0025] Step S21: adopt a discriminative self-supervised DINOv2 method to learn the spatial domain features of the face RGB image. The discriminative self-supervised DINOv2 method is based on DINO and iBOT losses and SwAV center. While learning the spatial domain features of the image, it considers the consistency and discrimination of the features, so as to obtain a more discriminative feature representation;
[0026] DINO is a self-supervised learning framework based on distillation. It performs two different random transformations on the input image and passes these transformed image features to the student and teacher networks. Each network outputs a K-dimensional feature. These features are normalized by temperature softmax in the feature dimension, and then the cross entropy loss is used to measure the similarity between features. The process specifically includes: training the student network Network with Faculty The output matches are respectively and Parameterization. Given an input image , passing the transformed image features to the student and teacher networks, each network outputs a dimensional features, the two networks The probability distribution on the dimension, these features are normalized by temperature softmax on the feature dimension, as follows:
[0027] (1)
[0028] in, is the temperature parameter, j is the index variable, is the dimension of the feature vector. Then the similarity between features is measured by minimizing the cross entropy loss as follows:
[0029] (2)
[0030] in, and Indicates that the two networks are dimensional probability distribution, H represents the cross entropy loss, are the parameters of the student network.
[0031] Step S22: iBOT is based on pixel-level targets and also uses the student-teacher model training method. It uses a visual encoder to divide the image into a series of small, fixed-size rectangular blocks, and randomly blocks a part of the image. By minimizing the reconstruction error of the blocked area of the image, it learns an effective image feature representation. The process may include: given a training set , uniformly sample the image and apply two random augmentations to produce two warped views and In the student-teacher network, knowledge is transferred from the teacher to the student by minimizing the cross entropy, which can be expressed as:
[0032] (3)
[0033] in, and is the distorted view generated by the image after different random enhancement operations, Is the teacher network view The predicted category distribution of Is the student network on view The predicted category distribution of . The teacher and the student share the backbone network and projection head architecture . Teacher Network Exponential Moving Average Study Student Network Parameters.
[0034] In the local area of the image, two enhanced views are used and Perform mask operation to get and , after passing through the teacher-student network, the student network outputs the local area occlusion view Projection , while the teacher network outputs the unobstructed local area Projection ,in It is the projection head of the student network. It is the projection head of the teacher network.
[0035] In iBOT, mask image modeling The training objective here can be defined as:
[0036] (4)
[0037] in, is the total number of image blocks, is an indicator function that assigns a value to each image patch. If the image patch is occluded, then is 1, otherwise it is 0. Is the teacher network for the first Unoccluded image patches The predicted probability distribution of Is the student network for the first Unoccluded image patches The predicted probability distribution of Represents the student network.
[0038] Step S3: Decompose the face RGB image using discrete cosine transform (DCT) to obtain several frequency domain coefficients and spectra, and flatten and reshape the obtained spectrum so that components of the same frequency are grouped into one channel to form a new input. Then, the encoder is used to finally convert the original color input of the face RGB image into a feature representation in the frequency domain.
[0039] In one embodiment, the above step S3 specifically includes:
[0040] Step S31: Transform the face RGB image Use discrete cosine transform (DCT) to decompose and obtain several frequency domain coefficients ,in , and is the height, width and number of channels of the image, Divide into a group Each patch is processed into a spectrum by DCT , each value represents the energy intensity of a specific frequency band;
[0041] Step S32: Spectrum Flatten and reshape to group components of the same frequency into one channel to form the new input:
[0042] (5)
[0043] Step S33: Encode the new input of step S32 through the encoder, thereby converting the original color input of the face RGB image into a feature representation in the frequency domain.
[0044] Step S4: Using the image spatial domain features obtained in step S2 and the frequency domain features obtained in step S3, they are respectively spliced and fused through the cross-attention module to form a new multi-domain representation of deep fake face images with spatial domain and frequency domain, and then input into the graph attention network to capture the complex relationship between different features. The graph attention network includes a graph convolutional network, an attention mechanism, attention pooling, and a feedforward neural network. The graph convolutional network captures the correlation between face labels through a reweighted correlation matrix, and the attention mechanism further enhances the relationship between image regions and labels and the correlation between labels. Subsequently, the attention pooling technique is used to aggregate the feature map and represent the features of each region as a single vector. This pooling method can retain the key information in the region and eliminate redundant information, thereby improving the robustness of the features. Finally, the feature vector is input into the feedforward neural network and the multi-label content forgery probability prediction and forgery part sorting are realized through the sigmoid function, ultimately providing a more fine-grained forgery content detection analysis result.
[0045] In one embodiment, the above step S4 specifically includes:
[0046] Step S41: The spatial domain features and frequency domain features are respectively used to learn the correlation and importance between different domain features through the cross-attention mechanism, and then the splicing operation is performed to fuse the features of the two domains to form a more comprehensive and comprehensive feature expression, providing effective support for subsequent multi-label sorting and positioning tasks. This process may specifically include:
[0047] The image spatial domain features obtained in step S2 and the frequency domain features obtained in step S3 are represented as the original visual feature map The cross-attention module is a Transformer encoder that uses a fixed position encoding to supplement the original visual feature map. .
[0048] In the Transformer encoder, the input sequence is treated as a sequence and each element in the sequence interacts with every other element. The self-attention mechanism implemented by generating keys and queries allows the model to focus on information at other locations when processing sequence data, which helps capture long-range dependencies and structural information in the sequence. , , A self-attention operation is performed to extract the relationship between positions and capture the forgery traces in the spatial and frequency domains. To facilitate the extraction of relational features in the spatial and frequency domains, the cross-attention module uses a multi-head self-attention mechanism to group position features along the channel dimension. The multi-head normalized attention based on dot product is as follows:
[0049] (6)
[0050] (7)
[0051] (8)
[0052] (9)
[0053] in, , , Representative the key, query, and value characteristics of the group, , , Representative The key, query, and value characteristics of a group, where is the dimension of query and key, generating d groups in total, and D is the number of attention heads. is the spatial relationship feature of the i-th group, is the frequency domain feature of the jth group. And all groups are concatenated to form The spatial relationship characteristics and The frequency domain characteristics.
[0054] Finally, the fused features are obtained by splicing , that is, a new multi-domain representation of deep fake face images with spatial and frequency domains:
[0055] (10)
[0056] Step S42: The fused features are input into the graph attention network. The nodes in the graph represent the facial region information that may be tampered with, and the edges represent the relationship information between the nodes. A shared linear transformation parameterized by a weight matrix is applied to each node, and then a self-attention operation is performed on the nodes to capture the correlation and importance between the nodes, thereby obtaining an updated node feature representation. The graph attention network can effectively capture the dependencies between different regions in the image and highlight the key information of each operation in the sequence. The process may include:
[0057] By calculating the fused features The similarity of N nodes in the graph is used to build a graph attention network. Indicates the facial area information that may be tampered with. Represents the relationship information between nodes, using a shared linear transformation parameterized by a weight matrix is applied to each node, and then performs self-attention on the node:
[0058] (11)
[0059] in, Representation Node and nodes The attention weights between is a learnable weight matrix that is used to linearly transform the features of each node. and Indicates the facial area information that may be tampered with. Represents learnable parameters. The node feature representation contains the facial region information and encodes the relationship features between nodes. The graph attention network can effectively capture the dependencies between different regions in the image and highlight the key information of each operation in the sequence;
[0060] Step S43: The node features are then input into the feedforward neural network (FFN) after attention pooling. The FFN learns to map the node features to the probability distribution of each operation category to represent the possibility of each operation category. By comparing the predicted probabilities of each operation category, the true order of the operations is obtained and combined with the operation category to locate and sort the forged operations.
[0061] In order to ensure that the predicted probability distribution can reflect the true order of operations, a multi-label sorting loss function is introduced. The loss function is calculated based on the difference between the true operation order and the predicted probability distribution, guiding FFN to learn the correct operation order.
[0062] The process may include:
[0063] The node features are input into the feedforward neural network after attention pooling ( )middle, By learning to map node features to the probability distribution of various operation categories to represent the probability of each operation category occurring, the deep fake localization task is converted into a multi-label ranking problem. is an ordered subset of length not exceeding N, Contains possibly modified Regions, each region Indicates that in the image A tampered facial component, Indicates the number of components that have been modified. When it is empty, it means that the image does not have any tampered components and is a real image;
[0064] Each node feature is reshaped to include an image block feature tag and a learnable category feature tag, and the similarity value of the two tags is calculated. :
[0065] (12)
[0066] in, All elements of are in the range [0,1]. The value after performing self-attention on the image block feature label is obtained. These image blocks that may have forgeries are arranged in ascending order according to the similarity value to obtain a forgery sorted list.
[0067] In the forgery sorted list, if the image is forged, then there is at least one value in the list whose value is greater than is close to 1. On the contrary, if there is no forged part, that is, the image is a real image, then all values will be less than , by comparing the average response of the positive and negative distributions to calculate the depth of the fake image before the sorted list The probability of the occurrence of the minimum similar response:
[0068] (13)
[0069] in, is in the fake sorted list, The probability that the minimum similarity response is less than a certain threshold, Indicates a forged image and is used to determine whether the similarity is low enough to indicate that the image has not been forged at that location.
[0070] By comparing different images The similarity between them is used to distinguish positive examples from negative examples. The loss function is defined as:
[0071] (14)
[0072] Among them, if is a fake image, =1, if If it is not forged, =0.
[0073] The loss function for deepfake classification is defined as:
[0074] (15)
[0075] in, is the regularization loss, Represents a forged dataset. In order to ensure that the predicted probability distribution can reflect the true order of operations, a multi-label sorting loss term is introduced. It is calculated based on the difference between the true operation order and the predicted probability distribution, guiding the network model to learn to output the predicted distribution of multiple labels. The multi-label sorting loss is defined as:
[0076] (16)
[0077] Among them, w is the ranking-aware weight vector, which is used to adjust the weight of each label according to the actual modification order. is a multi-label cross entropy loss function, which is used to measure the difference between the predicted label and the true label. is the ranking order, is the multi-label prediction probability.
[0078] Finally, by comparing the predicted probabilities of each operation category, the true order of operations is obtained, and the location information of the forged content is obtained. The overall loss function is defined as follows:
[0079] (17)
[0080] in, and is a parameter to measure the influence of the specific loss term, is the binary cross entropy loss function.
[0081] In summary, the present invention designs a deep fake multi-label sorting and positioning method based on multi-domain feature fusion. In the field of deep fake detection, there is a problem that the forged video is compressed, resulting in blurred identification artifacts, resulting in insufficient model generalization ability when facing data constructed by unknown forgery methods. In addition, most current methods cannot effectively process data that has undergone multiple steps of forgery operations, and have limitations in sequence detection tasks. A multi-domain feature fusion method based on DINOv2 and DCT transform is designed to capture image domain and frequency domain information respectively, and effectively capture the dependency between different features through a cross-attention mechanism, and use the complementarity between the two domains to improve the robustness and anti-interference ability of the model. At the same time, a forgery sorting and positioning method based on multi-label prediction is designed. The face relationship features are further learned through the graph attention network, the similarity of the regions in the image is analyzed, the specific facial components that have been modified are identified, and the forgery order is determined, so as to achieve more accurate deep fake analysis results.
[0082] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A deep fake multi-label sorting and positioning method based on multi-domain feature fusion, characterized in that: include: Step S1: Supplement the sequence deep fake dataset by preprocessing the face deep fake dataset; Step S2: Input the face RGB image into the spatial feature extractor based on DINOv2 to extract spatial domain features. DINOv2 consists of DINO and iBOT. DINO adopts the student-teacher model training method and uses the cross entropy between the probability distributions generated by category tokens based on image cropping to learn to extract global image features. iBOT randomly blocks part of the image area and improves the local information extraction ability of the network by minimizing the difference in fusion features in the occluded area of the image. Step S3: Decompose the face RGB image using discrete cosine transform to obtain several frequency domain coefficients and spectra, and flatten and reshape the obtained spectrum so that components with the same frequency are grouped into one channel to form a new input, and then finally convert the original color input of the face RGB image into a feature representation in the frequency domain through an encoder; Step S4: Using the spatial domain features obtained in step S2 and the frequency domain features obtained in step S3, they are respectively connected and fused through the cross-attention module to form a new multi-domain representation of deep fake face images with spatial domain and frequency domain, and then input into the graph attention network to capture the complex relationship between different features, including: the graph attention network includes a graph convolutional network, an attention mechanism, attention pooling and a feedforward neural network, wherein the graph convolutional network captures the correlation between face labels through a reweighted correlation matrix, and the attention mechanism further enhances the relationship between image regions and labels and the correlation between labels. Subsequently, the attention pooling technology is used to aggregate the feature map, and the features of each region are represented as a single feature vector. Finally, the feature vector is input into the feedforward neural network and the sigmoid function is used to realize the multi-label content forgery probability prediction and forged part sorting, ultimately providing a more fine-grained forged content detection and analysis result.
2. According to claim 1, a deep fake multi-label sorting and positioning method based on multi-domain feature fusion is characterized in that: Step S1 includes: obtaining a deep fake face dataset, extracting video frames from the video to be detected, and performing face cropping, normalization, horizontal image flipping, and denoising operations, and then adding it to the sequence deep fake dataset.
3. According to claim 2, a deep fake multi-label sorting and positioning method based on multi-domain feature fusion is characterized by: The step S1 comprises: Step S11: For the video to be detected, firstly extract the video frame through OpenCV, and then perform face cropping, normalization, image horizontal flipping, and denoising preprocessing operations on it; Step S12: Add the processed image data to the sequence deep fake dataset.
4. According to claim 1, a deep fake multi-label sorting and positioning method based on multi-domain feature fusion is characterized by: The step S2 comprises: Step S21: adopting a discriminative self-supervised DINOv2 method to learn the spatial domain features of the face RGB image, wherein the discriminative self-supervised DINOv2 method is based on DINO and iBOT losses and SwAV center; Step S22: iBOT is based on pixel-level targets and also adopts the training method of the student-teacher model. It uses a visual encoder to divide the image into a series of small, fixed-size rectangular area blocks, and randomly blocks part of the image area. By minimizing the reconstruction error of the occluded area of the image, it learns effective image feature representation.
5. According to claim 4, a deep fake multi-label sorting and positioning method based on multi-domain feature fusion is characterized by: DINO is a self-supervised learning framework based on distillation. It performs two different random transformations on the input image and passes these transformed image features to the student and teacher networks. Each network outputs a K-dimensional feature, which is temperature softmax normalized in the feature dimension, and then the cross entropy loss is used to measure the similarity between features.
6. According to claim 1, a deep fake multi-label sorting and positioning method based on multi-domain feature fusion is characterized by: The step S3 comprises: Step S31: decomposing the face RGB image by discrete cosine transform to obtain a number of frequency domain coefficients, dividing the frequency domain coefficients into a group of image blocks, and processing each image block into a spectrum by discrete cosine transform; Step S32: Flatten and reshape the spectrum so that components with the same frequency are grouped into one channel to form a new input; Step S33: Encode the new input of step S32 through the encoder, thereby converting the original color input of the face RGB image into a feature representation in the frequency domain.
7. According to claim 1, a deep fake multi-label sorting and positioning method based on multi-domain feature fusion is characterized by: The step S4 comprises: Step S41: The spatial domain features and the frequency domain features are respectively used to learn the correlation and importance between the features of different domains through the cross attention mechanism, and then the splicing operation is performed to fuse the features of the two domains; Step S42: Input the fused features into the graph attention network, where the nodes in the graph represent the facial region information that may be tampered with, and the edges represent the relationship information between the nodes. A shared linear transformation parameterized by a weight matrix is applied to each node, and then a self-attention operation is performed on the nodes to capture the correlation and importance between the nodes, thereby obtaining an updated node feature representation; Step S43: After attention pooling, the node features are input into the feedforward neural network FFN. FFN learns to map the node features to the probability distribution of each operation category to represent the possibility of each operation category. By comparing the predicted probabilities of each operation category, the true order of the operations is obtained and combined with the operation category to locate and sort the forged operations.
8. According to claim 7, a deep fake multi-label sorting and positioning method based on multi-domain feature fusion is characterized by: Step S41 includes: calculating the fused features of The similarity of nodes in the graph is used to build a graph attention network. Indicates the facial area information that may be tampered with. Represents the relationship information between nodes, using a shared linear transformation parameterized by a weight matrix is applied to each node, and then performs self-attention on the node: (11) in, Representation Node and nodes The attention weights between is a learnable weight matrix that is used to linearly transform the features of each node. and Indicates the facial area information that may be tampered with. represents the learnable parameter vector.
9. The deep fake multi-label sorting and positioning method based on multi-domain feature fusion according to claim 8 is characterized by: Step S43 includes: introducing a multi-label sorting loss function, calculating according to the difference between the actual operation sequence and the predicted probability distribution, and guiding the FFN to learn the correct operation sequence.
10. The deep fake multi-label sorting and positioning method based on multi-domain feature fusion according to claim 9 is characterized by: The multi-label ranking loss is defined as: (16) Among them, w is the sort-aware weight vector, which is used to adjust the weight of each label according to the actual modification order. is a multi-label cross entropy loss function, which is used to measure the difference between the predicted label and the true label. is the sort order, Represents the multi-label prediction probability.
Citation Information
Cited By
Immersive video quality evaluation method and device based on frequency characteristics
CN120220036A
Immersive Video Quality Evaluation Method and Device Based on Frequency Features
CN120220036B
Anaphora image segmentation method based on space-frequency dual tuning
CN120976550A
Reference image segmentation method based on space-frequency duality tuning
CN120976550B
Deep pseudo detection method for local and global self-supervised contrast learning
CN122176773A