Face forgery detection method and system based on inter-block discontinuous feature learning

By introducing low-rank matrices and feature redistribution strategies into the pre-trained visual model, the feature differences between real faces and forged faces are learned, which solves the problem of insufficient generalization ability of existing methods in unknown forged patterns and cross-dataset detection, and achieves efficient and robust face forgery detection.

CN120808454APending Publication Date: 2025-10-17SHANGHAI JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510924922.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing face forgery detection methods have poor generalization ability when facing unknown forgery patterns or cross-datasets, are complex in structure and have high computational cost. They rely on prior knowledge, resulting in a lack of adaptability to new attack methods. In addition, the fine-tuning of large-scale visual models is computationally complex and lacks cross-domain robustness.

Method used

Using a pre-trained visual model with frozen parameters and a low-rank matrix, the continuity features between real face image blocks and the discontinuity features between forged face image blocks are learned through feature space redistribution and feature enhancement function. The model is optimized with a small number of learnable parameters, and backpropagation training is performed in combination with cross entropy loss and weighted triplet loss.

Benefits of technology

It improves the generalization ability of the model, reduces the computational complexity, enhances the robustness against new forgery attacks, is easy to extend to other image forgery detection tasks, and has good cross-domain detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808454A_ABST
    Figure CN120808454A_ABST
Patent Text Reader

Abstract

The invention provides a face forgery detection method and system based on inter-block discontinuous feature learning, and the method comprises the steps: inputting a face image into a preset face forgery detection model, determining differential features, and obtaining a face forgery detection result; the preset face forgery detection model comprises a pre-training visual model with freezing parameters and a low-rank matrix with learnable parameters; performing feature redistribution optimization on the differentiated features by adopting a preset feature space redistribution strategy to determine first optimized features; performing enhancement processing on the first optimization feature by adopting a preset feature enhancement function, and determining a second optimization feature; classifying the second optimization features to determine total loss; and optimizing the low-rank matrix with learnable parameters by adopting total loss back propagation. According to the method and the device, the visual model is finely adjusted by adopting a small number of trainable parameters, the continuity representation between real face image blocks and the discontinuity representation between forged face image blocks are learned, the calculation complexity is reduced, and the model generalization is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of face forgery detection, in particular to a face forgery detection method and system based on inter-tile discontinuous feature learning. BACKGROUND

[0002] With the rapid development of generative artificial intelligence technology, the authenticity and diversity of forged face images have significantly improved. These deep learning-based image synthesis technologies, such as Deepfakes, Faceswap, and other face-swapping algorithms, while having positive application value in entertainment, advertising, and other fields, are also misused for malicious purposes such as creating fake news, political rumors, and other malicious purposes, seriously threatening personal privacy and public opinion safety.

[0003] To address the above problems, current research mainly focuses on detecting forged faces based on deep learning algorithms. Existing research usually relies on prior knowledge of forged clues in images, extracts local or global forgery features by designing complex neural network structures or specific modules. These methods perform well in the same domain scenario where the distribution of training and testing data is the same.

[0004] However, the existing technology still has the following outstanding problems: ① Poor generalization ability: when facing unknown forgery patterns or cross-dataset forgery detection tasks, the performance of existing models decreases significantly, making it difficult to adapt to the constantly emerging deep forgery technology. ② Complex structure, high design cost: Many methods introduce a large number of complex modules such as attention mechanisms, local discriminators, etc. to capture specific forgery traces, resulting in complex model structure, large training overhead, and difficulty in deploying to resource-constrained devices. ③ Over-reliance on prior knowledge: Most detection algorithms rely on the feature distribution of existing forgery types, which is prone to overfitting to training data and lacks adaptability to potential new attack methods.

[0005] In addition, recent face forgery detection methods use large-scale visual models in an unreasonable and poor performance manner. Although large-scale visual pre-training models such as Vision Transformer (ViT) have good image understanding capabilities, existing forgery detection methods either need to fine-tune all their parameters or introduce complex structure adapters to fine-tune these pre-training models to introduce prior knowledge of existing forgeries. The computational cost is high, and there is still a lack of cross-domain robustness.

[0006] In the patent "CN116188956A; a method for detecting a deep fake face image and related equipment", a deep neural network is trained through a real face image and a fake face image set, and an overall similarity loss value and an overall classification loss value are calculated. The sum of the loss values of the two is used as the overall loss, and the deep learning network is trained and the network parameters of the deep learning network are updated according to the overall loss through the back propagation method to obtain a fake face image detection model. The patent adjusts all parameters of the deep learning network through the training set and the loss function, which is complex to calculate, and still lacks cross-domain robustness. SUMMARY

[0007] In view of the defects in the prior art, the purpose of the present application is to provide a face forgery detection method and system based on inter-tile discontinuous feature learning.

[0008] The first aspect of the present application provides a face forgery detection method based on inter-tile discontinuous feature learning, comprising:

[0009] Inputting a face image into a preset face forgery detection model to determine a differential feature, the preset face forgery detection model comprising a pre-trained visual model with frozen parameters and a low-rank matrix with learnable parameters, the differential feature comprising continuity features between real face image blocks and discontinuity features between fake face image blocks;

[0010] Performing feature redistribution optimization on the differential feature using a preset feature space redistribution strategy to determine a first optimized feature;

[0011] Performing enhancement processing on the first optimized feature using a preset feature enhancement function to determine a second optimized feature;

[0012] Inputting the second optimized feature into a preset classification head to determine a total loss;

[0013] Optimizing the low-rank matrix with learnable parameters and the feature redistribution parameters using the total loss back propagation to determine a trained face forgery detection model;

[0014] Inputting an actual face image into the trained face forgery detection model to determine a face forgery detection result.

[0015] Optionally, the method for determining the preset face forgery detection model comprises:

[0016] Freezing the pre-training weights of the key matrix;

[0017] Freezing the pre-training weights of the query matrix and the pre-training weights of the value matrix;

[0018] A first low-rank matrix and a second low-rank matrix with learnable parameters are used to superimpose a first learnable correction term on the query matrix to determine a new query matrix;

[0019] A third low-rank matrix and a fourth low-rank matrix with learnable parameters are used to superimpose a second learnable correction term on the value matrix to determine a new value matrix;

[0020] According to the new query matrix and the new value matrix, the preset face forgery detection model is determined.

[0021] Optionally, the face image is input into the preset face forgery detection model to determine the differential feature, comprising:

[0022] According to the preset size, the face image is divided into a plurality of image blocks of the preset size;

[0023] The plurality of image blocks of the preset size are input into the preset face forgery detection model to determine the differential feature.

[0024] Optionally, the feature redistribution parameter includes a preset scaling factor and a preset noise.

[0025] Optionally, the differential feature is optimized by the preset feature space redistribution strategy to determine a first optimized feature, comprising:

[0026] According to the preset scaling factor and the preset noise, the differential feature is multiplied by the preset scaling factor and summed with the preset noise to determine the first optimized feature.

[0027] Optionally, the first optimized feature is enhanced by the preset feature enhancement function to determine a second optimized feature, comprising:

[0028] The first optimized feature is pooled to determine a pooled feature vector;

[0029] According to the real label and the fake label, the pooled feature vector is divided into a real face feature and a fake face feature;

[0030] According to the real face feature and the fake face feature, a classification direction from the real face feature to the fake face feature is determined;

[0031] The classification direction from the real face feature to the fake face feature is orthogonalized to determine an orthogonal direction of the classification direction;

[0032] According to a preset random disturbance factor and an orthogonal direction of the classification direction, the real face feature and the fake face feature are respectively enhanced to determine an enhanced real face feature and an enhanced fake face feature.

[0033] The enhanced real face feature and the enhanced fake face feature are spliced to determine the second optimization feature.

[0034] Optionally, the classification direction from the real face feature to the fake face feature is determined according to the real face feature and the fake face feature, comprising:

[0035] A vector from a real class centroid to a fake class centroid is determined according to the real face feature and the fake face feature.

[0036] The vector from the real class centroid to the fake class centroid is normalized to determine the classification direction from the real face feature to the fake face feature.

[0037] Optionally, the total loss is determined by inputting the second optimization feature into a preset classification head, comprising:

[0038] The prediction label of the face image is determined by inputting the second optimization feature into the preset classification head.

[0039] The cross-entropy loss is determined according to the ground truth label of the face image and the prediction label of the face image.

[0040] According to the ground truth label of the face image, an anchor sample, a positive sample and a negative sample are determined in the pooled feature vector.

[0041] The Euclidean distance between the anchor sample and the positive sample and the Euclidean distance between the anchor sample and the negative sample are determined.

[0042] According to the Euclidean distance between the anchor sample and the positive sample and the Euclidean distance between the anchor sample and the negative sample, the positive sample weight and the negative sample weight are respectively determined based on the dynamic weight mechanism of Softmax.

[0043] According to the positive sample weight, the Euclidean distance between the anchor sample and the positive sample, the negative sample weight and the Euclidean distance between the anchor sample and the negative sample, the weighted triplet loss is determined.

[0044] According to the second optimization feature and the first optimization feature, the feature enhancement-based discriminative optimization loss is determined.

[0045] The total loss is determined according to a weighted sum of the cross entropy loss, the weighted triplet loss, and the feature enhancement-based discriminant optimization loss.

[0046] In a second aspect, the present application provides a face forgery detection system based on learning discontinuous features between image blocks, comprising:

[0047] a feature extraction module configured to input a facial image into a preset face forgery detection model and determine differential features, wherein the preset face forgery detection model includes a pre-trained visual model with frozen parameters and a low-rank matrix with learnable parameters, and the differential features include continuity features between real face image blocks and discontinuity features between forged face image blocks;

[0048] A feature redistribution optimization module is used to perform feature redistribution optimization on the differentiated features using a preset feature space redistribution strategy to determine a first optimized feature;

[0049] a feature enhancement module, configured to enhance the first optimized feature using a preset feature enhancement function to determine a second optimized feature;

[0050] a loss calculation module, configured to input the second optimized feature into a preset classification head to determine a total loss;

[0051] a parameter optimization module, configured to optimize the low-rank matrix and feature redistribution parameters of the learnable parameters using the total loss back propagation to determine a trained face forgery detection model;

[0052] The face image authenticity detection module is used to input the actual face image into the trained face forgery detection model to determine the face forgery detection result.

[0053] The third aspect of the present application provides a non-temporary computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods provided in the first aspect of the present application.

[0054] According to a fourth aspect of the present application, an electronic device is provided, comprising:

[0055] a memory having a computer program stored thereon;

[0056] A processor is used to execute the computer program in the memory to implement the steps of any one of the methods provided in the first aspect of the present application.

[0057] The face forgery detection method based on inter-tile discontinuous feature learning of the present application forms a preset face forgery detection model by setting a low-rank matrix with learnable parameters in a pre-trained visual model with frozen parameters, fine-tunes the visual model with a small number of learnable parameters, learns the continuity representation between real face image tiles and the discontinuity representation between forged face image tiles, reduces the computational complexity, does not need to rely on a specific forgery prior model structure, has good generalization, is easy to extend and migrate to other image forgery detection tasks, also has the ability to resist new forgery attack methods, and has high robustness.

[0058] Other technical effects brought by the additional features will be further illustrated in the corresponding embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0059] Other features, objects and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments with reference to the attached drawings:

[0060] Figure 1 A flowchart of a face forgery detection method based on inter-tile discontinuous feature learning according to an exemplary embodiment is shown.

[0061] Figure 2 A structural schematic diagram of a face forgery detection framework according to an exemplary embodiment is shown.

[0062] Figure 3 A comparison between the model trainable parameter amount and the cross-domain generalization performance of the method of the present application and the existing method according to an exemplary embodiment is shown.

[0063] Figure 4 A visualization result schematic diagram of the method of the present application and the ViT model recognizing the discontinuity features between tiles of multiple forgery types according to an exemplary embodiment is shown.

[0064] Figure 5 A block diagram of a face forgery detection system based on inter-tile discontinuous feature learning according to an exemplary embodiment is shown. DETAILED DESCRIPTION

[0065] The present application will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that, for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made. These all belong to the protection scope of the present application.

[0066] The terms "first", "second", etc. are used only for the purpose of description and are not to be interpreted as indicating or implying relative importance or a number of indicated technical features. Thus, features defined with "first", "second" can explicitly or implicitly include one or more of the features.

[0067] The terms "comprise", "have" and any variations thereof in the embodiments of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally further include steps or units not listed or can optionally further include other steps or units inherent to these processes, methods, products or devices.

[0068] The existing method mainly focuses on detecting fake faces based on deep learning algorithms, but its generalization ability is poor. In the face of unknown fake modes or cross-dataset fake detection tasks, the performance of the existing model decreases significantly, the structure is complex, the design cost is high, and it relies on prior knowledge, and lacks adaptability to potential new fake ways. For the Vision Transformer visual model with good image understanding ability, it needs to fine-tune all parameters, the calculation cost is high, and it lacks cross-domain robustness. Based on the above problems, the embodiments of the present application provide a face forgery detection method based on the learning of discontinuous features between the blocks to solve the above problems.

[0069] Figure 1 A flowchart of a face forgery detection method based on learning of discontinuous features between blocks according to an exemplary embodiment is shown. Figure 2 A structural schematic diagram of a face forgery detection framework according to an exemplary embodiment is shown.

[0070] Referring to Figure 1 , Figure 2 An embodiment of the present application provides a face forgery detection method based on learning of discontinuous features between blocks, which includes S11-S16.

[0071] S11, inputting a face image into a preset face forgery detection model to determine a differential feature.

[0072] Specifically, the preset face forgery detection model includes a pre-trained visual model with frozen parameters and a low-rank matrix with learnable parameters, the differential feature represents the essential differential feature between the real face image and the fake face image, and the differential feature includes the continuity feature between the real face image blocks and the non-continuity feature between the fake face image blocks, that is, the continuity feature between the image blocks in the real face image and the non-continuity feature between the image blocks in the fake face image.

[0073] The pre-trained visual model with the low-rank matrix with learnable parameters fine-tunes the frozen parameters to obtain the differentiated features of the real face images and the fake face images.

[0074] The face image represents a face image sample in the training image data set.

[0075] S12, the feature re-distribution strategy is used to optimize the differentiated features, and the first optimized features are determined.

[0076] S13, the first optimized features are enhanced by using the preset feature enhancement function, and the second optimized features are determined.

[0077] S14, the second optimized features are input into the preset classification head, and the total loss is determined.

[0078] S15, the total loss is used to optimize the low-rank matrix with learnable parameters and the feature re-distribution parameters, and the face forgery detection model is determined.

[0079] S16, the actual face image is input into the trained face forgery detection model, and the face forgery detection result is determined.

[0080] The above embodiments of the present application introduce a small number of trainable parameters to fine-tune the pre-trained visual model trained on general images by setting a low-rank matrix with learnable parameters in the pre-trained visual model with frozen parameters, thereby learning the continuity features between real face image blocks and the non-continuity features between fake face image blocks, improving the generalization ability of the pre-trained visual model, and realizing effective recognition of unknown types of fake face images; the feature space re-distribution strategy and the feature enhancement function are used to optimize the features, and the low-rank matrix with learnable parameters and the feature re-distribution parameters are optimized based on the total loss, thereby improving the accuracy of face image classification and the parameter fine-tuning accuracy of the low-rank matrix with learnable parameters of the face forgery detection model.

[0081] In some specific embodiments of the present application, the preset face forgery detection model of the present application uses a pre-trained visual Transformer (ViT) model as a feature encoder, and introduces a trainable low-rank matrix decomposition (LoRA) in each layer of the multi-layer attention module, thereby realizing the transfer of the powerful feature representation ability of the ViT model on general images to the real and fake face image feature representation in a low-dimensional space.

[0082] To introduce trainable parameters in the pre-trained visual model, in some specific embodiments of the present application, the method for determining the preset face forgery detection model can use S101 to S105.

[0083] S101, freeze the pre-trained weight of the key matrix.

[0084] Specifically, the key matrix is represented as K.

[0085] S102, freezing the pre-training weights of the query matrix and the pre-training weights of the value matrix.

[0086] Specifically, the pre-training weights of the query matrix are represented as Wq, and the pre-training weights of the value matrix are represented as Wv.

[0087] S103, superimposing a learnable first correction term on the query matrix by using the first low-rank matrix and the second low-rank matrix with learnable parameters, to determine a new query matrix.

[0088] Specifically, the new query matrix is represented as:

[0089] Q=Wq·x+B1·A1·x

[0090] Wherein, Q represents the new query matrix, Wq represents the pre-training weights of the query matrix, x represents the feature extracted by the previous layer network structure, A1 represents the first low-rank matrix, and B1 represents the second low-rank matrix.

[0091] S104, superimposing a learnable second correction term on the value matrix by using the third low-rank matrix and the fourth low-rank matrix with learnable parameters, to determine a new value matrix.

[0092] Specifically, the new value matrix is represented as:

[0093] V=Wv·x+B2·A2·x

[0094] Wherein, V represents the new value matrix, Wv represents the pre-training weights of the value matrix, x represents the feature extracted by the previous layer network structure, A2 represents the third low-rank matrix, and B2 represents the fourth low-rank matrix.

[0095] S105, determining a preset face forgery detection model according to the new query matrix and the new value matrix.

[0096] Specifically, based on the above steps 101 to S105, the pre-training weights of the key matrix, the query matrix and the value matrix of the pre-training visual model are frozen, and a learnable first correction term is superimposed on the query matrix, and a learnable second correction term is superimposed on the value matrix, thereby determining a preset face forgery detection model.

[0097] The above embodiments of the present application keep the pre-training weights Wk of the key matrix K frozen, thereby preserving the original image understanding capability of the ViT model, keep the pre-training weights Wq of the query matrix Q and the pre-training weights Wv of the value matrix V frozen, and set trainable low-rank matrix parameters for the query matrix Q and the value matrix V, i.e., the first learnable correction term and the second learnable correction term, to realize the enhancement of the true and false face image feature representation capability of the original feature space, and the low-rank matrix is only optimized in the low-dimensional subspace, thereby reducing the training complexity of the face forgery detection model and improving the training efficiency of the face forgery detection model.

[0098] To extract the features of the input face image, in some specific embodiments of the present application, for S11, the face image is input into a pre-set face forgery detection model to determine the differential features, which can adopt S111 to S112.

[0099] S111, the face image is divided into a plurality of image blocks of a pre-set size according to the pre-set size.

[0100] S112, the plurality of image blocks of the pre-set size are input into the pre-set face forgery detection model to determine the differential features.

[0101] Specifically, the pre-trained visual model can be a ViT model pre-trained in a general image.

[0102] The above embodiments of the present application divide the face image into a plurality of fixed-size image blocks to meet the input requirements of the pre-set face forgery detection model and the subsequent feature modeling requirements, extract the image features of the input face image using the pre-set face forgery detection model, and fine-tune the low-rank matrix with learnable parameters in a low-dimensional space to learn the continuity features between image blocks in real images and the discontinuity features between image blocks in forged images, thereby capturing the differential features between real face images and forged face images.

[0103] The feature redistribution parameters of the present application include a pre-set scaling factor and a pre-set noise.

[0104] To optimize the differential features by feature redistribution, in some specific embodiments of the present application, for S12, a pre-set feature space redistribution strategy is adopted to optimize the differential features by feature redistribution to determine the first optimized features, which can adopt:

[0105] According to the pre-set scaling factor and the pre-set noise, the differential features are multiplied by the pre-set scaling factor and summed with the pre-set noise to determine the first optimized features.

[0106] Specifically, the calculation formula is as follows:

[0107] F = θ ⊙ z + ∈

[0108] wherein F represents the first optimized feature, 0 represents the preset scaling factor, z represents the differentiated feature, and e represents the preset noise.

[0109] In this embodiment, the preset scaling factor obeys a Gaussian distribution N(1, s 2 ), and the preset noise obeys N(0, se2).

[0110] The above embodiments of the present application are aimed at the feature distribution mismatching problem between general real images and face images, wherein the face images include real face images and fake face images. The scaling factor and the noise are used to simulate uncertainty, adjust the distribution range of the feature space, and optimize the feature distribution of the real face images and the fake face images, so as to enhance the representation ability of the face fake detection model for real and fake faces, improve the discriminability of the real and fake face feature spaces, and perform one-time optimization on the features output by the fine-tuned ViT model, avoiding complex adjustment layer by layer and improving the calculation efficiency.

[0111] To further enhance the differentiated features, in some specific embodiments of the present application, for S13, a preset feature enhancement function is used to enhance the first optimized feature to determine the second optimized feature, which can use S131 to S136.

[0112] S131, the first optimized feature is subjected to pooling processing to determine a pooling feature vector.

[0113] Specifically, the first optimized feature is F, and the pooling feature vector is f.

[0114] S132, according to the real label and the fake label, the pooling feature vector is divided into real face features and fake face features.

[0115] Specifically, the real face features are f real , and the fake face features are f fake .

[0116] S133, according to the real face features and the fake face features, the classification direction of the real face features to the fake face features is determined.

[0117] Specifically, according to the real face features and the fake face features, a vector from the real class centroid to the fake class centroid is determined.

[0118] wherein the vector from the real class centroid to the fake class centroid is represented as

[0119] The vector from the real class centroid to the fake class centroid is subjected to normalization processing to determine the classification direction of the real face features to the fake face features.

[0120] wherein the classification direction is denoted as d cls .

[0121] S134, orthogonalizing the classification direction from the real face features to the fake face features to determine an orthogonal direction of the classification direction.

[0122] Specifically, the Gram-Schmidt orthogonalization method can be used to orthogonalize the classification direction from the real face features to the fake face features to generate an orthogonal direction d cls of the classification direction d per .

[0123] S135, respectively enhancing the real face features and the fake face features according to a preset random disturbance factor and the orthogonal direction of the classification direction to determine enhanced real face features and enhanced fake face features.

[0124] Specifically, the calculation formula is as follows:

[0125]

[0126] wherein, denotes the enhanced real face features, F real denotes the real face features before enhancement, β denotes a random disturbance factor, and f real denotes the real face features before enhancement, f real pooling, d per denotes the orthogonal direction of the classification direction, denotes the enhanced fake face features, F fake denotes the fake face features before enhancement, f fake denotes the fake face features before enhancement, f fake pooling.

[0127] S136, splicing the enhanced real face features and the enhanced fake face features to determine second optimization features.

[0128] Specifically, the second optimization features are F aug .

[0129] The above embodiments of the present application introduce a disturbance orthogonal to the classification direction without introducing additional training overhead and affecting the discrimination ability, expand the distribution range of the real and fake feature spaces, and thus obtain features with better generalization ability.

[0130] To calculate the loss and optimize the network parameters of the low-rank matrix and the feature redistribution parameters of the learnable parameters, in some embodiments of the application, for S14, the second optimization feature is input into the preset classification head to determine the total loss, including: S141-S148.

[0131] S141, the second optimization feature is input into the preset classification head to determine the predicted label of the face image.

[0132] S142, according to the true value label of the face image and the predicted label of the face image, the cross-entropy loss is determined.

[0133] Specifically, referring to steps S141-S142, according to the true value label y and the predicted label of each training sample x, the cross-entropy loss L is calculated. ce .

[0134] Specifically, the calculation formula is as follows:

[0135]

[0136] where L ce represents the cross-entropy loss, N represents the number of training samples, y i represents the true value label of the training sample, represents the predicted label of the training sample.

[0137] S143, according to the sample label of the face image, the anchor sample, the positive sample and the negative sample are determined in the pooling feature vector.

[0138] Specifically, according to the sample label of the face image, the sample label represents its true value label, including the real image label and the fake image label, the anchor sample f a , the positive sample f p and the negative sample f n are determined in the pooling feature vector f.

[0139] S144, the Euclidean distance between the anchor sample and the positive sample and the Euclidean distance between the anchor sample and the negative sample are determined.

[0140] Specifically, the Euclidean distance between the anchor sample and the positive sample is represented as d(f a ,f p ), and the Euclidean distance between the anchor sample and the negative sample is represented as d(f a ,f n ).

[0141] S145, according to the Euclidean distance between the anchor sample and the positive sample and the Euclidean distance between the anchor sample and the negative sample, the positive sample weight and the negative sample weight are respectively determined based on the dynamic weight mechanism of Softmax.

[0142] Specifically, the calculation formula of the positive sample weight is as follows:

[0143] w p =softmax(d(f a ,f p ))

[0144] wherein w p denotes the positive sample weight, f a denotes the positive sample, f p denotes the anchor sample, d(f a ,f p ) denotes the Euclidean distance between the anchor sample and the positive sample, and softmax() denotes the dynamic weight mechanism based on Softmax.

[0145] The calculation formula of the negative sample weight is as follows:

[0146] w n =softmax(-d(f a ,f n ))

[0147] wherein w n denotes the negative sample weight, and d(f a ,f n ) denotes the Euclidean distance between the anchor sample and the negative sample.

[0148] S146, according to the positive sample weight, the Euclidean distance between the anchor sample and the positive sample, and the negative sample weight, the Euclidean distance between the anchor sample and the negative sample, determines the weighted triplet loss.

[0149] Specifically, the calculation formula of the weighted triplet loss is as follows:

[0150]

[0151] wherein L tri denotes the weighted triplet loss, and δ denotes the Softplus function.

[0152] Specifically, δ is defined as δ(x) = ln(1 + exp(x)), which is used to replace the ReLU function in the traditional triplet loss, and is used to prevent the optimization problem of the hard margin interval hyperparameter.

[0153] The weighted triplet loss optimizes the discriminability of the feature space by aggregating the real face image samples and pushing the fake face image samples away from the real face image samples, forms a clear classification boundary, and improves the detection ability of unknown fake attacks.

[0154] S147: Determine a discriminant optimization loss based on feature enhancement according to the second optimization feature and the first optimization feature.

[0155] Specifically, the discriminative optimization loss based on feature enhancement is expressed as L aug , for the feature F after feature enhancement aug , using the original cross entropy loss L ce The same form is used to perform discriminant optimization to ensure the discriminative ability of the enhanced features in the classifier. In order to take into account both the computational effect and the enhancement effect, the discriminant optimization loss based on feature enhancement is expressed as L aug Enabled with a probability of 50% to achieve randomness control of the category-independent feature augmentation strategy (CIFAug).

[0156] S148 , determining a total loss according to a weighted sum of the cross entropy loss, the weighted triplet loss, and the feature enhancement-based discriminant optimization loss.

[0157] Specifically, the total loss function is calculated as follows:

[0158] L=L ce +γ1L tri +γ2L aug

[0159] Among them, L represents the total loss function, L ce represents the cross entropy loss, L tri represents the weighted triplet loss, L aug represents the discriminant optimization loss based on feature enhancement, γ1 represents the first weight parameter, and γ2 represents the second weight parameter.

[0160] Specifically, the first weight parameter γ1 and the second weight parameter γ2 are used to balance different loss weights.

[0161] In the above embodiments of the present application, the total loss adopts the weighted sum of cross entropy loss, weighted triplet loss and discriminant optimization loss based on feature enhancement, wherein the cross entropy loss is used to distinguish between real and fake categories, while the weighted triplet loss is sampled to bring the sample features of real face images closer and push away the sample features of fake face images, thereby enhancing the feature discrimination ability. The feature enhancement loss function is introduced to assist the visual model in learning more generalized real and fake face features.

[0162] To train the low-rank matrix with learnable parameters and the feature redistribution parameters in the face forgery detection model, in some embodiments of the present application, the total loss is back-propagated to optimize the low-rank matrix with learnable parameters and the feature redistribution parameters, the network parameters of the pre-trained ViT are frozen, the parameters of the low-rank matrix with learnable parameters are updated, and the above steps S11 to S14 are repeated by randomly sampling new training samples until a preset training period is reached, obtaining the trained face forgery detection model and its parameters, and at the same time, the parameters of the feature redistribution, i.e., the preset scaling factor and the preset noise, are trained.

[0163] Exemplarily, the FF++ dataset is used as the training set, the batch size is 96, the training period is 40, the learning rate value is 3e-4, the training uses the AdamW optimizer, the low-rank matrix with learnable parameters is trained, and the trained face forgery detection model and its parameters, and the trained feature redistribution parameters are obtained.

[0164] Figure 3 To show the comparison between the model's trainable parameter amount and cross-domain generalization performance of the method of the present application and the existing methods Xception, UIA-ViT, UCF, RECCE, F3-Net, DE-Adapter, Deepfake+Adapter, and EfficientNet-B4 according to an exemplary embodiment.

[0165] Referring to Figure 3 As shown, the method provided by the present application can realize end-to-end deep forgery detection model training and reasoning with 0.28M trainable parameters, and has high reasoning efficiency.

[0166] Figure 4 To show the visualization result diagram of the method of the present application and the ViT model recognizing the discontinuity features between the blocks of various forgery types according to an exemplary embodiment.

[0167] Referring to Figure 4 As shown, based on the pre-trained ViT model, the present application can learn the diversified discontinuity features between various types of forged face image blocks and the continuity features between real face image blocks with a small amount of learnable parameters, which significantly improves the generalization performance of unknown forgery type face images compared with the existing methods.

[0168] In actual operation, the present application inputs the actual face image into the trained visual model, and can directly obtain the detection result of the actual face image, i.e., whether the actual face image is a real face image or a forged face image.

[0169] The preferred features of the above embodiments can be used alone in any embodiment, or in any combination without conflict. In addition, parts not described in detail in the embodiments can be implemented using existing technologies.

[0170] The following further illustrates the present application in conjunction with specific application examples / comparative examples to facilitate a better understanding of the above technical solutions of the present application. It should be understood that the following are merely partial examples and are not intended to limit the present application.

[0171] This application provides a face forgery detection method based on inter-tile discontinuous feature learning, which is experimentally verified on the FF++(HQ), Celeb-DF, DFDC and DFD face forgery datasets to verify the generalization performance of the method provided in this application for unknown forgery types in cross-dataset scenarios.

[0172] Table 1 shows a comparison table of the generalization performance and the number of trainable parameters of the method proposed in this application and the existing method, which are trained on the FF++(HQ) dataset and tested on the three datasets of Celeb-DF, DFDC and DFD.

[0173]

[0174]

[0175] Table 1

[0176] As can be seen from Table 1, the method provided in this application fine-tunes the pre-trained visual model with only a small number of trainable parameters, achieving the strongest generalization performance across data domains. Therefore, the method provided in this application is both simple and efficient, and can significantly improve the detection performance of unknown forgery types.

[0177] Table 2 shows a comparison of the generalization performance of the method provided by this application and the existing method on the four forgery types of DF, F2F, FS and NT in the FF++(HQ) dataset, where each method is trained on one forgery type and tested on the remaining three forgery types.

[0178]

[0179] Table 2

[0180] As can be seen from Table 2, the method provided in this application still has the best average generalization performance advantage in the cross-forgery type scenario on the FF++(HQ) dataset. Compared with other existing face forgery detection methods, when the method provided in this application is used to detect face images of unknown forgery types in the same domain, the average generalization performance across the four cross-forgery types is stable and close. This set of experimental results verifies that the method provided in this application has the best generalization performance in the same domain cross-forgery type scenario.

[0181] Table 3 represents a performance comparison table of the method provided in the present application and the existing method respectively trained and tested on FF++ (HQ), Celeb-DF, DFDC and DFD face forgery datasets to verify the performance of the method provided in the present application in the training and testing of the same domain scene within each dataset.

[0182]

[0183] Table 3

[0184] From Table 3, it can be seen that the method provided in the present application can achieve the best performance on each dataset, and greatly improve the forgery detection performance on the two datasets of DFDC and DFD, verifying the effectiveness of the method provided in the present application.

[0185] Table 4 represents a comparison table of the forgery detection results of the method provided in the present application and the existing method on the FF++ (HQ) dataset with four types of interference added.

[0186]

[0187] Table 4

[0188] From Table 4, it can be seen that the method provided in the present application has good robustness performance when adding contrast interference, saturation interference, pixelization interference and Gaussian blur interference to the face image, and shows obvious performance advantage compared with the existing method.

[0189] Therefore, the face forgery detection method based on inter-tile discontinuity feature learning provided in the present application is simple and efficient, which can introduce a small amount of trainable parameters to fine-tune the pre-trained ViT model on general images, so as to learn the continuity representation between real face image blocks and the discontinuity representation between forged face image blocks; it is superior to the existing technology in the cross-dataset and cross-forgery type scene, has strong generalization, can be used for effectively identifying unknown types of forged face images, and is especially suitable for practical application scenarios where the forgery mode is constantly evolving; without complex priori or fine design, it does not depend on a specific forgery priori design complex model structure, is easy to extend and migrate to other image forgery detection tasks; it can also adapt to diversified factors such as illumination, angle and race of different face images, and has the ability to resist new forgery attack means, has high robustness, and is suitable for strong diversity.

[0190] Figure 5 FIG. 1 shows a block diagram of a face forgery detection system based on inter-tile discontinuity feature learning according to an example embodiment.

[0191] Referring to Figure 5As shown, an embodiment of the present application provides a face forgery detection system 100 based on inter-block discontinuity feature learning, comprising: a feature extraction module 110, a feature redistribution optimization module 120, a feature enhancement module 130, a loss calculation module 140, a parameter optimization module 150, and a face image authenticity detection module 160.

[0192] The feature extraction module 110 is configured to input a face image into a preset face forgery detection model to determine differential features, wherein the preset face forgery detection model comprises a pre-trained visual model with frozen parameters and a low-rank matrix with learnable parameters, and the differential features comprise continuity features between real face image blocks and discontinuity features between forged face image blocks.

[0193] The feature redistribution optimization module 120 is configured to perform feature redistribution optimization on the differential features by using a preset feature space redistribution strategy to determine first optimized features.

[0194] The feature enhancement module 130 is configured to perform enhancement processing on the first optimized features by using a preset feature enhancement function to determine second optimized features.

[0195] The loss calculation module 140 is configured to input the second optimized features into a preset classification head to determine a total loss.

[0196] The parameter optimization module 150 is configured to optimize the low-rank matrix with learnable parameters and feature redistribution parameters by backpropagation using the total loss to determine a trained face forgery detection model.

[0197] The face image authenticity detection module 160 is configured to input an actual face image into the trained face forgery detection model to determine a face forgery detection result.

[0198] In the above embodiments of the present application, the feature extraction module 110 introduces a small number of trainable parameters to fine-tune the pre-trained visual model trained on general images by setting a low-rank matrix with learnable parameters in the pre-trained visual model, thereby learning the continuity features between real face image blocks and the discontinuity features between forged face image blocks, improving the generalization ability of the pre-trained visual model, and achieving effective identification of unknown types of forged face images. The feature redistribution optimization module 120 optimizes the differential features by using a feature space redistribution strategy, the feature enhancement module 130 optimizes the differential features by using a feature enhancement function, and the loss calculation module 140 and the parameter optimization module 150 optimize the learnable low-rank matrix by backpropagation based on the total loss, thereby improving the accuracy of face image classification and the accuracy of parameter fine-tuning of the low-rank matrix with learnable parameters in the preset face forgery detection model.

[0199] For the embodiments of the system described above, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be described in detail here.

[0200] Based on the same technical concept, in some embodiments of the present application, a terminal comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor is configured to execute the method when executing the program.

[0201] Based on the same technical concept, in some embodiments of the present application, a computer readable storage medium stores a computer program, and the program is configured to execute the method when executed by a processor.

[0202] Optionally, the memory is configured to store the program; the memory can include volatile memory (English: volatile memory), such as random access memory (English: random-access memory, abbreviated: RAM), such as static random access memory (English: static random access memory, abbreviated: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviated: DDR SDRAM) and the like; the memory can also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory is configured to store computer programs (such as application programs, functional modules and the like for implementing the above method), computer instructions and the like, and the above computer programs, computer instructions and the like can be stored in one or more memories. And the above computer program, computer instruction, data and the like can be called by the processor.

[0203] The above computer program, computer instruction and the like can be stored in one or more memories. And the above computer program, computer instruction, data and the like can be called by the processor.

[0204] The processor is configured to execute the computer program stored in the memory to implement each step in the method related to the above embodiments. For details, please refer to the related description in the method embodiments.

[0205] The processor and the memory can be an independent structure, or an integrated structure. When the processor and the memory are independent structures, the memory and the processor can be coupled and connected through a bus.

[0206] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0207] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0208] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0209] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0210] The above describes some specific embodiments of the present application. It should be understood that the present application is not limited to the specific embodiments described above, and those skilled in the art may make various variations or modifications within the scope of the claims, which do not affect the substantive content of the present application. The above preferred features may be used in any combination as long as they do not conflict with each other.

Claims

1. A face forgery detection method based on learning discontinuous features between image blocks, characterized in that: include: Inputting the facial image into a preset face forgery detection model to determine differential features, wherein the preset face forgery detection includes a pre-trained visual model with frozen parameters and a low-rank matrix with learnable parameters, and the differential features include continuity features between real face image blocks and discontinuity features between forged face image blocks; Performing feature redistribution optimization on the differentiated features using a preset feature space redistribution strategy to determine a first optimized feature; Using a preset feature enhancement function to enhance the first optimized feature to determine a second optimized feature; Inputting the second optimized feature into a preset classification head to determine the total loss; Optimizing the low-rank matrix with learnable parameters and feature redistribution parameters using the total loss back propagation to determine a trained face forgery detection model; The actual face image is input into the trained face forgery detection model to determine the face forgery detection result.

2. The face forgery detection method based on inter-block discontinuous feature learning according to claim 1 is characterized in that: The method for determining the preset face forgery detection model includes: Freeze the pre-trained weights of the bond matrix; Freeze the pre-trained weights of the query matrix and the pre-trained weights of the value matrix; Superimposing a learnable first correction term on the query matrix using a first low-rank matrix and a second low-rank matrix having learnable parameters to determine a new query matrix; Using a third low-rank matrix and a fourth low-rank matrix having learnable parameters to superimpose a learnable second correction term on the value matrix to determine a new value matrix; The preset face forgery detection model is determined according to the new query matrix and the new value matrix.

3. The face forgery detection method based on inter-block discontinuous feature learning according to claim 1, characterized in that: Inputting the facial image into a preset face forgery detection model to determine the differential features includes: Dividing the facial image into a plurality of blocks of the preset size according to a preset size; Input a plurality of image blocks of the preset size into the preset face forgery detection model to determine the differential features.

4. The face forgery detection method based on inter-block discontinuous feature learning according to claim 1, characterized in that: The feature redistribution parameters include a preset scaling factor and a preset noise; The adopting a preset feature space redistribution strategy to perform feature redistribution optimization on the differentiated features to determine the first optimized feature includes: According to the preset scaling factor and the preset noise, the product of the differentiated feature and the preset scaling factor is multiplied and then summed with the preset noise to determine the first optimized feature.

5. The face forgery detection method based on inter-block discontinuous feature learning according to claim 1, characterized in that: The step of enhancing the first optimized feature using a preset feature enhancement function to determine a second optimized feature includes: Performing pooling processing on the first optimized features to determine a pooled feature vector; According to the real label and the forged label, the pooled feature vector is divided into real face features and forged face features; Determining a classification direction from the real facial features to the forged facial features based on the real facial features and the forged facial features; orthogonalizing the classification direction from the real facial feature to the forged facial feature to determine the orthogonal direction of the classification direction; According to a preset random perturbation factor and an orthogonal direction of the classification direction, the real face feature and the forged face feature are enhanced respectively to determine an enhanced real face feature and an enhanced forged face feature; The enhanced real face feature and the enhanced forged face feature are spliced ​​together to determine the second optimized feature.

6. The face forgery detection method based on inter-block discontinuous feature learning according to claim 5, characterized in that: The determining, based on the real facial features and the forged facial features, a classification direction from the real facial features to the forged facial features includes: Determining a vector from a real category centroid to a forged category centroid based on the real face features and the forged face features; Normalizing the vector from the real category centroid to the forged category centroid to determine the classification direction from the real facial feature to the forged facial feature.

7. The face forgery detection method based on inter-block discontinuous feature learning according to claim 5, characterized in that: Inputting the second optimized feature into a preset classification head to determine the total loss includes: Inputting the second optimized feature into the preset classification head to determine a predicted label for the facial image; Determining a cross entropy loss based on the true value label of the face image and the predicted label of the face image; Determining, in the pooled feature vector, anchor samples, positive samples, and negative samples according to the sample labels of the face image; Determine the Euclidean distance between the anchor point sample and the positive sample and the Euclidean distance between the anchor point sample and the negative sample; According to the Euclidean distance between the anchor point sample and the positive sample and the Euclidean distance between the anchor point sample and the negative sample, a positive sample weight and a negative sample weight are determined based on a dynamic weight mechanism of Softmax; Determining a weighted triplet loss according to the positive sample weight, the Euclidean distance between the anchor sample and the positive sample, and the negative sample weight, the Euclidean distance between the anchor sample and the negative sample; Determining a discriminant optimization loss based on feature enhancement according to the second optimization feature and the first optimization feature; The total loss is determined according to a weighted sum of the cross entropy loss, the weighted triplet loss, and the feature enhancement-based discriminant optimization loss.

8. A face forgery detection system based on learning discontinuous features between image blocks, characterized by: include: a feature extraction module configured to input a facial image into a preset face forgery detection model and determine differential features, wherein the preset face forgery detection model includes a pre-trained visual model with frozen parameters and a low-rank matrix with learnable parameters, and the differential features include continuity features between real face image blocks and discontinuity features between forged face image blocks; A feature redistribution optimization module is used to perform feature redistribution optimization on the differentiated features using a preset feature space redistribution strategy to determine a first optimized feature; a feature enhancement module, configured to enhance the first optimized feature using a preset feature enhancement function to determine a second optimized feature; a loss calculation module, configured to input the second optimized feature into a preset classification head to determine a total loss; a parameter optimization module, configured to optimize the low-rank matrix with learnable parameters and feature redistribution parameters using the total loss back propagation to determine a trained face forgery detection model; The face image authenticity detection module is used to input the actual face image into the trained face forgery detection model to determine the face forgery detection result.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Deeply-forged face image detection method and related equipment

    CN116188956A