Face image depth forgery detection method, system, equipment and medium
The two-stage training mechanism for a deepfake detection model addresses the limitations of single-cue methods by capturing both macroscopic and microscopic forgery cues, improving detection precision and robustness through pseudo-force perception and classification, suitable for diverse scenarios.
Patent Information
- Application Number
- CN202510803790.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The prior art is difficult to fully capture weak forged pixel features in face images, resulting in poor depth forged detection.
A forged detection model is built, including forged force estimation branches and forged classification branches. The clues of strong forged and weak forged are extracted through a two-stage training mechanism, combined with the binary cross entropy loss function for model training, generate strong forged and complete forged estimation estimation estimation graphs, and perform loss integration to update model parameters.
It improves the detection accuracy and robustness of face image forgery, can adapt to a variety of complex scenarios, and enhances the perception of traces of faked at different abilities.
Smart Images

Figure CN120318894A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a method, system, device and medium for detecting deep fakes in face images. Background Art
[0002] With the rapid development of deep fake technology, deep learning-based facial forgery technology has significantly improved the authenticity of forged face images and has been widely used in image generation and video processing. The abuse of these technologies has caused security risks in fields such as payment security, visual communication, and identity authentication.
[0003] Current deep fake detection methods mainly rely on specific forgery clues. For example, by detecting boundary fusion features, inconsistent local region information, or frequency domain differences in forged face images to distinguish between real and fake face images. However, these methods usually rely on a single forgery clue. The traces generated by different forgery types vary greatly, and a single clue (strong forgery clue) cannot comprehensively capture all forgery features, often ignoring the forged pixels with lower forgery intensity in the forged face image, that is, ignoring weak forgery pixels. Although these weak forgery pixels are weak, the key forgery clues they contain are particularly important for improving detection accuracy, which results in some tampered areas not being recognized, causing poor detection effects. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, system, device and medium for detecting deep fakes in face images to solve the problems in the prior art in view of the above-mentioned deficiencies of the prior art.
[0005] The present invention specifically provides the following technical solutions: A method for detecting deep fakes in face images, comprising: Obtaining historical face images including real face images and corresponding forged face images; Constructing a forgery detection model, including a forgery intensity estimation branch and a forgery classification branch, and training the forgery detection model with the historical face images, and performing forgery detection on the face image to be detected with the trained forgery detection model; wherein training the forgery detection model includes the following steps: Extracting features of the historical face images through the forgery intensity estimation branch to obtain class labels and patch labels, and reshaping the patch labels into intermediate features, and obtaining a strong forgery estimation intensity map and a complete forgery estimation intensity map of the intermediate features through a two-stage training mechanism; wherein complete means including strong and weak forgery clues; Obtaining a first mask of the strong forgery estimation intensity map, and obtaining a second mask of the complete forgery estimation intensity map, and generating a forgery intensity perception loss of the first mask and the second mask during two-stage training; Classify the class labels by forging classification branches, and use the binary cross-entropy loss function to constrain the classification results to obtain a forgery classification loss for distinguishing real face images from forged face images. Integrate the forgery intensity perception loss and the forgery classification loss, and update the parameters of the forgery detection model through the integrated loss to obtain a trained forgery detection model.
[0006] Preferably, the feature extraction of the historical face image through the forgery intensity estimation branch to obtain class labels and patch labels is specifically as follows: Use the vision transformer backbone network of the forgery detection model as the forgery intensity estimation branch; Input the historical face image into the vision transformer backbone network to extract relevant features to obtain class labels and patch labels ; Among them, the historical face image , and respectively represent the height and width of the historical face image, represents the number of channels of the historical face image, is the dimension of the feature, , represents the number of patches into which the historical face image is divided.
[0007] Preferably, the reshaping of the patch labels into intermediate features and obtaining the strong forgery estimation intensity map and the complete forgery estimation intensity map of the intermediate features through a two-stage training mechanism includes: Reshape the patch labels into a feature map , and project the feature map into the forgery intensity space through a linear transformation to obtain the intermediate feature of the first stage; the specific expression is: ; Convert the intermediate feature into an estimated strong forgery estimation intensity map through a linear transformation , and the specific expression is: ; Fuse the intermediate feature of the first stage and the feature map , and generate a complete forgery estimation intensity map containing strong and weak forgery clues through a linear transformation , and the specific expression is: 。
[0008] Preferably, the first mask for obtaining the strong forgery estimation strength map and the second mask for obtaining the complete forgery estimation strength map are specifically as follows: By taking the absolute value of the pixel difference between the real face image and the corresponding forged face image , a forgery strength map is generated , and the specific expression is: ; wherein, represents the per-pixel absolute value operation; Convert the forgery strength map into the first mask and the second mask ; among them, the first mask is the strong forgery strength map mask, and the second mask is the patch-based forgery strength map mask; wherein, the conversion process of the first mask specifically includes: Set a threshold for the forgery strength map , remove the pixels with forgery strength lower than the threshold, and generate a strong forgery strength map , and the specific expression is: ; wherein, is the feature vector, is the true label of the face image; Sum the three channels of and reshape it into , where is the patch size of the image division, is the height of the face image data, is the number of patches; Sum the pixel values of the reshaped strong forgery strength map along the first dimension, obtain the sum of the forgery strength within each patch, and divide the sum of the forgery strength of each patch by the patch area , and generate the first mask ; among them, is normalized to the range of ; wherein, the conversion process of the second mask specifically includes: Sum the three channels of the forgery strength map and reshape it into ; Sum the pixel values of the reshaped strong forgery strength map Sum the pixel values, calculate the total forgery strength within each patch, and divide the total forgery strength of each patch by the patch area , to generate a second mask , where The value range of is normalized to .
[0009] Preferably, the forgery strength perception loss when generating the first mask and the second mask during two-stage training is specifically: Use the binary cross-entropy loss function to constrain the forgery strength estimation process to obtain the strong forgery strength estimation loss in the first stage and the comprehensive forgery strength estimation loss in the second stage , and the specific expression is: ; ; where represents the binary cross-entropy loss function; Through the strong forgery strength estimation loss in the first stage and the comprehensive forgery strength estimation loss in the second stage obtain the final forgery strength perception loss , and the specific expression is: .
[0010] Preferably, classify the category label through the forgery classification branch, and use the binary cross-entropy loss function to constrain the classification result to obtain the forgery classification loss for distinguishing real face images from forged face images, specifically: Input the category label into the binary classifier of the forgery classification branch, and use the binary cross-entropy loss function to constrain the classification result to obtain the specific expression of the forgery classification loss: ; where is the forgery probability predicted by the forgery detection model, represents the true label of the face image. When , the current face image is a real face image; when , the current face image is a forged face image, is the loss of the forgery classification branch, represents the binary cross-entropy loss function, The specific expression of is: ; where is the feature vector of the i th sample, The probability predicted for the i th sample; Distinguish real face images from forged face images through forged classification loss.
[0011] Preferably, after generating the forged intensity map, it further includes: Performing augmentation transformation on the forged intensity map by using shape transformation, intensity transformation, and frequency transformation to generate an augmented forged intensity map , and combining the augmented forged intensity map with the real face image to generate a new forged face image , and the generation formula is: = - ; Using the new forged face image as the training data of the forgery detection model.
[0012] The present invention provides a face image deep forgery detection system, including: An acquisition module, configured to acquire historical face images including real face images and corresponding forged face images; A model construction and detection module, configured to construct a forgery detection model, including a forged intensity estimation branch and a forged classification branch, and train the forgery detection model through the historical face images, and perform forgery detection on the face image to be detected through the trained forgery detection model; wherein training the forgery detection model includes: A forged intensity module, configured to extract features of the historical face images through the forged intensity estimation branch, obtain class labels and patch labels, reshape the patch labels into intermediate features, and obtain a strong forged estimation intensity map and a complete forged estimation intensity map of the intermediate features through a two-stage training mechanism; wherein complete means including strong and weak forged clues; A perceptual loss acquisition module, configured to obtain a first mask of the strong forged estimation intensity map, and obtain a second mask of the complete forged estimation intensity map, and generate a forged intensity perceptual loss of the first mask and the second mask during two-stage training; A classification loss acquisition module, configured to classify the class labels through the forged classification branch, and use a binary cross-entropy loss function to constrain the classification result, and obtain a forged classification loss for distinguishing real face images from forged face images; A model update module is configured to integrate the forgery intensity perception loss and the forgery classification loss, and update the parameters of the forgery detection model through the integrated loss to obtain a trained forgery detection model.
[0013] The present invention provides a computer device, including a memory and a processor. A program is stored in the memory. When the program is executed by the processor, the processor is caused to execute the steps of the above-mentioned method for detecting deep forgery of face images.
[0014] The present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for detecting deep forgery of face images are implemented.
[0015] Compared with the prior art, the present invention has the following remarkable advantages: By obtaining class labels and patch labels, the present invention captures the dual features of forged faces in terms of macroscopic attribute distribution and microscopic pixel anomalies. Through a two-stage training mechanism, a strong forgery estimation intensity map and a complete forgery estimation intensity map of intermediate features are obtained. Based on the strong forgery estimation intensity map and the complete forgery estimation intensity map, the forgery intensity perception loss during two-stage training is obtained, improving the sensitivity to local forgery traces, that is, improving the sensitivity to weak forgery clues. And the forgery classification loss is obtained through class labels, and the parameters of the forgery detection model are updated by combining the forgery intensity perception loss. The forgery detection of the face image to be detected is re-performed through the forgery detection model with updated parameters to distinguish real face images from forged face images. By perceiving forgery clues in the full range of faces, different intensities of forgery traces are effectively captured, enhancing the accuracy and robustness when detecting forged face images. And the detection accuracy is improved through multi-modal feature fusion, which can meet the requirements of various complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is the overall network model framework diagram provided by the present invention; Figure 2 is the flowchart of the data augmentation method based on forgery intensity transformation provided by the present invention; Figure 3 is the overall flowchart of a method for detecting deep forgery of face images provided by the present invention; Figure 4 is the detection flowchart of the detection model in a method for detecting deep forgery of face images provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0018] As Figure 3 and Figure 4 shown, a method for detecting deep forgery of face images in this embodiment includes the following steps: Step S1: Obtain historical face images containing real face images and corresponding forged face images.
[0019] Step S2: Construct a forgery detection model, including a forgery strength estimation branch and a forgery classification branch, and train the forgery detection model with historical face images, and perform forgery detection on the face images to be detected through the trained forgery detection model; where training the forgery detection model includes the following steps: Step S21: Extract the features of the historical face images through the forgery strength estimation branch, obtain class labels and patch labels, reshape the patch labels into intermediate features, and obtain a strong forgery estimation strength map and a complete forgery estimation strength map of the intermediate features through a two-stage training mechanism; where complete means including strong forgery and weak forgery clues.
[0020] Design a forgery strength estimation branch to extract the forgery strength map of the input face image. The forgery strength estimation is realized through a two-stage training mechanism. In the initial stage, regions with higher forgery strength are identified, and in the subsequent stage, the capture of weaker forgery clues is enhanced; the forgery strength map is generated based on the pixel difference between the real face image and the forged face image, and a patch-level forgery strength distribution is formed through a chunking operation to perceive the full range of forgery clues. That is, use the vision transformer (ViT, Vision Transformer) backbone network of the forgery detection model as the forgery strength estimation branch.
[0021] Input the historical face image into the vision transformer backbone network to extract relevant features, and obtain class labels and patch labels .
[0022] Among them, the historical face image , and respectively represent the height and width of the historical face image, represents the number of channels of the historical face image, is the dimension of the feature, , Indicates the number of patches into which the historical face image is divided.
[0023] Reshape the patch labels into intermediate features, and obtain a strong forgery estimation intensity map and a complete forgery estimation intensity map of the intermediate features through a two-stage training mechanism, including: The first stage: strong forgery intensity estimation. Reshape the patch labels into a feature map and project the feature map into the forgery intensity space through a linear transformation to obtain the intermediate features of the first stage ; The specific expression is: ; Convert the intermediate features into an estimated strong forgery estimation intensity map through a linear transformation ; The specific expression is: ; The second stage: comprehensive forgery intensity estimation. Fuse the intermediate features of the first stage and the feature map and generate a complete forgery estimation intensity map containing strong and weak forgery clues through a linear transformation .
[0024] Step S22: Obtain the first mask of the strong forgery estimation intensity map, and obtain the second mask of the complete forgery estimation intensity map, and generate the forgery intensity perception loss of the first mask and the second mask during two-stage training.
[0025] Obtain the first mask of the strong forgery estimation intensity map, and obtain the second mask of the complete forgery estimation intensity map, specifically: By taking the absolute value of the pixel difference between the real face image and the corresponding forged face image , generate a forgery intensity map ; The specific expression is: ; Among them, represents the per-pixel absolute value operation.
[0026] Convert the forgery intensity map into the first mask and the second mask ; Among them, the first mask is the strong forgery intensity map mask, and the second mask is the patch-based forgery intensity map mask.
[0027] Among them, the first mask The conversion process specifically includes: To generate a mask with strong forgery strength , set a threshold for the forgery strength map In one embodiment , remove the pixels with forgery strength lower than the threshold to generate a strong forgery strength map , and the specific expression is: ; ; Where is the feature vector, is the true label of the face image.
[0028] The pixel value range of the strong forgery strength map is , convert to a block-based strong forgery strength map mask , and quantify the forgery strength within each block to simplify the estimation task. Specifically: Sum the three channels of and reshape it into , where is the patch size for image partitioning, is the height of the face image data, is the number of patches. For the non-zero values of , directly set them to 1; sum the pixel values of the reshaped strong forgery strength map along the first dimension to obtain the total forgery strength within each patch, and divide the total forgery strength of each patch by the patch area to generate the first mask ; where 's value range is normalized to .
[0029] Among them, the pixel value range of the forgery strength map is , convert the forgery strength map to a block-based forgery strength map mask , and quantify the forgery strength within each block to simplify the estimation task. The conversion process of the second mask specifically includes: Sum the three channels of the forgery strength map and reshape it into , for the non-zero values of , directly set them to 1; sum the pixel values of the reshaped strong forgery strength map along the first dimension, calculate the total forgery strength within each patch, and divide the total forgery strength of each patch by the patch area , generate a second mask , where The value range of is normalized to .
[0030] Generate the forgery strength perception loss of the first mask and the second mask during two-stage training, specifically: Loss function design: To constrain the above forgery strength estimation process, a binary cross-entropy loss function is used to constrain the forgery strength estimation process to obtain the strong forgery strength estimation loss in the first stage and the comprehensive forgery strength estimation loss in the second stage , and the specific expression is: ; ; where represents the binary cross-entropy loss function.
[0031] Through the strong forgery strength estimation loss in the first stage and the comprehensive forgery strength estimation loss in the second stage Obtain the final forgery strength perception loss , and the specific expression is: .
[0032] Step S23: Classify the class label through the forgery classification branch, and use the binary cross-entropy loss function to constrain the classification result to obtain the forgery classification loss that distinguishes between real face images and forged face images.
[0033] Design a deepfake classification branch to extract the global features of the input face image and combine the forgery strength map to achieve the classification of face images; the classification branch integrates the forgery strength perception loss and the classification loss to distinguish between real face images and forged face images; the model uses this branch to enhance its generalization ability in cross-domain scenarios.
[0034] This step is specifically as follows: The class label extracted through the vision transformer backbone network , this class label is a global feature vector with a dimension of , which is used to represent the overall feature information of the input face image. Input the class label into the binary classifier of the forgery classification branch, and use the binary cross-entropy loss function to constrain the classification result to obtain the specific expression of the forgery classification loss: ; where is the forgery probability predicted by the forgery detection model, Represents the true label of the face image. When holds, the current face image is a real face image; when holds, the current face image is a forged face image. is the loss of the forgery classification branch. represents the binary cross-entropy loss function. The specific expression of is: Among them, is the i th eigenvector of the sample, is the i th probability predicted for the sample.
[0035] Distinguish real face images from forged face images through the forgery classification loss.
[0036] Step S24: Integrate the forgery intensity perception loss and the forgery classification loss, and update the parameters of the forgery detection model through the integrated loss to obtain the trained forgery detection model.
[0037] To comprehensively optimize the performance of the forgery intensity perception branch and the forgery classification branch, integrate the forgery intensity perception loss and the forgery classification loss to obtain the overall loss function , and update the forgery detection model through the overall loss function to obtain the updated forgery detection model, and distinguish real face images from forged face images through the updated forgery detection model. The overall loss function of the framework is defined as: ; Among them, is the loss of the forgery intensity perception branch, is the loss of the forgery classification branch. Through the above loss function , combining the forgery intensity perception loss and the forgery classification loss, the model updates the parameters using the loss gradient during the backpropagation training process, and gradually optimizes its performance in face forgery detection. and are respectively and weight coefficients, both default set to 1 to ensure the comprehensive performance of the model in forgery area intensity perception and classification.
[0038] Currently, most forgery detection models have insufficient generalization performance. Forgery detection models usually use known forged data during training, and the distribution of this data is significantly different from that of unknown forged samples. This leads to a significant decrease in the detection accuracy of existing methods when dealing with unknown forged samples, and the generated forged samples have a large difference from the actual distribution, further reducing the detection performance of the model.
[0039] Design a data augmentation method based on forgery intensity to transform the shape, intensity, and frequency of the forgery intensity map to generate diverse forged samples; the shape transformation adjusts the shape of the forgery intensity map through random grid deformation; the intensity transformation highlights weak cues by scaling the forgery intensity values; the frequency transformation adjusts the high-frequency features of the forged face image through Gaussian blur or Laplacian convolution; after the forgery intensity map is transformed, it is synthesized with the real face image to generate forged samples for model training.
[0040] The data augmentation method based on forgery intensity is as follows: Obtain a dataset containing real face images and forged face images, where the source dataset covers various face images, including both high-resolution images and low-resolution images. Generate forged samples through existing generative adversarial network (GAN) technology, and further adopt a variety of data augmentation operations to enrich the diversity of dataset samples.
[0041] During the data augmentation process, recalculate the forgery intensity map through the real face image and the forged face image . The forgery intensity map is used to characterize the intensity difference between the real face image and the forged face image. The forgery intensity map has the following calculation formula: ; The pixel value range of the forgery intensity map is .
[0042] As Figure 2 shown, in order to increase diversity, the following data augmentation method is adopted: (1) Shape transformation: Randomly sample the grid size , and perform a random deformation operation on the forgery intensity map . When the grid size is not 0, generate the corresponding deformation transformation. To avoid the negative impact of excessive deformation on the deep forgery detection performance, control the deformation degree within an appropriate range through reverse deformation operation, and finally generate the deformed forgery intensity map. (2) Intensity transformation: For the pixel regions with low forgery intensity, randomly sample the scaling factor , and scale the forgery intensity map Perform intensity scaling to reduce its overall intensity value to simulate the situation of subtle forgery in real scenarios. Through this operation, the deep forgery detection model can be guided to pay more attention to the detailed features with lower forgery intensity, thereby improving the model's perception ability of weak forgery features. (3) Frequency transformation: Based on the high-frequency information, randomly transform the forgery intensity map Specifically, randomly select Gaussian blur or Laplace convolution operations to process, enhancing or weakening the high-frequency information of the forged face image respectively, thereby changing its frequency characteristics. Through the above operations, generate the transformed forgery intensity map to increase the frequency diversity of the samples.
[0043] Use shape transformation, intensity transformation and frequency transformation to augment-transform the forgery intensity map to generate the augmented-transformed forgery intensity map . Combine the augmented-transformed forgery intensity map with the real face image to generate a new forged face image , and the generation formula is: = - ; Use the new forged face image as the training data for the forgery detection model.
[0044] The above data augmentation method significantly improves the diversity of forged samples through three methods: shape transformation, intensity transformation and frequency transformation, provides richer and more robust sample support for the training of the deep forgery detection model, and further improves the detection performance of the model. The overall network model framework of the present invention is as Figure 1 shown.
[0045] The present invention proposes a deep forgery detection system for face images, including: an acquisition module and a model construction and detection module.
[0046] Among them, the acquisition module is used to obtain historical face images containing real face images and corresponding forged face images; the model construction and detection module is used to construct a forgery detection model, including a forgery intensity estimation branch and a forgery classification branch, and train the forgery detection model through historical face images, and perform forgery detection on the face image to be detected through the trained forgery detection model; among them, training the forgery detection model includes: a forgery intensity module, a perceptual loss acquisition module, a classification loss acquisition module and a model update module.
[0047] Among them, the forgery intensity module is used to extract the features of historical face images through the forgery intensity estimation branch, obtain the class label and patch label, reshape the patch label into intermediate features, and obtain the strong forgery estimation intensity map and the complete forgery estimation intensity map of the intermediate features through a two-stage training mechanism; where "complete" means including strong forgery and weak forgery clues; the perceptual loss acquisition module is used to obtain the first mask of the strong forgery estimation intensity map, and obtain the second mask of the complete forgery estimation intensity map, and generate the forgery intensity perceptual loss of the first mask and the second mask during the two-stage training; the classification loss acquisition module is used to classify the class label through the forgery classification branch, and use the binary cross-entropy loss function to constrain the classification result to obtain the forgery classification loss for distinguishing real face images from forged face images; the model update module is used to integrate the forgery intensity perceptual loss and the forgery classification loss, and update the forgery detection model parameters through the integrated loss to obtain the trained forgery detection model.
[0048] The present invention also provides a computer device, including a memory and a processor. When a program stored in the memory is executed by the processor, the processor executes the steps of a method for detecting deep forgery of face images.
[0049] According to the disclosed embodiments, the computer device can communicate with one or more external devices (such as a keyboard, a pointing device, Bluetooth communication, etc.), or communicate with any device (such as a router, a demodulator, etc.) that enables the computing device to communicate with one or more other computing devices.
[0050] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a method for detecting deep forgery of face images are implemented.
[0051] According to the disclosed embodiments, the storage medium can be a non-volatile computer-readable storage medium, for example, it can include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0052] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for detecting deep fakes of face images, characterized in that, Including: Obtain historical face images including real face images and corresponding forged face images; Construct a forgery detection model, including a forgery intensity estimation branch and a forgery classification branch, and train the forgery detection model with the historical face images, and perform forgery detection on the face images to be detected through the trained forgery detection model; wherein training the forgery detection model includes the following steps: Extract the features of the historical face images through the forgery intensity estimation branch, obtain class labels and patch labels, reshape the patch labels into intermediate features, and obtain a strong forgery estimation intensity map and a complete forgery estimation intensity map of the intermediate features through a two-stage training mechanism; Wherein "complete" means including strong forgery and weak forgery clues; Obtain a first mask for the strong forgery estimation intensity map and a second mask for the complete forgery estimation intensity map, and generate a forgery intensity perception loss for the first mask and the second mask during two-stage training; Classify the class labels through the forgery classification branch, and use the binary cross-entropy loss function to constrain the classification results to obtain a forgery classification loss for distinguishing real face images from forged face images; Integrate the forgery intensity perception loss and the forgery classification loss, and update the parameters of the forgery detection model through the integrated loss to obtain a trained forgery detection model.
2. The face image deep forgery detection method according to claim 1, characterized in that, The extracting the features of the historical face images through the forgery intensity estimation branch to obtain class labels and patch labels is specifically: Use the vision transformer backbone network of the forgery detection model as the forgery intensity estimation branch; Input historical face images into the vision transformer backbone network to extract relevant features and obtain class tokens and patch tokens ; Among them, the historical face image , and respectively represent the height and width of the historical face image, represents the number of channels of the historical face image, is the dimension of the feature, , represents the number of patches into which the historical face image is divided.
3. The face image deep forgery detection method according to claim 2, characterized in that, The reshaping the patch labels into intermediate features and obtaining a strong forgery estimation intensity map and a complete forgery estimation intensity map of the intermediate features through a two-stage training mechanism includes: Mark the patch Reshape it into a feature map , and then Project the feature map onto the forgery strength space through a linear transformation to obtain the intermediate feature of the first stage ; The specific expression is: ; Convert the intermediate feature through a linear transformation into an estimated strong forgery estimation strength map , and the specific expression is: ; Fuse the intermediate features of the first stage and the feature map , and generate a complete forgery estimation strength map containing strong forgery and weak forgery clues through a linear transformation , and the specific expression is as follows: 。 4. The face image deep forgery detection method according to claim 3, wherein, The obtaining a first mask for the strong forgery estimation intensity map and a second mask for the complete forgery estimation intensity map is specifically: By taking the absolute value of the pixel difference between a real face image and the corresponding forged face image a forgery intensity map is generated , and the specific expression is: ; Among them, represents an absolute value operation per pixel; Convert the forgery intensity map into a first mask and a second mask ; wherein, the first mask is a strong forgery intensity map mask, and the second mask is a patch-based forgery intensity map mask; Among them, the first mask The conversion process specifically includes: Forgery intensity map Set a threshold , remove the pixels with forgery intensity lower than the threshold to generate a strong forgery intensity map , and the specific expression is: ; Among them, is the feature vector, is the true label of the face image; Sum the three channels of and reshape them into , where is the patch size for image partitioning, is the height of the face image data, is the number of patches; Sum the pixel values of the reshaped strong forgery intensity map along the first dimension to obtain the total forgery intensity within each patch, and divide the total forgery intensity of each patch by the patch area to generate the first mask ; where the value range of is normalized to ; Among them, the second mask The conversion process specifically includes: Sum the three channels of the forgery intensity map and reshape it into ; Sum the pixel values of the reshaped strong forgery intensity map along the first dimension, calculate the total forgery intensity within each patch, and divide the total forgery intensity of each patch by the patch area to generate the second mask where the value range of is normalized to . 5. The face image deep forgery detection method according to claim 4, characterized in that, The generating a forgery intensity perception loss for the first mask and the second mask during two-stage training is specifically: The binary cross-entropy loss function is used to constrain the forgery strength estimation process to obtain the strong forgery strength estimation loss in the first stage and the comprehensive forgery strength estimation loss in the second stage , and the specific expression is as follows: ; ; Among them, represents the binary cross-entropy loss function; Estimate the loss through the strong forgery intensity in the first stage and the comprehensive forgery intensity in the second stage to obtain the final forgery intensity perception loss , and the specific expression is: 。 6. The face image deep forgery detection method according to claim 5, wherein The classifying the class labels through the forgery classification branch and using the binary cross-entropy loss function to constrain the classification results to obtain a forgery classification loss for distinguishing real face images from forged face images is specifically: Label the category Input the binary classifier of the forgery classification branch, use the binary cross-entropy loss function to constrain the classification results, and obtain the specific expression of the forgery classification loss: ; Among them, is the forgery probability predicted by the forgery detection model, represents the true label of the face image. When , the current face image is a real face image; when , the current face image is a forged face image, is the loss of the forgery classification branch, represents the binary cross-entropy loss function, The specific expression of is: ; Among them, is the feature vector of the i th sample, is the probability predicted for the i th sample; Distinguish real face images from forged face images through the forgery classification loss.
7. The face image deep forgery detection method according to claim 4, wherein After generating the forgery intensity map, it further includes: Augment the forged force map by shape transformation, force transformation, and frequency transformation to generate an augmented forged force map , and combine the augmented forged force map with the real face image to generate a new forged face image , and the generation formula is: = - ; By means of newly forged face images as training data for a forgery detection model.
8. A face image deepfake detection system, characterized in that, Including: An acquisition module for obtaining historical face images including real face images and corresponding forged face images; A model construction and detection module for constructing a forgery detection model, including a forgery intensity estimation branch and a forgery classification branch, and training the forgery detection model with the historical face images, and performing forgery detection on the face images to be detected through the trained forgery detection model; Wherein training the forgery detection model includes: A forgery intensity module for extracting the features of the historical face images through the forgery intensity estimation branch, obtaining class labels and patch labels, reshaping the patch labels into intermediate features, and obtaining a strong forgery estimation intensity map and a complete forgery estimation intensity map of the intermediate features through a two-stage training mechanism; wherein "complete" means including strong forgery and weak forgery clues; A perception loss acquisition module, configured to obtain a first mask of a strong forgery estimation strength map, and obtain a second mask of a complete forgery estimation strength map, and generate a forgery strength perception loss of the first mask and the second mask during two-stage training; A classification loss acquisition module, configured to classify the class label through a forgery classification branch, and use a binary cross-entropy loss function to constrain the classification result, so as to obtain a forgery classification loss for distinguishing a real face image from a forged face image; A model update module, configured to integrate the forgery strength perception loss and the forgery classification loss, and update the forgery detection model parameters through the integrated loss, so as to obtain a trained forgery detection model.
9. A computer device, characterized in that, It includes a memory and a processor. A program is stored in the memory. When the program is executed by the processor, the processor is caused to execute the steps of the method for detecting deep forgery of face images according to any one of claims 1 to 7.
10. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, the steps of the method for detecting deep forgery of face images according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
High-robustness deep fake face detection method
CN115588226A
Face image forgery detection method and related equipment
CN115909445A
Face forgery detection method based on feature decoupling
CN116416686A
Deep forgery detection method and system based on graph convolution and multi-scale prompt fusion
CN118941936A
Forged face detecting method and apparatus thereof
US20100134250A1