A method and device for face deepfake detection
Through a one-stage multi-task collaboration method, multi-layer feature information is extracted and fused, and face classification, positioning and forgery detection are performed in collaboratively, which solves the problems of slow speed and limited accuracy in the existing technology, and realizes efficient face deep pseudo detection.
Patent Information
- Application Number
- CN202111569608.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-12-21
AI Technical Summary
The existing face deep pseudo detection technology usually adopts a two-stage process, which leads to slow inference speed and the second-stage analysis results are affected by the face positioning effect of the first-stage face, and it is impossible to effectively utilize the labeled data of the face deep pseudo detection dataset.
A one-stage multi-task collaboration method is adopted to extract multi-layer feature information from the input image, and input a multi-task deep pseudo-detection network through feature fusion and semantic fusion to collaborate on face classification, face positioning and fake face detection.
The speed and accuracy of face deep pseudo detection are improved, the limitations of the two-stage strategy is avoided, and the multi-task collaborative learning is used to provide richer supervision information, which improves the discriminant nature of the identification characteristics.
Smart Images

Figure CN114419693B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a technical solution for face deepfake detection with one-stage multi-task collaboration. Background Art
[0002] With the progress of machine learning and computer vision technologies, the Deepfake technology has also developed rapidly. Deepfake refers to the creation or synthesis of audiovisual content (such as images, audio and video, text, etc.) based on intelligent methods such as deep learning. The abuse of Deepfake data has brought a large number of security and privacy risks. Therefore, the detection task for Deepfake data has received increasing attention. Face Forgery specifically refers to the tampering technology for faces in Deepfake. Face deepfake detection is to determine whether the face contained in a given picture is forged.
[0003] In the prior art, the face deepfake detection technology usually adopts a two-stage process. Among them, in the first stage, an existing face detection tool is used to detect the face frame and crop out the face area; in the second stage, a face deepfake analysis tool is used to judge the face obtained in the previous step. Summary of the Invention
[0004] The purpose of this application is to provide a technical solution for face deepfake detection.
[0005] According to an embodiment of this application, a method for face deepfake detection is provided. The method includes:
[0006] Extracting multi-layer first feature information from the input image;
[0007] Obtaining multi-layer second feature information by performing feature fusion on the multi-layer first feature information;
[0008] Performing semantic fusion on each layer of the multi-layer second feature information respectively to obtain multi-layer third feature information, and inputting each layer of the third feature information into a multi-task deepfake detection network respectively to obtain the deepfake detection result output by the multi-task deepfake detection network. Among them, the multi-tasks executed collaboratively in the multi-task deepfake detection network include face classification, face localization, and forged face detection.
[0009] According to another embodiment of this application, a device for face deepfake detection is provided. The device includes:
[0010] A module for extracting multi-layer first feature information from the input image;
[0011] A module for obtaining multi-layer second feature information by performing feature fusion on the multi-layer first feature information;
[0012] A module for respectively performing semantic fusion on each layer of the multi-layer second feature information to obtain multi-layer third feature information, and respectively inputting each layer of the third feature information into a multi-task deepfake detection network to obtain a deepfake detection result output by the multi-task deepfake detection network. Among them, the multi-tasks jointly executed in the multi-task deepfake detection network include face classification, face localization, and forged face detection.
[0013] According to another embodiment of the present application, a computer device is further provided. Wherein, the computer device includes: a memory for storing one or more programs; one or more processors connected to the memory. When the one or more programs are executed by the one or more processors, the one or more processors perform the following operations:
[0014] Extract multi-layer first feature information from the input image;
[0015] Obtain multi-layer second feature information by performing feature fusion on the multi-layer first feature information;
[0016] Respectively perform semantic fusion on each layer of the multi-layer second feature information to obtain multi-layer third feature information, and respectively input each layer of the third feature information into a multi-task deepfake detection network to obtain a deepfake detection result output by the multi-task deepfake detection network. Among them, the multi-tasks jointly executed in the multi-task deepfake detection network include face classification, face localization, and forged face detection.
[0017] According to another embodiment of the present application, a computer-readable storage medium is further provided, on which a computer program is stored. The computer program can be executed by a processor to perform the following operations:
[0018] Extract multi-layer first feature information from the input image;
[0019] Obtain multi-layer second feature information by performing feature fusion on the multi-layer first feature information;
[0020] Respectively perform semantic fusion on each layer of the multi-layer second feature information to obtain multi-layer third feature information, and respectively input each layer of the third feature information into a multi-task deepfake detection network to obtain a deepfake detection result output by the multi-task deepfake detection network. Among them, the multi-tasks jointly executed in the multi-task deepfake detection network include face classification, face localization, and forged face detection.
[0021] Compared with the prior art, the present application has the following advantages: It can synchronously complete multiple tasks such as face classification, face localization, and fake face detection in one stage, which can greatly improve the speed of face deepfake detection; It can get rid of the limitation of the face localization effect in the first stage on the deepfake detection effect in the second stage in the two-stage strategy and improve the accuracy of face deepfake detection; Through the one-stage multi-task deepfake detection network and using the training method of multi-task collaborative learning, the discriminability of discriminative features is improved. The training data for the deepfake detection task can generate the annotation data for tasks such as classification, segmentation, and object detection at almost zero cost, so as to provide richer supervision information for the detection model to extract more robust and discriminative features. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Other features, objects, and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:
[0023] Figure 1 The flowchart shows the process of the method for face deepfake detection according to an embodiment of the present application;
[0024] Figure 2 The design flow framework shows an example of the face deepfake detection according to the present application;
[0025] Figure 3 The schematic diagram shows an example of the semantic fusion module according to the present application;
[0026] Figure 4 The schematic diagram shows the structure of the device for face deepfake detection according to an embodiment of the present application;
[0027] Figure 5 The exemplary system shows that it can be used to implement the various embodiments described in the present application.
[0028] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0030] As used herein, the term "device" refers to an intelligent electronic device that can perform a predetermined processing procedure such as numerical calculation and / or logical calculation by running a predetermined program or instruction. It may include a processor and a memory, and the processor executes the program instructions stored in the memory to perform the predetermined processing procedure, or the predetermined processing procedure is performed by hardware such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a digital signal processor (DSP), or is implemented by a combination of the above two.
[0031] The technical solution of this application is mainly implemented by computer devices. Among them, the computer devices include network devices and user devices. The network devices include, but are not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on cloud computing (Cloud Computing), where cloud computing is a type of distributed computing and is a super virtual computer composed of a group of loosely coupled computer sets. The user devices include, but are not limited to, PC machines, tablet computers, smart phones, IPTVs, PDAs, wearable devices, etc. Among them, the computer devices can run independently to implement this application, or can be connected to the network and implement this application through interactive operations with other computer devices in the network. Among them, the network where the computer devices are located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, VPN network, wireless ad hoc network (Ad Hoc network), etc.
[0032] It should be noted that the above computer devices are only examples, and other existing or future possible computer devices that are applicable to this application should also be included within the protection scope of this application and are included herein by reference.
[0033] The methods discussed later in this article (some of which are illustrated by flowcharts) can be implemented by hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segment for performing the necessary tasks can be stored in a machine or computer-readable medium (such as a storage medium). One or more processors can perform the necessary tasks.
[0034] The specific structures and functional details disclosed herein are merely representative and are for the purpose of describing the exemplary embodiments of this application. However, this application can be specifically implemented in many alternative forms and should not be construed as being limited only to the embodiments set forth herein.
[0035] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly, the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed related items.
[0036] The terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments. Unless the context clearly indicates otherwise, the singular forms "a", "an" used herein are also intended to include the plural. It should also be understood that the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, units and / or components, and do not preclude the presence or addition of one or more other features, integers, steps, operations, units, components and / or their combinations.
[0037] It should also be mentioned that in some alternative implementations, the functions / actions mentioned may occur in a different order than that indicated in the figures. For example, depending on the functions / actions involved, two consecutively shown figures may actually be executed substantially simultaneously or sometimes in the reverse order.
[0038] The applicant has found that the existing face deepfake detection process has the following disadvantages: 1. The inference speed of the two-stage operation is relatively slow, which does not meet the requirement of rapid judgment in the deepfake detection scenario; 2. The result of the second-stage analysis and judgment is affected by the face localization effect in the first stage, restricting the upper limit of the model, that is, inaccurate face localization will lead to deviations in the second-stage analysis and judgment; 3. The face deepfake detection dataset can generate labeled data for tasks such as classification, segmentation, and object detection at almost zero cost, and the two-stage strategy cannot utilize the features generated by these three tasks simultaneously. In view of the above disadvantages, the present application proposes a one-stage multi-task collaborative technical solution for face deepfake detection, which can complete face deepfake detection and localization through multi-task collaboration in only one stage.
[0039] The solution of the present application will be further described in detail below with reference to the accompanying drawings.
[0040] Figure 1The figure shows a schematic flowchart of a method for face deepfake detection according to an embodiment of the present application. The method according to this embodiment includes step S11, step S12, and step S13. In step S11, the computer device extracts multi-layer first feature information from the input image; in step S12, the computer device obtains multi-layer second feature information by performing feature fusion on the multi-layer first feature information; in step S13, the computer device performs semantic fusion on each layer of the multi-layer second feature information respectively to obtain multi-layer third feature information, and inputs each layer of the third feature information into a multi-task deepfake detection network respectively to obtain the deepfake detection result output by the multi-task deepfake detection network, where the multi-tasks executed collaboratively in the multi-task deepfake detection network include face classification, face localization, and forged face detection.
[0041] In step S11, the computer device extracts multi-layer first feature information from the input image. In some embodiments, the multi-layer first feature information includes, but is not limited to, any shallow features, such as color, texture, and other features. In some embodiments, the multi-layer first feature information is extracted from the input image through a feature extraction network (such as the basic network ResNet50, EfficientNet, Xception, etc.). For example, input an RGB image (i.e., the input image), the size of which is W×H×3. The basic network ResNet50 extracts the shallow first feature information such as color and texture in this image.
[0042] In step S12, the computer device obtains multi-layer second feature information by performing feature fusion on the multi-layer first feature information. In some embodiments, feature fusion is performed on the shallow features at each level extracted from the input image to obtain the multi-layer second feature information obtained by fusion. In some embodiments, feature fusion is performed through an FPN (Feature Pyramid Networks) layer. In some embodiments, step S12 further includes: performing feature fusion on the multi-layer first feature information to obtain the fused multi-layer second feature information; sampling the second feature information at the highest level in the fused multi-layer second feature information to obtain higher-level second feature information. As an example, input an RGB image with a size of W×H×3. The basic network ResNet50 extracts the shallow first feature information such as color and texture in this image. Then, first perform feature fusion on the extracted multi-layer first feature information to obtain the fused multi-layer second feature information. To increase the detection of larger target faces, further sample the second feature information at the highest level in the fused multi-layer second feature information to obtain higher-level second feature information.
[0043] In step S13, the computer device performs semantic fusion on each layer of the multi-layer second feature information respectively to obtain multi-layer third feature information, and inputs each layer of the third feature information into a multi-task deepfake detection network respectively to obtain the deepfake detection result output by the multi-task deepfake detection network. Among them, the multi-tasks executed collaboratively in the multi-task deepfake detection network include face classification, face localization, and forged face detection. Based on the multi-task deepfake detection network, multiple tasks can be executed collaboratively in one stage. Different tasks are sensitive to different features, and multiple features can promote each other. Therefore, multiple tasks can promote each other, thereby jointly helping to improve the speed, accuracy, and recall rate of face deepfake detection.
[0044] In some embodiments, for the multi-layer second feature information, semantic fusion is performed on each layer of the second feature information in the feature pyramid through a semantic fusion module to obtain third feature information, and the feature receptive field is increased to enhance the model's ability to model rigid objects. In some embodiments, the channel transformation in the semantic fusion module is completed by using a 3×3 deformable convolution with a stride of 1 to enhance the model's ability to model non-rigid objects (such as facial muscle expression movements).
[0045] In some embodiments, the multi-tasks executed collaboratively in the multi-task deepfake detection network further include forged area localization, so that the four tasks of face classification, face localization, face forgery detection, and forged area localization (segmentation) can be completed in one step. In some embodiments, each task in the multi-task deepfake detection network independently uses a 4-layer fully convolutional network to obtain the target output.
[0046] Figure 2The design process framework for face deepfake detection in an example of the present application is shown as follows: An RGB image (i.e., the input image) with a size of W×H×3 is input. The features C3, C4, and C5 are extracted from this image through the basic network ResNet50. Then, the FPN layer is used to further fuse the features to obtain P3, P4, and P5. To increase the detection of larger target faces, further sampling is performed on P5 (i.e., the feature at the highest level after fusion) to obtain a higher semantic representation P6. Among them, the channel dimensions of P3, P4, P5, and P6 are all 256. Next, the semantic fusion module based on deformable convolution is used to strengthen the features for the four feature maps respectively. By fusing multi-scale features and increasing the feature perception, the modeling ability of the model for rigid objects is enhanced, and the modeling ability of the model for non-rigid objects is enhanced through deformable convolution. After passing through the semantic fusion module, four feature maps of different sizes are obtained. Next, each feature map is connected to a multi-task detection head (i.e., the multi-task deepfake detection network) to jointly complete the execution of four tasks: face classification, face forgery detection, face localization, and forged area localization. Among them, each detection head consists of four 3×3 convolutions with a stride of 1 and the channel unchanged and a 1×1 convolution with a stride of 1. Each task independently uses a 4-layer fully convolutional network to obtain the target output. The fully convolutional target output of the face classification task is w’×h’×1; for the forged face detection task, the fully convolutional target output is w’×h’×1; for the face localization task, the fully convolutional target output is w’×h’×4; for the forged area localization task, the fully convolutional target output is w’×h’×1. It should be noted that the number of feature layers, the number of multi-tasks, etc. in the above example are only for illustration and not a limitation to the present application. Those skilled in the art can adjust some parameters in this framework based on actual needs.
[0047] Figure 3 The schematic diagram of the semantic fusion module in an example of the present application is shown. Among them, the channel transformation is all completed by a 3×3 deformable convolution with a stride of 1. The channel dimensions of the second feature information of each layer are all 256. After being input into the semantic fusion module, it is transformed from 256 to 128 (i.e., a 3×3 deformable convolution with a stride of 1, an input channel of 256, and an output channel of 128), and then from 128 to 64 (i.e., a 3×3 deformable convolution with a stride of 1, an input channel of 128, and an output channel of 64), from 64 to 64 (i.e., a 3×3 deformable convolution with a stride of 1, an input channel of 64, and an output channel of 64). Finally, the outputs of each channel transformation are merged (concat) to obtain an output with a dimension of 256.
[0048] In some embodiments, after the model training is completed, for the target image to be detected, the target image is input into the trained model. The trained model can predict and output the deepfake detection result corresponding to the target image by executing steps S11 - S13.
[0049] In some embodiments, steps S11 - S13 are operations during the model training process. During the model training process, for each sample image in the training sample set, the model can predict and output the deepfake detection result corresponding to the sample image by executing the above steps S11 - S13.
[0050] In some embodiments, the method further includes: according to the deepfake detection result, using a label assignment strategy to assign positive and negative sample labels to multiple anchor points or anchor boxes in the deepfake detection result, and calculating a loss function according to the labels. Among them, binary cross - entropy loss is used for face classification, and cross - entropy loss for face forgery detection is calculated based on the face being classified as a positive sample, and intersection - over - union loss is used for the face localization task. In some embodiments, for a face box (sparse ground - truth box, label) on the input image, during the model design process, many anchor points or anchor boxes are usually placed (only some of these anchor points or anchor boxes are positive samples, and a large number of others are negative samples) to cover as many possible faces in the picture. The label assignment strategy, that is, based on the ground - truth box of the picture label, determines which anchor points or anchor boxes on the picture are positive samples. In some embodiments, during the model training process, the deepfake detection result output by the model includes the prediction results of each task. Among them, the prediction result corresponding to the face localization task includes multiple anchor points or anchor boxes. The multiple anchor points or anchor boxes output by the model do not have ground - truth labels. Labels can be assigned to each anchor point or anchor box output by the model for the face localization task according to the deepfake detection result and the sample label, so as to calculate the loss function and measure the distance between the model prediction result and the true distribution. In some embodiments, since face authenticity classification is based on face classification and forgery area localization is based on face localization, this label assignment strategy is only for face classification and face localization tasks. Specifically, this solution adopts the dynamic sample assignment strategy proposed by ATSS (Adaptive Training Sample Selection) to select positive and negative samples; the process of the ATSS dynamic sample label assignment strategy: for the feature map corresponding to each layer of the second feature information, take the k anchor points with the smallest Euclidean distance between the anchor point and the center point of the ground - truth box (the anchor box and the anchor point are in one - to - one correspondence) as the candidate positive sample set, and then calculate the IoU (Intersection - over - Union) between the candidate anchor box and the ground - truth box. When the IoU of a certain sample is greater than the threshold of the candidate anchor point sample set (the mean + variance of the candidate set IoU, because the candidate set is dynamic, so the mean + variance is also dynamic), it is a positive sample, and this sample assignment strategy is either non - negative or positive.
[0051] In some embodiments, calculating the loss function based on the labels includes: first, weighting the positive samples according to the Euclidean distance from the center point to obtain the weighted labels, and then calculating the loss function according to the labels corresponding to the input images. In some embodiments, based on the set of positive samples assigned by ATSS (the set of negative samples remains unchanged), the distances of the set of positive samples from the center point of the true face bounding box are different. Given the prior assumption that the closer to the center point, the better the features, the set of positive samples can be further weighted according to the Euclidean distance from the center point (the closer the distance, the greater the weight). This weighting scheme can improve the model recall rate.
[0052] In some embodiments, the multiple tasks executed collaboratively in the multi-task deepfake detection network include face classification, face localization, and fake face detection. Among them, binary cross-entropy loss is used for face classification, and the cross-entropy loss for face forgery detection is calculated based on the face being classified as a positive sample (the loss of the model prediction result, that is, the loss of forgery detection only calculates the part where the face is classified as a positive sample, and the part where the face is classified as a negative sample does not consider authenticity). The IoU loss is used for the face localization task. In some embodiments, the multiple tasks executed collaboratively in the multi-task deepfake detection network also include fake region localization. Fake region localization belongs to the fine-grained face true / false classification problem and also uses binary cross-entropy loss (calculated based on face localization, that is, fake region localization only calculates the loss within the face bounding box predicted by the model, and other background regions are not included in the loss).
[0053] In some embodiments, the method further includes: weighting the loss function corresponding to each task according to a predetermined ratio to obtain a multi-task collaborative loss function. In some embodiments, the multi-task deepfake detection network collaboratively executes four tasks: face classification, face localization, fake face detection, and fake region localization. The loss function is divided into four parts. Binary cross-entropy loss is used for face classification, and the cross-entropy loss for face forgery detection is calculated based on the face being classified as a positive sample. IoU loss is used for face localization, and binary cross-entropy loss is used for fake region localization. Finally, the four parts of the loss are weighted in a certain ratio and ultimately used to guide the learning process of the model. In some embodiments, the gradient descent strategy is used for backpropagation to update and optimize the model parameters.
[0054] Figure 4The structural schematic diagram of the device for face deepfake detection according to an embodiment of the present application is shown. The device for face deepfake detection (hereinafter simply referred to as "face deepfake detection device 1") includes: a module for extracting multi-layer first feature information from an input image (hereinafter simply referred to as "module 11"), a module for obtaining multi-layer second feature information by performing feature fusion on the multi-layer first feature information (hereinafter simply referred to as "module 12"), and a module for performing semantic fusion on each layer of the multi-layer second feature information respectively to obtain multi-layer third feature information, and inputting each layer of the third feature information into a multi-task deepfake detection network respectively to obtain the deepfake detection result output by the multi-task deepfake detection network (hereinafter simply referred to as "module 13"), wherein the multi-tasks cooperatively executed in the multi-task deepfake detection network include face classification, face localization, and forged face detection.
[0055] Module 11 extracts multi-layer first feature information from the input image. In some embodiments, the multi-layer first feature information includes, but is not limited to, any shallow features, such as color, texture and other features. In some embodiments, the multi-layer first feature information is extracted from the input image through a feature extraction network (such as the basic network ResNet50, EfficientNet, Xception, etc.). For example, an RGB image (i.e., the input image) with a size of W×H×3 is input, and the basic network ResNet50 extracts the shallow first feature information such as color and texture in the image.
[0056] Module 12 obtains multi-layer second feature information by performing feature fusion on the multi-layer first feature information. In some embodiments, feature fusion is performed on the shallow features at each level extracted from the input image to obtain the multi-layer second feature information obtained by fusion. In some embodiments, feature fusion is performed through an FPN (Feature Pyramid Networks) layer. In some embodiments, module 12 is further configured to: perform feature fusion on the multi-layer first feature information to obtain the fused multi-layer second feature information; sample the second feature information at the highest level in the fused multi-layer second feature information to obtain the second feature information at a higher level. As an example, an RGB image with a size of W×H×3 is input, and the basic network ResNet50 extracts the shallow first feature information such as color and texture in the image. Then, first, feature fusion is performed on the extracted multi-layer first feature information to obtain the fused multi-layer second feature information. In order to increase the detection of larger target faces, further sampling is performed on the second feature information at the highest level in the fused multi-layer second feature information to obtain the second feature information at a higher level.
[0057] Module 13 performs semantic fusion on each layer of the multi-layer second feature information respectively to obtain multi-layer third feature information, and inputs each layer of the third feature information into the multi-task deepfake detection network respectively to obtain the deepfake detection result output by the multi-task deepfake detection network. Among them, the multi-tasks executed collaboratively in the multi-task deepfake detection network include face classification, face localization, and forged face detection. Based on the multi-task deepfake detection network, multiple tasks can be executed collaboratively in one stage. Different tasks are sensitive to different features, and multiple features can promote each other. Therefore, multiple tasks can promote each other, thereby jointly helping to improve the speed, accuracy, and recall rate of face deepfake detection.
[0058] In some embodiments, for the multi-layer second feature information, semantic fusion is performed on each layer of the second feature information in the feature pyramid through a semantic fusion module to obtain third feature information, and the feature receptive field is increased to enhance the model's ability to model rigid objects. In some embodiments, the channel transformation in the semantic fusion module is completed using 3×3 deformable convolution with a stride of 1 to enhance the model's ability to model non-rigid objects (such as facial muscle expression movements, etc.).
[0059] In some embodiments, the multi-tasks executed collaboratively in the multi-task deepfake detection network further include forged area localization, so that the four tasks of face classification, face localization, face forgery detection, and forged area localization (segmentation) can be completed in one step. In some embodiments, each task in the multi-task deepfake detection network independently uses a 4-layer fully convolutional network to obtain the target output.
[0060] Figure 2The design process framework for face deepfake detection in an example of this application is shown as follows: An RGB image (i.e., the input image) with a size of W×H×3 is input. Feature maps C3, C4, and C5 are extracted from this image through the basic network ResNet50. Then, the FPN layer is used to further fuse the features to obtain P3, P4, and P5. To increase the detection of larger target faces, P5 (i.e., the feature of the highest level after fusion) is further sampled to obtain a higher semantic representation P6. Among them, the channel dimensions of P3, P4, P5, and P6 are all 256. Next, the semantic fusion module based on deformable convolution is used to strengthen the features of the four feature maps respectively. By fusing multi-scale features and increasing the feature receptive field, the model's ability to model rigid objects is enhanced, and the model's ability to model non-rigid objects is enhanced through deformable convolution. After passing through the semantic fusion module, four feature maps of different sizes are obtained. Next, each feature map is connected to a multi-task detection head (i.e., a multi-task deepfake detection network) to jointly complete the execution of four tasks: face classification, face forgery detection, face localization, and forged area localization. Among them, each detection head consists of four 3×3 convolutions with a stride of 1 and the channel unchanged and a 1×1 convolution with a stride of 1. Each task independently uses a 4-layer fully convolutional network to obtain the target output. The fully convolutional target output of the face classification task is w’×h’×1; for the forged face detection task, the fully convolutional target output is w’×h’×1; for the face localization task, the fully convolutional target output is w’×h’×4; for the forged area localization task, the fully convolutional target output is w’×h’×1. It should be noted that the number of feature layers, the number of multi-tasks, etc. in the above example are only for illustration and not a limitation to this application. Those skilled in the art can adjust some parameters in this framework based on actual needs.
[0061] Figure 3 The schematic diagram of the semantic fusion module in an example of this application is shown. Among them, the channel transformation is completed by using a 3×3 deformable convolution with a stride of 1. The channel dimensions of the second feature information of each layer are all 256. After being input into the semantic fusion module, it is transformed from 256 to 128 (i.e., a 3×3 deformable convolution with a stride of 1, an input channel of 256, and an output channel of 128), and then from 128 to 64 (i.e., a 3×3 deformable convolution with a stride of 1, an input channel of 128, and an output channel of 64), from 64 to 64 (i.e., a 3×3 deformable convolution with a stride of 1, an input channel of 64, and an output channel of 64). Finally, the outputs of each channel transformation are merged (concat) to obtain an output with a dimension of 256.
[0062] In some embodiments, after completing the model training, for the target image to be detected, the target image is input into the trained model. The trained model can predict and output the deepfake detection result corresponding to the target image by triggering modules 11-13 to perform operations.
[0063] In some embodiments, during the model training process, for each sample image in the training sample set, the model can predict and output the deepfake detection result corresponding to the sample image by triggering modules 11-13 to perform operations.
[0064] In some embodiments, the face deepfake detection device 1 further includes: a module for assigning positive and negative sample labels to multiple anchor points or anchor boxes in the deepfake detection result according to the deepfake detection result by using a label assignment strategy, and calculating a loss function according to the labels, where face classification uses binary cross-entropy loss, and the cross-entropy loss of face forgery detection is calculated on the basis of the face being classified as a positive sample, and the intersection over union (IoU) loss is used for the face localization task. In some embodiments, for a face box (sparse ground truth box, label) on the input image, during the model design process, many anchor points or anchor boxes are usually placed (only some of these anchor points or anchor boxes are positive samples, and a large number of others are negative samples) to cover as many possible faces in the picture. The label assignment strategy, that is, based on the ground truth box of the picture label, determines which anchor points or anchor boxes on the picture are positive samples. In some embodiments, during the model training process, the deepfake detection result output by the model includes the prediction results of each task. Among them, the prediction result corresponding to the face localization task includes multiple anchor points or anchor boxes. The multiple anchor points or anchor boxes output by the model do not have ground truth labels. Labels can be assigned to each anchor point or anchor box output by the model for the face localization task according to the deepfake detection result and the sample label, so as to calculate the loss function and measure the distance between the model prediction result and the real distribution. In some embodiments, since face authenticity classification is based on face classification and forgery area localization is based on face localization, this label assignment strategy is only for face classification and face localization tasks. Specifically, this solution adopts the dynamic sample assignment strategy proposed by ATSS to select positive and negative samples; the process of the ATSS dynamic sample label assignment strategy: for the feature map corresponding to each layer of the second feature information, take the k anchor points with the smallest Euclidean distance between the anchor point and the center point of the ground truth box (the anchor box and the anchor point are in one-to-one correspondence) as the candidate positive sample set, and then calculate the IoU between the candidate anchor box and the ground truth box. When the intersection over union of a certain sample is greater than the threshold of the candidate anchor point sample set (the mean + variance of the candidate set IoU, because the candidate set is dynamic, so the mean + variance is also dynamic), it is a positive sample, and this sample assignment strategy is either non-negative or positive.
[0065] In some embodiments, calculating the loss function according to the label includes: first, weighting the positive samples according to the Euclidean distance from the center point to obtain the weighted label, and then calculating the loss function according to the label corresponding to the input image. In some embodiments, based on the set of positive samples (the set of negative samples remains unchanged) assigned by ATSS, the distances of the set of positive samples from the center point of the true face bounding box are different. Given the prior assumption that the closer to the center point, the better its features, so the set of positive samples can be further weighted according to the Euclidean distance from the center point (the closer the distance, the greater the weight). This weighting scheme can improve the model recall rate.
[0066] In some embodiments, the multiple tasks executed collaboratively in the multi-task deepfake detection network include face classification, face localization, and fake face detection. Among them, binary cross-entropy loss is used for face classification, and the cross-entropy loss of face forgery detection is calculated based on the face being classified as a positive sample (the loss of the model prediction result, that is, the loss of forgery detection only calculates the part where the face is classified as a positive sample, and the part where the face is classified as a negative sample does not consider authenticity). The Intersection-over-Union (IoU) loss is used for the face localization task. In some embodiments, the multiple tasks executed collaboratively in the multi-task deepfake detection network also include fake region localization. Fake region localization belongs to the fine-grained face authenticity classification problem and also uses binary cross-entropy loss (calculated based on face localization, that is, fake region localization only calculates the loss within the face bounding box predicted by the model, and other background regions are not included in the loss).
[0067] In some embodiments, the face deepfake detection device 1 further includes: a module for weighting the loss function corresponding to each task according to a predetermined ratio to obtain a multi-task collaborative loss function. In some embodiments, four tasks of face classification, face localization, fake face detection, and fake region localization are executed collaboratively in the multi-task deepfake detection network. The loss function is divided into four parts. Binary cross-entropy loss is used for face classification, and the cross-entropy loss of face forgery detection is calculated based on the face being classified as a positive sample. IoU loss is used for face localization, and binary cross-entropy loss is used for fake region localization. Finally, the four parts of the loss are weighted in a certain ratio and ultimately used to guide the learning process of the model. In some embodiments, the gradient descent strategy is used for backpropagation to update and optimize the model parameters.
[0068] According to the solution of the present application, multiple tasks such as face classification, face localization, and fake face detection can be synchronously completed in one stage, which can greatly improve the speed of face deepfake detection; it can get rid of the limitation of the face localization effect in the first stage on the deepfake detection effect in the second stage of the two-stage strategy and improve the accuracy of face deepfake detection; through the one-stage multi-task deepfake detection network, using the training method of multi-task collaborative learning, the discriminability of discriminative features is improved, and the training data for the deepfake detection task can generate annotation data for tasks such as classification, segmentation, and object detection at almost zero cost, so as to provide richer supervision information for the detection model to extract more robust and discriminative features.
[0069] The present application also provides a computer device, wherein the computer device includes: a memory for storing one or more programs; one or more processors connected to the memory, and when the one or more programs are executed by the one or more processors, the one or more processors execute the method for face deepfake detection described in the present application.
[0070] The present application also provides a computer-readable storage medium, on which a computer program is stored, and the computer program can be executed by a processor to execute the method for face deepfake detection described in the present application.
[0071] The present application also provides a computer program product, and when the computer program product is executed by a device, the device executes the method for face deepfake detection described in the present application.
[0072] Figure 5 An exemplary system that can be used to implement the various embodiments described in the present application is shown.
[0073] In some embodiments, the system 1000 can serve as any one of the processing devices in the embodiments of the present application. In some embodiments, the system 1000 may include one or more computer-readable media having instructions (such as a system memory or the NVM / storage device 1020) and one or more processors (such as (one or more) processors 1005) coupled to the one or more computer-readable media and configured to execute the instructions to implement modules to perform the actions described in the present application.
[0074] For one embodiment, the system control module 1010 may include any suitable interface controller to provide any suitable interface to at least one of the (one or more) processors 1005 and / or any suitable device or component communicating with the system control module 1010.
[0075] The system control module 1010 may include a memory controller module 1030 to provide an interface to the system memory 1015. The memory controller module 1030 may be a hardware module, a software module, and / or a firmware module.
[0076] The system memory 1015 may be used to load and store data and / or instructions for the system 1000, for example. For one embodiment, the system memory 1015 may include any suitable volatile memory, such as, for example, suitable DRAM. In some embodiments, the system memory 1015 may include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).
[0077] For one embodiment, the system control module 1010 may include one or more input / output (I / O) controllers to provide an interface to the NVM / storage device 1020 and the (one or more) communication interfaces 1025.
[0078] For example, the NVM / storage device 1020 may be used to store data and / or instructions. The NVM / storage device 1020 may include any suitable non-volatile memory (such as, for example, flash memory) and / or may include any suitable (one or more) non-volatile storage devices (such as, for example, one or more hard disk drives (HDDs), one or more compact discs (CDs) drives, and / or one or more digital versatile discs (DVDs) drives).
[0079] The NVM / storage device 1020 may include storage resources that are physically part of the device on which the system 1000 is installed, or it may be accessible by the device without being part of the device. For example, the NVM / storage device 1020 may be accessed via the (one or more) communication interfaces 1025 over a network.
[0080] (One or more) communication interfaces 1025 may provide an interface for the system 1000 to communicate through one or more networks and / or with any other suitable device. The system 1000 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols.
[0081] For one embodiment, at least one of the (one or more) processors 1005 may be logically encapsulated with one or more controllers of the system control module 1010 (e.g., the memory controller module 1030). For one embodiment, at least one of the (one or more) processors 1005 may be logically encapsulated with one or more controllers of the system control module 1010 to form a system-in-package (SiP). For one embodiment, at least one of the (one or more) processors 1005 may be logically integrated on the same die with one or more controllers of the system control module 1010. For one embodiment, at least one of the (one or more) processors 1005 may be logically integrated on the same die with one or more controllers of the system control module 1010 to form a system-on-chip (SoC).
[0082] In various embodiments, the system 1000 may be, but is not limited to: a server, a workstation, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the system 1000 may have more or fewer components and / or a different architecture. For example, in some embodiments, the system 1000 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit (ASIC), and speakers.
[0083] For those skilled in the art, it is obvious that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be construed as limiting the claimed claim. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. First, second, etc. are used to denote names and do not denote any particular order.
Claims
1. A method for face deepfake detection, wherein, The method includes: extracting multi-layer first feature information from an input image; obtaining multi-layer second feature information by performing feature fusion on the multi-layer first feature information; performing semantic fusion on each layer of the multi-layer second feature information respectively to obtain multi-layer third feature information, and inputting each layer of the third feature information into a multi-task deepfake detection network respectively to obtain a deepfake detection result output by the multi-task deepfake detection network, wherein the multi-tasks executed collaboratively in the multi-task deepfake detection network include face classification, face localization, and forged face detection.
2. The method according to claim 1, wherein The multi-tasks executed collaboratively in the multi-task deepfake detection network further include forged area localization.
3. The method according to claim 1, wherein, The obtaining multi-layer second feature information by performing feature fusion on the multi-layer first feature information includes: performing feature fusion on the multi-layer first feature information to obtain fused multi-layer second feature information; sampling the second feature information of the highest level in the fused multi-layer second feature information to obtain second feature information of a higher level.
4. The method according to any one of claims 1 to 3, wherein The method further includes: assigning positive and negative sample labels to multiple anchor points or anchor boxes in the deepfake detection result according to the deepfake detection result and sample labels by using a label assignment strategy, and calculating a loss function according to the labels, wherein binary cross-entropy loss is used for face classification, and cross-entropy loss for face forgery detection is calculated on the basis of the face classification being a positive sample, and intersection over union loss is used for the face localization task; wherein the label assignment strategy is used to determine which anchor points or anchor boxes in the input image are positive samples based on the ground truth boxes of the input image.
5. The method according to claim 4, wherein, The multi-tasks executed collaboratively in the multi-task deepfake detection network further include forged area localization, and the loss function for forged area localization uses binary cross-entropy loss.
6. The method according to claim 4, wherein, The method further includes: weighting the loss function corresponding to each task according to a predetermined ratio to obtain a multi-task collaborative loss function.
7. The method according to claim 4, wherein The calculating a loss function according to the labels includes: first weighting the positive samples according to the Euclidean distance from the center point to obtain weighted labels, and then calculating the loss function according to the labels corresponding to the input image.
8. An apparatus for face deepfake detection, wherein, The apparatus includes: a module for extracting multi-layer first feature information from an input image; a module for obtaining multi-layer second feature information by performing feature fusion on the multi-layer first feature information; a module for performing semantic fusion on each layer of the multi-layer second feature information respectively to obtain multi-layer third feature information, and inputting each layer of the third feature information into a multi-task deepfake detection network respectively to obtain a deepfake detection result output by the multi-task deepfake detection network, wherein the multi-tasks executed collaboratively in the multi-task deepfake detection network include face classification, face localization, and forged face detection.
9. The apparatus according to claim 8, wherein, The multi-tasks executed collaboratively in the multi-task deepfake detection network further include forged area localization.
10. The apparatus according to claim 8, wherein, The module for obtaining multi-layer second feature information by performing feature fusion on the multi-layer first feature information is configured to: perform feature fusion on the multi-layer first feature information to obtain fused multi-layer second feature information; Sample the second feature information at the highest level in the fused multi-layer second feature information to obtain second feature information at a higher level.
11. The device according to any one of claims 8 to 10, wherein, The apparatus further includes: a module configured to assign positive and negative sample labels to multiple anchors or anchor boxes in the deepfake detection result according to the deepfake detection result and the sample label by using a label assignment strategy, and calculate a loss function according to the labels, where face classification uses binary cross-entropy loss, and cross-entropy loss for face forgery detection is calculated based on the face being classified as a positive sample, and the intersection over union (IoU) loss is used for the face localization task; wherein the label assignment strategy is used to determine which anchors or anchor boxes in the input image are positive samples based on the ground truth boxes of the input image.
12. The device according to claim 11, wherein, The multi-tasks executed collaboratively in the multi-task deepfake detection network further include forged area localization, and the loss function for forged area localization uses binary cross-entropy loss.
13. The apparatus according to claim 11, wherein, The apparatus further includes: a module configured to weight the loss function corresponding to each task according to a predetermined ratio to obtain a multi-task collaborative loss function.
14. A computer device, wherein, The computer device includes: a memory for storing one or more programs; one or more processors connected to the memory, when the one or more programs are executed by the one or more processors, causing the one or more processors to execute the method according to any one of claims 1 to 7.
15. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Training method, recognition method and system based on multi-task deep learning network
CN111666873A
Detection apparatus and method of forged image, and recording medium storing program for executing method of the same in computer
KR1020130079729A