Face deep counterfeiting positioning detection method, model training method and device

Through the self-mixed module, a diversified self-mixed face image is generated and combined with the positioning module, the problems of insufficient generalization ability and low positioning accuracy in the existing technology are solved, and higher generalization ability and robustness are achieved.

CN120219933APending Publication Date: 2025-06-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510279869.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing facial deep forgery positioning detection technology faces the problems of insufficient generalization ability, strong data dependence, overfitting and low positioning accuracy, especially when facing unknown forgery methods or cross-data sets, the performance is significantly reduced.

Method used

The self-mixed face image with a sense of reality is generated through the self-mixed module, which includes different colors, frequency characteristics and deformation processing, making the face fake images more diverse, and the positioning module accurately detects whether the face image is fake and locates the forged area.

Benefits of technology

It reduces the requirements for data sets, reduces the collection, labeling and processing costs of forged data, improves the generalization ability and robustness of deep forged detection models, and makes the model perform more reliable in real application environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219933A_ABST
    Figure CN120219933A_ABST
Patent Text Reader

Abstract

The invention belongs to an image processing technology, and particularly relates to a face deep counterfeiting positioning detection method, a model training method and a model training device. The method comprises the following steps: acquiring a real face image; enabling the real face image to pass through a self-mixing module to generate a self-mixing face image; enabling the real face image and the self-mixing face image to pass through a positioning module, and generating a positioning region and a detection result; calculating face detection classification loss according to the real face label of the real face image, the prediction result, the forged face label of the self-mixing face image and the prediction result; face detection segmentation loss is calculated according to the positioning area of the real face image, the forged face area of the self-mixing face image and the positioning area; according to the face detection classification loss and the face detection segmentation loss, training a face deep counterfeiting positioning detection model; according to the method, the diversified self-mixing face image training model is utilized, so that the model can adapt to a complex and changeable counterfeit image scene, the adaptability and robustness of the model to different types of counterfeit means are enhanced, and the model is more reliable in performance in a real application environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to image processing technology, and in particular relates to a face deep fake positioning detection method, a model training method and a device. Background Art

[0002] With the continuous deepening of the application of artificial intelligence technology, visual deep fake technology continues to mature and the threshold for use is getting lower and lower. Visual deep fake refers to the use of deep learning to enable computers to create or synthesize fake facial images that are indistinguishable to the human eye. Deep fake technology has caused many problems, including the spread of false information. Therefore, how to effectively detect whether the target facial image has been tampered with has become an important issue that needs to be solved in the current multimedia information security field.

[0003] In this context, deep fake detection technology has emerged, among which the positioning of the fake area is particularly important. By performing binary classification at the pixel level, the positioning technology can clearly identify the subtle differences between the fake and real areas in the image. This meticulous analysis allows the detection system to focus on the characteristics of the fake part, thereby effectively improving the detection performance. In addition, accurate area positioning helps to reduce the false alarm rate, ensure that real images are not mistakenly judged as fakes, and further enhance the public's trust. This not only improves the reliability of deep fake detection, but also provides strong support for protecting personal privacy and security.

[0004] Image tampering localization refers to identifying and marking the tampered areas in the image. In this task, it is necessary to develop effective algorithms and models, take the image to be tested as input, and output a probability map of the same size as the input image. Each element value in the probability map represents the possibility that the pixel at the corresponding position has been tampered with. By thresholding the probability map, a binary image (mask) can be obtained as the final result. Therefore, image tampering localization is essentially a pixel-level binary classification problem. Tampering localization is closely related to tampering detection. Both need to determine whether the image has been tampered with; however, they differ in goals and methods: tampering detection only determines whether a given image is original or tampered, usually relying on global information for decision-making, and has low accuracy requirements; while tampering localization requires the use of local information of the image for more detailed judgments, so it requires a higher level of technology.

[0005] Currently, Patent CN114419693A discloses a method and device for face deepfake detection, which can synchronously complete multiple tasks such as face classification, face localization, and fake face detection in one stage, greatly improving the speed of face deepfake detection; it can get rid of the limitation of the face localization effect in the first stage on the deepfake detection effect in the second stage of the two-stage strategy and improve the accuracy of face forgery detection. However, face forgery localization detection technology still faces problems such as insufficient generalization ability, strong data dependence, overfitting, and low localization accuracy. Existing models perform excellently on specific datasets, but their performance drops significantly when facing unknown forgery methods or cross-datasets. The main reason is the over-reliance on the artifacts of specific forgery algorithms and the limitations of training data. In addition, existing methods need to rely on a large amount of real forgery data for training, but the acquisition cost of real forgery data is high, the covered scenarios are limited, and the timeliness is poor, which limits the training effect and practical application ability of the model. Summary of the Invention

[0006] Aiming at the problems of the existing technology, the present invention provides a method for face deepfake localization detection, a model training method, and a device. By using a self-mixing module to generate realistic self-mixed face images, and the generated self-mixed face images contain different color, frequency features, and deformation processing, making the face forgery images more diverse; through the localization module, it can accurately detect whether a face image is forged and can accurately locate the forged area of the forged image.

[0007] In the first aspect of the present invention, the present invention proposes a method for training a model for face deepfake localization detection, and the method includes:

[0008] Obtain real face images; the real face images have real face labels;

[0009] Pass the real face images through a self-mixing module to generate self-mixed face images; the self-mixed face images have forged face labels and forged face areas; the self-mixed face images are obtained by mixing a source image, a target image, and a grayscale mask image; the source image is obtained by performing data augmentation on the real face images, the target image is obtained by performing data fine-tuning on the real face images, and the grayscale mask image is obtained by a convex hull mask image constructed from the face key points of the real face images;

[0010] Pass the real face images and the self-mixed face images through a localization module to generate a localization area and a detection result; the localization area is used to indicate the forged area of the real face images or / and the self-mixed face images, and the detection result is used to indicate whether the real face images and the self-mixed face images are real images or forged images;

[0011] Calculate the face detection classification loss based on the true face label and prediction result of the true face image, and the forged face label and prediction result of the self-mixed face image;

[0012] Calculate the face detection segmentation loss based on the localization region of the true face image, the forged face region of the self-mixed face image, and the localization region of the self-mixed face image;

[0013] Train a face deepfake localization detection model based on the face detection classification loss and the face detection segmentation loss, where the face deepfake localization detection model includes a self-mixing module and a localization module.

[0014] Further, the self-mixed face image is obtained by mixing a source image, a target image, and a grayscale mask image, including:

[0015] Calculate a first mixed image according to the product of the source image and the grayscale mask image;

[0016] Calculate a second mixed image according to the product of the target image and the inverted mask image of the grayscale mask image;

[0017] Calculate the self-mixed face image according to the sum of the first mixed image and the second mixed image.

[0018] Further, the grayscale mask image is obtained by masking the convex hull constructed by the face key points of the true face image, including:

[0019] Perform face key point localization on the true face image to obtain multiple face key points;

[0020] Divide the multiple face key points according to semantic functions to obtain multiple face key point region groups;

[0021] Perform convex hull calculation on all face key point region groups to obtain a global mask image;

[0022] Perform convex hull calculation on randomly selected key points in each face key point region group to obtain a local mask image;

[0023] Perform Gaussian blur on the global mask image or / and the local mask image to obtain the grayscale mask image.

[0024] Further, the face key point region group includes at least an eyebrow region, an eye region, a nose region, and a mouth region.

[0025] Further, generate a localization region and a detection result by passing the true face image and the self-mixed face image through the localization module, including:

[0026] The real face image and the self-mixed face image are respectively passed through the BayarConv convolutional layer to extract the noise features of the real face image and the noise features of the self-mixed face image;

[0027] The noise features of the real face image and the noise features of the self-mixed face image are respectively passed through two parallel deep residual networks to extract the low-level features, global features of the real face image, and the low-level features, global features of the self-mixed face image;

[0028] The low-level features, global features of the real face image, and the low-level features, global features of the self-mixed face image are respectively passed through the pyramid pooling module to obtain the localization region and detection result of the real face image and the localization region and detection result of the self-mixed face image.

[0029] Further, passing the low-level features, global features of the real face image, and the low-level features, global features of the self-mixed face image through the pyramid pooling module respectively to obtain the localization region and detection result of the real face image and the localization region and detection result of the self-mixed face image includes:

[0030] The low-level features, global features of the real face image, and the low-level features, global features of the self-mixed face image are respectively passed through the detection head of the pyramid pooling module to obtain the detection result of the real face image and the detection result of the self-mixed face image;

[0031] The low-level features, global features of the real face image, and the low-level features, global features of the self-mixed face image are respectively passed through the classification head of the pyramid pooling module to obtain the localization region of the real face image and the localization region of the self-mixed face image.

[0032] In the second aspect of the present invention, the present invention also provides a method for face deepfake localization detection, the method includes:

[0033] Obtain a target face image; the target face image is a face image to be detected;

[0034] Input the target face image into the localization module of the trained face deepfake localization detection model to obtain the localization region and detection result of the target face image; wherein, the face deepfake localization detection model is obtained according to the training method for the face deepfake localization detection model described in the first aspect of the present invention.

[0035] In the third aspect of the present invention, the present invention also provides a training device for a face deepfake localization detection model, the device includes:

[0036] A first acquisition unit, configured to acquire a real face image; the real face image has a real face label;

[0037] A first processing unit, configured to generate a self-mixed face image by passing the real face image through a self-mixing module; the self-mixed face image has a forged face label and a forged face area; the self-mixed face image is obtained by mixing a source image, a target image, and a grayscale mask image; the source image is obtained by performing data augmentation on the real face image, the target image is obtained by performing data fine-tuning on the real face image, and the grayscale mask image is obtained by masking the convex hull of the real face image;

[0038] A first positioning unit, configured to generate a positioning area and a detection result by passing the real face image and the self-mixed face image through a positioning module; the positioning area is used to indicate the forged area of the real face image or / and the self-mixed face image, and the detection result is used to indicate whether the real face image and the self-mixed face image are real images or forged images;

[0039] A first training unit, configured to calculate a face detection classification loss according to the real face label and prediction result of the real face image, the forged face label and prediction result of the self-mixed face image; and calculate a face detection segmentation loss according to the positioning area of the real face image, the forged face area of the self-mixed face image, and the positioning area of the self-mixed face image; and train a face deepfake positioning detection model according to the face detection classification loss and the face detection segmentation loss, where the face deepfake positioning detection model includes a self-mixing module and a positioning module.

[0040] In a fourth aspect of the present invention, the present invention further provides a face deepfake detection device, where the device includes:

[0041] A second acquisition unit, configured to acquire a target face image; the target face image is a face image to be detected;

[0042] A second positioning unit, configured to input the target face image into the positioning module of the trained face deepfake positioning detection model to obtain the positioning area and the detection result of the target face image; where the face deepfake positioning detection model is obtained according to the training method for the face deepfake positioning detection model described in the first aspect of the present invention.

[0043] The present invention has the following advantages and beneficial effects compared with the prior art:

[0044] The embodiments of the present invention can generate realistic self-mixed face images for data augmentation and transformation by using only a single real face image without relying on external data, greatly reducing the requirements for the dataset and the costs of collecting, annotating, and processing forged data. Moreover, the generated self-mixed face images contain different color, frequency features, and deformation processing, making the face forged images more diverse, providing rich training artifacts for the face deep forgery detection model, and thus improving the generalization ability of the deep forgery detection model. The embodiments of the present invention use diverse self-mixed face images to train the model, enabling the model to adapt to complex and variable forged image scenarios, enhancing the adaptability and robustness of the model to different types of forgery methods, and making the performance of the model more reliable in real application environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the drawings required in the embodiments are briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope of the present invention. Those skilled in the art can derive other related drawings based on these drawings without creative efforts. In the drawings:

[0046] Figure 1 is the training architecture diagram of the face deep forgery localization detection model according to the embodiments of the present invention;

[0047] Figure 2 is the flowchart of the training method of the face deep forgery localization detection model according to the embodiments of the present invention;

[0048] Figure 3 is the schematic diagram of the generation of self-mixed face images according to the embodiments of the present invention;

[0049] Figure 4 is the schematic diagram of the localization module according to the embodiments of the present invention;

[0050] Figure 5 is the flowchart of the face deep forgery localization detection method according to the embodiments of the present invention;

[0051] Figure 6 is the structural diagram of the training device for the face deep forgery localization detection model according to the embodiments of the present invention;

[0052] Figure 7 is the structural diagram of the face deep forgery localization detection device according to the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] Face forgery localization and detection technology still faces problems such as insufficient generalization ability, strong data dependence, overfitting, and low localization accuracy. Existing face deep forgery localization and detection models perform excellently on specific datasets, but their performance drops significantly when facing unknown forgery methods or cross-datasets. The main reason is the over-reliance on the artifacts of specific forgery algorithms and the limitations of training data. In addition, existing methods need to rely on a large amount of real forgery data for training, but the acquisition cost of real forgery data is high, the covered scenarios are limited, and the timeliness is poor, which limits the training effect and practical application ability of the model. At the same time, the model is prone to overfitting to specific forgery patterns (such as fixed mixing boundaries), resulting in insufficient robustness to complex real scenarios. In terms of forgery area localization, existing methods are difficult to accurately locate small or irregular forgery areas (such as local expression modification, background replacement), mainly because the boundaries of forgery areas are blurred or irregular, and there is a lack of fine-grained modeling of local forgery features. These problems urgently need to be solved through technological innovation.

[0055] To address the above deficiencies, embodiments of the present invention propose a face deep forgery localization and detection method, a model training method, and a device, which can be implemented by a dedicated terminal with a face deep forgery localization and detection function, a server or a server cluster with a face deep forgery localization and detection model training function, and a terminal or a server for training and deploying a face deep forgery localization and detection model adapted to different localization detections.

[0056] It should be noted that the methods provided in this specification are mainly divided into two stages, one is the model training stage, and the other is the actual application stage. In the model training stage, the terminal device can obtain real face images and generate forged face images corresponding to the real face images through some means, and then train and optimize the face deep forgery localization and detection model, so that the face deep forgery localization and detection model can distinguish whether the input face image is a real face image or a forged face image and locate the corresponding forged area. In the actual application stage, the terminal device can obtain a target face image and perform localization detection on the obtained target face image through the face deep forgery localization and detection model optimized in the model training stage to determine whether the target face image is a real face image or a forged face image.

[0057] Please refer to Figure 1 ,as Figure 1 shown.Figure 1 It is the training architecture diagram of the face deepfake localization detection model according to an embodiment of the present invention. The face deepfake localization detection model mainly includes a self-mixing module and a localization module. During the training process, first, a real face image is input, and after data preprocessing of the real face image, a standard face image is generated. On the one hand, the standard face image serves as the original face image. On the other hand, after being processed by the self-mixing module, the standard face image can generate a self-mixed face image. The original face image and the self-mixed face image are input into the localization module together. The original face image serves as a reference image for the self-mixed face image to help the localization module identify whether the input face image is a real face image or a forged face image, as well as the forged area of the input face image. By constructing the differences between the real face image and the forged face image and between the real localization area and the forged area, corresponding losses are constructed, and the face deepfake localization detection model is optimized with the corresponding losses.

[0058] Please refer to Figure 2 ,such as Figure 2 shown in Figure 2 It is the flowchart of the training method of the face deepfake localization detection model according to an embodiment of the present invention. The method includes the following steps:

[0059] 101. Obtain a real face image; the real face image has a real face label;

[0060] In the embodiment of the present invention, operations such as the acquisition, storage, use, processing, transmission, provision, and disclosure of the real face image all comply with the provisions of relevant laws and regulations and do not violate public order and good customs. For example, the above operations are all executed on the premise of obtaining authorization.

[0061] In the embodiment of the present invention, the real face image can be some public face databases, such as the LFW (Labeled Faces in the Wild) dataset, the CAS-PEAL face database, etc. The real face image can include face images of different races, genders, and ages, and can also cover various variation factors, such as expressions, postures, lighting, etc.

[0062] In a preferred embodiment of the present invention, in order to better train the face deepfake localization detection model, this embodiment also performs operations on the input real face image including but not limited to standardization processing, including cropping, alignment, and resizing, etc.

[0063] 102. Generate a self-mixed face image from the real face image through a self-mixing module. The self-mixed face image has a forged face label and a forged face area. The self-mixed face image is obtained by mixing a source image, a target image, and a grayscale mask image. The source image is obtained by performing data augmentation on the real face image, the target image is obtained by performing data fine-tuning on the real face image, and the grayscale mask image is obtained by masking the convex hull constructed from the face key points of the real face image.

[0064] Please refer to Figure 3 , such as Figure 3 shown in Figure 3 FIG. is a schematic diagram of generating a self-mixed face image according to an embodiment of the present invention. During the generation process, first, a face image is input, and after positioning and removing background information from the face image, a standard face image is generated. On the one hand, the standard face image can generate a source image and a target image through some methods. On the other hand, through face key point positioning and convex hull processing on the standard face image, a global mask and a local mask can be generated. Then, by randomly selecting the global mask and the local mask, a grayscale mask image is generated. After mixing the source image, the target image, and the grayscale mask image, a self-mixed face image can be obtained.

[0065] In the embodiment of the present invention, the source image is a pseudo-image generated by performing a series of data augmentations on a single real face image. These data augmentation methods include operations such as color jitter (such as RGB channel offset, hue / saturation / brightness / contrast adjustment), frequency transformation (such as downsampling or sharpening), scaling, and translation, aiming to simulate the modified areas (such as forged facial features or backgrounds) during the forgery process and actively introduce statistical inconsistencies (such as color differences, frequency anomalies) and common forgery artifacts such as blending boundaries. The target image is usually a version that retains the real face image or only undergoes minor transformations, representing the unmodified parts (such as backgrounds or retained facial areas) during the forgery process. By mixing the source image, the target image, and the grayscale mask image, a self-mixed face image is generated, enabling the model to learn general forgery detection features, significantly improving the generalization ability across datasets and forgery methods, and since the source image and the target image are both generated from the same original image, it avoids the complex source-target image pair search process in traditional methods and has higher computational efficiency.

[0066] In an embodiment of the present invention, a first mixed image is calculated based on the product of a source image and a grayscale mask image; a second mixed image is calculated based on the product of a target image and an inverted mask image of the grayscale mask image; and a self-mixed face image is calculated based on the summation of the first mixed image and the second mixed image. In this way, the source image and the target image can be fused according to the grayscale mask image to control the mixing area and intensity, and finally a self-mixed face image with forged features is generated. This self-mixing process can simulate the edge transition characteristics and local inconsistencies of the forged area, enhance the authenticity and diversity of the forged samples, generate a natural forgery effect, effectively simulate the visual features of the forged face, and thus improve the focusing ability and detection accuracy of the model for the forged area.

[0067] In an embodiment of the present invention, face key points are located on the real face image to obtain multiple face key points; the multiple face key points are divided according to semantic functions to obtain multiple groups of face key point regions; wherein, the groups of face key point regions at least include an eyebrow region, an eye region, a nose region, and a mouth region. Convex hull calculations are performed on all groups of face key point regions to obtain a global mask image; convex hull calculations are performed on randomly selected key points of each group of face key point regions to obtain local mask images; and the global mask image or / local mask images are subjected to Gaussian blur to obtain a grayscale mask image. The edge transition phenomenon and local inconsistencies can be simulated using the grayscale mask, and the fusion of local features of the image and the enhancement of the forgery effect can be achieved through mask control. The generated forged image, that is, the self-mixed face image, has rich artifact features, such as key point misalignment and inconsistent boundary features, and has a high forgery fidelity. At the same time, the self-mixing process has a constant computational time complexity, which is independent of the scale of the face dataset, can quickly generate diverse forged samples, and provides effective support for subsequent forgery detection and region localization.

[0068] For example, based on each group of face key point regions, local area masks (such as only covering the eyes, mouth, or nose, etc.) can be flexibly generated, so as to achieve more refined control in the process of generating forged images. Through the global mask, a forged image covering the entire face can be generated, which is suitable for scenarios that require overall modification; while through the local mask, the range and intensity of the mixing area can be flexibly controlled, and a forged image that only modifies the local area can be generated, which can train the model to pay more attention to local areas such as eyebrows, eyes, and noses. Then, the mask is randomly scaled, deformed, and Gaussian blurred to adapt to the mixing requirements of the image. The source image and the target image are fused according to the mask mask to control the mixing area and intensity, and finally a self-mixed face image with forged features is generated. This mixing process can generate a natural forgery effect, thus effectively simulating the visual features of the forged face.

[0069] Exemplarily, the mixing formula adopted by the self-mixing process is expressed as:

[0070] Isb = Is ☉ M + It ☉ (1 - M)

[0071] Wherein, Isb represents the self-mixed face image, M represents the grayscale mask image, Is and It respectively represent the source image and the target image, and ☉ represents the Hadamard product.

[0072] Wherein, when performing self-mixing on each image, the significance level of the source image can be controlled by adjusting the mixing ratio. For example, 0.25, 0.5, 0.75, or 1 is used. When a more significant source image is required, the mixing ratio is appropriately increased, and when a less significant source image is required, the mixing ratio is appropriately decreased.

[0073] 103. Pass the real face image and the self-mixed face image through the positioning module to generate a positioning area and a detection result; the positioning area is used to indicate the forged area of the real face image or / and the self-mixed face image, and the detection result is used to indicate whether the real face image and the self-mixed face image are real images or forged images;

[0074] In the embodiment of the present invention, the real face image and the self-mixed face image are respectively passed through the BayarConv convolutional layer to extract the noise features of the real face image and the noise features of the self-mixed face image; the noise features of the real face image and the noise features of the self-mixed face image are respectively passed through two parallel deep residual networks to extract the low-level features, global features of the real face image, and the low-level features, global features of the self-mixed face image; the low-level features, global features of the real face image, and the low-level features, global features of the self-mixed face image are respectively passed through the pyramid pooling module to obtain the positioning area and detection result of the real face image and the positioning area and detection result of the self-mixed face image. In the embodiment of the present invention, high-frequency noise information is extracted through the BayarConv high-pass filter to enhance the subtle differences between the real area and the forged area; two independent ResNet50 networks are used to extract multi-level features, and the features are fused through a four-layer convolutional network to enhance the significance of the forged area; the pyramid pooling module is used to process multi-scale information to capture the detailed characteristics of the forged area, and finally a binary classification mask image is generated through the detection head to clearly label the forged area. This method effectively combines multiple technical means and significantly improves the accuracy and robustness of forged area positioning.

[0075] Such as Figure 4As shown, in order to improve the recognition ability of the positioning module for subtle face forgery traces, in this embodiment, a BayarConv layer is added to extract subtle noise features. The BayarConv layer is a special convolutional structure dedicated to extracting subtle noise features from images, especially suitable for forgery detection tasks. Forged images usually expose forgery traces in details, such as tiny color distortions, inconsistent textures, etc. These details are usually difficult to detect by conventional convolutional networks, while the BayarConv layer can more sensitively capture these subtle abnormalities in the forged areas through a specific convolutional method. Based on the conventional convolutional layer, the BayarConv layer suppresses the main image content during the convolutional calculation process and only retains the noise components, thereby more effectively locating the forged areas. By training the BayarConv high-pass filter, it can learn the high-frequency noise distribution of the input image, thereby enhancing the network's ability to identify the subtle differences between the real area and the forged area during feature extraction. The design of the BayarConv filter can identify the tiny noise changes in the image, helping the model find the noise feature differences in the forged areas of the forged image. BayarConv restricts the weight distribution of the convolutional kernel, making the convolutional layer only respond to certain specific frequency features. Compared with the conventional convolutional kernel, the BayarConv kernel reduces the interference of the main image content through a constraint method, making it easier to detect the forged details in the image.

[0076] After the preliminary noise feature extraction is completed, the positioning module further performs deep feature extraction and multi-level feature fusion through the ResNet50 structure. ResNet50 is a classic deep residual network. By introducing residual connections, it can effectively solve the gradient vanishing problem in deep networks, thereby being able to learn more complex image features. Moreover, the present invention uses two parallel ResNet50 network structures. The first ResNet50 network focuses on extracting low-level features, such as edge information and local texture, while the second ResNet50 network focuses on higher-level global features, such as shape, structure, etc. These two networks process the same input image respectively and fuse the features at a specific level to obtain a more comprehensive feature representation of the forged area. Using two independent ResNet50 networks as the backbone, multi-level features of the image are extracted to form a multi-level feature representation. The feature maps of each level are concatenated hierarchically to include different scale information and local features of the image, and the multi-level features are further fused through four convolutional layers to enhance the saliency of the forged area. By enhancing the saliency of the forged area, the positioning module can more easily distinguish between the forged and real areas, reduce the influence of interference information, and thus improve the recognition accuracy of the forged area. This saliency enhancement helps the model focus more on the forged area, improving the positioning accuracy and robustness, especially being able to accurately locate the forged area in complex scenarios.

[0077] Since the forgery regions may appear in different sizes and positions in the image, simple fixed-scale feature extraction is difficult to effectively handle all cases. The Pyramid Pooling Module (PPM) in the embodiments of the present invention integrates features of different scales through multi-scale pooling operations, enabling the model to accurately capture forgery features at different scales. The feature map output by the PPM is input into two branches, namely the detection head and the localization head, for two-class detection (real / fake) and forgery region localization. The detection head first extracts global features through a `3x3` convolutional kernel and a global average pooling layer. The number of channels in the convolutional layer is `256`, and then a fully connected layer is used to achieve two-class classification (real or fake). The Softmax function is used as the activation function, and the predicted class is finally output. The localization head converts the feature map into a single-channel heatmap representing the probability distribution of the forgery region through a `3x3` convolutional kernel. The stride of the convolutional kernel is `1`, and the padding is `1`. The size of the generated heatmap is `224x224` to visualize the forgery region.

[0078] Through this localization module, the position of the forgery region can be clearly identified and marked in the forged image, forming a binary classification mask map with a clear boundary from the real region to identify the forgery region.

[0079] 104. Calculate the face detection classification loss according to the real face label and prediction result of the real face image, and the forged face label and prediction result of the self-mixed face image.

[0080] In the embodiments of the present invention, a multi-task loss function is adopted, including the face detection classification loss and the face detection segmentation loss. The face detection classification loss is used to measure the difference between the real face image and the forged face image, and the face detection segmentation loss is used to measure the difference between the forgery region and the real region.

[0081] Specifically, for the face detection classification loss, in the embodiments of the present invention, the face detection classification loss is calculated by comparing the real label with the probability distribution predicted by the model, so as to optimize the model's discrimination ability for face categories.

[0082] In the embodiments of the present invention, the classification loss (Loss_cls) uses binary cross-entropy loss. This loss function is used to measure the difference between the model's predicted probability and the real label and is applicable to binary classification tasks (such as distinguishing real pictures and forged pictures). For each face image sample i, the real label is y i (real pictures are labeled as 1, and forged pictures are labeled as 0), and the model's predicted probability is p i , and the specific loss function is defined as:

[0083]

[0084] 105. Among them, N is the batch size. i can be either a real face image or a self-mixed face image. Correspondingly, yi can correspond to the real face label of the real face image and the forged face label of the self-mixed face image, and pi can correspond to the prediction result of the real face image and the prediction result of the self-mixed face image. According to the positioning region of the real face image, the forged face region of the self-mixed face image, and the positioning region of the self-mixed face image, calculate the face detection and segmentation loss;

[0085] Specifically, for the face detection and segmentation loss, in the embodiments of the present invention, by comparing the real label with the probability distribution predicted by the model, the face detection and segmentation loss is calculated, so as to optimize the discriminative ability of the model for the face forgery region.

[0086] Exemplarily, the binary cross-entropy loss is used to calculate the segmentation loss (Loss_seg) of the model. By reducing the difference between the predicted label of the forged region in the output and the real label of the forged region, the positioning accuracy of the model for the forged region is improved. The real label (ground-truth) of the forged region is the mixed mask generated in the image mixing stage. The mixed mask is obtained by detecting the face key points and calculating their convex hulls. This method can accurately define the forged region and provide accurate real labels for model training. In the training stage, these mixed masks are used as the real labels of the forged regions and are input into the model together with the self-mixed images, so as to supervise the model to locate the forged regions. It optimizes the model by calculating the binary cross-entropy loss of each pixel point and taking the average value of all pixel points. Specifically, for each pixel point j, the real label is q j , and the predicted label is r j , and the specific loss function is defined as:

[0087]

[0088] Among them, M is the total number of image pixels. Similar to the classification loss, for each pixel point j, it can be a pixel point in a real face image or a self-mixed face image.

[0089] 106. According to the face detection classification loss and the face detection and segmentation loss, train the face deep forgery localization and detection model, where the face deep forgery localization and detection model includes a self-mixed module and a localization module.

[0090] In the embodiments of the present invention, the weighted sum of the face detection classification loss and the face detection and segmentation loss is used as the final loss function, and the minimization of the final loss function is used as the optimization goal. This final loss function is defined as:

[0091] Loss = α·Loss cls +β·Loss seg

[0092] Among them, α and β are weight coefficients that adjust the influence degrees of the two tasks.

[0093] By adjusting the model parameters of the face deepfake localization detection model, the face deepfake localization detection model trained by this method integrates self-mixed face images with rich training artifacts, improves the generalization ability of the deepfake detection model, reduces the requirements for the dataset, and reduces the costs of collecting, annotating, and processing forged data. This enables the model to adapt to complex and changeable forged image scenarios, enhances the adaptability and robustness of the model to different types of forgery means, and makes the performance of the model more reliable in real application environments.

[0094] It can be understood that the self-mixing module and the localization module constitute the face deepfake localization detection model, and the self-mixing module and the localization module need to operate in coordination. The self-mixing module simulates the forged false face pictures, which contain richer artifacts, and then inputs the self-mixed pictures and themselves into the localization module to perform feature extraction and fusion in the localization module to locate the forged area.

[0095] In the embodiment of the present invention, through the combination of the self-mixing module and the localization module, an image self-mixing technology based on a gray mask is used to generate highly realistic forged samples, and at the same time, technologies such as the attention mechanism and high-frequency noise extraction are used to accurately locate the forged area. The self-mixing module can simulate the edge transition and local inconsistency in the forged image, improve the saliency of the forged area, while the localization module accurately identifies and locates the forged area through multi-layer feature extraction and pyramid pooling technology, thereby improving the accuracy and robustness of deepfake detection.

[0096] The face deepfake localization detection method provided by the embodiment of the present invention will be described below. Among them, the method can be implemented on a terminal or a server. Taking the implementation on a terminal as an example, a client of an application software for face deepfake localization detection function is set on the terminal, and a trained face deepfake localization detection model is deployed in the server to implement the inspection of the target face image on the terminal side. In this process, the terminal device uploads the target face image to be detected to the server, or the server directly calls the target face image to be detected in the database, and then uses the trained face deepfake localization detection model to verify the target face image to obtain the detection result and the localization area. The server can feedback the detection result to the terminal device, or keep the detection result locally for other business applications or processing.

[0097] Please refer to Figure 5 , as Figure 5 shown, Figure 5It is a flowchart of a method for detecting face deep forgery localization according to an embodiment of the present invention. The method includes the following steps:

[0098] 201. Obtain a target face image; the target face image is a face image to be detected.

[0099] In an embodiment of the present invention, the target face image is a face image to be detected for forged faces. For example, it can be a face image collected during face recognition for unlocking or a face image collected during face scan payment.

[0100] 202. Input the target face image into the localization module of the trained face deep forgery localization detection model to obtain the localization region and detection result of the target face image; wherein, the face deep forgery localization detection model is obtained according to the training method for the face deep forgery localization detection model described in the present invention.

[0101] After collecting the target face image, the trained face deep forgery localization detection model performs forgery localization detection on the target face image to obtain the forgery detection result and localization region corresponding to the target face image. Among them, the localization region is used to indicate the forged region of the target face image, and the detection result is used to indicate whether the target face image is a real image or a forged image; in particular, if the target face image is a real face image, then the localization region of the target face image is an image with all pixels being 0 or 1, indicating that none of the localization regions of the target face image are forged regions.

[0102] Please refer to Figure 6 As Figure 6 shown, Figure 6 It is a structural diagram of a training device for a face deep forgery localization detection model according to an embodiment of the present invention. The device includes:

[0103] The first acquisition unit 301 is used to acquire real face images; the real face images have real face labels.

[0104] The first processing unit 302 is used to generate self-mixed face images by passing the real face images through a self-mixing module; the self-mixed face images have forged face labels and forged face regions; the self-mixed face images are obtained by mixing a source image, a target image, and a grayscale mask image; the source image is obtained by performing data augmentation on the real face image, the target image is obtained by performing data fine-tuning on the real face image, and the grayscale mask image is obtained by masking the convex hull of the real face image.

[0105] The first positioning unit 303 is configured to generate a positioning region and a detection result for the real face image and the self-mixed face image through a positioning module; the positioning region is used to indicate the forged region of the real face image or / and the self-mixed face image, and the detection result is used to indicate whether the real face image and the self-mixed face image are real images or forged images;

[0106] The first training unit 304 is configured to calculate a face detection classification loss according to the real face label and prediction result of the real face image, the forged face label and prediction result of the self-mixed face image; and calculate a face detection segmentation loss according to the positioning region of the real face image, the forged face region of the self-mixed face image, and the positioning region of the self-mixed face image; and train a face deepfake positioning detection model according to the face detection classification loss and the face detection segmentation loss, where the face deepfake positioning detection model includes a self-mixing module and a positioning module.

[0107] Through the cooperation of the first acquisition unit 301, the first processing unit 302, the first positioning unit 303, and the first training unit 304 in the embodiments of the present invention, a model can be trained using diverse self-mixed face images, enabling the model to adapt to complex and changing forged image scenarios, enhancing the adaptability and robustness of the model to different types of forgery methods, and making the performance of the model more reliable in real application environments.

[0108] Please refer to Figure 7 , such as Figure 7 shown Figure 7 is a structural diagram of a face deepfake detection device according to an embodiment of the present invention. The device includes:

[0109] The second acquisition unit 401 is configured to acquire a target face image; the target face image is a face image to be detected;

[0110] The second positioning unit 402 is configured to input the target face image into the positioning module of the trained face deepfake positioning detection model to obtain the positioning region and the detection result of the target face image; wherein, the face deepfake positioning detection model is obtained according to the training method for the face deepfake positioning detection model of the present invention.

[0111] In a preferred embodiment of the present invention, the face deepfake detection device may also be a terminal with the function of detecting fake faces, and the terminal includes but is not limited to: wearable devices, handheld devices, personal computers, tablet computers, vehicle-mounted devices, smart phones, computing devices, or other processing devices connected to a wireless modem, etc. In different networks, the terminal may be called by different names. For example: user equipment, access terminal, user unit, user station, mobile station, mobile device, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), 5th Generation Mobile Communication Technology (5G) network, 4th Generation Mobile Communication Technology (4G) network, 3rd-Generation (3G) network, or a terminal in a future evolved network, etc.

[0112] It can be understood that the technical solution of the present application is mainly implemented by computer devices. Among them, the computer devices include network devices and user devices. The network devices include but are not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on cloud computing (Cloud Computing). Among them, cloud computing is a type of distributed computing, which consists of a super virtual computer composed of a group of loosely coupled computer sets. The user devices include but are not limited to PC machines, tablet computers, smart phones, IPTVs, PDAs, wearable devices, etc. Among them, the computer devices can run independently to implement the present application, or can be connected to the network and implement the present application through interaction with other computer devices in the network. Among them, the network where the computer devices are located includes but is not limited to the Internet, wide area network, metropolitan area network, local area network, VPN network, wireless ad hoc network (Ad Hoc network), etc.

[0113] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: ROM, RAM, magnetic disk, optical disk, etc.

[0114] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A training method for a face deep fake positioning detection model, characterized in that: The method comprises: Acquire a real face image; the real face image has a real face label; The real face image is passed through a self-mixing module to generate a self-mixing face image; the self-mixing face image has a forged face label and a forged face area; the self-mixing face image is obtained by mixing a source image, a target image and a grayscale mask image; the source image is obtained by performing data enhancement on the real face image, the target image is obtained by performing data fine-tuning on the real face image, and the grayscale mask image is obtained by a convex hull mask image constructed by face key points of the real face image; The real face image and the self-mixed face image are passed through a positioning module to generate a positioning area and a detection result; the positioning area is used to indicate the forged area of ​​the real face image and / or the self-mixed face image, and the detection result is used to indicate whether the real face image and the self-mixed face image are real images or forged images; Calculate the face detection classification loss according to the real face label and prediction result of the real face image and the forged face label and prediction result of the self-mixed face image; Calculating face detection segmentation loss according to the localization area of ​​the real face image, the forged face area of ​​the self-mixing face image, and the localization area of ​​the self-mixing face image; According to the face detection classification loss and the face detection segmentation loss, a face deep fake positioning detection model is trained, and the face deep fake positioning detection model includes a self-mixing module and a positioning module.

2. A training method for a face deep fake positioning detection model according to claim 1, characterized in that: The self-mixing face image is obtained by mixing the source image, the target image and the grayscale mask image, including: A first mixed image is obtained by calculating the product of the source image and the grayscale mask image; A second mixed image is obtained by calculating the product of the target image and the inverse mask image of the grayscale mask image; A self-mixed face image is obtained by calculation based on the sum of the first mixed image and the second mixed image.

3. A training method for a face deep fake positioning detection model according to claim 1 or 2, characterized in that: The grayscale mask image is obtained by constructing a convex hull mask image of the face key points of the real face image, including: Positioning facial key points of the real face image to obtain a plurality of facial key points; Dividing multiple facial key points according to semantic functions to obtain multiple facial key point region groups; Calculate the convex hull of all facial key point area groups to obtain a global mask map; For each facial key point area group, randomly select key points to calculate the convex hull and obtain a local mask map; Gaussian blur is performed on the global mask image or / and the local mask image to obtain a grayscale mask image.

4. A training method for a face deep fake positioning detection model according to claim 3, characterized in that: The facial key point region group includes at least an eyebrow region, an eye region, a nose region and a mouth region.

5. The training method for a face deep fake positioning detection model according to claim 1, characterized in that: Passing the real face image and the self-mixed face image through a positioning module to generate a positioning area and a detection result includes: Passing the real face image and the self-mixed face image through a BayarConv convolution layer respectively to extract noise features of the real face image and noise features of the self-mixed face image; The noise features of the real face image and the noise features of the self-mixed face image are respectively passed through two parallel deep residual networks to extract low-level features and global features of the real face image and low-level features and global features of the self-mixed face image; The low-level features and global features of the real face image and the low-level features and global features of the self-mixed face image are respectively passed through a pyramid pooling module to obtain the positioning area and detection results of the real face image and the positioning area and detection results of the self-mixed face image.

6. A training method for a face deep fake positioning detection model according to claim 5, characterized in that: The low-level features and global features of the real face image and the low-level features and global features of the self-mixed face image are respectively passed through a pyramid pooling module to obtain the positioning area and detection result of the real face image and the positioning area and detection result of the self-mixed face image, including: The low-level features and global features of the real face image and the low-level features and global features of the self-mixed face image are respectively passed through the detection head of the pyramid pooling module to obtain the detection results of the real face image and the detection results of the self-mixed face image; The low-level features and global features of the real face image and the low-level features and global features of the self-mixed face image are respectively passed through the classification head of the pyramid pooling module to obtain the positioning area of ​​the real face image and the positioning area of ​​the self-mixed face image.

7. A method for detecting deep fake facial features, characterized in that: The method comprises: Acquire a target face image; the target face image is a face image to be detected; The target facial image is input into the positioning module of the trained facial deep fake positioning detection model to obtain the positioning area and detection result of the target facial image; wherein the facial deep fake positioning detection model is obtained according to the training method for facial deep fake positioning detection model according to any one of claims 1-6.

8. A training device for a face deep fake positioning detection model, characterized in that: The device comprises: A first acquisition unit is used to acquire a real face image; the real face image has a real face label; A first processing unit is configured to generate a self-mixing face image by passing the real face image through a self-mixing module; the self-mixing face image has a forged face label and a forged face area; the self-mixing face image is obtained by mixing a source image, a target image and a grayscale mask image; the source image is obtained by performing data enhancement on the real face image, the target image is obtained by performing data fine-tuning on the real face image, and the grayscale mask image is obtained by a convex hull mask image of the real face image; A first positioning unit, configured to pass the real face image and the self-mixed face image through a positioning module to generate a positioning area and a detection result; the positioning area is used to indicate a forged area of ​​the real face image and / or the self-mixed face image, and the detection result is used to indicate whether the real face image and the self-mixed face image are real images or forged images; The first training unit is used to calculate the face detection classification loss according to the real face label of the real face image, the prediction result of the real face image, the forged face label of the self-mixed face image and the prediction result of the self-mixed face image; and calculate the face detection segmentation loss according to the positioning area of ​​the real face image, the forged face area of ​​the self-mixed face image and the positioning area of ​​the self-mixed face image; and train a face deep fake positioning detection model according to the face detection classification loss and the face detection segmentation loss, the face deep fake positioning detection model including a self-mixing module and a positioning module.

9. A facial deep fake detection device, characterized in that: The device comprises: A second acquisition unit is used to acquire a target face image; the target face image is a face image to be detected; A second positioning unit is used to input the target facial image into the positioning module of the trained facial deep fake positioning detection model to obtain the positioning area and detection result of the target facial image; wherein, the facial deep fake positioning detection model is obtained according to the training method for the facial deep fake positioning detection model according to any one of claims 1-6.