Face recognition method and device, equipment, storage medium and computer program product
By using a multi-scale and multi-branched deep forged face detection model in the face recognition system, combined with the authenticity classifier and mask regressor, the problem of existing systems being difficult to detect forged face images is solved, and higher detection accuracy and system security are achieved.
Patent Information
- Application Number
- CN202510276587.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-09
AI Technical Summary
Existing facial recognition systems are difficult to detect fake facial images, which leads to threatening personal privacy and property security.
A multi-scale and multi-branched deep fake face detection model is adopted, combining the authenticity classifier, mask regressor and gradient inversion domain class classifier, and the model parameters are updated to improve detection accuracy by training real and fake face images in the dataset.
Effectively detect and identify fake face images, improve the security and reliability of the face recognition system, and solve the problem of domain offset in fake face image detection.
Smart Images

Figure CN119964221A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of face recognition technology, and in particular to face recognition methods, devices, equipment, storage media and computer program products. Background Art
[0002] With the rapid progress of optical imaging technology and deep learning technology, face recognition technology has been widely used in many fields such as user authentication, mobile payment, access control systems, etc., and has become an important means of identity authentication in modern society. However, with the popularization of face recognition technology, deep fake face technology has also emerged and developed rapidly. This technology uses deep learning algorithms to generate highly realistic fake face images or videos, which can easily imitate the facial features of others, thereby bypassing the existing face recognition system, posing a serious threat to personal privacy and property security.
[0003] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention
[0004] The main purpose of this application is to provide a face recognition method, device, equipment, storage medium and product, aiming to solve the technical problem that existing face recognition systems are difficult to detect forged face images.
[0005] To achieve the above purpose, the present application proposes a face recognition method, which includes:
[0006] The face image to be recognized is input into a trained deep fake face detection model to obtain a face recognition result, wherein the deep fake face detection model includes a head module, and the head module includes a true and false classifier, a mask regressor, and a gradient reversal domain category classifier.
[0007] In one embodiment, before the step of inputting the face image to be recognized into the trained deep fake face detection model to obtain the face recognition result, the step further includes:
[0008] Collect real face images;
[0009] Generate a forged face image based on the real face image, and generate a mask for the forged face area in the forged face image;
[0010] Based on the real face image and the forged face image, construct a face image dataset;
[0011] Based on the facial image dataset, the model parameters of the deep fake face detection model are updated by a back propagation algorithm to obtain a trained deep fake face detection model.
[0012] In one embodiment, the step of updating the model parameters of the deep fake face detection model by a back propagation algorithm based on the face image dataset to obtain a trained deep fake face detection model includes:
[0013] Randomly select a target face image from the face image dataset;
[0014] Input the target face image into the deep fake face detection model to obtain a true and false binary classification loss value, a fake region mask loss value, and a domain category loss value;
[0015] Based on the loss weight coefficient, determining the overall loss value according to the true-false binary classification loss value, the forged area mask loss value and the domain category loss value;
[0016] Based on the overall loss value, the model parameters of the deep fake face detection model are updated through a back propagation algorithm to obtain a trained deep fake face detection model.
[0017] In one embodiment, the step of inputting the target face image into the deep fake face detection model to obtain the true and false binary classification loss value, the forged area mask loss value and the domain category loss value comprises:
[0018] Determining authenticity information, mask information, and domain category information of the target face image;
[0019] Inputting the target face image into the deep fake face detection model to obtain predicted authenticity information, predicted mask information, and predicted domain category information;
[0020] Based on the cross entropy loss function, determining a true-false binary classification loss value according to the true-false information and the predicted true-false information;
[0021] Based on the segmentation task loss function, determining a forged area mask loss value according to the mask information and the predicted mask information;
[0022] Based on a cross entropy loss function, a domain category loss value is determined according to the domain category information and the predicted domain category information.
[0023] In one embodiment, the deep fake face detection model further includes a key part cropping module, a feature extraction module, an attention module and a connection module, and the head module includes a separate head module and a comprehensive head module; wherein the step of inputting the target face image into the deep fake face detection model to obtain the true and false binary classification loss value, the forged area mask loss value and the domain category loss value comprises:
[0024] Inputting the target face image into the deep fake face detection model, and cropping the target face image using the key parts cropping module to obtain a key parts image;
[0025] Extracting features from the key part image by the feature extraction module to obtain a feature matrix;
[0026] Processing the feature matrix by the separate head module to obtain a first processing result;
[0027] The feature matrices output by different feature extraction modules are superimposed by the connection module to obtain a target feature matrix;
[0028] Processing the target feature matrix through the attention module to obtain an attention matrix;
[0029] Processing the attention matrix through a comprehensive head module to obtain a second output result;
[0030] Based on the first output result and / or the second output result, determine a true-false binary classification loss value, a forged area mask loss value, and a domain category loss value.
[0031] In one embodiment, the step of inputting the face image to be recognized into a trained deep fake face detection model to obtain a face recognition result includes:
[0032] Input the face image to be identified into the trained deep fake face detection model, output the target authenticity probability through the authenticity classifier, and output the target forged area through the mask regressor;
[0033] When the target authenticity probability is greater than a preset authenticity probability threshold and the target forged area is smaller than a preset area threshold, determining that the face recognition result is a real face;
[0034] When the target authenticity probability is less than or equal to the preset authenticity probability threshold or the target forgery area is greater than or equal to the preset area threshold, the face recognition result is determined to be a forged attack face.
[0035] In addition, to achieve the above-mentioned purpose, the present application also proposes a face recognition device, the face recognition device comprising:
[0036] An input module is used to input a face image to be recognized into a trained deep fake face detection model to obtain a face recognition result, wherein the deep fake face detection model includes a head module, and the head module includes a true and false classifier, a mask regressor, and a gradient reversal domain category classifier.
[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes a face recognition device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the face recognition method described above.
[0038] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the face recognition method described above are implemented.
[0039] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the face recognition method described above are implemented.
[0040] One or more technical solutions proposed in this application have at least the following technical effects:
[0041] The face recognition method, apparatus, device, storage medium and computer program product proposed in the present application obtain face recognition results by inputting the face image to be recognized into a trained deep fake face detection model, wherein the deep fake face detection model includes a head module, and the head module includes a true and false classifier, a mask regressor and a gradient reversal domain category classifier, which solves the technical problem that the existing face recognition system is difficult to detect forged face images. Compared with the existing technology, the present application adopts a multi-scale and multi-branch deep fake face detection model to extract more detailed texture information of the face image to be recognized, and then adopts a gradient reversal domain category classifier to improve the consistency of face image data of different categories, which helps to solve the domain shift problem in deep fake image detection, and jointly determines the face recognition result through the true and false classifier and the mask regressor, so that forged face images can be detected and identified more effectively and accurately, thereby improving the security and reliability of the face recognition system. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0044] Figure 1A schematic diagram of a flow chart provided for the first embodiment of the face recognition method of the present application;
[0045] Figure 2 A schematic diagram of the flow chart provided for the second embodiment of the face recognition method of the present application;
[0046] Figure 3 A schematic diagram of the structure of the deep fake face detection ink fragrance provided in the second embodiment of the face recognition method of the present application;
[0047] Figure 4 A schematic diagram of the structure of the integrated head module provided in the second embodiment of the face recognition method of the present application;
[0048] Figure 5 A schematic diagram of the structure of the attention module provided in Example 2 of the face recognition method of the present application;
[0049] Figure 6 This is a schematic diagram of the module structure of the face recognition device according to an embodiment of the present application;
[0050] Figure 7 Schematic diagram of the device structure of the hardware operating environment involved in the face recognition method in the embodiment of the present application.
[0051] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0052] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0053] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0054] The main solution of the embodiment of the present application is: input the face image to be recognized into a trained deep fake face detection model to obtain a face recognition result, wherein the deep fake face detection model includes a head module, and the head module includes a true and false classifier, a mask regressor, and a gradient reversal domain category classifier.
[0055] It can be seen from the above embodiments that the present application obtains a face recognition result by inputting the face image to be identified into a trained deep fake face detection model, wherein the deep fake face detection model includes a head module, and the head module includes a true and false classifier, a mask regressor, and a gradient reversal domain category classifier, which solves the technical problem that the existing face recognition system is difficult to detect forged face images. Compared with the prior art, the present application adopts a multi-scale and multi-branch deep fake face detection model to extract more detailed texture information of the face image to be identified, and then adopts a gradient reversal domain category classifier to improve the consistency of face image data of different categories, which helps to solve the domain shift problem in deep fake image detection, and jointly determines the face recognition result through the true and false classifier and the mask regressor, so that forged face images can be detected and identified more effectively and accurately, thereby improving the security and reliability of the face recognition system.
[0056] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a face recognition device, etc. The following takes face recognition as an example to illustrate this embodiment and the following embodiments.
[0057] Based on this, the present application embodiment provides a face recognition method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the face recognition method of the present application.
[0058] In this embodiment, the face recognition method includes steps S10 to S40:
[0059] Step S10, inputting the face image to be recognized into the trained deep fake face detection model to obtain a face recognition result, wherein the deep fake face detection model includes a head module, and the head module includes a true and false classifier, a mask regressor, and a gradient reversal domain category classifier.
[0060] It should be noted that the head module includes three main classifiers: the authenticity classifier is used to judge the authenticity of the face, the mask regressor is used to output which areas in the image are forged, and the gradient reversal domain category classifier refers to the domain category classifier that has passed the gradient reversal layer. It is used to distinguish which domain the image comes from (for example, the real image domain or a specific forged image domain). It can reduce the feature differences between data in different domains, so that the deep fake face detection model can maintain a consistent recognition level when facing different types of forged data.
[0061] In a feasible implementation, the step of inputting the face image to be identified into a trained deep fake face detection model to obtain a face recognition result includes: inputting the face image to be identified into a trained deep fake face detection model, outputting the target authenticity probability through the authenticity classifier, and outputting the target forged area area through the mask regressor; when the target authenticity probability is greater than a preset authenticity probability threshold and the target forged area area is less than a preset area threshold, determining that the face recognition result is a real face; when the target authenticity probability is less than or equal to the preset authenticity probability threshold or the target forged area area is greater than or equal to the preset area threshold, determining that the face recognition result is a forged attack face.
[0062] It should be noted that in the recognition stage of the deep fake face detection model, the true and false classifier in the head module is used mainly, and the mask regressor is used as a supplement to determine whether the face image to be recognized input into the model is a real face or a fake attack face.
[0063] In the specific implementation, the authenticity classifier will output a probability distribution of authenticity, and the mask regressor will output which areas in the image are forged areas. In actual application, since face images from different channels have different aspect ratios, clarity, image size, brightness, etc., and different face images have different tolerances for authenticity, the outputs of the authenticity classifier and the mask regressor can be combined to jointly judge the authenticity of the face, that is: when the probability of the real face output by the authenticity classifier (that is, the target authenticity probability) is greater than the preset authenticity threshold, and the target forged area output by the mask regressor is also required to be less than the preset area threshold, the face image to be identified can be determined as a real face. Otherwise, the face image to be identified is determined to be a forged attack face.
[0064] In this embodiment, the accuracy of the deep fake face detection model can be effectively improved by combining the output results of the authenticity classifier and the mask regressor of the comprehensive head module in the deep fake face detection model to jointly perform authenticity determination.
[0065] This embodiment obtains a face recognition result by inputting the face image to be identified into a trained deep fake face detection model, wherein the deep fake face detection model includes a head module, and the head module includes a true or false classifier, a mask regressor, and a gradient reversal domain category classifier, which solves the technical problem that the existing face recognition system is difficult to detect forged face images. Compared with the existing technology, the present application adopts a multi-scale and multi-branch deep fake face detection model to extract more detailed texture information of the face image to be identified, and then adopts a gradient reversal domain category classifier to improve the consistency of face image data of different categories, which helps to solve the domain shift problem in deep fake image detection, and jointly determines the face recognition result through the true or false classifier and the mask regressor, so that forged face images can be detected and identified more effectively and accurately, thereby improving the security and reliability of the face recognition system.
[0066] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 2 Before step S10, the face recognition method further includes steps S01 to S04:
[0067] Step S01, collecting real face images;
[0068] It should be noted that real face images are usually collected from a diverse group of people, including faces of different ages, genders, races, expressions, lighting conditions, and shooting angles. These images can be obtained through public datasets (such as CelebA, VGGFace2, etc.) or through custom collection methods. In addition, in order to improve the generalization ability of the model, the diversity and complexity of real face images should be ensured as much as possible.
[0069] Step S02, generating a forged face image according to the real face image, and generating a mask for the forged face area in the forged face image;
[0070] It should be noted that deep fake technology can be used to generate fake face images based on real face images. Currently, there are a variety of deep fake technologies that can be used to generate fake face images, such as methods based on generative adversarial networks, such as DCGAN, StyleGAN, etc. These algorithms can generate high-quality and realistic fake faces. In addition, there are also face-changing technologies based on autoencoders (such as Deepfakes, FaceSwap, etc.), which achieve face replacement by learning the feature representation of the face. When generating fake face images, you can also combine a variety of driving methods, such as expression driving, optical flow field driving, etc., to increase the diversity and authenticity of fake face images.
[0071] In a specific implementation, the generation of a fake face image usually includes: first, detecting and cropping the real face image to extract the face area; then generating a fake face area through a selected deep fake algorithm; and finally, fusing the fake face area into the target image to form a complete fake face image.
[0072] It should be noted that the mask is usually a binary image that is the same as the forged face image, which can help identify the forged face area in the forged face image, where the forged face area is marked as 1 (or any positive value), and other areas except the forged face area are marked as 0. The mask can be generated by an image segmentation algorithm or obtained through the intermediate output of the forgery algorithm. The role of the mask is to help the model better learn the characteristics of the forged area.
[0073] Step S03, constructing a face image dataset based on the real face image and the forged face image;
[0074] It should be noted that before constructing a face image dataset, face detection and face alignment need to be performed on all real face images and forged face images. Among them, face detection is the process of identifying the position of the face in the image. Commonly used detection algorithms include methods based on deep learning (such as MTCNN, YOLO, etc.); face alignment is to adjust the detected face to a uniform posture and size for subsequent processing. Face alignment is usually achieved through key point detection (such as eyes, nose, mouth, etc.), and then the face is adjusted to a standard position through affine transformation. Performing face detection and face alignment on all images can ensure that the size ratio of all images in the face image dataset remains consistent and aligned to a standard position, thereby reducing the impact of changes in face position, size or direction on model performance and improving model robustness.
[0075] Step S04: Based on the facial image dataset, the model parameters of the deep fake face detection model are updated by a back propagation algorithm to obtain a trained deep fake face detection model.
[0076] It should be noted that the deep fake face detection model can detect fake images by learning the feature differences between real faces and fake faces. During the training process, a variety of loss functions (such as adversarial loss, reconstruction loss, etc.) can be used to optimize model performance.
[0077] In a feasible implementation, the step of updating the model parameters of the deep fake face detection model through a back propagation algorithm based on the face image dataset to obtain a trained deep fake face detection model includes: randomly selecting a target face image from the face image dataset; inputting the target face image into the deep fake face detection model to obtain a true or false binary classification loss value, a forged area mask loss value, and a domain category loss value; based on a loss weight coefficient, determining an overall loss value according to the true or false binary classification loss value, the forged area mask loss value, and the domain category loss value; based on the overall loss value, updating the model parameters of the deep fake face detection model through a back propagation algorithm to obtain a trained deep fake face detection model.
[0078] It should be noted that one or more face images are randomly selected from the face image dataset as training samples. These images should include real face images and forged face images so that the model can learn to distinguish between the real and the fake.
[0079] It should be noted that the overall loss value can be used to update the model parameters through the back propagation algorithm. During each iterative training process, the model parameters will be adjusted according to the gradient of the loss function to reduce the prediction error.
[0080] It should be noted that during the training process of the model, it is necessary to continuously use the overall loss value obtained from each iteration to update the model parameters through the back propagation algorithm until the overall loss value converges, and finally obtain a trained deep fake face detection model that can accurately detect fake faces on new images to be identified.
[0081] In a feasible implementation, the step of inputting the target facial image into the deep fake face detection model to obtain the authenticity binary classification loss value, the forged area mask loss value and the domain category loss value includes: determining the authenticity information, mask information and domain category information of the target facial image; inputting the target facial image into the deep fake face detection model to obtain predicted authenticity information, predicted mask information and predicted domain category information; based on the cross entropy loss function, determining the authenticity binary classification loss value according to the authenticity information and the predicted authenticity information; based on the segmentation task loss function, determining the forged area mask loss value according to the mask information and the predicted mask information; based on the cross entropy loss function, determining the domain category loss value according to the domain category information and the predicted domain category information.
[0082] It should be noted that for each target face image, it is necessary to determine whether it is real or forged, which can be determined through manual annotation or authenticity labeling (i.e., authenticity information); mask information refers to the annotation of the forged area in the image, which is usually a binary image of the same size as the original image, in which the forged part is marked as 1 (or any positive value) and the non-forged part is marked as 0; domain category information refers to the category of the image source, such as different data sets, different generation methods or different forgery techniques.
[0083] It should be noted that the specific formula for calculating the loss value of the true and false binary classification using the cross entropy loss function is as follows:
[0084]
[0085] In the formula, y i Indicates true or false information, p i Represents the predicted true or false information, and N represents the number of training samples.
[0086] It should be noted that Dice Loss is mainly used for segmentation tasks. It is based on the Dice coefficient, which is used to measure the similarity between two sets. The Dice coefficient is defined as:
[0087]
[0088] Among them, X represents the predicted mask area (ie, predicted mask information), Y represents the actual mask area (ie, actual mask information), and the value of the Dice coefficient ranges from 0 to 1. The closer the value is to 1, the more similar the two areas are.
[0089] It should be noted that Dice Loss can be defined by converting the Dice coefficient into the form of a segmentation task loss function. The calculation formula for the forged area mask loss value is usually as follows:
[0090]
[0091] It should be noted that in practical applications, in order to facilitate calculation and optimization, the intersection and union operations in the above formula are usually converted into pixel-level product and sum operations. In addition, in order to prevent the denominator from being zero, a small constant such as epsilon is usually added to the denominator.
[0092] It should be noted that the specific formula for calculating the domain category loss value using the cross entropy loss function is as follows:
[0093]
[0094] In the formula, yi represents the domain category information, pi represents the predicted domain category information, and N represents the number of training samples.
[0095] It should be noted that there is a gradient reversal layer in front of the domain category classifier. Therefore, when training the model, the gradient returned by the loss of the true and false binary classifier will make the model more and more able to distinguish true and false categories, while the gradient returned by the domain category classifier loss will make the model less and less able to distinguish which domain the image belongs to. This helps the model improve its generalization and perform consistently in different domains.
[0096] In a feasible implementation, the deep fake face detection model also includes a connection module, and the head module includes a separate head module and a comprehensive head module; wherein, the step of inputting the target face image into the deep fake face detection model to obtain a true and false binary classification loss value, a forged area mask loss value, and a domain category loss value includes: inputting the target face image into the deep fake face detection model, and cropping the target face image through the key part cropping module to obtain a key part image; extracting features from the key part image through the feature extraction module to obtain a feature matrix; processing the feature matrix through the separate head module to obtain a first processing result; superimposing the feature matrices output by different feature extraction modules through the connection module to obtain a target feature matrix; processing the target feature matrix through the attention module to obtain an attention matrix; processing the attention matrix through the comprehensive head module to obtain a second output result; determining the true and false binary classification loss value, the forged area mask loss value, and the domain category loss value based on the first output result and / or the second output result.
[0097] It should be noted that the purpose of the key part cropping module is to crop out the key parts of the face image to be identified, such as the face, eyes, nose and lips. Cropping these key parts can help the model focus more on the detailed features of the face, thereby improving the accuracy of detection; the feature extraction module is responsible for extracting useful feature information from the cropped key part images. These features may include texture, shape, color, etc., which are crucial for identifying the authenticity of the face; the attention module is used to identify the most important areas in the image, which may contain key clues to forgery. By focusing on these areas, the model can use image information more effectively and improve the accuracy of detection; the connection module is used to superimpose the feature matrices output by different feature extraction modules to obtain the target feature matrix.
[0098] It should be noted that each individual head module includes a true or false classifier, a mask regressor and a gradient reversal domain category classifier, and the comprehensive head module also includes a true or false classifier, a mask regressor and a gradient reversal domain category classifier; the gradient reversal domain category classifier includes a GRL gradient reversal layer and a domain category classifier connected in sequence.
[0099] In the specific implementation, Figure 3 As shown, the target face image can be located through the key part cropping module, and then the whole body image, face image, double eye image, nose image and lip image are cropped, and then the feature matrices of these key part images are extracted respectively through multiple feature extraction modules (efcieninetv2), and these feature features are respectively input into the separate head module (head) for processing to obtain the first processing result (the first processing result includes three types of results, specifically: separate true and false binary classification information, separate forged area mask information and separate domain category information), and then the feature matrices output by all feature extraction modules are connected into the target feature matrix through the connection module, and then input into the attention module (AttentionModule) to obtain the attention matrix, and finally the attention matrix is input into the comprehensive head module to obtain the second output result (the second output result also includes three types of results, specifically: comprehensive true and false binary classification information, comprehensive forged area mask information and comprehensive domain category information).
[0100] In the specific implementation, the attention matrix is input into a separate head module, the feature matrix is processed by the true-false classifier of the separate head module to obtain separate true-false binary classification information, the feature matrix is processed by the mask regressor of the separate head module to obtain separate forgery mask information, and the feature matrix is processed by the gradient reversal domain category classifier of the separate head module to obtain separate domain category information.
[0101] In the specific implementation, Figure 4 As shown, the attention matrix is input into the comprehensive head module, and the attention matrix is processed by the true and false classifier of the comprehensive head module to obtain comprehensive true and false binary classification information. The attention matrix is processed by the mask regressor of the comprehensive head module to obtain comprehensive forgery mask information. The attention matrix is processed by the gradient reversal domain category classifier of the comprehensive head module to obtain comprehensive domain category information.
[0102] In the specific implementation, Figure 5As shown in the figure, the attention module includes a compression structure (Squeeze), an excitation structure (Excitation) and a scaling structure (Scale). The target feature matrix [c, h, w] is input into the Squeeze-Excitation structure, and an attention weight matrix of [c, 1, 1] is output. Then, the attention weight matrix is multiplied channel by channel with the target feature matrix through the scaling structure (i.e., channel-to-channel scaling) to obtain the final attention matrix. The compression structure compresses the spatial information (i.e., height and width) of each channel into a single value through global average pooling (GAP) to obtain a one-dimensional vector with a dimension of [C, 1, 1], where each element represents the global average of the corresponding channel. The excitation structure performs a nonlinear transformation on the one-dimensional vector through one or more fully connected layers (or dense layers) to learn the correlation between different channels. The ReLU activation function is usually used to introduce nonlinearity, and a sigmoid activation function is used to output the weight coefficient of each channel. The sum of these coefficients is 1, indicating the importance of each channel. The scaling structure can enhance important features and suppress unimportant features, thereby improving the performance of the model.
[0103] It can be understood that it is necessary to determine the separate authenticity loss value, the separate forged area loss value and the separate domain category loss value based on the first output result output by the separate head module (separate authenticity binary classification information, separate forged area mask information and separate domain category information), and then update the model parameters of the model branch of the separate head module based on these loss values; then, it is also necessary to determine the comprehensive authenticity loss value, the comprehensive forged area loss value and the comprehensive domain category loss value based on the second output result output by the comprehensive head module (comprehensive authenticity binary classification information, comprehensive forged area mask information and comprehensive domain category information), and then update the model parameters of the entire model based on these loss values, that is, it is necessary to update the model parameters of the model branch based on the first output result, and update the model parameters of the entire model based on the second output result (they can be updated simultaneously or one after another), which can improve the training efficiency of the model.
[0104] This embodiment collects real face images; generates fake face images based on the real face images, and generates masks for the fake face areas in the fake face images; constructs a face image dataset based on the real face images and the fake face images; based on the face image dataset, updates the model parameters of the deep fake face detection model through the back propagation algorithm to obtain a trained deep fake face detection model. By training with a dataset constructed using real face images and fake face images, the model can learn the feature differences between real and fake faces, thereby improving the generalization ability on unknown data.
[0105] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the face recognition method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0106] This application also provides a face recognition device, please refer to Figure 6 , the face recognition device comprises:
[0107] The input module 10 is used to input the face image to be recognized into the trained deep fake face detection model to obtain the face recognition result, wherein the deep fake face detection model includes a head module, and the head module includes a true and false classifier, a mask regressor and a gradient reversal domain category classifier.
[0108] The face recognition device provided by the present application adopts the face recognition method in the above embodiment, which can solve the technical problem that the existing face recognition system is difficult to detect forged face images. Compared with the prior art, the beneficial effects of the face recognition device provided by the present application are the same as the beneficial effects of the face recognition method provided by the above embodiment, and the other technical features of the face recognition device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0109] The present application provides a face recognition device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the face recognition method in the above-mentioned embodiment 1.
[0110] Reference below Figure 7 , which shows a schematic diagram of the structure of a face recognition device suitable for implementing the embodiment of the present application. The face recognition device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The facial recognition device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0111] like Figure 7As shown, the face recognition device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a ROM (Read Only Memory) 1002 or a program loaded from a storage device 1003 to a RAM (Random Access Memory) 1004. In RAM 1004, various programs and data required for the operation of the face recognition device are also stored. The processing device 1001, ROM 1002, and RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the face recognition device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a face recognition device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or provided instead.
[0112] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0113] The face recognition device provided by the present application adopts the face recognition method in the above embodiment, which can solve the technical problem that the existing face recognition system is difficult to detect forged face images. Compared with the prior art, the beneficial effects of the face recognition device provided by the present application are the same as the beneficial effects of the face recognition method provided by the above embodiment, and the other technical features in the face recognition device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0114] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0115] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0116] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, wherein the computer-readable program instructions are used to execute the face recognition method in the above-mentioned embodiment.
[0117] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0118] The computer-readable storage medium may be included in the face recognition device; or may exist independently without being assembled into the face recognition device.
[0119] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by a face recognition device, the face recognition device: inputs the face image to be recognized into a trained deep fake face detection model to obtain a face recognition result, wherein the deep fake face detection model includes a head module, and the head module includes a true and false classifier, a mask regressor, and a gradient reversal domain category classifier.
[0120] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0121] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0122] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0123] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned face recognition method, and can solve the technical problem that the existing face recognition system is difficult to detect forged face images. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as the beneficial effects of the face recognition method provided by the above-mentioned embodiment, and will not be repeated here.
[0124] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned face recognition method when executed by a processor.
[0125] The computer program product provided by the present application can solve the technical problem that the existing face recognition system is difficult to detect forged face images. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the face recognition method provided by the above embodiment, which will not be repeated here.
[0126] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A face recognition method, characterized in that: The method comprises: The face image to be recognized is input into a trained deep fake face detection model to obtain a face recognition result, wherein the deep fake face detection model includes a head module, and the head module includes a true and false classifier, a mask regressor, and a gradient reversal domain category classifier.
2. The method according to claim 1, characterized in that Before the step of inputting the face image to be recognized into the trained deep fake face detection model to obtain the face recognition result, the step further includes: Collect real face images; Generate a forged face image based on the real face image, and generate a mask for the forged face area in the forged face image; Based on the real face image and the forged face image, construct a face image dataset; Based on the facial image dataset, the model parameters of the deep fake face detection model are updated by a back propagation algorithm to obtain a trained deep fake face detection model.
3. The method according to claim 2, characterized in that The step of updating the model parameters of the deep fake face detection model by a back propagation algorithm based on the face image dataset to obtain a trained deep fake face detection model comprises: Randomly select a target face image from the face image dataset; Input the target face image into the deep fake face detection model to obtain a true and false binary classification loss value, a fake region mask loss value, and a domain category loss value; Based on the loss weight coefficient, determining the overall loss value according to the true-false binary classification loss value, the forged area mask loss value and the domain category loss value; Based on the overall loss value, the model parameters of the deep fake face detection model are updated through a back propagation algorithm to obtain a trained deep fake face detection model.
4. The method according to claim 3, characterized in that The step of inputting the target face image into the deep fake face detection model to obtain the true and false binary classification loss value, the forged area mask loss value and the domain category loss value comprises: Determining authenticity information, mask information, and domain category information of the target face image; Inputting the target face image into the deep fake face detection model to obtain predicted authenticity information, predicted mask information, and predicted domain category information; Based on the cross entropy loss function, determining a true-false binary classification loss value according to the true-false information and the predicted true-false information; Based on the segmentation task loss function, determining a forged area mask loss value according to the mask information and the predicted mask information; Based on a cross entropy loss function, a domain category loss value is determined according to the domain category information and the predicted domain category information.
5. The method according to claim 3, characterized in that The deep fake face detection model further includes a key part cropping module, a feature extraction module, an attention module and a connection module, and the head module includes a separate head module and a comprehensive head module; wherein the step of inputting the target face image into the deep fake face detection model to obtain the true and false binary classification loss value, the forged area mask loss value and the domain category loss value comprises: Inputting the target face image into the deep fake face detection model, and cropping the target face image using the key parts cropping module to obtain a key parts image; Extracting features from the key part image by the feature extraction module to obtain a feature matrix; Processing the feature matrix by the separate head module to obtain a first processing result; The feature matrices output by different feature extraction modules are superimposed by the connection module to obtain a target feature matrix; Processing the target feature matrix through the attention module to obtain an attention matrix; Processing the attention matrix through a comprehensive head module to obtain a second output result; Based on the first output result and / or the second output result, determine a true-false binary classification loss value, a forged area mask loss value, and a domain category loss value.
6. The method according to claim 1, characterized in that The step of inputting the face image to be recognized into the trained deep fake face detection model to obtain the face recognition result comprises: Input the face image to be identified into the trained deep fake face detection model, output the target authenticity probability through the authenticity classifier, and output the target forged area through the mask regressor; When the target authenticity probability is greater than a preset authenticity probability threshold and the target forged area is smaller than a preset area threshold, determining that the face recognition result is a real face; When the target authenticity probability is less than or equal to the preset authenticity probability threshold or the target forgery area is greater than or equal to the preset area threshold, the face recognition result is determined to be a forged attack face.
7. A face recognition device, characterized in that: The device comprises: An input module is used to input a face image to be recognized into a trained deep fake face detection model to obtain a face recognition result, wherein the deep fake face detection model includes a head module, and the head module includes a true and false classifier, a mask regressor, and a gradient reversal domain category classifier.
8. A face recognition device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the face recognition method according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the face recognition method according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the face recognition method according to any one of claims 1 to 6 are implemented.