Training Method, Device, Equipment and Storage Medium of Facial Anti-Spoofing Model
Through self-supervised learning, the training samples are generated by fusion of pseudo-true facial images with similar facial postures, which improves the accuracy and applicability of the facial pseudo-recognition model, and solves the problem that existing models rely on limited label information.
Patent Information
- Application Number
- CN202210804237.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-07-07
AI Technical Summary
The existing facial pseudo-recognition model based on deep learning relies on limited pseudo-face image training with label information, resulting in low accuracy in pseudo-recognition.
By obtaining pseudo-face images and real facial images with similar facial postures, the facial detection network is used to obtain gradient information, and fuse it through the sample generation network to generate fused facial images. The facial pseudo-face model is trained in combination with the pseudo-identification results to achieve self-supervised learning.
It improves the accuracy and generalization of the facial pseudo-identification model, reduces the cost and difficulty of obtaining training samples, enhances the robustness of interference factors, and supports multi-task detection.
Smart Images

Figure CN115188082B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method, device, equipment and storage medium for training a face anti-counterfeiting model. Background Art
[0002] With the development of artificial intelligence technology, its research and application in the field of face (i.e., facial) anti-counterfeiting are also increasing.
[0003] Currently, most of the face anti-counterfeiting models based on deep learning rely on pseudo-face images with label information for supervised training. However, in actual application scenarios, the pseudo-face images with label information are limited, and the anti-counterfeiting accuracy of the face anti-counterfeiting model trained based on the limited pseudo-face images is not high. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, equipment and storage medium for training a face anti-counterfeiting model, which can improve the anti-counterfeiting accuracy of the face anti-counterfeiting model. The technical solution may include the following contents.
[0005] According to one aspect of the embodiments of the present application, a method for training a face anti-counterfeiting model is provided. The face anti-counterfeiting model includes a sample generation network and a face detection network. The method includes:
[0006] Obtain a pseudo-face image and a real-face image; wherein, the difference degree between the face pose corresponding to the pseudo-face image and the face pose corresponding to the real-face image is less than or equal to a threshold value;
[0007] Through the face detection network, obtain the gradient information corresponding to the pseudo-face image, and the gradient information is used to indicate the face area in the pseudo-face image;
[0008] Through the sample generation network, based on the gradient information, fuse the pseudo-face image and the real-face image to obtain a fused face image;
[0009] Through the face detection network, obtain the anti-counterfeiting result corresponding to the fused face image; wherein, the anti-counterfeiting result includes a true / false prediction result, a fusion area prediction result and a fusion degree prediction result;
[0010] Based on the anti-counterfeiting result corresponding to the fused face image, train the face anti-counterfeiting model to obtain the trained face anti-counterfeiting model.
[0011] According to one aspect of the embodiments of the present application, a device for training a face anti-counterfeiting model is provided. The face anti-counterfeiting model includes a sample generation network and a face detection network. The device includes:
[0012] A facial image acquisition module, configured to acquire a pseudo-facial image and a real facial image; wherein, the difference degree between the facial pose corresponding to the pseudo-facial image and the facial pose corresponding to the real facial image is less than or equal to a threshold value;
[0013] A gradient information acquisition module, configured to acquire, through the facial detection network, the gradient information corresponding to the pseudo-facial image, where the gradient information is used to indicate the facial region in the pseudo-facial image;
[0014] A facial image fusion module, configured to fuse the pseudo-facial image and the real facial image based on the gradient information through the sample generation network to obtain a fused facial image;
[0015] A forgery detection result acquisition module, configured to acquire, through the facial detection network, the forgery detection result corresponding to the fused facial image; wherein, the forgery detection result includes a true / false prediction result, a fusion region prediction result, and a fusion degree prediction result;
[0016] A forgery detection model training module, configured to train the facial forgery detection model based on the forgery detection result corresponding to the fused facial image to obtain the trained facial forgery detection model.
[0017] According to one aspect of the embodiments of the present application, there is provided a computer device, where the computer device includes a processor and a memory, and a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned method for training a facial forgery detection model.
[0018] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium, where a computer program is stored in the readable storage medium, and the computer program is loaded and executed by a processor to implement the above-mentioned method for training a facial forgery detection model.
[0019] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method for training a facial forgery detection model.
[0020] The technical solutions provided by the embodiments of the present application at least include the following beneficial effects.
[0021] By randomly obtaining pseudo-face images and real-face images with similar facial postures, and fusing the pseudo-face images and real-face images, while generating the fused face images, determining the label data corresponding to the fused face images, and then based on the fused face images and label data, realizing self-supervised learning of the face anti-spoofing model, so that the training of the face anti-spoofing model is not limited by the limited training samples with labeled information, thereby improving the anti-spoofing accuracy of the face anti-spoofing model. At the same time, since it supports self-supervised learning of the face anti-spoofing model, the acquisition cost and difficulty of the samples required for training the face anti-spoofing model are reduced.
[0022] In addition, by using the fused face images as the training samples of the face anti-spoofing model, the training samples can be augmented near the pseudo-face images, thereby enriching the training samples and improving the diversity of the training samples, and further improving the generalization ability of the face anti-spoofing model and the robustness to interference factors.
[0023] In addition, based on the gradient information corresponding to the pseudo-face images, obtaining the fused face images can improve the accuracy of the fused face images, which is beneficial to further improving the anti-spoofing accuracy of the face anti-spoofing model.
[0024] In addition, the face anti-spoofing model in the embodiments of the present application supports multi-task detection, such as authenticity, fusion area, and fusion degree detection, thereby improving the applicability of the face anti-spoofing model. In addition, based on the prediction results corresponding to multiple complementary tasks, training the face anti-spoofing model is beneficial to further improving the anti-spoofing accuracy of the face anti-spoofing model. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0026] Figure 1 is a schematic diagram of the solution implementation environment provided by an embodiment of the present application;
[0027] Figure 2 is a schematic diagram of the face anti-spoofing model provided by an embodiment of the present application;
[0028] Figure 3 is a flowchart of the training method of the face anti-spoofing model provided by an embodiment of the present application;
[0029] Figure 4 is a flowchart of the method for obtaining gradient information provided by an embodiment of the present application;
[0030] Figure 5 It is a flowchart of a method for obtaining a fused facial image provided by an embodiment of the present application;
[0031] Figure 6 It is a block diagram of a training device for a facial anti-spoofing model provided by an embodiment of the present application;
[0032] Figure 7 It is a block diagram of a training device for a facial anti-spoofing model provided by another embodiment of the present application;
[0033] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0034] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0035] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0036] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0037] Computer Vision (CV) is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for object recognition, measurement, etc., and further performs graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, and map construction, etc.
[0038] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0039] The technical solution provided by the embodiments of this application relates to the computer vision technology and machine learning technology of artificial intelligence. It uses computer vision technology to obtain fused facial images, and authenticate the fused facial images to obtain the authentication results corresponding to the fused facial images. Then, it uses machine learning technology to train a facial authentication model (such as the facial detection network in the facial authentication model) based on the authentication results corresponding to the fused facial images, and obtains a trained facial authentication model.
[0040] For the method provided by the embodiments of this application, the execution subject of each step can be a computer device, which refers to an electronic device with data calculation, processing, and storage capabilities. The computer device can be terminals such as a PC (Personal Computer), tablet computer, smart phone, wearable device, intelligent robot, vehicle-mounted device, etc.; it can also be a server. Among them, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0041] The technical solutions provided in the embodiments of this application are applicable to any scenario that requires face anti-counterfeiting, such as forged face detection scenarios, content security review scenarios, intelligent transportation scenarios, assisted driving scenarios, image analysis scenarios, etc. The technical solutions provided in the embodiments of this application can improve the anti-counterfeiting accuracy of the face anti-counterfeiting model, as well as improve the generalization and robustness of the face anti-counterfeiting model.
[0042] The following will introduce in detail the model structure and training method of the face anti-counterfeiting model provided in the embodiments of this application.
[0043] Please refer to Figure 1 , which shows a schematic diagram of the solution implementation environment provided in an embodiment of this application. The solution implementation environment may include a model training device 10 and a model using device 20.
[0044] The model training device 10 may be an electronic device such as a PC, computer, tablet computer, intelligent robot, vehicle-mounted terminal, etc., or some other electronic device with strong computing power, or it may be a server. The model training device 10 is used to train the face anti-counterfeiting model 30.
[0045] In the embodiments of this application, the face anti-counterfeiting model 30 is a neural network model that can be used to perform face anti-counterfeiting on a face (i.e., the face). Exemplarily, the face anti-counterfeiting model 30 can be used to perform face anti-counterfeiting on objects in images, videos, photos, etc., as well as real objects. For example, for a forged face image obtained by various means, the face anti-counterfeiting model 30 can identify the authenticity, fusion area, and fusion degree of the forged face image. Among them, the various means include, but are not limited to, AI face swapping, fusing faces in multiple images, expression transfer, face modification, etc. The embodiments of this application do not limit the tasks applicable to the face anti-counterfeiting model 30.
[0046] Optionally, the model training device 10 can adopt a machine learning method to train the face anti-counterfeiting model 30 so that it has good anti-counterfeiting performance.
[0047] The above-trained face anti-counterfeiting model 30 can be deployed in the model using device 20 for use to provide face anti-counterfeiting services. The model using device 20 may be a terminal device such as a mobile phone, computer, smart TV, multimedia playback device, wearable device, vehicle-mounted terminal, intelligent robot, etc., or it may be a server. The embodiments of this application do not limit this.
[0048] In some embodiments, as Figure 1 shown, the face anti-counterfeiting model 30 may include a sample generation network 310 and a face detection network 320.
[0049] The sample generation network 310 is used to fuse the pseudo-face image and the real-face image to obtain a fused face image. Optionally, the sample generation network 310 may be constructed based on the Mixup algorithm (a data augmentation algorithm). It may also be constructed based on methods such as Alpha fusion, Poisson fusion, and deep learning fusion. The embodiments of this application do not limit this.
[0050] Among them, there are similar facial postures between the pseudo-face image and the real-face image. Exemplarily, the difference degree between the facial posture corresponding to the pseudo-face image and the facial posture corresponding to the real-face image is less than or equal to a threshold value, and this threshold value can be set and adjusted according to empirical values. For example, a pseudo-face image is randomly obtained, and based on the facial posture of this pseudo-face image, a real-face image with a similar facial posture is retrieved from the real-face database, so as to obtain the training input of the face anti-spoofing model. It is also possible to randomly obtain a real-face image, and based on the facial posture of this real-face image, a pseudo-face image with a similar facial posture is retrieved from the real-face database, so as to obtain the training input of the face anti-spoofing model. The embodiments of this application do not limit this.
[0051] The pseudo-face image refers to the image corresponding to the forged face, and the real-face image refers to the image corresponding to the real face. In the embodiments of this application, the face can be understood as the face, and the facial posture can be understood as the face posture. Since the embodiments of this application do not need to obtain training samples and label data based on prior knowledge, nor do they need to collect corresponding pseudo-face image and real-face image pairs (such as faces for the same object), only pseudo-face images and real-face images with similar facial postures need to be collected, and the acquisition of training samples can be completed in this way. This is beneficial to reducing the difficulty and cost of obtaining training samples. At the same time, it is also beneficial to improve the diversity of the fused face image, thereby improving the generalization and robustness of the face anti-spoofing model.
[0052] Optionally, the facial posture can be characterized by three angles: yaw angle, roll angle, and pitch angle. For example, during the retrieval process, first, the facial posture vector corresponding to the pseudo-face image and the facial posture vector corresponding to the real-face image are constructed respectively with the yaw angle, roll angle, and pitch angle, and then the difference degree between the facial posture vector corresponding to the pseudo-face image and the facial posture vector corresponding to the real-face image is calculated. The real-face image with a difference degree less than or equal to the threshold value is determined as the target real-face image corresponding to the pseudo-face image. Among them, the difference degree can be calculated using algorithms such as cosine similarity, Euclidean distance, Manhattan distance, and Chebyshev distance.
[0053] Optionally, methods based on LBP (Local Binary Pattern) or LTP (Local Ternary Pattern) features or feature point matching can also be used to obtain the facial pose, which is not limited in the embodiments of the present application.
[0054] The face detection network 320 is used to authenticate the face in the face image. In the embodiments of the present application, the face detection network 320 has two tasks. One task is to authenticate the face in the fake face image, and the other task is to authenticate the face in the fused face image. Optionally, the authentication results output by the face detection network 320 include the true / false prediction result, the fused area prediction result, and the fusion degree prediction result. Among them, the true / false prediction result is used to indicate the authenticity of the face in the face image, the fused area prediction result is used to indicate the area belonging to face fusion in the face image, and the fusion degree prediction result is used to indicate the face fusion degree in the face image.
[0055] Exemplarily, in the process of obtaining the fused face image, the first training loss corresponding to the fake face image can be obtained based on the true / false prediction result in the authentication result corresponding to the fake face image. The gradient information corresponding to the fake face image can be obtained by taking the derivative of each pixel value in the fake face image based on the first training loss. The above sample generation network 310 fuses the fake face image and the real face image based on the gradient information corresponding to the fake face image, and then the fused face image can be obtained. Among them, the first training loss is used to indicate the prediction accuracy of the face authentication model for the authenticity of the face image, and the gradient information can reflect the important regions and features of the face in the fake face image.
[0056] In the process of authenticating the fused face image, the authentication result corresponding to the fused face image can be obtained by processing the fused face image through the face detection network 320.
[0057] Optionally, the face detection network 320 can be constructed based on the Xception network structure (a depthwise separable convolutional network architecture), or can also be constructed based on ResNet (Residual Network), which is not limited in the embodiments of the present application.
[0058] In the process of obtaining the fused face image, the label data corresponding to the fused face image can be obtained, such as the true / false real result, the fused area real result, and the fusion degree real result corresponding to the fused face image. Based on the label data corresponding to the fused face image and the authentication result, the training loss of the face authentication model can be obtained, and then the face authentication model can be trained based on the training loss of the face authentication model to obtain the trained face authentication model.
[0059] In an exemplary embodiment, with reference to Figure 2 , first, through the face detection network 206, the gradient of the fake face image 202 is calculated to obtain the gradient information 203 corresponding to the fake face image 202. Then, based on the gradient information 203 and a random fusion degree, the sample generation network 204 fuses the real face image 201 and the fake face image 202 to obtain a fused face image 205 (such as fusing regions like the mouth, eyes, nose, etc.). Subsequently, the face detection network 206 performs target forgery detection on the fused face image 205 to obtain the forgery detection result corresponding to the fused face image 205. Finally, based on the forgery detection result corresponding to the fused face image 205 and the label information, the face forgery detection model is trained to obtain a trained face forgery detection model.
[0060] The above introduced the model architecture of the face forgery detection model. The following will elaborate on the training method of the face forgery detection model.
[0061] Please refer to Figure 3 , which shows a flowchart of the training method of the face forgery detection model provided by an embodiment of the present application. The execution subject of each step of this method can be the model training device introduced above. This method may include the following steps (301 - 305).
[0062] Step 301, obtain a fake face image and a real face image; wherein, the difference degree between the face pose corresponding to the fake face image and the face pose corresponding to the real face image is less than or equal to a threshold value.
[0063] A fake face image refers to an image corresponding to a forged face, and a real face image refers to an image corresponding to a real face. A face image may refer to an image, video frame, photo, etc. including the face of one or more target objects. Optionally, a fake face image may refer to a face image synthesized by various means, including but not limited to AI face swapping, fusing faces in multiple images, expression transfer, face modification, etc. In the embodiments of the present application, a face can be understood as a face, and a face pose can be understood as a face pose.
[0064] Optionally, a fake face image may refer to any fake face image in the fake face image library, and a real face image may refer to a real face image in the real face image library whose difference degree between any face pose and the face pose corresponding to the fake face image is less than or equal to the threshold value. The acquisition methods of the fake face image and the real face image are the same as those introduced in the above embodiments and will not be elaborated here.
[0065] Step 302, through the face detection network, obtain the gradient information corresponding to the fake face image, and this gradient information is used to indicate the face region in the fake face image.
[0066] In the embodiments of the present application, the gradient information corresponding to the pseudo-face image can reflect the important regions and features of the face in the pseudo-face image, and can be used to guide the generation of the fused face image.
[0067] In one example, the gradient information corresponding to the pseudo-face image can be obtained through a face detection network. Refer to Figure 4 , and the specific process may include the following steps.
[0068] Step 302a: Obtain the anti-counterfeiting result corresponding to the pseudo-face image through the face detection network.
[0069] The sample generation network and the face detection network in the embodiments of the present application are the same as those introduced in the above embodiments, and will not be elaborated here.
[0070] Optionally, a series of processes such as convolution, feature map acquisition, and result calculation can be performed on the pseudo-face image through the face detection network to obtain the anti-counterfeiting result corresponding to the pseudo-face image. For example, refer to Figure 2 , and the authenticity prediction result corresponding to the pseudo-face image is obtained by processing the feature map corresponding to the pseudo-face image through the authenticity detector in the face detection network.
[0071] Step 302b: Based on the anti-counterfeiting result corresponding to the pseudo-face image and the label data corresponding to the pseudo-face image, obtain the first training loss corresponding to the pseudo-face image, where the first training loss is used to indicate the prediction accuracy of the face anti-counterfeiting model for the authenticity of the face image.
[0072] Optionally, the cross-entropy loss function can be used to obtain the first training loss corresponding to the pseudo-face image based on the authenticity prediction result corresponding to the pseudo-face image and the true authenticity result corresponding to the pseudo-face image. The first training loss can be expressed as follows:
[0073] L o =-∑ i q i logp i ;
[0074] where L o is the first training loss, p i is the true authenticity result corresponding to the pseudo-face image (i.e., the label data), q i is the authenticity prediction result corresponding to the pseudo-face image, and i is the i-th pseudo-face image.
[0075] Step 302c: Based on the first training loss corresponding to the pseudo-face image, take the derivative of each pixel value corresponding to the pseudo-face image to obtain the gradient information corresponding to the pseudo-face image.
[0076] Optionally, the pixel values corresponding to each pixel point in the pseudo-face image can be obtained first, and then the gradient information corresponding to the pseudo-face image can be obtained by taking the derivative of the pixel values corresponding to each pixel point based on the first training loss corresponding to the pseudo-face image. This process can be expressed as follows:
[0077] where x is the input pseudo-face image.
[0078] By using the first training loss to obtain the gradient information of the pseudo-face image and then obtaining the fused face image based on the gradient information, the correlation between the fused face image and the face anti-spoofing model can be improved, which is beneficial to further improving the anti-spoofing accuracy of the face anti-spoofing model.
[0079] Step 303, through the sample generation network, fuse the pseudo-face image and the real face image based on the gradient information to obtain a fused face image.
[0080] In the embodiments of the present application, the fused face image may refer to the pseudo-face image obtained by fusing the pseudo-face image into the real face image. Among them, the fusion area and the fusion intensity can be randomly generated. The fusion area is used to indicate which area in the real face image needs to be fused, such as areas like the mouth, eyes, nose, etc. The fusion intensity is used to indicate the proportion of the pseudo-face image and the real face image. For example, for a certain pixel point on the real face image, the greater the fusion intensity, the greater the proportion of the pseudo-face image. Optionally, before fusion, the facial area in the pseudo-face image can be scaled according to the facial area of the real face image.
[0081] In one example, referring to Figure 5 , step 303 may further include the following sub-steps.
[0082] Step 303a, adjust the gradient information to obtain a mask corresponding to the pseudo-face image.
[0083] Optionally, the gradient information can be converted into a binary image (i.e., a mask) through a threshold. This process can be expressed as follows:
[0084] g′ = 0, if g < γ; g′ = 1, if g ≥ γ; where g is the gradient value image (i.e., gradient information), and γ is the threshold, which can be adaptively set and adjusted according to empirical values.
[0085] Exemplarily, taking γ as 0.6 as an example, if the gradient value corresponding to a certain pixel point is less than 0.6, the mask value corresponding to this pixel point is set to 0. If the gradient value corresponding to a certain pixel point is greater than or equal to 0.6, the mask value corresponding to this pixel point is set to 1. In this way, the facial area in the pseudo-facial image can be made more obvious and clear. By using gradient information or masks to guide the generation of the fused facial image, compared with randomly or using a fixed method to generate the fused facial image, the fusion accuracy is higher and the fusion effect is better, which is conducive to improving the authenticity verification accuracy of the facial anti-spoofing model.
[0086] Optionally, the gradient information can also be converted into a non-binary image, such as a ternary image, a quaternary image, etc., and the embodiments of the present application do not limit this.
[0087] Step 303b, obtain the target fusion area and the target fusion intensity corresponding to the real facial image.
[0088] Optionally, during the acquisition process of a certain fused facial image, the target fusion area and the target fusion intensity are randomly generated. The target fusion area can refer to the corresponding part or all of the area of the face in the real face image, and the target fusion intensity can be any fusion intensity. For example, any value between 0% and 100% can be determined as the fusion intensity value of the pseudo-facial image.
[0089] The target fusion area can be determined as the real result of the fusion area corresponding to this fused facial image, and the target fusion intensity can be determined as the real result of the fusion degree corresponding to this fused facial image, so as to obtain the label data corresponding to this fused facial image. In addition, the authenticity real result of the fused facial image is defaulted to be fake (represented by 0 for example).
[0090] Step 303c, through the sample generation network, in the target fusion area, fuse the pseudo-facial image and the real facial image according to the target fusion intensity and the mask to obtain the fused facial image.
[0091] In one example, the acquisition process of the fused facial image can be as follows: obtain the difference image between the pseudo-facial image and the real facial image; through the sample generation network, multiply the target fusion intensity, the mask and the difference image to obtain the transition image; through the sample generation network, in the target fusion area, fuse the transition image and the real facial image to obtain the fused facial image.
[0092] Exemplarily, this process can be expressed by the following formula:
[0093] I a =β*g′*(I f -I r )+I r ;
[0094] Among them, Ia For the fused face image, I f For the pseudo-face image, I r For the real face image, β is the target fusion intensity, and g′ is the mask corresponding to the pseudo-face image. For each pixel point in the target fusion region, a formula is used to calculate the new pixel value, thereby obtaining the fused face image.
[0095] Step 304, obtain the anti-counterfeiting result corresponding to the fused face image through the face detection network; wherein, the anti-counterfeiting result includes the true / false prediction result, the fusion region prediction result, and the fusion degree prediction result.
[0096] Optionally, a series of processes such as convolution, feature map acquisition, and result calculation can be performed on the fused face image through the face detection network to obtain the anti-counterfeiting result corresponding to the fused face image. Among them, the true / false prediction result corresponding to the fused face image is used to indicate the authenticity of the face in the fused face image, the fusion region prediction result corresponding to the fused face image is used to indicate the region belonging to face fusion in the fused face image, and the fusion degree prediction result corresponding to the fused face image is used to indicate the face fusion degree in the fused face image.
[0097] For example, referring to Figure 2 , the true / false detector in the face detection network 206 can be used to process the feature map corresponding to the fused face image to obtain the true / false prediction result corresponding to the fused face image.
[0098] The fusion region detector in the face detection network 206 can be used to process the feature map corresponding to the fused face image to obtain the fusion region prediction result corresponding to the fused face image.
[0099] The fusion intensity detector in the face detection network 206 can be used to process the feature map corresponding to the fused face image to obtain the fusion intensity prediction result corresponding to the fused face image. In this way, the embodiments of the present application can achieve multi-task detection, thereby improving the applicability of the face anti-counterfeiting model.
[0100] Step 305, train the face anti-counterfeiting model based on the anti-counterfeiting result corresponding to the fused face image to obtain the trained face anti-counterfeiting model.
[0101] Optionally, based on the difference between the true / false prediction result corresponding to the fused face image and the true / false ground truth result corresponding to the fused face image, the first training loss corresponding to the fused face image can be obtained. The method for obtaining the first training loss corresponding to the fused face image is the same as that introduced in the above embodiments and will not be elaborated here.
[0102] Obtain the second training loss corresponding to the fused face image based on the difference between the predicted result of the fusion region corresponding to the fused face image and the true result of the fusion region corresponding to the fused face image.
[0103] Exemplarily, L1 regularization can be adopted to obtain the second training loss corresponding to the fused face image, and this process can be expressed as follows:
[0104] L r = ||g′ - M P ||1;
[0105] where M p is the mask corresponding to the predicted result of the fusion region (i.e., the above-mentioned target fusion region), g′ is the mask corresponding to the pseudo-face image, and the smaller the difference between the two, the better. For the real face image, g′ = 0, that is, each mask value in the binary image corresponding to the real face image is 0.
[0106] Obtain the third training loss corresponding to the fused face image based on the difference between the predicted result of the fusion degree corresponding to the fused face image and the true result of the fusion degree corresponding to the fused face image.
[0107] Exemplarily, L2 regularization can be adopted to obtain the third training loss corresponding to the fused face image, and this process can be expressed as follows:
[0108]
[0109] where is the predicted result of the fusion degree corresponding to the fused face image, and α is the true result of the fusion degree (i.e., the above-mentioned target fusion degree).
[0110] Train the face anti-spoofing model based on the first training loss, the second training loss, and the third training loss corresponding to the fused face image to obtain a trained face anti-spoofing model.
[0111] Optionally, the second training loss and the third training loss corresponding to the fused face image can be weighted and summed to obtain the intermediate training loss corresponding to the fused face image; then, the first training loss and the intermediate training loss corresponding to the fused face image are summed to obtain the training loss of the face anti-spoofing model; with the goal of minimizing the training loss of the face anti-spoofing model, train the face anti-spoofing model to obtain a trained face anti-spoofing model.
[0112] Exemplarily, this process can be expressed as follows:
[0113] L = L o + λ1L r + λ2L t ;
[0114] Among them, λ1 and λ2 are hyperparameters used to adjust the weights corresponding to the second loss and the third loss respectively.
[0115] With the goal of minimizing L, the face anti-spoofing model is iteratively trained to obtain a trained face anti-spoofing model. Optionally, with the goal of minimizing L, only the face detection network in the face anti-spoofing model can be iteratively trained to obtain a trained face anti-spoofing model. Based on the training losses corresponding to the three complementary tasks, the face anti-spoofing model is iteratively trained, further improving the anti-spoofing accuracy of the face anti-spoofing model.
[0116] In one example, during the use of the trained face anti-spoofing model, the sample generation network in the face anti-spoofing model can be removed, and the face image can be directly input into the face detection network to obtain the authenticity prediction result, fusion area prediction result, and fusion degree prediction result corresponding to the face image.
[0117] Optionally, a face detection model can be reconstructed based on the face anti-spoofing network in the face anti-spoofing model to authenticate the face image.
[0118] In summary, the technical solution provided by the embodiments of the present application randomly obtains fake face images and real face images with similar face postures, fuses the fake face images and real face images, determines the label data corresponding to the fused face image while generating the fused face image, and then based on the fused face image and label data, realizes the self-supervised learning of the face anti-spoofing model, so that the training of the face anti-spoofing model is not limited by limited training samples with labeled information, thereby improving the anti-spoofing accuracy of the face anti-spoofing model. At the same time, since it supports the self-supervised learning of the face anti-spoofing model, the acquisition cost and difficulty of the samples required for training the face anti-spoofing model are reduced.
[0119] In addition, by using the fused face image as the training sample of the face anti-spoofing model, the training samples can be expanded near the fake face image, thereby enriching the training samples and improving the diversity of the training samples, and further improving the generalization ability of the face anti-spoofing model and the robustness to interference factors.
[0120] In addition, based on the gradient information corresponding to the fake face image, the fused face image can be obtained, which can improve the accuracy of the fused face image and is beneficial to further improving the anti-spoofing accuracy of the face anti-spoofing model.
[0121] In addition, the face anti-spoofing model in the embodiments of the present application supports multi-task detection, such as authenticity, fusion area, and fusion degree detection, thereby improving the applicability of the face anti-spoofing model. In addition, training the face anti-spoofing model based on the prediction results corresponding to multiple complementary tasks is beneficial to further improve the anti-spoofing accuracy of the face anti-spoofing model.
[0122] In addition, by using the first training loss to obtain the gradient information of the fake face image, and then obtaining the fused face image based on the gradient information, the correlation between the fused face image and the face anti-spoofing model can be improved, which is beneficial to further improve the anti-spoofing accuracy of the face anti-spoofing model.
[0123] The following are the device embodiments of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0124] Please refer to Figure 6 , which shows a block diagram of a training device for a face anti-spoofing model provided by an embodiment of the present application. This device can be used to implement the above-mentioned training method for the face anti-spoofing model. The device 600 may include: a face image acquisition module 601, a gradient information acquisition module 602, a face image fusion module 603, an anti-spoofing result acquisition module 604, and an anti-spoofing model training module 605.
[0125] The face image acquisition module 601 is configured to acquire a fake face image and a real face image; wherein, the difference degree between the face pose corresponding to the fake face image and the face pose corresponding to the real face image is less than or equal to a threshold value.
[0126] The gradient information acquisition module 602 is configured to obtain the gradient information corresponding to the fake face image through the face detection network, and the gradient information is used to indicate the face area in the fake face image.
[0127] The face image fusion module 603 is configured to fuse the fake face image and the real face image based on the gradient information through the sample generation network to obtain a fused face image.
[0128] The anti-spoofing result acquisition module 604 is configured to obtain the anti-spoofing result corresponding to the fused face image through the face detection network; wherein, the anti-spoofing result includes an authenticity prediction result, a fusion area prediction result, and a fusion degree prediction result.
[0129] The anti-spoofing model training module 605 is configured to train the face anti-spoofing model based on the anti-spoofing result corresponding to the fused face image to obtain the trained face anti-spoofing model.
[0130] In an exemplary embodiment, the gradient information acquisition module 602 is configured to:
[0131] Obtain an anti-counterfeiting result corresponding to the pseudo-face image through the face detection network;
[0132] Based on the anti-counterfeiting result corresponding to the pseudo-face image and the label data corresponding to the pseudo-face image, obtain a first training loss corresponding to the pseudo-face image, where the first training loss is used to indicate the prediction accuracy of the face anti-counterfeiting model for the authenticity of the face image;
[0133] Derive each pixel value corresponding to the pseudo-face image based on the first training loss corresponding to the pseudo-face image to obtain gradient information corresponding to the pseudo-face image.
[0134] In an exemplary embodiment, as Figure 7 shown, the face image fusion module 603 further includes: a gradient information adjustment sub-module 603a, a fusion information acquisition sub-module 603b, and a face image fusion sub-module 603c.
[0135] The gradient information adjustment sub-module 603a is configured to adjust the gradient information to obtain a mask corresponding to the pseudo-face image.
[0136] The fusion information acquisition sub-module 603b is configured to obtain a target fusion region and a target fusion intensity corresponding to the real face image.
[0137] The face image fusion sub-module 603c is configured to fuse the pseudo-face image and the real face image in the target fusion region according to the target fusion intensity and the mask through the sample generation network to obtain the fused face image.
[0138] In an exemplary embodiment, the face image fusion sub-module 603c is configured to:
[0139] Obtain a difference image between the pseudo-face image and the real face image;
[0140] Multiply the target fusion intensity, the mask, and the difference image through the sample generation network to obtain a transition image;
[0141] Fuse the transition image and the real face image in the target fusion region through the sample generation network to obtain the fused face image.
[0142] In an exemplary embodiment, the anti-counterfeiting model training module 605 is configured to:
[0143] Obtain the first training loss corresponding to the fused face image based on the difference between the authenticity prediction result corresponding to the fused face image and the true authenticity result corresponding to the fused face image;
[0144] Obtain the second training loss corresponding to the fused face image based on the difference between the fusion area prediction result corresponding to the fused face image and the true fusion area result corresponding to the fused face image;
[0145] Obtain the third training loss corresponding to the fused face image based on the difference between the fusion degree prediction result corresponding to the fused face image and the true fusion degree result corresponding to the fused face image;
[0146] Train the face anti-spoofing model based on the first training loss, the second training loss, and the third training loss corresponding to the fused face image to obtain the trained face anti-spoofing model.
[0147] In an exemplary embodiment, the anti-spoofing model training module 605 is further configured to:
[0148] Perform a weighted sum of the second training loss and the third training loss corresponding to the fused face image to obtain the intermediate training loss corresponding to the fused face image;
[0149] Sum the first training loss and the intermediate training loss corresponding to the fused face image to obtain the training loss of the face anti-spoofing model;
[0150] Train the face anti-spoofing model with the goal of minimizing the training loss of the face anti-spoofing model to obtain the trained face anti-spoofing model.
[0151] In summary, the technical solution provided by the embodiments of the present application randomly obtains fake face images and real face images with similar face postures, and fuses the fake face images and the real face images. While generating the fused face image, the label data corresponding to the fused face image is determined. Furthermore, based on the fused face image and the label data, self-supervised learning of the face anti-spoofing model is realized, so that the training of the face anti-spoofing model is not limited by limited training samples with labeled information, thereby improving the anti-spoofing accuracy of the face anti-spoofing model. At the same time, since self-supervised learning of the face anti-spoofing model is supported, the acquisition cost and difficulty of the samples required for training the face anti-spoofing model are reduced.
[0152] In addition, by using the fused face image as a training sample of the face anti-spoofing model, the training samples can be expanded near the fake face image, thereby enriching the training samples and improving the diversity of the training samples, and further improving the generalization ability of the face anti-spoofing model and the robustness to interference factors.
[0153] In addition, based on the gradient information corresponding to the pseudo-face image, a fused face image can be obtained, which can improve the accuracy of the fused face image and is beneficial to further improving the authenticity verification accuracy of the face anti-spoofing model.
[0154] In addition, the face anti-spoofing model in the embodiments of the present application supports multi-task detection, such as authenticity, fusion area, and fusion degree detection, thereby improving the applicability of the face anti-spoofing model. In addition, based on the prediction results corresponding to multiple complementary tasks, training the face anti-spoofing model is beneficial to further improving the authenticity verification accuracy of the face anti-spoofing model.
[0155] It should be noted that when the device provided in the above embodiments realizes its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be elaborated here.
[0156] Please refer to Figure 8 , which shows a schematic structural diagram of a computer device provided in an embodiment of the present application. The computer device can be any electronic device with data calculation, processing, and storage functions, and the computer device can be implemented as Figure 1 the model training device 10 and / or the model using device 20 in the implementation environment of the solution shown. Specifically, it can include the following contents.
[0157] The computer device 800 includes a central processing unit (such as a CPU (Central Processing Unit, central processor), a GPU (Graphics Processing Unit, graphics processor), and an FPGA (Field Programmable Gate Array, field programmable logic gate array), etc.) 801, a system memory 804 including a RAM (Random-Access Memory, random access memory) 802 and a ROM (Read-Only Memory, read-only memory) 803, and a system bus 805 connecting the system memory 804 and the central processing unit 801. The computer device 800 also includes a basic input / output system (Input Output System, I / O system) 806 for helping to transmit information between various devices in the server, and a mass storage device 807 for storing an operating system 813, an application program 814, and other program modules 815.
[0158] In some embodiments, the basic input / output system 806 includes a display 808 for displaying information and input devices 809 such as a mouse, keyboard, etc. for user input of information. Among them, both the display 808 and the input devices 809 are connected to the central processing unit 801 through an input / output controller 810 connected to the system bus 805. The basic input / output system 806 may also include an input / output controller 810 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 810 also provides outputs to a display screen, printer, or other types of output devices.
[0159] The mass storage device 807 is connected to the central processing unit 801 through a mass storage controller (not shown) connected to the system bus 805. The mass storage device 807 and its associated computer-readable medium provide non-volatile storage for the computer device 800. That is to say, the mass storage device 807 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0160] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc), or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art know that the computer storage media is not limited to the above several. The above system memory 804 and mass storage device 807 may be collectively referred to as memory.
[0161] According to an embodiment of the present application, the computer device 800 can also run on a remote computer on the network through a network such as the Internet. That is, the computer device 800 can be connected to the network 812 through the network interface unit 811 connected to the system bus 805. Or rather, the network interface unit 811 can also be used to connect to other types of networks or remote computer systems (not shown).
[0162] The memory further includes a computer program, which is stored in the memory and is configured to be executed by one or more processors to implement the above-mentioned method for training the face anti-spoofing model.
[0163] In an exemplary embodiment, a computer-readable storage medium is also provided. A computer program is stored in the storage medium, and when the computer program is executed by a processor, the above-mentioned method for training the face anti-spoofing model is implemented.
[0164] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0165] In an exemplary embodiment, a computer program product or a computer program is also provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method for training the face anti-spoofing model.
[0166] It should be noted that the information (including but not limited to object device information, object personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the object or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the fake face images, real face images, etc. involved in the present application are all obtained under full authorization.
[0167] It should be understood that the "plurality" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In addition, the step numbers described herein only exemplarily show a possible execution sequence between steps. In some other embodiments, the above steps may not be executed in the numbered order. For example, two steps with different numbers can be executed simultaneously, or two steps with different numbers can be executed in the reverse order of the illustration. The embodiments of the present application do not make any limitations in this regard.
[0168] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A training method for a face anti-spoofing model, characterized in that The facial anti-counterfeiting model includes a sample generation network and a facial detection network, and the method includes: Obtain a fake facial image and a real facial image; wherein, the difference degree between the facial pose corresponding to the fake facial image and the facial pose corresponding to the real facial image is less than or equal to a threshold value; Through the facial detection network, obtain the anti-counterfeiting result corresponding to the fake facial image; Based on the anti-counterfeiting result corresponding to the fake facial image and the label data corresponding to the fake facial image, obtain the first training loss corresponding to the fake facial image, and the first training loss is used to indicate the prediction accuracy of the facial anti-counterfeiting model for the authenticity of the facial image; Based on the first training loss corresponding to the fake facial image, take the derivative of each pixel value corresponding to the fake facial image to obtain the gradient information corresponding to the fake facial image, and the gradient information is used to indicate the facial area in the fake facial image; Through the sample generation network, fuse the fake facial image and the real facial image based on the gradient information to obtain a fused facial image; Through the facial detection network, obtain the anti-counterfeiting result corresponding to the fused facial image; wherein, the anti-counterfeiting result includes a authenticity prediction result, a fusion area prediction result, and a fusion degree prediction result; Based on the anti-counterfeiting result corresponding to the fused facial image, train the facial anti-counterfeiting model to obtain the trained facial anti-counterfeiting model.
2. The method according to claim 1, wherein The step of fusing the fake facial image and the real facial image based on the gradient information through the sample generation network to obtain the fused facial image includes: Adjust the gradient information to obtain a mask corresponding to the fake facial image; Obtain the target fusion area and the target fusion intensity corresponding to the real facial image; Through the sample generation network, fuse the fake facial image and the real facial image in the target fusion area according to the target fusion intensity and the mask to obtain the fused facial image.
3. The method according to claim 2, wherein The step of fusing the fake facial image and the real facial image in the target fusion area according to the target fusion intensity and the mask through the sample generation network to obtain the fused facial image includes: Obtain the difference image between the fake facial image and the real facial image; Through the sample generation network, multiply the target fusion intensity, the mask, and the difference image to obtain a transition image; Through the sample generation network, fuse the transition image and the real facial image in the target fusion area to obtain the fused facial image.
4. The method according to claim 1, wherein The step of training the facial anti-counterfeiting model based on the anti-counterfeiting result corresponding to the fused facial image to obtain the trained facial anti-counterfeiting model includes: Based on the difference between the authenticity prediction result corresponding to the fused facial image and the true authenticity result corresponding to the fused facial image, obtain the first training loss corresponding to the fused facial image; Obtain the second training loss corresponding to the fused face image based on the difference between the predicted result of the fusion region corresponding to the fused face image and the true result of the fusion region corresponding to the fused face image; Obtain the third training loss corresponding to the fused face image based on the difference between the predicted result of the fusion degree corresponding to the fused face image and the true result of the fusion degree corresponding to the fused face image; Train the face anti-spoofing model based on the first training loss, the second training loss, and the third training loss corresponding to the fused face image to obtain the trained face anti-spoofing model.
5. The method according to claim 4, wherein The training of the face anti-spoofing model based on the first training loss, the second training loss, and the third training loss corresponding to the fused face image to obtain the trained face anti-spoofing model includes: Perform weighted summation on the second training loss and the third training loss corresponding to the fused face image to obtain the intermediate training loss corresponding to the fused face image; Sum the first training loss and the intermediate training loss corresponding to the fused face image to obtain the training loss of the face anti-spoofing model; Train the face anti-spoofing model with the goal of minimizing the training loss of the face anti-spoofing model to obtain the trained face anti-spoofing model.
6. A training device for a face anti-spoofing model, characterized in that The face anti-spoofing model includes a sample generation network and a face detection network, and the device includes: A face image acquisition module, configured to acquire a fake face image and a real face image; wherein, the difference degree between the face pose corresponding to the fake face image and the face pose corresponding to the real face image is less than or equal to a threshold value; A gradient information acquisition module, configured to obtain the anti-spoofing result corresponding to the fake face image through the face detection network; based on the anti-spoofing result corresponding to the fake face image and the label data corresponding to the fake face image, obtain the first training loss corresponding to the fake face image, and the first training loss is used to indicate the prediction accuracy of the face anti-spoofing model for the authenticity of the face image; based on the first training loss corresponding to the fake face image, take the derivative of each pixel value corresponding to the fake face image to obtain the gradient information corresponding to the fake face image, and the gradient information is used to indicate the face region in the fake face image; A face image fusion module, configured to fuse the fake face image and the real face image based on the gradient information through the sample generation network to obtain a fused face image; An anti-spoofing result acquisition module, configured to obtain the anti-spoofing result corresponding to the fused face image through the face detection network; wherein, the anti-spoofing result includes a authenticity prediction result, a fusion region prediction result, and a fusion degree prediction result; An anti-spoofing model training module, configured to train the face anti-spoofing model based on the anti-spoofing result corresponding to the fused face image to obtain the trained face anti-spoofing model.
7. A computer device, characterized in that, The computer device includes a processor and a memory, and a computer program is stored in the memory. The computer program is loaded and executed by the processor to implement the method for training a face anti-spoofing model according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. The computer program is loaded and executed by a processor to implement the method for training a face anti-spoofing model according to any one of claims 1 to 5.
9. A computer program product, characterized in that, The computer program product includes computer instructions. The computer instructions are stored in a computer-readable storage medium, and a processor reads and executes the computer instructions to implement the method for training a face anti-spoofing model according to any one of claims 1 to 5.
Citation Information
Patent Citations
False face detection method and device, electronic equipment and storage medium
CN111753729A
Facial image forgery detection
CN113128271A