Liveness detection model training method, liveness detection method and related device

Through the multi-grained feature extraction model and combined with the feature extraction network of different categories, the problem of low recognition accuracy in the existing live detection methods is solved, effective recognition of fake faces is achieved, and the security of the face recognition system is improved.

CN114333019BActive Publication Date: 2025-08-19HUNDSUN TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111647626.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-08-19
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

The existing live detection methods cannot effectively extract edge information of images, resulting in low recognition accuracy, especially powerless against fake face attack methods.

Method used

A multi-grained feature extraction model is adopted, including a feature extraction network with the first type of biased texture features, a feature extraction network with the second type of biased edge features, and a third type of feature extraction network with the third type of unbiased feature extraction model. The live detection model is constructed through training sample sets, and the feature fusion network is combined to achieve the extraction and classification of multi-grained feature.

Benefits of technology

It improves the recognition accuracy and detection accuracy of the live detection model, can effectively distinguish between real faces and fake faces, and enhances the security of the face recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333019B_ABST
    Figure CN114333019B_ABST
Patent Text Reader

Abstract

The present invention provides a liveness detection model training method, liveness detection method, and related devices, comprising: obtaining a face training sample set; the face training sample set includes multiple face training samples; the face training samples are marked with an identifier indicating whether the face training samples are positive samples or negative samples; constructing an initial liveness detection model based on a constructed initial feature fusion model and a trained feature extraction model; and training the model parameters of the initial liveness detection model based on the face training sample set to obtain a liveness detection model. The present invention constructs a liveness detection model using a trained feature extraction model, which can solve the problems of single feature extraction and the inability to extract or obtain edge features in the prior art, thereby improving the recognition accuracy and detection precision of the trained liveness detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a training method for a liveness detection model, a liveness detection method, and related devices. Background Art

[0002] Facial recognition is now widely used in various fields. To ensure the high security of facial recognition systems, liveness detection has become indispensable. The purpose of liveness detection is to determine whether a detected face is real or forged. Only when it is determined to be real will it proceed to the subsequent facial recognition process. Due to the diverse forms of forgeries, high-precision liveness detection is a very challenging task.

[0003] Existing liveness detection methods are based on images and videos. Regardless of whether they are video-based or image-based methods, the liveness detection models used mainly extract texture features of the image, but are unable to obtain images of certain attack methods, including effective edge information (such as moiré patterns on electronic screens and facial features on masks) and other attack features, resulting in incorrect recognition and low recognition accuracy. Summary of the Invention

[0004] One of the objectives of the present invention is to provide a method for training a liveness detection model, a liveness detection method, and related devices to solve the above-mentioned technical problems. The embodiments of the present invention can be implemented as follows:

[0005] In a first aspect, the present invention provides a training method for a liveness detection model, the method comprising: obtaining a face training sample set; the face training sample set comprising multiple face training samples; the face training samples being marked with an identifier indicating whether the face training samples are positive samples or negative samples; constructing an initial liveness detection model based on a constructed initial feature fusion model and a trained feature extraction model; wherein the feature extraction model comprises at least a first type of feature extraction network, a second type of feature extraction network, and a third type of feature extraction network; the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second type of feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third type of feature extraction network extracting texture features is equal to the probability of extracting edge features; and training the model parameters of the initial liveness detection model based on the face training sample set to obtain a liveness detection model.

[0006] In a second aspect, the present invention provides a liveness detection method, which includes: obtaining a facial image of a face to be detected; inputting the facial image into a trained liveness detection model, and outputting a detection result; wherein the feature extraction model in the liveness detection model includes at least a first type of feature extraction network, a second type of feature extraction network, and a third type of feature extraction network; the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second type of feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third type of feature extraction network extracting texture features is equal to the probability of extracting edge features.

[0007] In a third aspect, the present invention provides a training device for a liveness detection model, comprising: an acquisition module for acquiring a face training sample set; the face training sample set comprises multiple face training samples; the face training sample is marked with an identifier indicating whether the training sample is a positive sample or a negative sample; a construction module for constructing an initial liveness detection model based on the constructed initial feature fusion model and the trained feature extraction model; wherein the feature extraction model comprises at least a first type of feature extraction network, a second type of feature extraction network and a third type of feature extraction network; the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second type of feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third type of feature extraction network extracting texture features is equal to the probability of extracting edge features; a training module for training the model parameters of the initial liveness detection model based on the face training sample set to obtain a liveness detection model.

[0008] In a fourth aspect, the present invention provides a liveness detection device, comprising: an acquisition module for acquiring a facial image of a face to be detected; a detection module for inputting the facial image into a trained liveness detection model and outputting a detection result; wherein the feature extraction model in the liveness detection model includes at least a first type of feature extraction network, a second type of feature extraction network and a third type of feature extraction network; the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second type of feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third type of feature extraction network extracting texture features is equal to the probability of extracting edge features.

[0009] In a fifth aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the method described in the first aspect and / or implement the method described in the second aspect.

[0010] In a sixth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect and / or implements the method described in the second aspect.

[0011] The present invention provides a training method for a liveness detection model, a liveness detection method, and related devices, the method comprising: obtaining a face training sample set; the face training sample set comprising multiple face training samples; the face training samples being marked with an identifier indicating whether the face training samples are positive samples or negative samples; constructing an initial liveness detection model based on a constructed initial feature fusion model and a trained feature extraction model; wherein the feature extraction model comprises at least a first type of feature extraction network, a second type of feature extraction network, and a third type of feature extraction network; the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second type of feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third type of feature extraction network extracting texture features is equal to the probability of extracting edge features; and training model parameters of the initial liveness detection model based on the face training sample set to obtain a liveness detection model. The present invention constructs a liveness detection model through a trained feature extraction model. Since the trained feature extraction model can extract multi-granularity texture features, that is, through different categories of feature extraction networks, it can extract texture-biased features, edge-biased features, and unbiased features. This can solve the problems of single feature extraction and the inability to extract or even obtain edge features in the existing technology, so that the trained liveness detection model improves recognition accuracy and detection precision. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0013] Figure 1 A schematic diagram of an application environment for liveness detection provided by an embodiment of the present invention;

[0014] Figure 2 A schematic diagram of a feature extraction model provided by an embodiment of the present invention;

[0015] Figure 3 A schematic flow chart of a method for training a feature extraction model provided by an embodiment of the present invention;

[0016] Figure 4 A schematic diagram of obtaining a second training image according to an embodiment of the present invention;

[0017] Figure 5 A schematic flow chart of a method for training a liveness detection model provided in an embodiment of the invention;

[0018] Figure 6 A schematic diagram of an initial liveness detection model provided by an embodiment of the present invention;

[0019] Figure 7 A schematic flowchart of step S503 provided in an embodiment of the present invention;

[0020] Figure 8 A schematic diagram of feature cascade of a feature extraction model provided by an embodiment of the present invention;

[0021] Figure 9 A schematic flow chart of a liveness detection method provided in an embodiment of the present invention;

[0022] Figure 10 A functional module diagram of a training device for a liveness detection model provided by an embodiment of the invention;

[0023] Figure 11 A functional module diagram of a living body detection device provided by an embodiment of the invention;

[0024] Figure 12 A schematic structural diagram of an electronic device provided in an embodiment of the invention. DETAILED DESCRIPTION

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0026] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.

[0027] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0028] In the description of the present invention, it should be noted that if the terms "upper", "lower", "inside", "outside", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the accompanying drawings, or is the orientation or position relationship in which the product of the invention is usually placed when in use. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be understood as a limitation on the present invention.

[0029] In addition, the terms "first", "second", etc., if used, are merely used to distinguish and describe, and should not be understood as indicating or implying relative importance.

[0030] It should be noted that, in the absence of conflict, the features in the embodiments of the present invention may be combined with each other.

[0031] Currently, facial recognition has been applied in many fields, such as ID verification at train stations, online and even offline business processing. However, facial data is extremely easy to obtain. All it takes is a clear image taken with a mobile phone or other camera. This image can evade detection by facial recognition systems without liveness detection capabilities, making the face, the "key," easily stolen.

[0032] To ensure high security, facial recognition systems must verify that faces in video and images are real, not photos. This is where liveness detection comes in. Liveness detection involves a computer determining whether a detected face is real or a forged face attack, such as a legitimate user photo or a pre-recorded video. Traditionally, this problem is considered a binary classification problem: "live" versus "fake." However, it can also be considered a multi-classification problem, such as real-person, image attacks, video playback attacks, and mask attacks.

[0033] See Figure 1 , Figure 1 A schematic diagram of an application environment for liveness detection provided by an embodiment of the present invention includes a terminal 110, a first service device 120, and a second service device 130, wherein the terminal 110, the first service device 120, and the second service device 130 can be connected via a network.

[0034] The terminal 110 can send a model training instruction to the first service device 120. The first service device 120 responds to the model training instruction and performs model training on the liveness detection model based on image samples of different application scenarios. Then, the first service device 120 deploys the trained liveness detection model to the second service device 130. The terminal 110 collects the image to be tested and sends the image to be tested to the second service device 130, so that the second service device 130 performs liveness detection on the image to be tested through the deployed liveness detection model, and obtains a detection result of whether the image to be tested is a liveness image. The second service device 130 can perform corresponding operations based on the detection result. Alternatively, the second service device 130 can also send the detection result to the terminal 110 to display the detection result, and / or perform corresponding operations based on the detection result.

[0035] Among them, the terminal 110 is a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch or shooting device (such as a camera), etc. In addition, the terminal 110 can also be a display with the first service device 120 or the second service device 130, and the display can be integrated with a camera or connected to a camera.

[0036] The first service device 120 and the second service device 130 can be independent physical devices or a service device cluster composed of multiple physical service devices. The fact that the first service device 120 has a model training function and the second service device 130 has a liveness detection function is merely an example and does not limit their functions. They can also be servers that integrate both model training and liveness detection functions.

[0037] Please continue to see Figure 1 As a key component of face recognition systems, liveness detection must effectively prevent attacks from non-live faces. Currently, liveness detection is divided into image-based and video-based methods based on the form of input data. The following describes the shortcomings of these two methods:

[0038] Video-based liveness detection first detects faces in a video and analyzes whether the face meets certain criteria, such as whether it exhibits specific movements, to determine if the face is real. However, this detection method requires user cooperation, is time-consuming, offers a poor user experience, and is unsuitable for some scenarios. It also presents the risk of video attacks. Furthermore, video-based liveness detection is performed only once in the facial recognition process, making it prone to the possibility of using fake faces for recognition after a liveness detection has passed.

[0039] The image-based liveness detection method first preprocesses the image, then sends it to the model for detection, and directly outputs whether it is a real face. The images used are useful single-modal images, such as RGB images, depth maps, etc., and multi-modal images, such as RGB images + depth maps + infrared imaging, etc. The method using single-modal images has a single data source, and the features extracted by the model are not sufficient to describe the attack characteristics, resulting in inaccurate detection results. The method using multi-modal images uses more data sources and is significantly better than the single-modal image method. However, both methods perform detection on data from a single source, and then combine multiple detection results in some way (such as logical operations, weighted averages, etc.) to obtain the final result. There is also the problem of poor recognition accuracy.

[0040] On the other hand, regardless of whether multimodal or single-modal images are used, regardless of whether video or image methods are used, conventional methods are used to train conventional convolutional neural network models when calculating image features. This type of neural network model mainly extracts the texture features of the image, while images of certain attack methods include effective edge information (such as moiré patterns on electronic screens and facial features of masks). This type of network cannot obtain effective edge information, resulting in incorrect recognition.

[0041] Research has found that there are differences in local textures between re-shot face photos or videos and real faces. For example, in a face video shot by a mobile phone, due to the difference in light reflection between the phone screen and the face skin, the phone screen imaged by the camera will have moiré patterns. In view of this, the present invention aims to obtain multi-granularity features by designing a feature extraction model: biased texture features, biased edge features and unbiased features. Therefore, the feature extraction model provided in the embodiment of the present application can be formed as follows: Figure 2 , Figure 2 A schematic diagram of a feature extraction model provided by an embodiment of the present invention.

[0042] It should be noted that Figure 2 There is no special restriction on the structure and number of networks of the feature extraction model shown. It only needs to ensure that there is at least one feature extraction network of each type and that each feature extraction network can output features with consistent width and height.

[0043] in, Figure 2 The feature extraction model shown can include at least: a first type of feature extraction network, a second type of feature extraction network and a third type of feature extraction network; the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second type of feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third type of feature extraction network extracting texture features is equal to the probability of extracting edge features.

[0044] Among them, the texture features of the face are highly discriminative, so the biometric recognition technology for identity authentication based on facial features has a good classification and recognition effect; the above-mentioned edge features include: any one or combination of photo frame information, bracket information, screen information, and the intersection edge information of the mask and the human body. These edge information are information other than facial information, which is convenient for identifying information of facial attacks, that is, information that cannot be formed by a real face.

[0045] The following first introduces the three types of feature extraction networks in the feature extraction model in the embodiment of the present application.

[0046] The first type of feature extraction network: This type of feature extraction network is mainly a texture-biased network, that is, in the process of feature extraction, the texture features are extracted more favorably, which means that the edge information in the texture-biased features is not as obvious as the edge information in the edge-biased features, that is, the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features. In possible embodiments, the first type of feature extraction network can be but is not limited to using ordinary classification convolutional neural networks, such as ResNet, EfficientNet, GoogleNet, MobileNet V2, VGGNet, etc.

[0047] The second type of feature extraction network: The basic principle of this type of feature extraction network is to identify texture areas and edge areas based on the local self-information of the image. It can be considered that the self-information of the texture area is smaller than the self-information of the edge area. Based on this, during the training process, the output of each layer of the network is decorrelated with the input texture, thereby realizing the edge bias of the feature, which means that the texture information in the feature biased towards the edge is not as obvious as the texture information in the feature biased towards the texture, that is, the probability of extracting texture features is lower than the probability of extracting edge features.

[0048] The third type of feature extraction network: This type of feature extraction network is an unbiased feature extraction network that not only extracts edge information and texture information, that is, the probability of the extraction network extracting texture features is equal to the probability of extracting edge features, but also preserves the spatial relationship between edge and texture information and the corresponding strength relationship. This is because in real face images, the eye contour not only has edge information, but also relatively rich texture information. In fake face images, such as printed pictures or faces on electronic displays, the edges of the eye contour are relatively clear, while the texture information may not be obvious or may differ significantly from that of real faces. For example, many LCD screens have moiré patterns that are either obvious or not. The unbiased network extracts these features, so that both texture and edge information are represented as much as possible, making the features of real and fake faces as different as possible. In possible embodiments, the third type of feature extraction network can be, but is not limited to, a common classification convolutional neural network, such as ResNet, EfficientNet, GoogleNet, MobileNet V2, VGGNet, etc.

[0049] By designing feature extraction models with different feature extraction granularities, we can obtain multi-granular texture features of the image: biased texture features, biased edge features and unbiased features. By cascading these features and combining them with the feature fusion network, we can achieve autonomous fusion and classification of features. This allows the obtained liveness detection model to accurately identify the authenticity of human faces and solve the problems of weak feature expression and low recognition rate.

[0050] In order to achieve the above effect, the present application embodiment first provides a method for training a feature extraction model. Figure 3 , Figure 3 A schematic flow chart of a method for training a feature extraction model provided by an embodiment of the present invention. It should be noted that the method for training a feature extraction model can be applied to Figure 1 In the first service device 120, of course, if the second service device 130 also integrates a model training function, the training method of the feature extraction model can be applied to any of the above service devices.

[0051] Step 1: Acquire a first training image set, and obtain a second training image set based on the first training image set; wherein any second training image in the second training set is synthesized based on two first training images in the first training image set;

[0052] Step 2: Based on the first training image set, the model parameters of the initial first-class feature extraction network are trained to obtain the first-class feature extraction network, and the model parameters of the initial second-class feature extraction network are trained to obtain the second-class feature extraction network.

[0053] Step 3: Based on the second training image set, the model parameters of the initial third-category feature extraction network are trained to obtain the third-category feature extraction network.

[0054] Step 4: Obtain a feature extraction model based on the trained first-category feature extraction network, second-category feature extraction network, and third-category feature extraction network.

[0055] The training method of the feature extraction model provided by the embodiment of the present invention first obtains a first training image set and a second training image set, and respectively trains an initial first-class feature extraction network and a second-class feature extraction network through the first training image set, so that the first-class feature extraction network is biased to extract texture features, and the second-class feature extraction network is biased to extract edge features, and trains a third-class feature extraction network through the second training image set, so that the third-class feature extraction network is unbiased in feature extraction. The feature extraction network finally generated can realize the extraction of multi-granularity texture features, and can overcome the problem of insufficient feature extraction in the existing model leading to the accuracy of liveness detection.

[0056] The following is an introduction to steps 1 to 4 in the above training method.

[0057] In step 1, firstly, a first training image set and a second training image set are obtained.

[0058] In the embodiment of the present application, a large number of images can be collected from different application environments. The category and number of images are not limited, and there are no duplicate images in each category. To achieve better training results, if a target object appears in each image, the target object can be a person, an animal, a plant, etc., which is not limited here. To verify the classification results of the model, the location, size, and category of the target can also be marked, thereby obtaining a first training image set. After obtaining the first training image set, the second training image set can be determined from the first training image set.

[0059] In one possible embodiment, the second training image set may be obtained by sampling from the first training image set multiple times to obtain multiple groups of first training image pairs; for each pair of first training images, migrating the texture features of one of the first training images to the other first training image, and generating a second training image based on the texture identifier of one of the first training images and the shape identifier of the other first training image, and all the second training images constitute the second training image set.

[0060] Among them, the method of sampling multiple times from the first training image set to obtain multiple groups of first training image pairs can be to perform uniform random sampling among the labeled images in the first training image set. That is to say, the target image set can be first determined from the first training image set, and then uniform random sampling can be performed in the target image set, sampling two images each time, to obtain multiple first training image pairs.

[0061] For easier understanding, see Figure 4 , Figure 4 This is a schematic diagram of obtaining a second training image according to an embodiment of the present invention. It can be seen that the target object of one of the first training images is a flower, and the target object of the other first training image is a seal. The texture features of the flower in the first training image with the flower as the target object are transferred to the first training image with the seal as the target object. The synthesized second training image has the texture of the flower and the shape of the seal, and then the label of the synthesized image is generated. The label generation method is as follows: the original label uses the one-hot method, and the shape label uses the y s Indicates that the texture label is y t Represented by, and the label of the synthesized second training image can be expressed as:

[0062]

[0063] Among them, γ is a hyperparameter, the theoretical value range is [0,1], and the optimal value range is generally between [0.5,0.9].

[0064] The first training image set and the second training image set obtained in the above manner can be used to train a feature extraction model, so that the trained feature extraction model can extract multi-granularity features.

[0065] In step 2 and step 3, based on the first training image set, the initial first type feature extraction network and the second type feature extraction network are trained respectively, and based on the second training image set, the third type feature extraction network is trained.

[0066] In the embodiment of the present application, the first type of feature extraction network and the second type of feature extraction network are both trained using the first training image set as training samples, and the third type of feature extraction network is trained using the second training image set as training samples. Among them, the training methods of the first type of feature extraction network and the third type of feature extraction network can adopt existing mature training methods, which will not be repeated here.

[0067] The following is a detailed introduction to the training process of the second type of feature extraction network in step 2 above. Training the second type of feature extraction network may include the following steps:

[0068] a1, input the first training image into the initial second type feature extraction network for feature extraction.

[0069] a2. During the feature extraction process, the first texture feature and the second texture feature corresponding to each convolutional layer in the second type of feature extraction network are decorrelated.

[0070] Among them, the first texture feature is the texture feature output by the previous convolution layer of each convolution layer. It can be considered that the first texture feature is the input of each convolution layer, and the second texture feature is the texture feature output by each convolution layer.

[0071] a3, based on each convolutional layer after decorrelation processing, the second type of feature extraction network is obtained.

[0072] In the embodiment of the present application, the second type of feature extraction network (hereinafter referred to as the network biased towards edge information) is based on texture decorrelation of self-information in different areas of the image, thereby achieving edge bias of the feature.

[0073] That is to say, in step a2, it is necessary to implement decorrelation processing of the output texture features and the input texture features for each convolution layer. Assume that: and They represent the output of the l-1th layer (that is, the input of the lth layer) and the output of the lth layer, c, w, and h represent the channel, width, and height of the image, respectively. The decorrelation processing can be done as follows:

[0074] a21, dividing the first feature map corresponding to the first texture feature into blocks according to a preset size to obtain multiple image blocks.

[0075] In the embodiment of the present invention, the preset size may be 3 times 3, 5 times 5, etc., which is not limited here.

[0076] a22, calculate the autocorrelation information of each image block, and calculate the discard probability of the neurons corresponding to the multiple channels output by each convolutional layer based on the autocorrelation information and the preset distribution function.

[0077] In one possible implementation, the autocorrelation information may be calculated as follows:

[0078] First, according to a preset radius parameter, a neighborhood image block corresponding to each image block is determined from the first feature map.

[0079] Then, the probability distribution corresponding to each of the multiple image blocks is determined according to the preset radius parameter, the neighborhood image block and the preset Gaussian kernel function. The probability distribution is specifically expressed as follows:

[0080]

[0081] Where R represents a preset radius parameter, which can be a Manhattan radius, and p′ represents the neighborhood image block of the image block p. The domain image block representing the jth image block; the Gaussian kernel function is expressed as follows: 2h 2 ),h represents bandwidth.

[0082] Finally, based on the probability distribution and preset parameters corresponding to each image block, the autocorrelation information of each image block is calculated.

[0083] Assume that the first feature map of the l-th layer input is divided into N image blocks, then the self-information of image block p can be expressed as:

[0084]

[0085] in, The self-information of the j-th image block p corresponding to the feature map of the first feature of the l-th layer input, p′ represents the neighborhood image block of the image block p; h represents the bandwidth; const is a preset parameter.

[0086] After obtaining the self-information corresponding to each image block, we can use the following relationship to obtain the self-information to adjust the dropout probability of neurons corresponding to the multiple channels output by each convolutional layer:

[0087]

[0088] in, Represents the dropout probability of the jth neuron in the cth channel of the output of the lth layer; The self-information of the image block p corresponding to the feature map of the first feature of the l-th layer input; T represents the temperature.

[0089] a23, according to the discarding probability of the neurons corresponding to the multiple channels, sampling is performed in all neurons corresponding to each convolutional layer, and the values of the sampled neurons are set to preset values to achieve decorrelation processing of the first texture feature and the second texture feature.

[0090] In this embodiment, the preset value is 0, and the collected neuron values are set to 0, so that the model output and texture can be decorrelated. Based on each convolution layer after the decorrelation process, the second type of feature extraction network is obtained.

[0091] Finally, in step 4, a feature extraction model can be obtained based on the trained first-category feature extraction network, second-category feature extraction network, and third-category feature extraction network.

[0092] In the implementation of this application, the trained model includes two parts: feature extraction and classification. By removing the classification part, the feature extraction model is obtained.

[0093] It should be noted that, in one scenario, the above-mentioned feature extraction model can be pre-trained and stored in the memory of the training device, and when the feature extraction model is needed (for example, to train a liveness detection model), it can be directly loaded from the memory; in another scenario, the above-mentioned steps 1 to 4 can also be executed in real time according to the obtained model training instructions to obtain the trained feature extraction model.

[0094] In order to realize the above embodiment Figure 2 The training method of the feature extraction model provided in the embodiment of the present invention can be executed in a hardware device or in the form of a software module. When the training method of the feature extraction model is implemented in the form of a software module, the embodiment of the present invention also provides a training device for the feature extraction model. The training device for the feature extraction model can be used to perform the above Figure 2 The various steps in order to achieve the training effect.

[0095] In the embodiment of the present application, the training device of the feature extraction model can be stored in the memory or fixed in the form of software or firmware. Figure 1 The modules are stored in the operating system (OS) of the first service device 120 and can be executed by the processor of the first service device 120. Meanwhile, the data and program codes required to execute the modules can be stored in the memory of the first service device 120.

[0096] In order to obtain a more accurate liveness detection model, the embodiment of the present invention provides a training method for the liveness detection model based on the obtained feature extraction model. Figure 5 , Figure 5 A schematic flow chart of a training method for a liveness detection model provided in an embodiment of the invention. It should be noted that the training method for the liveness detection model can be applied to Figure 1 In the first service device 120, of course, if the second service device 130 also integrates a model training function, the training method of the feature extraction model can be applied to any of the above service devices.

[0097] S501: Obtain a face training sample set.

[0098] The face training sample set includes multiple face training samples, and each face training sample is marked with an identifier indicating whether the face training sample is a positive sample or a negative sample.

[0099] S502: construct an initial living body detection model based on the constructed initial feature fusion model and the trained feature extraction model.

[0100] In the embodiment of the present application, the feature extraction model trained in step S502 is the model obtained according to the training method of steps 1 to 4 above, and has the characteristics of the feature extraction model trained above, which will not be repeated here.

[0101] It should be noted that if there is no trained feature extraction model, it can be performed before step S501 or step S502. Figure 2 The various steps in are used to obtain the trained feature extraction model; if a trained feature extraction model exists, it can be directly loaded to obtain the trained feature extraction model. The user can determine the order of implementation according to the actual situation, which is not limited here.

[0102] S503: Training the model parameters of the initial liveness detection model according to the face training sample set to obtain the liveness detection model.

[0103] According to an embodiment of the present invention, a method for training a liveness detection model is provided. First, a face training sample set is obtained, and a liveness detection model is constructed based on the trained feature extraction model and the initial feature model. The initial liveness detection model is trained based on the face training sample set. Since the feature extraction model can extract multi-granularity texture features, the liveness detection model constructed based on this feature extraction model solves the problem of low detection accuracy caused by insufficient feature extraction in existing models.

[0104] It should be noted that the location of the database used to store the face training sample set complies with the laws and regulations of the country / region where the above-mentioned record-related behaviors occur, including but not limited to: authorization, generation, use, storage, etc.

[0105] It should also be noted that in the process of obtaining the above-mentioned face training sample set, the collection, use and storage instructions are provided to the relevant users in a public form and the user's authorization is obtained. The obtained face sample images do not include face information unrelated to the services provided by this embodiment.

[0106] The following is for Figure 5 Steps S501 to S503 in FIG. 5 are introduced.

[0107] In step S501, a face training sample set is obtained.

[0108] The face training sample set includes multiple face training samples, and the face training samples are marked with an identifier indicating whether the face training samples are positive samples or negative samples.

[0109] In this embodiment, the above-mentioned face training samples can come from real face image samples and forged face image samples from different application scenarios, and can be but not limited to multimodal face images, where the real face image samples can be obtained by taking real people using a camera, a depth camera and an infrared camera; the forged face image samples include printed face images, face images played by electronic devices, faces wearing masks, video screenshots, etc.

[0110] In a possible embodiment, image samples for different application scenarios may include at least one of the following: images captured in different shooting environments, images captured using different shooting devices, and images captured using a certain shooting device capturing different objects. Different shooting environments may include: indoor environments with strong lighting, indoor environments with weak lighting, outdoor environments with strong lighting (e.g., daytime), and outdoor environments with weak lighting (e.g., nighttime), as well as backlit and frontlit shooting environments, such as outdoor environments with strong lighting and backlit environments. Furthermore, a mask or a three-dimensional human body model may also be used.

[0111] In order to evaluate the accuracy of the liveness detection model, the following annotations can also be completed for the training samples: face position and size information, marking the face training samples as positive samples or negative samples, etc.

[0112] In step S502, an initial living body detection model is constructed based on the constructed initial feature fusion model and the trained feature extraction model.

[0113] In this embodiment, the trained feature extraction model can be found in Figure 3 As shown in the figure, the training process is not described here. The initial liveness detection model is composed of the trained feature extraction model and the initial feature fusion model. The initial liveness detection model can be found in Figure 6 , Figure 6 A schematic diagram of an initial living body detection model provided by an embodiment of the present invention.

[0114] The initial feature fusion model in the embodiments of the present application may include, but is not limited to, the following layers: convolutional layers, activation layers, batch normalization layers, softmax layers, and other network layers. The input of the feature fusion network is the concatenated features of the features output by the first, second, and third feature extraction networks in the feature extraction model. The output is two probability values, representing the probability of a real face or a fake face.

[0115] In step S503, the model parameters of the initial liveness detection model are trained according to the face training sample set to obtain the liveness detection model.

[0116] By combining the trained feature extraction model with the trained feature fusion network, an initial liveness detection model can be obtained. The initial liveness detection model is then trained using the face training sample set to obtain the final liveness detection model.

[0117] It should be noted that, in one scenario, the above-mentioned liveness detection model can be pre-trained and stored in the memory of the training device, and can be directly loaded from the memory when liveness detection is required; in another scenario, the above-mentioned liveness detection model can also be executed in real time according to the obtained model training instructions. Figure 5 , and obtain the trained feature extraction model.

[0118] Optionally, an implementation of the above step S503 can be found in Figure 7 , Figure 7 A schematic flow chart of step S503 provided in an embodiment of the present invention:

[0119] S503-1: Input the face training sample set into the first feature extraction network, the second feature extraction network, and the third feature extraction network respectively, and concatenate the features extracted by the first feature extraction network, the second feature extraction network, and the third feature extraction network to obtain the concatenated features. Figure 8 , Figure 8 A schematic diagram of feature cascade of a feature extraction model provided in an embodiment of the present invention.

[0120] S503-2, keeping the model parameters of the feature extraction model unchanged, inputting the cascaded features into the initial feature fusion model, training the model parameters of the initial feature fusion model, and obtaining a trained feature fusion model.

[0121] In this embodiment, the parameters of the feature extraction model remain unchanged in this step, and are used to extract the features of the true and false face data. Each model feature is cascaded in a certain order and input into the feature fusion network. The feature fusion model obtains the result and compares it with the real label to obtain the error. The feature fusion network parameters are optimized and adjusted through back propagation and optimizer (such as Adam, SGD) until the end condition is met.

[0122] S503-3, obtaining a living body detection model based on the trained feature extraction model and the trained feature fusion model.

[0123] In order to verify the accuracy of the liveness detection model provided by the embodiment of the present invention, experiments were conducted on an open source multimodal dataset. The experimental results are shown in Table 1. By comparison, it can be seen that the use of multi-granularity texture feature fusion can not only obtain rich features, but also learn the relationship between the different granularity texture features of each modality, thereby achieving better results.

[0124]

[0125] Among them, net in Table 1 represents the basic network feathernetB, net+RGB represents the detection method of a single basic network combined with RGB images; net+depth represents the detection method of a single basic network combined with depth images; net+ir represents the detection method of a single basic network combined with infrared images; result_fusion represents the detection method of the liveness detection model of this scheme performing feature extraction and classification and fusing the detection results; feat_fusion represents the detection method of the liveness detection model of the scheme performing feature fusion; TPR (True Positive Rate) represents the probability of true samples, that is, the recall rate; FPR (False Positive Rate) represents the probability of false positive samples; APCER represents the classification error rate of attack performance (forged face); NPCER represents the classification error rate of normal performance (real face); ACER represents the average classification error rate, which is the evaluation result of the proposed model.

[0126] Based on the above-trained liveness detection model, the embodiment of the present invention also provides a liveness detection method, which can be applied to Figure 1 In the second service device 130, see Figure 9 , Figure 9 A schematic flow chart of a liveness detection method provided in an embodiment of the present invention:

[0127] S901: Acquire a facial image of a face to be detected.

[0128] In this embodiment, the facial images come from facial images collected in different application scenarios, such as mobile phone face recognition unlocking, face recognition payment, identity authentication and other application scenarios.

[0129] S902: Input the face image into the trained liveness detection model and output the detection result.

[0130] The liveness detection model in the above step S902 is trained by the liveness detection model training method provided in the embodiment of the present application, which will not be repeated here.

[0131] It should be noted that before executing step S902, the trained liveness detection model can be loaded first. If the trained liveness detection model does not exist, the above steps S501 to S503 can be executed to obtain the trained liveness detection model.

[0132] It should also be noted that before executing the above steps S501 to S503, it is necessary to determine whether there is a trained feature extraction model. If not, it is necessary to first execute Figure 2 Steps 1 to 4 in the above are used to obtain a trained feature extraction model, and then the feature extraction model is used to perform the above steps S501 to S503 to obtain a trained liveness detection model.

[0133] After obtaining and loading the model, the facial image is input into the model, and the model will give the corresponding prediction result, that is, whether the face in the facial image is a real face or a fake face. The prediction result can be used for subsequent business operations.

[0134] For example, in one embodiment, when the object to be tested approaches the access control device, the liveness detection terminal (including the gate) in the access control device is triggered to perform liveness detection on the object to be tested, and the data acquisition terminal in the access control device collects the facial image of the object to be tested. When the facial image is collected, the data acquisition terminal sends the facial image to the liveness detection terminal, and the liveness detection terminal determines whether the facial image is a live image through the built-in liveness detection model to obtain the detection result. When the image to be tested is a live image, an indication message of passing the test is sent to the data acquisition terminal of the access control device. The data acquisition terminal will display the words "Please pass" in the result display area or output the voice of "Please pass" through the voice device. In addition, the liveness detection terminal 132 opens the gate to allow the object to be tested to enter.

[0135] In order to realize the above embodiment Figure 5 The training method of the liveness detection model provided by the embodiment of the present invention can be executed in a hardware device or in the form of a software module. When the training method of the feature extraction model is implemented in the form of a software module, the embodiment of the present invention also provides a training device for the liveness detection model. Figure 10 The training method 200 of the living body detection model may include:

[0136] The acquisition module 210 is used to acquire a face training sample set; the face training sample set includes multiple face training samples; the face training sample marking training sample is a positive sample or a negative sample.

[0137] A construction module 220 is used to construct an initial living body detection model based on the constructed initial feature fusion model and the trained feature extraction model;

[0138] Among them, the feature extraction model includes at least a first type of feature extraction network, a second type of feature extraction network and a third type of feature extraction network; the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second type of feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third type of feature extraction network extracting texture features is equal to the probability of extracting edge features.

[0139] The training module 230 is used to train the model parameters of the initial liveness detection model according to the face training sample set to obtain the liveness detection model.

[0140] It is understandable that the acquisition module 210, the construction module 220 and the training module 230 can be executed in a coordinated manner. Figure 5 、 Figure 7 Each step in the process is performed to achieve the corresponding technical effects.

[0141] In an optional embodiment, the training module 230 can also be used to: obtain a first training image set, and obtain a second training image set based on the first training image set; wherein, any second training image in the second training set is synthesized based on two first training images in the first training image set; based on the first training image set, train the model parameters of the initial first-category feature extraction network to obtain the first-category feature extraction network, and train the model parameters of the initial second-category feature extraction network to obtain the second-category feature extraction network; based on the second training image set, train the model parameters of the initial third-category feature extraction network to obtain the third-category feature extraction network; obtain a feature extraction model based on the trained first-category feature extraction network, the second-category feature extraction network and the third-category feature extraction network.

[0142] In an optional embodiment, the training module 230 is specifically used to: sample multiple times from the first training image set to obtain multiple pairs of first training image pairs; for each pair of first training image pairs, migrate the texture features of one of the first training images to the other first training image, and generate a second training image based on the texture identifier of one of the first training images and the shape identifier of the other first training image, and all the second training images constitute the second training image set.

[0143] In an optional embodiment, the training module 230 is specifically used to: input the first training image into the initial second-class feature extraction network for feature extraction; in the process of feature extraction, decorrelate the first texture features and the second texture features corresponding to each convolution layer in the second-class feature extraction network; wherein the first texture features are the texture features output by the previous convolution layer of each convolution layer; the second texture features are the texture features output by each convolution layer; based on each convolution layer after decorrelation processing, a second-class feature extraction network is obtained.

[0144] In an optional embodiment, the training module 230 is specifically used to divide the first feature map corresponding to the first texture feature into blocks according to a preset size to obtain multiple image blocks; calculate the autocorrelation information of each image block, and calculate the discarding probability of neurons corresponding to each of the multiple channels output by each convolutional layer based on the autocorrelation information of each image block and a preset distribution function; sample all neurons corresponding to each convolutional layer based on the discarding probability of neurons corresponding to the multiple channels, and set the values of the sampled neurons to preset values to achieve decorrelation processing of the first texture feature and the second texture feature.

[0145] In an optional embodiment, the training module 230 is specifically used to determine the neighborhood image block corresponding to each image block from the first feature map according to a preset radius parameter; determine the probability distribution corresponding to each image block according to the preset radius parameter, the neighborhood image block and a preset Gaussian kernel function; and calculate the autocorrelation information of each image block based on the probability distribution corresponding to each image block and the preset parameters.

[0146] In an optional implementation, the training module 230 is specifically used to input the face training sample set into the first type of feature extraction network, the second type of feature extraction network and the third type of feature extraction network respectively, and cascade the features extracted by the first type of feature extraction network, the second type of feature extraction network and the third type of feature extraction network to obtain the cascaded features; keep the model parameters of the feature extraction model unchanged, input the cascaded features into the initial feature fusion model, train the model parameters of the initial feature fusion model, and obtain the trained feature fusion model; obtain the liveness detection model based on the trained feature extraction model and the trained feature fusion model.

[0147] In order to realize the above embodiment Figure 9 The liveness detection method provided by the embodiment of the present invention can be executed in a hardware device or in the form of a software module. When the liveness detection method is implemented in the form of a software module, the embodiment of the present invention also provides a liveness detection device. Figure 11 , the living body detection device 300 may include:

[0148] The acquisition module 310 is used to acquire a facial image of a face to be detected.

[0149] The detection module 320 is used to input the face image into the trained liveness detection model and output the detection result.

[0150] Among them, the feature extraction model in the liveness detection model includes at least a first-type feature extraction network, a second-type feature extraction network and a third-type feature extraction network; the probability of the first-type feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second-type feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third-type feature extraction network extracting texture features is equal to the probability of extracting edge features.

[0151] It is understandable that the acquisition module 310 and the detection module 320 can be executed in a coordinated manner. Figure 9 Each step in the process is performed to achieve the corresponding technical effects.

[0152] It should be noted that the various functional modules in the liveness detection model training device 600 and / or the liveness detection device 300 provided in the embodiment of the present invention can be stored in the memory in the form of software or firmware or fixed in the operating system (OS) of the electronic device 400, and can be executed by the processor in the electronic device 400. At the same time, the data and program code required to execute the above modules can also be stored in the memory.

[0153] like Figure 12 , Figure 12 Schematic diagram of a block diagram of an electronic device provided by an embodiment of the present invention. The electronic device 400 can be Figure 1 The first service device 120 or the second service device 130 in the embodiment of the present invention, the electronic device 400 includes a communication interface 401, a processor 402 and a memory 403. The processor 402, the memory 403 and the communication interface 401 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines. The memory 403 can be used to store software programs and modules, such as program instructions / modules corresponding to the model training method and / or liveness detection method provided in the embodiment of the present invention. The processor 402 executes various functional applications and data processing by executing the software programs and modules stored in the memory 403. The communication interface 401 can be used for signaling or data communication with other node devices. In the present invention, the electronic device 40 can have multiple communication interfaces 401.

[0154] Among them, the memory 403 can be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0155] Processor 402 may be an integrated circuit chip with signal processing capabilities. The processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0156] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the liveness detection model training method and / or liveness detection method described in any of the aforementioned embodiments. The computer-readable storage medium may be, but is not limited to, a USB flash drive, a mobile hard drive, ROM, RAM, PROM, EPROM, EEPROM, a magnetic disk, or an optical disk, among other media capable of storing program code.

[0157] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A training method for a liveness detection model, characterized in that: The method comprises: Obtaining a face training sample set; the face training sample set includes a plurality of face training samples; the face training samples are marked with an identifier indicating whether the face training samples are positive samples or negative samples; Construct an initial liveness detection model based on the constructed initial feature fusion model and the trained feature extraction model; The feature extraction model includes at least a first type of feature extraction network, a second type of feature extraction network, and a third type of feature extraction network; the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second type of feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third type of feature extraction network extracting texture features is equal to the probability of extracting edge features; The feature extraction model is trained in the following manner: obtaining a first training image set, and obtaining a second training image set based on the first training image set; wherein any second training image in the second training image set is synthesized based on two first training images in the first training image set; based on the first training image set, training the model parameters of the initial first-category feature extraction network to obtain the first-category feature extraction network, and training the model parameters of the initial second-category feature extraction network to obtain the second-category feature extraction network; based on the second training image set, training the model parameters of the initial third-category feature extraction network to obtain the third-category feature extraction network; and obtaining the feature extraction model based on the trained first-category feature extraction network, the trained second-category feature extraction network, and the trained third-category feature extraction network; The model parameters of the initial liveness detection model are trained according to the face training sample set to obtain a liveness detection model.

2. The training method according to claim 1, characterized in that Obtaining a second training image set based on the first training image set includes: Sampling multiple times from the first training image set to obtain multiple pairs of first training images; For each pair of first training images, the texture features of one of the first training images are transferred to the other first training image, and a second training image is generated based on the texture identifier of the one of the first training images and the shape identifier of the other first training image, and all the second training images constitute the second training image set.

3. The training method according to claim 1, characterized in that Based on the first training image set, training the initial model parameters of the second type feature extraction network to obtain the second type feature extraction network includes: Inputting the first training image into the initial second-category feature extraction network for feature extraction; During the feature extraction process, decorrelation processing is performed on the first texture feature and the second texture feature corresponding to each convolutional layer in the second type feature extraction network; wherein the first texture feature is the texture feature output by the convolutional layer before each convolutional layer; and the second texture feature is the texture feature output by each convolutional layer; Based on each convolutional layer after decorrelation processing, the second type of feature extraction network is obtained.

4. The training method according to claim 3, characterized in that During the feature extraction process, decorrelation processing is performed on the first texture feature and the second texture feature corresponding to each convolutional layer in the second type of feature extraction network, including: Divide the first feature map corresponding to the first texture feature into blocks according to a preset size to obtain a plurality of image blocks; Calculating the autocorrelation information of each image block, and calculating the discarding probability of neurons corresponding to each of the multiple channels output by each convolutional layer based on the autocorrelation information of each image block and a preset distribution function; According to the discarding probability of the neurons corresponding to each of the multiple channels, sampling is performed in all neurons corresponding to each convolutional layer, and the values of the sampled neurons are set to preset values to achieve decorrelation processing of the first texture feature and the second texture feature.

5. The training method according to claim 4, characterized in that Calculate the autocorrelation information for each image block, including: Determining, from the first feature map, a neighborhood image block corresponding to each image block according to a preset radius parameter; Determining the probability distribution corresponding to each of the image blocks according to the preset radius parameter, the neighborhood image block and a preset Gaussian kernel function; Based on the probability distribution corresponding to each image block and preset parameters, the autocorrelation information of each image block is calculated.

6. The training method according to claim 1, characterized in that: Training the model parameters of the initial liveness detection model according to the face training sample set to obtain the liveness detection model includes: Inputting the face training sample set into the first feature extraction network, the second feature extraction network, and the third feature extraction network respectively, and cascading the features extracted by the first feature extraction network, the second feature extraction network, and the third feature extraction network to obtain cascaded features; Keeping the model parameters of the feature extraction model unchanged, inputting the cascaded features into the initial feature fusion model, training the model parameters of the initial feature fusion model, and obtaining the trained feature fusion model; The living body detection model is obtained according to the trained feature extraction model and the trained feature fusion model.

7. A method for detecting a living body, characterized in that: The method comprises: Obtain a facial image of a face to be detected; Inputting the face image into the trained liveness detection model and outputting the detection result; The feature extraction model in the liveness detection model includes at least a first type of feature extraction network, a second type of feature extraction network, and a third type of feature extraction network; the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second type of feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third type of feature extraction network extracting texture features is equal to the probability of extracting edge features; The feature extraction model is trained in the following manner: obtaining a first training image set, and obtaining a second training image set based on the first training image set; wherein, any second training image in the second training image set is synthesized based on two first training images in the first training image set; based on the first training image set, training the model parameters of the initial first type of feature extraction network to obtain the first type of feature extraction network, and training the model parameters of the initial second type of feature extraction network to obtain the second type of feature extraction network; based on the second training image set, training the model parameters of the initial third type of feature extraction network to obtain the third type of feature extraction network; and obtaining the feature extraction model based on the trained first type of feature extraction network, the second type of feature extraction network and the third type of feature extraction network.

8. A training device for a liveness detection model, characterized in that: include: An acquisition module is used to obtain a face training sample set; The face training sample set includes multiple face training samples; The face training sample is marked with an identifier indicating whether the face training sample is a positive sample or a negative sample; A construction module is used to construct an initial liveness detection model based on the constructed initial feature fusion model and the trained feature extraction model; The feature extraction model includes at least a first type of feature extraction network, a second type of feature extraction network, and a third type of feature extraction network; the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second type of feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third type of feature extraction network extracting texture features is equal to the probability of extracting edge features; The feature extraction model is trained in the following manner: obtaining a first training image set, and obtaining a second training image set based on the first training image set; wherein any second training image in the second training image set is synthesized based on two first training images in the first training image set; based on the first training image set, training the model parameters of the initial first-category feature extraction network to obtain the first-category feature extraction network, and training the model parameters of the initial second-category feature extraction network to obtain the second-category feature extraction network; based on the second training image set, training the model parameters of the initial third-category feature extraction network to obtain the third-category feature extraction network; and obtaining the feature extraction model based on the trained first-category feature extraction network, the trained second-category feature extraction network, and the trained third-category feature extraction network; The training module is used to train the model parameters of the initial liveness detection model according to the face training sample set to obtain the liveness detection model.

9. A living body detection device, characterized in that: include: An acquisition module is used to acquire a facial image of a face to be detected; A detection module, configured to input the face image into a trained liveness detection model and output a detection result; The feature extraction model in the liveness detection model includes at least a first type of feature extraction network, a second type of feature extraction network, and a third type of feature extraction network; the probability of the first type of feature extraction network extracting texture features is greater than the probability of extracting edge features; the probability of the second type of feature extraction network extracting texture features is less than the probability of extracting edge features; the probability of the third type of feature extraction network extracting texture features is equal to the probability of extracting edge features; The feature extraction model is trained in the following manner: obtaining a first training image set, and obtaining a second training image set based on the first training image set; wherein, any second training image in the second training image set is synthesized based on two first training images in the first training image set; based on the first training image set, training the model parameters of the initial first type of feature extraction network to obtain the first type of feature extraction network, and training the model parameters of the initial second type of feature extraction network to obtain the second type of feature extraction network; based on the second training image set, training the model parameters of the initial third type of feature extraction network to obtain the third type of feature extraction network; and obtaining the feature extraction model based on the trained first type of feature extraction network, the second type of feature extraction network and the third type of feature extraction network.

10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the method according to any one of claims 1 to 6 and / or the method according to claim 7.

11. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 and / or the method according to claim 7 are implemented.

Citation Information

Patent Citations

  • Training method and device and detection method and device for human face living body detection model, and electronic equipment

    CN111597918A

  • Living body face detection model training method and device, apparatus and storage medium

    CN113052144A