A training method, apparatus, equipment, and medium for a liveness attack detection model.

CN115512446BActive Publication Date: 2026-05-26ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
Filing Date
2022-09-28
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing liveness detection models are not robust enough to effectively handle the changing attributes, backgrounds, and scenarios of people using facial recognition, leading to a decrease in detection accuracy.

Method used

By obtaining the intermediate data features of the intermediate residual network output between the first and last residual networks in the liveness attack detection model, and using them as input to the auxiliary attribute classifier, the first model recognition loss value is calculated. Based on this, the model parameters are adjusted, and multiple attribute recognition loss values ​​are fused to improve the robustness of the model.

Benefits of technology

This improves the robustness and accuracy of the liveness attack detection model, reduces the influence of external factors, and enhances the ability to identify liveness attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115512446B_ABST
    Figure CN115512446B_ABST
Patent Text Reader

Abstract

This specification discloses a training method, apparatus, device, and medium for a liveness attack detection model. The scheme may include: acquiring intermediate data features from the output of an intermediate residual network between the first and last residual networks in the liveness attack detection model; using the intermediate data features as input to an auxiliary attribute classifier to obtain a first model recognition loss value for the auxiliary attribute classifier; adjusting the parameters of the liveness attack detection model based on the first model recognition loss value; and obtaining a trained liveness attack detection model based on the adjusted liveness attack detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of model training technology, and in particular to a training method, apparatus, device and medium for a liveness attack detection model. Background Technology

[0002] With the continuous development of facial recognition systems in recent years, "liveness attack detection" has become an indispensable part of these systems. This step can effectively intercept non-liveness attack samples, including attacks using mobile phones, paper, and head models. As the volume of facial recognition services continues to increase, the number of people using facial recognition and the scenarios they are in are also increasing. This leads to a greater variety of attributes and combinations for those using facial recognition, as well as an increase in the types of backgrounds. This places higher demands on the robustness of liveness attack detection algorithms.

[0003] Therefore, how to provide a robust liveness attack detection model is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] This specification provides a training method, apparatus, device, and medium for a liveness attack detection model to address the problem of poor robustness in existing liveness attack detection models.

[0005] To solve the above-mentioned technical problems, the embodiments in this specification are implemented as follows:

[0006] This specification provides an embodiment of a training method for a liveness attack detection model, comprising:

[0007] Obtain intermediate data features from the intermediate residual network output between the first and last residual networks in the liveness attack detection model;

[0008] The intermediate data features are used as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier.

[0009] Based on the loss value identified by the first model, the parameters of the liveness attack detection model are adjusted;

[0010] Based on the adjusted liveness detection model, the trained liveness detection model is obtained.

[0011] This specification provides an embodiment of a training device for a liveness attack detection model, comprising:

[0012] The feature acquisition module is used to acquire intermediate data features from the intermediate residual network output between the first and last residual networks in the liveness attack detection model.

[0013] The recognition loss value determination module is used to take the intermediate data features as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier.

[0014] The parameter adjustment module is used to adjust the parameters of the liveness attack detection model based on the loss value identified by the first model.

[0015] The model training module is used to obtain a trained live attack detection model based on the adjusted live attack detection model.

[0016] This specification provides an embodiment of a training device for a liveness detection model, comprising:

[0017] At least one processor; and,

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:

[0020] Obtain intermediate data features from the intermediate residual network output between the first and last residual networks in the liveness attack detection model;

[0021] The intermediate data features are used as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier.

[0022] Based on the loss value identified by the first model, the parameters of the liveness attack detection model are adjusted;

[0023] Based on the adjusted liveness detection model, the trained liveness detection model is obtained.

[0024] This specification provides an embodiment of a computer-readable medium storing computer-readable instructions that can be executed by a processor to implement a training method for a liveness attack detection model.

[0025] One embodiment of this specification achieves the following beneficial effects: It obtains intermediate data features from the output of the intermediate residual network between the first and last residual networks in a liveness attack detection model; uses these intermediate data features as input to an auxiliary attribute classifier to obtain a first model recognition loss value for the auxiliary attribute classifier; adjusts the parameters of the liveness attack detection model based on the first model recognition loss value; and obtains a trained liveness attack detection model based on the adjusted liveness attack detection model. This results in a more robust liveness attack detection model and improves its detection accuracy. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a schematic diagram of the overall architecture of a training method for a liveness attack detection model provided in the embodiments of this specification;

[0028] Figure 2 This is a flowchart illustrating a training method for a liveness detection model provided in the embodiments of this specification.

[0029] Figure 3 This is a schematic diagram illustrating the training process of a liveness attack detection model provided in the embodiments of this specification;

[0030] Figure 4 This is a schematic diagram of the structure of a training device for a liveness detection model provided in the embodiments of this specification;

[0031] Figure 5 This is a schematic diagram of the structure of a training device for a liveness detection model provided in the embodiments of this specification. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of one or more embodiments of this specification clearer, the technical solutions of one or more embodiments of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of one or more embodiments of this specification.

[0033] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0034] In existing technologies, with the continuous development of facial recognition systems in recent years, "liveness attack detection" has become an indispensable part of facial recognition systems. This module can effectively intercept non-liveness attack samples, including mobile phone attacks, paper attacks, head models, etc. As the volume of facial recognition services continues to increase, there are more and more people using facial recognition and in more and more facial recognition scenarios. The types of attributes and combinations of people using facial recognition are also increasing, and the types of backgrounds are also increasing. This places higher demands on the robustness of liveness attack detection algorithms.

[0035] In existing technologies, the commonly used training method for liveness attack detection models is to input the features of the original training images into the liveness attack detection model, and to perform binary classification of liveness and attack based on the received data in an end-to-end manner, and then train the model based on the classification results.

[0036] To address the shortcomings of existing technologies, this solution provides the following embodiments:

[0037] Figure 1 This is a schematic diagram illustrating the overall architecture of a training method for a liveness detection model, as described in this specification, in a practical application scenario. Figure 1 As shown, the scheme mainly includes: intermediate data features 1, auxiliary attribute classifier 2, and liveness detection model 3. Intermediate data features can represent intermediate products obtained from a certain network structure layer of the liveness detection model during training. These can be data features output by the intermediate residual network between the first and last residual networks of the liveness detection model. The auxiliary attribute classifier can represent a classifier capable of classifying attributes related to liveness detection, such as face attributes, image background attributes, etc. In practical applications, intermediate data features 1 can be input into auxiliary attribute classifier 2 to obtain the first model recognition loss value of the auxiliary attribute classifier. Based on the first model recognition loss value, the parameters of the liveness detection model 3 are adjusted to obtain the trained liveness detection model. This improves the robustness of the liveness detection model, reducing the influence of some external factors during liveness detection and also improving the accuracy of liveness detection.

[0038] Next, the training method for a liveness detection model provided in the embodiments of the specification will be described in detail with reference to the accompanying drawings:

[0039] Figure 2This is a flowchart illustrating a training method for a liveness detection model provided in an embodiment of this specification. From a programming perspective, the execution entity of the process can be a program hosted on an application server or an application client.

[0040] like Figure 2 As shown, the process may include the following steps:

[0041] Step 202: Obtain intermediate data features from the output of the intermediate residual network between the first and last residual networks in the liveness attack detection model.

[0042] In the embodiments of this specification, the liveness attack detection model can be a convolutional neural network model, which may include multiple layers of convolutional neural networks. Specifically, it may include multiple layers of residual networks. Each residual network extracts features from the input image layer by layer, so that the liveness attack classifier can classify liveness and attacks, thus completing the liveness attack detection. The first layer of the residual network can represent the residual network closest to the model input. The last layer of the residual network can represent the residual network closest to the model output.

[0043] Step 204: Use the intermediate data features as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier.

[0044] In the embodiments of this specification, the auxiliary attribute classifier may include a model capable of classifying attributes related to liveness detection. Attributes related to liveness detection may include at least one of attributes such as facial expression, age, lighting conditions, and mobile phone attack. The first model recognition loss value can be calculated using existing loss functions or other methods for calculating model loss during model training.

[0045] In practical applications, the auxiliary attribute classifier used to classify attributes related to liveness detection can be the same or multiple. Furthermore, the auxiliary attribute classifier can be part of the liveness attack detection model itself or a separate auxiliary attribute classifier. Auxiliary classifiers can include simple classifiers or trained classification models, etc., and the specific form is not specifically limited here.

[0046] Step 206: Based on the loss value identified by the first model, adjust the parameters of the liveness attack detection model.

[0047] In practical applications, training a neural network model can mean adjusting the parameters used by the model during business processing, enabling the model to better perform business tasks. In this embodiment, the parameters can be those used by the model during computation, such as weights used in image feature extraction. The liveness detection model can adjust its parameters based on the first model's recognition loss value, ensuring the adjusted model meets preset conditions, such as not exceeding a preset threshold for the recognition loss value, thereby improving its robustness.

[0048] Step 208: Based on the adjusted liveness attack detection model, obtain the trained liveness attack detection model.

[0049] It should be understood that the order of some steps in the methods described in one or more embodiments of this specification may be interchanged according to actual needs, or some steps may be omitted or deleted.

[0050] Figure 2 The method described herein involves obtaining intermediate data features from the output of the intermediate residual network between the first and last residual networks in a liveness attack detection model; using these intermediate data features as input to an auxiliary attribute classifier to obtain a first model recognition loss value for the auxiliary attribute classifier; adjusting the parameters of the liveness attack detection model based on the first model recognition loss value; and obtaining a trained liveness attack detection model based on the adjusted liveness attack detection model. This improves the robustness of the liveness attack detection model, reducing the influence of some external factors during liveness attack detection, and simultaneously improving the accuracy of the liveness attack detection model in detecting liveness attacks.

[0051] based on Figure 2 In addition to the method described herein, this specification also provides some specific implementation schemes of the method, which will be described below.

[0052] Optionally, in the embodiments of this specification, obtaining the intermediate data features of the intermediate residual network output between the first and last residual networks in the liveness attack detection model may specifically include:

[0053] Obtain training data;

[0054] The training data is input into the liveness attack detection model to train the liveness attack detection model;

[0055] The intermediate data features of the intermediate residual network between the first and last residual networks in the liveness attack detection model are obtained based on the training data input to the liveness attack detection model.

[0056] The training data obtained in the embodiments of this specification can be existing historical liveness detection data. For example, it can include historical recognition images that have been recognized by the initial liveness detection model, historical images that have been recognized by other liveness detection models, images obtained by processing actual liveness detection images through image processing models, adversarial images generated based on generative adversarial networks, and so on.

[0057] In practical applications, a network model can include multiple network modules. The training data input into the liveness attack detection model can be processed by each network module to obtain processed data features. After processing, each network module can transmit the processed data to the next network module for further processing. In the embodiments of this specification, the liveness attack detection model can include multiple residual networks. Training data can pass through the feature data of each residual network layer by layer, starting from the input layer. The output features of one residual network can be used as the input of the next adjacent residual network, which can then perform further feature processing. The intermediate data features can represent the data features output by the intermediate residual networks between the first and last residual networks of the liveness attack detection model.

[0058] In order to train the liveness attack detection model more effectively and improve its robustness while ensuring the accuracy of liveness attack detection, the liveness attack detection model described in the embodiments of this specification may include a liveness attribute classifier.

[0059] The method may further include:

[0060] Obtain the second model recognition loss value obtained by the liveness attribute classifier based on the end data features output by the end residual network;

[0061] The step of adjusting the parameters of the liveness detection model based on the loss value identified by the first model may specifically include:

[0062] The first model recognition loss value and the second model recognition loss value are fused to obtain the fused loss value;

[0063] Based on the fusion loss value, the parameters of the liveness attack detection model are adjusted.

[0064] In the embodiments of this specification, the second model recognition loss value obtained from the end data features output by the end residual network can include the loss value of the liveness attribute. The liveness attribute classifier can detect whether the data input to the model is live or inactive based on the end data features. Inactive data can also be understood as an attack. The loss function is then used to calculate the loss value for liveness recognition. In the embodiments of this specification, the second model recognition loss value representing the accuracy of liveness attack recognition can be fused with the first model recognition loss value representing the accuracy of auxiliary attribute recognition, which was previously calculated using the loss function, to obtain a fused loss value. The parameters of the liveness attack detection model are adjusted based on the fused loss value. In practical applications, the model can be trained iteratively, specifically by performing multiple rounds of training according to the above training steps. When the fused loss value of the trained liveness attack detection model meets the preset loss value conditions, the training can be terminated, and a liveness attack detection model that meets the requirements can be obtained.

[0065] To obtain a fused loss value of the first model recognition loss value and the second model recognition loss value, the fusion of the first model recognition loss value and the second model recognition loss value as described in the embodiments of this specification may specifically include:

[0066] The loss values ​​of the first model and the second model are weighted and summed.

[0067] In the embodiments of this specification, the first model identification loss value and the second model identification loss value are weighted and summed to obtain the fusion loss value. Different weights can be assigned based on the implementation of the main function of the liveness attack detection model; however, these weights can be similar, or the weights of each attribute identification loss value can be assigned the same weight. For example, since the liveness attack detection model is mainly for detecting liveness attributes, the weight of the second model identification loss value can be set to 0.6, and the weight of the first model identification loss value can be set to 0.4 to calculate the fusion loss value. Alternatively, the weights of both the second and first model identification loss values ​​can be set to 0.5 to calculate the fusion loss value.

[0068] The auxiliary attribute classifier described in the embodiments of this specification may include at least one of a face attribute classifier, a background attribute classifier, and an attack attribute classifier; the face attribute classifier may be used to identify at least one of gender, age, expression, glasses, and hairstyle; the background attribute classifier may be used to identify at least one of light intensity, environment, and number of people; and the attack attribute classifier may be used to identify at least one of mobile phone attack, paper attack, high-definition screen attack, and head model attack.

[0069] The attribute classifier used in the embodiments of this specification can be a binary attribute classifier. For example, a binary facial attribute classifier can be used to identify age, and the identification result can be young person, elderly person, etc.; the identification result for facial expression can be expressionless, expressiony, etc. The facial expression identification result can also be a specific facial expression type identification result, such as happy, sad, afraid, angry, disgusted, surprised, contemptuous, etc. A binary background attribute classifier can be used to identify the environment, and the identification result can be indoor, outdoor, etc.; the identification result for light intensity can be strong light, weak light, etc.; the identification result for the number of people can be whether there are multiple people, etc. A mobile phone attack can mean that the object of the liveness detection is an image displayed on a mobile phone screen, a paper attack can mean that the object of the liveness detection is an image existing in the form of paper, a high-definition screen attack can mean that the object of the liveness detection is an image displayed on a high-definition screen, and a head model attack can mean that the object of the liveness detection is an object wearing a head model. Among them, the attack attribute classifier can identify whether the object to be identified is a mobile phone attack, a paper attack, a high-definition screen attack, a head model attack, etc.

[0070] In practical applications, facial attributes, background attributes, and attack attributes can be further subdivided. For example, a facial attribute classifier can classify expressions into specific emotions such as calm, smiling, laughing, and crying. A background attribute classifier can classify lighting intensity into several scenarios, including no lighting, 25% lighting, 50% lighting, 75% lighting, and 100% lighting. An attack attribute classifier can classify attacks based on the phone model. The auxiliary attribute classifier can further subdivide different attributes into several categories, which can be selected based on the user's specific application needs for the liveness detection model; no specific limitations are imposed here.

[0071] The auxiliary attribute classifier described in the embodiments of this specification may include a face attribute classifier;

[0072] The step of using the intermediate data features as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier may specifically include:

[0073] The first intermediate data feature is used as the input to the face attribute classifier to obtain the face attribute recognition loss value of the face attribute classifier; the first intermediate data feature is the data feature output by the first intermediate residual network between the first layer residual network and the last layer residual network in the liveness attack detection model.

[0074] The face attribute classifier in this embodiment requires detailed face attribute feature information to classify face attributes. The first intermediate residual network can be a residual network layer relatively close to the input. In the liveness detection model, this residual network can include fewer other residual networks before it, and the output first intermediate data features can retain more image features. The face attribute classifier can classify preset face attributes based on the output first intermediate data features, and then calculate the loss value of face attribute recognition through a loss function. The preset face attributes can include at least one of gender, age, expression, glasses, and hairstyle.

[0075] The auxiliary attribute classifier described in this specification may include a background attribute classifier;

[0076] The step of using the intermediate data features as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier may specifically include:

[0077] The second intermediate data feature is used as the input to the background attribute classifier to obtain the background attribute recognition loss value of the background attribute classifier; the second intermediate data feature is the data feature output by the second intermediate residual network between the first layer residual network and the last layer residual network in the liveness attack detection model.

[0078] In the embodiments described in the specification, the background attribute classifier may require detailed background feature information to complete the identification of background attributes. Therefore, the second intermediate residual network can be a residual network layer closer to the input. In the liveness detection model, this residual network may include fewer other residual networks before it, and the output second intermediate data features can retain more image features. The background attribute classifier can identify the background of the data based on the second intermediate data features, and then calculate the identification result using a loss function to obtain the background attribute loss value. Background attributes may include at least one of the following: illumination intensity, environment, and number of people.

[0079] The auxiliary attribute classifier described in the embodiments of this specification may include an attack attribute classifier;

[0080] The step of using the intermediate data features as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier may specifically include:

[0081] The third intermediate data feature is used as the input to the attack attribute classifier to obtain the attack attribute identification loss value of the attack attribute classifier; the third intermediate data feature is the data feature output by the third intermediate residual network between the first layer residual network and the last layer residual network in the liveness attack detection model.

[0082] The attack attributes in the embodiments of this specification may include at least one of the following: mobile phone attack, paper attack, high-definition screen attack, and head model attack. The attack attributes identified by the attack attribute classifier are obvious features. To reduce data transmission and recognition, and improve the training efficiency of the liveness detection model, the third intermediate residual network can be a residual network layer relatively close to the output. In the liveness detection model, this residual network may be preceded by many other residual networks, and the output third intermediate data features can retain fewer image features. The attack attribute classifier can identify attacks based on the fewer but obvious features contained in the third intermediate data features.

[0083] In the embodiments of this specification, the liveness detection model can be trained based on intermediate feature data obtained from intermediate residual networks of different layers. Specifically:

[0084] The auxiliary attribute classifier may include a first auxiliary attribute classifier for performing face attribute recognition and / or background attribute recognition and a second auxiliary attribute classifier for performing attack attribute recognition;

[0085] The acquisition of intermediate data features from the intermediate residual network output between the first and last residual networks in the liveness attack detection model may specifically include:

[0086] Obtain the fourth intermediate data feature output by the fourth intermediate residual network between the first and last residual networks in the liveness attack detection model;

[0087] Obtain the fifth intermediate data feature output by the fifth intermediate residual network between the first and last layer residual networks in the liveness attack detection model; the fifth intermediate residual network is the residual network located after the fourth intermediate residual network in the liveness attack detection model.

[0088] The step of using the intermediate data features as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier may specifically include:

[0089] The fourth intermediate data feature is used as the input to the first auxiliary attribute classifier to obtain the first auxiliary recognition loss value of the first auxiliary attribute classifier.

[0090] The fifth intermediate data feature is used as input to the second auxiliary attribute classifier to obtain the second auxiliary recognition loss value of the second auxiliary attribute classifier.

[0091] This specification describes an example where the first auxiliary attribute classifier used for face attribute recognition and / or background attribute recognition requires a large number of data features for accurate identification. The fourth intermediate residual network can be a residual network layer closer to the input. In the liveness detection model, this residual network can include fewer other residual networks before it, and the output fourth intermediate data features can retain more image features. The fourth intermediate data features can be used to recognize both face attributes and background attributes. The second auxiliary attribute classifier used for attack attribute recognition can identify attack attributes using only a few obvious features. To accelerate data transmission and recognition and improve the training efficiency of the liveness detection model, the fourth intermediate data features can be used as input to the fifth intermediate residual network for convolutional neural network processing to obtain the fifth intermediate data features. The fifth intermediate data features have fewer features and a smaller data volume than the fourth intermediate data features, but the features carrying attack attributes are more obvious and easier to identify. Each recognition loss value can be calculated using a loss function based on the recognition results of each classifier. This can be understood as the fourth intermediate data feature undergoing fewer convolution operations compared to the fifth intermediate data feature, and thus retaining more data features compared to the fifth intermediate data feature.

[0092] In practical applications, expert experience can be used to determine which specific residual network layer in the liveness attack detection model should be used as the intermediate residual network. Alternatively, the intermediate residual network can be determined based on the data features required for the network to learn the model based on the auxiliary classifier.

[0093] In the embodiments of this specification, the first auxiliary recognition loss value may include the face attribute recognition loss value corresponding to the face attribute classifier used for face attribute recognition and / or the background attribute recognition loss value corresponding to the background attribute classifier used for background attribute recognition; the second auxiliary recognition loss value may include the attack attribute recognition loss value corresponding to the attack attribute classifier used for attack attribute recognition.

[0094] The step of adjusting the parameters of the liveness detection model based on the loss value identified by the first model may specifically include:

[0095] The face attribute recognition loss value and / or the background attribute recognition loss value and the attack attribute recognition loss value are fused with the liveness attribute recognition loss value to obtain the fused model recognition loss value; the liveness attribute recognition loss value is the model loss value obtained based on the liveness attribute classifier in the liveness attack detection model;

[0096] The fused model identification loss value is backpropagated to the liveness attack detection model to adjust the parameters of the liveness attack detection model.

[0097] In the embodiments of this specification, the backpropagation of the loss value can be achieved by backpropagating the loss value obtained by fusing the face attribute recognition loss value, background attribute recognition loss value, attack attribute recognition loss value and liveness attribute recognition loss value to the liveness attack detection model for parameter adjustment.

[0098] In practical applications, backpropagation of the loss value can also represent backpropagation starting from the intermediate residual network. Specifically, the fused model recognition loss value obtained by fusing the face attribute recognition loss value and / or the background attribute recognition loss value with the liveness attribute recognition loss value can be backpropagated to the fourth intermediate residual network. This is used to adjust the parameters of the residual network between the fourth intermediate residual network and the first layer residual network (including the fourth intermediate residual network and the first layer residual network) through backpropagation. Alternatively, the fused model recognition loss value obtained by fusing the attack attribute recognition loss value and the liveness attribute recognition loss value can be backpropagated to the fifth intermediate residual network. This is used to adjust the parameters of the residual network between the fifth intermediate residual network and the first layer residual network (including the fifth intermediate residual network and the first layer residual network) through backpropagation. Furthermore, the liveness attribute recognition loss value can be backpropagated to the last layer residual network. This is used to adjust the parameters of the residual network between the last layer residual network and the first layer residual network (including the last layer residual network and the first layer residual network) through backpropagation. The specific method for backpropagating the loss value is not specifically limited. Furthermore, the first, second, third, fourth, and fifth intermediate residual networks can be the same or different intermediate residual networks; no specific limitations are imposed here.

[0099] In practical applications, the loss values ​​for face attribute recognition, background attribute recognition, and attack attribute recognition are fused with the loss value for liveness attribute recognition to obtain a fused loss value. This fused loss value is then backpropagated to the liveness attack detection model to adjust its parameters, such as the weights used in image feature extraction. The weights used in image feature extraction, i.e., the weights of each loss value, can be the same or different, but they should be similar and not significantly different. For example, the fused loss value might be calculated as follows: face attribute recognition loss value 0.05, weight 0.2; background attribute recognition loss value 0.06, weight 0.2; attack attribute recognition loss value 0.04, weight 0.25; liveness attribute recognition loss value 0.01, weight 0.35. Since the liveness attack detection model primarily aims to improve the robustness of liveness detection, the weights of the liveness attribute recognition loss value and the attack attribute recognition loss value can be appropriately increased, but not higher than the weight of the liveness attribute recognition loss value. (Fused Loss Value)

[0100] = 0.05*0.2 + 0.06*0.2 + 0.04*0.25 + 0.01*0.35. When the weights are the same, a weighted calculation is performed to fuse the loss values.

[0101] =0.05*0.25+0.06*0.25+0.04*0.25+0.01*0.25.

[0102] In practical applications, joint optimization using multiple attribute classifications and losses enables the model to adaptively focus on sensitive regions for different inputs, further improving the robustness of liveness attack detection. For example, if glasses are detected on a face, the eyes can be de-emphasized as the primary feature for liveness detection, reducing the weight of the eye region and instead selecting features such as the nose or lips as the liveness detection features.

[0103] The training data described in the embodiments of this specification may include at least one of face image data and face video data.

[0104] In practical applications, training data can be not only image data but also video data. When the training data is video data, temporal sub-classification can be introduced as auxiliary recognition attribute information, which will further contribute to robust feature extraction.

[0105] The training data described in the embodiments of this specification may include a first label for representing liveness attributes and a second label for representing auxiliary recognition attribute information.

[0106] In practical applications, the first label can indicate whether the training data is live; for example, a live object can be marked as 1, and an attack can be marked as 2. The second label can represent facial attributes, background attributes, attack attributes, etc. For example, the second label can represent the gender, age, expression, glasses, hairstyle, etc., corresponding to the training data. It can also represent the corresponding light intensity, environment, number of people, etc., and can also represent attack types such as mobile phone attacks, paper attacks, high-definition screen attacks, and head model attacks.

[0107] The robustness of the obtained liveness attack detection model is improved by the above method, which can more accurately identify liveness and reduce the influence of some unnecessary factors.

[0108] Based on the above training methods Figure 3 This specification also provides a schematic diagram illustrating the training process of a liveness attack detection model as an embodiment. For example... Figure 3 As shown, the training process may include:

[0109] Step 302: Input the training data into the liveness attack detection model.

[0110] The liveness detection model may include a multi-layer residual network, as illustrated in the embodiments of this specification. Figure 3 In this context, STM can represent the data input layer of the liveness attack detection model, Res Block1 can represent the first layer residual network, Res Block2 and Res Block3 can represent the intermediate residual networks, and Res Block4 can represent the last layer residual network.

[0111] Step 304: Use the intermediate data features output by the intermediate residual network as input to the face attribute classifier, and then obtain the face attribute recognition loss value based on the classification result.

[0112] In practical applications, when facial attributes include multiple attributes such as gender and age, the recognition loss value corresponding to each attribute can be calculated separately, and then the loss values ​​can be fused together using methods such as weighted summation to obtain the fusion loss. The aforementioned facial attribute recognition loss value can represent the facial attribute fusion loss value obtained after fusing the loss values ​​corresponding to each facial attribute.

[0113] Step 306: Use the intermediate data features output by the intermediate residual network as input to the background attribute classifier, and then obtain the background attribute recognition loss value based on the classification result.

[0114] The inputs to the face attribute classifier and the background attribute classifier can be intermediate data features obtained from the intermediate residual network of the same layer. For example... Figure 3 As shown, the output features of Res Block2 can be selected as inputs to the face attribute classifier and the background attribute classifier. It is understood that in practical applications, intermediate residual networks from different layers can also be selected; no specific limitations are made here.

[0115] In practical applications, when background attributes include multiple attributes such as lighting, environment, and number of people, the recognition loss value corresponding to each attribute can be calculated separately, and then the loss values ​​can be fused together using methods such as weighted summation to obtain the fusion loss. The aforementioned background attribute recognition loss value can represent the background attribute fusion loss value obtained after fusing the loss values ​​corresponding to each background attribute.

[0116] Step 308: Use the intermediate data features output by the intermediate residual network as input to the attack attribute classifier, and then obtain the attack attribute identification loss value based on the classification result.

[0117] In this context, the intermediate residual network whose output serves as the input intermediate data features for the attack attribute classifier can be a residual network located after the intermediate residual network whose output serves as the input intermediate data features for the face attribute classifier. For example... Figure 3As shown, the output features of Res Block3 can be selected as inputs to both the attack attribute classifier and the background attribute classifier. It is understood that in practical applications, the same intermediate residual network can also be chosen; no specific limitations are made here.

[0118] In practical applications, when attack attributes include multiple attributes such as mobile phone attack, paper attack, high-definition screen attack, and head model attack, the recognition loss value corresponding to each attribute can be calculated separately. Then, the loss values ​​can be fused together using methods such as weighted summation to obtain the fusion loss. The aforementioned attack attribute recognition loss value can represent the attack attribute fusion loss value obtained after fusing the loss values ​​corresponding to each attack attribute.

[0119] Step 310: Use the data features output by the last residual network as input to the liveness attack classifier to obtain the classification result; then, the main loss value for liveness attack identification can be obtained based on the classification result.

[0120] Step 312: Fuse the obtained loss values ​​to obtain the fused loss value of the liveness attack detection model; backpropagate the obtained fused loss value to the liveness attack detection model to obtain the trained liveness attack detection model.

[0121] The liveness attack loss value obtained by fusing the various loss values ​​is backpropagated to the liveness attack detection model. The liveness attack detection model adjusts its parameters according to the liveness attack loss value to obtain the trained liveness attack detection model.

[0122] Based on the same idea, this specification also provides a schematic diagram of the structure of a training device for a liveness attack detection model. For example... Figure 4 As shown, the device may include:

[0123] The feature acquisition module 402 is used to acquire intermediate data features from the intermediate residual network output between the first-layer residual network and the last-layer residual network in the liveness attack detection model.

[0124] The identification loss value determination module 404 is used to take the intermediate data features as input to the auxiliary attribute classifier to obtain the first model identification loss value of the auxiliary attribute classifier.

[0125] The parameter adjustment module 406 is used to adjust the parameters of the liveness attack detection model based on the loss value identified by the first model.

[0126] The model training module 408 is used to obtain a trained live attack detection model based on the adjusted live attack detection model.

[0127] Based on the same idea, this specification also provides a schematic diagram of the structure of a training device for a liveness detection model. For example... Figure 5 As shown, device 500 may include:

[0128] At least one processor 510; and,

[0129] Memory 530 communicatively connected to the at least one processor; wherein,

[0130] The memory 530 stores instructions 520 that can be executed by the at least one processor 510, the instructions being executed by the at least one processor 510 to enable the at least one processor 510 to:

[0131] Obtain intermediate data features from the intermediate residual network output between the first and last residual networks in the liveness attack detection model;

[0132] The intermediate data features are used as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier.

[0133] Based on the loss value identified by the first model, the parameters of the liveness attack detection model are adjusted;

[0134] Based on the adjusted liveness detection model, the trained liveness detection model is obtained.

[0135] Based on the same approach, embodiments of this specification also provide a computer-readable medium corresponding to the above-described method. The computer-readable medium stores computer-readable instructions that can be executed by a processor to implement the training method for the above-described liveness detection model.

[0136] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, for... Figure 5 As the device shown is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0137] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0138] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0139] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0140] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0141] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0142] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0143] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0144] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0145] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0146] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0147] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital character versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0148] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0149] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0150] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0151] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A training method for a liveness attack detection model, comprising: The intermediate residual network between the first and last residual networks in the liveness attack detection model is obtained based on the intermediate data features output by the training data input to the liveness attack detection model; the training data includes a first label representing the liveness attribute and a second label representing auxiliary recognition attribute information; The intermediate data features are used as input to an auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier; the auxiliary attribute classifier includes at least one of a face attribute classifier, a background attribute classifier, and an attack attribute classifier; the face attribute classifier is used to identify at least one attribute of gender, age, expression, glasses, and hairstyle; the background attribute classifier is used to identify at least one attribute of light intensity, environment, and number of people; the attack attribute classifier is used to identify at least one attribute of mobile phone attack, paper attack, high-definition screen attack, and head model attack; Based on the loss value identified by the first model, the parameters of the liveness attack detection model are adjusted; Based on the adjusted liveness detection model, the trained liveness detection model is obtained.

2. The method according to claim 1, wherein obtaining the intermediate residual network between the first and last residual networks in the liveness attack detection model is based on the intermediate data features output from the training data input to the liveness attack detection model, specifically includes: Obtain training data; The training data is input into the liveness attack detection model to train the liveness attack detection model; The intermediate data features of the intermediate residual network between the first and last residual networks in the liveness attack detection model are obtained based on the training data input to the liveness attack detection model.

3. The method according to claim 1, wherein the liveness attack detection model includes a liveness attribute classifier; The method further includes: Obtain the second model recognition loss value obtained by the liveness attribute classifier based on the end data features output by the last layer residual network; The step of adjusting the parameters of the liveness attack detection model based on the loss value identified by the first model specifically includes: The first model recognition loss value and the second model recognition loss value are fused to obtain the fused loss value; Based on the fusion loss value, the parameters of the liveness attack detection model are adjusted.

4. The method according to claim 3, wherein fusing the first model recognition loss value and the second model recognition loss value specifically includes: The loss values ​​of the first model and the second model are weighted and summed.

5. The method according to claim 1, wherein the auxiliary attribute classifier includes a face attribute classifier; The step of using the intermediate data features as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier specifically includes: The first intermediate data features are used as input to the face attribute classifier to obtain the face attribute recognition loss value of the face attribute classifier. The first intermediate data feature is the data feature output by the first intermediate residual network between the first layer residual network and the last layer residual network in the liveness attack detection model.

6. The method according to claim 1, wherein the auxiliary attribute classifier includes a background attribute classifier; The step of using the intermediate data features as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier specifically includes: The second intermediate data feature is used as the input to the background attribute classifier to obtain the background attribute recognition loss value of the background attribute classifier. The second intermediate data feature is the data feature output by the second intermediate residual network between the first-layer residual network and the last-layer residual network in the liveness attack detection model.

7. The method according to claim 1, wherein the auxiliary attribute classifier includes an attack attribute classifier; The step of using the intermediate data features as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier specifically includes: The third intermediate data feature is used as the input to the attack attribute classifier to obtain the attack attribute identification loss value of the attack attribute classifier; The third intermediate data feature is the data feature output by the third intermediate residual network between the first-layer residual network and the last-layer residual network in the liveness attack detection model.

8. The method according to claim 1, further comprising: The auxiliary attribute classifier includes a first auxiliary attribute classifier for performing face attribute recognition and / or background attribute recognition, and a second auxiliary attribute classifier for performing attack attribute recognition; The acquisition of intermediate data features from the intermediate residual network output between the first and last residual networks in the liveness attack detection model specifically includes: Obtain the fourth intermediate data feature output by the fourth intermediate residual network between the first and last layer residual networks in the liveness attack detection model; Obtain the fifth intermediate data feature output by the fifth intermediate residual network between the first and last layer residual networks in the liveness attack detection model; the fifth intermediate residual network is the residual network located after the fourth intermediate residual network in the liveness attack detection model. The step of using the intermediate data features as input to the auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier specifically includes: The fourth intermediate data feature is used as the input to the first auxiliary attribute classifier to obtain the first auxiliary recognition loss value of the first auxiliary attribute classifier. The fifth intermediate data feature is used as input to the second auxiliary attribute classifier to obtain the second auxiliary recognition loss value of the second auxiliary attribute classifier.

9. The method according to claim 8, wherein the first auxiliary recognition loss value includes the face attribute recognition loss value corresponding to the face attribute classifier used for face attribute recognition and / or the background attribute recognition loss value corresponding to the background attribute classifier used for background attribute recognition; the second auxiliary recognition loss value includes the attack attribute recognition loss value corresponding to the attack attribute classifier used for attack attribute recognition; The step of adjusting the parameters of the liveness attack detection model based on the loss value identified by the first model specifically includes: The face attribute recognition loss value and / or the background attribute recognition loss value and the attack attribute recognition loss value are fused with the liveness attribute recognition loss value to obtain the fused model recognition loss value; the liveness attribute recognition loss value is the model loss value obtained based on the liveness attribute classifier in the liveness attack detection model; The fused model identification loss value is backpropagated to the liveness attack detection model to adjust the parameters of the liveness attack detection model.

10. The method according to claim 1, wherein the training data includes at least one of face image data and face video data.

11. A training device for a liveness attack detection model, comprising: The feature acquisition module is used to acquire intermediate data features of the intermediate residual network between the first and last residual networks in the liveness attack detection model, based on the training data input to the liveness attack detection model; the training data includes a first label representing the liveness attribute and a second label representing auxiliary recognition attribute information; The recognition loss value determination module is used to take the intermediate data features as input to an auxiliary attribute classifier to obtain a first model recognition loss value for the auxiliary attribute classifier; the auxiliary attribute classifier includes at least one of a face attribute classifier, a background attribute classifier, and an attack attribute classifier; the face attribute classifier is used to recognize at least one attribute of gender, age, expression, glasses, and hairstyle; the background attribute classifier is used to recognize at least one attribute of light intensity, environment, and number of people; the attack attribute classifier is used to recognize at least one attribute of mobile phone attack, paper attack, high-definition screen attack, and head model attack; The parameter adjustment module is used to adjust the parameters of the liveness attack detection model based on the loss value identified by the first model. The model training module is used to obtain a trained live attack detection model based on the adjusted live attack detection model.

12. A training device for a liveness attack detection model, comprising: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: The intermediate residual network between the first and last residual networks in the liveness attack detection model is obtained based on the intermediate data features output by the training data input to the liveness attack detection model; the training data includes a first label representing the liveness attribute and a second label representing auxiliary recognition attribute information; The intermediate data features are used as input to an auxiliary attribute classifier to obtain the first model recognition loss value of the auxiliary attribute classifier; the auxiliary attribute classifier includes at least one of a face attribute classifier, a background attribute classifier, and an attack attribute classifier; the face attribute classifier is used to identify at least one attribute of gender, age, expression, glasses, and hairstyle; the background attribute classifier is used to identify at least one attribute of light intensity, environment, and number of people; the attack attribute classifier is used to identify at least one attribute of mobile phone attack, paper attack, high-definition screen attack, and head model attack; Based on the loss value identified by the first model, the parameters of the liveness attack detection model are adjusted; Based on the adjusted liveness detection model, the trained liveness detection model is obtained.

13. A computer-readable medium having stored thereon computer-readable instructions that can be executed by a processor to implement the training method of the liveness attack detection model according to any one of claims 1 to 10.