Training Method of Live Detection Model, Live Detection Method and Related Devices

By training the field discrimination loss of image pairs, the parameters of the live detection model are adjusted, and the problem of inaccurate live detection results in the prior art is solved, and the accuracy and generalization ability of live detection are improved.

CN116052289BActive Publication Date: 2025-07-18ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310030053.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-05
Publication Date
2025-07-18
Estimated Expiration
2043-01-05

AI Technical Summary

Technical Problem

The results of the live body detection method obtained by the prior art in vivo detection are not accurate enough, and there is a risk of camouflage target attack.

Method used

By acquiring the training image pairs, the first and second feature extraction branches of the live detection model extract the features of the first and second mode images, the domain discrimination is performed based on the domain discriminative features, the domain discrimination loss is constructed, and the parameters of the live detection model are adjusted to improve the domain generalization ability of the live detection model.

Benefits of technology

The accuracy of the in vivo detection model based on characteristics during the application stage is improved, the attention to the relevant information of the field is reduced, and the accuracy of the in vivo detection is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052289B_ABST
    Figure CN116052289B_ABST
Patent Text Reader

Abstract

The present application discloses a training method for a live detection model, a live detection method, an electronic device, and a computer-readable storage medium. The training method includes: obtaining a training image pair; extracting first-modal image features using a first feature extraction branch of the live detection model to obtain a first feature, where the first feature includes a first domain discrimination feature; extracting second-modal image features using a second feature extraction branch of the live detection model to obtain a second feature, where the second feature includes a second domain discrimination feature; performing domain discrimination based on the first domain discrimination feature and / or the second domain discrimination feature to obtain a domain discrimination result; constructing a domain discrimination loss based on the difference between the domain discrimination result and the domain label; and adjusting the parameters of the live detection model at least based on the domain discrimination loss. Through the above method, the accuracy of the live discrimination result based on features in the application stage can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of target recognition, and in particular, to a method for training a live detection model, a live detection method, an electronic device, and a computer-readable storage medium. Background Art

[0002] Target recognition technology can be applied to many fields such as turnstile access control, work attendance machines, and mobile phone face recognition payment. The general process of target recognition technology is to obtain an image of the target and perform recognition based on the image of the target. However, target recognition technology has the risk of being attacked by a disguised target, that is, the obtained image of the target is not an image taken of a real target (a live body), but an image taken of a disguised target (a non-live body). The disguised target is, for example, a paper photo of the target, an electronic image of the target, a three-dimensional mold, and so on.

[0003] Based on this, it is necessary to perform live detection on the image of the target during the target recognition process to determine whether the target is a live body. However, the live detection methods in the prior art have inaccurate live discrimination results. Summary of the Invention

[0004] The present application provides a method for training a live detection model, a live detection method, an electronic device, and a computer-readable storage medium, which can solve the problem that the live discrimination results obtained by the live detection methods in the prior art are inaccurate.

[0005] To solve the above technical problems, a technical solution adopted by the present application is: to provide a method for training a live detection model. The training method includes: obtaining a training image pair, the training image pair including a first-modal image and a second-modal image, and a domain label characterizing the domain to which the first-modal image and / or the second-modal image belong, the first-modal image and the second-modal image being from a training object of the same live category; using a first feature extraction branch of the live detection model to perform feature extraction on the first-modal image to obtain a first feature, the first feature including a first domain discrimination feature; using a second feature extraction branch of the live detection model to perform feature extraction on the second-modal image to obtain a second feature, the second feature including a second domain discrimination feature; performing domain discrimination based on the first domain discrimination feature and / or the second domain discrimination feature to obtain a domain discrimination result; constructing a domain discrimination loss based on the difference between the domain discrimination result and the domain label; and adjusting the parameters of the live detection model based at least on the domain discrimination loss.

[0006] To solve the above technical problems, a technical solution adopted in this application is: to provide a living body detection method. The living body detection method includes: obtaining an image of an object to be detected, where the image of the object to be detected includes a to-be-detected image in a first modality and / or a second modality; using a living body detection model to extract features from the image of the object to be detected to obtain living body discrimination features; obtaining a living body discrimination result of the object to be detected based on the living body discrimination features; where the living body detection model is trained by the aforementioned training method.

[0007] To solve the above technical problems, another technical solution adopted in this application is: to provide an electronic device, which includes a processor and a memory connected to the processor, where the memory stores program instructions; the processor is configured to execute the program instructions stored in the memory to implement the above method.

[0008] To solve the above technical problems, yet another technical solution adopted in this application is: to provide a computer-readable storage medium storing program instructions, and when the program instructions are executed, they can implement the above method.

[0009] In the above manner, this application performs domain discrimination based on the first domain discrimination features and / or the second domain discrimination features to obtain a domain discrimination result; constructs a domain discrimination loss based on the difference between the domain discrimination result and the domain label, and adjusts the parameters of the living body detection model based on the domain discrimination loss, thereby realizing the training of the living body detection model. Among them, since the greater the difference between the domain discrimination result and the domain label, the smaller the domain discrimination loss, therefore, by training the living body detection model through the domain discrimination loss, the parameters of the living body detection model can be adjusted through the domain discrimination loss, the domain generalization ability of the living body detection model can be improved, that is, the attention of the living body detection model to domain-related information during feature extraction can be reduced, thereby improving the expression ability of the features extracted by the living body detection model for the living body category, and further improving the accuracy of the living body discrimination result obtained based on the features in the application stage. Description of the Drawings

[0010] Figure 1 It is a schematic flowchart of an embodiment of the training method of the living body detection model in this application;

[0011] Figure 2 It is a schematic flowchart of another embodiment of the training method of the living body detection model in this application;

[0012] Figure 3 It is a schematic flowchart of yet another embodiment of the training method of the living body detection model in this application;

[0013] Figure 4 It is a schematic structural diagram of the training of the living body detection model;

[0014] Figure 5It is a schematic diagram of the processing of central difference convolution;

[0015] Figure 6 It is a schematic flowchart of an embodiment of the living body detection method of the present application;

[0016] Figure 7 It is a schematic structural diagram of an embodiment of the electronic device of the present application;

[0017] Figure 8 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners

[0018] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0019] The terms "first", "second", and "third" in the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., and the meaning of "several" is one or more, unless otherwise clearly and specifically defined.

[0020] Referring to "embodiment" herein means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that, without conflict, the embodiments described herein may be combined with other embodiments.

[0021] Figure 1 It is a schematic flowchart of an embodiment of the training method of the living body detection model of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 1 the process sequence shown. As Figure 1 shown, this embodiment may include:

[0022] S11: Obtain training image pairs.

[0023] The training image pair includes a first-modal image and a second-modal image, as well as a domain label characterizing the domain to which the first-modal image and / or the second-modal image belong. The first-modal image and the second-modal image are from training objects of the same living body category.

[0024] The execution subject of the training method embodiment of this application is a training device, and the training device can be an electronic device in the form of a computer, a mobile phone, a server, etc.

[0025] The training object can be a living body, such as a person or an animal (such as a dog), the face of a person or an animal; the training object can also be a non-living body, such as a paper photo, an electronic image, a three-dimensional mold, etc. that disguises as a living body. The image modalities available for living body detection include but are not limited to the near-infrared modality, the visible light modality, and the depth modality. The first modality to which the first-modal image belongs can be one of the near-infrared modality, the visible light modality, and the depth modality, and the second modality to which the second-modal image belongs can be one of the near-infrared modality, the visible light modality, and the depth modality. For example, the first-modal image is a visible light image, and the second-modal image is a near-infrared image. The first modality and the second modality are different, and the first modality and the second modality are relative and can be swapped according to actual needs.

[0026] The domain corresponding to the visible light image can be divided according to the external environmental factors affecting imaging and other factors that are expected to be generalized. The external environmental factors can include the shooting scene (indoors, outdoors), lighting (strong light, weak light). The domain corresponding to the infrared light image can be divided according to the type of living body attack, such as paper photo attack, electronic image attack, three-dimensional head mold attack.

[0027] There are multiple training image pairs for training the living body detection model. The first-modal image and the second-modal image in the same training image pair are from training objects of the same living body category. The living body category includes living bodies and non-living bodies. The training objects of the same living body category can be further divided into training objects of the same living body category and the same ID, and training objects of the same living body category and different IDs. If the first-modal image and the second-modal image are from training objects of the same living body category and the same ID, then the first-modal image and the second-modal image can be taken at the same time or at different times. At least two training image pairs are from training objects of different living body categories. For the sake of simplicity of description, the embodiments of this application only illustrate one training image pair.

[0028] In some embodiments, the training image pair may further include a living body label characterizing the true living body category of the training object, depending on the training requirements.

[0029] S12: Extract the first-modal image features using the first feature extraction branch of the living body detection model to obtain the first features.

[0030] The first feature includes a first domain discrimination feature. The first domain discrimination feature can be used for the training task of domain discrimination of the first-modal image.

[0031] S13: Use the second feature extraction branch of the live detection model to extract features from the second-modal image to obtain a second feature.

[0032] The second feature includes a second domain discrimination feature. The second domain discrimination feature can be used for the training task of domain discrimination of the second-modal image.

[0033] S14: Perform domain discrimination based on the first domain discrimination feature and / or the second domain discrimination feature to obtain a domain discrimination result.

[0034] The domain label of the training image pair can include a first domain label and / or a second domain label. The first domain label represents the domain to which the first-modal image belongs, and the second domain label represents the domain to which the second-modal image belongs. Correspondingly, the domain discrimination result can include a first domain discrimination result and / or a second domain discrimination result.

[0035] If the domain label includes the first domain label, the domain discrimination result includes the first domain discrimination result, and domain discrimination can be performed based on the first domain discrimination feature to obtain the first domain discrimination result. If the domain label includes the second domain label, domain discrimination can be performed based on the second domain discrimination feature to obtain the second domain discrimination result.

[0036] The format of the first domain discrimination result is the same as that of the first domain label, including the probabilities of each first-modal domain category. The format of the second domain discrimination result is the same as that of the second domain label, including the probabilities of each second-modal domain category.

[0037] S15: Construct a domain discrimination loss based on the difference between the domain discrimination result and the domain label.

[0038] The greater the difference between the domain discrimination result and the domain label, the smaller the domain discrimination loss.

[0039] The domain discrimination result can include a first domain discrimination loss corresponding to the first feature extraction branch and / or a second domain discrimination loss corresponding to the second feature extraction branch. The first domain discrimination loss can be constructed based on the difference between the first domain discrimination result and the first domain label, and the second domain discrimination loss can be constructed based on the difference between the second domain discrimination result and the second domain label. The greater the difference between the first domain discrimination result and the first domain label, the smaller the first domain discrimination loss; the greater the difference between the second domain discrimination result and the second domain label, the smaller the second domain discrimination loss.

[0040] S16: Adjust the parameters of the live detection model based on at least the domain discrimination loss.

[0041] The parameters of the first feature extraction branch can be adjusted based on the first domain discrimination loss, and the parameters of the second feature extraction branch can be adjusted based on the second domain discrimination loss.

[0042] Through the implementation of this embodiment, the present application performs domain discrimination based on the first domain discrimination feature and / or the second domain discrimination feature to obtain a domain discrimination result; constructs a domain discrimination loss based on the difference between the domain discrimination result and the domain label, and adjusts the parameters of the live detection model based on the domain discrimination loss, thereby realizing the training of the live detection model. Among them, since the greater the difference between the domain discrimination result and the domain label, the smaller the domain discrimination loss, therefore, by training the live detection model with the domain discrimination loss, the parameters of the live detection model can be adjusted through the domain discrimination loss, and the domain generalization ability of the live detection model can be improved, that is, the attention of the live detection model to domain-related information during feature extraction is reduced, thereby improving the expression ability of the features extracted by the live detection model for the live category, and further improving the accuracy of the live discrimination result obtained based on the features in the application stage.

[0043] Furthermore, in some embodiments, the first domain discrimination feature can also be used for training tasks such as live discrimination of training objects, cross-modal supervision training tasks, and so on.

[0044] In some embodiments, the first feature may further include at least one of a first cross-modal supervision feature and a first live discrimination feature. The first cross-modal supervision feature is used for the training task of cross-modal supervision, the first live discrimination feature is used for the training task of live discrimination, and the first domain discrimination feature focuses on the training task of domain discrimination.

[0045] In some embodiments, the second domain discrimination feature can also be used for training tasks such as live discrimination of training objects, cross-modal supervision training tasks, and so on.

[0046] In some embodiments, the second feature may further include at least one of a second cross-modal supervision feature and a second live discrimination feature. The second cross-modal supervision feature is used for the training task of cross-modal supervision, the first live discrimination feature is used for the training task of live discrimination, and the second domain discrimination feature focuses on the training task of domain discrimination.

[0047] The following introduces the training tasks of cross-modal supervision and live discrimination:

[0048] Figure 2 It is a schematic flowchart of another embodiment of the training method of the live detection model of the present application. It should be noted that if there are substantially the same results, this embodiment does not Figure 2It is limited to the process sequence shown. In this embodiment, S21 is a step that can be included before S16, and S22 is a further extension of S16. In this embodiment, the content that is not repeated in the above embodiments will not be elaborated. The first feature further includes a first cross-modal supervision feature, and the second feature further includes a second cross-modal supervision feature. As Figure 2 shown, this embodiment may include:

[0049] S21: Construct a cross-modal supervision loss based on the first cross-modal supervision feature and the second cross-modal supervision feature.

[0050] The first cross-modal supervision feature and the first domain discriminant feature may be the same feature or different features. The same applies to the second cross-modal supervision feature and the second domain discriminant feature.

[0051] The first cross-modal supervision feature may include at least one of the supervision feature of the first modality and / or the supervised feature of the first modality. In some embodiments, the supervised feature of the first modality is a frequency-domain feature, and the supervision feature of the first modality is a convolutional feature.

[0052] The second cross-modal supervision feature may include the supervision feature of the second modality and / or the supervised feature of the second modality. In some embodiments, both the supervision feature of the second modality and the supervised feature of the second modality are convolutional features.

[0053] The cross-modal supervision loss may be, but is not limited to, an L2 regression loss. The cross-modal supervision loss may include the cross-modal supervision loss corresponding to the first feature extraction branch and / or the cross-modal supervision loss corresponding to the second feature extraction branch. The cross-modal supervision loss corresponding to the first feature extraction branch may be constructed based on the supervised feature of the first modality and the supervision feature of the second modality, and the cross-modal supervision loss corresponding to the second feature extraction branch may be constructed based on the supervision feature of the first modality and the supervised feature of the second modality.

[0054] In some embodiments, when the supervision feature of the first modality is a frequency-domain feature, S21 may include: converting the supervised feature of the second modality to the frequency domain; constructing the cross-modal supervision loss corresponding to the second feature extraction branch based on the difference between the supervision feature of the first modality and the converted supervised feature of the second modality. Thus, the differential supervision of the supervised feature of the second modality by the supervision feature of the first modality is realized.

[0055] In some embodiments, S21 may include: translating the supervised feature of the first modality from the first modality to the second modality; constructing the cross-modal supervision loss corresponding to the first feature extraction branch based on the difference between the translated supervised feature of the first modality and the supervision feature of the second modality. Thus, the differential supervision of the supervised feature of the first modality by the supervision feature of the second modality is realized.

[0056] S22: Adjust the parameters of the live detection model based on the cross-modal supervision loss and the domain discrimination loss.

[0057] The training of the live detection model through the training task of cross-modal supervision (cross-modal supervision loss) and the training task of domain discrimination (domain discrimination loss) can be carried out synchronously or in stages, and the order is not limited. For example, in the case of synchronous execution, weights can be assigned to the cross-modal supervision loss and the domain discrimination loss, and the parameters of the live detection model can be adjusted based on the weighted results. The weights of the supervision loss and the domain discrimination loss can be the same or different, which are specifically set according to requirements. Another example is that in the case of staged execution, the parameters of the live detection model can be adjusted based on the cross-modal supervision loss to obtain a live detection model that meets the cross-modal supervision conditions; the parameters of the live detection model that meets the cross-modal supervision conditions can be adjusted based on the domain discrimination loss to obtain a live detection model that meets the domain discrimination conditions. Among them, when adjusting the parameters of the live detection model based on the cross-modal supervision loss, the parameters of the first feature extraction branch can be adjusted based on the cross-modal supervision loss corresponding to the first feature extraction branch, and the parameters of the second feature extraction branch can be adjusted based on the cross-modal supervision loss corresponding to the second feature extraction branch. The cross-modal supervision conditions and the domain discrimination conditions can include but are not limited to the number of training times reaching the expectation, the training time reaching the expectation, and the training effect reaching the expectation.

[0058] Different from other embodiments, in this embodiment, considering that the first-modal image and the second-modal image contain information of training objects of the same live category in different modalities. For example, the first-modal image is a visible light image (including information such as color, which can express the contour of the training object, etc.), and the second-modal image is a near-infrared image (including information such as reflection intensity, etc.). Using the cross-modal supervision loss to train the live detection model can enable the feature extraction processes of the first feature extraction branch and the second feature extraction branch to be cross-modally mutually supervised, so as to realize the mutual complementation and enhancement of the information between the features extracted by the first feature extraction branch and the second feature extraction branch, further improving the expression ability of the features extracted by the live detection model for the live category, and further improving the accuracy of the live discrimination result obtained based on the features during live detection in the application stage.

[0059] Figure 3 It is a schematic flowchart of another embodiment of the training method of the live detection model of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 3 the process sequence shown. In this embodiment, S31-S33 are steps that can be included before S22, and S34 is a further expansion of S22. In this embodiment, the content that is repeated in the above embodiments will not be elaborated. In this embodiment, the first feature further includes a first live discrimination feature, and the second feature further includes a second live discrimination feature, such as Figure 3As shown, this embodiment may include:

[0060] S31: Fuse the first living body discrimination feature and the second living body discrimination feature to obtain a living body discrimination feature.

[0061] The first living body discrimination feature and the first domain discrimination feature may be the same feature or different features. The same applies to the second living body discrimination feature and the second domain discrimination feature. The fusion methods include but are not limited to multiplication and addition.

[0062] S32: Based on the living body discrimination feature, perform living body discrimination on the training object to obtain the living body discrimination result of the training object.

[0063] The living body discrimination result represents the predicted living body category of the training object and is in the same format as the living body label. Specifically, the living body discrimination result may include the probability that the training object is a living body and the probability that it is not a living body. If the probability of being a living body is greater than the probability of not being a living body, it represents that the training object is a living body; otherwise, it represents that the training object is a non-living body.

[0064] S33: Construct a living body discrimination loss based on the difference between the living body discrimination result and the living body label.

[0065] S34: Adjust the parameters of the living body detection model based on the living body discrimination loss, the cross-modal supervision loss, and the domain discrimination loss.

[0066] The training of the living body detection model through the training task of living body discrimination (living body discrimination loss), the training task of cross-modal supervision loss (cross-modal supervision loss), and the training task of domain discrimination (domain discrimination loss) can be carried out synchronously or in stages, and the sequence is not limited. For example, in the case of synchronous execution, weights can be assigned to the living body discrimination loss, the cross-modal supervision loss, and the domain discrimination loss, and the parameters of the living body detection model can be adjusted based on the weighted result. The weights of the living body discrimination loss, the cross-modal supervision loss, and the domain discrimination loss can be the same or different. For example, the weight of the living body discrimination loss is larger. Another example is that in the case of staged execution, the parameters of the living body detection model can be adjusted based on the cross-modal supervision loss to obtain a living body detection model that meets the cross-modal supervision conditions, and the parameters of the living body detection model that meets the cross-modal supervision conditions can be adjusted based on the domain discrimination loss to obtain a living body detection model that meets the domain discrimination conditions; the parameters of the living body detection model that meets the domain discrimination conditions can be adjusted based on the living body discrimination loss to obtain a living body detection model that meets the living body discrimination conditions. The cross-modal supervision conditions, the living body discrimination conditions, and the domain discrimination conditions may include but are not limited to the number of training times reaching the expectation, the training time reaching the expectation, and the training effect reaching the expectation.

[0067] Different from other embodiments, in this embodiment, information in the features of two different modality images is combined for live detection, and a live detection loss is constructed based on the live detection result. Based on the live detection loss, training of the live detection model is realized, which can further improve the expression ability of the features extracted by the live detection model for the live category, and further improve the accuracy of the live detection result obtained based on the features during live detection in the application stage.

[0068] In some embodiments, the first feature extraction branch includes two separate first feature extraction sub-branches. At least one first feature extraction sub-branch includes a plurality of first feature extraction layers, and the first feature extraction layers in the first feature extraction sub-branch are connected in sequence. The subsequent first feature extraction layer performs feature extraction based on the previous first feature extraction layer. Different features included in the first feature can be extracted by different first feature extraction sub-branches or different first feature extraction layers. Exemplarily, the supervised feature of the first modality and the supervised-by feature of the first modality are extracted by different first feature extraction sub-branches, and the supervised feature of the first modality is respectively extracted by different first feature extraction layers of the same first feature extraction sub-branch as the first domain discrimination feature and the first live detection feature. When the supervised feature of the first modality is a frequency domain feature, the corresponding first feature extraction sub-branch (frequency domain feature extraction sub-branch) can extract the frequency domain feature by means such as Fourier transform. For example, the frequency domain feature extraction sub-branch can perform interpolation and then downsampling on the first modality image, or can first perform LBP or Hog feature extraction and then downsample, and perform Fourier transform on the downsampling result to obtain the frequency domain feature.

[0069] In some embodiments, the second feature extraction branch includes a plurality of second feature extraction layers, and different features included in the second feature can be extracted by different or the same second feature extraction layers. Exemplarily, the supervised-by feature of the second modality, the supervised feature of the second modality, the second domain discrimination feature, and the second live detection feature are respectively extracted by different second feature extraction layers.

[0070] In some embodiments, the network structures of the first feature extraction branch and the second feature extraction branch can be based on VGG, ResNet, etc. Exemplarily, the network structure of the first feature extraction branch is ResNet50, and the network structure of the second feature extraction branch is VGG16.

[0071] The training method provided by the present application is described below by way of an example:

[0072] With reference to Figure 4 , Figure 4 is a schematic diagram of the training structure of the live detection model. As shown in Figure 4As shown, the live detection model includes a first feature extraction branch, a second feature extraction branch, a translation layer, a conversion layer, a first domain discriminator, a second domain discriminator, and a live discriminator. The first feature extraction branch includes two separate first feature extraction sub-branches, namely a frequency domain feature extraction sub-branch and a convolutional feature extraction sub-branch. The convolutional feature extraction sub-branch includes several first convolutional feature extraction layers M1i, and the second feature extraction branch includes several second convolutional feature extraction layers M2i.

[0073] 1) Obtain training image pairs, including visible light images and near-infrared images from the same face, live labels of the face, first domain labels of the visible light images, and second domain labels of the near-infrared images. The first modality is the visible light modality, and the second modality is the near-infrared modality.

[0074] 2) Input the visible light images into the frequency domain feature extraction sub-branch and the convolutional feature extraction sub-branch of the first feature extraction branch respectively, and input the near-infrared images into the second feature extraction branch.

[0075] 3) Processing of the first feature extraction branch: The frequency domain feature extraction sub-branch extracts the frequency domain features of the visible light images (supervised features of the first modality); M11 of the convolutional feature extraction sub-branch processes the visible light images, and the output is the supervised feature of the first modality; then M12 continues to process the supervised feature of the first modality, and the output is the first domain discriminant feature and the first live discriminant feature. Exemplarily, M11 includes a basic feature extraction layer and a higher-level feature extraction layer. The basic feature extraction layer can process the visible light images using central difference convolution to obtain fine-grained features, and the higher-level feature extraction layer can process the fine-grained features to obtain the convolutional features of the first modality. For examples of the processing of central difference convolution, refer to Figure 5 。

[0076] 4) Processing of the second feature extraction branch: M21 processes the near-infrared images, and the output is the supervised feature of the second modality; then M22 continues to process the supervised feature of the second modality, and the output is the supervised feature of the second modality; then M23 processes the supervised feature of the second modality, and the output is the second domain discriminant feature and the second live discriminant feature. The network structure of M21 can be similar to that of M11.

[0077] 5) Processing of the translation layer: Translate the supervised feature of the first modality from the first modality to the second modality to obtain the translated supervised feature of the first modality.

[0078] 6) Processing of the conversion layer: Convert the supervised feature of the second modality to the frequency domain to obtain the converted supervised feature of the second modality.

[0079] 7) Based on the differences between the supervised features of the first modality and the supervised features of the transformed second modality, construct the cross-modal supervision loss corresponding to the second feature extraction branch; based on the differences between the supervised features of the translated first modality and the supervised features of the second modality, construct the cross-modal supervision loss corresponding to the first feature extraction branch. The specific formulas can be as follows:

[0080]

[0081]

[0082] Among them, L GD1 and L GD2 respectively represent the cross-modal supervision losses corresponding to the second feature extraction branch and the first feature extraction branch, n represents the total number of dimensions, represents the j-th dimension of the supervised features of the transformed second modality, represents the j-th dimension of the supervised features of the first modality, represents the j-th dimension of the supervised features of the translated first modality, represents the j-th dimension of the supervised features of the second modality.

[0083] 8) Fuse the first living discriminant feature and the second living discriminant feature to obtain the living discriminant feature, and the living discriminator obtains the living discriminant result based on the living discriminant feature; based on the differences between the living discriminant result and the living label, construct the living discriminant loss. The specific formulas can be as follows:

[0084]

[0085] Among them, L cls represents the living discriminant loss, C represents the total number of living categories, represents the probability of the j-th living category in the living label, and p j represents the probability of the j-th living category in the living discriminant result.

[0086] 9) The first domain discriminator obtains the first domain discriminant result based on the first domain discriminant feature; based on the differences between the first domain discriminant result and the first domain label, construct the domain discriminant loss corresponding to the first feature extraction branch. The second domain discriminator obtains the second domain discriminant result based on the second domain discriminant feature; based on the differences between the second domain discriminant result and the second domain label, construct the domain discriminant loss corresponding to the second feature extraction branch. The specific formulas can be as follows:

[0087]

[0088]

[0089] Among them, LDG1 and L DG2 respectively represent the losses corresponding to the first feature extraction branch and the second feature extraction branch, C represents the total number of domain categories, and p m respectively represent the first domain label and the probability of the m-th visible domain category in the first domain discrimination result. and p n respectively represent the second domain label and the probability of the n-th infrared domain category in the second domain discrimination result.

[0090] 8) Weight each loss to obtain the final loss of the live detection model; adjust the parameters of the live detection model based on the final loss. The formula for the final loss can be as follows:

[0091] L all =γ1L GD1 +γ2L GD2 +γ3L cls +γ4L DG1 +γ5L DG2 ;

[0092] where γ1, γ2, γ3, γ4, and γ5 represent weights and can be set according to requirements. For example, set γ1 and γ2 to 0.2, γ3 to 0.4, and γ4 and γ5 to 0.2.

[0093] It should be noted that in some embodiments, the number of modalities in the training image pair can be adaptively increased, that is, the training image pair is extended to a training image group, and the training image group includes images of at least three modalities. Correspondingly, the feature extraction branches of the live detection model can be adaptively increased, that is, the live detection model includes at least three feature extraction branches. The subsequent processing after the increase can be deduced according to the situation of two modalities and will not be elaborated here.

[0094] The live detection model trained through any of the above embodiments can be used in the live detection method. Specifically, it can be as follows:

[0095] Figure 6 is a schematic flowchart of an embodiment of the live detection method of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 6 the process sequence shown. As Figure 6 shown, this embodiment may include:

[0096] S41: Obtain an image of the object to be detected.

[0097] The image of the object to be detected includes the image to be detected in the first modality and / or the second modality.

[0098] The execution subject of the embodiment of the in-vivo detection method of this application is an application device, and the application device can be an electronic device in the form of a computer, a mobile phone, a server, etc.

[0099] S42: Use the in-vivo detection model to extract features from the image of the object to be detected to obtain in-vivo discrimination features.

[0100] The in-vivo detection model can use the first feature extraction branch to extract features from the to-be-detected image of the first modality to obtain the first in-vivo discrimination feature, and / or use the second feature extraction branch to extract features from the to-be-detected image of the second modality to obtain the second in-vivo discrimination feature; if one of the first in-vivo discrimination feature and the second in-vivo discrimination feature is extracted, the extracted one is the in-vivo discrimination feature; if both are extracted, the first in-vivo discrimination feature and the second in-vivo discrimination feature can also be fused to obtain the in-vivo discrimination feature.

[0101] S43: Obtain the in-vivo discrimination result of the object to be detected based on the in-vivo discrimination feature.

[0102] The in-vivo discriminator of the in-vivo detection model can be used to discriminate based on the in-vivo discrimination feature to obtain the in-vivo discrimination result of the object to be detected. The in-vivo discrimination result of the object to be detected represents the in-vivo category of the object to be detected.

[0103] For other detailed descriptions of this embodiment, refer to the previous embodiments, and details are not described herein.

[0104] Through the implementation of this embodiment, since the in-vivo detection model is trained by the training method of the previous embodiment, the in-vivo discrimination feature extracted by the in-vivo detection model has strong expression ability and high accuracy for the in-vivo category, and the in-vivo discrimination result obtained based on this has high accuracy.

[0105] Figure 7 It is a schematic structural diagram of an embodiment of an electronic device of this application. As Figure 7 shown, the electronic device includes a processor 21 and a memory 22 coupled to the processor 21.

[0106] Among them, the memory 22 stores program instructions for implementing the method of any of the above embodiments; the processor 21 is configured to execute the program instructions stored in the memory 22 to implement the steps of the above method embodiments. Among them, the processor 21 may also be referred to as a CPU (Central Processing Unit). The processor 21 may be an integrated circuit chip with signal processing capabilities. The processor 21 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0107] In this embodiment, the electronic device may be the aforementioned training device or application device.

[0108] Figure 8 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. As Figure 8 shown, the computer-readable storage medium 30 of the embodiment of the present application stores program instructions 31, and when the program instructions 31 are executed, they implement the method provided in the above embodiments of the present application. Among them, the program instructions 31 may form a program file and be stored in the above computer-readable storage medium 30 in the form of a software product, so that a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor can execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned computer-readable storage medium 30 includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0109] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be through some interfaces, and the indirect couplings or communication connections of devices or units may be in electrical, mechanical, or other forms.

[0110] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units. The above are only the implementation manners of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A training method for a live detection model, characterized in that, Including: Obtaining training image pairs, where the training image pairs include a first-modal image and a second-modal image, as well as a domain label characterizing the field to which the first-modal image and / or the second-modal image belong, and the first-modal image and the second-modal image are from training objects of the same living body category; Extracting features from the first-modal image by using a first feature extraction branch of the living body detection model to obtain a first feature, where the first feature includes a first domain discrimination feature; Extracting features from the second-modal image by using a second feature extraction branch of the living body detection model to obtain a second feature, where the second feature includes a second domain discrimination feature; Performing domain discrimination based on the first domain discrimination feature and / or the second domain discrimination feature to obtain a domain discrimination result; Constructing a domain discrimination loss based on the difference between the domain discrimination result and the domain label, where the greater the difference between the domain discrimination result and the domain label, the smaller the domain discrimination loss; Adjusting parameters of the living body detection model based at least on the domain discrimination loss.

2. The method according to claim 1, characterized in that, The first feature further includes a first cross-modal supervision feature, the second feature further includes a second cross-modal supervision feature, and the method further includes: Constructing a cross-modal supervision loss based on the first cross-modal supervision feature and the second cross-modal supervision feature; The adjusting the parameters of the living body detection model based at least on the domain discrimination loss includes: Adjusting the parameters of the living body detection model based on the cross-modal supervision loss and the domain discrimination loss.

3. The method according to claim 2, characterized in that The first cross-modal supervision feature includes a supervision feature of the first modality, the second cross-modal supervision feature includes a supervised feature of the second modality, the supervision feature of the first modality is a frequency-domain feature, and constructing the cross-modal supervision loss based on the first cross-modal supervision feature and the second cross-modal supervision feature includes: Converting the supervised feature of the second modality to the frequency domain; Constructing the cross-modal supervision loss corresponding to the second feature extraction branch based on the difference between the supervision feature of the first modality and the converted supervised feature of the second modality.

4. The method according to claim 2, wherein The first cross-modal supervision feature includes a supervised feature of the first modality, the second cross-modal supervision feature includes a supervision feature of the second modality, and constructing the cross-modal supervision loss based on the first cross-modal supervision feature and the second cross-modal supervision feature includes: Translating the supervised feature of the first modality from the first modality to the second modality; Constructing the cross-modal supervision loss corresponding to the first feature extraction branch based on the difference between the translated supervised feature of the first modality and the supervision feature of the second modality.

5. The method according to claim 3 or 4, wherein The first-modal image is a visible light image, and the second-modal image is a near-infrared image.

6. The method according to claim 2, wherein The first feature further includes a first living body discrimination feature, the second feature further includes a second living body discrimination feature, the training image pair further includes a living body label characterizing the true living body category of the training object, and before adjusting the parameters of the living body detection model based on the cross-modal supervision loss and the domain discrimination loss, the method further includes: Fuse the first living body discrimination feature and the second living body discrimination feature to obtain a living body discrimination feature; Based on the living body discrimination feature, perform living body discrimination on the training object to obtain the living body discrimination result of the training object; Construct a living body discrimination loss based on the difference between the living body discrimination result and the living body label; Adjusting the parameters of the living body detection model based on the cross-modal supervision loss and the domain discrimination loss includes: Adjust the parameters of the living body detection model based on the living body discrimination loss, the cross-modal supervision loss, and the domain discrimination loss.

7. The method according to claim 6, wherein The first cross-modal supervision feature includes a supervised feature of the first modality and a supervision feature of the first modality, and the second cross-modal supervision feature includes a supervised feature of the second modality and a supervision feature of the second modality; The first feature extraction branch includes two separate first feature extraction sub-branches, at least one of the first feature extraction sub-branches includes a plurality of first feature extraction layers, the supervised feature of the first modality and the supervised feature of the first modality are extracted by different first feature extraction sub-branches, and the supervised feature of the first modality is respectively extracted from different first feature extraction layers of the same first feature extraction sub-branch as the first domain discrimination feature and the first living body discrimination feature; and / or The second feature extraction branch includes a plurality of second feature extraction layers, and the supervised feature of the second modality, the supervision feature of the second modality, the second domain discrimination feature, and the second living body discrimination feature are respectively extracted by different second feature extraction layers.

8. The method according to claim 6, wherein Adjusting the parameters of the living body detection model based on the living body discrimination loss, the cross-modal supervision loss, and the domain discrimination loss includes: Adjust the parameters of the living body detection model based on the cross-modal supervision loss to obtain a living body detection model that meets the cross-modal supervision conditions; Adjust the parameters of the living body detection model that meets the cross-modal supervision conditions based on the domain discrimination loss to obtain a living body detection model that meets the domain discrimination conditions; Adjust the parameters of the living body detection model that meets the domain discrimination conditions based on the living body discrimination loss to obtain a living body detection model that meets the living body discrimination conditions.

9. A method for live detection, characterized in that, Includes: Obtain an image of the object to be detected, and the image of the object to be detected includes a to-be-detected image of the first modality and / or the second modality; Use the living body detection model to extract features from the image of the object to be detected to obtain a living body discrimination feature; Obtain the living body discrimination result of the object to be detected based on the living body discrimination feature; Wherein, the living body detection model is trained by the method according to any one of claims 1-8.

10. An electronic device, characterized in that, Includes a processor and a memory connected to the processor, wherein, The memory stores program instructions; The processor is configured to execute the program instructions stored in the memory to implement the method according to any one of claims 1-9.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, and the program instructions can be executed by a processor, and when executed, implement the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Face living body detection method, device and equipment and storage medium

    CN113128481A

  • Living body detection network training method, living body detection method, and devices

    CN113723215A