Living body detection model training method, living body detection method, living body detection device and living body detection equipment
Through the method of feature fusion and multimodal data training, the efficiency and accuracy of video live detection are improved, the problems of low efficiency and insufficient accuracy in the prior art are solved, and effective live detection under various conditions are achieved.
Patent Information
- Application Number
- CN202510336414.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The existing video live detection methods based on deep learning technology have insufficient efficiency and accuracy, especially when facing plane attacks of multiple mode videos, and are susceptible to light conditions.
The pre-trained feature processing network and the prompt module to be trained are adopted to improve the efficiency of video frame processing through feature fusion and description information generation, and the prompt module to be trained with multimodal data to generate more accurate live detection results.
It improves the efficiency and accuracy of live detection of videos, enhances the generalization ability of the model, and can perform effective detection under different lighting conditions.
Smart Images

Figure CN120260144A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to technical fields such as computer vision, deep learning, and large models. Background Art
[0002] Video liveness detection technology is the core of a liveness recognition system, which mainly uses deep learning technology to distinguish between genuine and fake videos to ensure system security. However, current video liveness detection methods based on deep learning technology have deficiencies in terms of efficiency and accuracy. Therefore, there is an urgent need for an efficient and highly accurate video liveness detection method. Summary of the Invention
[0003] The present disclosure provides a training method for a liveness detection model, a liveness detection method, and their devices and equipment.
[0004] According to one aspect of the present disclosure, there is provided a training method for a liveness detection model, including:
[0005] Using a pre-trained first feature processing network in an initial liveness detection model to perform feature processing on multiple first-modal data to obtain a first fusion feature; wherein, the initial liveness detection model at least includes the pre-trained first feature processing network and a to-be-trained hint module;
[0006] Using the to-be-trained hint module to perform feature description on the multiple first-modal data to obtain first description information; wherein, the first description information is used to describe the liveness features of a target object in at least part of the multiple first-modal data;
[0007] Based on the first fusion feature and the first description information, predicting a first liveness detection result for the multiple first-modal data;
[0008] At least using the first liveness detection result to train the to-be-trained hint module to obtain a target liveness detection model including a target hint module.
[0009] According to another aspect of the present disclosure, there is provided a liveness detection method, including:
[0010] Obtaining multiple target video frames including an object to be detected; wherein, the multiple target video frames include at least one first video frame of a first modality, and / or at least one second video frame of a second modality;
[0011] Inputting the multiple target video frames into a target liveness detection model; wherein, the target liveness detection model includes a first detection branch for detecting data of the first modality and a second detection branch for detecting data of the second modality;
[0012] Obtain a target detection result; wherein, the target detection result is obtained based on the first detection branch and / or the second detection branch.
[0013] According to another aspect of the present disclosure, there is provided a training device for a live detection model, including:
[0014] A feature processing unit, configured to use a pre-trained first feature processing network in an initial live detection model to perform feature processing on a plurality of first-modal data to obtain a first fusion feature; wherein, the initial live detection model at least includes the pre-trained first feature processing network and a to-be-trained prompt module;
[0015] An information generation unit, configured to use the to-be-trained prompt module to perform feature description on the plurality of first-modal data to obtain first description information; wherein, the first description information is used to describe the live features of a target object in at least part of the plurality of first-modal data.
[0016] A prediction unit, configured to predict a first live detection result for the plurality of first-modal data based on the first fusion feature and the first description information.
[0017] A training unit, configured to train the to-be-trained prompt module at least using the first live detection result to obtain a target live detection model including a target prompt module.
[0018] According to another aspect of the present disclosure, there is provided a live detection device, including:
[0019] An acquisition unit, configured to acquire a plurality of target video frames including an object to be detected; wherein, the plurality of target video frames include at least one first video frame of a first modality and / or at least one second video frame of a second modality.
[0020] An input unit, configured to input the plurality of target video frames into a target live detection model; wherein, the target live detection model includes a first detection branch for detecting data of the first modality and a second detection branch for detecting data of the second modality.
[0021] An output unit, configured to obtain a target detection result; wherein, the target detection result is obtained based on the first detection branch and / or the second detection branch.
[0022] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0023] At least one processor; and
[0024] A memory communicatively connected to the at least one processor; wherein,
[0025] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.
[0026] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0027] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, which implements any method in the embodiments of the present disclosure when executed by a processor.
[0028] The solution of the present disclosure uses a pre-trained first feature processing network to obtain first fusion features for a plurality of first modality data. Compared with traditional manual feature processing and classification methods, the processing efficiency of the plurality of first modality data is effectively improved. Moreover, the solution of the present disclosure uses a to-be-trained prompt module to obtain first description information for describing the liveness features of the target object in at least part of the first modality data. Thus, it provides strong reference information for the prediction stage, and further provides strong support for improving the accuracy of the liveness detection result.
[0029] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0031] Figure 1 is a schematic flowchart of a method for training a liveness detection model according to an embodiment of the present application Figure 1 ;
[0032] Figure 2 is a schematic flowchart of a method for training a liveness detection model according to an embodiment of the present application Figure 2 ;
[0033] FIG. 3(a) is a schematic diagram of feature fusion according to an embodiment of the present application Figure 1 ;
[0034] FIG. 3(b) is a schematic diagram of feature fusion according to an embodiment of the present application Figure 2 ;
[0035] Figure 4Schematic flowchart III of a method for training a live detection model according to an embodiment of the present application;
[0036] Figure 5 Schematic diagram of obtaining a first live detection result by predicting using a first detection branch according to an embodiment of the present application;
[0037] Figure 6 Schematic process of a method for training a live detection model according to an embodiment of the present application Figure 4 ;
[0038] Figure 7 Schematic flowchart of the training process of a preset detection model for processing RGB video frames according to an embodiment of the present application;
[0039] Figure 8 Schematic diagram of the structure of an initial live detection model constructed according to an embodiment of the present application;
[0040] Figure 9 Schematic flowchart of a live detection method according to an embodiment of the present application;
[0041] Figure 10 Schematic diagram of the structure of a live detection model training device 1000 according to an embodiment of the present disclosure;
[0042] Figure 11 Schematic diagram of the structure of a live detection device 1100 according to an embodiment of the present disclosure;
[0043] Figure 12 Schematic block diagram of an example electronic device 1200 that can be used to implement the embodiments of the present disclosure. Detailed implementation manners
[0044] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0045] As used herein, the term "and / or" merely describes an associated relationship between associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. As used herein, the term "at least one" means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C. As used herein, the terms "first" and "second" represent multiple similar technical terms for differentiation and do not mean to limit the order or limit to only two. For example, the first feature and the second feature refer to two types / two features. The first feature may be one or more, and the second feature may also be one or more.
[0046] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can still be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.
[0047] The related technologies of the embodiments of the present disclosure are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present disclosure as optional solutions, and all of them fall within the protection scope of the embodiments of the present disclosure.
[0048] Video live detection technology is usually used to distinguish whether video content is captured in real time by a real person, thereby ensuring the security of the entire system. Deep learning technology is the mainstream method for realizing video live detection. Compared with traditional methods, it has achieved a significant improvement in detection accuracy. However, in actual applications, video live detection methods based on deep learning technology still face many challenges.
[0049] On the one hand, since the deep learning model needs to process a large number of video frames with similar features, this leads to a huge consumption of computing resources and a long processing time, thereby reducing the efficiency of video live detection.
[0050] On the other hand, the deep learning model highly depends on the richness of training data. If the amount of training data is insufficient, it may weaken the generalization ability of the model in actual applications. Especially when facing planar attacks of multi-modal videos, this lack of generalization ability makes the model easily deceived, thus having an adverse impact on the actual application effect.
[0051] In addition, deep learning algorithms are easily affected by lighting conditions, thereby causing fluctuations in the performance of video live detection.
[0052] Based on this, the present disclosure provides a method for training a live detection model and a method for performing video live detection using the trained live detection model. The training method of the present disclosure makes full use of the feature fusion technology for multiple video frames, effectively improving the processing efficiency of multiple video frames, and thus effectively improving the model detection efficiency. Moreover, the present disclosure also introduces a hint module in the live detection model to generate description information for multiple video frames by using the hint module, so as to assist the model in video detection through the description information, effectively improving the accuracy of video live detection.
[0053] Specifically, Figure 1 is a schematic flowchart of a method for training a live detection model according to an embodiment of the present application. Figure 1 This method is optionally applied to an electronic device, such as a personal computer, a server, a server cluster, and other electronic devices.
[0054] Furthermore, this method at least includes at least part of the following content. As Figure 1 shown, it includes:
[0055] Step S101: Use the pre-trained first feature processing network in the initial live detection model to perform feature processing on multiple first-modal data to obtain a first fusion feature.
[0056] In other words, the first fusion feature is the result of feature fusion of multiple first-modal data.
[0057] Here, the initial live detection model at least includes the pre-trained first feature processing network and the to-be-trained hint module. It can be understood that the pre-trained first feature processing network can be a feature processing network after pre-training; correspondingly, the to-be-trained hint module is the network that needs to be trained currently.
[0058] It should be noted that the "pre-training" can be to train the corresponding module, network, or model using an existing training method, and the present disclosure does not make specific limitations on the specific pre-training method.
[0059] Here, it should be noted that the first-modal data can be understood as the to-be-trained data of the first modality, and the present disclosure simply refers to the to-be-trained data of the first modality as the first-modal data.
[0060] Furthermore, in an example, each first-modal data contains a target object.
[0061] Step S102: Use the to-be-trained hint module to perform feature description on the multiple first-modal data to obtain first description information.
[0062] Here, the first description information is used to describe the live body characteristics of the target object in at least some of the multiple first modality data.
[0063] For example, in one example, the first modality data is a video frame. At this time, the N1 first modality data can be multiple consecutive or non - consecutive video frames containing the target object in the video data.
[0064] Furthermore, the first description information can describe the live body characteristics of the target object included in each of the multiple first modality data, or can describe the live body characteristics of some of the multiple first modality data. For example, for the scenario where the multiple first modality data are consecutive multiple video frames, at this time, the first description information can describe the live body characteristics of the target object in at least one video frame with specific typical features among the multiple video frames. In this way, redundant description content can be effectively avoided, providing strong support for further improving the detection efficiency while ensuring the detection accuracy.
[0065] In other words, the first description information in the present disclosure only needs to be able to describe the live body characteristics of the corresponding target object in the input multiple first modality data. The present disclosure does not limit the specific description content, which can be determined based on the implementation scenario requirements.
[0066] Furthermore, in one example, the live body characteristics include information that can reflect the life state or biological characteristics of the target object. For example, the live body characteristics of the first modality data can include the skin texture characteristics, motion characteristics, etc. of the target object in the first modality data. The present disclosure does not make specific restrictions on this.
[0067] Step S103: Based on the first fusion feature and the first description information, predict the first live body detection result for the multiple first modality data.
[0068] For example, in one example, the first live body detection result can represent whether the target object corresponding to the multiple first modality data is a live body.
[0069] Furthermore, in one example, the first prediction module (such as the first modality classification head) pre - trained in the initial live body detection model can be used to process the first fusion feature and the first description information to predict the first live body detection result.
[0070] Step S104: At least use the first live body detection result to train the to - be - trained prompt module to obtain a target live body detection model including a target prompt module.
[0071] In this way, the disclosed solution uses a pre-trained first feature processing network to obtain first fusion features for multiple first-modal data. Compared with traditional manual feature processing and classification methods, the processing efficiency of multiple first-modal data is effectively improved. Moreover, the disclosed solution uses a to-be-trained prompt module to obtain first description information for describing the liveness features of the target object in at least part of the first-modal data. Thus, it provides strong reference information for the prediction stage, and further provides strong support for improving the accuracy of the liveness detection result.
[0072] It should be noted that during the model training stage, the to-be-trained prompt module can be initialized based on the prompt template pool. For example, in one example, a prompt template can be randomly selected from the prompt template pool and used as the initial first description information.
[0073] It should be pointed out that in the disclosed solution, the to-be-trained prompt module can be trained by using contrastive learning. In this way, the understanding ability of the to-be-trained prompt module for the input data is improved. At the same time, the accuracy of the description information output by it is improved, and further strong support is provided for improving the accuracy of the liveness detection result in the future. For example, real samples (i.e., positive samples) and false samples (i.e., negative samples) for the target object are set, and the to-be-trained prompt module is used to understand the positive and negative samples to predict the description information of the positive and negative samples. Then, the to-be-trained prompt module is trained by using the predicted description information of the positive and negative samples.
[0074] It should be noted that the above contrastive learning is only an exemplary illustration. In actual applications, other training methods can also be used, and the disclosed solution does not make specific limitations on this.
[0075] Furthermore, in one example, training the to-be-trained prompt module by using at least the first liveness detection result can specifically include: training the to-be-trained prompt module by using at least the first liveness detection result while freezing the network parameters of the pre-trained first feature processing network. For example, without changing the network parameters of the pre-trained first feature processing network, the adjustable parameters in the to-be-trained prompt module are adjusted.
[0076] Here, in one example, the to-be-trained prompt module can be specifically a liveness video description model. The input of this liveness video description model can be specifically video data or video frames, and its output is description information.
[0077] Figure 2 is a schematic flow of a method for training a liveness detection model according to an embodiment of the present application Figure 2 . This method is optionally applied to an electronic device, such as a personal computer, a server, a server cluster, and other electronic devices. It can be understood that the above Figure 1The relevant content of the method shown can also be applied to this example, and the relevant associated content will not be elaborated further in this example.
[0078] Furthermore, the method at least includes at least some of the following content. As Figure 2 shown, it includes:
[0079] Step S201: Use the first feature extraction network in the pre-trained first feature processing network to extract features from multiple first-modal data to obtain the first initial features of each first-modal data.
[0080] Step S202: Use the first fusion network in the pre-trained first feature processing network to perform feature fusion on the first initial features of each first-modal data to obtain the first fusion feature.
[0081] Here, in the pre-trained first feature processing network, it at least includes a pre-trained first feature extraction network and a pre-trained first fusion network.
[0082] Here, the pre-trained first feature extraction network can be the first feature extraction network after pre-training. Correspondingly, the pre-trained first fusion network can also be the first fusion network after pre-training.
[0083] Furthermore, in one example, the first feature extraction network can be a Vision Transformer Base (ViT-Base) model. At this time, the ViT-Base model can be used to extract the features of each first-modal data to obtain the first initial features of each first-modal data.
[0084] Furthermore, based on the first initial features of each first-modal data and using the first fusion network, the first fusion feature can be obtained.
[0085] Furthermore, in one example, freeze the pre-trained first feature processing network and train the to-be-trained prompt module, which can specifically include: training the to-be-trained prompt module under the condition of freezing the network parameters in the first feature extraction network and freezing the network parameters in the first fusion network.
[0086] Step S203: Use the to-be-trained prompt module to describe the features of the multiple first-modal data to obtain the first description information.
[0087] Here, the first description information is used to describe the liveness features of the target object in at least some of the multiple first-modal data.
[0088] Step S204: Based on the first fusion feature and the first description information, a first liveness detection result for the multiple first-modal data is predicted.
[0089] It should be noted that for relevant examples of the to-be-trained prompt module, the first description information, etc., reference can be made to the above description, which will not be elaborated here.
[0090] Step S205: At least using the first liveness detection result, the to-be-trained prompt module is trained to obtain a target liveness detection model including a target prompt module.
[0091] In this way, the solution of the present disclosure extracts features from each first-modal data by using the first feature extraction network. Thus, the first initial features of each first-modal data are efficiently obtained, improving the accuracy and efficiency of feature extraction.
[0092] Furthermore, the solution of the present disclosure uses the first fusion network to fuse the first initial features of each first-modal data, thereby generating a first fusion feature that can comprehensively and accurately reflect the multiple first-modal data. In other words, this first fusion feature carries the features of the target object in different first-modal data. Therefore, it provides richer data for prediction, and further provides data support for efficiently and accurately predicting the first liveness detection result.
[0093] Furthermore, in a specific example, the following method can be used to perform feature fusion on the first initial features of each first-modal data; specifically, the feature fusion of the first initial features of each first-modal data to obtain the first fusion feature (such as in step S202) can specifically include:
[0094] Step S202-1: Concatenate the first initial features of each first-modal data to obtain a first concatenated feature.
[0095] Step S202-2: Obtain the first initial feature values of each feature unit in the first concatenated feature.
[0096] Here, the first initial feature value can represent the importance degree of the feature unit in the first concatenated feature.
[0097] Here, it should be noted that the feature unit can refer to the smallest feature unit in the first concatenated feature. In other words, it is the smallest component of the first concatenated feature. For example, it is a token.
[0098] Step S202-3: Based on the first initial feature values of each feature unit, perform pruning processing on the first concatenated feature to obtain the pruned first concatenated feature.
[0099] Here, the first fusion feature is the first spliced feature after pruning.
[0100] For example, FIGS. 3(a) and 3(b) are schematic diagrams of feature fusion according to an embodiment of the present application. As shown in FIG. 3(a), the first fusion network 301 may include a pre-trained first splicing module 3011 and a pre-trained first pruning module 3012. At this time, first use the pre-trained first splicing module 3011 to splice the first initial features of each first-modal data. For example, splice the tokens head to tail to obtain the first spliced feature, and then use the pre-trained first pruning module 3012 to perform pruning processing on the first spliced feature based on the first initial feature values of each feature unit to obtain the first spliced feature after pruning.
[0101] Further, in another example, as shown in FIG. 3(b), the first fusion network 301 may further include a pre-trained class-attention module 3013. At this time, the pre-trained class-attention module 3013 can be used to determine the first initial feature values of each feature unit in the first spliced feature.
[0102] Specifically, using the pre-trained class-attention module 3013, calculate the attention weights of each feature unit (for example, calculate each token). At this time, the attention weight can be used as the first initial feature value of the feature unit.
[0103] In this way, the present disclosure provides a refined solution for feature fusion. In this refined solution, first obtain the first spliced feature that splices all the first initial features of the first-modal data, and then perform pruning processing on the first spliced feature based on the importance degree of the feature units. In this way, remove the feature units with lower importance or redundancy, and then obtain the first fusion feature with more concise and diverse expressions. In other words, the information carried by the first fusion feature is large but not redundant. In this way, it provides data support for subsequent model prediction. Moreover, since the first fusion feature is concise and diverse, this refined solution also provides strong support for improving the model prediction efficiency.
[0104] Further, in a specific example, the following method can be used to implement the pruning processing of the first spliced feature; specifically, the above-mentioned pruning processing of the first spliced feature based on the first initial feature values of each feature unit (for example, step S202-3) may specifically include:
[0105] Remove the feature units from the first spliced feature whose first initial feature values are less than or equal to the first preset threshold.
[0106] That is to say, in this example, based on the first preset threshold, feature units with relatively low importance or redundancy can be determined from the first spliced features, that is, feature units whose first initial feature values are less than or equal to the first preset threshold. Furthermore, the feature units with relatively low importance or redundancy are removed from the first spliced features to obtain the first fused feature.
[0107] Here, in one example, the first preset threshold can be a preset value or can be dynamically calculated through an adaptive algorithm. The present disclosure does not specifically limit the determination method of the first preset threshold.
[0108] In this way, the present disclosure provides a refined pruning solution. This method is simple and efficient, can quickly identify and remove feature units with relatively low importance or redundancy, and thus obtain a more concise and diverse first fused feature. In this way, it provides data support for subsequent model prediction and also provides strong support for improving the model prediction efficiency.
[0109] Specifically, Figure 4 FIG. 3 is a schematic flowchart III of a method for training a living body detection model according to an embodiment of the present application. This method is optionally applied to an electronic device, such as a personal computer, a server, a server cluster, and other electronic devices. It can be understood that the relevant content of the method shown in FIG. 3 above can also be applied to this example, and the relevant associated content will not be elaborated in this example. Figure 1 To FIG. 3, the relevant content of the method shown in FIG. 3 above can also be applied to this example, and the relevant associated content will not be elaborated in this example.
[0110] Furthermore, this method at least includes at least part of the following content. As Figure 4 shown, it includes:
[0111] Step S401: Using a first feature processing network pre-trained in an initial living body detection model, perform feature processing on a plurality of first modality data to obtain a first fused feature.
[0112] Step S402: Using a to-be-trained prompt module, perform feature description on the plurality of first modality data to obtain first description information.
[0113] Here, the first description information is used to describe the living body features of the target object in at least part of the plurality of first modality data.
[0114] Step S403: Based on the first fused feature and the first description information, predict a first living body detection result for the plurality of first modality data.
[0115] It should be noted that for relevant examples of the to-be-trained prompt module and the first description information, etc., reference can be made to the above description, and details will not be elaborated here.
[0116] Here, in one example, the pre-trained first feature processing network and the to-be-trained prompt module are on the first detection branch of the initial living body detection model, and the first detection branch is used to detect the first modality data. That is to say, in this example, the initial living body detection model includes a branch dedicated to detecting the first modality. In this way, the first detection branch is used to detect the first modality data, so as to efficiently extract, process, and analyze the key feature information in the first modality, and then obtain the detection result in the first modality.
[0117] For example, Figure 5 is a schematic diagram of predicting the first living body detection result using the first detection branch according to an embodiment of the present application. As Figure 5 shown, the first detection branch 500 includes: a to-be-trained prompt module 501, a pre-trained first feature extraction network 502, and a pre-trained first fusion network 503. Further, in one example, it may further include a pre-trained first modality classification head 504.
[0118] As Figure 5 shown, the input data may specifically include multiple first modality data. At this time, the multiple first modality data can be respectively input into the pre-trained first feature extraction network 502 and the to-be-trained prompt module 501. Further, by using the pre-trained first feature extraction network 502, the first initial features corresponding to each first modality data can be obtained, and by using the to-be-trained prompt module 501, the first description information for the multiple first modality data can be generated.
[0119] Further, by using the pre-trained first fusion network 503, the multiple first initial features output by the pre-trained first feature extraction network 502 are fused into a complete feature to obtain the first fusion feature.
[0120] Further, by using the pre-trained first modality classification head 504, and based on the first fusion feature and the first description information, the first living body detection result is predicted.
[0121] Step S404: Determine the second living body detection result.
[0122] Here, in this example, the initial living body detection model further includes a second detection branch for detecting the second modality data. Further, the second living body detection result is obtained by detecting multiple second modality data using the second detection branch. In other words, the second living body detection result is based on multiple second modality data.
[0123] Here, in this example, the target object is also included in each second-modal data. At this time, the second living body detection result can be used to predict whether the target object corresponding to the multiple second-modal data is a living body.
[0124] It should be noted that the number of the first-modal data and the number of the second-modal data may be the same or different, and the present disclosure solution does not limit this.
[0125] In addition, it should be noted that the second-modal data can be understood as the training data to be trained in the second modality, and the present disclosure solution simply refers to the training data to be trained in the second modality as the second-modal data.
[0126] For example, in an example, the second-modal data can also be video frames. At this time, the multiple second-modal data can be consecutive or non-consecutive video frames containing the target object in the video data.
[0127] Here, it should be pointed out that the first modality and the second modality are different. At this time, the initial living body detection model of the present disclosure solution is a multi-modal detection model. In this way, the generalization ability of the model is effectively improved, thereby laying a foundation for improving the user experience.
[0128] Step S405: Use the first living body detection result and the second living body detection result to train the to-be-trained prompt module, so as to obtain a target living body detection model including a target prompt module.
[0129] In this way, the present disclosure solution combines the first living body detection result and the second living body detection result to train the to-be-trained prompt module, so that the to-be-trained prompt module can output more accurate description information, thereby improving the accuracy of the living body detection result.
[0130] Moreover, since the first living body detection result and the second living body detection result are predicted based on two different modal data, the target prompt module trained based on the first living body detection result and the second living body detection result can take into account the data in multiple modalities when making a description. In other words, the description information output can be made more accurate. Furthermore, compared with the single-modal video living body detection method based on deep learning, the present disclosure solution can take into account the living body detection in multiple modalities, thereby effectively improving the generalization ability of the target living body detection model.
[0131] Further, in a specific example, the second detection branch includes a pre-trained second feature processing network; further, the second living body detection result can be obtained in the following manner:
[0132] Use the pre-trained second feature processing network in the second detection branch to perform feature processing on the multiple second-modal data to obtain a second fusion feature;
[0133] Based on the second fusion feature, the second liveness detection result for the multiple second-modal data is predicted.
[0134] Further, in an example, the second detection branch may further include a pre-trained second prediction module, such as a second-modal classification head, for predicting the second liveness detection result of the target object corresponding to the multiple second-modal data based on the second fusion feature.
[0135] That is to say, in this example, the pre-trained second feature processing network can be used to process the multiple second-modal data to obtain the second fusion feature. Using the second-modal classification head and based on the second fusion feature, the second liveness detection result is predicted.
[0136] In this way, the solution of the present disclosure uses the pre-trained second feature processing network to obtain the second fusion feature for the multiple second-modal data. Compared with the traditional manual feature processing and classification methods, the processing efficiency of the multiple second-modal data is effectively improved. Moreover, the information carried by the second fusion feature of the solution of the present disclosure is richer. Therefore, the accuracy of predicting the second liveness detection result using the second fusion feature is effectively improved, and further provides data support for effectively improving the generalization ability of the target liveness detection model.
[0137] Further, in a specific example, the following method can be used to obtain the second fusion feature for the multiple second-modal data; specifically, the above-mentioned use of the pre-trained second feature processing network in the second detection branch to perform feature processing on the multiple second-modal data to obtain the second fusion feature may specifically include:
[0138] Using the second feature extraction network in the pre-trained second feature processing network to perform feature extraction on the multiple second-modal data to obtain the second initial features of the respective second-modal data;
[0139] Using the second fusion network in the pre-trained second feature processing network to perform feature fusion on the second initial features of the respective second-modal data to obtain the second fusion feature.
[0140] Here, in the pre-trained second feature processing network, at least a pre-trained second feature extraction network and a pre-trained second fusion network are included.
[0141] Here, the pre-trained second feature processing network can be a pre-trained second feature processing network. Similarly, the pre-trained second feature extraction network can also be a pre-trained second feature extraction network, and the pre-trained second fusion network can be a pre-trained second fusion network.
[0142] Further, in one example, the second feature extraction network can be a ViT-Base model. At this time, the ViT-Base model can be used to extract the features of each second-modal data to obtain the second initial features of each second-modal data.
[0143] Further, based on the second initial features of each second-modal data and using the pre-trained second fusion network, second fusion features can be obtained.
[0144] At this time, in one example, training the to-be-trained prompt module as described above can specifically include: training the to-be-trained prompt module while freezing the network parameters in the pre-trained first feature processing network and freezing the network parameters in the pre-trained second feature processing network.
[0145] Further, training the to-be-trained prompt module while freezing the network parameters in the pre-trained first feature processing network and freezing the network parameters in the pre-trained second feature processing network can specifically include:
[0146] Training the to-be-trained prompt module while freezing the network parameters in the pre-trained first feature extraction network, freezing the network parameters in the pre-trained first fusion network, freezing the network parameters in the pre-trained second feature extraction network, and freezing the network parameters in the pre-trained second fusion network.
[0147] In this way, the solution of the present disclosure extracts the features of each second-modal data by using the pre-trained second feature extraction network. In this way, the second initial features of each second-modal data are efficiently obtained, improving the accuracy and efficiency of feature extraction.
[0148] Further, the second initial features of each second-modal data are fused by using the pre-trained second fusion network, thereby generating second fusion features that can comprehensively and accurately reflect the characteristics of multiple second-modal data. In other words, the second fusion features carry the features of the target object in different second-modal data. Therefore, more abundant data is provided for prediction, and thus a data foundation is laid for predicting the results of the second live detection subsequently.
[0149] Further, in one specific example, the following method can be adopted to perform feature fusion on the second initial features of each second-modal data; specifically, the above-mentioned feature fusion of the second initial features of each second-modal data to obtain the second fusion features can specifically include:
[0150] Concatenate the second initial features of each second-modal data to obtain a second concatenated feature;
[0151] Obtain the second initial feature values of each feature unit in the second splicing feature; wherein, the second initial feature value can characterize the importance degree of the feature unit in the second splicing feature;
[0152] Based on the second initial feature values of each feature unit, perform pruning processing on the second splicing feature to obtain the pruned second splicing feature. Here, the second fusion feature is the pruned second splicing feature.
[0153] Here, it should be noted that the feature unit can refer to the smallest feature unit in the second splicing feature, in other words, it is the smallest component in the second splicing feature. For example, it is a token.
[0154] For example, the pre-trained second fusion network can include a pre-trained second splicing module and a pre-trained second pruning module. At this time, first use the pre-trained second splicing module to splice the second initial features of each second-modal data. For example, perform head-to-tail splicing on the tokens to obtain the second splicing feature, and then use the pre-trained second pruning module to perform pruning processing on the second splicing feature based on the second initial feature values of each feature unit to obtain the pruned second splicing feature.
[0155] Furthermore, in another example, the pre-trained second fusion network can also include a pre-trained category attention module. At this time, the pre-trained category attention module can be used to determine the second initial feature values of each feature unit in the second splicing feature.
[0156] Specifically, use the pre-trained category attention module to calculate the attention weights of each feature unit (for example, calculate each token). At this time, the attention weight can be used as the second initial feature value of the feature unit.
[0157] Here, the structural schematic diagram of the second fusion network is similar to those in FIGS. 3(a) and 3(b), and will not be elaborated here.
[0158] In this way, the solution of the present disclosure provides a refined solution for feature fusion. This refined solution first obtains the second splicing feature that splices all the second initial features of the second-modal data, and then performs pruning processing on the second splicing feature based on the importance degree of the feature unit, so as to remove the feature units with lower importance or redundancy, and further obtain the second fusion feature with more concise and diverse expressions. In other words, the information carried by the second fusion feature is large but not redundant. In this way, it provides data support for subsequent model prediction, and moreover, since the second fusion feature is concise and diverse, this refined solution also provides strong support for improving the model prediction efficiency.
[0159] Further, in a specific example, the following method can be used to implement the pruning process of the second splicing feature; specifically, the pruning process of the second splicing feature based on the second initial feature values of each feature unit described above may specifically include:
[0160] Remove the feature units with the second initial feature value less than or equal to the second preset threshold from the second splicing feature.
[0161] That is to say, in this example, the feature units with relatively low importance or redundancy can be determined from the second splicing feature based on the second preset threshold, that is, the feature units with the second initial feature value less than or equal to the second preset threshold. Furthermore, the feature units with relatively low importance or redundancy are removed from the second splicing feature to obtain the second fusion feature.
[0162] Here, in an example, the second preset threshold can be a preset value or can be dynamically calculated through an adaptive algorithm. The present disclosure does not specifically limit the determination method of the second preset threshold.
[0163] In this way, the present disclosure provides a refined pruning scheme. This method is simple and efficient, can quickly identify and remove the feature units with relatively low importance or redundancy, and further obtain a more concise and diverse second fusion feature. In this way, it provides data support for subsequent model prediction and also provides strong support for improving the model prediction efficiency.
[0164] Further, in a specific example, the above-mentioned multiple first-modal data are obtained based on the first-modal video; the multiple second-modal data are obtained based on the second-modal video.
[0165] Here, both the first-modal video and the second-modal video are obtained after video acquisition of the target object. That is to say, in this example, the same target object is included in the first-modal video and the second-modal video.
[0166] Further, in an example, the target object can be a face, a limb, a torso, etc. The present disclosure does not specifically limit the definition of the target object.
[0167] In this way, the present disclosure obtains multiple first-modal data through the first-modal video and obtains multiple second-modal data through the second-modal video. In this way, it provides rich multi-modal training materials for the training of the initial liveness detection model, and further lays a foundation for improving the generalization ability of the initial liveness detection model.
[0168] Further, in a specific example, the first-modal video described above may specifically be a Near Infrared (NIR) video. And / or, the second-modal video is a Red Green Blue (RGB) video.
[0169] That is to say, in one example, the first-modal video may specifically be an NIR video. Correspondingly, the first-modal data is then an NIR video frame. Or, in another example, the second-modal video is RGB. Correspondingly, the second-modal data is then an RGB video frame. Or, in another example, the first-modal video may specifically be an NIR video, and the second-modal data is an RGB video frame. At this time, the first-modal data is an NIR video frame, and the second-modal data is an RGB video frame.
[0170] That is to say, in one example, the first detection branch can be used to process NIR video frames. Correspondingly, the second detection branch can be used to process RGB video frames. In this way, the initial live detection model can be made to have multi-modal processing capabilities, thereby laying a foundation for improving the generalization ability of the initial live detection model.
[0171] In this way, the solution of the present disclosure can obtain corresponding NIR data and RGB data according to the NIR video and the RGB video, and then use the NIR data and the RGB data to train the initial live detection model, so that the initial live detection model can perform live detection on the NIR video and the RGB video. In this way, a foundation is laid for improving the generalization ability of the initial live detection model.
[0172] In addition, the NIR video is usually insensitive to light changes and can clearly image under low light conditions, while the RGB video can provide more detailed information about the color, texture, etc. of the target object. By combining the two to train the initial live detection model, the initial live detection model can be made to have stronger environmental adaptability and detection accuracy.
[0173] Further, in a specific example, the first description information used in the solution of the present disclosure is obtained based on the first-modal video, or based on at least some of the first-modal data among multiple first-modal data.
[0174] For example, in one example, the first description information can be directly obtained based on the first-modal video and by using the to-be-trained prompt module. In other words, the first description information is directly obtained based on the first-modal video. Or, in another example, after obtaining multiple first-modal data based on the first-modal video, based on the obtained multiple first-modal data and by using the to-be-trained prompt module, the first description information is obtained. In other words, the first description information is obtained based on multiple first-modal data.
[0175] In this way, the present disclosure provides a refinement solution for obtaining the first description information. This solution is simple, efficient, and has a more flexible acquisition method, enhancing the robustness of the first description information and providing a data basis for improving the accuracy of subsequent detection results.
[0176] In practical applications, in order to obtain high-quality first-modal data and second-modal data, an image processing process can also be performed. In this way, it lays a foundation for improving the training efficiency and training effect of the initial living body detection model in the future.
[0177] For example, in one example, a preset strategy can be used to sample the first-modal video (for example, NIR video) to obtain multiple first-modal sampling data corresponding to the first-modal video; further, image processing can be performed on each of the obtained first-modal sampling data to obtain multiple first-modal data; similarly, a preset strategy is used to sample the second-modal video (for example, RGB video) to obtain multiple second-modal sampling data corresponding to the second-modal video, and image processing is performed on each of the obtained second-modal sampling data to obtain multiple second-modal data.
[0178] Here, taking the specified video data (for example, the specified video data can specifically be the first-modal video or the second-modal video) as an example, the specific steps for obtaining the specified-modal data for model training are described in detail:
[0179] The first step: Extract candidate video frames from the specified video data according to the preset frame extraction interval.
[0180] The second step: Randomly determine the sampling data from the candidate video frames according to the preset sampling step. It can be understood that for the specified video data being the first-modal video, multiple first-modal sampling data can be obtained, and for the specified video data being the second-modal video, multiple second-modal sampling data can be obtained.
[0181] Further, for example, the preset frame extraction interval is 5. At this time, one video frame can be extracted from the specified video data every 5 frames as a candidate video frame. Further, assuming that the total number of frames of the specified video data is 75 frames, according to this frame extraction interval, 15 candidate video frames can be extracted from the specified video data.
[0182] Further, assuming that the sampling step is 5, at this time, 3 video frames can be randomly selected from the 15 candidate video frames as the determined sampling data.
[0183] It should be noted that in actual applications, in different batches of initial live detection model training, the present disclosure solution can also adopt the above sampling method to re-sample, so as to ensure that all candidate video frames extracted from the specified video data are likely to be learned by the initial live detection model, thereby enhancing the robustness of the initial live detection model to data noise.
[0184] In addition, the present disclosure solution can also set an upper limit for single acquisition, so as to effectively prevent overlearning of videos containing long or repetitive information, and thus avoid unnecessary waste of computing power.
[0185] In the third step, after obtaining the sampled data, perform face detection and image preprocessing on the sampled data to obtain specified modal data for model training. It can be understood that for the specified video data being the first-modal video, after image preprocessing, multiple first-modal data can be obtained; for the specified video data being the second-modal video, after image preprocessing, multiple second-modal data can be obtained.
[0186] Specifically, perform face detection on the sampled data to intercept the face region. For example, in one example, in order to improve the accuracy of face detection, a face detection model can be used to perform face detection on the sampled data, such as the first-modal sampled data or the second-modal sampled data. For example, the first-modal sampled data can be input into a pre-trained first-modal face detection model (such as a NIR face detection model) to obtain the face position of the first-modal sampled data, and after interception, the initial face region of the first-modal sampled data can be obtained. Similarly, the second-modal sampled data can be input into a pre-trained second-modal face detection model (such as an RGB face detection model) to obtain the face position of the second-modal sampled data, and after interception, the initial face region of the second-modal sampled data can be obtained.
[0187] Furthermore, perform face key point detection on the face region of the first-modal sampled data to obtain the first-modal face key point coordinates; similarly, perform face key point detection on the face region of the second-modal sampled data to obtain the second-modal face key point coordinates. For example, M face key points can be predefined in advance. Furthermore, perform face key point detection on the face region of the first-modal sampled data to obtain M first target key point coordinates for the face region of the first-modal sampled data. Similarly, perform face key point detection on the face region of the second-modal sampled data to obtain M second target key point coordinates for the face region of the second-modal sampled data.
[0188] Here, in one example, the facial key points may include but are not limited to features such as eyes, mouth, nose, etc. Further, M is a positive integer. For example, in one example, M can take the value of 72, that is, 72 facial key points are defined, and the coordinates of these 72 facial key points can be respectively denoted as: (x1, y1), (x2, y2), …, (x72, y72).
[0189] Further, in another example, after obtaining the coordinates of M first target key points in the facial region of the first-modal sampling data, the facial region of each first-modal sampling data can also be aligned according to the coordinates of the M first target key points; similarly, the facial region of each second-modal sampling data is aligned according to the coordinates of the M second target key points.
[0190] Further, in one example, after facial alignment, an affine transformation can be used and intercepted to obtain a first-modal facial image. Similarly, a second-modal facial image obtained by using an affine transformation and interception is obtained.
[0191] Here, operations such as translation, rotation, and scaling of the facial region can be realized through affine transformation, effectively adjusting the facial pose and size. For example, in one example, the image sizes of the first-modal facial image and the second-modal facial image intercepted by affine transformation are both 224 pixels × 224 pixels.
[0192] Further, in one example, normalization processing, data augmentation processing, etc. can be performed on the facial image of the specified modality.
[0193] Here, normalization processing can simplify the data to the greatest extent. For example, in one example, normalization processing can be performed on each first-modal facial image and each second-modal facial image. For example, normalization processing means: each pixel in the image (such as the first-modal facial image or the second-modal facial image) is normalized in turn. Specifically, the pixel value of each pixel can be first subtracted by 128 and then divided by 256 so that the pixel value of each pixel is between [-0.5, 0.5].
[0194] Further, in one example, after normalization processing, random data augmentation processing such as random rotation, random scaling, color perturbation, etc. can be performed to enhance the detectability of relevant information in the image, and then the specified-modal data for model training is obtained.
[0195] It should be noted that in practical applications, before model training, the specified-modal data can also be cut into multiple image patches. For example, cut into 196 image patches with a size of 16×16.
[0196] It is understandable that the above upsampling and image preprocessing processes are only for illustrative purposes. In actual applications, other methods can also be adopted, and the solution of the present disclosure is not limited thereto.
[0197] Figure 6 is a schematic flowchart of a method for training a living body detection model according to an embodiment of the present application Figure 4 This method is optionally applied to an electronic device, such as a personal computer, a server, a server cluster, and other electronic devices. It is understandable that the relevant content of the above Figures 1 to 5 shown method can also be applied to this example, and the relevant associated content will not be elaborated herein.
[0198] Furthermore, this method at least includes at least part of the following content. As Figure 6 shown, it includes:
[0199] Step S601: Using the pre-trained first feature processing network in the initial living body detection model, perform feature processing on multiple pieces of first-modal data to obtain a first fusion feature.
[0200] Here, the initial living body detection model includes a first detection branch for detecting the first-modal data. The first detection branch at least includes the pre-trained first feature processing network and a to-be-trained prompt module.
[0201] Step S602: Using the to-be-trained prompt module, perform feature description on the multiple pieces of first-modal data to obtain first description information.
[0202] Here, the first description information is used to describe the living body features of the target object in at least part of the multiple pieces of first-modal data.
[0203] It should be noted that for relevant examples of the to-be-trained prompt module and the first description information, reference can be made to the above description, and details will not be elaborated herein.
[0204] Step S603: Based on the first fusion feature and the first description information, predict a first living body detection result for the multiple pieces of first-modal data.
[0205] It should be noted that for relevant examples of the first living body detection result, reference can be made to the above description, and details will not be elaborated herein.
[0206] Step S604: Determine a second living body detection result.
[0207] Here, the initial living body detection model further includes a second detection branch for detecting second-modal data; the second living body detection result is obtained by using the second detection branch to detect multiple pieces of second-modal data.
[0208] It should be noted that for relevant examples of the second live detection result, reference can be made to the above description, which will not be elaborated here.
[0209] Step S605: Using the first live detection result and the second live detection result, while at least freezing the network parameters of the pre-trained first feature processing network, fine-tune the adjustable parameters in the to-be-trained prompt module.
[0210] Here, in one example, the adjustable parameters in the to-be-trained prompt module can be fine-tuned while freezing the network parameters in the pre-trained first feature processing network in the first detection branch. Further, in the pre-trained first detection branch, there may also be a pre-trained first prediction module. At this time, the adjustable parameters in the to-be-trained prompt module can be fine-tuned while freezing the pre-trained first feature processing network and the pre-trained first prediction module in the first detection branch.
[0211] Alternatively, in another example, the adjustable parameters in the to-be-trained prompt module can be fine-tuned while freezing the pre-trained first feature processing network in the first detection branch and the pre-trained second feature processing network in the second detection branch. Further, there may also be a pre-trained first prediction module in the first detection branch. Similarly, there may also be a pre-trained second prediction module in the second detection branch. At this time, the adjustable parameters in the to-be-trained prompt module can be fine-tuned while freezing the pre-trained first feature processing network and the pre-trained first prediction module in the first detection branch, and freezing the pre-trained second feature processing network and the pre-trained second prediction module in the second detection branch.
[0212] In this way, the solution of the present disclosure provides a refined training scheme, which can effectively avoid the problem of catastrophic forgetting of existing knowledge while the model is learning new knowledge. In other words, this training scheme can improve the accuracy of the description information output by the to-be-trained prompt module on the basis of effectively avoiding the problem of catastrophic forgetting. Thus, it lays a foundation for providing strong reference information in the prediction stage and also provides strong support for improving the accuracy of the live detection result.
[0213] In addition, since this training scheme is carried out while freezing some network parameters, it can also effectively improve the training efficiency of the model.
[0214] It should be noted that the "pre-training" in the solution of the present disclosure can be to train the corresponding module or network or model using an existing training method, and the present disclosure does not make specific limitations on the specific pre-training method.
[0215] The following is a detailed description of the present disclosure solution with specific examples. Specifically, taking the first-modal video as an NIR video and the second-modal video as an RGB video as an example, a detailed description of the training method of the initial live detection model is provided, which may include the following three parts:
[0216] The first part: Pretraining.
[0217] Figure 7 It is a schematic diagram of the training process of a preset detection model for processing RGB video frames in an embodiment of the present application.
[0218] It should be noted that multiple RGB sampled video frames can be determined from the RGB video based on a preset strategy. Here, the above-mentioned sampling method can be used for sampling to obtain the multiple RGB sampled video frames, and the present disclosure solution does not make specific limitations on this.
[0219] Furthermore, as Figure 7 shown, face detection (for example, using the face detection module 701 for face detection) and image preprocessing (for example, using the image preprocessing module 702 for image preprocessing) can be performed on each RGB sampled video frame among the multiple RGB sampled video frames to obtain each RGB video frame for model training. Here, it should be noted that the above-mentioned methods can be used for face detection and image preprocessing, and the present disclosure solution does not make specific limitations on this.
[0220] Furthermore, the RGB video frames are input into the feature extraction network to be trained to obtain the image features of each RGB video frame. Here, in one example, the feature extraction network can specifically be ViT-Base 703.
[0221] Furthermore, the image features of each RGB video frame are input into the feature fusion network 704 to be trained to obtain the RGB fusion feature.
[0222] Here, the structure of the feature fusion network is the same as that of the first fusion network or the second fusion network described above. For details, please refer to the above description and will not be elaborated here.
[0223] Furthermore, after the RGB fusion feature is input into the RGB data classification head 705, the RGB live detection result is obtained.
[0224] Furthermore, in one example, the loss value of the first loss can be calculated using the RGB live detection result and the RGB category label, and then the preset detection model can be optimized using the loss value of the first loss. Here, the first loss can specifically be the binary cross-entropy loss function.
[0225] The second part: Constructing the initial live detection model.
[0226] Figure 8 It is a schematic structural diagram of an initial live detection model constructed in an embodiment of the present application.
[0227] Here, after training a preset detection model using the first part of the training method, an initial live detection model as shown in Figure 8 can be constructed based on the trained preset detection model. For example, taking the trained preset detection model directly as a branch (e.g., the RGB detection branch 802), adding a to-be-trained prompt module to the trained preset detection model, and using the preset detection model containing the to-be-trained prompt module as another branch (e.g., the NIR detection branch 801) to construct an initial live detection model as shown in Figure 8 is shown.
[0228] Furthermore, as shown in Figure 8 , the two branches of the initial live detection model can be respectively called the NIR detection branch 801 and the RGB detection branch 802. Among them, the NIR detection branch 801 can be used to detect NIR video frames. Correspondingly, the RGB detection branch 802 can be used to detect RGB video frames.
[0229] Here, it should be noted that for the sake of distinction, the ViT-Base in the NIR detection branch 801 can be called the first ViT-Base 8011, and the feature fusion network can be called the first fusion network 8012. Similarly, the ViT-Base in the RGB detection branch 802 can be called the second ViT-Base 8021, and the feature fusion network can be called the second fusion network 8022.
[0230] Part Three: Training the constructed initial live detection model.
[0231] Here, the NIR sampled video frames are sampled from the NIR video. Similarly, the RGB sampled video frames are sampled from the RGB video. It should be noted that in one example, the above-mentioned sampling method can be used to obtain multiple NIR sampled video frames and multiple RGB sampled video frames.
[0232] Further, after sampling multiple NIR sampled video frames, face detection can be performed on each NIR sampled video frame (for example, using the first face detection module 803 for face detection) and image preprocessing (for example, using the first image preprocessing module 804 for image preprocessing) to obtain NIR video frames for model training. Similarly, after sampling multiple RGB sampled video frames, face detection can be performed on each RGB sampled video frame (for example, using the second face detection module 805 for face detection) and image preprocessing (for example, using the second image preprocessing module 806 for image preprocessing) to obtain RGB video frames for model training.
[0233] Here, it should be noted that in one example, the above-described method can be used for face detection and image preprocessing, and then multiple NIR video frames for model training and multiple RGB video frames for model training can be obtained.
[0234] Further, for the NIR detection branch 801, multiple NIR video frames can be input into the first ViT-Base 8011 to obtain image features of each NIR video frame. After processing the image features of each NIR video frame using the first fusion network 8012, NIR fusion features (such as NIR class token features) can be obtained, and the first description information of the live features of the target object in the multiple NIR video frames can be obtained using the to-be-trained prompt module 8013. Further, using the NIR data classification head 8014 and based on the first description information and the NIR fusion features, the NIR liveness detection result can be obtained. Based on the NIR liveness detection result, the loss value of the second loss can be obtained.
[0235] Here, in one example, the first description information and the NIR fusion features can be concatenated and the concatenated result can be input into the NIR data classification head 8014 to predict the NIR liveness detection result.
[0236] Alternatively, in another example, the first description information and the NIR fusion features can be separately input into the NIR data classification head 8014 to predict the NIR liveness detection result.
[0237] It should be noted that the above is only an exemplary display, and the present disclosure does not specifically limit the specific form of inputting the first description information and the NIR fusion features into the NIR data classification head.
[0238] Similarly, for the RGB detection branch 802, multiple RGB video frames can be input into the second ViT-Base 8021 to obtain the image features of each RGB video frame. After processing the image features of each RGB video frame using the second fusion network 8022, RGB fusion features (such as RGB class token features) are obtained. Further, using the RGB data classification head 8023 and based on the RGB fusion features, an RGB liveness detection result is obtained. Here, in one example, the loss value of the third loss can be directly obtained based on the RGB liveness detection result. Or, in another example, the loss value of the third loss can be obtained based on the RGB liveness detection result and the NIR liveness detection result.
[0239] Further, in one example, after directly obtaining the loss value of the third loss based on the RGB liveness detection result, the loss value of the second loss and the loss value of the third loss can be used, for example, to obtain the total loss value of the two through a weighted method. Then, based on the total loss value and with other network parameters frozen except for the adjustable parameters of the to-be-trained prompt module, the to-be-trained prompt module is trained, and thus the trained target liveness detection model is obtained.
[0240] Or, in another example, after obtaining the loss value of the third loss based on the RGB liveness detection result and the NIR liveness detection result, the loss value of the third loss can be used, and with other network parameters frozen except for the adjustable parameters of the to-be-trained prompt module, the to-be-trained prompt module is trained, and thus the trained target liveness detection model is obtained.
[0241] It should be noted that in this training method, the essence of training the initial liveness detection model is to fine-tune the adjustable parameters of the to-be-trained prompt module, so as to improve the liveness description ability of the to-be-trained prompt module.
[0242] That is to say, the solution of the present disclosure can use the to-be-trained prompt module to describe multiple input NIR video frames and obtain the first description information. For example, in one example, the first description information can be represented by n prompt tokens. At this time, during the training phase, the corresponding RGB video frames are added, and the RGB class token can be used to supervise the NIR class token features, so that the NIR features are similar to the trained RGB features, and at the same time, the description ability of the to-be-trained prompt module is improved.
[0243] In addition, the solution of the present disclosure can add the processing ability of NIR modal data on the premise of not degrading the performance of the RGB model, thus greatly improving the accuracy and generalization of the NIR modality in the hybrid modality.
[0244] The present disclosure also provides a method for live detection. Specifically, Figure 9 is a schematic flowchart of a method for live detection according to an embodiment of the present application. This method is optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices. It can be understood that the relevant content of the above Figures 1 to 8 The content of the method shown can also be applied to this example, and the relevant associated content will not be elaborated in this example.
[0245] Furthermore, this method at least includes at least some of the following content. As Figure 9 shown, it includes:
[0246] Step S901: Obtain a plurality of target video frames including the object to be detected.
[0247] Here, the plurality of target video frames include at least one first video frame of the first modality, and / or at least one second video frame of the second modality.
[0248] In other words, in one example, the plurality of target video frames may include at least one first video frame of the first modality. Or, in another example, the plurality of target video frames may include at least one second video frame of the second modality. Or, in another example, the plurality of target video frames may also include at least one first video frame of the first modality and at least one second video frame of the second modality at the same time.
[0249] Step S902: Input the plurality of target video frames into the target live detection model.
[0250] Here, the target live detection model includes a first detection branch for detecting data of the first modality, and a second detection branch for detecting data of the second modality.
[0251] Here, the target live detection model is obtained by training an initial live detection model. For example, it is obtained by training using the above training method. It should be noted that for relevant examples of the initial live detection model, reference can be made to the above description, and details will not be elaborated here.
[0252] Step S903: Obtain the target detection result.
[0253] Here, the target detection result is obtained based on the first detection branch and / or the second detection branch.
[0254] In this way, since the target live detection model in the present disclosure includes detection branches for different modalities, the target live detection model can adapt to the detection requirements under different conditions. Thus, the usage scenarios are enriched, and the user experience is improved.
[0255] In addition, since the target living body detection model can process video frames of different modalities simultaneously, it can effectively fuse the information differences between different modalities, thereby effectively improving the accuracy and reliability of the living body detection results.
[0256] Further, in a specific example, at least one first video frame of the first modality is obtained based on an NIR video;
[0257] and / or, at least one second video frame of the second modality is obtained based on an RGB video.
[0258] For example, in one example, the first video frame can specifically be an NIR video frame. Further, in another example, the second video frame can specifically be an RGB video frame. Or, in yet another example, the first video frame can specifically be an NIR video frame, and the second video frame can specifically be an RGB video frame.
[0259] In this way, the solution of the present disclosure can utilize the target living body detection model to achieve multi-modal living body detection, improving the generalization ability of the living body detection model, so that the living body detection model can be applicable to a wider range of application scenarios.
[0260] In addition, it should be noted that NIR videos are usually insensitive to light changes and can clearly image under low light conditions, while RGB videos can provide more detailed information about the color, texture, etc. of the target object. Since the solution of the present disclosure has the processing capabilities of both modalities, the living body detection model has stronger environmental adaptability and detection accuracy. For example, the living body detection model of the solution of the present disclosure reduces the sensitivity to a single light condition and can complete the living body detection task under different lighting environments.
[0261] For example, the living body detection method provided by the solution of the present disclosure can be applied to various fields such as attendance, access control, security, and financial payment in the field of living body recognition. In other words, the solution of the present disclosure can be widely applied to many current services.
[0262] Further, in a specific example, the target detection result can be obtained based on at least one of the following:
[0263] When all of the multiple target video frames belong to the first modality, the target detection result is obtained based on the first detection branch;
[0264] When all of the multiple target video frames belong to the second modality, the target detection result is obtained based on the second detection branch;
[0265] When some of the multiple target video frames belong to the first modality and the other part of the video frames belong to the second modality, the target detection result is obtained based on the first detection branch and the second detection branch.
[0266] That is to say, the solution of the present disclosure can select the corresponding detection branch for detection according to the modality to which the target video frame belongs. In this way, the flexibility of multi-modal live detection is effectively improved, and further, the target live detection model can provide accurate target detection results in a variety of different scenarios, thereby further improving the user experience.
[0267] In addition, when all of the multiple target video frames belong to a specific modality, the corresponding detection branch can be directly called, avoiding unnecessary computational overhead and improving the detection efficiency.
[0268] Further, in a specific example, the first detection branch at least includes a first feature processing network, a target prompt module, and a first prediction module;
[0269] The first feature processing network is used to obtain the first target fusion feature of the at least one first video frame;
[0270] The target prompt module is used to obtain the target description information for the multiple target video frames or for the at least one first video frame;
[0271] The first prediction module is used to predict the target detection result based on the first target fusion feature and the target description information.
[0272] Here, the target prompt module can be obtained by training the to-be-trained prompt module using the above training method.
[0273] In this way, the solution of the present disclosure can use the first feature processing network to obtain the first target fusion feature of the at least one first video frame, providing a data basis for live detection. At the same time, by combining the target description information provided by the target prompt module, the first prediction module can accurately determine whether the target object in the first video frame is a live body, thereby improving the accuracy of live detection.
[0274] Further, in a specific example, the first feature processing network includes a first feature extraction network and a first fusion network; wherein,
[0275] The first feature extraction network is used to obtain the image features of each first video frame in the at least one first video frame;
[0276] The first fusion network is used to obtain the first target fusion feature based on the image features of each first video frame in the at least one first video frame.
[0277] Here, the first fusion network can be used to splice the image features of each first video frame in the first video frame to obtain the first target splicing feature. Further, according to the importance degree of each feature unit in the first target splicing feature, the first splicing feature is pruned to obtain the first target fusion feature.
[0278] Here, it should be noted that the implementation manners of splicing the image features and pruning the splicing features can be referred to the above description and will not be elaborated here.
[0279] In addition, it should be noted that when there is only one first video frame, at this time, the image features of the extracted first video frame can be directly used as the first target fusion feature.
[0280] In this way, the solution of the present disclosure extracts features from each first video frame by using the first feature extraction network, so as to efficiently obtain the image features of each first video frame, improving the accuracy and efficiency of feature extraction.
[0281] Further, the solution of the present disclosure fuses the image features of each first video frame by using the first fusion network, thereby generating the first target fusion feature that can comprehensively and accurately reflect at least one first video frame. That is to say, the first target fusion feature carries the features of the target object in different first video frames. Therefore, it provides richer data for live detection, and further provides data support for efficient and accurate live detection.
[0282] Further, in a specific example, the second detection branch at least includes a second feature processing network and a second prediction module; wherein,
[0283] The second feature processing network is used to obtain the second target fusion feature of the at least one second video frame;
[0284] The second prediction module is used to predict the target detection result based on the second target fusion feature.
[0285] In this way, the solution of the present disclosure can obtain the second target fusion feature of the at least one second video frame by using the second feature processing network, and then predict the target detection result based on the second target fusion feature, providing strong support for realizing multi-modal live detection.
[0286] Further, in a specific example, the second feature processing network includes a second feature extraction network and a second fusion network; wherein,
[0287] The second feature extraction network is used to obtain the image features of each second video frame in the at least one second video frame;
[0288] The second fusion network is configured to obtain the second target fusion feature based on the image features of each second video frame in the at least one second video frame.
[0289] Here, the second fusion network can be used to splice the image features of each second video frame in the second video frame to obtain the second target splicing feature. Further, according to the importance of each feature unit in the second target splicing feature, pruning is performed on the second splicing feature to obtain the second target fusion feature.
[0290] Here, it should be noted that the implementation manners of splicing the image features and pruning the spliced features can be referred to the above description and will not be elaborated here.
[0291] In addition, it should be noted that when there is only one second video frame, at this time, the image features of the extracted second video frame can be directly used as the second target fusion feature.
[0292] In this way, the solution of the present disclosure extracts features from each second video frame by using the second feature extraction network, so as to efficiently obtain the image features of each second video frame, improving the accuracy and efficiency of feature extraction.
[0293] Further, the solution of the present disclosure uses the second fusion network to fuse the image features of each second video frame, thereby generating the second target fusion feature that can comprehensively and accurately reflect at least one second video frame. That is to say, the second target fusion feature carries the features of the target object in different second video frames. Therefore, it provides richer data for live detection, and further provides data support for efficient and accurate live detection.
[0294] The solution of the present disclosure also provides a training device 1000 for a live detection model, as Figure 10 shown, including:
[0295] A feature processing unit 1001, configured to perform feature processing on multiple first-modal data by using a first feature processing network pre-trained in an initial live detection model, to obtain a first fusion feature; wherein, the initial live detection model at least includes the pre-trained first feature processing network and a to-be-trained prompt module;
[0296] An information generation unit 1002, configured to perform feature description on the multiple first-modal data by using the to-be-trained prompt module, to obtain first description information; wherein, the first description information is used to describe the live features of the target object in at least some of the multiple first-modal data;
[0297] A prediction unit 1003, configured to predict a first live detection result for the multiple first-modal data based on the first fusion feature and the first description information;
[0298] A training unit 1004, configured to train the to-be-trained prompt module at least using the first live detection result, so as to obtain a target live detection model including a target prompt module.
[0299] In a specific example of the present disclosure solution, the feature processing unit 1001 is specifically configured to:
[0300] Use the first feature extraction network in the pre-trained first feature processing network to perform feature extraction on the multiple first-modal data, so as to obtain first initial features of the respective first-modal data;
[0301] Use the first fusion network in the pre-trained first feature processing network to perform feature fusion on the first initial features of the respective first-modal data, so as to obtain the first fusion feature.
[0302] In a specific example of the present disclosure solution, the feature processing unit 1001 is specifically configured to:
[0303] Concatenate the first initial features of the respective first-modal data to obtain a first concatenated feature;
[0304] Obtain first initial feature values of the respective feature units in the first concatenated feature, where the first initial feature value can characterize the importance degree of the feature unit in the first concatenated feature;
[0305] Perform pruning processing on the first concatenated feature based on the first initial feature values of the respective feature units, so as to obtain a pruned first concatenated feature, where the first fusion feature is the pruned first concatenated feature.
[0306] In a specific example of the present disclosure solution, the feature processing unit 1001 is specifically configured to:
[0307] Remove the feature units in the first concatenated feature whose first initial feature values are less than or equal to a first preset threshold.
[0308] In a specific example of the present disclosure solution, the pre-trained first feature processing network and the to-be-trained prompt module are on a first detection branch of the initial live detection model, and the first detection branch is used to detect the first-modal data.
[0309] In a specific example of the present disclosure solution, the initial live detection model further includes a second detection branch for detecting second-modal data;
[0310] The training unit 1004 is specifically configured to:
[0311] Train the to-be-trained prompt module by using the first live detection result and the second live detection result, where the second live detection result is obtained by detecting a plurality of second-modal data by using the second detection branch.
[0312] In a specific example of the present disclosure solution, the second detection branch includes a pre-trained second feature processing network;
[0313] The feature processing unit 1001 is further configured to perform feature processing on the plurality of second-modal data by using the pre-trained second feature processing network in the second detection branch to obtain a second fusion feature;
[0314] The prediction unit 1003 is further configured to predict the second live detection result for the plurality of second-modal data based on the second fusion feature.
[0315] In a specific example of the present disclosure solution, the feature processing unit 1001 is specifically configured to:
[0316] Perform feature extraction on the plurality of second-modal data by using a second feature extraction network in the pre-trained second feature processing network to obtain second initial features of the respective second-modal data;
[0317] Perform feature fusion on the second initial features of the respective second-modal data by using a second fusion network in the pre-trained second feature processing network to obtain the second fusion feature.
[0318] In a specific example of the present disclosure solution, the feature processing unit 1001 is specifically configured to:
[0319] Concatenate the second initial features of the respective second-modal data to obtain a second concatenated feature;
[0320] Obtain second initial feature values of the respective feature units in the second concatenated feature, where the second initial feature value can represent the importance degree of the feature unit in the second concatenated feature;
[0321] Perform pruning processing on the second concatenated feature based on the second initial feature values of the respective feature units to obtain a pruned second concatenated feature, where the second fusion feature is the pruned second concatenated feature.
[0322] In a specific example of the present disclosure solution, the feature processing unit 1001 is specifically configured to:
[0323] Remove the feature units in the second splicing feature whose second initial feature values are less than or equal to the second preset threshold.
[0324] In a specific example of the solution of the present disclosure, the training unit 1004 is specifically configured to:
[0325] Fine-tune the adjustable parameters in the to-be-trained prompt module while at least freezing the network parameters of the pre-trained first feature processing network.
[0326] In a specific example of the solution of the present disclosure, the multiple first-modal data are obtained based on a first-modal video; the multiple second-modal data are obtained based on a second-modal video; both the first-modal video and the second-modal video are obtained after video acquisition of the target object.
[0327] For the specific functions and example descriptions of the units of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated here.
[0328] The solution of the present disclosure further provides a live detection device 1100, as Figure 11 shown, including:
[0329] An acquisition unit 1101, configured to acquire a plurality of target video frames including an object to be detected; wherein, the plurality of target video frames include at least one first video frame of a first modality and / or at least one second video frame of a second modality;
[0330] An input unit 1102, configured to input the plurality of target video frames into a target live detection model; wherein, the target live detection model includes a first detection branch for detecting data of the first modality and a second detection branch for detecting data of the second modality;
[0331] An output unit 1103, configured to obtain a target detection result; wherein, the target detection result is obtained based on the first detection branch and / or the second detection branch.
[0332] In a specific example of the solution of the present disclosure, the target detection result is obtained based on one of the following:
[0333] When all of the plurality of target video frames belong to the first modality, the target detection result is obtained based on the first detection branch;
[0334] When all of the plurality of target video frames belong to the second modality, the target detection result is obtained based on the second detection branch;
[0335] In the case that some of the multiple target video frames belong to the first modality and some other video frames belong to the second modality, the target detection result is obtained based on the first detection branch and the second detection branch.
[0336] In a specific example of the present disclosure solution, the first detection branch at least includes a first feature processing network, a target prompt module, and a first prediction module;
[0337] The first feature processing network is configured to obtain a first target fusion feature of the at least one first video frame;
[0338] The target prompt module is configured to obtain target description information for the multiple target video frames or for the at least one first video frame;
[0339] The first prediction module is configured to predict the target detection result based on the first target fusion feature and the target description information.
[0340] In a specific example of the present disclosure solution, the first feature processing network includes a first feature extraction network and a first fusion network; wherein,
[0341] The first feature extraction network is configured to obtain image features of each first video frame in the at least one first video frame;
[0342] The first fusion network is configured to obtain the first target fusion feature based on the image features of each first video frame in the at least one first video frame.
[0343] In a specific example of the present disclosure solution, the second detection branch at least includes a second feature processing network and a second prediction module;
[0344] The second feature processing network is configured to obtain a second target fusion feature of the at least one second video frame;
[0345] The second prediction module is configured to predict the target detection result based on the second target fusion feature.
[0346] In a specific example of the present disclosure solution, the second feature processing network includes a second feature extraction network and a second fusion network; wherein,
[0347] The second feature extraction network is configured to obtain image features of each second video frame in the at least one second video frame;
[0348] The second fusion network is configured to obtain the second target fusion feature based on the image features of each second video frame in the at least one second video frame.
[0349] For the description of the specific functions and examples of the units of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.
[0350] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0351] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0352] Figure 12 FIG. shows a schematic block diagram of an exemplary electronic device 1200 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0353] As Figure 12 shown, the device 1200 includes a computing unit 1201, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1202 or the computer program loaded from the storage unit 1208 into the random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other through a bus 1204. The input / output (I / O) interface 1205 is also connected to the bus 1204.
[0354] A plurality of components in the device 1200 are connected to the I / O interface 1205, including: an input unit 1206, such as a keyboard, a mouse, etc.; an output unit 1207, such as various types of displays, speakers, etc.; a storage unit 1208, such as a disk, an optical disc, etc.; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 allows the device 1200 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0355] The computing unit 1201 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 executes the various methods and processes described above, such as the training method or the live detection method of the live detection model. For example, in some embodiments, the training method or the live detection method of the live detection model may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1208. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the training method or the live detection method of the live detection model described above may be executed. Alternatively, in other embodiments, the computing unit 1201 may be configured to execute the training method or the live detection method of the live detection model by any other suitable means (e.g., by means of firmware).
[0356] Various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor may be a dedicated or general-purpose programmable processor, which can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0357] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0358] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0359] For providing interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0360] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0361] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0362] It should be understood that various forms of the processes shown above may be used, steps may be reordered, added, or deleted. For example, the steps described in this disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0363] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A training method for a live detection model, comprising: Using a pre-trained first feature processing network in an initial live detection model to perform feature processing on multiple first-modal data to obtain a first fused feature; wherein, the initial live detection model at least includes the pre-trained first feature processing network and a to-be-trained hint module; Using the to-be-trained hint module to perform feature description on the multiple first-modal data to obtain first description information; wherein, the first description information is used to describe the live features of a target object in at least part of the multiple first-modal data; Based on the first fused feature and the first description information, predicting a first live detection result for the multiple first-modal data; Training the to-be-trained hint module at least using the first live detection result to obtain a target live detection model including a target hint module.
2. The method according to claim 1, wherein, The step of using a pre-trained first feature processing network in an initial live detection model to perform feature processing on multiple first-modal data to obtain a first fused feature includes: Using a first feature extraction network in the pre-trained first feature processing network to perform feature extraction on the multiple first-modal data to obtain first initial features of each first-modal data; Using a first fusion network in the pre-trained first feature processing network to perform feature fusion on the first initial features of each first-modal data to obtain the first fused feature.
3. The method according to claim 2, wherein The step of performing feature fusion on the first initial features of each first-modal data to obtain the first fused feature includes: Concatenating the first initial features of each first-modal data to obtain a first concatenated feature; Obtaining first initial feature values of each feature unit in the first concatenated feature, wherein the first initial feature value can represent the importance degree of the feature unit in the first concatenated feature; Based on the first initial feature values of each feature unit, performing pruning processing on the first concatenated feature to obtain a pruned first concatenated feature, wherein the first fused feature is the pruned first concatenated feature.
4. The method according to claim 3, wherein, The step of performing pruning processing on the first concatenated feature based on the first initial feature values of each feature unit includes: Removing feature units with first initial feature values less than or equal to a first preset threshold from the first concatenated feature.
5. The method according to any one of claims 1-4, wherein The pre-trained first feature processing network and the to-be-trained hint module are on a first detection branch of the initial live detection model, and the first detection branch is used to detect the first-modal data.
6. The method according to claim 5, wherein, The initial live detection model further includes a second detection branch for detecting second-modal data; Wherein, the step of training the to-be-trained hint module at least using the first live detection result includes: Training the to-be-trained hint module using the first live detection result and a second live detection result; wherein, the second live detection result is obtained by detecting multiple second-modal data using the second detection branch.
7. The method according to claim 6, wherein, The second detection branch includes a pre-trained second feature processing network; the method further includes: Using the pre-trained second feature processing network in the second detection branch, perform feature processing on the multiple second-modal data to obtain second fusion features; Based on the second fusion features, predict the second liveness detection results for the multiple second-modal data.
8. The method according to claim 7, wherein, The using the pre-trained second feature processing network in the second detection branch to perform feature processing on the multiple second-modal data to obtain second fusion features includes: Using the second feature extraction network in the pre-trained second feature processing network to perform feature extraction on the multiple second-modal data to obtain second initial features of each second-modal data; Using the second fusion network in the pre-trained second feature processing network to perform feature fusion on the second initial features of each second-modal data to obtain the second fusion features.
9. The method according to claim 8, wherein The performing feature fusion on the second initial features of each second-modal data to obtain the second fusion features includes: Concatenate the second initial features of each second-modal data to obtain a second concatenated feature; Obtain the second initial feature values of each feature unit in the second concatenated feature, where the second initial feature value can characterize the importance of the feature unit in the second concatenated feature; Based on the second initial feature values of each feature unit, perform pruning processing on the second concatenated feature to obtain a pruned second concatenated feature, where the second fusion features are the pruned second concatenated feature.
10. The method according to claim 9, wherein, The based on the second initial feature values of each feature unit to perform pruning processing on the second concatenated feature includes: Remove the feature units with second initial feature values less than or equal to a second preset threshold from the second concatenated feature.
11. According to the method according to any one of claims 1-10, wherein The training the to-be-trained prompt module includes: Fine-tune the adjustable parameters in the to-be-trained prompt module while at least freezing the network parameters of the pre-trained first feature processing network.
12. The method according to any one of claims 6-10, wherein, The multiple first-modal data are obtained based on a first-modal video; the multiple second-modal data are obtained based on a second-modal video; both the first-modal video and the second-modal video are obtained after video capturing of the target object.
13. A liveness detection method, including: Obtain multiple target video frames including a to-be-detected object; wherein, the multiple target video frames include at least one first video frame of a first modality and / or at least one second video frame of a second modality; Input the multiple target video frames into a target liveness detection model; wherein, the target liveness detection model includes a first detection branch for detecting data of the first modality and a second detection branch for detecting data of the second modality; Obtain target detection results; wherein, the target detection results are obtained based on the first detection branch and / or the second detection branch.
14. The method according to claim 13, wherein, The target detection results are obtained based on one of the following: When all of the multiple target video frames belong to the first modality, the target detection results are obtained based on the first detection branch; When all of the multiple target video frames belong to the second modality, the target detection result is obtained based on the second detection branch; When some of the multiple target video frames belong to the first modality and the other part of the video frames belong to the second modality, the target detection result is obtained based on the first detection branch and the second detection branch.
15. The method according to claim 14, wherein, The first detection branch at least includes a first feature processing network, a target prompt module, and a first prediction module; The first feature processing network is configured to obtain a first target fusion feature of the at least one first video frame; The target prompt module is configured to obtain target description information for the multiple target video frames or for the at least one first video frame; The first prediction module is configured to predict the target detection result based on the first target fusion feature and the target description information.
16. The method according to claim 15, wherein, The first feature processing network includes a first feature extraction network and a first fusion network; wherein, The first feature extraction network is configured to obtain image features of each first video frame in the at least one first video frame; The first fusion network is configured to obtain the first target fusion feature based on the image features of each first video frame in the at least one first video frame.
17. The method according to claim 13, wherein, The second detection branch at least includes a second feature processing network and a second prediction module; The second feature processing network is configured to obtain a second target fusion feature of the at least one second video frame; The second prediction module is configured to predict the target detection result based on the second target fusion feature.
18. The method according to claim 17, wherein The second feature processing network includes a second feature extraction network and a second fusion network; wherein, The second feature extraction network is configured to obtain image features of each second video frame in the at least one second video frame; The second fusion network is configured to obtain the second target fusion feature based on the image features of each second video frame in the at least one second video frame.
19. A training device for a living body detection model, comprising: A feature processing unit, configured to use a pre-trained first feature processing network in an initial living body detection model to perform feature processing on multiple first modality data to obtain a first fusion feature; wherein, the initial living body detection model at least includes the pre-trained first feature processing network and a to-be-trained prompt module; An information generation unit, configured to use the to-be-trained prompt module to perform feature description on the multiple first modality data to obtain first description information; wherein, the first description information is used to describe the living body features of target objects in at least part of the multiple first modality data; A prediction unit, configured to predict a first living body detection result for the multiple first modality data based on the first fusion feature and the first description information; A training unit, configured to at least use the first living body detection result to train the to-be-trained prompt module to obtain a target living body detection model including a target prompt module.
20. A living body detection device, comprising: An acquisition unit, configured to acquire a plurality of target video frames including an object to be detected; wherein, the plurality of target video frames include at least one first video frame in a first modality and / or at least one second video frame in a second modality; An input unit, configured to input the plurality of target video frames into a target liveness detection model; wherein, the target liveness detection model includes a first detection branch for detecting data in the first modality and a second detection branch for detecting data in the second modality; An output unit, configured to obtain a target detection result; wherein, the target detection result is obtained based on the first detection branch and / or the second detection branch.
21. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-18.
22. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-18.
23. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-18.