A living body detection method, device, storage medium and electronic device
By combining a cross-modal liveness detection model and a large-scale language tuning model, interpretable detection explanation information is generated, which solves the problem of uninterpretable liveness detection results in existing biometric systems, and improves user experience and detection success rate.
Patent Information
- Application Number
- CN202310571297.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-18
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2043-05-18
AI Technical Summary
In existing biometric systems, liveness detection methods are mainly black-box models, which cannot provide interpretable feedback. This leaves users unable to understand the reasons for liveness detection results or how to adjust them to improve the pass rate.
By employing a cross-modal liveness detection model combined with a large-scale language tuning model, liveness attack detection is performed by acquiring target facial images, generating detection explanation information, and outputting it to the user to guide condition adjustment.
It improves the user experience and pass rate of liveness detection, optimizes the liveness detection process, and provides interpretable feedback to help users effectively adjust to improve the success rate of detection.
Smart Images

Figure CN116778586B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a method, apparatus, storage medium and electronic device for detecting liveness. Background Technology
[0002] With the rapid development of computer technology, biometric technology has been widely applied to people's production and daily lives. For example, facial recognition payment, facial access control, facial attendance, and facial recognition station entry all rely on biometrics. However, as biometric technology becomes more widely used, the need for liveness detection in biometric scenarios is becoming increasingly prominent. While biometrics provides convenience, it also brings new risks and challenges. The most common means of threatening the security of biometric systems is liveness attacks, which involve attempting to bypass image biometric verification through means such as device screens or printed photos. Therefore, liveness detection is particularly important in biometric scenarios. Summary of the Invention
[0003] This specification provides a method, apparatus, storage medium, and electronic device for liveness detection, the technical solution of which is as follows:
[0004] Firstly, this specification provides a method for detecting live organisms, the method comprising:
[0005] Obtain the target facial image corresponding to the target object;
[0006] Liveness detection results are obtained by performing liveness attack detection based on the target facial image, and detection interpretation information is obtained by performing detection interpretation processing on the liveness detection results.
[0007] The detection explanation information is output to the target object to prompt the target object to adjust the liveness detection conditions.
[0008] Secondly, this specification provides a method for training a cross-modal liveness detection model, the method comprising:
[0009] Create an initial cross-modal liveness detection model and obtain a first sample training dataset, which includes at least one sample face image labeled with a detection type label and a text description label;
[0010] Based on the sample facial images, the initial cross-modal liveness detection model is trained for at least one round to obtain the sample liveness detection results and sample image description text features for the sample facial images;
[0011] Based on the sample liveness detection results, the detection type label, the sample image description text features, and the text description label, the model parameters of the initial cross-modal liveness detection model are adjusted until the initial cross-modal liveness detection model completes model training, thus obtaining the cross-modal liveness detection model.
[0012] Thirdly, this specification provides a method for interpreting and evaluating model training, the method comprising:
[0013] Create an initial interpretation and evaluation model to obtain the textual features of sample images generated by the cross-modal liveness detection model during the model training phase;
[0014] The textual features describing the sample images are input into a large-scale language model to obtain the reference image detection and explanation information.
[0015] Determine the evaluation label for the detection interpretation information corresponding to the detection interpretation information of the reference image;
[0016] Based on the reference image detection interpretation information, the initial interpretation evaluation model is trained for at least one round to determine the sample perturbation detection interpretation information corresponding to the reference image detection interpretation information, and to obtain the image detection interpretation features, sample perturbation detection interpretation features, and detection interpretation information evaluation results.
[0017] Based on the image detection interpretation features, the sample perturbation detection interpretation features, the detection interpretation information evaluation labels, and the detection interpretation information evaluation results, the model parameters of the initial interpretation evaluation model are adjusted until the initial interpretation evaluation model completes model training, thus obtaining the interpretation evaluation model.
[0018] Fourthly, this specification provides a method for training a large-scale language tuning model, the method comprising:
[0019] Create an initial large-scale language tuning model, which is based on an existing large-scale language model;
[0020] The sample image description text features generated by the cross-modal liveness detection model during the model training phase are obtained, and the detection interpretation information evaluation labels are generated by the interpretation evaluation model during the model training phase.
[0021] The sample image description text features and the detection explanation information evaluation labels are input into the initial large-scale language tuning model for at least one round of model training to obtain the sample image detection explanation information and the reference tuning network layer.
[0022] The model parameters of the reference tuning network layer of the initial large-scale language tuning model are adjusted until the initial large-scale language tuning model completes model training, thus obtaining the large-scale language tuning model.
[0023] Fifthly, this specification provides a liveness detection device, the device comprising:
[0024] The image acquisition module is used to acquire the target facial image corresponding to the target object;
[0025] The liveness detection module is used to perform liveness attack detection based on the target facial image to obtain liveness detection results, and to perform detection interpretation processing on the liveness detection results to obtain detection interpretation information;
[0026] The information prompting module is used to output the detection explanation information to the target object, so as to prompt the target object to adjust the liveness detection conditions through the detection explanation information.
[0027] Sixthly, this specification provides a cross-modal liveness detection model training device, the device comprising:
[0028] The data processing module is used to create an initial cross-modal liveness detection model and obtain a first sample training dataset, which includes at least one sample face image labeled with a detection type label and a text description label.
[0029] The model training module is used to train the initial cross-modal liveness detection model at least once based on the sample facial images to obtain the sample liveness detection results and sample image description text features for the sample facial images.
[0030] The parameter adjustment module is used to adjust the model parameters of the initial cross-modal liveness detection model based on the sample liveness detection results, the detection type label, the sample image description text features, and the text description label, until the initial cross-modal liveness detection model completes model training and a cross-modal liveness detection model is obtained.
[0031] Seventhly, this specification provides an apparatus for interpreting and evaluating model training, the apparatus comprising:
[0032] The model creation module is used to create an initial interpretation and evaluation model and obtain the sample image description text features generated by the cross-modal liveness detection model during the model training phase.
[0033] The model training module is used to input the sample image description text features into a large-scale language model to obtain reference image detection explanation information, determine the detection explanation information evaluation label corresponding to the reference image detection explanation information, and perform at least one round of model training based on the reference image detection explanation information input into the initial explanation evaluation model to determine the sample perturbation detection explanation information corresponding to the reference image detection explanation information, and obtain image detection explanation features, sample perturbation detection explanation features and detection explanation information evaluation results;
[0034] The parameter adjustment module is used to adjust the model parameters of the initial interpretation evaluation model based on the image detection interpretation features, the sample perturbation detection interpretation features, the detection interpretation information evaluation labels, and the detection interpretation information evaluation results, until the initial interpretation evaluation model completes model training and an interpretation evaluation model is obtained.
[0035] Eighthly, this specification provides a large-scale language tuning model training device, the device comprising:
[0036] The model creation module is used to create an initial large-scale language tuning model, which is created based on an existing large-scale language model.
[0037] The data acquisition module is used to acquire the sample image description text features generated by the cross-modal liveness detection model during the model training phase, and the detection interpretation information evaluation labels for the reference image detection interpretation information. The reference image detection interpretation information is the reference image detection interpretation information generated by the interpretation evaluation model during the model training phase.
[0038] The model training module is used to input the sample image description text features and the detection explanation information evaluation labels into the initial large-scale language tuning model for at least one round of model training, so as to obtain the sample image detection explanation information and the reference tuning network layer.
[0039] The parameter adjustment module is used to adjust the model parameters of the reference tuning network layer of the initial large-scale language tuning model until the initial large-scale language tuning model completes model training and obtains the large-scale language tuning model.
[0040] Ninthly, this specification provides a computer storage medium storing at least one instruction adapted to be loaded by a processor and to execute method steps of one or more embodiments of this specification.
[0041] Tenthly, this specification provides a computer program product storing at least one instruction adapted to be loaded by a processor and to execute method steps of one or more embodiments of this specification.
[0042] Eleventhly, this specification provides an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute method steps of one or more embodiments of this specification.
[0043] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:
[0044] In one or more embodiments of this specification, a liveness detection result is obtained by performing liveness attack detection based on the target facial image of the target object. The liveness detection result is then processed to obtain detection interpretation information, which can be output to the target object. This allows the target object to be prompted to adjust liveness detection conditions, avoiding the limitation of only knowing the liveness detection result after liveness detection. By forming interpretable feedback of the liveness detection process through detection interpretation information, the liveness detection process is optimized, resulting in better liveness detection performance. The formation and output of detection interpretation information improves the user experience of liveness detection, and to some extent, increases the pass rate of liveness detection. It also plays a positive role in the subsequent optimization of the liveness detection method. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a schematic diagram of a liveness detection system provided in this manual;
[0047] Figure 2 This is a flowchart illustrating a liveness detection method provided in this manual;
[0048] Figure 3 This is a schematic diagram of a liveness detection scenario provided in this manual;
[0049] Figure 4 This is a schematic flowchart of another embodiment of the liveness detection method provided in this specification;
[0050] Figure 5 This is a flowchart illustrating one embodiment of a cross-modal liveness detection model training method provided in this specification;
[0051] Figure 6This is a flowchart illustrating one embodiment of the explanatory evaluation model training method provided in this specification;
[0052] Figure 7 This is a flowchart illustrating one embodiment of the large-scale language tuning model training method provided in this specification;
[0053] Figure 8 This is a schematic diagram of the structure of a liveness detection device provided in this specification;
[0054] Figure 9 This is a schematic diagram of the structure of a model training device provided in this manual;
[0055] Figure 10 This is a schematic diagram of the structure of a model training device provided in this specification;
[0056] Figure 11 This is a schematic diagram of the structure of a model training device provided in this specification;
[0057] Figure 12 This is a schematic diagram of the structure of an electronic device provided in this specification. Detailed Implementation
[0058] The technical solutions in this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0059] In the description of this specification, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this specification based on the specific circumstances. Furthermore, in the description of this specification, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0060] Liveness detection attacks are attacks where attackers attempt to bypass facial recognition systems using facial materials of a "victim" (such as mobile phone photos, printed photos, masks, etc.). With the widespread use of facial recognition systems, the threat of liveness detection attacks has become increasingly serious. Therefore, various liveness detection methods have been proposed to address the complex and ever-changing nature of these attacks. However, almost all current liveness detection methods are black-box models; after performing liveness detection on an image, only the result of the detection is known, which has certain limitations.
[0061] The present specification will now be described in detail with reference to specific embodiments.
[0062] Please see Figure 1 This is a schematic diagram of a liveness detection system provided in this specification. Figure 1 As shown, the liveness detection system may include at least a client cluster and a service platform 100.
[0063] The client cluster may include at least one client, such as Figure 1 As shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.
[0064] Each client in a client cluster can be an electronic device with communication capabilities, including but not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolved networks.
[0065] The service platform 100 can be a standalone server device, such as a rack-mount, blade, tower, or cabinet-type server device, or a workstation, mainframe, or other hardware device with strong computing power; or it can be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, wherein each server is functionally and hierarchically equivalent in the transaction chain, and each server can provide services independently. The independent provision of services can be understood as not requiring the assistance of other servers.
[0066] In one or more embodiments of this specification, the service platform 100 may establish a communication connection with at least one client in the client cluster, and complete the data interaction during the liveness detection process based on the communication connection.
[0067] It should be noted that the service platform 100 establishes a communication connection with at least one client in the client cluster via a network for interactive communication. This network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0068] The liveness detection system embodiments provided in this specification and the liveness detection methods described in one or more embodiments belong to the same concept. The execution entity corresponding to the liveness detection method involved in one or more embodiments of this specification can be the aforementioned service platform 100; the execution entity corresponding to the liveness detection method involved in one or more embodiments of this specification can be the service platform or the electronic device corresponding to the client, depending on the actual application environment. The implementation process of the liveness detection system embodiments can be detailed in the following method embodiments, and will not be repeated here.
[0069] based on Figure 1 The following is a detailed description of the liveness detection method provided by one or more embodiments of this specification, as illustrated in the schematic diagram of the scenario.
[0070] Please see Figure 2This specification provides a flowchart illustrating a liveness detection method according to one or more embodiments. This method can be implemented using a computer program and can run on a liveness detection device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. The liveness detection device can be a client application.
[0071] Specifically, the liveness detection method includes:
[0072] Understandably, liveness detection is a method used in identity verification scenarios to determine the true physiological characteristics of an object. In image liveness detection applications, it is necessary to verify whether the acquired target liveness detection image represents a real living object. Image liveness detection needs to effectively resist common liveness attack methods such as photos, face swaps, masks, occlusions, and screen captures, thereby helping to identify fraudulent activities and protect users' rights.
[0073] S102: Obtain the target facial image corresponding to the target object;
[0074] As an illustration, in actual liveness detection scenarios, the target facial image can be acquired using devices such as RGB cameras, monocular cameras, and infrared cameras, based on the corresponding liveness detection task.
[0075] Please see Figure 3 This is a schematic diagram of a liveness detection scenario provided in an embodiment of this specification. Figure 3 As shown, the client is a device used for identity authentication in the target scenario. It is equipped with an image acquisition device. When a target object approaches the client and is within the image acquisition range of the image acquisition device, the image acquisition device will capture a facial image of the target object (e.g., ...). Figure 3 The user's facial image is shown, and liveness detection is performed based on the acquired target facial image.
[0076] In this specification, the image modality type of the target facial image is not limited. The image modality type can be a fit of one or more of the following image modality types: video modality type, color image modality type (RGB) type, short video modality type, animation modality type, depth image modality type, infrared image modality type, near-infrared modality type, etc.
[0077] S104: Perform liveness detection based on the target facial image to obtain liveness detection results, and perform detection interpretation processing on the liveness detection results to obtain detection interpretation information;
[0078] To combat liveness detection attacks, it's crucial to address the issue of black-box models used in current technologies, enabling interpretable feedback for liveness detection. Interpretable feedback is essential for improving the user experience (providing guidance on how to adjust settings to bypass the attack) and optimizing liveness detection.
[0079] The detection interpretation information can be understood as providing a reasonable semantic interpretation of the liveness detection result, that is, result interpretation information. For example, the detection interpretation information is similar to: For example, because there is a mask in this image, the liveness detection result is a liveness attack type. For example, the main object in this image belongs to the printed photo type, so the liveness detection result is a liveness attack type.
[0080] In one feasible implementation, a liveness detection model can be pre-built and trained based on a machine learning model. The liveness detection model is then used to perform liveness attack detection processing based on multimodal image combination to obtain liveness detection results.
[0081] In some embodiments, the liveness detection result may carry one or more image description text features for the target facial image. The image description text features can provide detectable explanatory information during the liveness detection process. Subsequently, based on these image description text features, detection explanatory information is automatically generated for the liveness detection process.
[0082] Indicatively, the process of interpreting the liveness detection results to obtain detection interpretation information can be as follows:
[0083] Determine the target detection type corresponding to the liveness detection result;
[0084] If the target detection type is an attack type, then facial recognition interception processing is performed on the target object, and detection interpretation processing is performed on the liveness detection result to obtain detection interpretation information.
[0085] If the target detection type is liveness, then the facial recognition is successful.
[0086] Optionally, the liveness detection results can be interpreted by using a large-scale language model to process the image description text features carried by the liveness detection results, and outputting a reasonable semantic interpretation, i.e., detection interpretation information, for each liveness detection decision.
[0087] S106: Output the detection explanation information to the target object to prompt the target object to adjust the liveness detection conditions through the detection explanation information.
[0088] Specifically, the client can output the detection explanation information to the target object, for example, the detection explanation information can be output as text information, the detection explanation information can be output as audio information, and so on.
[0089] Indicatively, outputting detection interpretation information to the target object can serve as a prompt, reminding the target object to adjust the current liveness detection conditions. For example, if the detection interpretation information indicates that the target is identified as an attack type because the subject is wearing a mask, the target object can adjust the liveness detection conditions based on this detection interpretation information, such as removing the mask and performing liveness detection again.
[0090] The liveness detection conditions can be the state conditions of the subject (such as whether it is wearing a mask, whether its face is covered, etc.), the environmental conditions for liveness detection (such as lighting, background brightness, etc.).
[0091] As an illustration, after prompting the target object to adjust the liveness detection conditions through the detection interpretation information, the following steps may also be performed:
[0092] Considering that target users can further adjust the liveness detection conditions based on the detection interpretation information to meet the requirements for passing the liveness detection;
[0093] The client can obtain the next target facial image corresponding to the target object, use the next target facial image as the target facial image, and perform the step of obtaining the liveness detection result based on the target facial image for liveness attack detection.
[0094] In one or more embodiments of this specification, the limitation of only knowing the liveness detection result after liveness detection is avoided. Interpretive feedback on the liveness detection process is formed in the form of detection explanation information, optimizing the liveness detection process and achieving better liveness detection results. The generation and output of detection explanation information improves the user experience of liveness detection, and to some extent, it can also increase the pass rate of liveness detection, and also plays a positive role in the subsequent optimization of the liveness detection method. In the liveness detection scenario, detection explanation information is determined for the detection decision process of each liveness detection result, so that users can adjust the liveness detection conditions based on some prompts in the detection explanation information, such as adjusting the environment, which can improve the liveness detection effect and streamline the liveness detection process.
[0095] Optionally, the process of performing liveness detection based on the target facial image to obtain liveness detection results and performing detection interpretation processing on the liveness detection results to obtain detection interpretation information can be carried out using the following steps;
[0096] Please see Figure 4 , Figure 4This is a schematic flowchart of another embodiment of a liveness detection method proposed in one or more embodiments of this specification. Specifically:
[0097] S202: Based on the target facial image, a cross-modal liveness detection model is used to perform liveness detection processing to obtain liveness detection results, and the image description text features for the target facial image are determined through the cross-modal liveness detection model;
[0098] In some scenarios, existing liveness detection models simply extract image features from images, and these features have no correlation or correspondence with the descriptive text accompanying the image. Due to this characteristic, they are typically unsuitable as direct input to large-scale language models. To mitigate this limitation, this specification proposes a cross-modal liveness detection model. During model training, while performing liveness detection training, the model leverages both a cross-modal pre-trained model and the output of a large-scale language model to enhance the textual relevance of the image features output by the liveness detection model.
[0099] Understandably, cross-modal liveness detection models integrate the capabilities of large-scale language models, cross-modal image-to-text translation, liveness feature extraction, and feature encoding during the model training phase.
[0100] In one feasible implementation, a cross-modal liveness detection model can be pre-built and trained based on a machine learning model. The liveness detection model is then used to perform liveness attack detection processing on the target facial image to obtain the liveness detection result.
[0101] Optionally, the liveness detection result may carry one or more image description text features for the target facial image. The image description text features can provide detectable explanatory information during the liveness detection process. Subsequently, based on these image description text features, detection explanatory information is automatically generated for the liveness detection process.
[0102] Furthermore, in the network structure of the cross-modal liveness detection model, the image description text features can be of various types, such as image description text features characterized from the dimension of cross-modal image-text association characteristics, or image description text features characterized from the dimension of text generation characteristics of large-scale language models.
[0103] Optionally, the cross-modal liveness detection model includes a liveness feature processing network, a cross-modal pre-trained model association network, a large-scale language model association network, and a cross-modal feature transformation network;
[0104] The large-scale language model association network can be a model network layer obtained based on an existing large-scale language model. Examples of large-scale language models include ChatGPT and the Large Language Model (LLM).
[0105] In one feasible implementation, the step of performing liveness detection processing based on the target facial image using a cross-modal liveness detection model to obtain liveness detection results, and determining image description text features for the target facial image using the cross-modal liveness detection model, can be:
[0106] The target facial image is input into a cross-modal liveness detection model, and the target facial image is processed by a liveness feature processing network to obtain liveness detection results and image liveness features;
[0107] Furthermore, text association transformation is performed based on image liveness features through a cross-modal pre-trained model association network, a large-scale language model association network, and a cross-modal feature transformation network to obtain a first image description text feature based on the cross-modal pre-trained model association network and a second image description text feature based on the large-scale language model association network.
[0108] Furthermore, the first image description text feature and the second image description text feature are image description text features extracted from different image description dimensions. The first image description text feature is extracted from the dimension of cross-modal text-image association, which represents the image's liveness detection characteristics, because the cross-modal liveness detection model is configured with a cross-modal association part, that is, a machine learning model part that associates images with corresponding text. The second image description text feature is generated by the cross-modal liveness detection model, which is configured with a description text generation part based on an existing large-scale language model, by converting the image liveness features across modalities into text descriptions of the image, thereby generating image description text features representing the image's liveness detection characteristics with the help of a large-scale language model.
[0109] In this specification, the image-text correlation of the model output is enhanced by using a cross-modal pre-trained model that spans from image to text dimensions, as well as the output of a large-scale language model.
[0110] S204: Based on the image description text features, a large-scale language optimization model is used to generate detection explanation information. The large-scale language optimization model is a large-scale language optimization model obtained by adapting a large-scale language model to a liveness detection scenario.
[0111] The large-scale language optimization model is the model obtained by adapting an existing large-scale language model to the liveness detection scenario. The existing large-scale language model has not been optimized for the liveness detection task, so it cannot fit the liveness detection scenario and cannot produce good detection interpretation information output.
[0112] Indicatively, based on the output of the cross-modal liveness detection model from the previous step, a large-scale language model can be used to output detection explanation information, that is, the relevant detection cause analysis of this liveness detection.
[0113] In one feasible implementation, the step of generating detection explanation information based on the image description text features using a large-scale language tuning model can be:
[0114] The client can input the first image description text features and the second image description text features into a large-scale language tuning model for interpretation and generation, so as to obtain the first detection interpretation information corresponding to the first image description text features and the second detection interpretation information corresponding to the second image description text features through the large-scale language tuning model;
[0115] The client can perform interpretation and evaluation processing based on the first and second detection interpretation information using an interpretation evaluation model to obtain the detection interpretation information.
[0116] To illustrate, if the input of a large-scale language tuning model is two image-description text features, then it typically outputs two detection explanations. Here, we use an explanation evaluation model trained based on a machine learning model to further make text decisions, and the better evaluation among these detection explanations is taken as the final detection explanation.
[0117] In one or more embodiments of this specification, a method is proposed for optimizing a cross-modal liveness detection model by combining it with an evaluation model. During the model training process, while training for liveness detection, the textual relevance of the image features output by the liveness detection model is enhanced by using the output of a cross-modal pre-trained model and a large-scale language tuning model. This allows the large-scale language model to be further optimized for the liveness detection task, resulting in a better interpretable output.
[0118] Please see Figure 5 , Figure 5 This is a flowchart illustrating one embodiment of a cross-modal liveness detection model training method proposed in one or more embodiments of this specification. The method can be implemented using a computer program and can run on a liveness detection device based on the von Neumann architecture. This computer program can be integrated into an application or run as a standalone utility application. The liveness detection device can be a service platform or a client. Specifically:
[0119] S302: Create an initial cross-modal liveness detection model and obtain a first sample training dataset, wherein the first sample training data includes at least one sample face image labeled with a detection type label and a text description label;
[0120] In one or more embodiments of this specification, in response to a liveness detection task, an initial cross-modal liveness detection model can be created based on a machine learning model, a model can be trained to obtain a cross-modal liveness detection model, and liveness detection can be performed using the cross-modal liveness detection model;
[0121] During the data annotation phase, the first sample training dataset is obtained, and each sample facial image in the first sample training dataset is labeled with a detection type label, that is, whether the sample facial image belongs to the attack type or the liveness type. In addition, the expert-end service is used to customize a brief text description for each sample facial image as a text description label.
[0122] In some implementations, existing liveness detection models simply extract features from images, and these features have no correlation or correspondence with the text. Due to this characteristic, they cannot be directly used as input for subsequent large-scale language models. This paper innovatively proposes a cross-modal compatible liveness detection model training method. While training for liveness detection, it enhances the textual correlation of image features output by the liveness detection model through cross-modal pre-trained models and the output of large-scale language models.
[0123] Optionally, the initial cross-modal liveness detection model may include, but is not limited to, a liveness feature processing network, a cross-modal pre-trained model association network, a large-scale language model association network, and a cross-modal feature transformation network;
[0124] S304: Based on the sample facial image, perform at least one round of model training on the initial cross-modal liveness detection model to obtain the sample liveness detection result and the sample image description text features for the sample facial image;
[0125] Indicatively, sample facial images can be input into an initial cross-modal liveness detection model for at least one round of model training. Each round of model training can generate sample liveness detection results and sample image description text features for the sample facial images.
[0126] A2: The sample facial image is processed by the liveness feature processing network to obtain the sample liveness detection result and sample image liveness features;
[0127] The input to the liveness feature processing network is a sample facial image. The network extracts liveness features from the sample facial image to obtain liveness features, and performs liveness detection based on the liveness features to obtain the sample liveness detection result.
[0128] Among them, the liveness features of the sample image can include fitting one or more of the feature types of the object image, such as texture features, color features, shape features, and spatial relationship features.
[0129] A4: The cross-modal pre-trained model association network is used to perform cross-modal association processing on the sample facial image and the text description label to obtain the first sample image features and the first sample image description text features;
[0130] Optionally, the cross-modal pre-trained model association network can be based on the association network of a pre-trained text encoder and image encoder;
[0131] Optionally, the model structure and parameters of the cross-modal pre-trained model's interconnected network can be kept unchanged during the model training phase;
[0132] The input to the cross-modal pre-trained model association network is a sample facial image and a text description label. The cross-modal pre-trained model association network performs cross-modal association between the image features of the sample facial image and the text features of the text description label. Here, the text description label directly instructs the network to extract text features that conform to the expected direction of the sample facial image during the forward propagation training process. Since the text description label is an accurate description of the sample facial image, it can enhance the correlation between cross-modal image features and text features. The output of the cross-modal pre-trained model association network is the first sample image features and the first sample image description text features.
[0133] The first sample image feature is the image feature corresponding to the sample face image, and the first sample image description text feature is the text feature of the text description label;
[0134] A6: The text description labels are processed by feature extraction through the large-scale language model association network to obtain the text features of the second sample image description;
[0135] The large-scale language model association network can be a model module created based on an existing large-scale language model, such as ChatGPT.
[0136] Since the large-scale language model association network is a pre-trained model network, its input and output can be configured. The text description label is used as the input of the large-scale language model association network, and the second image description text features are extracted from the text description label through the large-scale language model association network.
[0137] A8: Based on the liveness features of the sample images, cross-modal transformation processing is performed through the cross-modal feature transformation network to obtain the third sample image description text features for the cross-modal pre-trained model association network and the fourth sample image description text features for the large-scale language model association network.
[0138] The input of the cross-modal feature conversion network is the liveness features of the sample image. Here, the cross-modal feature conversion network is used to convert the liveness features of the sample image into image description text features for the sample image. The output is the third sample image description text features for the cross-modal pre-trained model association network and the fourth sample image description text features for the large-scale language model association network.
[0139] S306: Based on the sample liveness detection results, the detection type label, the sample image description text features, and the text description label, adjust the model parameters of the initial cross-modal liveness detection model until the initial cross-modal liveness detection model completes model training, and obtain the cross-modal liveness detection model.
[0140] Optionally, the model comprehensive loss of the initial cross-modal liveness detection model is calculated based on the sample liveness detection results, the detection type label, the sample image description text features, and the text description label, and the model parameters are adjusted during backpropagation training based on the model comprehensive loss;
[0141] Optionally, the model's overall loss may include liveness classification loss, liveness feature consistency loss, and cross-modal consistency loss;
[0142] Furthermore, the liveness classification loss is determined based on the sample liveness detection results and the detection type label;
[0143] The liveness classification loss serves as a supervisory signal for liveness detection classification, supervising the model training effect from the perspective of liveness classification. The liveness classification loss can be obtained by using a correlation loss calculation function based on the sample liveness detection results and the detection type label.
[0144] For example, a binary classification loss function can be used to calculate the liveness classification loss, as follows:
[0145]
[0146] Wherein, the Loss cls The loss can be defined as the liveness classification loss, where pred is the liveness detection result of the sample and y is the detection type label; CrossEntropy() represents the binary classification loss calculation.
[0147] Furthermore, a liveness feature consistency loss is determined based on the liveness features of the sample image and the features of the first sample image;
[0148] The liveness feature consistency loss serves as a consistency supervision signal. It is used to ensure that the features output by the liveness feature extraction module and the image features output by the cross-modal pre-trained model are as consistent as possible. The liveness feature consistency loss can be obtained by using a correlation loss calculation function based on the liveness features of the sample image and the features of the first sample image.
[0149] For example, the Euclidean distance loss function can be used to calculate the liveness feature consistency loss, as follows:
[0150]
[0151] Wherein, the Loss img This can be a loss of consistency in living characteristics, where f is... liveness-img and f crossm-img These are the liveness features of the sample image and the features of the first sample image, respectively.
[0152] Furthermore, based on the first sample image description text features, the third sample image description text features, the second sample image description text features, and the fourth sample image description text features, a cross-modal consistency loss is determined;
[0153]
[0154] Wherein, the Loss tex It is the cross-modal consistency loss, f in the two fractions. liveness-text These are the first sample image description text features and the second sample image description text features output by the cross-modal transformation model, respectively, where f crossm-text The third sample image output by the cross-modal pre-trained model association network describes the text features, wherein f llm-text It is the fourth sample image describing the text features output by the large-scale language model association network;
[0155] Furthermore, the model parameters of the initial cross-modal liveness detection model are adjusted based on the liveness classification loss, the liveness feature consistency loss, and the cross-modal consistency loss.
[0156] In one or more embodiments of this specification, a training method for a liveness detection model based on cross-modal compatibility is proposed. While training for liveness detection, the textual correlation of the image features output by the liveness detection model is enhanced by using the output of a cross-modal pre-trained model and a large-scale language model. Based on the output of the cross-modal liveness detection model in the previous step, the relevant causes can be analyzed.
[0157] Please see Figure 6 , Figure 6This is a flowchart illustrating one embodiment of an explanatory evaluation model training method proposed in one or more embodiments of this specification. The method can be implemented using a computer program and can run on a von Neumann-based liveness detection device. This computer program can be integrated into an application or run as a standalone utility application. The liveness detection device can be a service platform or a client. Specifically:
[0158] In some scenarios, existing large-scale language models, lacking optimization for liveness detection tasks, fail to produce satisfactory outputs. To address this, this specification trains an interpretation evaluation model. This model ranks one or more outputs from large-scale language models to obtain superior outputs as detection interpretation information.
[0159] S402: Create an initial interpretation and evaluation model to obtain the sample image description text features generated by the cross-modal liveness detection model during the model training phase;
[0160] In one or more embodiments of this specification, in response to a liveness detection task, an initial interpretation evaluation model can be created based on a machine learning model, the model can be trained to obtain an interpretation evaluation model, and the interpretation evaluation model can be used to sort one or more outputs of a large-scale language model.
[0161] Optionally, the initial interpretation evaluation model includes a text perturbation module, a feature extraction module, and a result comparison module;
[0162] S404: Input the sample image description text features into a large-scale language model to obtain reference image detection explanation information, and determine the detection explanation information evaluation label corresponding to the reference image detection explanation information;
[0163] Optionally, the number of sample image description text features corresponds to the number of reference image detection explanation information. For example, if the number of sample image description text features is 1, then the number of reference image detection explanation information is 1.
[0164] In some embodiments, the sample image description text features are obtained by inputting the training image data of the initial interpretation evaluation model into the cross-modal liveness detection model after the model training phase is completed, thereby obtaining the third sample image description text features for the cross-modal pre-trained model association network and the fourth sample image description text features for the large-scale language model association network.
[0165] In a schematic way, the text features describing the third sample image and the text features describing the fourth sample image are input into a large-scale language model to obtain two reference image detection explanations, namely the first reference image detection explanation and the second reference image detection explanation.
[0166] Data annotation: The detection interpretation information of the first reference image and the detection interpretation information of the second reference image are provided to the expert side for selection and adjustment to determine the description that is more in line with the liveness detection task (two texts are displayed, and the expert chooses which one is more appropriate), thereby obtaining the detection interpretation information evaluation label;
[0167] S406: Based on the reference image detection interpretation information, input the initial interpretation evaluation model to perform at least one round of model training, so as to determine the sample perturbation detection interpretation information corresponding to the reference image detection interpretation information, and obtain the image detection interpretation features, sample perturbation detection interpretation features and detection interpretation information evaluation results;
[0168] During the forward propagation of model training, reference image detection interpretation information is input into the initial interpretation evaluation model for model training. The initial interpretation evaluation model randomly perturbs the reference image detection interpretation information to obtain sample perturbation detection interpretation information and extracts the sample perturbation detection interpretation features corresponding to the sample perturbation detection interpretation information. On the other hand, it extracts the image detection interpretation features corresponding to the reference image detection interpretation information. Furthermore, it obtains the detection interpretation information evaluation result based on the reference image detection interpretation information. Typically, there are multiple reference image detection interpretation information, for example, the final detection interpretation information evaluation result is determined from the aforementioned first reference image detection interpretation information and second reference image detection interpretation information.
[0169] In one feasible implementation, the initial training process for the explanatory evaluation model is as follows:
[0170] B2: Based on the reference image detection interpretation information, input the initial interpretation evaluation model for at least one round of model training, and obtain sample perturbation detection interpretation information by perturbing the reference image detection interpretation information through the text perturbation module;
[0171] The input to the text perturbation module is the reference image detection interpretation information. The text perturbation module performs random perturbation processing on the reference image detection interpretation information (such as adding random Gaussian noise) to obtain the corresponding sample perturbation detection interpretation information.
[0172] Perturbation of the detection and interpretation information of the input reference image during model training can improve the training effect of the model against interference from complex environments.
[0173] B4: The feature extraction module extracts the image detection interpretation features corresponding to the reference image detection interpretation information and the sample perturbation detection interpretation features corresponding to the sample perturbation detection interpretation information, respectively.
[0174] The input to the feature extraction module is reference image detection interpretation information and sample perturbation detection interpretation information. The feature extraction module extracts features from the reference image detection interpretation information to obtain corresponding image detection interpretation features, and extracts features from the sample perturbation detection interpretation information to obtain corresponding sample perturbation detection interpretation features.
[0175] Optionally, the feature extraction module may employ a relevant feature encoder;
[0176] B6: The image detection interpretation features are evaluated and compared using the result comparison module to obtain the detection interpretation information evaluation result.
[0177] The result comparison module performs evaluation and comparison processing based on image detection interpretation features to obtain the evaluation result of detection interpretation information.
[0178] Schematic, the reference image detection interpretation information includes first reference image detection interpretation information and second reference image detection interpretation information, and the image detection interpretation features include first image detection interpretation features corresponding to the first reference image detection interpretation information and second image detection interpretation features corresponding to the second reference image detection interpretation information;
[0179] Optionally, the step of evaluating and comparing the image detection interpretation features through the result comparison module to obtain the detection interpretation information evaluation result can be:
[0180] The result comparison module evaluates and compares the first image detection interpretation features and the second image detection interpretation features to obtain a detection interpretation information evaluation result, which is one of the first reference image detection interpretation information and the second reference image detection interpretation information.
[0181] S408: Based on the image detection interpretation features, the sample perturbation detection interpretation features, the detection interpretation information evaluation label, and the detection interpretation information evaluation result, adjust the model parameters of the initial interpretation evaluation model until the initial interpretation evaluation model completes model training to obtain the interpretation evaluation model.
[0182] Optionally, the model comprehensive loss of the initial interpretation evaluation model is calculated based on the image detection interpretation features, the sample perturbation detection interpretation features, the detection interpretation information evaluation label, and the detection interpretation information evaluation result, and the model parameters are adjusted during the backpropagation training process based on the model comprehensive loss;
[0183] Optionally, the model integration loss may include perturbation consistency loss and result comparison loss;
[0184] C2: Determine the perturbation consistency loss based on the image detection interpretation features and the sample perturbation detection interpretation features;
[0185] The perturbation consistency loss serves as a supervisory signal for evaluating classification, supervising the model training effect from the dimension of perturbation consistency. The liveness classification loss can be obtained by using the correlation loss calculation function based on the image detection explanatory features and the sample perturbation detection explanatory features.
[0186] For example, the Euclidean distance loss function can be used to calculate the perturbation consistency loss;
[0187]
[0188] LOSS dis It is a perturbation consistency constraint (for the same text, the features before and after the perturbation should be as consistent as possible), f ori f represents the image detection interpretation features. dir This indicates the perturbation detection interpretation features of the perturbated samples;
[0189] C4: Determine the result comparison loss based on the evaluation labels and the evaluation results of the detection interpretation information;
[0190]
[0191] The second part represents the classification loss. The model's classification result is consistent with the human-annotated result.
[0192] C6: Adjust the model parameters of the initial interpretation evaluation model based on the perturbation consistency loss and the result comparison loss.
[0193] The scene properties are described to generate a model, and the model parameters are adjusted.
[0194] Understandably, the model comprehensive loss for the initial interpretation and evaluation model can be obtained based on the perturbation consistency loss and the result comparison loss. The model parameters of the initial interpretation and evaluation model are then adjusted using the model backpropagation method based on the model comprehensive loss until the model ends training conditions, thus obtaining the trained interpretation and evaluation model.
[0195] The model's training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. Specific training termination conditions can be determined based on actual circumstances and are not specifically limited here.
[0196] In one or more embodiments of this specification, considering that large-scale language models cannot produce good output because they have not been optimized for liveness detection tasks, an interpretation evaluation model and its training method are presented. The purpose of this model is to evaluate the output of large-scale language models, determine the optimal output, and improve the model's interpretation performance.
[0197] Please see Figure 7 , Figure 7 This is a flowchart illustrating one embodiment of a large-scale language tuning model training method proposed in one or more embodiments of this specification. The method can be implemented using a computer program and can run on a liveness detection device based on the von Neumann architecture. This computer program can be integrated into an application or run as a standalone utility application. The liveness detection device can be a service platform or a client. Specifically:
[0198] S602: Create an initial large-scale language tuning model, which is created based on an existing large-scale language model;
[0199] In one or more embodiments of this specification, in response to a liveness detection task, an initial large-scale language tuning model can be created based on a machine learning model and the aforementioned trained interpretation and evaluation model and an existing large-scale language tuning model. The model can then be trained to obtain a large-scale language tuning model, which can be used for targeted tuning in new scenarios.
[0200] This section utilizes an evaluation model to perform targeted optimization of a large-scale language model. Because large-scale language models require a large amount of training resources to train, an adaptive progressive local large-scale model training optimization method is proposed to optimize them more efficiently and with lower resource consumption.
[0201] In one feasible implementation, the initial large-scale language tuning model may include a large-scale language model layer, an explanation evaluation model layer, and a network layer screening layer.
[0202] The large-scale language model layer can be a model layer built on an existing large-scale language model;
[0203] The explanation and evaluation model layer can be a model layer built based on the aforementioned trained explanation and evaluation model;
[0204] S604: Obtain the sample image description text features generated by the cross-modal liveness detection model during the model training phase, and the detection interpretation information evaluation labels for the reference image detection interpretation information, wherein the reference image detection interpretation information is the reference image detection interpretation information generated by the interpretation evaluation model during the model training phase;
[0205] In illustrative terms, the sample image descriptive text features can be the first sample image descriptive text features and the second sample image descriptive text features generated by the cross-modal liveness detection model during the model training phase;
[0206] S606: Input the sample image description text features and the detection explanation information evaluation labels into the initial large-scale language tuning model for at least one round of model training to obtain the sample image detection explanation information and the reference tuning network layer;
[0207] Indicatively, during each round of forward propagation model training, the sample image description text features and the detection explanation information evaluation labels can be input into the initial large-scale language tuning model for model processing, which can obtain the sample image detection explanation information and the reference tuning network layer;
[0208] In one feasible implementation, the following illustrates a model training method:
[0209] D2: The large-scale language model layer is used to detect and interpret the text features describing the sample image to obtain the first sample image detection and interpretation information;
[0210] The large-scale language model layer can be a model layer built on existing large-scale language models, such as ChatGPT, large language model LLM, GPT-3 model, etc.
[0211] D4: The first sample image detection interpretation information and the detection interpretation information evaluation label are interpreted and evaluated through the interpretation evaluation model layer to obtain the sample image detection interpretation information;
[0212] The explanation and evaluation model layer can be a model layer built based on the aforementioned trained explanation and evaluation model, and the model structure and model parameters of the explanation and evaluation model layer can be kept unchanged during the model training phase.
[0213] D6: The reference tuning network layer is determined from the large-scale language model layer based on the sample image detection and interpretation information through the network layer screening layer.
[0214] The network layer selection layer can be regarded as an adaptive layer selection model to select or determine the reference network layer to be trained from the large-scale language model layers, so as to perform targeted model parameter tuning on the reference network layer, while the model structure or model parameters of the other network layers remain unchanged.
[0215] S608: Adjust the model parameters of the reference tuning network layer of the initial large-scale language tuning model until the initial large-scale language tuning model completes model training, and obtain the large-scale language tuning model.
[0216] The model integration loss of the initial large-scale language tuning model can be composed of two loss dimensions: evaluation quality loss and sparse selection constraint loss.
[0217] In one feasible implementation, the model loss can be calculated in the following form:
[0218] 1) Determine the evaluation quality loss based on the sample image detection interpretation information and the evaluation label of the detection interpretation information;
[0219]
[0220] Among them, the LOSS compare To evaluate the quality loss, CrossEntropy represents the calculation of cross-entropy loss, where pred represents the sample image detection explanation information. CrossEntropy(pred,1) means that the output of the large language model - the sample image detection explanation information - should be classified into a better class each time, so that the output description of the large language model is better than the description of the manually labeled detection explanation information.
[0221] 2) Determine the sparse selection constraint loss based on the reference tuning network layer;
[0222]
[0223] Among them, the LOSS sparse For the sparse selection constraint loss, weights represent the network weights of the selected reference tuned network layer, and LOSS sparse As a sparse selection supervision signal, the L1 norm of the selected reference tuning network layer (which can be one or more layers, referring to the sum of weights here) should be as small as possible, so as to indicate that the adaptive layer selection model selects as few layers as possible during backpropagation.
[0224] 3) Adjust the model parameters of the reference tuning network layer of the initial large-scale language tuning model based on the evaluation quality loss and the sparse selection constraint loss.
[0225] Understandably, the model comprehensive loss for the initial large-scale language tuning model can be obtained based on the evaluation quality loss and the sparse selection constraint loss. Based on the model comprehensive loss, the model parameters of the reference tuning network layer of the initial large-scale language tuning model are adjusted by the model backpropagation method until the model ends training conditions are met, and then the trained large-scale language tuning model can be obtained.
[0226] The model's training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. Specific training termination conditions can be determined based on actual circumstances and are not specifically limited here.
[0227] In one or more embodiments of this specification, an evaluation model is used to perform targeted tuning of a large-scale language model. Because large-scale language models require a large amount of training resources to train, an adaptive progressive local large-scale model training optimization method is proposed in order to optimize them more efficiently and with low cost. The large-scale language tuning model trained based on this method can improve the model's processing performance and enhance its robustness.
[0228] The following will combine Figure 8 This manual provides a detailed description of the live animal detection device provided. It should be noted that... Figure 8 The liveness detection device shown is used to perform the functions described in this manual. Figures 1-7 The methods of the embodiments shown are illustrated only in the parts relevant to this specification for ease of explanation. For specific technical details not disclosed, please refer to this specification. Figures 1-7 The example shown.
[0229] Please see Figure 8 This diagram illustrates the structure of the liveness detection device described in this specification. The liveness detection device 1 can be implemented as all or part of a device through software, hardware, or a combination of both. According to some embodiments, the liveness detection device 1 includes an image acquisition module 11, a liveness detection module 12, and an information prompting module 13, specifically used for:
[0230] Image acquisition module 11 is used to acquire the target facial image corresponding to the target object;
[0231] The liveness detection module 12 is used to perform liveness attack detection based on the target facial image to obtain liveness detection results, and to perform detection interpretation processing on the liveness detection results to obtain detection interpretation information;
[0232] The information prompting module 13 is used to output the detection explanation information to the target object, so as to prompt the target object to adjust the liveness detection conditions through the detection explanation information.
[0233] Optionally, the liveness detection module 12 is used for:
[0234] Based on the target facial image, a cross-modal liveness detection model is used to perform liveness detection processing to obtain liveness detection results, and the image description text features for the target facial image are determined through the cross-modal liveness detection model.
[0235] Based on the image description text features, a large-scale language optimization model is used to generate detection explanation information. The large-scale language optimization model is a large-scale language optimization model obtained by adapting a large-scale language model to a liveness detection scenario.
[0236] Optionally, the cross-modal liveness detection model includes a liveness feature processing network, a cross-modal pre-trained model association network, a large-scale language model association network, and a cross-modal feature transformation network. The liveness detection module 12 is used for:
[0237] The target facial image is processed by a liveness detection network to obtain liveness detection results and image liveness features;
[0238] By performing text association transformation based on the image liveness features through the cross-modal pre-trained model association network, the large-scale language model association network, and the cross-modal feature transformation network, a first image description text feature based on the cross-modal pre-trained model association network and a second image description text feature based on the large-scale language model association network are obtained.
[0239] Optionally, the liveness detection module 12 is used for:
[0240] The first image description text features and the second image description text features are input into a large-scale language tuning model for interpretation and generation, so as to obtain the first detection interpretation information corresponding to the first image description text features and the second detection interpretation information corresponding to the second image description text features;
[0241] Based on the first detection interpretation information and the second detection interpretation information, an interpretation evaluation model is used to perform interpretation evaluation processing to obtain the detection interpretation information.
[0242] Optionally, the liveness detection module 12 is used for:
[0243] Determine the target detection type corresponding to the liveness detection result;
[0244] If the target detection type is an attack type, then facial recognition interception processing is performed on the target object, and detection interpretation processing is performed on the liveness detection result to obtain detection interpretation information.
[0245] If the target detection type is liveness, then the facial recognition is successful.
[0246] Optionally, the device 1 is further configured to: acquire a next target facial image corresponding to the target object, use the next target facial image as the target facial image, and perform the step of obtaining a liveness detection result by performing liveness attack detection based on the target facial image.
[0247] Please see Figure 9 This diagram illustrates the structure of a model training device for liveness detection as described in this specification. The liveness detection device 2 can be implemented as all or part of a device through software, hardware, or a combination of both. According to some embodiments, the liveness detection device 2 includes a data processing module 21, a model training module 22, and a parameter adjustment module 23, specifically used for:
[0248] Data processing module 21 is used to create an initial cross-modal liveness detection model and obtain a first sample training dataset, wherein the first sample training data includes at least one sample face image labeled with a detection type label and a text description label;
[0249] Model training module 22 is used to train the initial cross-modal liveness detection model at least once based on the sample facial images to obtain sample liveness detection results and sample image description text features for the sample facial images.
[0250] The parameter adjustment module 23 is used to adjust the model parameters of the initial cross-modal liveness detection model based on the sample liveness detection results, the detection type label, the sample image description text features, and the text description label, until the initial cross-modal liveness detection model completes model training and a cross-modal liveness detection model is obtained.
[0251] Optionally, the initial cross-modal liveness detection model includes a liveness feature processing network, a cross-modal pre-trained model association network, a large-scale language model association network, and a cross-modal feature transformation network.
[0252] The model training module 22 is used for:
[0253] The liveness detection process is performed on the sample facial image using the liveness feature processing network to obtain the sample liveness detection result and sample image liveness features;
[0254] The cross-modal pre-trained model association network is used to perform cross-modal association processing on the sample facial image and the text description label to obtain the first sample image features and the first sample image description text features;
[0255] The text description labels are processed by feature extraction through the large-scale language model association network to obtain the text features of the second sample image description.
[0256] Based on the liveness features of the sample images, cross-modal transformation processing is performed through the cross-modal feature transformation network to obtain the third sample image descriptive text features for the cross-modal pre-trained model association network and the fourth sample image descriptive text features for the large-scale language model association network.
[0257] Optionally, the model training module 23 is used for:
[0258] The liveness classification loss is determined based on the sample liveness detection results and the detection type label;
[0259] The liveness feature consistency loss is determined based on the liveness features of the sample image and the features of the first sample image.
[0260] Based on the text features describing the first sample image, the text features describing the third sample image, the text features describing the second sample image, and the text features describing the fourth sample image, the cross-modal consistency loss is determined.
[0261] The model parameters of the initial cross-modal liveness detection model are adjusted based on the liveness classification loss, the liveness feature consistency loss, and the cross-modal consistency loss.
[0262] Please see Figure 10 This diagram illustrates the structure of the model training device for liveness detection as described in this specification. The liveness detection device 3 can be implemented as all or part of a device through software, hardware, or a combination of both. According to some embodiments, the liveness detection device 3 includes a model creation module 31, a model training module 32, and a parameter adjustment module 33, specifically used for:
[0263] Model creation module 31 is used to create an initial interpretation and evaluation model and obtain sample image description text features generated by the cross-modal liveness detection model during the model training phase.
[0264] The model training module 32 is used to input the sample image description text features into a large-scale language model to obtain reference image detection explanation information, determine the detection explanation information evaluation label corresponding to the reference image detection explanation information, and perform at least one round of model training based on the reference image detection explanation information input into an initial explanation evaluation model to determine the sample perturbation detection explanation information corresponding to the reference image detection explanation information, and obtain image detection explanation features, sample perturbation detection explanation features and detection explanation information evaluation results;
[0265] The parameter adjustment module 33 is used to adjust the model parameters of the initial interpretation evaluation model based on the image detection interpretation features, the sample perturbation detection interpretation features, the detection interpretation information evaluation label, and the detection interpretation information evaluation result, until the initial interpretation evaluation model completes model training and an interpretation evaluation model is obtained.
[0266] Optionally, the initial interpretation and evaluation model includes a text perturbation module, a feature extraction module, and a result comparison module.
[0267] The model training module 32 is used for
[0268] Based on the reference image detection interpretation information, the initial interpretation evaluation model is trained for at least one round. The reference image detection interpretation information is perturbed by the text perturbation module to obtain sample perturbation detection interpretation information.
[0269] The feature extraction module extracts the image detection interpretation features corresponding to the reference image detection interpretation information and the sample perturbation detection interpretation features corresponding to the sample perturbation detection interpretation information, respectively.
[0270] The image detection interpretation features are evaluated and compared using the result comparison module to obtain the evaluation result of the detection interpretation information.
[0271] Optionally, the reference image detection interpretation information includes first reference image detection interpretation information and second reference image detection interpretation information, and the image detection interpretation features include first image detection interpretation features corresponding to the first reference image detection interpretation information and second image detection interpretation features corresponding to the second reference image detection interpretation information.
[0272] The model training module 32 is used for:
[0273] The result comparison module evaluates and compares the first image detection interpretation features and the second image detection interpretation features to obtain a detection interpretation information evaluation result, which is one of the first reference image detection interpretation information and the second reference image detection interpretation information.
[0274] Optionally, the parameter adjustment module 33 is used for:
[0275] The perturbation consistency loss is determined based on the image detection interpretation features and the sample perturbation detection interpretation features;
[0276] Based on the evaluation labels of the detection interpretation information and the evaluation results of the detection interpretation information, the result comparison loss is determined;
[0277] The model parameters of the initial interpretation evaluation model are adjusted based on the perturbation consistency loss and the result comparison loss.
[0278] Please see Figure 11 This diagram illustrates the structure of a model training device for liveness detection as described in this specification. The liveness detection device 4 can be implemented as all or part of a device through software, hardware, or a combination of both. According to some embodiments, the liveness detection device 3 includes a model creation module 41, a data acquisition module 42, a model training module 43, and a parameter adjustment module 44, specifically used for:
[0279] Model creation module 41 is used to create an initial large-scale language tuning model, which is created based on an existing large-scale language model;
[0280] Data acquisition module 42 is used to acquire sample image description text features generated by the cross-modal liveness detection model during the model training phase, and detection interpretation information evaluation labels for the reference image detection interpretation information. The reference image detection interpretation information is the reference image detection interpretation information generated by the interpretation evaluation model during the model training phase.
[0281] The model training module 43 is used to input the sample image description text features and the detection explanation information evaluation labels into the initial large-scale language tuning model for at least one round of model training, so as to obtain the sample image detection explanation information and the reference tuning network layer.
[0282] The parameter adjustment module 44 is used to adjust the model parameters of the reference tuning network layer of the initial large-scale language tuning model until the initial large-scale language tuning model completes model training and obtains the large-scale language tuning model.
[0283] Optionally, the initial large-scale language tuning model includes a large-scale language model layer, an explanation and evaluation model layer, and a network layer selection layer.
[0284] The model training module 43 is used for:
[0285] The large-scale language model layer is used to detect and interpret the text features describing the sample image to obtain the first sample image detection and interpretation information.
[0286] The first sample image detection interpretation information and the detection interpretation information evaluation label are interpreted and evaluated through the interpretation evaluation model layer to obtain the sample image detection interpretation information.
[0287] The reference tuning network layer is determined from the large-scale language model layer based on the sample image detection and interpretation information through the network layer screening layer.
[0288] Optionally, the parameter adjustment module 44 is used for:
[0289] The evaluation quality loss is determined based on the sample image detection interpretation information and the evaluation label of the detection interpretation information.
[0290] The sparse selection constraint loss is determined based on the aforementioned reference tuning network layer;
[0291] The model parameters of the reference tuning network layer of the initial large-scale language tuning model are adjusted based on the evaluation quality loss and the sparse selection constraint loss.
[0292] It should be noted that the liveness detection device provided in the above embodiments is only illustrated by the division of the above functional modules when performing the liveness detection method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the liveness detection device and the liveness detection method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.
[0293] The serial numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0294] This specification also provides a computer storage medium capable of storing multiple instructions adapted to be loaded and executed by a processor as described above. Figures 1-7 The liveness detection method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-7 The specific details of the illustrated embodiments will not be elaborated here.
[0295] This specification also provides a computer program product that stores at least one instruction, said at least one instruction being loaded and executed by the processor as described above. Figures 1-7 The liveness detection method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-7 The specific details of the illustrated embodiments will not be elaborated here.
[0296] Please refer to Figure 12 This is a structural block diagram of an electronic device provided in an embodiment of this specification. The electronic device in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, memory 120, input device 130, and output device 140 can be connected via the bus 150.
[0297] Processor 110 may include one or more processing cores. Processor 110 connects to various parts of the terminal using various interfaces and lines, and performs various functions and processes data of terminal 100 by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). Processor 110 may integrate one or more of a central processing unit (CPU), graphics processing unit (GPU), and modem. The CPU mainly handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem is used for wireless communication. It is understood that the modem may also not be integrated into processor 110, but implemented separately through a communication chip.
[0298] The memory 120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 120 may include non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets.
[0299] The input device 130 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 140 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In this embodiment, the input device 130 can be a temperature sensor to obtain the operating temperature of the terminal. The output device 140 can be a speaker to output audio signals.
[0300] In addition, those skilled in the art will understand that the structure of the terminal shown in the above figures does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the terminal may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WIFI) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.
[0301] In the embodiments of this specification, the executing entity for each step can be the terminal described above. Optionally, the executing entity for each step is the terminal's operating system. The operating system can be Android, iOS, or other operating systems; this specification does not limit this.
[0302] exist Figure 12 In the electronic device, the processor 110 can be used to call the liveness detection program stored in the memory 120 and execute it to implement the liveness detection method and / or model training method as described in the various method embodiments of this specification.
[0303] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.
[0304] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, prompt data, etc.) and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the object characteristics, interaction behavior characteristics and user information involved in this specification were all obtained under full authorization.
[0305] The above-disclosed embodiments are merely preferred embodiments of this specification and should not be construed as limiting the scope of this specification. Therefore, any equivalent variations made in accordance with the claims of this specification shall still fall within the scope of this specification.
Claims
1. A live detection method, the method comprising: obtaining a target face image corresponding to a target object; performing live detection processing based on the target face image using a cross-modal live detection model comprising a live feature processing network, a cross-modal pre-training model association network, a large-scale language model association network, and a cross-modal feature conversion network to obtain a live detection result; obtaining an image live feature through the live feature processing network; performing text association conversion based on the image live feature to obtain a first image description text feature based on the cross-modal pre-training model association network and a second image description text feature based on the large-scale language model association network, the first image description text feature being an image description text feature characterized from a cross-modal image text association characteristic dimension through the cross-modal pre-training model association network, and the second image description text feature being an image description text feature characterized from a large-scale language model text generation characteristic dimension through the large-scale language model association network; using a large-scale language tuning model based on the first image description text feature and the second image description text feature to obtain first detection explanation information and second detection explanation information, respectively, and performing explanation evaluation processing based on the first detection explanation information and the second detection explanation information to obtain detection explanation information, the large-scale language tuning model being a large-scale language tuning model obtained by adapting a large-scale language model to a live detection scenario; outputting the detection explanation information to the target object to prompt the target object to adjust live detection conditions through the detection explanation information.
2. The method of claim 1, wherein using a large-scale language tuning model based on the first image description text feature and the second image description text feature to obtain first detection explanation information and second detection explanation information, respectively, and performing explanation evaluation processing based on the first detection explanation information and the second detection explanation information to obtain detection explanation information comprises: inputting the first image description text feature and the second image description text feature into a large-scale language tuning model to generate explanations to obtain first detection explanation information corresponding to the first image description text feature and second detection explanation information corresponding to the second image description text feature; performing explanation evaluation processing based on the first detection explanation information and the second detection explanation information using an explanation evaluation model to obtain detection explanation information.
3. The method of claim 2, wherein performing explanation evaluation processing based on the first detection explanation information and the second detection explanation information using an explanation evaluation model to obtain detection explanation information comprises: performing text decision based on the first detection explanation information and the second detection explanation information to obtain more optimal evaluation detection explanation information from the first detection explanation information and the second detection explanation information.
4. The method of claim 1, wherein after obtaining the live detection result, the method further comprises: determining a target detection type corresponding to the live detection result. If the target detection type is a living type, it is determined that face recognition passes; If the target detection type is an attack type, face recognition interception processing is performed on the target object, and the following steps are executed: Obtaining image living features through the living feature processing network, performing text association conversion based on the image living features, obtaining first image description text features based on the cross-modal pre-training model association network and second image description text features based on the large-scale language model association network, the first image description text features being image description text features represented from the cross-modal image text association characteristic dimension through the cross-modal pre-training model association network, the second image description text features being image description text features represented from the large-scale language model text generation characteristic dimension through the large-scale language model association network, using a large-scale language tuning model based on the first image description text features and the second image description text features to obtain first detection explanation information and second detection explanation information respectively, performing explanation evaluation processing based on the first detection explanation information and the second detection explanation information to obtain detection explanation information, the large-scale language tuning model being a large-scale language tuning model obtained by adapting a large-scale language model to a living detection scene, and outputting the detection explanation information to the target object to prompt the target object to adjust the living detection condition through the detection explanation information.
5. The method of claim 1, further comprising, after prompting the target object to adjust the living detection condition through the detection explanation information: Obtaining a next target face image corresponding to the target object, taking the next target face image as the target face image, and performing the step of obtaining a living detection result based on the target face image for living attack detection.
6. The method of claim 1, comprising: Creating an initial cross-modal living detection model, obtaining a first sample training data set, the first sample training data including at least one sample face image with a labeled detection type label and a text description label; Based on the sample face image, performing at least one round of model training on the initial cross-modal living detection model to obtain a sample living detection result and a sample image description text feature for the sample face image; Based on the sample living detection result, the detection type label, the sample image description text feature, and the text description label, adjusting the model parameters of the initial cross-modal living detection model until the model training of the initial cross-modal living detection model is completed, obtaining a cross-modal living detection model.
7. The method of claim 6, wherein the initial cross-modal living detection model includes a living feature processing network, a cross-modal pre-training model association network, a large-scale language model association network, and a cross-modal feature conversion network. The sample face image in the first sample training data set is used to perform at least one round of model training on the initial cross-modal liveness detection model to obtain a sample liveness detection result and a sample image description text feature for the sample face image, including: performing liveness detection processing on the sample face image through the liveness feature processing network to obtain a sample liveness detection result and a sample image liveness feature; performing cross-modal association processing on the sample face image and the text description label through the cross-modal pre-training model association network to obtain a first sample image feature and a first sample image description text feature; performing feature extraction processing on the text description label through the large-scale language model association network to obtain a second sample image description text feature; performing cross-modal conversion processing on the sample image liveness feature through the cross-modal feature conversion network to obtain a third sample image description text feature for the cross-modal pre-training model association network and a fourth sample image description text feature for the large-scale language model association network.
8. The method of claim 7, wherein the model parameter adjustment of the initial cross-modal liveness detection model based on the sample liveness detection result, the detection type label, the sample image description text feature, and the text description label comprises: determining a liveness classification loss based on the sample liveness detection result and the detection type label; determining a liveness feature consistency loss based on the sample image liveness feature and the first sample image feature; determining a cross-modal consistency loss based on the first sample image description text feature and the third sample image description text feature, the second sample image description text feature, and the fourth sample image description text feature; adjusting the model parameters of the initial cross-modal liveness detection model based on the liveness classification loss, the liveness feature consistency loss, and the cross-modal consistency loss.
9. The method of claim 1, wherein the method comprises: creating an initial explanation evaluation model to obtain a sample image description text feature generated by a cross-modal liveness detection model during model training; inputting the sample image description text feature into a large-scale language model to obtain reference image detection explanation information; determining a detection explanation information evaluation label corresponding to the reference image detection explanation information; inputting the reference image detection explanation information into the initial explanation evaluation model to perform at least one round of model training to determine a sample perturbation detection explanation information corresponding to the reference image detection explanation information, and to obtain image detection explanation features, sample perturbation detection explanation features, and a detection explanation information evaluation result; adjusting the model parameters of the initial explanation evaluation model based on the image detection explanation features, the sample perturbation detection explanation features, the detection explanation information evaluation label, and the detection explanation information evaluation result until the model training of the initial explanation evaluation model is completed to obtain an explanation evaluation model.
10. The method of claim 9, wherein the initial interpretation evaluation model comprises a text perturbation module, a feature extraction module, and a result comparison module. The method further comprises: inputting the reference image detection interpretation information into the initial interpretation evaluation model to perform at least one round of model training, to determine sample perturbation detection interpretation information corresponding to the reference image detection interpretation information, and to obtain image detection interpretation features, sample perturbation detection interpretation features, and detection interpretation information evaluation results, comprising: inputting the reference image detection interpretation information into the initial interpretation evaluation model to perform at least one round of model training, and obtaining sample perturbation detection interpretation information by perturbing the reference image detection interpretation information through the text perturbation module; extracting image detection interpretation features corresponding to the reference image detection interpretation information and sample perturbation detection interpretation features corresponding to the sample perturbation detection interpretation information through the feature extraction module; and comparing and evaluating the image detection interpretation features through the result comparison module to obtain detection interpretation information evaluation results.
11. The method of claim 10, wherein the reference image detection interpretation information comprises first reference image detection interpretation information and second reference image detection interpretation information, and the image detection interpretation features comprise first image detection interpretation features corresponding to the first reference image detection interpretation information and second image detection interpretation features corresponding to the second reference image detection interpretation information. The method further comprises: comparing and evaluating the first image detection interpretation features and the second image detection interpretation features through the result comparison module to obtain detection interpretation information evaluation results, wherein the detection interpretation information evaluation results are one of the first reference image detection interpretation information and the second reference image detection interpretation information.
12. The method of claim 9, wherein the model parameter adjustment of the initial interpretation evaluation model based on the image detection interpretation features, the sample perturbation detection interpretation features, the detection interpretation information evaluation labels, and the detection interpretation information evaluation results comprises: determining a perturbation consistency loss based on the image detection interpretation features and the sample perturbation detection interpretation features; determining a result comparison loss based on the detection interpretation information evaluation labels and the detection interpretation information evaluation results; and adjusting model parameters of the initial interpretation evaluation model based on the perturbation consistency loss and the result comparison loss.
13. The method of claim 1, wherein the method comprises: creating an initial large-scale language tuning model based on an existing large-scale language model; obtaining sample image description text features generated by a cross-modal liveness detection model during a model training stage, and detection interpretation information evaluation labels annotated for reference image detection interpretation information, wherein the reference image detection interpretation information is reference image detection interpretation information generated by an interpretation evaluation model during the model training stage. inputting the sample image description text features and the detection explanation information evaluation label into an initial large-scale language tuning model for at least one round of model training to obtain sample image detection explanation information and a reference tuning network layer; performing model parameter adjustment on the reference tuning network layer of the initial large-scale language tuning model until the initial large-scale language tuning model completes model training to obtain a large-scale language tuning model.
14. The method of claim 13, wherein the initial large-scale language tuning model comprises a large-scale language model layer, an explanation evaluation model layer, and a network layer screening layer. The inputting the sample image description text features and the detection explanation information evaluation label into an initial large-scale language tuning model for at least one round of model training to obtain sample image detection explanation information and a reference tuning network layer comprises: performing detection explanation processing on the sample image description text features through the large-scale language model layer to obtain first sample image detection explanation information; performing explanation evaluation processing on the first sample image detection explanation information and the detection explanation information evaluation label through the explanation evaluation model layer to obtain sample image detection explanation information; determining a reference tuning network layer from the large-scale language model layer based on the sample image detection explanation information and the sample image detection explanation information through the network layer screening layer.
15. The method of claim 14, wherein the performing model parameter adjustment on the reference tuning network layer of the initial large-scale language tuning model comprises: determining an evaluation quality loss based on the sample image detection explanation information and the detection explanation information evaluation label; determining a sparse selection constraint loss based on the reference tuning network layer; performing model parameter adjustment on the reference tuning network layer of the initial large-scale language tuning model based on the evaluation quality loss and the sparse selection constraint loss.
16. A living body detection apparatus, comprising: an image acquisition module configured to acquire a target face image corresponding to a target object; a living body detection module configured to perform living body detection processing based on the target face image by using a cross-modal living body detection model comprising a living body feature processing network, a cross-modal pre-training model association network, a large-scale language model association network, and a cross-modal feature conversion network to obtain a living body detection result, wherein the image living body feature is obtained by using the living body feature processing network; performing text association conversion based on the image living body feature to obtain first image description text features based on the cross-modal pre-training model association network and second image description text features based on the large-scale language model association network, wherein the first image description text features are image description text features represented from a cross-modal image text association characteristic dimension by using the cross-modal pre-training model association network, and the second image description text features are image description text features represented from a large-scale language model text generation characteristic dimension by using the large-scale language model association network. adopt a large-scale language tuning model based on the first image description text feature and the second image description text feature to obtain first detection explanation information and second detection explanation information respectively, perform explanation evaluation processing based on the first detection explanation information and the second detection explanation information to obtain detection explanation information, the large-scale language tuning model being a large-scale language tuning model obtained by adapting a large-scale language model to a live body detection scene; an information prompting module configured to output the detection explanation information to the target object to prompt the target object to adjust a live body detection condition through the detection explanation information.
17. A computer storage medium storing a plurality of instructions, the instructions being adapted to be loaded and executed by a processor to perform the method of any one of claims 1-15.
18. A computer program product storing at least one instruction, the at least one instruction being loaded and executed by a processor to perform the method of any one of claims 1-15.
19. An electronic device comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the method of any one of claims 1-15. a processor and a memory; wherein the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the method of any one of claims 1-15.
Citation Information
Patent Citations
Face recognition-based doctor seeing method, apparatus and device, and storage medium
CN113111846A
Intelligent question and answer method, device and equipment and storage medium
CN114416927A
Image description generation method and device, storage medium and electronic equipment
CN115810068A
Method and device for determining picture description information generation model, medium and equipment
CN116050496A
Fine adjustment method of vehicle classification model, vehicle classification method, device and equipment
CN116091824A