Binocular human face in-vivo detection method based on three modes, medium and equipment

By adopting a three-modal (RGB, infrared, and depth) face detection method in facial vivo detection, using the mutual information module and ViT model, the problems of insufficient generalization ability and poor recognition performance in the existing technology are solved, and higher classification accuracy and generalization performance are achieved.

CN120108048AActive Publication Date: 2025-06-06SOUTH CHINA UNIV OF TECH
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510115354.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-06
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The existing multimodal face live detection technology has insufficient generalization capabilities and poor recognition performance when dealing with unknown attack modes, multi-type camera sensors and complex environment changes, especially under interference from 3D head-mode faces and complex light.

Method used

The binocular face live detection method based on three modes is adopted to adaptively enhance the favorable mode by maximizing the mutual information between different modes, while suppressing unfavorable modes, and use the ViT model and the mutual information module to mine more abundant and effective classification features.

Benefits of technology

It improves the classification accuracy and generalization performance of face fraud detection methods, improves the system's adaptability and customer operation experience, and significantly improves the recognition performance in cross-domain testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108048A_ABST
    Figure CN120108048A_ABST
Patent Text Reader

Abstract

The invention provides a binocular human face living body detection method based on three modes, a medium and equipment. The method comprises the following steps: acquiring an RGB image and an infrared image of a binocular camera, and generating a depth map; inputting the three modal images into a human face living body detection model; the face living body detection model comprises a ViT model and a mutual information module; after the three modal images are respectively processed by a layer normalization module I and a multi-head self-attention module, two modal features form a modal pair, and the modal pair is input into a mutual information module; carrying out splicing, convolution and activation on the input features of the modal pairs, carrying out re-weighting, and then carrying out splicing and convolution on the input features and interaction features; and performing integration to obtain model output so as to obtain a human face living body detection result. According to the method, mutual information between different modals is maximized to adaptively enhance favorable modals and suppress unfavorable modals, more effective classification features are mined, and the classification accuracy and generalization performance of the face fraud detection method are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically to a binocular face liveness detection method, medium and device based on tri-modality. Background Art

[0002] Face liveness detection is a very important part of face recognition tasks. It can ensure the reliability of face recognition in the financial industry or security scenarios, such as face-swiping deposits and withdrawals at self-service machines, face payment, etc. Due to the rapid increase in the types of attacks on facial recognition, researchers believe that single-modal face liveness detection has a weak ability to resist attacks. Fortunately, multimodal face images can provide additional and supplementary information, which can greatly improve the robustness of face liveness detection. Specifically, print-based 2D attacks are very easy to distinguish in the depth mode, but difficult to distinguish in the visible light mode. Under complex light interference, RBG images are not clear, overexposed or too dark, while infrared images are not affected by complex light.

[0003] However, the current multimodal face liveness detection has the following limitations, which prevents it from being widely used in certain application scenarios: (1) Visible light and infrared cameras have simple hardware designs and are easy to obtain, but depth maps are difficult to obtain due to complex optical path designs, interference from ambient light, and high costs. (2) The camera with the depth map is too large to be deployed on a mobile terminal; (III) Due to the complexity of application scenarios, the current multimodal liveness detection algorithm has poor recognition performance. There are two main reasons for this: (1) Previous invention patents believe that using multimodality can reveal more intrinsic deception traces, and most of them simply splice the features of the three modalities; however, the reliability of each modality can fluctuate depending on the type of attack, and it is not advisable to strictly or uniformly treat each modality; therefore, the model must adaptively prioritize specific modalities or image feature areas based on their reliability, and a simple strategy of fusing features of each modality may not meet the requirements; for example, some modalities (e.g., depth) may be non-defensive to specific attacks (e.g., 3D masks); (2) Customers are more casual when using the system. Their faces may be too far or too close, or not necessarily in the center of the camera screen. This results in poor facial photos being input into the model and a decrease in recognition rate.

[0004] (IV) Some application terminals explicitly require customers to cooperate with the camera and ensure that their faces are directly in front of the camera, resulting in a poor customer experience.

[0005] At present, there are the following documents on using binocular stereo vision technology for face liveness detection: In 2015, Beijing Haixin Kejin Co., Ltd. filed an invention patent entitled "A method, device and system for face detection based on binocular stereo vision" (publication number: CN104834901A). This method generates a depth map through binocular imaging, and then determines whether the face image constitutes a three-dimensional face structure map according to a preset three-dimensional structure classification rule to determine whether the face is alive or not. The misjudgment rate is 100% for 3D head model faces, and the effect is extremely poor. In 2015, the Institute of Semiconductors of the Chinese Academy of Sciences filed an invention patent entitled "A method and system for face liveness detection" (publication number: CN105023010A). This method is basically the same as Beijing Haixin Kejin Co., Ltd.'s invention patent CN104834901A. Both methods generate a depth map through binocular imaging, and then determine whether the face is alive based on the depth information. For 3D head model faces, the misjudgment rate is high and the effect is poor. The invention patent "A method and device for detecting staff fatigue" (publication number: CN115880757A) proposed by Jiangsu Jicui Intelligent Optoelectronics Co., Ltd. in 2022 uses binoculars to achieve liveness detection, obtains the width of the face through binocular stereo vision, does not generate a depth map, and the recognition rate of binocular liveness detection is limited.

[0006] The following documents use multimodal technology to detect human face liveness: The invention patent "A multimodal face liveness detection method and system" (publication number: CN112487922A) proposed by Orbbec in 2021, this method performs face detection on three modalities (color image, infrared image and depth image) respectively, and sends them into three neural networks for face feature extraction, and then merges the features to determine whether it is alive. The invention patent "A single-modal face liveness detection method based on multimodal face training" (publication number: CN113705400A) proposed by Sun Yat-sen University at the end of 2021. The idea of ​​this method is to generate infrared and depth maps from visible light images, and then use the features of the three modalities to splice them together; because the gap between the image generated by the generative adversarial network and the image actually collected by the camera is still very large, the actual test performance is very poor and the generator is unstable. Hangzhou Qiyuan Vision Co., Ltd. proposed an invention patent in 2023, "A multimodal fusion face liveness detection model generation method and device, electronic equipment" (publication number: CN116311451A). The multimodality of this method only includes two modes, namely color image and infrared image, and the method is implemented by cross-comparison, that is, visible light generates infrared light image, infrared light image generates visible light image, and the authenticity is identified by generating images and original images. The biggest defect of this method is the instability of the generated network. Zhejiang University of Technology proposed an invention patent in 2023, "A multimodal face anti-fraud detection method based on AR-MLP" (publication number: CN117011911A). This method processes the input face RGB image, depth image and infrared image to realize the multimodal feature fusion of the face. The fusion method is the feature fusion of the same face area patch of the three modalities. This method is better than CN112487922A and CN116311451A, but actual tests have found that the recognition performance is limited. The invention patent "A multimodal face liveness detection method based on attention mechanism" (publication number: CN117894082A) proposed by Beijing University of Technology in 2024, this method performs image preprocessing, feature extraction, and multi-level feature fusion on three modalities (color images, infrared images, and depth images) respectively; judging from the experimental results given in the patent, the BPCER obtained in a single database is 1.82%, which is a good effect, but it is still out of practical application, because practical applications must encounter multiple domains, that is, cross-domain testing, and the actual cross-domain performance test is still not good. Summary of the invention

[0007] In order to overcome the shortcomings and deficiencies in the prior art, the purpose of the present invention is to provide a binocular face liveness detection method, medium and device based on three modalities; the method proposes a method of maximizing the mutual information between different modalities to adaptively enhance the favorable modality while suppressing the unfavorable modality, thereby mining more abundant and effective classification features, and improving the classification accuracy and generalization performance of the existing face fraud detection method.

[0008] In order to achieve the above object, the present invention is implemented by the following technical solution: a binocular face liveness detection method based on tri-modality, comprising the following steps: Step S1, obtaining the RGB image and infrared image of the binocular camera, and generating a depth map; Step S2: RGB image, infrared image and depth Figure 3 The modality images are input into the face liveness detection model; The face liveness detection model includes three ViT models and two mutual information modules; the three ViT models process three modal images respectively; the three ViT models all include a layer normalization module 1 and a multi-head self-attention module; the two ViT models processing RGB images and infrared images also include a layer normalization module 2 and a multi-layer perceptron; The RGB image, infrared image and depth map are processed by the layer normalization module 1 and the multi-head self-attention module respectively to obtain the feature , , ; The feature and Features Form a mode pair, feature and Features Form another modality pair and input them into two mutual information modules respectively; The mutual information module combines the modality into two modalities m1 and m2 Input features , Perform concatenation and convolution to obtain interactive features f fused ; By the interaction features f fused Perform Modal m1 and m2 Activation, get two information masks m m1 , m m2 ; Use information masking m m1 , m m2 For the input features , Re-weight the features z m1_new , z m2_new , and then respectively with the interaction features f fused Perform concatenation and convolution to obtain the output features of the mutual information module z out_m1 ,z out_m2 ; Features corresponding to RGB images Features corresponding to infrared images After being processed by layer normalization module 2 and multi-layer perceptron, the output features of the corresponding modes of the mutual information module are superimposed respectively, and then integrated to obtain the output of the face liveness detection model; based on the output of the face liveness detection model, the face liveness detection result is obtained.

[0009] Preferably, the interactive feature f fused for: ; Among them, Conv(·) represents convolution calculation; Cat represents feature connection; Two information masks m m1 , m m2 for: ; Among them, sigmoid(·) is the activation function; The features z m1_new , z m2_new They are: ; ; Output features of the mutual information module z out_m1 , z out_m2 They are: ; .

[0010] Preferably, the RGB image, infrared image and depth map are processed by a layer normalization module and a multi-head self-attention module respectively, which means that for each modality image of the RGB image, infrared image and depth map, each modality image is first , where H, W, and C represent the height, width, and number of channels of the image, respectively. The image is evenly divided into b blocks, and the width and height of each block are s. ; and the image blocks are represented as a sequence through a layer of two-dimensional convolution ; then in the sequence E I Add a category tag to the header , and after adding the position encoding, it is input into the multi-head self-attention module for processing.

[0011] Preferably, in the ViT model for processing the depth map, the layer normalization module 1 and the multi-head self-attention module are connected in sequence; the multi-head self-attention module is connected to the mutual information module; In the two ViT models for processing RGB images and infrared images, the layer normalization module 1 and the multi-head self-attention module are connected in sequence; the multi-head self-attention module is connected to the mutual information module; the output of the multi-head self-attention module is added to the layer normalization module 1; then it is connected to the layer normalization module 2 and the multi-layer perceptron in sequence, and added to the output of the multi-layer perceptron and the output of the mutual information module to obtain the output of the ViT model; the outputs of the two ViT models are spliced ​​and feature fused, and then connected to the Softmax layer to obtain the output of the face liveness detection model.

[0012] Preferably, the binocular camera has an adjustment module for adjusting the angle; In the step S1, during the process of acquiring the RGB image and the infrared image of the binocular camera, the adjustment module adjusts the angle of the binocular camera so that the human face is in the center of the RGB image and the infrared image.

[0013] Preferably, in step S1, the method for generating the depth map is, first, calculating the distance from each point on the face to the two camera sensor planes of the binocular camera according to the RGB image and the infrared image acquired by the binocular camera. z : ; ; in, f represents the focal length, which is the distance from the lens to the sensor; b Represents the distance between two camera sensors, that is, the physical distance between two camera sensor chips; u L and u R Represents the horizontal coordinates of the pixel points on the face imaged by the left camera and the right camera respectively; d stands for parallax; According to distance z , generate a depth map.

[0014] A readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the three-modality based binocular face liveness detection method.

[0015] A computer device comprises a processor and a memory for storing a program executable by the processor. When the processor executes the program stored in the memory, the tri-modal based binocular face liveness detection method is implemented.

[0016] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. In view of the problem that the current face liveness detection technology has insufficient generalization ability when dealing with unknown attack modes, multiple types of camera sensors and complex environmental changes, the method of the present invention proposes a method of maximizing the mutual information between different modalities to adaptively enhance the favorable modality while suppressing the unfavorable modality, digging out more abundant and effective classification features, and improving the classification accuracy and generalization performance of the existing face fraud detection method; 2. The integrated face fraud module will be superimposed with the target intelligent tracking technology that can measure distance and is equipped with a binocular camera adjustment module to achieve system adaptive face recognition, which will not only improve the overall security and reliability of my country's bank terminal equipment, but also greatly enhance the actual operation experience of customers. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flow chart of the binocular face liveness detection method based on three modalities of the present invention; FIG2(a) and FIG2(b) are respectively the geometrical schematic diagram and the calculation principle diagram of the triangle formed by the image planes of the two image sensors and the measured object; Figure 3 It is a structural schematic diagram of the face liveness detection model of the present invention; Figure 4 It is a schematic diagram of the internal implementation of the mutual information module of the present invention. DETAILED DESCRIPTION

[0018] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0019] Embodiment 1 This embodiment is a binocular face liveness detection method based on three modalities, such as Figure 1 As shown, the following steps are included: Step S1, the binocular camera hardware is two independent monocular cameras, the wavelength of the cameras are both 380nm~1200nm, and the parameters are set to collect visible light and infrared light, that is, collect RGB images and infrared images respectively. The binocular camera has an adjustment module for adjusting the angle; the adjustment module can use a micro motor, etc.

[0020] The hardware structure of the face video data collected by the binocular camera ensures that the face is within the preset common field of view of the binocular camera. Face detection is performed based on the RGB image to obtain the distance between the current face and the camera and the angle of the upper and lower positions; the adjustment module adjusts the angle of the binocular camera so that the face is in the center of the RGB image and the infrared image.

[0021] According to the face video data of the binocular camera, the RGB image and infrared image are obtained, and the depth map is generated. Based on the parallax of the left eye camera and the right eye camera of the binocular camera, the three-dimensional information is obtained by the triangulation principle, that is, a triangle is formed between the image planes of the two image sensors and the object being measured, as shown in Figure 2 (a). Knowing the positional relationship between the two image sensors, the three-dimensional coordinates of the object in the common field of view of the two image sensors can be obtained. The calculation principle is shown in Figure 2 (b); the geometric relationship formula is: ; in, z Represents the distance between each point on the subject's face and the two camera sensor planes of the binocular camera; f represents focal length, which is the distance from the lens to the sensor; b Represents the distance between two camera sensors, that is, the physical distance between two camera sensor chips; u L and u R Represents the horizontal coordinates of the pixel points on the face imaged by the left camera and the right camera respectively; According to mathematical operations, z for: ; ; in, d Represents parallax; based on distance z , generate a depth map.

[0022] Step S2: RGB image, infrared image and depth Figure 3 The images of different modalities are input into the face liveness detection model.

[0023] Face liveness detection model, such as Figure 3 As shown in the figure, the face liveness detection model is built on ViT (Visual Transformer) and fine-tuned using frozen pre-trained weights. Specifically, the face liveness detection model includes three ViT models and two mutual information modules (MI); the three ViT models process three modal images respectively; the three ViT models all include layer normalization module 1 (Norm) and multi-head self-attention module (MHA); the two ViT models that process RGB images and infrared images also include layer normalization module 2 (Norm) and multi-layer perceptron (MLP).

[0024] The face liveness detection model of the present invention adopts a mutual information module, which is obtained by the inventors of the present invention through the following research: Assume that random variables A and B represent the possible values ​​of RGB image feature representation, infrared image feature representation and depth map feature representation respectively; the mutual information value is calculated by the following formula: ; in, is the mutual information value, a and b represent the possible values ​​of RGB image features, infrared image features and depth map features respectively. is the joint probability distribution of two different modality images, , is a modal edge probability distribution of RGB image features, infrared image features and depth map features; set the threshold and judge the degree of dependence according to the mutual information value: when When , it is determined that the dependence between modes is strong; when When , the dependence between the modalities is judged to be weak; the complementarity between the modalities is determined by analyzing the unique information provided by each modality and its potential contribution to the fusion task.

[0025] Specifically: 1) Depth map features can effectively defend against various material attacks on two-dimensional planes, but are not defensive against specific attacks (for example, 3D masks). When the attack target is defined as a 3D masked 3D face portrait, the depth map and RGB image (or depth map and infrared image) The value is very small, that is, the mutual information of the depth map and the RGB image or infrared image is independent of each other; when the attack object is a two-dimensional paper image or electronic tablet image, The value is very large, that is, the mutual information between the depth map and the RGB image or infrared image is very large, and the depth map can clearly show that it is a fake attack; 2) The characteristic of infrared images is that they are independent of the ambient light intensity. When the ambient light is overexposed or too dark (such as at night), the infrared image imaging is not affected; while the RGB image is just the opposite, which is related to the ambient light intensity. When the ambient light is overexposed or too dark, the infrared image and RGB image are The value is very large, that is, the mutual information between the infrared image and the RGB image is large. The current lighting environment can be known through the RGB image and the infrared image. The current model is suitable for using the features of the infrared image instead of the RGB image features. When the ambient light is of normal brightness, The value is very small, that is, the mutual information between the infrared image and the RGB image is large and independent of each other. The current model needs to use the features of both to jointly judge the authenticity of the living body.

[0026] They typically maximize the mutual information between features extracted from different views, modalities, or images, which come from data augmentation and aim to capture high-level factors that affect spanning different viewpoints - for example, the presence of certain traces of deception from different viewpoints or certain inconsistencies in the data. This capability is particularly valuable in multimodal FAS, where each modality has unique strengths or weaknesses in combating specific types of attacks. By maximizing the mutual information between modalities, the model can adaptively emphasize task-relevant information, thereby enhancing reliable modalities while mitigating the impact of unreliable modalities.

[0027] In general, the reliability of each modality can fluctuate depending on the type of attack, and it is not advisable to treat each modality strictly or uniformly. Therefore, the model must adaptively prioritize specific modes or regions based on their reliability. To achieve this, the present invention proposes a mutual information module (MI) based on the ViT model. The mutual information module dynamically emphasizes reliable modes and suppresses unreliable modes by maximizing mutual information, thereby improving the classification accuracy and generalization performance of existing face fraud detection methods.

[0028] In the ViT model for processing depth maps, the layer normalization module 1 and the multi-head self-attention module are connected in sequence; the multi-head self-attention module is connected to the mutual information module; In the two ViT models for processing RGB images and infrared images, the layer normalization module 1 and the multi-head self-attention module are connected in sequence; the multi-head self-attention module is connected to the mutual information module; the output of the multi-head self-attention module is added to the layer normalization module 1; then it is connected to the layer normalization module 2 and the multi-layer perceptron in sequence, and added to the output of the multi-layer perceptron and the output of the mutual information module to obtain the output of the ViT model; the outputs of the two ViT models are spliced ​​and feature fused, and then connected to the Softmax layer to obtain the output of the face liveness detection model.

[0029] The RGB image, infrared image and depth map are processed by the layer normalization module 1 and the multi-head self-attention module respectively to obtain the feature , , .

[0030] Specifically, the RGB image, infrared image and depth map are processed by the layer normalization module and the multi-head self-attention module respectively, which means that for each modality image of the RGB image, infrared image and depth map, each modality image is first , where H, W, and C represent the height, width, and number of channels of the image, respectively. The image is evenly divided into b blocks, and the width and height of each block are s. ; and the image blocks are represented as a sequence through a layer of two-dimensional convolution ; then in the sequence EI Add a category tag to the header , and after adding the position encoding, it is input into the multi-head self-attention module for processing.

[0031] The characteristics and Features Form a mode pair, feature and Features Another modality pair is formed and respectively input into two mutual information modules; the present invention feeds the output of each multi-head self-attention module into the mutual information module to adjust the weight so that the reliability of each modality can fluctuate according to the attack type.

[0032] In each mutual information module, the two modes of the mode pair are m1 and m2 If m1 In RGB mode, m2 is infrared mode ir or depth mode d; Pair the modal with the two modal m1 and m2 Input features , Concatenate and convolve to obtain interactive features by directly connecting features along the channel dimension and feeding them into a lightweight interactive convolution block f fused : ; Among them, Conv(·) represents convolution calculation; Cat represents feature connection; By using the interactive features f fused Perform Modal m1 and m2 Activation, get two information masks m m1 , m m2 : ; Among them, sigmoid(·) is the activation function.

[0033] Two information masks m m1 , m m2 It reflects the information with the most semantic information as the mutual information (MI) rich area. We believe that the higher the weight of the area (information point), the more reliable the information it involves. On the contrary, the area associated with a lower weight may carry redundant or negative information.

[0034] Then, using information mask mm1 , m m2 For the input features , Re-weight the features z m1_new , z m2_new : ; ; Then, the interaction features f fused Perform concatenation and convolution to obtain the output features of the mutual information module z out_m1 , z out_m2 : ; ; Figure 4 This is a schematic diagram of the internal implementation of the mutual information module (MI), taking the interaction between RGB images and infrared images as an example.

[0035] Features corresponding to RGB images Features corresponding to infrared images After being processed by layer normalization module 2 and multi-layer perceptron, the output features of the corresponding modes of the mutual information module are superimposed respectively, and then integrated to obtain the output of the face liveness detection model; based on the output of the face liveness detection model, the face liveness detection result is obtained.

[0036] This patent of the present invention achieved the best HTER and accuracy on four public datasets: CASIA-CeFA, PADISI-Face, CASIA-SURF and WMCA.

[0037] The CASIA-SURF dataset has a total of 1,000 volunteers participating in the recording, totaling 21,000 multimodal videos. The data comes from multiple channels (visible light, depth map, and near infrared), using flat printing attacks or curled printing attacks, and randomly removing areas such as eyes, noses, and mouths. The upgraded version of this dataset, CASIA-SURF CeFA, was released in 2020, adding cross-racial volunteer identities to test the generalization of the algorithm across races. The Multi-Channel Presentation Attack (WMCA) dataset contains 1,941 short video records of real people and presentation disguise attacks from 72 different identities. The data is recorded from several channels, including color, depth, infrared, and thermal imaging. The PADISI-Face dataset has a total of 360 volunteers participating in the recording, totaling 1,105 real-person videos, 924 presentation attack videos, and up to 37 types of attacks.

[0038] In order to verify the beneficial effects of the method of the present invention, the experimental comparison results of the method of the present invention and the existing method on four public data sets CASIA-CeFA (abbreviated as C), PADISI-Face (abbreviated as P), CASIA-SURF (abbreviated as S) and WMCA (abbreviated as W) are shown in Table 1; wherein, CPS->W represents the training set is C, P, S, and the test set is W; CPW->S represents the training set is C, P, W, and the test set is S; CSW->P represents the training set is C, S, W, and the test set is P; PSW->C represents the training set is P, S, W, and the test set is C.

[0039] Table 1 Experimental results of this paper (unit: %)

[0040] As can be seen from Table 1, on the four data sets, the present invention achieves adaptive enhancement of favorable modes and suppression of unfavorable modes based on the method of maximizing the mutual information between different modes, and has achieved quite competitive results. The average half error rate (HTER, the lower the better) and accuracy (AUC, the higher the better) have achieved the best. ViTAF, ViT+AMA and MMDG in Table 1 are three algorithms based on the ViT model. The method of the present invention is significantly better than them in terms of cross-domain testing effects. The above experimental results show that the method of the present invention can better process and classify various attack methods that the model has not seen.

[0041] Embodiment 2 This embodiment provides a readable storage medium, wherein the readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the binocular face liveness detection method based on three modalities described in Embodiment 1.

[0042] Embodiment 3 The present embodiment provides a computer device, including a processor and a memory for storing a program executable by the processor. When the processor executes the program stored in the memory, the three-modality-based binocular face liveness detection method described in the first embodiment is implemented.

[0043] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A binocular face liveness detection method based on three modalities, characterized by: The steps include: Step S1, obtaining the RGB image and infrared image of the binocular camera, and generating a depth map; Step S2: inputting the three modal images of RGB image, infrared image and depth map into the face liveness detection model; The face liveness detection model includes three ViT models and two mutual information modules; the three ViT models process three modal images respectively; the three ViT models all include a layer normalization module 1 and a multi-head self-attention module; the two ViT models processing RGB images and infrared images also include a layer normalization module 2 and a multi-layer perceptron; The RGB image, infrared image and depth map are processed by the layer normalization module 1 and the multi-head self-attention module respectively to obtain the feature , , ; The feature and Features Form a mode pair, feature and Features Form another modality pair and input them into two mutual information modules respectively; The mutual information module combines the modality into two modalities m1 and m2 Input features , Perform concatenation and convolution to obtain interactive features f fused ; By the interaction features f fused Perform Modal m1 and m2 Activation, get two information masks m m1 , m m2 ; Use information masking m m1 , m m2 For the input features , Re-weight the features z m1_new , z m2_new , and then respectively with the interaction features f fused Perform concatenation and convolution to obtain the output features of the mutual information module z out_m1 , z out_m2 ; Features corresponding to RGB images Features corresponding to infrared images After being processed by layer normalization module 2 and multi-layer perceptron, the output features of the corresponding modes of the mutual information module are superimposed respectively, and then integrated to obtain the output of the face liveness detection model; based on the output of the face liveness detection model, the face liveness detection result is obtained.

2. The method for binocular face liveness detection based on tri-modality according to claim 1, characterized in that: The interactive features f fused for: ; Among them, Conv(·) represents convolution calculation; Cat represents feature connection; Two information masks m m1 , m m2 for: ; Among them, sigmoid(·) is the activation function; The features z m1_new , z m2_new They are: ; ; Output features of the mutual information module z out_m1 , z out_m2 They are: ; 。 3. The method for binocular face liveness detection based on tri-modality according to claim 1, characterized in that: The RGB image, infrared image and depth map are processed by the layer normalization module and the multi-head self-attention module respectively, which means that for each modality image of the RGB image, infrared image and depth map, each modality image is first , where H, W, and C represent the height, width, and number of channels of the image, respectively. The image is evenly divided into b blocks, and the width and height of each block are s. ; and the image blocks are represented as a sequence through a layer of two-dimensional convolution ; then in the sequence E I Add a category tag to the header , and after adding the position encoding, it is input into the multi-head self-attention module for processing.

4. The method for binocular face liveness detection based on tri-modality according to claim 1, characterized in that: In the ViT model for processing depth maps, the layer normalization module 1 and the multi-head self-attention module are connected in sequence; the multi-head self-attention module is connected to the mutual information module; In the two ViT models for processing RGB images and infrared images, the layer normalization module 1 and the multi-head self-attention module are connected in sequence; the multi-head self-attention module is connected to the mutual information module; the output of the multi-head self-attention module is added to the layer normalization module 1; then it is connected to the layer normalization module 2 and the multi-layer perceptron in sequence, and added to the output of the multi-layer perceptron and the output of the mutual information module to obtain the output of the ViT model; the outputs of the two ViT models are spliced ​​and feature fused, and then connected to the Softmax layer to obtain the output of the face liveness detection model.

5. The method for binocular face liveness detection based on tri-modality according to claim 1, characterized in that: The binocular camera has an adjustment module for adjusting the angle; In the step S1, during the process of acquiring the RGB image and the infrared image of the binocular camera, the adjustment module adjusts the angle of the binocular camera so that the human face is in the center of the RGB image and the infrared image.

6. The tri-modal binocular face liveness detection method according to claim 1, characterized in that: In step S1, the method for generating the depth map is, first, calculating the distance from each point on the face to the two camera sensor planes of the binocular camera based on the RGB image and infrared image acquired by the binocular camera. z : ; ; in, f represents the focal length; b Represents the distance between two camera sensors; u L and u R Represents the horizontal coordinates of the pixel points on the face imaged by the left camera and the right camera respectively; d stands for parallax; According to distance z , generate a depth map.

7. A readable storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, enables the processor to execute the tri-modal binocular face liveness detection method according to any one of claims 1 to 6.

8. A computer device comprising a processor and a memory for storing a program executable by the processor, characterized in that: When the processor executes the program stored in the memory, the binocular face liveness detection method based on three modalities described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Binocular stereo vision-based human face detection method, device and system

    CN104834901A

  • Face living body detection method and system

    CN105023010A

  • Multi-mode human face living body detection method and system

    CN112487922A

  • Single-mode face living body detection method based on multi-mode face training

    CN113705400A

  • Staff fatigue detection method and device

    CN115880757A