Liveness detection method, device, medium and equipment based on multi-wavelength feature fusion

Through the multi-wavelength feature fusion method, combining the semantic and texture features of visible light, infrared and vein facial images, the problem of insufficient accuracy of face recognition systems in liveness detection is solved, defense against diverse attacks is achieved, and the accuracy of liveness detection is improved.

CN116778588BActive Publication Date: 2025-09-05RECONOVA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310798606.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2025-09-05
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

Existing face recognition systems are unable to cope with diverse attack methods in liveness detection, resulting in low accuracy in liveness judgment.

Method used

A multi-wavelength feature fusion method is adopted to obtain visible light, infrared and vein face images, and use the pre-trained semantic feature extraction model and texture feature extraction model to fuse the visible light face liveness semantic features, infrared face liveness semantic features, vein face liveness semantic features and multi-wavelength face texture features for liveness recognition.

Benefits of technology

It improves the accuracy of liveness detection, can effectively defend against electronic screen, paper printing and 3D face attacks, and enhances the robustness of liveness detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116778588B_ABST
    Figure CN116778588B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a method, device, medium and equipment for liveness detection based on multi-wavelength feature fusion. The method includes: obtaining a visible light face image, an infrared face image and a vein face image corresponding to the target face, wherein the vein face image contains vein information of the target face; extracting corresponding visible light face liveness semantic features, infrared face liveness semantic features, vein face liveness semantic features and multi-wavelength face texture features based on the visible light face image, infrared face image and vein face image; fusing the visible light face liveness semantic features, infrared face liveness semantic features, vein face liveness semantic features and multi-wavelength face texture features to obtain face liveness fusion features; performing recognition based on the face liveness fusion features to determine the liveness recognition result. The technical solution of the embodiments of the present application can adapt to the diversity of attack methods, reduce the occurrence of liveness judgment errors, and improve the accuracy of liveness detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of liveness detection, and in particular to a liveness detection method, apparatus, medium, and equipment based on multi-wavelength feature fusion. Background Art

[0002] In the technical field of identity recognition, facial recognition systems use the face as a unique biometric feature for identification. In practice, the face, as an open biometric feature, is easily accessible to third parties through photos or videos, which can then be used to launch attacks against the person. Therefore, facial recognition systems must not only "identify the person" but also "identify the person." This means that the system must not only verify that the face is the person's face, but also that it is a live person, not just a picture or video. Current technical solutions can detect liveness based on visible light images, infrared images, or depth images. However, liveness detection based on a single image cannot cope with the diversity of current attack methods, is prone to false liveness determinations, and has low accuracy. Summary of the Invention

[0003] The embodiments of the present application provide a liveness detection method, apparatus, medium, and device based on multi-wavelength feature fusion, which can adapt to the diversity of attack methods, at least to a certain extent, reduce the occurrence of liveness judgment errors, and improve the accuracy of liveness detection.

[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.

[0005] According to one aspect of an embodiment of the present application, a method for liveness detection based on multi-wavelength feature fusion is provided, comprising:

[0006] Acquire a visible light face image, an infrared face image, and a vein face image corresponding to a target face, wherein the vein face image includes vein information of the target face;

[0007] Based on the pre-trained visible light face liveness semantic feature extraction model, infrared face liveness semantic feature extraction model, and vein face liveness semantic feature extraction model, semantic feature extraction is performed on the visible light face image, the infrared face image, and the vein face image, respectively, to determine the corresponding visible light face liveness semantic features, infrared face liveness semantic features, and vein face liveness semantic features, wherein the visible light face liveness semantic feature extraction model, the infrared face liveness semantic feature extraction model, and the vein face liveness semantic feature extraction model are respectively trained using real face data and attack data of other possible methods;

[0008] Performing texture feature extraction based on the visible light face image, the infrared face image, and the vein face image to determine corresponding multi-wavelength face texture features, wherein the multi-wavelength face texture features include visible light face texture features, infrared face texture features, and vein face texture features;

[0009] Fusing the visible light face liveness semantic feature, the infrared face liveness semantic feature, the vein face liveness semantic feature, and the multi-wavelength face texture feature to obtain a face liveness fusion feature;

[0010] Identification is performed based on the live face fusion features to determine a liveness recognition result.

[0011] According to one aspect of an embodiment of the present application, a living body detection device based on multi-wavelength feature fusion is provided, comprising:

[0012] An acquisition module is used to acquire a visible light face image, an infrared face image, and a vein face image corresponding to a target face, wherein the vein face image includes vein information of the target face;

[0013] a semantic feature extraction module for performing semantic feature extraction on the visible light face image, the infrared face image, and the vein face image, respectively, based on pre-trained visible light face liveness semantic feature extraction models, infrared face liveness semantic feature extraction models, and vein face liveness semantic feature extraction models, to determine corresponding visible light face liveness semantic features, infrared face liveness semantic features, and vein face liveness semantic features, wherein the visible light face liveness semantic feature extraction model, the infrared face liveness semantic feature extraction model, and the vein face liveness semantic feature extraction model are respectively trained using real face data and other possible attack data;

[0014] a texture feature extraction module, configured to extract texture features based on the visible light face image, the infrared face image, and the vein face image, and determine corresponding multi-wavelength face texture features, wherein the multi-wavelength face texture features include visible light face texture features, infrared face texture features, and vein face texture features;

[0015] a feature fusion module, configured to fuse the visible light face liveness semantic feature, the infrared face liveness semantic feature, the vein face liveness semantic feature, and the multi-wavelength face texture feature to obtain a face liveness fusion feature;

[0016] The processing module is used to perform recognition based on the face liveness fusion feature to determine the liveness recognition result.

[0017] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for liveness detection based on multi-wavelength feature fusion as described in the above embodiment is implemented.

[0018] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the liveness detection method based on multi-wavelength feature fusion as described in the above embodiments.

[0019] According to one aspect of an embodiment of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the liveness detection method based on multi-wavelength feature fusion provided in the above-described embodiments.

[0020] In the technical solutions provided in some embodiments of the present application, by obtaining visible light face images, infrared face images and vein face images corresponding to the target face, the visible light face liveness semantic features, infrared face liveness semantic features and vein face liveness semantic features of each image are extracted based on the pre-trained visible light face liveness semantic feature extraction model, infrared face liveness semantic feature model and vein face liveness semantic feature extraction model, respectively. Each model is trained using real face data and other possible attack data, thereby ensuring the targeted nature of feature extraction and thus ensuring the accuracy of subsequent liveness detection results. Next, texture features are extracted from the visible light face image, infrared face image, and vein face image to determine the corresponding multi-wavelength face texture features. These multi-wavelength face texture features include visible light face texture features, infrared face texture features, and vein face texture features. The visible light face liveness semantic features, infrared face liveness semantic features, vein face liveness semantic features, and multi-wavelength face texture features are then fused to obtain a face liveness fusion feature. This face liveness fusion feature is then used to determine the liveness recognition result. Thus, liveness detection is performed by extracting corresponding semantic and texture features from the three images and performing feature fusion. The features of different images can be complementary, thus adapting to the diversity of attack methods, reducing the occurrence of liveness misjudgments, and improving the accuracy of liveness detection.

[0021] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, explaining the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:

[0023] Figure 1 A schematic diagram of a flow chart of a method for liveness detection based on multi-wavelength feature fusion according to an embodiment of the present application is shown;

[0024] Figure 2 A block diagram of a living body detection device based on multi-wavelength feature fusion according to an embodiment of the present application is shown;

[0025] Figure 3 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0026] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.

[0027] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0028] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0029] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0030] Figure 1 A flow chart of a method for liveness detection based on multi-wavelength feature fusion according to an embodiment of the present application is shown. The method can be applied to a terminal device or a server, wherein the terminal device may include but is not limited to one or more of a smartphone, a tablet computer, a laptop computer, a desktop computer, an access control device, and a gate device. The server can be a physical server or a cloud server, and this application does not specifically limit this. The following description will be given using the method applied to a terminal device (hereinafter referred to as a "terminal") as an example.

[0031] Please refer to Figure 1 The liveness detection method based on multi-feature fusion includes at least steps S110 to S150, which are described in detail as follows:

[0032] In step S110 , a visible light face image, an infrared face image, and a vein face image corresponding to a target face are acquired, wherein the vein face image includes vein information of the target face.

[0033] In this embodiment, the terminal can obtain a visible light face image, infrared face image, and vein face image corresponding to the target face through pre-configured visible light face recognition module, infrared face module, and vein face recognition module, respectively. Each module can calibrate the camera's internal and external parameters, and correct the visible light face image, infrared face image, and vein face image using the calibrated internal and external parameters.

[0034] Taking a turnstile as an example, a turnstile can be equipped with a visible light facial recognition module, an infrared facial recognition module, and a vein facial recognition module. When a person approaches the detection location, the visible light facial recognition module, infrared facial recognition module, and vein facial recognition module respectively capture the person's corresponding visible light facial image, infrared facial image, and vein facial image, and transmit them to the processor on the access control device for subsequent processing.

[0035] In step S120, based on the pre-trained visible light face liveness semantic feature extraction model, infrared face liveness semantic feature extraction model and vein face liveness semantic feature extraction model, semantic feature extraction is performed on the visible light face image, the infrared face image and the vein face image respectively to determine the corresponding visible light face liveness semantic features, infrared face liveness semantic features and vein face liveness semantic features. The visible light face liveness semantic feature extraction model, the infrared face liveness semantic feature extraction model and the vein face liveness semantic feature extraction model are respectively trained using real face data and attack data of other possible types.

[0036] In this embodiment, those skilled in the art can pre-build and train a visible light face liveness semantic feature extraction model, an infrared face liveness semantic feature extraction model, and a vein face liveness semantic feature extraction model, and the terminal can call the corresponding model to perform semantic feature extraction on the corresponding image. That is, the visible light face liveness semantic feature extraction model is used to perform semantic feature extraction on the visible light face image to determine the corresponding visible light face liveness semantic features, the infrared face liveness semantic feature extraction model is used to perform feature extraction on the infrared face image to determine the corresponding infrared face liveness semantic features, and the vein face liveness semantic feature extraction model is used to perform feature extraction on the vein face image to determine the corresponding vein face liveness semantic features.

[0037] It's worth noting that different feature extraction models are trained using real face data and other possible attack data, meaning they use different training data. It's important to understand that liveness features vary across images. By training different feature extraction models, we can fully extract the most distinguishing liveness features of faces in each image, thereby improving the accuracy of subsequent liveness recognition results.

[0038] In one embodiment of the present application, the training data for the visible light face liveness semantic feature extraction model includes real visible light face images, conventional black and white paper face attack data, simulated black and white paper face attack data with vein patterns added, simulated 3D face attack data with vein patterns added, and corresponding label information. In one example, real face data is assigned label 1, while other presented attack data is assigned label 0. The training data for the infrared face liveness semantic feature extraction model includes real infrared face images, conventional color paper face attack data, simulated color paper face attack data with vein patterns added, and simulated 3D face attack data with vein patterns added, as well as corresponding label information. In one example, real face data is assigned label 1, while other presented attack data is assigned label 0. The training data for the vein face liveness semantic feature extraction model includes real vein face images and conventional 3D face attack data, as well as corresponding label information. The real vein face images contain vein information, while the conventional 3D face attack data does not. In one example, real face data is assigned label 1, while other presented attack data is assigned label 0. Therefore, by selecting the above training data, the accuracy of the extracted facial features can be improved, thereby ensuring the training effect of each feature extraction model.

[0039] In one embodiment of the present application, the visible light face live semantic feature extraction model, the infrared face live semantic feature extraction model, and the vein face live semantic feature extraction model all include connected convolutional layers and fully connected layers, and the fully connected layer is used to perform dimensionality conversion on the features output by the convolutional layer.

[0040] In this example, the classic ResNet network structure can be used as the backbone. In other examples, network structures such as Inception and MobileNet can also be used as the backbone. The basic network backbone serves as the feature extraction module (i.e., the convolutional layer). The feature extraction module is followed by a fully connected layer to convert the living body feature dimension to 512*1*1.

[0041] Based on the aforementioned embodiments, when training the visible light face liveness semantic feature extraction model, the infrared face liveness semantic feature extraction model, and the vein face liveness semantic feature extraction model, a classic cross-entropy loss function is used to calculate the loss between the classification output and the label. The error is then back-propagated via the chain rule to update the network parameters, leading to gradual model convergence. It should be understood that the most significant characteristic of vein face images is that real faces have vein patterns, while fake faces do not. This is the most obvious liveness characteristic of vein face images that distinguish live faces from those presented as attacks. The most significant characteristic of infrared face images is that real faces do not display vein information and have a bright pupil effect, while faces printed on color paper appear to have blurred texture and a whitish appearance. This is the most obvious liveness characteristic of live faces from those presented as attacks under infrared light. The most significant characteristic of visible light face images is that real faces do not display vein information and have color information, while faces printed on black and white paper lack color information. This is the most obvious liveness characteristic of live faces from those presented as attacks under visible light.

[0042] After the above training, the vein face liveness semantic feature extraction model can extract the obvious liveness features of vein faces and accurately distinguish whether the input facial data contains vein information; the visible light face liveness semantic feature extraction model can extract the obvious liveness features of visible light faces and accurately identify black and white paper printing attacks and simulated black and white faces and 3D face attacks with added vein patterns; the infrared face liveness semantic feature extraction model can extract the obvious liveness features of infrared faces and accurately identify conventional color paper printing attacks and simulated color faces and 3D face attacks with added vein patterns, thereby improving the accuracy of subsequent liveness recognition results.

[0043] In step S130, texture feature extraction is performed based on the visible light facial image, the infrared facial image, and the vein facial image to determine corresponding multi-wavelength facial texture features, where the multi-wavelength facial texture features include visible light facial texture features, infrared facial texture features, and vein facial texture features.

[0044] In this embodiment, multi-wavelength facial texture features can be extracted using traditional image processing methods. For visible light facial images, in order to highlight the skin color of the face, the RGB color space is first converted to the YCrCb color space, which is more sensitive to the color information of the face. The facial image in this area is then converted into a 512-dimensional LBP (Local Binary Pattern) feature (i.e., visible light facial texture feature). For infrared facial images, in order to highlight the bright pupil and detail information of the face, the infrared facial image is first subjected to a Laplacian transform to highlight the detail information of the image. The facial image in this area is then converted into a 512-dimensional LBP feature (i.e., infrared facial texture feature). For vein facial images, in order to highlight the texture information of the vein, the infrared facial image is first subjected to a Laplacian transform to highlight the vein information of the image. The facial image in this area is then converted into a 512-dimensional LBP feature (i.e., vein facial texture feature). The three extracted LBP features are spliced ​​to obtain a 3*512-dimensional multi-wavelength facial texture feature.

[0045] In step S140, the visible light face liveness semantic feature, the infrared face liveness semantic feature, the vein face liveness semantic feature and the multi-wavelength face texture feature are fused to obtain a face liveness fusion feature.

[0046] In one embodiment, the extracted visible light face liveness semantic features, infrared face liveness semantic features, vein face liveness semantic features, and multi-wavelength face texture features can be pre-combined, where each face liveness semantic feature has a dimension of 512, resulting in a combined dimension of 6*512. By setting an Alive token vector as the extracted face liveness fusion feature, the Alive token vector with an initial dimension of 1*512 is spliced ​​into the above combined features, resulting in a combined feature dimension of 7*512.

[0047] The combined features to be fused are transferred to the transformer encoder module, which contains a Multi-Head Attention module. The Multi-Head Attention module is composed of 8 Self-Attention modules. The Multi-Head Attention module can weightedly fuse the features of each token with those of the remaining tokens by learning the attention score. Therefore, the Alive token fuses the token features of vein face, infrared face and visible light face. The Multi-Head Attention module globally associates and fuses all multi-wavelength live features. After passing through the transformer's Multi-Head Attention module, the output feature dimension is still 7*512, and the feature 1*512 corresponding to the Alive token position is extracted as the face liveness fusion feature.

[0048] In step S150, recognition is performed based on the face liveness fusion feature to determine a liveness recognition result.

[0049] Based on the aforementioned embodiment, after determining the face liveness fusion feature, it can be input into the fully connected layer connected to the multi-head attention module, so that the fully connected layer outputs the corresponding liveness recognition result.

[0050] In one example, when training a network for feature fusion and liveness recognition, the classification classic cross-loss entropy function can be selected to calculate the loss between the network classification output and the label, and the error is back-propagated to the network layer through the chain rule to update the network parameters. It is worth noting that at this stage, only the network parameters of the transformer architecture are updated, while the parameters of the vein face liveness semantic feature extraction model, the infrared face liveness semantic feature extraction model, and the visible light face liveness semantic feature extraction model are not updated. The trained transformer model can fuse multi-wavelength face liveness features, and the fused face liveness fusion features integrate the most obvious liveness features of multiple wavelength bands (i.e., visible light face, infrared face, and vein face), which can fully defend against electronic screen face attacks, paper print face attacks, and 3D face attacks.

[0051] In one embodiment of the present application, there are multiple target faces. After obtaining the visible light face image, infrared face image, and vein face image corresponding to the target face, the method further includes:

[0052] performing image face detection based on the visible light face image, the infrared face image, and the vein face image, and determining position information of each target face in the visible light face image, the infrared face image, and the vein face image;

[0053] Based on the position information of each target face, the overlap between the target faces in different images is calculated to determine the visible light face, infrared face and vein face corresponding to the same human body.

[0054] In this embodiment, when multiple people are being enrolled, the same image may contain multiple faces. Therefore, image face detection can be performed based on each image to determine the position information of each target face contained in each image. The position information can be the position information of the detection box corresponding to each face.

[0055] Based on the positional information of each target face, the degree of overlap between target faces in different images can be calculated, such as the intersection-over-union ratio. When the overlap between target faces in different images exceeds a certain threshold, it can be determined that the two belong to the same person. This allows the identification of visible light faces, infrared faces, and vein-based faces from visible light, infrared, and vein-based face images as belonging to the same person, preparing for subsequent liveness recognition and preventing the incorrect association of faces from different individuals when multiple people are being recorded.

[0056] The following describes an embodiment of the device of the present application, which can be used to perform the liveness detection method based on multi-feature fusion described in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the liveness detection method based on multi-feature fusion described in the above embodiment of the present application.

[0057] Figure 2 A block diagram of a living body detection device based on multi-wavelength feature fusion according to an embodiment of the present application is shown.

[0058] Reference Figure 2 As shown, a liveness detection device based on multi-feature fusion according to an embodiment of the present application includes:

[0059] An acquisition module 210 is configured to acquire a visible light face image, an infrared face image, and a vein face image corresponding to a target face, wherein the vein face image includes vein information of the target face;

[0060] a semantic feature extraction module 220 for performing feature extraction on the visible light face image, the infrared face image, and the vein face image, respectively, based on pre-trained visible light face liveness semantic feature extraction models, infrared face liveness semantic feature extraction models, and vein face liveness semantic feature extraction models, to determine corresponding visible light face liveness semantic features, infrared face liveness semantic features, and vein face liveness semantic features, wherein the visible light face liveness semantic feature extraction model, the infrared face liveness semantic feature extraction model, and the vein face liveness semantic feature extraction model are respectively trained using real face data and other possible attack data;

[0061] a texture feature extraction module 230 for extracting texture features based on the visible light face image, the infrared face image, and the vein face image to determine corresponding multi-wavelength face texture features, where the multi-wavelength face texture features include visible light face texture features, infrared face texture features, and vein face texture features;

[0062] A feature fusion module 240 is configured to fuse the visible light face liveness semantic feature, the infrared face liveness semantic feature, and the vein face liveness semantic feature to obtain a face liveness fusion feature;

[0063] The processing module 250 is used to perform recognition based on the face liveness fusion feature to determine a liveness recognition result.

[0064] In one embodiment of the present application, the feature fusion module 240 is used to: based on the multi-head attention mechanism in the Transformer, perform feature fusion on the visible light face liveness semantic features, the infrared face liveness semantic features, the vein face liveness semantic features and the multi-wavelength face texture features to obtain the face liveness fusion features.

[0065] In one embodiment of the present application, the processing module 250 is configured to input the face liveness fusion feature into a fully connected layer, so that the fully connected layer outputs a corresponding liveness recognition result.

[0066] In one embodiment of the present application, the training data of the visible light face liveness semantic feature extraction model includes real visible light face pictures, conventional black and white paper face attack data, simulated black and white paper face attack data with vein patterns added, simulated 3D face attack data with vein patterns added, and corresponding label information; the training data of the infrared face liveness semantic feature extraction model includes real infrared face pictures, conventional color paper face attack data, simulated color paper face attack data with vein patterns added, simulated 3D face attack data with vein patterns added, and corresponding label information; the training data of the vein face liveness semantic feature extraction model includes real vein face pictures and conventional 3D face attack data and corresponding label information, wherein the real vein face pictures have vein information, and the conventional 3D face attack data has no vein information.

[0067] In one embodiment of the present application, the visible light face live semantic feature extraction model, the infrared face live semantic feature extraction model, and the vein face live semantic feature extraction model all include connected convolutional layers and fully connected layers, and the fully connected layer is used to perform dimensionality conversion on the features output by the convolutional layer.

[0068] In one embodiment of the present application, when training the visible light face live semantic feature extraction model, the infrared face live semantic feature extraction model, and the vein face live semantic feature extraction model, the classification classic cross entropy loss function is used to calculate the loss between the classification output and the label, and the error is back-propagated through the chain rule to update the network parameters, so that the model gradually converges.

[0069] In one embodiment of the present application, there are multiple target faces. After obtaining the visible light face image, infrared face image and vein face image corresponding to the target face, the acquisition module 210 is further used to: perform image face detection based on the visible light face image, the infrared face image and the vein face image, and determine the position information of each target face in the visible light face image, the infrared face image and the vein face image; calculate the overlap between the target faces in different images based on the position information of each target face, so as to determine the visible light face, infrared face and vein face corresponding to the same human body.

[0070] Figure 3 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown.

[0071] It should be noted that Figure 3 The computer system of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0072] like Figure 3 As shown, the computer system includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage part 308 to the random access memory (RAM) 303, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the RAM 303. The CPU 301, ROM 302 and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0073] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, and the like; an output section 307 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 308 including a hard disk and the like; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. Removable media 311, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 310 as needed, so that computer programs read therefrom can be installed into the storage section 308 as needed.

[0074] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 309, and / or installed from a removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, the various functions defined in the system of the present application are executed.

[0075] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable computer program. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0077] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.

[0078] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device implements the method described in the above embodiments.

[0079] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0080] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0081] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.

[0082] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for liveness detection based on multi-feature fusion, characterized in that: include: Acquire a visible light face image, an infrared face image, and a vein face image corresponding to a target face, wherein the vein face image includes vein information of the target face; Based on the pre-trained visible light face liveness semantic feature extraction model, infrared face liveness semantic feature extraction model, and vein face liveness semantic feature extraction model, semantic feature extraction is performed on the visible light face image, the infrared face image, and the vein face image, respectively, to determine the corresponding visible light face liveness semantic features, infrared face liveness semantic features, and vein face liveness semantic features, wherein the visible light face liveness semantic feature extraction model, the infrared face liveness semantic feature extraction model, and the vein face liveness semantic feature extraction model are respectively trained using real face data and attack data of other possible methods; Performing texture feature extraction based on the visible light face image, the infrared face image, and the vein face image to determine corresponding multi-wavelength face texture features, wherein the multi-wavelength face texture features include visible light face texture features, infrared face texture features, and vein face texture features; Fusing the visible light face liveness semantic feature, the infrared face liveness semantic feature, the vein face liveness semantic feature, and the multi-wavelength face texture feature to obtain a face liveness fusion feature; Identification is performed based on the live face fusion features to determine a liveness recognition result.

2. The method according to claim 1, characterized in that The visible light face liveness semantic feature, the infrared face liveness semantic feature, the vein face liveness semantic feature, and the multi-wavelength face texture feature are fused to obtain a face liveness fusion feature, including: Based on the multi-head attention mechanism in Transformer, the visible light face liveness semantic features, the infrared face liveness semantic features, the vein face liveness semantic features and the multi-wavelength face texture features are fused to obtain the face liveness fusion features.

3. The method according to claim 2, characterized in that Performing recognition based on the face liveness fusion feature to determine a liveness recognition result includes: The face liveness fusion feature is input into the fully connected layer so that the fully connected layer outputs the corresponding liveness recognition result.

4. The method according to any one of claims 1 to 3, characterized in that The training data of the visible light face living body semantic feature extraction model includes real visible light face pictures, conventional black and white paper face attack data, simulated black and white paper face attack data with vein patterns added, simulated 3D face attack data with vein patterns added, and corresponding label information; The training data of the infrared face living body semantic feature extraction model includes real infrared face images, conventional color paper face attack data, simulated vein pattern added color paper face attack data, simulated vein pattern added 3D face attack data and corresponding label information; The training data of the vein face living semantic feature extraction model includes real vein face pictures and conventional 3D face attack data and corresponding label information, wherein the real vein face pictures have vein information, and the conventional 3D face attack data does not have vein information.

5. The method according to claim 4, characterized in that The visible light face live semantic feature extraction model, the infrared face live semantic feature extraction model, and the vein face live semantic feature extraction model all include connected convolutional layers and fully connected layers, and the fully connected layer is used to perform dimensionality conversion on the features output by the convolutional layer.

6. The method according to claim 5, characterized in that When training the visible light face live semantic feature extraction model, the infrared face live semantic feature extraction model, and the vein face live semantic feature extraction model, the classification classic cross entropy loss function is used to calculate the loss between the classification output and the label, and the error is back-propagated through the chain rule to update the network parameters, so that the model gradually converges.

7. The method according to claim 1, characterized in that If there are multiple target faces, after obtaining the visible light face image, infrared face image, and vein face image corresponding to the target faces, the method further includes: performing image face detection based on the visible light face image, the infrared face image, and the vein face image, and determining position information of each target face in the visible light face image, the infrared face image, and the vein face image; Based on the position information of each target face, the overlap between the target faces in different images is calculated to determine the visible light face, infrared face and vein face corresponding to the same human body.

8. A living body detection device based on multi-feature fusion, characterized in that: include: An acquisition module is used to acquire a visible light face image, an infrared face image, and a vein face image corresponding to a target face, wherein the vein face image includes vein information of the target face; a semantic feature extraction module for performing semantic feature extraction on the visible light face image, the infrared face image, and the vein face image, respectively, based on pre-trained visible light face liveness semantic feature extraction models, infrared face liveness semantic feature extraction models, and vein face liveness semantic feature extraction models, to determine corresponding visible light face liveness semantic features, infrared face liveness semantic features, and vein face liveness semantic features, wherein the visible light face liveness semantic feature extraction model, the infrared face liveness semantic feature extraction model, and the vein face liveness semantic feature extraction model are respectively trained using real face data and other possible attack data; a texture feature extraction module, configured to extract texture features based on the visible light face image, the infrared face image, and the vein face image, and determine corresponding multi-wavelength face texture features, wherein the multi-wavelength face texture features include visible light face texture features, infrared face texture features, and vein face texture features; a feature fusion module, configured to fuse the visible light face liveness semantic feature, the infrared face liveness semantic feature, the vein face liveness semantic feature, and the multi-wavelength face texture feature to obtain a face liveness fusion feature; The processing module is used to perform recognition based on the face liveness fusion feature to determine the liveness recognition result.

9. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for liveness detection based on multi-feature fusion according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the liveness detection method based on multi-feature fusion as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multispectral-based fingerprint anti-counterfeiting method and system

    CN110443217A

  • Face living body detection method and device, electronic equipment and storage medium

    CN115147937A