A face-based live detection method, system, device and medium

By introducing a lightweight convolutional neural network and feature pyramid network, and combining face frames and key point information for multi-scale feature extraction, the problem of difficulty and lack of robustness of live detection is solved, and efficient live detection is achieved.

CN114399843BActive Publication Date: 2025-07-04UNIVERSAL UBIQUITOUS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111564532.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-20
Publication Date
2025-07-04
Estimated Expiration
2041-12-20

AI Technical Summary

Technical Problem

In the prior art, face-based live detection has problems such as difficult detection and poor robustness of the detection model.

Method used

The lightweight convolutional neural network mobilenetV2 and the feature pyramid network BiFPN are used to combine the face frame position information and key point position information, and through multi-scale feature extraction and feature mapping, the effective face feature value index is calculated, softmax processing and average value calculation is performed, and the final living score is obtained for live detection.

Benefits of technology

It improves the robustness and generalization ability of the in vivo detection model, reduces the complexity of the model, and improves the accuracy of in vivo detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114399843B_ABST
    Figure CN114399843B_ABST
Patent Text Reader

Abstract

This application relates to a face-based liveness detection method, system, device, and medium. Among them, the method includes: obtaining an image to be detected, performing face detection and key point detection on the image to be detected to obtain face frame position information and key point position information; then, according to the face frame position information, cropping and normalizing the image to be detected to obtain a target input image, and performing multi-scale feature extraction and feature mapping on the target input image through a lightweight convolutional network to obtain a target feature map; finally, processing the target feature map through the convolutional network downsampling scale and key point position information to obtain an effective face feature value index, performing softmax processing and average value calculation on the target feature map according to the effective face feature value index to obtain a final liveness score, and performing liveness detection according to the liveness score. Through this application, the robustness of the liveness model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of face recognition, and particularly to a face-based liveness detection method, system, device, and medium. Background Art

[0002] As an important machine vision technology, face recognition plays a significant role in the field of artificial intelligence. For example, face recognition technology is used to achieve access control, door locks, face payment, etc. With the wide use of face recognition technology, its security issues have attracted more and more attention. Among them, particular attention is paid to the technology of determining whether the face captured by the camera in the recognition device is a real face or a forged live body.

[0003] In related technologies, silent liveness detection technology can be divided into different detection technologies based on different hardware, such as monocular visible light, near-infrared NIR, or 3D structured light. Liveness detection based on near-infrared NIR has a large discrimination degree for screen attacks, but a small discrimination degree for high-definition printed paper, especially for high-definition black and white paper, and it is very difficult to distinguish and detect. Liveness detection based on 3D structured light can relatively accurately obtain the point cloud map and depth map of the face and background at close range, and can perform more accurate liveness detection. However, this method generally has a high cost and limited usage scenarios. Finally, based on a monocular visible light camera is the most common and convenient way to obtain face images, and the device cost is also low. However, liveness detection based on monocular visible light is difficult and has the problem of poor robustness.

[0004] Currently, for face-based liveness detection in related technologies, there are problems of large detection difficulty and poor robustness of the detection model, and no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of this application provide a face-based liveness detection method, system, device, and medium to at least solve the problems of large detection difficulty and poor robustness of the detection model in face-based liveness detection in related technologies.

[0006] In a first aspect, embodiments of this application provide a face-based liveness detection method, and the method includes:

[0007] Obtain an image to be detected, and perform face detection and key point detection on the image to be detected to obtain face frame position information and key point position information;

[0008] According to the face frame position information, crop and normalize the image to be detected to obtain a target input image, and perform multi-scale feature extraction and feature mapping on the target input image through a lightweight convolutional network to obtain a target feature map;

[0009] Process the target feature map through convolutional network downsampling and the key point position information to obtain valid face feature value indexes, perform softmax processing and average value calculation on the target feature map according to the valid face feature value indexes to obtain the final liveness score, and perform liveness detection according to the liveness score.

[0010] In some embodiments, performing multi-scale feature extraction and feature mapping on the target input image through a lightweight convolutional network to obtain a target feature map includes:

[0011] Performing feature extraction on the target input image through a Feature Pyramid Network (BiFPN) to obtain multiple feature maps of different scales, and respectively performing deconvolution operations and pointwise convolution operations on the multiple feature maps to obtain the target feature map that fuses multi-scale features.

[0012] In some embodiments, processing the target feature map through the convolutional network downsampling scale and the key point position information to obtain valid face feature value indexes includes:

[0013] Calculate the downsampling multiple according to the size of the target input image and the size of the target feature map, and respectively map each pixel point in the target feature map to each region of the target input image according to the downsampling multiple, and calculate the coordinates of each region, where each pixel point in the target feature map corresponds to each region in the target input image;

[0014] After calculating the maximum circumscribed matrix of the key points according to the key point position information, calculate the intersection over union of the coordinates of each region in the target input image and the maximum circumscribed matrix to obtain the calculated value of each region;

[0015] Compare the calculated value of each region with a preset threshold, and record the index of the region when the calculated value of the region is greater than the preset threshold to obtain the valid face feature value index.

[0016] In some embodiments, performing softmax processing and average value calculation on the target feature map according to the valid face feature value indexes to obtain the final liveness score includes:

[0017] Perform softmax processing on the target feature map to obtain a feature matrix, expand the feature matrix into a one-dimensional array, obtain the corresponding feature values in the one-dimensional array according to the valid face feature value indexes, and perform mean value calculation on the feature values to obtain the final liveness score.

[0018] In some of these embodiments, performing live detection and determination based on the live score includes:

[0019] Presetting a live threshold, comparing the live score with the preset live threshold. When the live score is greater than the preset live threshold, the image to be detected is a live body; otherwise, it is an attack body.

[0020] In a second aspect, an embodiment of the present application provides a live detection system based on a human face. The system includes:

[0021] An acquisition module, configured to acquire an image to be detected, perform face detection and key point detection on the image to be detected, and obtain face frame position information and key point position information;

[0022] A feature extraction module, configured to crop and normalize the image to be detected according to the face frame position information to obtain a target input image, and perform multi-scale feature extraction and feature mapping on the target input image through a lightweight convolutional network to obtain a target feature map;

[0023] A detection module, configured to process the target feature map through the downsampling scale of the convolutional network and the key point position information to obtain an effective face feature value index, perform softmax processing and average value calculation on the target feature map according to the effective face feature value index to obtain a final live score, and perform live detection according to the live score.

[0024] In some of these embodiments, the feature extraction module is further configured to perform feature extraction on the target input image through a Feature Pyramid Network (BiFPN) to obtain multiple feature maps of different scales, and perform deconvolution operations and pointwise convolution operations on the multiple feature maps respectively to obtain the target feature map that fuses multi-scale features.

[0025] In some of these embodiments, the detection module is further configured to calculate a downsampling multiple according to the size of the target input image and the size of the target feature map, and respectively map each pixel point in the target feature map to each region of the target input image according to the downsampling multiple, and calculate the coordinates of each region, where each pixel point in the target feature map corresponds to each region in the target input image.

[0026] After calculating the maximum circumscribed matrix of the key points according to the key point position information, calculate the intersection over union of the coordinates of each region in the target input image and the maximum circumscribed matrix to obtain a calculated value for each region.

[0027] Compare the calculated values of the respective regions with a preset threshold. When the calculated value of a region is greater than the preset threshold, record the index of the region to obtain the index of the effective face feature value.

[0028] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the face-based liveness detection method as described in the first aspect above.

[0029] In a fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the face-based liveness detection method as described in the first aspect above.

[0030] Compared with the related art, the face-based liveness detection method provided by the embodiment of the present application obtains a to-be-detected image, performs face detection and key point detection on the to-be-detected image to obtain face frame position information and key point position information; then, according to the face frame position information, crops and normalizes the to-be-detected image to obtain a target input image, and performs multi-scale feature extraction and feature mapping on the target input image through a lightweight convolutional network to obtain a target feature map; finally, processes the target feature map through the downsampling scale of the convolutional network and the key point position information to obtain an effective face feature value index, performs softmax processing and average value calculation on the target feature map according to the effective face feature value index to obtain a final liveness score, and performs liveness detection according to the liveness score.

[0031] The beneficial effects of the present application are as follows: 1. Introduce the anchor mechanism in the face detection technology into the liveness detection network, effectively improving the performance of the liveness detection model.

[0032] 2. Adopt the lightweight convolutional neural network mobilenetV2 and the feature pyramid network BiFPN to extract features of different scales, which can not only reduce the model complexity, but also improve the generalization ability of the liveness detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0034] Figure 1 is a flowchart of the face-based liveness detection method according to an embodiment of the present application;

[0035] Figure 2 is a schematic diagram of the liveness detection process according to an embodiment of the present application;

[0036] Figure 3 is a structural block diagram of a face-based live detection system according to an embodiment of the present application;

[0037] Figure 4 is an internal structural schematic diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0038] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without creative efforts fall within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as insufficient disclosure of the content of the present application.

[0039] Referring to "embodiment" in the present application means that the specific features, structures or characteristics described in combination with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in the present application can be combined with other embodiments without conflict.

[0040] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "an", "one kind", "the" and the like involved in this application do not indicate a quantity limitation and can represent a singular or plural number. The terms "comprising", "including", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may further include steps or units not listed, or may further include other steps or units inherent to these processes, methods, products or devices. The words such as "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application means greater than or equal to two. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0041] This embodiment provides a face-based live detection method. Figure 1 It is a flowchart of the face-based live detection method according to the embodiment of this application, as Figure 1 shown, and this process includes the following steps:

[0042] Step S101, obtain the image to be detected, and perform face detection and key point detection on the image to be detected to obtain the face frame position information and key point position information;

[0043] Figure 2 It is a schematic diagram of the live detection process according to the embodiment of this application, as Figure 2 shown, obtain the image to be detected through the monocular visible light camera in the face recognition device, and perform face detection operation and key point detection operation on the image to be detected to obtain the face frame position information and key point position information. In this embodiment, the face image is obtained through the monocular visible light camera, which is not only simple and convenient, but also has a lower device cost;

[0044] Step S102, according to the face frame position information, crop and normalize the image to be detected to obtain the target input image, and perform multi-scale feature extraction and feature mapping on the target input image through a lightweight convolutional network to obtain the target feature map;

[0045] As Figure 2As shown, in this embodiment, according to the position information of the face bounding box, a preset ratio is expanded around it to determine a larger rectangular area containing the face. Then, this rectangular area is cropped from the image to be detected and normalized to a specified size to obtain the target input image.

[0046] Next, a lightweight convolutional network is used to perform multi-scale feature extraction and feature mapping on the target input image to obtain the target feature map (FeatureMap). Preferably, in this embodiment, a Feature Pyramid Network (BiFPN) is used to perform feature extraction on the target input image to obtain multiple feature maps of different scales, and deconvolution operations and pointwise convolution operations are respectively performed on the multiple feature maps to obtain the target feature map that fuses multi-scale features.

[0047] For example, after scaling the size of the processed target input image to 224*224, feature processing of a four-layer Feature Pyramid Network (BiFPN) is performed on it to obtain multiple feature maps of different scales, namely P3, P4, P5, and P6. Among them, the sizes of these feature maps are: 56*56, 28*28, 14*14, and 7*7 respectively. Then, a deconvolution operation is performed on P6 to obtain a 14*14 feature map P5', and then a PW (Point Wise) operation is performed on P5' and P5 to increase the dimension of the feature without changing the size of the feature. Next, a deconvolution operation is performed on the feature map after dimension increase to obtain a 28*28 feature map P4', and then a PW operation is performed on P4' and P4 to increase the dimension of the feature without changing the size of the feature. And so on, corresponding operations are performed on other feature maps. Finally, the processed feature maps obtained above are input into a convolutional layer for convolution operation to reduce the feature dimension of the feature map to the number of target categories. It should be noted that the liveness detection in this embodiment is essentially a binary classification problem, so the number of target categories is two. After the above series of operations, finally, a 2*28*28 target feature map (FeatureMap) that fuses multi-scale features can be obtained;

[0048] Step S103: Process the target feature map through the downsampling scale of the convolutional network and the key point position information to obtain the effective face feature value index. Perform softmax processing and average value calculation on the target feature map according to the effective face feature value index to obtain the final liveness score, and perform liveness detection based on the liveness score.

[0049] As Figure 2 shown, the target feature map is processed through the downsampling scale of the convolutional network and the key point position information to obtain the effective face feature value index. Among them, the specific processing steps are as follows:

[0050] S1: Calculate the downsampling factor sz based on the size of the target input image and the size of the target feature map. The calculation formula is as shown in Equation (1) below:

[0051] sz = INP_SZ / MAP_SZ, (1)

[0052] where INP_SZ represents the size of the target input image, and MAP_SZ represents the size of the target feature map;

[0053] S2: According to the downsampling factor sz, map each pixel point in the target feature map to each region of the target input image, and calculate the coordinates of each region. Specifically, the coordinates of each region in the target input image: [xs:xe, ys:ye], and the calculation formulas are as shown in Equations (2)-(5) below:

[0054] xs = col_idx * sz - sz / 2 (2)

[0055] ys = row_idx * sz - sz / 2 (3)

[0056] xe = xs + sz (4)

[0057] ye = ys + sz (5)

[0058] where both col_idx and row_idx take values in the range of [0, 27].

[0059] It should be noted that the size of each region in the target input image is 8 * 8, and each pixel point in the target feature map corresponds to a 8 * 8 region in the target input image.

[0060] Therefore, in this embodiment, the coordinates of 28 * 28 8 * 8 regions in the target input image can be calculated;

[0061] S3: After calculating the maximum circumscribed matrix of the key points based on the key point position information, calculate the intersection over union (IoU) between the coordinates of each region in the target input image and the maximum circumscribed matrix to obtain the calculated value of each region. Specifically, according to the face key point position information, calculate the maximum circumscribed matrix M of the key points, and then calculate the intersection over union (IoU) between the coordinates [xs:xe, ys:ye] of the 28 * 28 8 * 8 regions in the target input image calculated in S2 and the maximum circumscribed matrix M to obtain the calculated value of each region;

[0062] S4: Compare the calculated values of each region obtained above with a preset threshold. When the calculated value of a region is greater than the preset threshold, record the index (idx) of this region, and finally obtain the index of the effective face feature value. That is, the regions recorded by these idxes contain the effective face feature regions to a certain extent.

[0063] Through the above steps, the effective face feature region can be more focused on, the interference of some background regions in the feature map can be removed, and the generalization performance of the model can be improved.

[0064] Furthermore, as Figure 2 shown, after obtaining the effective face feature value index, the target feature map is subjected to softmax processing and average value calculation according to the effective face feature value index, and the final liveness score is obtained. Then, liveness detection and judgment are performed based on this liveness score. The specific steps are as follows:

[0065] First, perform softmax processing on the target feature map to obtain a feature matrix; then expand the feature matrix into a one-dimensional array, and obtain the corresponding feature values in the one-dimensional array according to the effective face feature value index (idx); next, perform average calculation on these feature values to obtain the final liveness score.

[0066] Finally, preset a liveness threshold, and compare the obtained liveness score with the preset liveness threshold. When the liveness score is greater than the preset liveness threshold, it can be determined that the image to be detected is a live body, that is, a real face, otherwise it is an attack body, that is, not a real face.

[0067] Through the above steps S101 to S103, in this embodiment, the anchor mechanism in the face detection technology is introduced into the liveness detection network for liveness detection. Among them, the lightweight convolutional neural network mobilenetV2 and the feature pyramid network BiFPN are also used, which not only reduces the model complexity, but also improves the generalization ability of the liveness detection model. It solves the problems of difficult detection and weak robustness of the detection model in face-based liveness detection, and improves the robustness of the liveness model.

[0068] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0069] This embodiment also provides a face-based liveness detection system, which is used to implement the above embodiment and the preferred implementation manner, and those that have been described will not be repeated. As used below, terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0070] Figure 3It is a structural block diagram of a face-based liveness detection system according to an embodiment of the present application. As Figure 3 shown, the system includes an acquisition module 31, a feature extraction module 32, and a detection module 33:

[0071] The acquisition module 31 is used to acquire the image to be detected, perform face detection and key point detection on the image to be detected, and obtain the face frame position information and key point position information; the feature extraction module 32 is used to crop and normalize the image to be detected according to the face frame position information to obtain the target input image, and perform multi-scale feature extraction and feature mapping on the target input image through a lightweight convolutional network to obtain the target feature map; the detection module 33 is used to process the target feature map through the convolutional network downsampling scale and key point position information to obtain the effective face feature value index, perform softmax processing and average value calculation on the target feature map according to the effective face feature value index to obtain the final liveness score, and perform liveness detection according to the liveness score.

[0072] Through the above system, in this embodiment, the anchor mechanism in face detection technology is introduced into the liveness detection network for liveness detection, and the lightweight convolutional neural network mobilenetV2 and the feature pyramid network BiFPN are also used, which not only reduces the model complexity, but also improves the generalization ability of the liveness detection model. It solves the problems of large detection difficulty and weak robustness of the detection model in face-based liveness detection, and improves the robustness of the liveness model.

[0073] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be repeated here.

[0074] In addition, it should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combination form.

[0075] This embodiment also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0076] Optionally, the above-mentioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.

[0077] In addition, in combination with the face-based liveness detection method in the above embodiments, an embodiment of the present application can provide a storage medium to implement. A computer program is stored on the storage medium; when the computer program is executed by a processor, any one of the face-based liveness detection methods in the above embodiments is implemented.

[0078] In one embodiment, a computer device is provided. The computer device may be a terminal. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a face-based liveness detection method is implemented. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0079] In one embodiment, Figure 4 is a schematic internal structure diagram of an electronic device according to an embodiment of the present application, as Figure 4 shown, an electronic device is provided. The electronic device may be a server, and its internal structure diagram may be as Figure 4 shown. The electronic device includes a processor, a network interface, an internal memory, and a non-volatile memory connected through an internal bus. Among them, the non-volatile memory stores an operating system, a computer program, and a database. The processor is used to provide computing and control capabilities. The network interface is used to communicate with an external terminal through a network connection. The internal memory is used to provide an environment for the operation of the operating system and the computer program. When the computer program is executed by the processor, a face-based liveness detection method is implemented. The database is used to store data.

[0080] Those skilled in the art can understand that Figure 4 the structure shown in

[0081] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0082] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0083] The above embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A face-based live detection method, characterized in that, The method includes: Obtain the image to be detected, and perform face detection and key point detection on the image to be detected to obtain the face box position information and key point position information; According to the face box position information, crop and normalize the image to be detected to obtain a target input image, and perform multi-scale feature extraction and feature mapping on the target input image through a lightweight convolutional network to obtain a target feature map; Process the target feature map through the convolutional network downsampling scale and the key point position information to obtain an effective face feature value index, perform softmax processing and average value calculation on the target feature map according to the effective face feature value index to obtain a final liveness score, and perform liveness detection according to the liveness score. Processing the target feature map through the convolutional network downsampling scale and the key point position information to obtain an effective face feature value index includes: Calculate the downsampling multiple according to the size of the target input image and the size of the target feature map, and according to the downsampling multiple, map each pixel point in the target feature map to each area of the target input image respectively, and calculate the coordinates of each area, where each pixel point in the target feature map corresponds to each area in the target input image; After calculating the maximum circumscribed matrix of the key points according to the key point position information, calculate the intersection over union of the coordinates of each area in the target input image and the maximum circumscribed matrix to obtain the calculated value of each area; Compare the calculated value of each area with a preset threshold, and record the index of the area when the calculated value of the area is greater than the preset threshold to obtain the effective face feature value index.

2. The method according to claim 1, wherein Performing multi-scale feature extraction and feature mapping on the target input image through a lightweight convolutional network to obtain a target feature map includes: Through the Feature Pyramid Network BiFPN, perform feature extraction on the target input image to obtain multiple feature maps of different scales, and perform deconvolution operations and pointwise convolution operations on the multiple feature maps respectively to obtain the target feature map that fuses multi-scale features.

3. The method according to claim 1, characterized in that Performing softmax processing and average value calculation on the target feature map according to the effective face feature value index to obtain a final liveness score includes: Perform softmax processing on the target feature map to obtain a feature matrix, expand the feature matrix into a one-dimensional array, obtain the corresponding feature values in the one-dimensional array according to the effective face feature value index, and perform mean value calculation on the feature values to obtain the final liveness score.

4. The method according to claim 1, characterized in that Judging liveness detection according to the liveness score includes: Preset a liveness threshold, compare the liveness score with the preset liveness threshold, and when the liveness score is greater than the preset liveness threshold, the image to be detected is a live body, otherwise it is an attack body.

5. A face-based live detection system, characterized in that, The system includes: An acquisition module, configured to acquire the image to be detected, and perform face detection and key point detection on the image to be detected to obtain the face box position information and key point position information; A feature extraction module, configured to crop and normalize the image to be detected according to the face frame position information to obtain a target input image, and perform multi-scale feature extraction and feature mapping on the target input image through a lightweight convolutional network to obtain a target feature map; A detection module, configured to process the target feature map according to the downsampling scale of the convolutional network and the key point position information to obtain an effective face feature value index, perform softmax processing and average value calculation on the target feature map according to the effective face feature value index to obtain a final liveness score, and perform liveness detection according to the liveness score. The detection module is further configured to calculate a downsampling multiple according to the size of the target input image and the size of the target feature map, and respectively map each pixel point in the target feature map to each region of the target input image according to the downsampling multiple, and calculate the coordinates of each region, where each pixel point in the target feature map corresponds to each region in the target input image. After calculating the maximum circumscribed matrix of the key points according to the key point position information, calculate the intersection over union of the coordinates of each region in the target input image and the maximum circumscribed matrix to obtain a calculated value for each region. Compare the calculated value of each region with a preset threshold, and record the index of the region when the calculated value of the region is greater than the preset threshold to obtain the effective face feature value index.

6. The system according to claim 5, wherein The feature extraction module is further configured to perform feature extraction on the target input image through a Feature Pyramid Network (BiFPN) to obtain multiple feature maps of different scales, and perform deconvolution operations and pointwise convolution operations on the multiple feature maps respectively to obtain the target feature map that fuses multi-scale features.

7. An electronic device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to execute the face-based liveness detection method according to any one of claims 1 to 4.

8. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the face-based liveness detection method according to any one of claims 1 to 4 when running.

Citation Information

Patent Citations

  • Living body detection method based on infrared characteristics

    CN111767877A

  • Face detection method and device, electronic equipment and storage medium

    CN111783749A