Living body detection method, device, electronic device and storage medium

By using dense fixed points and constructing an accumulated adjacency matrix to detect face videos, the problem of low accuracy in vivo detection in the prior art is solved, and higher robustness and accuracy are achieved, especially when facing attack types such as 3D printed masks.

CN119360457BActive Publication Date: 2025-05-27JINAN BOGUAN INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411933466.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-27
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

In the prior art, the accuracy of live detection is low, especially the robustness of changes in the face image content and lighting environment, and there is a misjudgment of attack types such as 3D printed masks.

Method used

By performing dense fixed-pointing of the user's face video to be detected, the dense fixed-point results are determined, and an accumulated adjacency matrix is ​​constructed based on these results to characterize the face change pattern. The accumulated adjacency matrix and target dense fixed-point results are input into the graph neural network to determine the live detection results.

Benefits of technology

It improves the robustness of the graph neural network to the face image content and lighting environment, improves the accuracy of live detection, and enhances the protection effect of attack types such as 3D printed masks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360457B_ABST
    Figure CN119360457B_ABST
Patent Text Reader

Abstract

The present invention provides a live detection method, device, electronic device and storage medium, relating to the field of artificial intelligence technology. The method includes: obtaining a face video to be detected corresponding to a user, and determining the dense fixed-point results of each frame of face image to be detected in the face video to be detected; based on each dense fixed-point result, determining an accumulated adjacency matrix corresponding to the face video to be detected; the accumulated adjacency matrix is used to characterize the face change law in the face video to be detected; inputting the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into a graph neural network to determine the live detection result corresponding to the user. The present invention can improve the robustness of the graph neural network to face image content and illumination environment, and improve the accuracy of live detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to a method, device, electronic device, and storage medium for living body detection. Background Art

[0002] With the development of technology, living body detection plays an important role in identity recognition and has become an important checkpoint for information security protection. For example, when a user inputs an account in an electronic device to log in and use it, the living body detection function needs to be used for authentication to ensure that the account is used by the real person, rather than a printed photo, electronic video, or 3D (Three Dimensional) printed mask of someone else, providing security protection for logging in to the account.

[0003] In the prior art, when using a silent living body detection method, generally, the features of a face image are extracted through a deep learning network, and a discrimination result is output according to the features of the face image. However, this method is relatively sensitive to changes in the content of the face image and the lighting environment, has weak robustness, and misjudges attack types such as 3D printed masks, resulting in low detection accuracy. Summary of the Invention

[0004] The present invention provides a method, device, electronic device, and storage medium for living body detection to solve the defect of low accuracy in living body detection in the prior art.

[0005] The present invention provides a method for living body detection, including the following steps.

[0006] Obtain a face video to be detected corresponding to a user, and determine the dense fixed-point results of each frame of face image to be detected in the face video to be detected.

[0007] Based on each of the dense fixed-point results, determine an accumulated adjacency matrix corresponding to the face video to be detected; the accumulated adjacency matrix is used to characterize the face change rule in the face video to be detected.

[0008] Input the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into a graph neural network to determine the living body detection result corresponding to the user.

[0009] According to the method for living body detection provided by the present invention, the step of determining the accumulated adjacency matrix corresponding to the face video to be detected based on each of the dense fixed-point results includes:

[0010] Based on each of the dense fixed-point results, determine a weighted adjacency matrix corresponding to each face image to be detected; the weighted adjacency matrix is used to characterize the positional relationship between each face partition in the face area, and each face partition is obtained by dividing the face area according to the positions of facial features;

[0011] Based on a preset association relationship, perform sparse processing on each of the weighted adjacency matrices to obtain a sparse adjacency matrix corresponding to each of the to-be-detected face images; the preset association relationship is used to characterize the change correlation between each face partition in the face region.

[0012] Based on the sparse adjacency matrices corresponding to each to-be-detected face image in each sliding window within the to-be-detected face video, determine an accumulated adjacency matrix corresponding to each of the sliding windows; the sliding window includes multiple consecutive frames of to-be-detected face images.

[0013] According to the live detection method provided by the present invention, the step of determining an accumulated adjacency matrix corresponding to each of the sliding windows based on the sparse adjacency matrices corresponding to each to-be-detected face image in each sliding window within the to-be-detected face video includes:

[0014] For multiple consecutive frames of to-be-detected face images in each of the sliding windows, based on the sparse adjacency matrices corresponding to two adjacent frames of to-be-detected face images, determine the mean square error corresponding to the two adjacent frames of to-be-detected face images; determine the sum of all mean square errors as the accumulated adjacency matrix corresponding to the sliding window.

[0015] According to the live detection method provided by the present invention, each of the dense fixed-point results includes multiple feature points in the face region and the fixed-point positions corresponding to each of the feature points.

[0016] According to the live detection method provided by the present invention, the step of determining a weighted adjacency matrix corresponding to each of the to-be-detected face images based on each of the dense fixed-point results includes:

[0017] For each of the dense fixed-point results, based on the fixed-point positions corresponding to any two feature points in the face region, determine the distance between the any two feature points; based on all the distances, determine the weighted adjacency matrix corresponding to the to-be-detected face image.

[0018] According to the live detection method provided by the present invention, the step of performing sparse processing on each of the weighted adjacency matrices based on the preset association relationship to obtain a sparse adjacency matrix corresponding to each of the to-be-detected face images includes:

[0019] For each of the to-be-detected face images, based on the preset association relationship, determine at least one sparse partition corresponding to each of the face partitions; for each of the face partitions, in the weighted adjacency matrix corresponding to the to-be-detected face image, set the distances between each feature point in the face partition and each feature point in each of the sparse partitions to zero to obtain the sparse adjacency matrix corresponding to the to-be-detected face image.

[0020] According to the living body detection method provided by the present invention, inputting the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into a graph neural network to determine the living body detection result corresponding to the user includes:

[0021] Input the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into the graph neural network to obtain the face discrimination results corresponding to each of the sliding windows sequentially output by the graph neural network; the target dense fixed-point result is the dense fixed-point result corresponding to the last frame of the face image to be detected in the sliding window corresponding to each accumulated adjacency matrix.

[0022] Judge whether each of the face discrimination results meets a preset condition, and determine the living body detection result corresponding to the user based on the judgment result.

[0023] The preset condition includes any one of the following:

[0024] The face discrimination result is true.

[0025] The face discrimination results of the first preset number of consecutive ones are all true.

[0026] Among the face discrimination results of the second preset number of consecutive ones, the proportion of the face discrimination results that are true reaches a preset threshold.

[0027] The present invention also provides a living body detection device, including the following modules.

[0028] An acquisition module, configured to acquire a face video to be detected corresponding to a user, and determine the dense fixed-point results of each frame of the face image to be detected in the face video to be detected.

[0029] A determination module, configured to determine the accumulated adjacency matrix corresponding to the face video to be detected based on each of the dense fixed-point results; the accumulated adjacency matrix is used to characterize the face change rule in the face video to be detected.

[0030] A detection module, configured to input the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into a graph neural network to determine the living body detection result corresponding to the user.

[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the living body detection method as described in any one of the above is implemented.

[0032] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the living body detection method as described in any one of the above is implemented.

[0033] The present invention also provides a computer program product, including a computer program, which when executed by a processor implements the living body detection method as described in any one of the above.

[0034] The living body detection method, device, electronic device and storage medium provided by the present invention determine the dense fixed-point result by performing dense fixed-point on each frame of the face image to be detected in the face video to be detected corresponding to the user, and determine the cumulative adjacency matrix for characterizing the face change law corresponding to the face video to be detected according to each dense fixed-point result, and input the cumulative adjacency matrix and the target dense fixed-point result corresponding to the cumulative adjacency matrix into the graph neural network, and then determine the living body detection result corresponding to the user. In the present invention, each frame of the face image to be detected is not directly used as the input of the graph neural network, but the cumulative adjacency matrix reflecting the face change law determined according to the dense fixed-point result and reflecting the positions of face muscles and facial features is used as the input of the graph neural network, which improves the robustness of the graph neural network to the face image content and the illumination environment, improves the accuracy of the graph neural network and the protection effect against attack types such as 3D printed masks, and further improves the accuracy of living body detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0036] Figure 1 is a flowchart of the living body detection method provided by an embodiment of the present invention.

[0037] Figure 2 is a schematic diagram of the fixed face area provided by an embodiment of the present invention.

[0038] Figure 3 is a schematic diagram of the dense fixed-point result provided by an embodiment of the present invention.

[0039] Figure 4 is a schematic diagram of the partition of the dense fixed-point result provided by an embodiment of the present invention.

[0040] Figure 5 is a schematic diagram of the sliding window provided by an embodiment of the present invention.

[0041] Figure 6 is a schematic diagram of the structure of the living body detection device provided by an embodiment of the present invention.

[0042] Figure 7 is a schematic diagram of the structure of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0043] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0044] Aiming at the problem of low accuracy in live detection in the prior art, an embodiment of the present invention provides a live detection method. Figure 1 It is a schematic flowchart of the live detection method provided by the embodiment of the present invention. As Figure 1 shown, the method includes the following steps 110 to 130.

[0045] Step 110: Obtain a face video to be detected corresponding to a user, and determine the dense fixed-point results of each frame of face image to be detected in the face video to be detected.

[0046] Specifically, Figure 2 It is a schematic diagram of a fixed face area provided by the embodiment of the present invention. As Figure 2 shown, when the user performs login authentication, place the face within the fixed face area on the login interface and silently gaze for a preset duration. The acquisition device can acquire the face video to be detected that lasts for the preset duration and is within the fixed face area. During the acquisition process, there is no need for the user to cooperate to make actions such as opening the mouth, blinking, and shaking the head. Only the user needs to silently gaze for the preset duration, which greatly reduces the usage complexity. After the acquisition device acquires the face video to be detected, it can send the face video to be detected to an electronic device, and the electronic device can then obtain the face video to be detected corresponding to the user. Then, for each frame of face image in the face video to be detected, the electronic device can detect the face area in each frame of face image to be detected and perform dense fixed-pointing on each face area, that is, locate a large number of dense feature points in each face area. These feature points can more precisely describe the contour, facial features, and muscles of the face area, so as to obtain the dense fixed-point results of each frame of face image to be detected.

[0047] It should be noted that each of the dense fixed-point results includes multiple feature points in the face area and the fixed-point positions corresponding to each of the feature points. The number of feature points obtained by dense fixed-pointing is relatively large. For example, the number of feature points can be 20,000, 30,000, or 50,000, etc. The embodiments of the present invention do not limit this. For example, Figure 3 It is a schematic diagram of the dense fixed-point result provided by the embodiment of the present invention. As Figure 3 shown, for Figure 3Dense fixed points are performed on the face region in (a) to obtain a large number of dense feature points and the corresponding fixed-point positions of each feature point. According to these feature points and the corresponding fixed-point positions, it is possible to construct Figure 3 the dense face feature point map shown in (b). Adding the face feature points to the corresponding face regions can obtain Figure 3 the updated face image shown in (c).

[0048] Optionally, dense fixed points can be performed on each face image by methods based on feature points or deep learning methods, etc. Among them: The method based on feature points can be: by detecting feature points in the face region (such as the corners of the eyes, the corners of the mouth, and the tips of the eyebrows, etc.) and interpolating between these feature points to achieve the effect of dense fixed points. The method based on deep learning can be: using deep learning methods such as neural networks to identify dense feature points in the face region.

[0049] Optionally, each frame of the face image to be detected is obtained by frame extraction from the face video to be detected. The frame rate can be set according to experience, and the embodiments of the present invention do not limit this. In addition, the sampling frequency of the face video to be detected can be set according to experience. For example, the sampling frequency can be 15fps or 20fps, etc., and the embodiments of the present invention do not limit this.

[0050] It should be noted that the preset duration is the acquisition duration of the face video to be detected. This preset duration can be set according to experience, but this preset duration does not need to be set too long to avoid affecting the detection efficiency. For example, the preset duration can be 3s, 4s or 5s, and the embodiments of the present invention do not limit this.

[0051] In addition, the face fixed area can be located at the center position or above the central axis of the login interface. The shape of the face fixed area can be circular, elliptical or square, etc., to ensure that the user's face area can be covered at a relatively close distance, and the embodiments of the present invention do not limit this.

[0052] It should be noted that the face region can be divided into 7 face sub-regions. For example, the left eye and left eyebrow region of the face are divided into face sub-region 1, the right eye and right eyebrow region of the face are divided into face sub-region 2, the nose region of the face is divided into face sub-region 3, the left face region is divided into face sub-region 4, the right face region is divided into face sub-region 5, the mouth region is divided into face sub-region 6, and the other regions of the face are divided into face sub-region 7. After determining the dense fixed-point results, each feature point can be divided into a feature point subset corresponding to each face sub-region according to the face sub-region where each feature point is located in each dense fixed-point result. That is, the feature points in face sub-region 1 in the dense fixed-point result are divided into feature point subset 1, the feature points in face sub-region 2 in the dense fixed-point result are divided into feature point subset 2, the feature points in face sub-region 3 in the dense fixed-point result are divided into feature point subset 3, the feature points in face sub-region 4 in the dense fixed-point result are divided into feature point subset 4, the feature points in face sub-region 5 in the dense fixed-point result are divided into feature point subset 5, the feature points in face sub-region 6 in the dense fixed-point result are divided into feature point subset 6, and the feature points in face sub-region 7 in the dense fixed-point result are divided into feature point subset 7. For example, Figure 4 is a schematic diagram of the dense fixed-point result partitioning provided by an embodiment of the present invention. After partitioning the Figure 3 corresponding dense fixed-point result, the schematic diagrams corresponding to the feature point subsets of each frame of the face image to be detected as shown in Figure 4 can be obtained.

[0053] Step 120: Based on each of the dense fixed-point results, determine an accumulated adjacency matrix corresponding to the face video to be detected; the accumulated adjacency matrix is used to characterize the face change rule in the face video to be detected.

[0054] Specifically, after determining the dense fixed-point results corresponding to each frame of the face image to be detected, an accumulated adjacency matrix corresponding to multiple consecutive frames of the face image to be detected in the face video to be detected can be determined. The accumulated adjacency matrix is used to characterize the face change rule in multiple consecutive frames of the face image to be detected. That is, through this accumulated adjacency matrix, the subtle change rules of the facial features and muscles in multiple consecutive frames of the face image to be detected can be clarified.

[0055] Further, the determining the accumulated adjacency matrix corresponding to the face video to be detected based on each of the dense fixed-point results includes:

[0056] Based on each of the dense fixed-point results, determine a weighted adjacency matrix corresponding to each of the face images to be detected; the weighted adjacency matrix is used to characterize the positional relationship corresponding to each face sub-region in the face region, and each of the face sub-regions is divided according to the positions of the facial features in the face region;

[0057] Based on a preset association relationship, perform sparse processing on each of the weighted adjacency matrices to obtain a sparse adjacency matrix corresponding to each of the to-be-detected face images; the preset association relationship is used to characterize the change correlation between each face partition in the face region.

[0058] Based on the sparse adjacency matrices corresponding to each of the to-be-detected face images in each sliding window within the to-be-detected face video, determine an accumulated adjacency matrix corresponding to each of the sliding windows; the sliding window includes multiple consecutive frames of to-be-detected face images.

[0059] Specifically, after determining the dense fixed-point results corresponding to each frame of the to-be-detected face images, the weighted adjacency matrix corresponding to each to-be-detected face image can be calculated according to the positional relationship between the feature points. This weighted adjacency matrix is used to characterize the position and structure of each face partition in a single frame of the to-be-detected face image. For example, the elements corresponding to feature point subset 1 and feature point subset 2 in this weighted adjacency matrix are used to describe the opening and closing of the eyes and the shape of the eyebrows, the elements corresponding to feature point subset 4 and feature point subset 5 are used to describe the slight movement of the facial muscles and expressions, and the element corresponding to feature point subset 6 is used to describe the opening and closing and slight movement of the mouth, etc. After determining each weighted adjacency matrix, according to the change correlation between each face partition in the preset association relationship, perform sparse processing on the element values in each weighted adjacency matrix, so as to obtain a sparse adjacency matrix corresponding to each to-be-detected face image. Then, when the sliding window is at the initial moment, the accumulated adjacency matrix corresponding to this sliding window can be determined according to the sparse adjacency matrices corresponding to multiple consecutive frames of to-be-detected face images included in this sliding window. After that, with each movement of the sliding window, the accumulated adjacency matrix can be re-determined once. Each accumulated adjacency matrix can be used to describe the subtle change rules of the facial features and muscles in multiple consecutive frames of to-be-detected face images.

[0060] Optionally, the length of the sliding window is the number of frames of multiple consecutive frames of to-be-detected face images included in the sliding window. This number of frames can be set according to experience. For example, the length of the sliding window is 40 frames, 45 frames, or 50 frames, etc. Figure 5 is a schematic diagram of the sliding window provided by an embodiment of the present invention. As Figure 5 shown, taking the length of the sliding window as 45 frames as an example, at the initial moment, the sliding window includes to-be-detected face image 1 to to-be-detected face image 45. After the sliding window slides once, it includes to-be-detected face image 2 to to-be-detected face image 46. After the sliding window slides twice, it includes to-be-detected face image 3 to to-be-detected face image 47, and so on. With each movement of the sliding window, the accumulated adjacency matrix including 45 frames of to-be-detected face images can be re-determined once.

[0061] Further, the determining, based on each of the dense fixed-point results, a weighted adjacency matrix corresponding to each of the to-be-detected face images includes:

[0062] For each of the dense fixed-point results, based on the fixed-point positions corresponding to any two feature points in the face region, determine the distance between the two feature points; based on all the distances, determine the weighted adjacency matrix corresponding to the face image to be detected.

[0063] Specifically, for each dense fixed-point result, according to the fixed-point positions corresponding to any two feature points, the distance between the two feature points can be calculated, and this distance is used as an element in the weighted adjacency matrix. After determining the distances between all feature points, the weighted adjacency matrix corresponding to the face image to be detected can be obtained. This weighted adjacency matrix can simultaneously represent the connectivity and connection distance between feature points, and can be automatically adjusted according to the loss during the training of the graph neural network, improving the model accuracy of the graph neural network. The weighted adjacency matrix is shown in Equation (1), and Equation (1) is:

[0064] ,

[0065] where M represents the weighted adjacency matrix, v i,j represents the distance between the i-th feature point and the j-th feature point in the dense fixed-point result, and n represents the number of all feature points in the dense fixed-point result.

[0066] Further, the step of performing sparse processing on each of the weighted adjacency matrices based on a preset association relationship to obtain a sparse adjacency matrix corresponding to each of the face images to be detected includes:

[0067] For each of the face images to be detected, based on the preset association relationship, determine at least one sparse partition corresponding to each of the face partitions; for each of the face partitions, in the weighted adjacency matrix corresponding to the face image to be detected, set the distances between each feature point in the face partition and each feature point in each of the sparse partitions to zero, to obtain the sparse adjacency matrix corresponding to the face image to be detected.

[0068] Specifically, the preset association relationship includes the association relationship between each face partition and the corresponding sparse partition. For each face partition, each association relationship is used to characterize that the change correlation between this face partition and the corresponding sparse partition is weak, that is, in this face partition and the corresponding sparse partition, a subtle change in the muscles of one partition will not cause a muscle change in the other partition. For example, the preset association relationship may include: the sparse partition corresponding to face partition 1 includes face partitions 3, 5, 6, and 7; the sparse partition corresponding to face partition 2 includes face partitions 3, 4, 6, and 7; the sparse partition corresponding to face partition 3 includes face partitions 1, 2, 6, and 7; the sparse partition corresponding to face partition 4 includes face partitions 2, 5, 6, and 7; the sparse partition corresponding to face partition 5 includes face partitions 1, 4, 6, and 7; the sparse partition corresponding to face partition 6 includes face partitions 1, 2, 4, 5, and 7; the sparse partition corresponding to face partition 7 includes face partitions 1, 2, 3, 4, 5, and 7.

[0069] Therefore, after determining the weighted adjacency matrix corresponding to each frame of the face image to be detected, at least one sparse partition corresponding to each face partition can be determined from the preset association relationship, and the adjacency relationship value between the feature points corresponding to each face partition and the feature points corresponding to each sparse partition is set to zero, that is, the distance value between the feature points corresponding to each face partition and the feature points corresponding to each sparse partition in the weighted adjacency matrix is set to zero. For example, for face partition 1, the distance between the feature points in subset 1 of the feature points corresponding to face partition 1 and the feature points in subset 3 of the feature points corresponding to face partition 3 is set to zero. If subset 1 of the feature points includes feature points 1 to 5000 and subset 3 of the feature points includes feature points 15000 to 20000, then the element v in the weighted adjacency matrix p,q is set to zero, and is any natural number, and Any natural number among them. Similarly, the distance between the feature points in the feature point subset 1 corresponding to the face partition 1 and the feature points in the feature point subset 5 corresponding to the face partition 5 can be set to zero, the distance between the feature points in the feature point subset 1 corresponding to the face partition 1 and the feature points in the feature point subset 6 corresponding to the face partition 6 can be set to zero, and the distance between the feature points in the feature point subset 1 corresponding to the face partition 1 and the feature points in the feature point subset 7 corresponding to the face partition 7 can be set to zero. After traversing each face partition, the weighted adjacency matrix can be sparsified to obtain the sparse adjacency matrix corresponding to the weighted adjacency matrix, and each frame of the face image to be detected corresponds to a sparse adjacency matrix. The sparse matrix can represent the regions associated with facial micro-movements and remove the influence of irrelevant regions, thereby improving the accuracy of subsequent live detection.

[0070] Further, determining the cumulative adjacency matrix corresponding to each sliding window based on the sparse adjacency matrices corresponding to the face images to be detected in each sliding window of the face video to be detected includes:

[0071] For consecutive multiple frames of face images to be detected in each sliding window, based on the sparse adjacency matrices corresponding to two adjacent frames of face images to be detected, determine the mean square error corresponding to the two adjacent frames of face images to be detected; determine the sum of all mean square errors as the cumulative adjacency matrix corresponding to the sliding window.

[0072] Specifically, after determining the sparse adjacency matrices corresponding to the face images to be detected, the cumulative adjacency matrix corresponding to each sliding window can be calculated using Equation (2), and Equation (2) is:

[0073] ,

[0074] where M c represents the cumulative adjacency matrix corresponding to the sliding window, represents the i-th frame of the face image to be detected in the sliding window, f represents the length of the sliding window, that is, the number of frames of the face images to be detected included in the sliding window, and i is an integer and .

[0075] Step 130: Input the cumulative adjacency matrix and the target dense fixed-point result corresponding to the cumulative adjacency matrix into a Graph Neural Network (GNN) to determine the live detection result corresponding to the user.

[0076] Specifically, after determining each dense fixed-point result and each accumulated adjacency matrix, the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix are input into a pre-trained graph neural network. The graph neural network is trained based on dense fixed-point samples, the sample accumulated adjacency matrices corresponding to the dense fixed-point samples, and the classification labels corresponding to the sample accumulated adjacency matrices, so as to obtain the liveness detection result corresponding to the user, that is, to determine whether the face image to be detected is an image collected from a real face or an image collected from a prosthesis such as a printed photo, an electronic video, or a 3D printed mask. This liveness detection result can be applied to application scenarios such as access control detection or user login authentication. The embodiments of the present invention do not limit this. For example, taking the user login authentication scenario as an example, if the liveness detection result is true, the user login authentication passes; if the liveness detection result is false, the user login authentication fails.

[0077] Further, the step of inputting the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into the graph neural network to determine the liveness detection result corresponding to the user includes:

[0078] Input the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into the graph neural network to obtain the face discrimination results corresponding to each of the sliding windows sequentially output by the graph neural network; the target dense fixed-point result is the dense fixed-point result corresponding to the last frame of the face image to be detected in each sliding window corresponding to the accumulated adjacency matrix;

[0079] Judge whether each of the face discrimination results meets a preset condition, and determine the liveness detection result corresponding to the user based on the judgment result;

[0080] The preset condition includes any one of the following:

[0081] The face discrimination result is true;

[0082] The face discrimination results of the first preset number of consecutive ones are all true;

[0083] Among the face discrimination results of the second preset number of consecutive ones, the proportion of the face discrimination results that are true reaches a preset threshold.

[0084] Specifically, for the cumulative adjacency matrix corresponding to each sliding window, the dense fixed-point result corresponding to the last frame of the face image to be detected in the sliding window is determined as the target dense fixed-point result of the cumulative adjacency matrix corresponding to the sliding window. For example, taking the length of the sliding window as 45 frames, when the sliding window 1 includes the face images to be detected from image 1 to image 45, the face image 45 can be determined as the target dense fixed-point result of the cumulative adjacency matrix corresponding to the sliding window 1. Another example is that when the sliding window 2 includes the face images to be detected from image 2 to image 46, the face image 46 can be determined as the target dense fixed-point result of the cumulative adjacency matrix corresponding to the sliding window 2. After determining the target dense fixed-point results corresponding to each cumulative adjacency matrix, when each cumulative adjacency matrix and the target dense fixed-point result corresponding to each cumulative adjacency matrix are input into the graph neural network, the graph neural network can determine the face discrimination result corresponding to the current sliding window through the subtle change rules of facial features and muscles reflected in each cumulative adjacency matrix to judge whether the face is real or fake. Among them, the subtle change rules reflected by the cumulative adjacency matrix corresponding to a real person are non-rigid, while the subtle change rules reflected by the cumulative adjacency matrix corresponding to a prosthesis are rigid. Then, the live detection result corresponding to the user is related to the preset conditions. The first one is: when the preset condition is that the face discrimination result of a single frame is true, if the face discrimination result is false, the sliding window slides once, and a new face discrimination result is determined according to the next cumulative adjacency matrix. If the face discrimination result is true, it is determined that the live detection result corresponding to the user is true. The second one is: when the preset condition is that the face discrimination results of the first preset number of consecutive frames are all true, that is, only when the face discrimination results of the first preset number of consecutive frames are all true, it is determined that the live detection result corresponding to the user is true. There is no false face discrimination result allowed among the face discrimination results of the first preset number of consecutive frames. If the face discrimination result is false, the sliding window slides once, and a new face discrimination result is determined according to the next cumulative adjacency matrix, and the count is restarted to determine the live detection result. The third one is: when the preset condition is that among the face discrimination results of the second preset number of consecutive frames, the proportion of true face discrimination results reaches the preset threshold. Among the face discrimination results of the second preset number of consecutive frames, there can be false face discrimination results. As long as the number of true face discrimination results among the face discrimination results of the second preset number reaches the preset threshold in the second preset number, it can be considered that the live detection result is true. For example, the second preset number is 10 and the preset threshold is 80%. Among the 10 consecutive face discrimination results, as long as there are 8 or more face discrimination results that are true, it can be considered that the live detection result of the user is true, and the position of the true face discrimination results among the 10 face discrimination results is not limited. This preset condition is related to the application scenario. If the application scenario is the user login authentication scenario and the logged-in account is an entertainment account, the first preset condition can be adopted to achieve fast login authentication.When the logged-in account involves personal privacy information or amounts, the second or third preset condition can be adopted, and by setting a high threshold, the detection accuracy and authentication security can be ensured. The embodiments of the present invention do not make limitations in this regard.

[0085] It should be noted that before performing step 130, an initial graph neural network needs to be constructed first. The nodes of this initial graph neural network are dense fixed-point results, and the edges between the nodes can be the distances or adjacency relationships or other high-dimensional features between the nodes. The embodiments of the present invention do not make limitations in this regard. After constructing the initial graph neural network, real human face video samples and prosthetic face video samples can be obtained, and dense fixed-point operations are respectively performed on the real human face video samples and the prosthetic face video samples. Then, according to the dense fixed-point result samples, the cumulative adjacency sample matrices corresponding to the real human face video samples and the prosthetic face video samples are respectively determined with reference to the above content, and according to the categories corresponding to the real human face video samples and the prosthetic face video samples, the classification labels of the cumulative adjacency sample matrices corresponding to them are determined. For example, if the classification label is 1, it indicates that the cumulative adjacency sample matrix is the adjacency matrix corresponding to the real human face video sample; if the classification label is 0, it indicates that the cumulative adjacency sample matrix is the adjacency matrix corresponding to the prosthetic face video sample. After that, various types of dense fixed-point result samples, each cumulative adjacency sample matrix, and the corresponding classification labels are input into the initial graph neural network to train the initial graph neural network. The initial graph neural network can learn the association relationships between various types of cumulative adjacency sample matrices and the corresponding dense fixed-point result samples, and distinguish the differences in the subtle change rules of various types of face regions until the initial graph neural network converges, and then the trained graph neural network can be obtained.

[0086] The living body detection method provided by the embodiments of the present invention determines the dense fixed-point result by performing dense fixed-point operations on each frame of the face image to be detected in the face video to be detected corresponding to the user. According to each dense fixed-point result, the cumulative adjacency matrix for characterizing the face change rule corresponding to the face video to be detected is determined, and the cumulative adjacency matrix and the target dense fixed-point result corresponding to the cumulative adjacency matrix are input into the graph neural network, and then the living body detection result corresponding to the user is determined. In the present invention, each frame of the face image to be detected is not directly used as the input of the graph neural network, but the cumulative adjacency matrix reflecting the face change rule of the positions of facial muscles and facial features determined according to the dense fixed-point result is used as the input of the graph neural network, which improves the robustness of the graph neural network to the face image content and the illumination environment, improves the accuracy of the graph neural network and the protection effect against attack types such as 3D printed masks, and further improves the accuracy of living body detection.

[0087] Next, the living body detection device provided by the present invention will be described. The living body detection device described below can be mutually referred to with the living body detection method described above.

[0088] An embodiment of the present invention also provides a live detection device. Figure 6 It is a schematic structural diagram of the live detection device provided by the embodiment of the present invention. As Figure 6 shown, the live detection device 600 includes an acquisition module 610, a determination module 620, and a detection module 630.

[0089] The acquisition module 610 is configured to acquire a to-be-detected face video corresponding to a user, and determine the dense fixed-point results of each frame of the to-be-detected face image in the to-be-detected face video.

[0090] The determination module 620 is configured to determine an accumulated adjacency matrix corresponding to the to-be-detected face video based on each of the dense fixed-point results; the accumulated adjacency matrix is used to characterize the face change rule in the to-be-detected face video.

[0091] The detection module 630 is configured to input the accumulated adjacency matrix and a target dense fixed-point result corresponding to the accumulated adjacency matrix into a graph neural network, and determine a live detection result corresponding to the user.

[0092] For the live detection device provided by the embodiment of the present invention, by performing dense fixed-point on each frame of the to-be-detected face image in the to-be-detected face video corresponding to the user to determine the dense fixed-point result, and based on each dense fixed-point result, determining an accumulated adjacency matrix corresponding to the to-be-detected face video for characterizing the face change rule, and inputting the accumulated adjacency matrix and a target dense fixed-point result corresponding to the accumulated adjacency matrix into a graph neural network, and then determining a live detection result corresponding to the user. In the present invention, each frame of the to-be-detected face image is not directly used as the input of the graph neural network, but the accumulated adjacency matrix reflecting the face change rule of the positions of facial muscles and facial features determined according to the dense fixed-point result is used as the input of the graph neural network, which improves the robustness of the graph neural network to the face image content and the illumination environment, improves the accuracy of the graph neural network and the protection effect against attack types such as 3D printed masks, and thus improves the accuracy of live detection.

[0093] Optionally, the determination module 620 is specifically configured to:

[0094] Based on each of the dense fixed-point results, determine a weighted adjacency matrix corresponding to each of the to-be-detected face images; the weighted adjacency matrix is used to characterize the position relationship between each face partition in the face area, and each face partition is obtained by dividing the face area according to the positions of facial features;

[0095] Based on a preset association relationship, perform sparse processing on each of the weighted adjacency matrices to obtain a sparse adjacency matrix corresponding to each of the to-be-detected face images; the preset association relationship is used to characterize the change correlation between each face partition in the face area;

[0096] Based on the sparse adjacency matrices corresponding to each face image to be detected in each sliding window of the face video to be detected, determine the cumulative adjacency matrix corresponding to each sliding window; multiple consecutive frames of face images to be detected are included in the sliding window.

[0097] Optionally, the determining module 620 is specifically configured to:

[0098] For multiple consecutive frames of face images to be detected in each sliding window, based on the sparse adjacency matrices corresponding to two adjacent frames of face images to be detected, determine the mean square error corresponding to the two adjacent frames of face images to be detected; determine the sum of all mean square errors as the cumulative adjacency matrix corresponding to the sliding window.

[0099] Optionally, each of the dense fixed-point results includes multiple feature points in the face area and the fixed-point positions corresponding to each feature point.

[0100] Optionally, the determining module 620 is specifically configured to:

[0101] For each of the dense fixed-point results, based on the fixed-point positions corresponding to any two feature points in the face area, determine the distance between the any two feature points; based on all the distances, determine the weighted adjacency matrix corresponding to the face image to be detected.

[0102] Optionally, the determining module 620 is specifically configured to:

[0103] For each face image to be detected, based on the preset association relationship, determine at least one sparse partition corresponding to each face partition; for each face partition, in the weighted adjacency matrix corresponding to the face image to be detected, set the distance between each feature point in the face partition and each feature point in each sparse partition to zero to obtain the sparse adjacency matrix corresponding to the face image to be detected.

[0104] Optionally, the detecting module 630 is specifically configured to:

[0105] Input the cumulative adjacency matrix and the target dense fixed-point result corresponding to the cumulative adjacency matrix into the graph neural network to obtain the face discrimination results corresponding to each sliding window sequentially output by the graph neural network; the target dense fixed-point result is the dense fixed-point result corresponding to the last frame of face image to be detected in the sliding window corresponding to each cumulative adjacency matrix;

[0106] Judge whether each face discrimination result meets a preset condition, and determine the live detection result corresponding to the user based on the judgment result;

[0107] The preset condition includes any one of the following:

[0108] The face discrimination result is true;

[0109] The first preset number of consecutive face discrimination results are all true;

[0110] Among the second preset number of consecutive face discrimination results, the proportion of face discrimination results that are true reaches a preset threshold.

[0111] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 7 shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute a live detection method, and the method includes:

[0112] Obtain a face video to be detected corresponding to the user, and determine the dense fixed-point results of each frame of face image to be detected in the face video to be detected;

[0113] Based on each of the dense fixed-point results, determine an accumulated adjacency matrix corresponding to the face video to be detected; the accumulated adjacency matrix is used to characterize the face change rule in the face video to be detected;

[0114] Input the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into a graph neural network to determine the live detection result corresponding to the user.

[0115] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0116] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the living body detection method provided by each of the above methods, and the method includes:

[0117] Obtain a face video to be detected corresponding to a user, and determine the dense fixed-point results of each frame of face image to be detected in the face video to be detected;

[0118] Based on each of the dense fixed-point results, determine an accumulated adjacency matrix corresponding to the face video to be detected; the accumulated adjacency matrix is used to characterize the face change rule in the face video to be detected;

[0119] Input the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into a graph neural network to determine the living body detection result corresponding to the user.

[0120] In yet another aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the living body detection method provided by each of the above methods, and the method includes:

[0121] Obtain a face video to be detected corresponding to a user, and determine the dense fixed-point results of each frame of face image to be detected in the face video to be detected;

[0122] Based on each of the dense fixed-point results, determine an accumulated adjacency matrix corresponding to the face video to be detected; the accumulated adjacency matrix is used to characterize the face change rule in the face video to be detected;

[0123] Input the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into a graph neural network to determine the living body detection result corresponding to the user.

[0124] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0125] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting a living body, characterized in that: include: Obtain a face video to be detected corresponding to the user, and determine a dense fixed-point result of each frame of the face image to be detected in the face video to be detected; Each of the dense fixed-point results includes a plurality of feature points in the face region and a fixed-point position corresponding to each of the feature points; Based on each of the dense fixed-point results, determining a cumulative adjacency matrix corresponding to the face video to be detected; the cumulative adjacency matrix is ​​used to characterize the face change pattern in the face video to be detected; Inputting the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into a graph neural network to determine a liveness detection result corresponding to the user; The target dense fixed-point result is a dense fixed-point result corresponding to the last frame of the face image to be detected in the sliding window corresponding to the accumulated adjacency matrix; The step of determining the cumulative adjacency matrix corresponding to the face video to be detected based on each of the dense fixed-point results includes: Based on the dense fixed-point results, a weighted adjacency matrix corresponding to each of the face images to be detected is determined; the weighted adjacency matrix is ​​used to characterize the positional relationship corresponding to each face partition in the face region, and each face partition is a division of the face region according to the position of the facial features; Based on a preset association relationship, sparse processing is performed on each of the weighted adjacency matrices to obtain a sparse adjacency matrix corresponding to each of the face images to be detected; the preset association relationship is used to characterize the change correlation between each face partition in the face region; Based on the sparse adjacency matrix corresponding to each face image to be detected in each sliding window in the face video to be detected, the cumulative adjacency matrix corresponding to each sliding window is determined; the sliding window includes multiple consecutive frames of face images to be detected.

2. The method for detecting living body according to claim 1, characterized in that: The step of determining the accumulated adjacency matrix corresponding to each sliding window based on the sparse adjacency matrix corresponding to each face image to be detected in each sliding window in the face video to be detected comprises: For the continuous multiple frames of face images to be detected in each of the sliding windows, based on the sparse adjacency matrices corresponding to the two adjacent frames of face images to be detected, the mean square error corresponding to the two adjacent frames of face images to be detected is determined; and the sum of all mean square errors is determined as the cumulative adjacency matrix corresponding to the sliding window.

3. The method for detecting living body according to claim 1, characterized in that: The step of determining a weighted adjacency matrix corresponding to each of the face images to be detected based on each of the dense fixed-point results includes: For each of the dense fixed-point results, based on the fixed-point positions corresponding to any two feature points in the face area, the distance between the any two feature points is determined; based on all the distances, the weighted adjacency matrix corresponding to the face image to be detected is determined.

4. The method for liveness detection according to claim 3, characterized in that: The step of performing sparse processing on each of the weighted adjacency matrices based on the preset association relationship to obtain a sparse adjacency matrix corresponding to each of the face images to be detected includes: For each of the face images to be detected, based on the preset association relationship, determine at least one sparse partition corresponding to each of the face partitions; for each of the face partitions, in the weighted adjacency matrix corresponding to the face image to be detected, set the distance between each feature point in the face partition and each feature point in the sparse partitions to zero, so as to obtain the sparse adjacency matrix corresponding to the face image to be detected.

5. The method for detecting a living body according to any one of claims 1 to 4, characterized in that: The step of inputting the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into a graph neural network to determine the liveness detection result corresponding to the user includes: Inputting the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into the graph neural network, and obtaining the face recognition results corresponding to each of the sliding windows sequentially output by the graph neural network; Determine whether each of the face recognition results meets a preset condition, and determine a liveness detection result corresponding to the user based on the judgment result; The preset condition includes any one of the following: The face recognition result is true; The first preset number of consecutive face recognition results are all true; Among the second preset number of consecutive face recognition results, the proportion of true face recognition results reaches a preset threshold.

6. A device for implementing the liveness detection method according to any one of claims 1 to 5, characterized in that: include: An acquisition module is used to acquire a face video to be detected corresponding to the user, and determine a dense fixed-point result of each frame of the face image to be detected in the face video to be detected; A determination module, used to determine the cumulative adjacency matrix corresponding to the face video to be detected based on each of the dense fixed-point results; The accumulated adjacency matrix is ​​used to characterize the face change rules in the face video to be detected; The detection module is used to input the accumulated adjacency matrix and the target dense fixed-point result corresponding to the accumulated adjacency matrix into the graph neural network to determine the liveness detection result corresponding to the user.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the living body detection method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the living body detection method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Face detection method and device based on space-time diagram convolutional network

    CN112381064A