Face living body detection method and device, storage medium and equipment
By combining action liveness detection and silent liveness detection in face liveness detection and using weighted multimodal fusion, the problem of insufficient protection against video synthesis and image synthesis attacks in existing technologies is solved, and a more efficient face liveness detection effect is achieved.
Patent Information
- Application Number
- CN202210100938.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-27
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-01-27
AI Technical Summary
Existing face liveness detection methods are not sufficiently resistant to attacks involving video synthesis and motion attacks involving image synthesis, resulting in poor face liveness detection performance.
By prompting the user to perform several specified actions, the system obtains the facial image corresponding to each specified action, extracts features, calculates the liveness detection score, performs weighted summation and decision fusion using weights, and combines action liveness detection and silent liveness detection methods to perform multimodal fusion judgment.
It improves the ability to defend against video recording attacks and synthetic motion attacks, enhances the accuracy and effectiveness of face liveness detection, and reduces the adverse effects of performance differences of different specified actions on the detection results.
Smart Images

Figure CN116563903B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biometrics, and in particular to a method, apparatus, storage medium, and device for detecting liveness of human faces. Background Technology
[0002] In recent years, facial recognition has become increasingly popular due to its convenience and contactless nature, and is widely used in fields such as finance and security. However, precisely because it is easily accessible, it is also easily exploited by others. People can create fake faces by printing photos or recording videos to attack facial recognition systems and impersonate others. Therefore, liveness detection based on facial images is extremely important and is a prerequisite for ensuring the success of facial recognition.
[0003] Action liveness detection is a commonly used method for detecting facial liveness. It determines whether a person is real or an imposter by judging whether the eyes, mouth, or head follow specific instructions to perform actions. Action liveness detection performs liveness detection by judging whether specified actions in a video image are correct. These actions are primarily determined by the relative changes of key points such as the eyes or mouth. However, action liveness detection is not effective against video synthesis attacks or image synthesis attacks. Figure 1 The image synthesis actions shown can be used to attack actions such as opening the mouth, blinking, shaking the head left and right, and nodding up and down. This can successfully bypass action liveness detection, resulting in poor face liveness detection performance. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a method, apparatus, storage medium, and device for face liveness detection, thereby improving the effectiveness of face liveness detection.
[0005] The technical solution provided by this invention is as follows:
[0006] In a first aspect, the present invention provides a method for detecting human face liveness, the method comprising:
[0007] The system prompts the user to perform several specified actions and performs a liveness detection.
[0008] After the liveness detection of the action is passed, the face image corresponding to each specified action is acquired;
[0009] Feature extraction is performed on the face image corresponding to each specified action to obtain the image features of each specified action, and the liveness detection score of each specified action is calculated based on the image features of each specified action.
[0010] The liveness detection scores for each specified action are weighted and summed according to the weights set for each specified action to obtain the fusion score;
[0011] The weight of each specified action is determined based on the distribution of the liveness detection scores in the training sample set for that specified action and the liveness rejection rate in the training sample set for that specified action.
[0012] The liveness detection result of each specified action is determined based on the liveness detection score of each specified action and the set liveness detection threshold of each specified action. The liveness detection result of each specified action is then weighted and voted on according to the set weight of each specified action to obtain the decision fusion result.
[0013] The user is determined to be either a live or a fake based on the fusion score and / or the decision fusion result.
[0014] Furthermore, the weight of each specified action is determined using the following method:
[0015] Feature extraction is performed on the training sample set for each specified action to obtain the image feature set for each specified action, and the liveness detection score set for each specified action is calculated based on the image feature set for each specified action.
[0016] The training sample set includes live samples and spur samples, and the live detection score set includes live scores and spur scores.
[0017] The liveness detection score set for each specified action is normalized based on the non-overlapping regions of the distribution of liveness scores and the distribution of spur scores for each specified action.
[0018] The weight of each specified action is calculated based on the normalized liveness detection score set for each specified action and the liveness rejection rate of the training sample set for each specified action.
[0019] The normalization of the liveness detection score set for each specified action based on the non-overlapping regions of the distribution of liveness scores and the distribution of spurious scores for each specified action includes:
[0020] The non-overlapping regions of the distribution of live body scores and the distribution of prosthetic scores for each specified action are calculated using the following formula, and are used as the normalized dimensions for each specified action.
[0021] q(k)=μ a (k)-min a (k)+max r (k)-μ r (k)
[0022] Where q(k) is the normalized dimension of the k-th specified action, k = 1, 2, ..., m, m is the total number of specified actions, and max r(k) represents the maximum liveness score in the liveness detection score set for the k-th specified action, μ r (k) represents the average liveness score in the liveness detection score set for the k-th specified action, min a (k) represents the minimum spur score in the liveness detection score set for the k-th specified action, μ a (k) is the average value of the prosthesis scores in the liveness detection score set for the k-th specified action;
[0023] The liveness detection score set for each specified action is normalized according to the normalized dimensions of each specified action;
[0024]
[0025] in, and Let i be the liveness detection score in the liveness detection score set for the k-th specified action before and after normalization, respectively, where i = 1, 2, 3, ..., N. k N k Let j = 0 represent the total number of liveness detection scores in the liveness detection score set for the k-th specified action. The score represents the prosthesis score, and j=1 represents the liveness detection score. The number of living individuals;
[0026] The step of calculating the weight of each specified action based on the normalized liveness detection score set for each specified action and the liveness rejection rate of the training sample set for each specified action includes:
[0027] The weight of each specified action is calculated using the following formula;
[0028]
[0029] Where w(k) is the weight of the k-th specified action;
[0030]
[0031] FRR k The liveness rejection rate is the training sample set for the k-th specified action.
[0032] Furthermore, the fusion score is calculated using the following formula;
[0033]
[0034] Where s is the fusion score, and s(k) is the liveness detection score for the k-th specified action;
[0035] The decision fusion result is calculated using the following formula;
[0036]
[0037] Wherein, δ is the decision fusion result, and δ(k) is the liveness detection result of k specified actions;
[0038]
[0039] T(k) is the liveness detection threshold for the k-th specified action.
[0040] Furthermore, determining whether the user is a live or spoofed based on the fusion score and / or the decision fusion result includes:
[0041] If the fusion score is greater than a set first threshold, the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0042] or,
[0043] If the decision fusion result is greater than the set second threshold, the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0044] or,
[0045] If the fusion score is greater than a set first threshold and the decision fusion result is greater than a set second threshold, then the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0046] Furthermore, the specified action includes one or more of the following: blinking, opening the mouth, turning the head left and right, nodding up and down, and remaining still in the front.
[0047] A specified action acquires a face image, performs feature extraction on the face image, and obtains the image features of the specified action; or, a specified action acquires multiple face images, performs feature extraction on the multiple face images, and fuses the extracted features of the multiple face images to obtain the image features of the specified action.
[0048] Furthermore, the multiple face images include a first face image and a second face image. Feature extraction is performed on the first face image and the second face image to obtain first face features and second face features. Feature fusion is then performed on the first face features and the second face features to obtain the image features:
[0049] Principal component analysis was performed on the first facial feature and the second facial feature using the following formula;
[0050]
[0051] Where l = 0 and 1 represent the numbers of the first facial feature and the second facial feature, respectively, and x l Representing the first facial feature and the second facial feature, z l This represents the feature obtained after performing principal component analysis on the first and second facial features, μ. l W represents the mean of the first facial feature and the mean of the second facial feature. l Represents the principal component analysis matrix of the first facial feature and the principal component analysis matrix of the second facial feature; μ l and W l It is obtained by training on the first face feature training set and the second face feature training set;
[0052] The image features are obtained by fusing the features obtained from principal component analysis of the first and second facial features using the following formula;
[0053]
[0054] Where Z represents the image feature, P o and P 1 Let P be the projection matrix. o and P 1 It is obtained by training using the first facial feature training set and the second facial feature training set.
[0055] Furthermore, the μ is obtained through training using the following method. l W l P o and P 1 :
[0056] The mean value of the first facial feature and the mean value of the second facial feature are calculated using the following formulas;
[0057]
[0058] in, The i-th first face feature sample / second face feature sample in the first face feature training set / second face feature training set, i = 1, 2, 3, ..., n, where n is the total number of first face feature samples / second face feature samples;
[0059] The covariance matrix of the first face feature sample and the second face feature sample is calculated using the following formula;
[0060]
[0061] Among them, S lLet be the covariance matrix of the first face feature sample and the second face feature sample;
[0062] The covariance matrix of the first face feature sample and the second face feature sample is decomposed into eigenvalues using the following formula;
[0063]
[0064] in, and S l Eigenvalues and eigenvectors, t = 0, 1, 2, ..., d, where d is the eigenvalue of S. l The number of rows and columns;
[0065] Based on the percentage of the main component The principal component analysis matrix W is constructed by selecting the top r eigenvectors, which account for 95% of the total. l ;
[0066] in,
[0067] Principal component analysis was performed on the first and second facial feature samples using the following formula;
[0068]
[0069] in, This represents the features obtained after performing principal component analysis on the first and second face feature samples;
[0070] The projection matrix P is obtained by optimizing the following formula. o and P 1 ;
[0071]
[0072] in, and Z 0 and Z 1 The within-class variance matrix, For Z 0 and Z 1 The inter-class variance matrix.
[0073] Secondly, the present invention provides a face liveness detection device, the device comprising:
[0074] The motion liveness detection module is used to prompt the user to perform several specified actions and to perform motion liveness detection.
[0075] The data acquisition module is used to acquire the face image corresponding to each specified action after the action liveness detection is passed;
[0076] The liveness detection score calculation module is used to extract features from the face image corresponding to each specified action, obtain the image features of each specified action, and calculate the liveness detection score of each specified action based on the image features of each specified action.
[0077] The fusion score calculation module is used to perform a weighted summation of the liveness detection scores of each specified action according to the set weight of each specified action, so as to obtain the fusion score;
[0078] The weight of each specified action is determined based on the distribution of the liveness detection scores in the training sample set for that specified action and the liveness rejection rate in the training sample set for that specified action.
[0079] The decision fusion result calculation module is used to determine the liveness detection result of each specified action based on the liveness detection score of each specified action and the set liveness detection threshold of each specified action, and to perform weighted voting on the liveness detection result of each specified action according to the set weight of each specified action to obtain the decision fusion result.
[0080] The liveness detection module is used to determine whether the user is a live or spoofed based on the fusion score and / or the decision fusion result.
[0081] Furthermore, the weight of each specified action is determined through the following module:
[0082] The training data preparation module is used to extract features from the training sample set for each specified action, obtain the image feature set for each specified action, and calculate the liveness detection score set for each specified action based on the image feature set for each specified action.
[0083] The training sample set includes live samples and spur samples, and the live detection score set includes live scores and spur scores.
[0084] The normalization module is used to normalize the liveness detection score set for each specified action based on the non-overlapping regions of the distribution of liveness scores and the distribution of spur scores for each specified action.
[0085] The weight calculation module is used to calculate the weight of each specified action based on the normalized liveness detection score set of each specified action and the liveness rejection rate of the training sample set of each specified action.
[0086] Furthermore, the normalization module includes:
[0087] The normalized dimension calculation unit is used to calculate the non-overlapping region of the distribution of the live body fraction and the distribution of the prosthesis fraction for each specified action using the following formula, as the normalized dimension for each specified action.
[0088] q(k)=μ a (k)-min a (k)+max r (k)-μ r (k)
[0089] Where q(k) is the normalized dimension of the k-th specified action, k = 1, 2, ..., m, m is the total number of specified actions, and max r (k) represents the maximum liveness score in the liveness detection score set for the k-th specified action, μ r (k) represents the average liveness score in the liveness detection score set for the k-th specified action, min a (k) represents the minimum spur score in the liveness detection score set for the k-th specified action, μ a (k) is the average value of the prosthesis scores in the liveness detection score set for the k-th specified action;
[0090] The normalization unit is used to normalize the liveness detection score set for each specified action according to the normalized dimension of each specified action.
[0091]
[0092] in, and Let i be the liveness detection score in the liveness detection score set for the k-th specified action before and after normalization, respectively, where i = 1, 2, 3, ..., N. k N k Let j = 0 represent the total number of liveness detection scores in the liveness detection score set for the k-th specified action. The score represents the prosthesis score, and j=1 represents the liveness detection score. The number of living individuals;
[0093] The weight calculation module is used for:
[0094] The weight of each specified action is calculated using the following formula;
[0095]
[0096] Where w(k) is the weight of the k-th specified action;
[0097]
[0098] FRR kThe liveness rejection rate is the training sample set for the k-th specified action.
[0099] Furthermore, the fusion score is calculated using the following formula;
[0100]
[0101] Where s is the fusion score, and s(k) is the liveness detection score for the k-th specified action;
[0102] The decision fusion result is calculated using the following formula;
[0103]
[0104] Wherein, δ is the decision fusion result, and δ(k) is the liveness detection result of k specified actions;
[0105]
[0106] T(k) is the liveness detection threshold for the k-th specified action.
[0107] Furthermore, the liveness detection module is used for:
[0108] If the fusion score is greater than a set first threshold, the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0109] or,
[0110] If the decision fusion result is greater than the set second threshold, the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0111] or,
[0112] If the fusion score is greater than a set first threshold and the decision fusion result is greater than a set second threshold, then the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0113] Furthermore, the specified action includes one or more of the following: blinking, opening the mouth, turning the head left and right, nodding up and down, and remaining still in the front.
[0114] A specified action acquires a face image, performs feature extraction on the face image, and obtains the image features of the specified action; or, a specified action acquires multiple face images, performs feature extraction on the multiple face images, and fuses the extracted features of the multiple face images to obtain the image features of the specified action.
[0115] Furthermore, the multiple face images include a first face image and a second face image. Feature extraction is performed on the first face image and the second face image to obtain first face features and second face features. Feature fusion is then performed on the first face features and the second face features through the following module to obtain the image features:
[0116] The first principal component analysis module is used to perform principal component analysis on the first facial feature and the second facial feature using the following formula;
[0117]
[0118] Where l = 0 and 1 represent the numbers of the first facial feature and the second facial feature, respectively, and x l Representing the first facial feature and the second facial feature, z l This represents the feature obtained after performing principal component analysis on the first and second facial features, μ. l W represents the mean of the first facial feature and the mean of the second facial feature. l Represents the principal component analysis matrix of the first facial feature and the principal component analysis matrix of the second facial feature; μ l and W l It is obtained by training on the first face feature training set and the second face feature training set;
[0119] The feature fusion module is used to fuse the features obtained by principal component analysis of the first face feature and the second face feature using the following formula to obtain the image features;
[0120]
[0121] Where Z represents the image feature, P o and P 1 Let P be the projection matrix. o and P 1 It is obtained by training using the first facial feature training set and the second facial feature training set.
[0122] Furthermore, the μ is obtained through training using the following modules. l W l P o and P 1 :
[0123] The mean calculation module is used to calculate the mean of the first facial feature and the mean of the second facial feature using the following formula;
[0124]
[0125] in, The i-th first face feature sample / second face feature sample in the first face feature training set / second face feature training set, i = 1, 2, 3, ..., n, where n is the total number of first face feature samples / second face feature samples;
[0126] The covariance matrix calculation module is used to calculate the covariance matrix of the first face feature sample and the second face feature sample using the following formula;
[0127]
[0128] Among them, S l Let be the covariance matrix of the first face feature sample and the second face feature sample;
[0129] The eigenvalue decomposition module is used to perform eigenvalue decomposition on the covariance matrix of the first face feature sample and the second face feature sample using the following formula;
[0130]
[0131] in, and S l Eigenvalues and eigenvectors, t = 0, 1, 2, ..., d, where d is the eigenvalue of S. l The number of rows and columns;
[0132] The principal component analysis matrix acquisition module is used to obtain the matrix based on the percentage of principal components. The principal component analysis matrix W is constructed by selecting the top r eigenvectors, which account for 95% of the total. l ;
[0133] in,
[0134] The second principal component analysis module is used to perform principal component analysis on the first face feature sample and the second face feature sample using the following formula;
[0135]
[0136] in, This represents the features obtained after performing principal component analysis on the first and second face feature samples;
[0137] The projection matrix calculation module is used to optimize and obtain the projection matrix P using the following formula. o and P 1 ;
[0138]
[0139] in, and Z 0 and Z 1 The within-class variance matrix, Let Z0 and Z1 be the inter-class variance matrices.
[0140] Thirdly, the present invention provides a computer-readable storage medium for face liveness detection, including a memory for storing processor-executable instructions, which, when executed by the processor, implement the steps of the face liveness detection method described in the first aspect.
[0141] Fourthly, the present invention provides an apparatus for face liveness detection, comprising at least one processor and a memory storing computer-executable instructions, wherein the processor executes the instructions to implement the steps of the face liveness detection method described in the first aspect.
[0142] The present invention has the following beneficial effects:
[0143] This invention combines action-based liveness detection and silent liveness detection, providing effective protection against video recording attacks and synthetic action attacks. Furthermore, this invention performs multimodal silent liveness detection on the facial image corresponding to each specified action, fusing the silent liveness detection results from different actions to provide a final judgment of whether the target is a live or fake. This makes silent liveness detection more accurate and, by performing liveness detection on the facial image within each specified action, prevents synthetic action attacks.
[0144] This invention fuses liveness detection scores for multiple specified actions using weights and performs weighted voting on the liveness detection results for these actions. This achieves both liveness detection score fusion and multimodal decision fusion. Furthermore, the fusion weights are related to the distribution of liveness detection scores for each specified action and the liveness rejection rate. By utilizing the distribution of liveness detection scores and the liveness rejection rate to perform multimodal fusion of liveness detection scores and decisions, and by unifying the dimensions and distribution of liveness detection scores through weights, the adverse effects of differences in liveness detection performance for different specified actions on the fused liveness detection are reduced. This makes the face liveness detection method more effective and feasible, achieving better face liveness detection results. Attached Figure Description
[0145] Figure 1 Example diagram of the image synthesis process;
[0146] Figure 2 This is a flowchart of the face liveness detection method of the present invention;
[0147] Figure 3A schematic diagram showing the distribution of prosthetic and live body scores for the k-th specified action;
[0148] Figure 4 This is a schematic diagram of the face liveness detection device of the present invention. Detailed Implementation
[0149] To make the technical problems, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. The components of the embodiments of this invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.
[0150] Example 1:
[0151] This invention provides a method for face liveness detection, such as... Figure 2 As shown, the method includes:
[0152] S100: Prompts the user to perform several specified actions and performs liveness detection.
[0153] This step is used to perform a liveness detection. The system prompts the user to perform several specified actions and determines whether the specified actions performed by the user are correct. If they are correct, the liveness detection passes; otherwise, the liveness detection fails.
[0154] The specified actions may include one or more of the following: blinking, opening the mouth, turning the head left and right, nodding up and down, and remaining still in the front. The specified actions may be indicated in a set order or in a random order.
[0155] S200: After the liveness detection of the action is passed, obtain the face image corresponding to each specified action.
[0156] Action liveness detection is not effective against video recording attacks or synthetic motion attacks. Many existing video motion synthesis software programs, such as AE (Adobe After Effects), can synthesize actions such as opening the mouth and closing the eyes in images. The actions synthesized by the software can break the action liveness detection algorithm.
[0157] Figure 1 The synthesized blinking, mouth opening, left and right head turning, and up and down head nodding actions are given. Figure 1In the diagram, 1A and 1B represent blinking, 1C and 1D represent opening the mouth, 1E and 1F represent turning the head left and right, and 1G and 1H represent nodding up and down. Figure 1 The synthesized actions can break the action liveness detection algorithm.
[0158] However, although the actions synthesized by the software satisfy the local muscle movements of the face, they modify the mouth or eye images, which will cause the texture of the face image to be distorted. Therefore, a liveness detection method that relies on texture features can be used to further detect the face image when performing the specified action.
[0159] The liveness detection method that relies on texture features is also known as the silent liveness detection method. It does not require active user cooperation. Its execution process is as follows: first, the face image to be detected is acquired and the image features are extracted; then, the liveness detection score is calculated based on the image features; and finally, the liveness detection score is used to determine whether the person is alive.
[0160] This invention combines motion liveness detection and silent liveness detection, which can effectively prevent attacks from video recording or synthetic motion.
[0161] Silent liveness detection methods require acquiring a face image first. Existing technologies only acquire one qualified face image for liveness detection, typically a frontal face image. Unlike existing technologies, this invention does not acquire only one face image, but acquires a corresponding face image representing each specified action.
[0162] A specified action can acquire one or more corresponding facial images. For example, blinking acquires a facial image with closed eyes, opening mouth acquires a facial image with open mouth, turning head left and right acquires a facial image with a tilt greater than 30 degrees to the left and right, nodding head up and down acquires a facial image with an upward and downward tilt angle greater than 15 degrees, and remaining still in front acquires a frontal facial image of acceptable quality. A frontal facial image of acceptable quality requires that the eyes are open, the mouth is closed, and the head posture is upright.
[0163] This invention acquires the facial image corresponding to each specified action for subsequent multimodal silent liveness detection, with each specified action constituting one modality. Through multimodal silent liveness detection, on the one hand, silent liveness detection becomes more accurate, and on the other hand, liveness detection is performed on the facial image within each specified action, preventing attacks using synthetic actions.
[0164] S300: Extract features from the face image corresponding to each specified action to obtain the image features of each specified action, and calculate the liveness detection score of each specified action based on the image features of each specified action.
[0165] This invention does not limit the method of obtaining the liveness detection score. For example, silent liveness detection methods such as CNN can be used to extract features from face images of multiple specified actions to obtain image features, and then a discriminator such as SVM can be used to classify the image features to obtain the liveness detection score.
[0166] The liveness detection score for multiple specified actions can be denoted as s(k), where s(k) is the liveness detection score for the k-th specified action, k = 1, 2, ..., m, and m is the total number of specified actions.
[0167] S400: The liveness detection scores of each specified action are weighted and summed according to the weight of each specified action to obtain the fusion score.
[0168] For example, the fusion score can be calculated using the following formula;
[0169]
[0170] Where s is the fusion score, and w(k) is the weight of the k-th specified action.
[0171] Since the dimensions and distribution of liveness detection scores obtained from different specified actions are not uniform, and the liveness detection performance of different specified actions is also different, directly fusing the liveness detection scores obtained from different specified actions will reduce the performance of liveness detection.
[0172] This invention uses weighted summation of liveness detection scores for multiple specified actions to obtain a fusion score. The weight of each specified action is determined based on the distribution of liveness detection scores in the training sample set for that action and the liveness rejection rate in that training sample set. The distribution of liveness detection scores in the training sample set for the specified action reflects the performance of the liveness detection algorithm, and this distribution also allows for the standardization of the liveness detection scores.
[0173] The liveness rejection rate (LPR) is a metric for evaluating liveness detection, representing the proportion of live samples that are incorrectly identified as fakes. It's an assessment of the number of false positives for live samples and also reflects the performance of the liveness detection algorithm. Incorporating the LPR into the determination of weights can guide weight generation, directing optimization towards reducing the LPR.
[0174] Therefore, the weight of the specified action, determined by the distribution of the liveness detection score in the training sample set of each specified action and the liveness rejection rate in the training sample set of the specified action, can unify the dimensions and distribution of the liveness detection score, and also reduce the adverse effects of the difference in liveness detection performance of different specified actions on the fused liveness detection, thereby obtaining better liveness detection results.
[0175] S500: Determine the liveness detection result of each specified action based on the liveness detection score of each specified action and the set liveness detection threshold of each specified action, and perform weighted voting on the liveness detection result of each specified action according to the set weight of each specified action to obtain the decision fusion result.
[0176] The decision fusion result is the result of a weighted vote on the liveness detection results of multiple specified actions. It reflects the role of the liveness detection result of each specified action in the final liveness detection, as well as the proportion of the liveness detection result of each specified action in the final liveness detection.
[0177] For example, the decision fusion result can be calculated using the following formula;
[0178]
[0179] Wherein, δ is the decision fusion result, and δ(k) is the liveness detection result of k specified actions;
[0180]
[0181] T(k) is the set liveness detection threshold for the k-th specified action.
[0182] S600: Determine whether the user is a live or spoofed based on the fusion score and / or the decision fusion result.
[0183] After processing multiple facial images for specified actions as described above, two judgment factors can be obtained: the fusion score and the decision fusion result. Both the fusion score and the decision fusion result are important indicators in liveness detection, and they can be judged separately or in combination.
[0184] For example, if the fusion score s is greater than a set first threshold T s If the user is alive, then the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0185] or,
[0186] If the decision fusion result δ is greater than the set second threshold δ s If the user is alive, then the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0187] or,
[0188] If the fusion score s is greater than the set first threshold T s Furthermore, the decision fusion result δ is greater than the set second threshold δ. s If the user is alive, then the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0189] This invention combines action-based liveness detection and silent liveness detection, providing effective protection against video recording attacks and synthetic action attacks. Furthermore, this invention performs multimodal silent liveness detection on the facial image corresponding to each specified action, fusing the silent liveness detection results from different actions to provide a final judgment of whether the target is a live or fake. This makes silent liveness detection more accurate and, by performing liveness detection on the facial image within each specified action, prevents synthetic action attacks.
[0190] This invention fuses liveness detection scores for multiple specified actions using weights and performs weighted voting on the liveness detection results for these actions. This achieves both liveness detection score fusion and multimodal decision fusion. Furthermore, the fusion weights are related to the distribution of liveness detection scores for each specified action and the liveness rejection rate. By utilizing the distribution of liveness detection scores and the liveness rejection rate to perform multimodal fusion of liveness detection scores and decisions, and by unifying the dimensions and distribution of liveness detection scores through weights, the adverse effects of differences in liveness detection performance for different specified actions on the fused liveness detection are reduced. This makes the face liveness detection method more effective and feasible, achieving better face liveness detection results.
[0191] The present invention includes a training process and an inference process. The ultimate goal of the training process is to obtain the weight w(k) of each specified action fusion through the labeled training sample set. The inference process is to determine whether the face images of the multiple specified actions belong to the live or fake category by using the weights obtained through training.
[0192] The inference process is as described in S100 to S600 above. The method for determining the weight of each specified action during the training process is as follows:
[0193] S10: Extract features from the training sample set for each specified action to obtain the image feature set for each specified action, and calculate the liveness detection score set for each specified action based on the image feature set for each specified action.
[0194] The training sample set includes live samples and spoofed samples, and the liveness detection score set includes liveness scores and spoofed scores. The detection score obtained by silent liveness detection of live samples is the liveness score, and the detection score obtained by silent liveness detection of spoofed samples is the spoofed score.
[0195] For example, with This represents the i-th liveness detection score in the liveness detection score set for the k-th specified action, where i = 1, 2, 3, ..., N. k N kLet j = 0, 1, where j = 0 represents the total number of liveness detection scores in the liveness detection score set for the k-th specified action. The score represents the prosthesis score, and j=1 represents the liveness detection score. This represents the number of living individuals.
[0196] The liveness score obtained by performing silent liveness detection on the live sample using the k-th specified action is denoted as r(k);
[0197]
[0198] The prosthesis score obtained by performing silent liveness detection on the prosthesis sample with the k-th specified action is denoted as a(k);
[0199]
[0200] The distributions of the live and prosthetic scores generally follow a Gaussian distribution. The distributions of the prosthetic and live scores for the k-th specified action are as follows: Figure 3 As shown.
[0201] in,
[0202] max r (k) represents the maximum liveness score in the liveness detection score set for the k-th specified action;
[0203] min r (k) represents the minimum liveness score in the liveness detection score set for the k-th specified action;
[0204] μ r (k) is the average of the liveness scores in the liveness detection score set for the k-th specified action;
[0205] max a (k) represents the maximum value of the prosthesis score in the liveness detection score set for the k-th specified action;
[0206] min a (k) represents the minimum value of the prosthesis score in the liveness detection score set for the k-th specified action;
[0207] μ a (k) is the average value of the prosthesis scores in the liveness detection score set for the k-th specified action.
[0208] Clearly, the overlap between the liveness score and the spurious score distributions can reflect the performance of the liveness detection algorithm for a given action. The larger the overlap, the less distinct the difference between live and spurious samples, and the worse the performance. Conversely, the smaller the overlap, the higher the distinction between live and spurious samples, and the better the performance. Therefore, it is necessary to reduce the overlap of scores and increase the difference between the two through multimodal fusion.
[0209] S20: Normalize the liveness detection score set for each specified action based on the non-overlapping regions of the distribution of liveness scores for each specified action and the distribution of spurious scores for each specified action.
[0210] As mentioned above, the larger the overlap area (i.e., the smaller the non-overlapping area) between the distributions of liveness scores and spurious scores, the less distinct the distinction between live and spurious samples; conversely, the smaller the overlap area, the higher the distinction between live and spurious samples. Therefore, this invention uses the non-overlapping area to normalize the liveness detection score set for each specified action, which can unify the dimensions and distribution of the liveness detection scores and increase the difference in the distribution of liveness and spurious scores, thus achieving better liveness detection results.
[0211] The specific implementation method for this step can be as follows:
[0212] S21: Calculate the non-overlapping regions of the distribution of the live fraction and the distribution of the prosthesis fraction for each specified action using the following formula, as the normalized dimension for each specified action.
[0213] q(k)=μ a (k)-min a (k)+max r (k)-μ r (k)
[0214] Where q(k) is the normalized dimension of the k-th specified action.
[0215] The fusion of liveness detection scores requires normalization of the liveness detection scores judged under different specified actions. The normalization operation needs to use the parameters of the non-overlapping region in the distribution of liveness detection scores in the training sample set as the normalization dimension to operate on the liveness detection scores. The calculation of the normalization dimension is as shown in the above formula.
[0216] The larger q(k) is, the greater the difference between the distribution of the liveness score and the spoof score for the k-th specified action, the higher the discrimination and the better the effect. This normalization dimension can reduce the overlapping area of the distribution when normalizing the liveness detection score, thereby obtaining a better liveness detection effect.
[0217] S22: Normalize the liveness detection score set for each specified action according to the normalized dimension of each specified action;
[0218]
[0219] in, It is the i-th liveness detection score in the normalized liveness detection score set for the k-th specified action.
[0220] S30: Calculate the weight of each specified action based on the normalized liveness detection score set for each specified action and the liveness rejection rate of the training sample set for each specified action.
[0221] In one example, the specific formula for calculating the weight of each specified action is as follows:
[0222]
[0223] Where w(k) is the weight of the k-th specified action;
[0224]
[0225] FRR k The liveness rejection rate is the training sample set for the k-th specified action.
[0226] Two metrics for evaluating liveness detection performance are the liveness rejection rate (FRR) and the fake pass rate (FAR). FRR represents the proportion of live samples that are incorrectly identified as fakes, while FAR represents the proportion of fake samples that are incorrectly identified as live. In liveness detection performance evaluation, both FRR and FAR are desirable to be as low as possible. In the face liveness detection method of this invention, it is necessary to consider not only the influence of the dimensions of the liveness detection components but also the potential for differences in liveness detection performance under different specified actions to degrade the performance of the fused algorithm.
[0227] Therefore, this invention not only introduces the normalized dimension representing the distribution of liveness detection scores into the weight calculation, but also introduces FRR into the weight calculation, and calculates the weight w(k) of each specified action through the above formula.
[0228] As described above, this invention uses the distribution of liveness detection scores from the training sample set to determine weights. Using this distribution increases the difference between liveness and spurious scores, improving liveness detection classification performance. Furthermore, introducing a false rejection rate when determining weights guides weight generation, directing optimization towards a goal of reducing the false rejection rate.
[0229] As mentioned above, a specified action can acquire one or more corresponding face images. When a specified action acquires a face image, silent liveness detection methods such as CNNs are used to extract features from the face image to obtain the image features of the specified action.
[0230] When a specified action acquires multiple face images, a silent liveness detection method such as CNN is used to extract features from the multiple face images, and the extracted features from the multiple face images are fused to obtain the image features of the specified action.
[0231] When the multiple face images include two images, they are a first face image and a second face image. For example, a face image obtained by turning the head left and right by more than 30 degrees is obtained. Silent liveness detection methods such as CNN are used to extract features from the first face image and the second face image to obtain first face features and second face features. The first face features and the second face features are then fused using the following method to obtain the image features:
[0232] S1: Principal component analysis is performed on the first facial feature and the second facial feature using the following formula.
[0233]
[0234] Where l = 0 and 1 represent the numbers of the first facial feature and the second facial feature, respectively, and x l Representing the first facial feature and the second facial feature, z l This represents the feature obtained after performing principal component analysis on the first and second facial features, μ. l W represents the mean of the first facial feature and the mean of the second facial feature. l Represents the principal component analysis matrix of the first facial feature and the principal component analysis matrix of the second facial feature; μ l and W l It is obtained by training on the first face feature training set and the second face feature training set.
[0235] S2: The image features are obtained by fusing the features obtained from the principal component analysis of the first face features and the second face features using the following formula.
[0236]
[0237] Where Z represents the image feature, P o and P 1 Let P be the projection matrix. o and P 1 It is obtained by training using the first facial feature training set and the second facial feature training set.
[0238] This invention requires the simultaneous use of first and second facial features, resulting in high data dimensionality and introducing additional computational and storage space requirements. Therefore, principal component analysis is used to reduce the dimensionality of the first and second facial features, extracting more significant features from the original features, reducing the amount of data required for classification, and increasing the classification capability of the features.
[0239] Principal component analysis (PCA) projects high-dimensional data onto a low-dimensional space through linear mapping, maximizing the variance of the projected data and thus preserving the characteristics of the original data.
[0240] After fusing the first and second facial features into image features, and combining them with image features of other specified actions, subsequent multimodal silent liveness detection is performed.
[0241] The present invention can obtain the μ by training using the following method. l W l P o and P 1 :
[0242] S11: The mean value of the first facial feature and the mean value of the second facial feature are calculated using the following formula.
[0243]
[0244] in, Let i be the i-th first face feature sample / second face feature sample in the first face feature training set / second face feature training set, where i = 1, 2, 3, ..., n, and n is the total number of the first face feature samples / second face feature samples.
[0245] S12: Calculate the covariance matrix of the first face feature sample and the second face feature sample using the following formula.
[0246]
[0247] Among them, S l Let be the covariance matrix of the first face feature sample and the second face feature sample.
[0248] S13: Perform eigenvalue decomposition on the covariance matrix of the first face feature sample and the second face feature sample using the following formula.
[0249]
[0250] in, Know S l Eigenvalues and eigenvectors, t = 0, 1, 2, ..., d, where d is the eigenvalue of S. l The number of rows and columns.
[0251] S14: Based on the percentage of the principal component The principal component analysis matrix W is constructed by selecting the top r eigenvectors, which account for 95% of the total. l .
[0252] in,
[0253] S15: Perform principal component analysis on the first facial feature sample and the second facial feature sample using the following formula.
[0254]
[0255] in, This represents the features obtained after performing principal component analysis on the first and second face feature samples.
[0256] S16: The projection matrix P is obtained by optimizing the following formula. o and P 1 .
[0257]
[0258] in, and Z 0 and Z 1 The within-class variance matrix, Let Z0 and Z1 be the inter-class variance matrices.
[0259] Example 2:
[0260] This invention provides a face liveness detection device, such as... Figure 4 As shown, the device includes:
[0261] The motion liveness detection module 1 is used to prompt the user to perform several specified actions and to perform motion liveness detection.
[0262] The data acquisition module 2 is used to acquire the face image corresponding to each specified action after the action liveness detection is passed.
[0263] The liveness detection score calculation module 3 is used to extract features from the face image corresponding to each specified action, obtain the image features of each specified action, and calculate the liveness detection score of each specified action based on the image features of each specified action.
[0264] The fusion score calculation module 4 is used to perform a weighted summation of the liveness detection scores of each specified action according to the weight of each specified action, so as to obtain the fusion score;
[0265] The weight of each specified action is determined based on the distribution of the liveness detection scores in the training sample set for that specified action and the liveness rejection rate in the training sample set for that specified action.
[0266] The decision fusion result calculation module 5 is used to determine the liveness detection result of each specified action based on the liveness detection score of each specified action and the set liveness detection threshold of each specified action, and to perform weighted voting on the liveness detection result of each specified action according to the set weight of each specified action to obtain the decision fusion result.
[0267] The liveness detection module 6 is used to determine whether the user is a live person or a fake person based on the fusion score and / or the decision fusion result.
[0268] This invention combines action-based liveness detection and silent liveness detection, providing effective protection against video recording attacks and synthetic action attacks. Furthermore, this invention performs multimodal silent liveness detection on the facial image corresponding to each specified action, fusing the silent liveness detection results from different actions to provide a final judgment of whether the target is a live or fake. This makes silent liveness detection more accurate and, by performing liveness detection on the facial image within each specified action, prevents synthetic action attacks.
[0269] This invention fuses liveness detection scores for multiple specified actions using weights and performs weighted voting on the liveness detection results for these actions. This achieves both liveness detection score fusion and multimodal decision fusion. Furthermore, the fusion weights are related to the distribution of liveness detection scores for each specified action and the liveness rejection rate. By utilizing the distribution of liveness detection scores and the liveness rejection rate to perform multimodal fusion of liveness detection scores and decisions, and by unifying the dimensions and distribution of liveness detection scores through weights, the adverse effects of differences in liveness detection performance for different specified actions on the fused liveness detection are reduced. This makes the face liveness detection method more effective and feasible, achieving better face liveness detection results.
[0270] As an improvement to this embodiment of the invention, the weight of each specified action is determined by the following module:
[0271] The training data preparation module is used to extract features from the training sample set for each specified action, obtain the image feature set for each specified action, and calculate the liveness detection score set for each specified action based on the image feature set for each specified action.
[0272] The training sample set includes live samples and spur samples, and the live detection score set includes live scores and spur scores.
[0273] The normalization module is used to normalize the liveness detection score set for each specified action based on the non-overlapping regions of the distribution of liveness scores and the distribution of spur scores for each specified action.
[0274] The weight calculation module is used to calculate the weight of each specified action based on the normalized liveness detection score set of each specified action and the liveness rejection rate of the training sample set of each specified action.
[0275] Furthermore, the normalization module includes:
[0276] The normalized dimension calculation unit is used to calculate the non-overlapping region of the distribution of the live body fraction and the distribution of the prosthesis fraction for each specified action using the following formula, as the normalized dimension for each specified action.
[0277] q(k)=μ a (k)-min a (k)+max r (k)-μ r (k)
[0278] Where q(k) is the normalized dimension of the k-th specified action, k = 1, 2, ..., m, m is the total number of specified actions, and max r (k) represents the maximum liveness score in the liveness detection score set for the k-th specified action, μ r (k) represents the average liveness score in the liveness detection score set for the k-th specified action, min a (k) represents the minimum spur score in the liveness detection score set for the k-th specified action, μ a (k) is the average value of the prosthesis scores in the liveness detection score set for the k-th specified action.
[0279] The normalization unit is used to normalize the liveness detection score set for each specified action according to the normalized dimension of each specified action.
[0280]
[0281] in, and Let i be the liveness detection score in the liveness detection score set for the k-th specified action before and after normalization, respectively, where i = 1, 2, 3, ..., N. k N k Let j = 0 represent the total number of liveness detection scores in the liveness detection score set for the k-th specified action. The score represents the prosthesis score, and j=1 represents the liveness detection score. This represents the number of living individuals.
[0282] The weight calculation module is used for:
[0283] The weight of each specified action is calculated using the following formula;
[0284]
[0285] Where w(k) is the weight of the k-th specified action;
[0286]
[0287] FRR k The liveness rejection rate is the training sample set for the k-th specified action.
[0288] Based on the weights of each specified action calculated above, the fusion score is calculated using the following formula;
[0289]
[0290] Where s is the fusion score, and s(k) is the liveness detection score for the k-th specified action.
[0291] The decision fusion result is calculated using the following formula;
[0292]
[0293] Wherein, δ is the decision fusion result, and δ(k) is the liveness detection result of k specified actions;
[0294]
[0295] T(k) is the liveness detection threshold for the k-th specified action.
[0296] As another improvement to this embodiment of the invention, the liveness detection module is used for:
[0297] If the fusion score is greater than a set first threshold, the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0298] or,
[0299] If the decision fusion result is greater than the set second threshold, the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0300] or,
[0301] If the fusion score is greater than a set first threshold and the decision fusion result is greater than a set second threshold, then the user is determined to be a live user; otherwise, the user is determined to be a fake user.
[0302] The aforementioned specified actions may include one or more of the following: blinking, opening the mouth, turning the head left and right, nodding up and down, and remaining still in the front.
[0303] One of the methods is to acquire a face image by a specified action, extract features from the face image to obtain the image features of the specified action; or, one of the methods is to acquire multiple face images by a specified action, extract features from the multiple face images, and fuse the extracted features of the multiple face images to obtain the image features of the specified action.
[0304] In one example, the multiple face images include a first face image and a second face image. Feature extraction is performed on the first face image and the second face image to obtain first face features and second face features. The first face features and the second face features are then fused using the following module to obtain the image features:
[0305] The first principal component analysis module is used to perform principal component analysis on the first facial feature and the second facial feature using the following formula;
[0306]
[0307] Where l = 0 and 1 represent the numbers of the first facial feature and the second facial feature, respectively, and x l Representing the first facial feature and the second facial feature, z l This represents the feature obtained after performing principal component analysis on the first and second facial features, μ. l W represents the mean of the first facial feature and the mean of the second facial feature. l Represents the principal component analysis matrix of the first facial feature and the principal component analysis matrix of the second facial feature; μ l and W l It is obtained by training on the first face feature training set and the second face feature training set.
[0308] The feature fusion module is used to fuse the features obtained by principal component analysis of the first face feature and the second face feature using the following formula to obtain the image features;
[0309]
[0310] Where Z represents the image feature, P o and P 1 Let P be the projection matrix. o and P 1 It is obtained by training using the first facial feature training set and the second facial feature training set.
[0311] The aforementioned μ l W l P o and P 1The following modules were trained to obtain the following:
[0312] The mean calculation module is used to calculate the mean of the first facial feature and the mean of the second facial feature using the following formula;
[0313]
[0314] in, Let i be the i-th first face feature sample / second face feature sample in the first face feature training set / second face feature training set, where i = 1, 2, 3, ..., n, and n is the total number of the first face feature samples / second face feature samples.
[0315] The covariance matrix calculation module is used to calculate the covariance matrix of the first face feature sample and the second face feature sample using the following formula;
[0316]
[0317] Among them, S l Let be the covariance matrix of the first face feature sample and the second face feature sample.
[0318] The eigenvalue decomposition module is used to perform eigenvalue decomposition on the covariance matrix of the first face feature sample and the second face feature sample using the following formula:
[0319]
[0320] in, and S l Eigenvalues and eigenvectors, t = 0, 1, 2, ..., d, where d is the eigenvalue of S. l The number of rows and columns.
[0321] The principal component analysis matrix acquisition module is used to obtain the matrix based on the percentage of principal components. The principal component analysis matrix W is constructed by selecting the top r eigenvectors, which account for 95% of the total. l ;
[0322] in,
[0323] The second principal component analysis module is used to perform principal component analysis on the first face feature sample and the second face feature sample using the following formula;
[0324]
[0325] in, This represents the features obtained after performing principal component analysis on the first and second face feature samples.
[0326] The projection matrix calculation module is used to optimize and obtain the projection matrix P using the following formula. o and P 1 ;
[0327]
[0328] in, and Z 0 and Z 1 The within-class variance matrix, Let Z0 and Z1 be the inter-class variance matrices.
[0329] The device provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiment 1. For the sake of brevity, any parts not mentioned in this device embodiment can be referred to the corresponding content in the aforementioned method embodiment 1. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the aforementioned device and unit can all be referred to the corresponding processes in the aforementioned method embodiment 1, and will not be repeated here.
[0330] Example 3:
[0331] The method described in Embodiment 1 of this invention can implement business logic through a computer program and record it on a storage medium. This storage medium can be read and executed by a computer, achieving the effects of the solution described in Embodiment 1 of this specification. Therefore, this invention also provides a computer-readable storage medium for face liveness detection, including a memory for storing processor-executable instructions. When executed by a processor, the instructions implement the steps of the face liveness detection method of Embodiment 1.
[0332] The storage medium may include a physical device for storing information, typically digitizing the information and then storing it using electrical, magnetic, or optical methods. The storage medium may include: devices that store information using electrical energy, such as various types of memory, like RAM and ROM; devices that store information using magnetic energy, such as hard disks, floppy disks, magnetic tapes, magnetic core memory, bubble memory, and USB flash drives; and devices that store information using optical methods, such as CDs or DVDs. Of course, there are other readable storage media, such as quantum memories and graphene memories.
[0333] The storage medium described above may also include other implementation methods according to the description of method embodiment 1. The implementation principle and technical effects of this embodiment are the same as those of the aforementioned method embodiment 1. For details, please refer to the description of the relevant method embodiment 1, which will not be repeated here.
[0334] Example 4:
[0335] The present invention also provides a device for face liveness detection. The device may be a standalone computer, or it may include an actual operating device that uses one or more of the methods or embodiments described in this specification. The face liveness detection device may include at least one processor and a memory storing computer-executable instructions. When the processor executes the instructions, it implements the steps of any one or more of the face liveness detection methods described in Embodiment 1.
[0336] The device described above may include other implementation methods according to the description of method embodiment 1. The implementation principle and technical effects of this embodiment are the same as those of the aforementioned method embodiment 1. For details, please refer to the description of the relevant method embodiment 1, which will not be repeated here.
[0337] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the scope of the technology disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. All should be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for detecting human face liveness, characterized in that, The method includes: The system prompts the user to perform several specified actions and performs a liveness detection. After the liveness detection of the action is passed, the face image corresponding to each specified action is acquired; Feature extraction is performed on the face image corresponding to each specified action to obtain the image features of each specified action, and the liveness detection score of each specified action is calculated based on the image features of each specified action. The liveness detection scores for each specified action are weighted and summed according to the weights set for each specified action to obtain the fusion score; The weight of each specified action is determined based on the distribution of the liveness detection scores in the training sample set for that specified action and the liveness rejection rate in the training sample set for that specified action. The liveness detection result of each specified action is determined based on the liveness detection score of each specified action and the set liveness detection threshold of each specified action. The liveness detection result of each specified action is then weighted and voted on according to the set weight of each specified action to obtain the decision fusion result. The user is determined to be either a live or a fake based on the fusion score and / or the decision fusion result.
2. The face liveness detection method according to claim 1, characterized in that, The weight of each specified action is determined using the following method: Feature extraction is performed on the training sample set for each specified action to obtain the image feature set for each specified action, and the liveness detection score set for each specified action is calculated based on the image feature set for each specified action. The training sample set includes live samples and spur samples, and the live detection score set includes live scores and spur scores. The liveness detection score set for each specified action is normalized based on the non-overlapping regions of the distribution of liveness scores and the distribution of spur scores for each specified action. The weight of each specified action is calculated based on the normalized liveness detection score set for each specified action and the liveness rejection rate of the training sample set for each specified action.
3. The face liveness detection method according to claim 2, characterized in that, The normalization of the liveness detection score set for each specified action based on the non-overlapping regions of the distribution of liveness scores and the distribution of spurious scores for each specified action includes: The non-overlapping regions of the distribution of live body scores and the distribution of prosthetic scores for each specified action are calculated using the following formula, and are used as the normalized dimensions for each specified action. q(k)=μ a (k)-min a (k)+max r (k)-μ r (k) Where q(k) is the normalized dimension of the k-th specified action, k = 1, 2, ..., m, m is the total number of specified actions, and max r (k) represents the maximum liveness score in the liveness detection score set for the k-th specified action, μ r (k) represents the average liveness score in the liveness detection score set for the k-th specified action, min a (k) represents the minimum spur score in the liveness detection score set for the k-th specified action, μ a (k) is the average value of the prosthesis scores in the liveness detection score set for the k-th specified action; The liveness detection score set for each specified action is normalized according to the normalized dimensions of each specified action; in, and Let i be the liveness detection score in the liveness detection score set for the k-th specified action before and after normalization, respectively, where i = 1, 2, 3, ..., N. k N k Let j = 0 represent the total number of liveness detection scores in the liveness detection score set for the k-th specified action. The score represents the prosthesis score, and j=1 represents the liveness detection score. The number of living individuals; The step of calculating the weight of each specified action based on the normalized liveness detection score set for each specified action and the liveness rejection rate of the training sample set for each specified action includes: The weight of each specified action is calculated using the following formula; Where w(k) is the weight of the k-th specified action; FRR k The liveness rejection rate is the training sample set for the k-th specified action.
4. The face liveness detection method according to claim 3, characterized in that, The fusion score is calculated using the following formula; Where s is the fusion score, and s(k) is the liveness detection score for the k-th specified action; The decision fusion result is calculated using the following formula; Wherein, δ is the decision fusion result, and δ(k) is the liveness detection result of k specified actions; T(k) is the liveness detection threshold for the k-th specified action.
5. The face liveness detection method according to any one of claims 1-4, characterized in that, The specified actions include one or more of the following: blinking, opening the mouth, turning the head left and right, nodding up and down, and remaining still in the front. A specified action acquires a face image, performs feature extraction on the face image, and obtains the image features of the specified action; or, a specified action acquires multiple face images, performs feature extraction on the multiple face images, and fuses the extracted features of the multiple face images to obtain the image features of the specified action.
6. The face liveness detection method according to claim 5, characterized in that, The multiple face images include a first face image and a second face image. Feature extraction is performed on the first face image and the second face image to obtain first face features and second face features. Feature fusion is then performed on the first face features and the second face features to obtain the image features: Principal component analysis was performed on the first facial feature and the second facial feature using the following formula; Where l = 0 and 1 represent the numbers of the first facial feature and the second facial feature, respectively, and x l Representing the first facial feature and the second facial feature, z l This represents the feature obtained after performing principal component analysis on the first and second facial features, μ. l W represents the mean of the first facial feature and the mean of the second facial feature. l Represents the principal component analysis matrix of the first facial feature and the principal component analysis matrix of the second facial feature; μ l and W l It is obtained by training on the first face feature training set and the second face feature training set; The image features are obtained by fusing the features obtained from principal component analysis of the first and second facial features using the following formula; Where Z represents the image feature, P o and P 1 Let P be the projection matrix. o and P 1 It is obtained by training using the first facial feature training set and the second facial feature training set.
7. The face liveness detection method according to claim 6, characterized in that, The μ is obtained through training using the following method. l W l P o and P 1 : The mean value of the first facial feature and the mean value of the second facial feature are calculated using the following formulas; in, The i-th first face feature sample / second face feature sample in the first face feature training set / second face feature training set, i = 1, 2, 3, ..., n, where n is the total number of first face feature samples / second face feature samples; The covariance matrix of the first face feature sample and the second face feature sample is calculated using the following formula; Among them, S l Let be the covariance matrix of the first face feature sample and the second face feature sample; The covariance matrix of the first face feature sample and the second face feature sample is decomposed into eigenvalues using the following formula; in, and S l Eigenvalues and eigenvectors, t = 0, 1, 2, ..., d, where d is the eigenvalue of S. l The number of rows and columns; Based on the percentage of the main component The principal component analysis matrix W is constructed by selecting the top r eigenvectors, which account for 95% of the total. l ; in, Principal component analysis was performed on the first and second facial feature samples using the following formula; in, This represents the features obtained after performing principal component analysis on the first and second face feature samples; The projection matrix P is obtained by optimizing the following formula. o and P 1 ; in, and Z 0 and Z 1 The within-class variance matrix, For Z 0 and Z 1 The inter-class variance matrix.
8. A face liveness detection device, characterized in that, The device includes: The motion liveness detection module is used to prompt the user to perform several specified actions and to perform motion liveness detection. The data acquisition module is used to acquire the face image corresponding to each specified action after the action liveness detection is passed; The liveness detection score calculation module is used to extract features from the face image corresponding to each specified action, obtain the image features of each specified action, and calculate the liveness detection score of each specified action based on the image features of each specified action. The fusion score calculation module is used to perform a weighted summation of the liveness detection scores of each specified action according to the set weight of each specified action, so as to obtain the fusion score; The weight of each specified action is determined based on the distribution of the liveness detection scores in the training sample set for that specified action and the liveness rejection rate in the training sample set for that specified action. The decision fusion result calculation module is used to determine the liveness detection result of each specified action based on the liveness detection score of each specified action and the set liveness detection threshold of each specified action, and to perform weighted voting on the liveness detection result of each specified action according to the set weight of each specified action to obtain the decision fusion result. The liveness detection module is used to determine whether the user is a live or spoofed based on the fusion score and / or the decision fusion result.
9. A computer-readable storage medium for face liveness detection, characterized in that, It includes a memory for storing processor-executable instructions, which, when executed by the processor, implement the steps of the face liveness detection method according to any one of claims 1-7.
10. A device for face liveness detection, characterized in that, It includes at least one processor and a memory storing computer-executable instructions, wherein the processor executes the instructions to implement the steps of the face liveness detection method according to any one of claims 1-7.
Citation Information
Patent Citations
Human face living body detection method and device
CN110751069A
Human face silence living body detection method and device, readable storage medium and equipment
CN111860078A