Face living body detection method and device, intelligent equipment and computer program product
Through the combination of near-infrared camera and multi-frame judgment strategy, the problems of low accuracy and high cost of face live detection in the prior art are solved, and efficient and low-cost live detection without user cooperation are achieved.
Patent Information
- Application Number
- CN202510336603.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-11
AI Technical Summary
The existing facial live detection methods have problems such as low accuracy, user active cooperation or high hardware costs, and it is difficult to achieve efficient and low-cost live detection without user cooperation.
A near-infrared camera is used to collect infrared images of human faces, combined with the trained live detection model and multi-frame judgment strategy, and by analyzing the current and historical live detection results, the target live detection results are determined, reducing hardware costs and improving detection accuracy.
Without the need for active cooperation of users, through multi-frame judgment strategy and flexible historical image analysis, the accuracy of face live detection is improved, hardware costs are reduced, and user operations are simplified.
Smart Images

Figure CN120299096A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to a face liveness detection method, a face liveness detection device, an intelligent device, and a computer program product. Background Art
[0002] With the development of computer technology, face recognition technology has been continuously matured. Due to its advantages in terms of security and stability, as well as non-contact and concealment, face recognition technology has been widely applied in fields such as access control, attendance punching, system login, and payment. However, due to problems such as easy leakage / counterfeiting of faces, face recognition systems must be able to cope with fake face attacks.
[0003] Based on this, face liveness detection technology has been currently proposed, the purpose of which is to distinguish real faces from counterfeited faces. As an important part of ensuring system security in face recognition technology, this technology has received increasing attention in the industry. Currently, common face liveness detection methods have their respective drawbacks, some of which require active cooperation from users, and some are difficult to achieve a balance between detection accuracy and cost. Summary of the Invention
[0004] This application provides a face liveness detection method, a face liveness detection device, an intelligent device, and a computer program product, which can improve the accuracy of face liveness detection at a relatively low cost without the need for active cooperation from users.
[0005] In a first aspect, this application provides a face liveness detection method, including:
[0006] Obtain a face infrared image, where the face infrared image is collected by a near-infrared camera;
[0007] Detect the face infrared image with a pre-trained liveness detection model to obtain a liveness detection result;
[0008] When the liveness detection result indicates passing the liveness detection, determine a target liveness detection result according to the liveness detection result and historical liveness detection results, where the historical liveness detection results are obtained based on historical face infrared images, the historical face infrared images are the first N frames of face infrared images adjacent to the face infrared image, the historical face infrared images and the face infrared image include the same face, and N is determined based on the input frame sequence of the face infrared image in the pre-trained liveness detection model.
[0009] In a second aspect, this application provides a face liveness detection device, including:
[0010] An obtaining module, configured to obtain a face infrared image, where the face infrared image is collected by a near-infrared camera;
[0011] A detection module, configured to detect a face infrared image with a trained live detection model to obtain a live detection result;
[0012] A first determination module, configured to determine a target live detection result according to the live detection result and the historical live detection result when the live detection result indicates passing the live detection, where the historical live detection result is obtained based on historical face infrared images, the historical face infrared images are the first N frames of face infrared images adjacent to the face infrared image, the historical face infrared images and the face infrared image include the same face, and N is determined based on the input frame sequence of the face infrared image in the trained live detection model.
[0013] In a third aspect, the present application provides an intelligent device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method in the first aspect are implemented.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method in the first aspect are implemented.
[0015] In a fifth aspect, the present application provides a computer program product, which includes a computer program. When the computer program is executed by one or more processors, the steps of the method in the first aspect are implemented.
[0016] The beneficial effects of the present application compared with the prior art are as follows: The solution of the present application designs a multi-frame determination strategy. When the live detection result for the current face infrared image indicates passing the live detection, the live detection results of historical face infrared images can be combined for auxiliary discrimination, so as to obtain a more accurate target live detection result; moreover, the number of historical face infrared images to be considered can also be flexibly set according to the input frame sequence of the face infrared image in the trained live detection model, making the multi-frame determination strategy more flexible based on different time periods. On this basis, based on the solution of the present application, there is no need for the user to cooperate to make an indication action, simplifying the user operation; nor is there a requirement for the number of lenses of the near-infrared camera, reducing the hardware cost. In summary, the solution of the present application can improve the accuracy of face live detection at a low cost without the user's active cooperation.
[0017] It can be understood that the beneficial effects of the second to fifth aspects can refer to the relevant descriptions in the first aspect, and will not be repeated here. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 It is a schematic flow chart of the face liveness detection method provided by the embodiments of the present application;
[0020] Figure 2 It is a schematic diagram of the liveness detection model to be trained provided by the embodiments of the present application;
[0021] Figure 3 It is a structural block diagram of the face liveness detection device provided by the embodiments of the present application;
[0022] Figure 4 It is a schematic structural diagram of the intelligent device provided by the embodiments of the present application. Specific embodiments
[0023] The following will describe in detail the embodiments of the technical solutions of the present application with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present application, so they are only examples and cannot be used to limit the protection scope of the present application.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion.
[0025] In the description of the embodiments of the present application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features.
[0026] Referring to "embodiments" herein means that the specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0027] In the description of the embodiments of the present application, the term "and / or" is merely an association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0028] In the description of the embodiments of the present application, the term "plurality" refers to two or more (including two), unless otherwise specifically defined.
[0029] With the development of computer technology, face recognition technology has been continuously matured. Due to its advantages in terms of security and stability, as well as non-contact and concealment, face recognition technology has been widely applied in fields such as access control, attendance punching, system login, and payment. However, due to problems such as easy leakage / counterfeiting of faces, face recognition systems must be able to handle fake face attacks.
[0030] Based on this, currently, face liveness detection technology has been proposed, whose purpose is to distinguish real faces from counterfeited faces. As an important part of ensuring system security in face recognition technology, this technology has attracted increasing attention from the industry. Currently, there are three common face liveness detection methods as follows:
[0031] The first one: Use a conventional camera to obtain a color image, and extract features through traditional image processing methods and deep learning methods to determine whether it is a real face;
[0032] The second one: Instruct the user to complete a specified action to determine whether it is a real face;
[0033] The third one: Use a near-infrared camera to capture an infrared image, and combine the infrared image and depth information to determine whether it is a real face.
[0034] The above three methods each have their disadvantages. Specifically, the accuracy rate of the first method is relatively low; the second method requires user cooperation and takes a relatively long time; the third method usually requires a binocular or multi-view near-infrared camera to obtain depth information, resulting in a relatively high cost.
[0035] Based on this, the present application proposes a face liveness detection method, a face liveness detection device, an intelligent device, and a computer program product, which can achieve face liveness detection based on a relatively low cost without the need for active user cooperation, and designs a flexible multi-frame determination strategy to improve the accuracy of the liveness detection result. The following is illustrated through specific embodiments.
[0036] An embodiment of the present application proposes a method for face liveness detection. Among them, this face liveness detection method can be applied to any intelligent device capable of performing image processing operations, and the embodiment of the present application does not limit the type of this intelligent device. Please refer to Figure 1 , Figure 1 which gives the implementation process of this face liveness detection method, and the details are as follows:
[0037] Step 101, obtain a face infrared image.
[0038] The intelligent device can be equipped with or connected to a near-infrared camera. Compared with an ordinary camera, the near-infrared camera has the following advantages: on the one hand, it can resist ambient light interference and can stably image at night or under low-light conditions; on the other hand, it can prevent spoofing attacks. For example, a paper photo usually cannot correctly reflect infrared light, and a mask may have abnormal infrared reflection characteristics. And the embodiment of the present application does not limit the number of lenses of this near-infrared camera. Therefore, under the condition of controlling costs, this near-infrared camera can specifically be a monocular near-infrared camera. Compared with a binocular or multiocular near-infrared camera, the number of lenses of a monocular near-infrared camera is significantly reduced, and the hardware cost is lower. On this basis, the intelligent device can also be equipped with or connected to a single fill light to further save hardware costs.
[0039] The near-infrared camera can always be in the working state; or, it can also be default in the standby state and be woken up to the working state when an external event is detected, which is not limited here. Thus, the near-infrared camera can collect an infrared image when there is a suspected humanoid target approaching, and transmit the infrared image to the intelligent device. The intelligent device can perform face detection on the infrared image to locate the possible face area in the infrared image. After cropping based on this face area, it is the face infrared image.
[0040] Step 102, detect the face infrared image with a trained liveness detection model to obtain a liveness detection result.
[0041] The intelligent device is pre-deployed with a trained liveness detection model. Among them, this liveness detection model can be trained by other devices first and then deployed on this intelligent device after training; or, this liveness detection model can also be directly trained by this intelligent device, which is not limited here.
[0042] The currently obtained face infrared image can be used as the input of the trained liveness detection model to obtain the liveness detection result output by the trained liveness detection model. Specifically, there are two possible liveness detection results, namely: passing the liveness detection (that is, detecting a real face), and not passing the liveness detection (that is, detecting a forged face).
[0043] Step 103: When the live detection result indicates passing the live detection, determine the target live detection result according to the live detection result and the historical live detection result.
[0044] Due to factors such as ambient light changes and sensor noise, errors may occasionally occur when performing live detection on a single frame. Based on this, to improve the detection accuracy, the embodiments of this application propose a segmented multi-frame determination strategy to reduce misjudgment. Based on this strategy, the intelligent device can maintain the live detection results of the most recent N frames, that is, the historical live detection results.
[0045] Among them, the historical live detection result is obtained based on the historical face infrared images, and the historical face infrared images are the first N face infrared images adjacent to the face infrared image in this round of face live detection. In the actual application scenario, the value of N is an integer greater than or equal to 0, and can be specifically determined based on the input frame sequence of the face infrared image in the trained live detection model.
[0046] Assume that for the live detection model, the current face infrared image is the Xth face infrared image input to the live detection model in this round of face live detection. Then the historical face infrared images refer to: the (X - 1)th face infrared image input to the live detection model until the (X - N)th face infrared image; that is, the historical face infrared images refer to: in this round of face live detection process, the last N face infrared images input to the live detection model before the Xth face infrared image.
[0047] When the live detection result in this time indicates passing the live detection, the target live detection result can be determined by combining this live detection result and the historical live detection result. If the target live detection result indicates not passing the live detection, the historical live detection result is updated based on this live detection result in this time, and wait for the next live detection. If the target live detection result indicates passing the live detection, the processing flow of the live detection can be ended.
[0048] Specifically, when N is 0, the intelligent device can directly take the live detection result in this time as the standard, that is, determine that the target live detection result is: passing the live detection.
[0049] Specifically, when N is greater than 0, the liveness detection results of historical face infrared images need to be considered (that is, the historical liveness detection results need to be considered). The intelligent device can first count the number of target historical results, where the number of target historical results is the number of historical liveness detection results indicating passing the liveness detection. When the number of target historical results is greater than a preset number threshold, the intelligent device can determine that the target liveness detection result is passing the liveness detection. Conversely, when the number of target historical results is less than or equal to the number threshold, the intelligent device can determine that the target liveness detection result is failing the liveness detection. The number threshold is determined based on N.
[0050] In some examples, the number threshold can be obtained by multiplying N by a specified proportionality coefficient, and the specified proportionality coefficient is a decimal greater than 0 and less than 1. For example, the specified proportionality coefficient can be 0.5. Correspondingly, the number threshold can take the value of N / 2 (or the ceiling value of N / 2).
[0051] Conversely, when the current liveness detection result indicates failing the liveness detection, there is no need to obtain the target liveness detection result. The historical liveness detection results can be directly updated based on the current liveness detection result, and wait for the next liveness detection in this round of face liveness detection process. In addition, if the number of historical face infrared images is less than N frames, the intelligent device can also directly update the historical liveness detection results based on the current liveness detection result without obtaining the target liveness detection result, and wait for the next liveness detection in this round of face liveness detection process.
[0052] Herein, this round of face liveness detection refers to the process of continuously performing liveness detection on the same face. That is, the face infrared images obtained in this process are continuous and point to the same person. In practical application scenarios, the intelligent device can determine whether different face infrared images include the same face through methods such as target tracking, Hungarian algorithm, or feature matching, which is not limited herein. Of course, to avoid brute force cracking and reduce the power consumption of the intelligent device and the near-infrared camera, an upper limit value Y of the number of liveness detections and a preset duration T can also be set in this round of face liveness detection process. That is, in this round of face liveness detection process, if the number of liveness detections has reached Y times (that is, the liveness detection model has recognized Y frames of face infrared images), and the obtained target liveness detection result is always failing the liveness detection, this round of face liveness detection can be forced to exit, a prompt message indicating the failure of face recognition can be output, and the user is restricted to wait for at least the preset duration T before starting the next round of face recognition (that is, the next round of face liveness detection).
[0053] In some embodiments, to flexibly implement a segmented multi-frame determination strategy, the intelligent device can determine the value of N in the following manner:
[0054] A1. Determine the target frame sequence interval in which the input frame sequence of the face infrared image is located among at least two preset frame sequence intervals.
[0055] A2. Determine the number of frames for judgment corresponding to the target frame sequence interval as N.
[0056] The intelligent device can be preset with at least two frame sequence intervals, and each frame sequence interval can respectively correspond to its own number of frames for judgment. Generally speaking, the number of frames for judgment corresponding to a smaller frame sequence interval is also smaller. The reason is that when a real person (real face) appears in the picture, it will probably pass the live detection in the first few frames, while a fake person (imitated face) will try to break through the live detection model for a long time. On this basis, each frame sequence interval can also respectively correspond to its own specified proportional coefficient. Similar to the number of frames for judgment, the specified proportional coefficient corresponding to a smaller frame sequence interval can also be smaller (that is, the corresponding quantity threshold is smaller). In some examples, the corresponding relationship between the frame sequence interval, the number of frames for judgment, and the proportional coefficient can be shown in Table 1 below:
[0057] Frame sequence interval Judgment frame number Specified scale factor 1-5 0 0 6-30 5 0.6 30 or more 15 0.8
[0058] Table 1
[0059] The intelligent device can first determine the input frame sequence of the current face infrared image in the live detection model during this round of face liveness detection; that is, determine which frame the face infrared image is in the current round of face liveness detection input to the live detection model. After that, the intelligent device can know the frame sequence interval in which the input frame sequence is located. For the convenience of description, this frame sequence interval can be denoted as the target frame sequence interval. The intelligent device can thus take N as the number of frames for judgment corresponding to this target frame sequence interval.
[0060] Still taking Table 1 as an example, when the current face infrared image is the first 5 frames input to the live detection model during this round of face liveness detection, N can be taken as 0, that is, the target liveness detection result can be directly determined based on the liveness detection result corresponding to this current face infrared image in the future; when the current face infrared image is the 6th to 30th frames input to the live detection model during this round of face liveness detection, N can be taken as 5, that is, after obtaining the liveness detection result of the current face infrared image, it is necessary to combine this current liveness detection result and the historical liveness detection result (obtained based on the last 5 face infrared images before) to determine the target liveness detection result; and so on, which will not be elaborated here.
[0061] Based on the above flexible value-taking method for N (and for the specified proportional coefficient), the probability that a fake person breaks through the live detection model can be effectively reduced while ensuring that a real person quickly passes the live detection.
[0062] In some embodiments, in order to enable intelligent devices with limited computing power to also run a living body detection model with good performance and effects, the living body detection model can be trained in the following manner:
[0063] B1. Construct a living body detection model to be trained.
[0064] The living body detection model to be trained includes a feature extraction module and at least two branches for performing living body detection tasks. Among them, the feature extraction module can adopt a conventional network structure, including but not limited to mobilenet or resnet, etc., which is not limited here. Each branch is deployed with a classifier, and initially, different classifiers may have different classifier parameters. Please refer to Figure 2 , Figure 2 for a schematic illustration of the living body detection model to be trained.
[0065] B2. Calculate the model loss based on at least two branches.
[0066] The intelligent device can obtain training samples through conventional methods, including positive samples (samples belonging to real human faces) and negative samples (samples belonging to forged human faces), etc., which is not limited here. Of course, in this process, a data augmentation algorithm can also be applied to augment the collected training samples to improve the robustness and generalization ability of the trained model, which is not limited here.
[0067] After each training sample is input into the living body detection model to be trained, the image features can be first extracted by the feature extraction module of the living body detection model to be trained, and the image features can be used as the input of each branch; that is, during the training process, each branch (equivalent to each classifier) shares the same input. For each classifier, it can first extract classification features for classification from the image features, and then output the corresponding classification result based on the classification features, that is, the prediction result. It can be seen that each branch can correspondingly obtain the following two types of data: one is the classification features as intermediate products, and the other is the prediction result (classification result) as the final output.
[0068] Based on the above two types of data, the intelligent device can calculate the model loss, where the model loss includes: similarity loss and classification loss.
[0069] Specifically, the purpose of training includes widening the gap between the classification features extracted by different branches, so the similarity loss can be used to describe the difference degree between classifiers, and it can be obtained based on the classification features obtained by each branch; for example, it can be obtained by calculating the cosine similarity of the classification features obtained by each branch. Of course, it can also be considered to use schemes such as mean square error to calculate the similarity loss, which is not limited in the embodiments of this application.
[0070] In addition, the purpose of training also includes making the prediction results of each branch as close as possible to the true labels. Therefore, the classification loss can be used to describe the degree of difference between the prediction results and the preset true labels, and it can be obtained based on the prediction results of each branch. For example, for any branch, the classification loss of a single branch can be obtained by calculating the cross-entropy loss for the prediction results of that branch. Of course, other classification loss functions can also be considered to calculate the classification loss, and the embodiments of the present application do not limit this. Finally, the weighted sum of the classification losses of each branch can be used as the total classification loss. Among them, the purpose of weighting is to increase the penalty for misclassified samples. For example, for two branches, a larger weight coefficient will be used when both predict incorrectly, and the more misclassified branches, the greater the weight, which can accelerate the model convergence speed and help the results of different branches to tend to be consistent.
[0071] Only by way of example, the model loss can be expressed by the following formula:
[0072] Model loss = similarity loss + total classification loss * weight coefficient. Total classification loss = sum of the classification losses of each branch
[0073] Among them, the weight coefficient is positively correlated with the number of misclassified branches.
[0074] B3. Optimize the model parameters of the to-be-trained live detection model according to the model loss until the model loss converges, and complete the training of the to-be-trained live detection model to obtain the trained live detection model, where the trained live detection model only includes one branch.
[0075] The intelligent device can, based on the obtained model loss, backpropagate the gradients of all branches into the next round of training until the model loss converges. At this point, the training of the to-be-trained live detection model is completed. Due to the optimization of the classification loss, the prediction results of all branches will eventually tend to be consistent. Therefore, when deploying, the intelligent device only needs to select one classifier (branch) for inference, rather than running multiple classifiers (branches) simultaneously, which can reduce the computational overhead and improve the inference speed.
[0076] It can be understood that during the training process proposed in the present application, on the one hand, the parameter differences of each branch are widened through the similarity loss, and on the other hand, the prediction results of each branch are made consistent through the classification loss, thereby improving the final performance of the live detection model while limiting the size of the live detection model.
[0077] In some embodiments, in order to alleviate the problem that the performance of the live detection model significantly degrades in side light, backlight, or strong light scenarios, before step 102, the following operations can also be performed:
[0078] Obtain the image brightness based on the face infrared image.
[0079] Among them, the image brightness refers to the average brightness of the infrared face image, which can be specifically obtained by statistical means; or it can also be calculated through a brightness calculation model or traditional machine learning methods. The embodiments of the present application do not limit the acquisition method of the image brightness.
[0080] Considering that the infrared face image may include a certain background area, and the brightness of the background area usually varies greatly from that of the face area, the embodiments of the present application may consider statistically or calculating the image brightness only for a certain proportion of the area within the infrared face image (this area necessarily belongs to the face area), so as to exclude the interference of the background area. Of course, the face area and the background area in the infrared face image can also be accurately labeled through a segmentation scheme, and the image brightness is only statistically or calculated for the face area, so as to exclude the interference of the background area. In the embodiments of the present application, the method for excluding the interference of the background area is not limited.
[0081] Correspondingly, step 102 can be optimized as: when the image brightness is within a preset brightness range, the infrared face image is detected by a trained live detection model to obtain a live detection result. That is, after obtaining the image brightness, the intelligent device can first determine whether the image brightness is within the preset brightness range; if the image brightness is within this brightness range, it is considered that the brightness of the face area meets the brightness requirement, and the infrared face image can be input into the trained live detection model.
[0082] On the contrary, if the image brightness is within this brightness range, it is considered that the brightness of the face area does not meet the brightness requirement, and the infrared face image will not be input into the trained live detection model, but will be directly excluded. On this basis, the intelligent device can also adjust the image signal processing (ISP) parameters of the near-infrared camera. It can be understood that if the image brightness is less than the minimum value of this brightness range, the ISP parameters can be adjusted to increase the overall brightness of the subsequent captured images; if the image brightness is greater than the maximum value of this brightness range, the ISP parameters can be adjusted to decrease the overall brightness of the subsequent captured images. In some examples, the adjustment step size can be determined by the ratio of the image brightness to the preset target brightness. For example, if the current image brightness is 1 / 10 of the target brightness, then the ISP parameters can be adjusted to 10 times the original; or a parameter mapping table can be pre-constructed, which is used to express the amount of change required from different image brightnesses to the target brightness, and the adjustment step size can be directly obtained by looking up the table during adjustment.
[0083] In some embodiments, in order to reduce the misjudgment problem caused by the low-quality face images input to the live detection model, before step 102, the following operations can also be performed:
[0084] Evaluate the face infrared image based on at least one preset quality dimension, and determine whether to retain the face infrared image according to the evaluation result. That is, before inputting the face infrared image into the liveness detection model, the face infrared image can be filtered based on at least one preset quality dimension first.
[0085] In some examples, the at least one preset quality dimension may include, but is not limited to: whether it belongs to a face, the face angle, and the face quality score. For example, the intelligent device can first input the face infrared image into a preset classification model, and the classification model can distinguish between the two types of belonging to a face and not belonging to a face, so as to filter out non-faces and side faces with extremely large angles; only when the output of the classification model indicates that the face infrared image belongs to a face, further calculate the face angle of the face infrared image; only when the face angle is within a preset angle range, continue to calculate the face quality score of the face infrared image through an image quality scoring model; finally, when the face quality score is lower than a preset score threshold, the face infrared image is excluded, so that low-quality images that are too blurred, too small, and / or have occlusions will not be used as the input of the liveness detection model.
[0086] Correspondingly, step 102 can be optimized as: when the face infrared image is retained, use the trained liveness detection model to detect the face infrared image to obtain the liveness detection result.
[0087] In some embodiments, in an application scenario where both the image brightness and quality of the face infrared image need to be considered, since the statistics / calculation of the image brightness is relatively easier to implement, therefore, the filtering based on the image brightness can be performed first, and the filtering based on the quality can be performed later. That is, the execution order of the above steps can be optimized as: after step 101, first obtain the image brightness based on the face infrared image; then, when the image brightness is within a preset brightness range, evaluate the face infrared image based on at least one preset quality dimension, and determine whether to retain the face infrared image according to the evaluation result; thus, step 102 can use the trained liveness detection model to detect the face infrared image to obtain the liveness detection result when the face infrared image is retained. Of course, the execution order of the above steps can also be adjusted according to actual needs, as long as the input of the liveness detection model is a face infrared image that meets the brightness requirements and quality requirements. The embodiments of the present application do not limit this.
[0088] As can be seen from the above, the embodiment of the present application designs a multi-frame determination strategy. When the live detection result of the current face infrared image indicates passing the live detection, the live detection results of historical face infrared images can be combined for auxiliary discrimination, so as to obtain a more accurate target live detection result; moreover, the number of historical face infrared images to be considered can also be flexibly set according to the input frame sequence of the face infrared image in the trained live detection model, making the multi-frame determination strategy more flexible based on different time periods. On this basis, based on the embodiment of the present application, there is no need for the user to cooperate to make an indication action, which simplifies the user operation; nor is there a requirement for the number of lenses of the near-infrared camera, reducing the hardware cost. In summary, the embodiment of the present application can improve the accuracy of face live detection based on a lower cost without the user's active cooperation.
[0089] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiment of the present application.
[0090] Corresponding to the face live detection method provided above, the embodiment of the present application further provides a face live detection device. Please refer to Figure 3 , the face live detection device 3 in the embodiment of the present application includes:
[0091] An acquisition module 301, configured to acquire a face infrared image, where the face infrared image is acquired by a near-infrared camera;
[0092] A detection module 302, configured to detect the face infrared image with a trained live detection model to obtain a live detection result;
[0093] A first determination module 303, configured to determine a target live detection result according to the live detection result and the historical live detection result when the live detection result indicates passing the live detection, where the historical live detection result is obtained based on historical face infrared images, the historical face infrared images are the first N frames of face infrared images adjacent to the face infrared image, the historical face infrared images and the face infrared image include the same face, and N is determined according to the input frame sequence of the face infrared image in the trained live detection model.
[0094] In some embodiments, the first determination module 303 includes:
[0095] A first determination unit, configured to determine that the target live detection result is: passing the live detection when the live detection result indicates passing the live detection and N is 0;
[0096] A statistical unit, configured to count the number of target historical results when the in-living detection result indicates passing the in-living detection and N is greater than 0, where the number of target historical results is the number of historical in-living detection results indicating passing the in-living detection.
[0097] A second determination unit, configured to determine that the target in-living detection result is passing the in-living detection when the number of target historical results is greater than a preset number threshold, where the number threshold is determined based on N.
[0098] In some embodiments, the face in-living detection device 3 further includes:
[0099] A second determination module, configured to determine the target frame sequence interval in which the input frame sequence of the face infrared image is located among at least two preset frame sequence intervals, where each frame sequence interval corresponds to its own judgment number of frames.
[0100] A third determination module, configured to determine the judgment number of frames corresponding to the target frame sequence interval as N.
[0101] In some embodiments, the face in-living detection device 3 further includes: a training module for training a to-be-trained in-living detection model; where the training module includes:
[0102] A construction unit, configured to construct a to-be-trained in-living detection model, where the to-be-trained in-living detection model includes at least two branches for performing in-living detection tasks, and each branch is deployed with a classifier.
[0103] A calculation module, configured to calculate a model loss based on at least two branches, where the model loss includes: a similarity loss and a classification loss, where the similarity loss is used to describe the difference degree between classifiers, and the classification loss is used to describe the difference degree between the prediction result of the branch and a preset true label.
[0104] An optimization module, configured to optimize the model parameters of the to-be-trained in-living detection model according to the model loss until the model loss reaches convergence, complete the training of the to-be-trained in-living detection model, and obtain a trained in-living detection model, where the trained in-living detection model only includes one branch.
[0105] In some embodiments, the face in-living detection device 3 further includes:
[0106] A statistics module, configured to obtain an image brightness based on the face infrared image.
[0107] Correspondingly, the detection module 302 is specifically configured to detect the face infrared image with the trained in-living detection model to obtain an in-living detection result when the image brightness is within a preset brightness interval.
[0108] In some embodiments, the face liveness detection device 3 further includes:
[0109] An adjustment module, configured to adjust the image signal processing parameters of the near-infrared camera when the image brightness is not within the brightness range.
[0110] In some embodiments, the face liveness detection device 3 further includes:
[0111] An evaluation module, configured to evaluate the face infrared image based on at least one preset quality dimension;
[0112] A fourth determination module, configured to determine whether to retain the face infrared image according to the evaluation result;
[0113] Correspondingly, the detection module 302 is specifically configured to, when the face infrared image is retained, detect the face infrared image with the trained liveness detection model to obtain a liveness detection result.
[0114] As can be seen from the above, the embodiments of the present application design a multi-frame determination strategy. When the liveness detection result of the current face infrared image indicates passing the liveness detection, the liveness detection results of historical face infrared images can be combined for auxiliary discrimination, so as to obtain a more accurate target liveness detection result; moreover, the number of historical face infrared images to be considered can also be flexibly set according to the input frame sequence of the face infrared image in the trained liveness detection model, so that the multi-frame determination strategy is more flexible based on different time periods. On this basis, based on the embodiments of the present application, neither the user needs to cooperate to make an indication action, which simplifies the user operation; nor is there a requirement for the number of lenses of the near-infrared camera, which reduces the hardware cost. In summary, the embodiments of the present application can improve the accuracy of face liveness detection based on a low cost without the need for the user's active cooperation.
[0115] Corresponding to the face liveness detection method provided above, the embodiments of the present application further provide an intelligent device. Please refer to Figure 4 , the intelligent device 4 in the embodiments of the present application includes: a memory 401, one or more processors 402 ( Figure 4 only one is shown in ) and a computer program stored on the memory 401 and executable on the processor. Among them: the memory 401 is used to store software programs and modules, and the processor 402 executes various functional applications and data processing by running the software programs and units stored in the memory 401 to obtain the resources corresponding to the above preset events. Specifically, when the processor 402 runs the above computer program stored in the memory 401, the following steps are implemented:
[0116] Obtain a face infrared image, where the face infrared image is collected by a near-infrared camera;
[0117] Detect the face infrared image with the trained liveness detection model to obtain the liveness detection result;
[0118] When the liveness detection result indicates passing the liveness detection, determine the target liveness detection result according to the liveness detection result and the historical liveness detection result, where the historical liveness detection result is obtained based on the historical face infrared image, the historical face infrared image is the first N frames of face infrared images adjacent to the face infrared image, the historical face infrared image and the face infrared image include the same face, and N is determined based on the input frame sequence of the face infrared image in the trained liveness detection model.
[0119] Assume the above is the first possible implementation manner. Then, in the second possible implementation manner provided based on the first possible implementation manner, when the liveness detection result indicates passing the liveness detection, determining the target liveness detection result according to the liveness detection result and the historical liveness detection result includes:
[0120] When the liveness detection result indicates passing the liveness detection and N is 0, determine that the target liveness detection result is: passing the liveness detection;
[0121] When the liveness detection result indicates passing the liveness detection and N is greater than 0, count the number of target historical results, where the number of target historical results is: the number of historical liveness detection results indicating passing the liveness detection;
[0122] When the number of target historical results is greater than the preset number threshold, determine that the target liveness detection result is: passing the liveness detection, where the number threshold is determined based on N.
[0123] In the third possible implementation manner provided based on the above first possible implementation manner or the above second possible implementation manner, when the processor 402 runs the above computer program stored in the memory 401, the following steps are implemented:
[0124] In at least two preset frame sequence intervals, determine the target frame sequence interval where the input frame sequence of the face infrared image is located, where each frame sequence interval corresponds to its own judgment number of frames;
[0125] Determine the judgment number of frames corresponding to the target frame sequence interval as N.
[0126] In the fourth possible implementation manner provided based on the above first possible implementation manner or the above second possible implementation manner, when the processor 402 runs the above computer program stored in the memory 401, the following steps are implemented:
[0127] Build a live detection model to be trained. Among them, the live detection model to be trained includes at least two branches for performing live detection tasks, and each branch is deployed with a classifier;
[0128] Calculate the model loss based on at least two branches. The model loss includes: similarity loss and classification loss. Among them, the similarity loss is used to describe the degree of difference between classifiers, and the classification loss is used to describe the degree of difference between the prediction result of the branch and the preset true label;
[0129] Optimize the model parameters of the live detection model to be trained according to the model loss until the model loss converges, and complete the training of the live detection model to be trained to obtain a trained live detection model. Among them, the trained live detection model only includes one branch.
[0130] In the fifth possible implementation manner provided on the basis of the first possible implementation manner above, or the second possible implementation manner above, after obtaining the face infrared image, when the processor 402 runs the above computer program stored in the memory 401, the following steps are implemented:
[0131] Obtain the image brightness based on the face infrared image;
[0132] Correspondingly, detect the face infrared image with the trained live detection model to obtain a live detection result, including:
[0133] When the image brightness is within the preset brightness interval, detect the face infrared image with the trained live detection model to obtain a live detection result.
[0134] In the sixth possible implementation manner provided on the basis of the fifth possible implementation manner above, after obtaining the image brightness based on the face infrared image, when the processor 402 runs the above computer program stored in the memory 401, the following steps are implemented:
[0135] When the image brightness is not within the brightness interval, adjust the image signal processing parameters of the near-infrared camera.
[0136] In the seventh possible implementation manner provided on the basis of the first possible implementation manner above, or the second possible implementation manner above, after obtaining the face infrared image, when the processor 402 runs the above computer program stored in the memory 401, the following steps are implemented:
[0137] Evaluate the face infrared image based on at least one preset quality dimension;
[0138] Determine whether to retain the face infrared image according to the evaluation result;
[0139] Accordingly, the trained live detection model is used to detect the face infrared image, and the live detection result is obtained, including:
[0140] While retaining the face infrared image, the trained live detection model is used to detect the face infrared image, and the live detection result is obtained.
[0141] It should be understood that in the embodiment of the present application, the so-called processor 402 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0142] The memory 401 may include a read-only memory and a random access memory, and provide instructions and data to the processor 402. A part or all of the memory 401 may also include a non-volatile random access memory. For example, the memory 401 may also store information about the device type.
[0143] As can be seen from the above, the embodiment of the present application designs a multi-frame determination strategy. When the live detection result of the current face infrared image indicates passing the live detection, the live detection results of historical face infrared images can be combined for auxiliary discrimination, so as to obtain a more accurate target live detection result; moreover, the number of historical face infrared images to be considered can also be flexibly set according to the input frame sequence of the face infrared image in the trained live detection model, making the multi-frame determination strategy more flexible based on different time periods. On this basis, based on the solution of the present application, neither the user's cooperation to make an indication action is required, which simplifies the user operation; nor is there a requirement for the number of lenses of the near-infrared camera, which reduces the hardware cost. In summary, the solution of the present application can improve the accuracy of face live detection at a low cost without the user's active cooperation.
[0144] The embodiment of the present application also provides a computer program product. When the computer program product runs on an intelligent device, the intelligent device can implement the steps in the above-mentioned various method embodiments.
[0145] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be described in detail here.
[0146] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0147] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of external device software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0148] In the embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the above division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0149] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0150] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above embodiment methods of this application, it can also be completed by a computer program instructing related hardware. The above computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the above computer program includes computer program code, and the above computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The above computer-readable storage medium can include: any entity or device capable of carrying the above computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer-readable memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the above computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0151] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A face liveness detection method, characterized in that, Including: Obtain an infrared face image, where the infrared face image is collected by a near-infrared camera; Detect the infrared face image with a trained live detection model to obtain a live detection result; When the live detection result indicates passing the live detection, determine a target live detection result according to the live detection result and historical live detection results, where the historical live detection results are obtained based on historical infrared face images, the historical infrared face images are the first N frames of infrared face images adjacent to the infrared face image, the historical infrared face images and the infrared face image include the same face, and the N is determined based on the input frame sequence of the infrared face image in the trained live detection model.
2. The face liveness detection method according to claim 1, wherein, The step of determining a target live detection result according to the live detection result and historical live detection results when the live detection result indicates passing the live detection includes: When the live detection result indicates passing the live detection and N is 0, determine that the target live detection result is: passing the live detection; When the live detection result indicates passing the live detection and N is greater than 0, count the number of target historical results, where the number of target historical results is the number of historical live detection results indicating passing the live detection; When the number of target historical results is greater than a preset number threshold, determine that the target live detection result is: passing the live detection, where the number threshold is determined based on N.
3. The face liveness detection method according to claim 1 or 2, characterized in that, The face live detection method further includes: Determine a target frame sequence interval in which the input frame sequence of the infrared face image is located among at least two preset frame sequence intervals, where each frame sequence interval corresponds to a respective number of judgment frames; Determine the number of judgment frames corresponding to the target frame sequence interval as N.
4. The face liveness detection method according to claim 1 or 2, characterized in that, The training process of the live detection model includes: Construct a live detection model to be trained, where the live detection model to be trained includes at least two branches for performing live detection tasks, and each branch is deployed with a classifier; Calculate a model loss based on at least two branches, where the model loss includes: a similarity loss and a classification loss, where the similarity loss is used to describe the degree of difference between the classifiers, and the classification loss is used to describe the degree of difference between the prediction result of the branch and a preset true label; Optimize the model parameters of the live detection model to be trained according to the model loss until the model loss converges, and complete the training of the live detection model to be trained to obtain the trained live detection model, where the trained live detection model only includes one of the branches.
5. The face liveness detection method according to claim 1 or 2, characterized in that, After obtaining the infrared face image, the face live detection method further includes: Obtain the image brightness based on the infrared face image; Correspondingly, detecting the infrared face image with a trained live detection model to obtain a live detection result includes: When the image brightness is within a preset brightness range, the trained live detection model is used to detect the face infrared image to obtain the live detection result.
6. The face liveness detection method according to claim 5, wherein After obtaining the image brightness based on the face infrared image, the face liveness detection method further includes: When the image brightness is not within the brightness range, adjusting the image signal processing parameters of the near-infrared camera.
7. The face liveness detection method according to claim 1 or 2, characterized in that, After obtaining the face infrared image, the face liveness detection method further includes: Evaluating the face infrared image based on at least one preset quality dimension; Determining whether to retain the face infrared image according to the evaluation result; Accordingly, the detecting the face infrared image with the trained live detection model to obtain the live detection result includes: When the face infrared image is retained, the trained live detection model is used to detect the face infrared image to obtain the live detection result.
8. A face liveness detection device, characterized in that, Includes: An acquisition module, configured to acquire a face infrared image, where the face infrared image is acquired by a near-infrared camera; A detection module, configured to detect the face infrared image with a trained live detection model to obtain a live detection result; A first determination module, configured to determine a target live detection result according to the live detection result and historical live detection results when the live detection result indicates passing the live detection, where the historical live detection result is obtained based on historical face infrared images, the historical face infrared images are the first N frame face infrared images adjacent to the face infrared image, the historical face infrared images and the face infrared image include the same face, and the N is determined based on the input frame sequence of the face infrared image in the trained live detection model.
9. An intelligent device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by one or more processors, the method according to any one of claims 1 to 7 is implemented.