Living body detection method, electronic device and storage medium
By collecting and processing face data, forming a training data set and training a live detection model, the problem of low near-infrared live detection efficiency in the existing technology is solved, and efficient and accurate live detection is achieved.
Patent Information
- Application Number
- CN202011014555.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-09-24
AI Technical Summary
In the prior art, when performing near-infrared live detection, 3D Align is usually used to train the face area in the face detection frame, resulting in longer time consumption and lower efficiency.
By collecting face data, using the face detection frame to intercept and similar transformation to obtain the first face data and the second face data, forming a training data set, and through data preprocessing and enhancement, the live detection model is trained for live detection.
The process of 3D Align is reduced, the efficiency and accuracy of live detection is improved, and the detection experience is enhanced, and it is suitable for scenarios such as financial payment and access control.
Smart Images

Figure CN114332967B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to a method for detecting a living body, an electronic device, and a storage medium. Background Art
[0002] Living body detection needs to determine whether the captured face is a real face or a forged face attack (such as a printed face on colored paper, a face image on the screen of an electronic device, and a mask, etc.). Currently, living body detection is often divided into silent living body detection and action-based living body detection according to whether cooperation is required. Cooperative living body detection requires the tester to make corresponding actions according to instructions, and has high living body detection performance but low efficiency and poor experience effect. Non-cooperative living body detection, on the other hand, performs living body detection without the tester making command actions, with a better experience effect but average performance.
[0003] Living body detection can be divided into visible light, near-infrared, and depth according to the input modality Figure 3 types. The cost of depth maps is relatively high, so generally monocular visible light living body detection or binocular (visible light + near-infrared) is used for living body detection. However, when using deep learning for near-infrared living body detection, the 3D Align method is usually used to train the face region part in the detected face bounding box, which takes a long time and has low efficiency. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] In view of the above problems, the main purpose of the present disclosure is to provide a method for detecting a living body, an electronic device, and a storage medium, so as to at least partially solve at least one of the above-mentioned technical problems.
[0006] (2) Technical Solutions
[0007] According to one aspect of the present disclosure, a method for detecting a living body is provided, including:
[0008] Collecting face data;
[0009] Processing the face data to obtain training data;
[0010] Training a living body detection model using the training data;
[0011] And performing living body detection using the trained living body detection model;
[0012] Wherein, the training data includes first face data and second face data, the first face data is obtained by intercepting the collected face data with a face detection bounding box, and the second face data is obtained by performing a similarity transformation on the collected face data.
[0013] Further, processing the face data to obtain training data includes:
[0014] Preprocessing the collected face data;
[0015] And performing data augmentation on the preprocessed data to obtain training data.
[0016] Further, preprocessing the collected face data includes:
[0017] For the collected face data, using a face detection box to intercept the face rectangular area, thereby obtaining face data with background information as the first face data;
[0018] For the collected face data, according to the key point positions of the left eye, right eye, nose, left mouth corner, and right mouth corner of the face, performing similarity transformation to the frontal face area, thereby obtaining face data without background information as the second face data.
[0019] Further, collecting face data includes:
[0020] Presetting different fill light shapes and fill light powers;
[0021] Under different fill light shapes and fill light powers, collecting live face image data and prosthetic face image data at different poses.
[0022] Further, the fill light shape is ring fill light or strip fill light, and the fill light power is 0.5 - 1.0 times the rated power.
[0023] Further, performing live detection using the trained live detection model includes:
[0024] Inputting the face data to be detected into the trained live detection model;
[0025] Determining the key point information and model live detection probability data in the face data to be detected;
[0026] Using the key point information and model live detection probability data to perform live detection on the face data to be detected and determining the live detection result;
[0027] Wherein the key point information includes the eye tilt angle theta1 of the face, the mouth tilt angle theta2, the horizontal distance lenx_dis from the left eye to the nose, the horizontal distance renx_dis from the nose to the right eye, the horizontal distance lmnx_dis from the left mouth corner to the nose, and the horizontal distance rmnx_dis from the nose to the right mouth corner; the model liveness detection probability data includes the liveness detection probability Prob_bb of the liveness detection model for the face in the face detection frame and the liveness detection probability Prob_st of the liveness detection model for the face of the similarity transformation.
[0028] Further, the liveness detection result satisfies the following relational expression:
[0029] Alive = (Prob_bb >= Prob_thresh) && (Prob_st >= Prob_thresh) && (theta1 >= 0) && (theta1 <= theta_thresh) && (theta2 >= 0) && (theta2 <= theta_thresh) && (x_ratio >= x_ratio_thresh_min) && (x_ratio <= x_ratio_thresh_max);
[0030] Wherein, Prob_thresh is a preset probability threshold, theta_thresh is a preset angle threshold, x_ratio_thresh is a preset ratio threshold, x_ratio_thresh_min is the minimum value of the preset ratio threshold, and x_ratio_thresh_max is the maximum value of the preset ratio threshold.
[0031] Further, the eye tilt angle theta1 satisfies the following relational expression:
[0032] theta1 = 180 / math.pi * math.atan2((re[1] - le[1]), (re[0] - le[0]));
[0033] The mouth tilt angle theta2 satisfies the following relational expression:
[0034] theta2 = 180 / math.pi * math.atan2((r_mouth[1] - l_mouth[1]), (r_mouth[0] - l_mouth[0]));
[0035] lenx_dis = nose[0] - le[0] + 1;
[0036] renx_dis = re[0] - nose[0] + 1;
[0037] lmnx_dis = nose[0] - 1_mouth[0] + 1;
[0038] rmnx_dis = r_mouth[0] - nose[0] + 1;
[0039] lrmnx_ratio = lenx_dis / rmnx_dis;
[0040] lrenx_ratio = renx_dis / lmnx_dis;
[0041] x_ratio = lrmnx_ratio / lrenx_ratio;
[0042] Among them, nose[0] represents the abscissa of the nose, le[0] represents the abscissa of the left eye, re[0] represents the abscissa of the right eye, l_mouth[0] represents the abscissa of the left mouth corner, and r_mouth[0] represents the abscissa of the right mouth corner.
[0043] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0044] One or more processors;
[0045] A storage device for storing one or more programs,
[0046] wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method described above.
[0047] According to still another aspect of the present disclosure, there is provided a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the method described above.
[0048] (III) Beneficial effects
[0049] It can be seen from the above technical solutions that the living body detection method, electronic device and storage medium of the present disclosure have at least one of the following beneficial effects:
[0050] (1) The present disclosure uses similarity transformation to replace face extraction without background information, reducing the 3D Align process.
[0051] (2) The present disclosure selects a corresponding prediction method according to different acquisition methods of the training data set, with higher accuracy.
[0052] (3) The present disclosure proposes a complete live detection solution, including data collection, preprocessing, training, preprocessing during testing, comprehensive live detection results, etc., improving the efficiency and experience of live detection, and having broad application prospects in financial payment, access control and other scenarios.
[0053] (4) The present disclosure simultaneously uses the key point information and the model live detection probability data to perform live detection on the face data to be detected, and determines the live detection result, thereby improving the accuracy of detection and avoiding misjudgments caused by simply using the model live detection probability data.
[0054] (5) The performance differences between real people, high-definition videos, and high-definition printed pictures in a monocular RGB camera are small. Near-infrared can naturally prevent screen attacks, and the absorption and reflection intensities of real faces and non-living carriers for the near-infrared band are also different. The present disclosure improves the accuracy of live detection through near-infrared. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure. In the drawings:
[0056] Figure 1 is a flowchart of the live detection method of the present disclosure.
[0057] Figure 2 is a flowchart of the process of obtaining training data by processing the face data of the present disclosure.
[0058] Figure 3 is a flowchart of collecting face data of the present disclosure.
[0059] Figure 4 is a flowchart of performing live detection using the trained live detection model of the present disclosure.
[0060] Figure 5 is a schematic diagram of the structure of the near-infrared fill light of the present disclosure.
[0061] Figure 6 is a schematic diagram of pictures collected of real people under different fill light shapes and different powers of the present disclosure.
[0062] Figure 7 is a schematic diagram of pictures collected of paper under different fill light shapes and different powers of the present disclosure.
[0063] Figure 8 is a schematic diagram of different processing methods of A4 photos of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the following provides a further detailed description of the present disclosure in conjunction with specific embodiments and with reference to the accompanying drawings.
[0065] Generally, when using deep learning for near-infrared in vivo detection, the 3D Align method is generally used to train the face region part in the extracted face detection box. However, 3D Align consumes a long time and thus has low efficiency.
[0066] In view of this, the present disclosure provides a method for in vivo detection. As Figure 1 shown, the method for in vivo detection includes:
[0067] Collecting face data;
[0068] Processing the face data to obtain training data;
[0069] Training an in vivo detection model using the training data;
[0070] And performing in vivo detection using the trained in vivo detection model;
[0071] Wherein, the training data includes first face data and second face data. The first face data is obtained by intercepting the collected face data through a face detection box, and the second face data is obtained by performing a similarity transformation on the collected face data.
[0072] The present disclosure uses similarity transformation to replace the extraction of faces without background information, reducing the 3D Align process.
[0073] According to an embodiment of the present disclosure, face data can be collected through a collection device such as a camera or a video camera. The collected face data includes in vivo face data of a living body and prosthetic face data. The in vivo face data of the living body is the face data of a real person, that is, the face data obtained by collecting face images of a real person. The prosthetic face data is non-real-person face data, such as a wax figure face diagram, a printed face image on colored paper, a face image on an electronic device screen, and a mask.
[0074] As Figure 2 shown, processing the face data to obtain training data includes:
[0075] Preprocessing the collected face data;
[0076] And performing data augmentation on the preprocessed data to thereby obtain training data.
[0077] According to an embodiment of the present disclosure, the data preprocessing method is specifically as follows:
[0078] For the case with background information: The rectangular area of the face is cropped according to the Bounding box (detection box) of the face, and thus the first face data is obtained. For the case without background information: The key point positions of the left eye, right eye, nose, left mouth corner, and right mouth corner of the face are transformed by similarity transformation to the face area of the frontal face, and thus the second face data is obtained.
[0079] According to the embodiments of the present disclosure, the data augmentation method is specifically as follows: The preprocessed data is augmented by random cropping.
[0080] Further, as Figure 3 shown, face data is collected, including:
[0081] Preset different fill light shapes and fill light powers;
[0082] Under different fill light shapes and fill light powers, live face image data and prosthetic face image data in different poses are collected.
[0083] According to the embodiments of the present disclosure, near-infrared fill lights are used to provide fill light. Exemplarily, the near-infrared fill lights are selected as fill lights with a switchable light source shape and adjustable power. The different poses are, for example, the poses corresponding to different rotation angles of the face, where the face rotation of 0° is the frontal pose. The range of the face rotation angle is preferably in the angle range of (0° - 30°). Here, the frontal pose means that the face is facing the acquisition device, and the face plane is parallel to the shooting plane of the acquisition device. The face can rotate in different directions such as up, down, left, and right relative to the frontal pose.
[0084] According to the embodiments of the present disclosure, during the training process, the live detection model is trained using the training data. Specifically, the training data is input into the live detection model, and the parameters of the live detection model are adjusted according to the output result of the live detection model until the output of the live detection model meets the requirements.
[0085] As Figure 4 shown, live detection is performed using the trained live detection model, including:
[0086] Input the face data to be detected into the trained live detection model;
[0087] Determine the key point information and the model live detection probability data in the face data to be detected;
[0088] Use the key point information and the model live detection probability data to perform live detection on the face data to be detected, and determine the live detection result.
[0089] According to an embodiment of the present disclosure, the key point information includes the eye tilt angle theta1 of the face, the mouth corner tilt angle theta2, the horizontal distance lenx_dis from the left eye to the nose, the horizontal distance renx_dis from the nose to the right eye, the horizontal distance lmnx_dis from the left mouth corner to the nose, and the horizontal distance rmnx_dis from the nose to the right mouth corner; the model liveness detection probability data includes the liveness detection probability Prob_bb of the liveness detection model for the face in the face detection box and the liveness detection probability Prob_st of the liveness detection model for the face of the similarity transformation. Among them, Prob_bb and Prob_st are forward prediction results.
[0090] The present disclosure simultaneously uses the key point information and the model liveness detection probability data to perform liveness detection on the face data to be detected, and determines the liveness detection result, thereby improving the detection accuracy and avoiding misjudgment caused by simply using the model liveness detection probability data.
[0091] The liveness detection method of the present disclosure will be introduced in detail below in conjunction with embodiments.
[0092] The liveness detection process of this embodiment mainly includes three parts: data collection, training, and prediction.
[0093] (1) Data collection
[0094] In the data collection process of this embodiment, supplementary light is provided by an infrared supplementary light. The selection of the shape of the near-infrared supplementary light (shape selection and power selection) is as follows:
[0095] As Figure 5 shown, the near-infrared supplementary light 1 of this embodiment includes a first light source array 11 and a second light source array 12. The first light source array includes a plurality of light sources 111 arranged in a ring, thereby providing a ring-shaped supplementary light shape. The second light source array includes a plurality of light sources 121 arranged in a rectangle, thereby providing a strip-shaped supplementary light shape. The near-infrared supplementary light is a power-adjustable supplementary light.
[0096] The near-infrared supplementary light may include a light source switching switch and a power adjustment switch, and can be switched to a mode where the ring-shaped supplementary light is on and the strip-shaped supplementary light is off, or a mode where the ring-shaped supplementary light is off and the strip-shaped supplementary light is on, or a mode where the ring-shaped supplementary light and the strip-shaped supplementary light are both on, or a mode where the ring-shaped supplementary light and the strip-shaped supplementary light are both off through the light source switching switch. The power adjustment switch can adjust the power of the supplementary light to 0.5 times the rated power, 0.6 times the rated power, 0.7 times the rated power, 0.8 times the rated power, 0.9 times the rated power, 1.0 times the rated power, etc.
[0097] When collecting data, the ring-shaped fill light can be turned on and the strip-shaped fill light can be turned off first. By adjusting the power adjustment switch respectively, the power of the fill light can be adjusted to 0.5 times the rated power, 0.6 times the rated power, 0.7 times the rated power, 0.8 times the rated power, 0.9 times the rated power, and 1.0 times the rated power. Under the above-mentioned power conditions of the fill light, the face images of real people are collected respectively. Then, the strip-shaped fill light is turned on and the ring-shaped fill light is turned off. By adjusting the power adjustment switch respectively, the power of the fill light can be adjusted to 0.5 times the rated power, 0.6 times the rated power, 0.7 times the rated power, 0.8 times the rated power, 0.9 times the rated power, and 1.0 times the rated power. Under the above-mentioned power conditions of the fill light, the face images of real people are collected respectively, as Figure 6 shown.
[0098] Next, the ring-shaped fill light is turned on and the strip-shaped fill light is turned off. By adjusting the power adjustment switch respectively, the power of the fill light can be adjusted to 0.5 times the rated power, 0.6 times the rated power, 0.7 times the rated power, 0.8 times the rated power, 0.9 times the rated power, and 1.0 times the rated power. Under the above-mentioned power conditions of the fill light, the face images on the paper are collected respectively. Finally, the strip-shaped fill light is turned on and the ring-shaped fill light is turned off. By adjusting the power adjustment switch respectively, the power of the fill light can be adjusted to 0.5 times the rated power, 0.6 times the rated power, 0.7 times the rated power, 0.8 times the rated power, 0.9 times the rated power, and 1.0 times the rated power. Under the above-mentioned power conditions of the fill light, the face images on the paper are collected respectively, as Figure 7 shown.
[0099] When collecting the face images of real people, for the collection pose, a frontal face or a pose with the rotation angle less than 30° in the up, down, left, and right directions can be selected; for the collection distance, a distance of 0.3 - 1.0 meters from the lens can be selected; for the collection environment, an outdoor environment or an indoor environment can be selected; for the collection time, day or night can be selected. When collecting the face images on the paper (A4 paper photos), select the original A4 paper photos, A4 paper photos with eyes removed, A4 paper photos with eyes and nose removed, and A4 paper photos with eyes, nose, and mouth removed, as Figure 8 shown; the A4 paper can be laid flat on the face of a real person or bent and laid on the face of a real person; for the collection pose, a frontal face or a pose with the rotation angle less than 30° in the up, down, left, and right directions can be selected; for the collection distance, a distance of 0.3 - 1.0 meters from the lens can be selected; for the collection environment, an outdoor environment or an indoor environment can be selected; for the collection time, day or night can be selected.
[0100] (2) Data preprocessing
[0101] Living body datasets should try to avoid interference from background information, but since the background information of the collected datasets is rich and not single, face images with background information and face images without background information can be used separately. Specifically, the face rectangular area can be cut out according to the face Bounding box as the face image with background information, and the key points of the left eye, right eye, nose, left corner of the mouth, and right corner of the mouth can be transformed to the face area of the front face as the face image without background information.
[0102] (3) Data enhancement method
[0103] Since near-infrared fill lights at different powers are used for fill lighting, there is no need to perform data enhancement on brightness information; since near-infrared images are similar to black and white images and in order to ensure that real-person information is not enhanced into photo attacks as much as possible, there is no need to perform data enhancement on chromaticity contrast and saturation, and only simple random cropping is used for data enhancement.
[0104] (4) Selection of training data ratio
[0105] An important indicator of liveness detection is the false rejection rate. In this public training data set, the ratio of real people to fake people is about 1:4.
[0106] (5) Prediction
[0107] In theory, the model should have a good liveness detection effect on the faces in the face detection frame and the faces after similar transformation. The prediction process is as follows:
[0108] 1. The calculation model determines the liveness probability of the face in the face detection frame as Prob_bb (probility bounding box);
[0109] 2. The calculation model calculates the probability of the live body of the face with similar transformation as Prob_st (probility similarity transformation);
[0110] Since the face rotation angle in the data set is within 30°, the model theoretically has a better effect on liveness detection of the front face, so (assuming that the left eye coordinate is le, the right eye coordinate is re, the nose coordinate is nose, the left corner of the mouth coordinate is 1_mouth, the right corner of the mouth coordinate is r_mouth, and the coordinate format is [x, y]);
[0111] 3. Calculate the eye tilt angle theta1:
[0112] theta1=180 / math.pi*math.atan2((re[1]-le[1]), (re[0]-le[0]));
[0113] 4. Calculate the inclination angle theta2 of the mouth corners:
[0114] theta2 = 180 / math.pi * math.atan2((r_mouth[1] - l_mouth[1]), (r_mouth[0] - l_mouth[0]));
[0115] When the face is at a frontal face angle, the horizontal distance lenx_dis from the left eye to the nose in the horizontal direction should be approximately equal to the horizontal distance renx_dis from the nose to the right eye, and the horizontal distance lmnx_dis from the left mouth corner to the nose should be approximately equal to the horizontal distance rmnx_dis from the nose to the right mouth corner;
[0116] 5. Calculate:
[0117] lenx_dis = nose[0] - le[0] + 1;
[0118] renx_dis = re[0] - nose[0] + 1;
[0119] lmnx_dis = nose[0] - l_mouth[0] + 1;
[0120] rmnx_dis = r_mouth[0] - nose[0] + 1;
[0121] lrmnx_ratio = lenx_dis / rmnx_dis;
[0122] lrenx_ratio = renx_dis / lmnx_dis;
[0123] x_ratio = lrmnx_ratio / lrenx_ratio;
[0124] Among them, nose[0] represents the abscissa of the nose, le[0] represents the abscissa of the left eye, re[0] represents the abscissa of the right eye, l_mouth[0] represents the abscissa of the left mouth corner, and r_mouth[0] represents the abscissa of the right mouth corner.
[0125] 6. Set Prob_thresh (preferably 0.9 in the present disclosure), theta_thresh (preferably 15° in the present disclosure), x_ratio_thresh (preferably 1 / 3 - 3 in the present disclosure), of course, it is not limited thereto.
[0126] 7. Comprehensive prediction result:
[0127] Alive = (Prob_bb >= Prob_thresh) && (Prob_st >= Prob_thresh) && (theta1 >= 0) && (theta1 <= theta_thresh) && (theta2 >= 0) && (theta2 <= theta_thresh) && (x_ratio >= x_ratio_thresh_min) && (x_ratio <= x_ratio_thresh_max).
[0128] Among them, && represents logical AND. When all the expressions of the above logical AND hold, the prediction result is a living body; otherwise, it is a prosthesis. The living body detection result Alive in the embodiments of the present disclosure is determined based on both the key point information (theta1, theta2, lenx_dis, renx_dis, lmnx_dis, rmnx_dis, etc.) and the model living body detection probability data (Prob_bb, Prob_st), thereby improving the accuracy of detection and avoiding misjudgment caused by solely using the model living body detection probability data.
[0129] According to the living body detection result Alive of the present disclosure, it can be determined whether it is a living body or a prosthesis. Specifically, a threshold can be preset. If the Alive value is greater than the threshold, it is determined as a living body; otherwise, it is determined as a prosthesis.
[0130] The present disclosure selects the corresponding prediction method according to different acquisition methods of the training data set, and the accuracy is higher.
[0131] So far, the present disclosure has been described in detail with reference to the accompanying drawings. Based on the above description, those skilled in the art should have a clear understanding of the present disclosure.
[0132] It should be noted that in the accompanying drawings or the text of the specification, the implementation manners that are not illustrated or described are all forms known to those of ordinary skill in the art and are not described in detail. In addition, the above definitions of each component are not limited to the various specific structures, shapes or manners mentioned in the embodiments, and those of ordinary skill in the art can make simple changes or substitutions to them.
[0133] Of course, according to actual needs, the present disclosure may also include other parts, which are not described herein because they have nothing to do with the innovative points of the present disclosure.
[0134] Similarly, it should be understood that, in order to streamline the present disclosure and assist in understanding one or more of the various disclosed aspects, in the above description of the exemplary embodiments of the present disclosure, the various features of the present disclosure are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed present disclosure requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the disclosed aspects lie in less than all the features of the single preceding disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present disclosure.
[0135] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.
[0136] The various component embodiments of the present disclosure can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the relevant devices according to the embodiments of the present disclosure. The present disclosure can also be implemented as a device or apparatus program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present disclosure can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or in any other form.
[0137] Furthermore, the ordinal numbers used in the specification and claims, such as "first", "second", "third", etc., are used to modify the corresponding elements. They do not imply or represent any ordinal number of the element itself, nor do they represent the order of one element relative to another element or the order in the manufacturing method. The use of these ordinal numbers is only to clearly distinguish an element with a certain name from another element with the same name.
[0138] In addition, in the drawings or the description of the specification, similar or identical parts are denoted by the same reference numerals. The technical features in each of the embodiments illustrated in the specification can be freely combined to form a new solution on the premise of no conflict. Additionally, each claim can be regarded as an embodiment alone, or the technical features in each claim can be combined to form a new embodiment. Moreover, in the drawings, the shape or thickness of the embodiments can be enlarged, and they can be simplified or conveniently labeled. Furthermore, the elements or implementation manners not depicted or described in the drawings are in forms known to those of ordinary skill in the art. Additionally, although this document may provide examples of parameters including specific values, it should be understood that the parameters need not exactly equal the corresponding values, but may approximate the corresponding values within an acceptable error tolerance or design constraint.
[0139] Unless there are technical obstacles or contradictions, the above various embodiments of the present disclosure can be freely combined to form additional embodiments, and these additional embodiments are all within the protection scope of the present disclosure.
[0140] Although the present disclosure has been described in conjunction with the accompanying drawings, the embodiments disclosed in the drawings are intended to exemplarily illustrate the preferred embodiments of the present disclosure and should not be construed as a limitation to the present disclosure. The dimensional ratios in the drawings are merely illustrative and should not be construed as a limitation to the present disclosure.
[0141] Although some embodiments of the general concept of the present disclosure have been shown and described, those of ordinary skill in the art will understand that changes can be made to these embodiments without departing from the principles and spirit of the general concept of the present disclosure. The scope of the present disclosure is defined by the claims and their equivalents.
[0142] The above are only the preferred embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for detecting a living body, characterized in that, Including: Collecting face data; Processing the face data to obtain training data, where the training data includes first face data and second face data. The first face data is obtained by intercepting the collected face data with a face detection frame, and the second face data is obtained by performing a similarity transformation on the collected face data. Processing the face data to obtain training data includes: preprocessing the collected face data; and performing data augmentation on the preprocessed data to obtain training data. Preprocessing the collected face data includes: for the collected face data, intercepting the face rectangular area with a face detection frame to obtain face data with background information as the first face data; for the collected face data, performing a similarity transformation on the key point positions of the left eye, right eye, nose, left mouth corner, and right mouth corner of the face to the face area of the frontal face to obtain face data without background information as the second face data. Training a live detection model using the training data; And performing live detection using the trained live detection model. Performing live detection using the trained live detection model includes: inputting the face data to be detected into the trained live detection model; determining the key point information and the model live detection probability data in the face data to be detected; using the key point information and the model live detection probability data to perform live detection on the face data to be detected and determining the live detection result. The key point information includes the eye tilt angle theta1, mouth tilt angle theta2, horizontal distance lenx_dis from the left eye to the nose, horizontal distance renx_dis from the nose to the right eye, horizontal distance lmnx_dis from the left mouth corner to the nose, and horizontal distance rmnx_dis from the nose to the right mouth corner of the face. The model live detection probability data includes the live detection probability Prob_bb of the live detection model for the face in the face detection frame and the live detection probability Prob_st of the live detection model for the face after the similarity transformation.
2. The living body detection method according to claim 1, characterized in that, Collecting face data includes: Presetting different light supplement shapes and light supplement powers; Collecting live face image data and prosthetic face image data in different poses under different light supplement shapes and light supplement powers.
3. The living body detection method according to claim 2, wherein The light supplement shape is annular light supplement or strip light supplement, and the light supplement power is 0.5 - 1.0 times the rated power.
4. The living body detection method according to claim 1, characterized in that, The live detection result satisfies the following relational expression: Alive = (Prob_bb >= Prob_thresh) && (Prob_st >= Prob_thresh) && (theta1 >=0) &&(theta1 <= theta_thresh) && (theta2 >=0) && (theta2<= theta_thresh) &&(x_ratio >= x_ratio_thresh_min) && (x_ratio <= x_ratio_thresh_max); lenx_dis = nose[0] - le[0] + 1; renx_dis = re[0] - nose[0] + 1; lmnx_dis = nose[0] - l_mouth[0] + 1; rmnx_dis = r_mouth[0] - nose[0] + 1; lrmnx_ratio = lenx_dis / rmnx_dis; lrenx_ratio = renx_dis / lmnx_dis; x_ratio = lrmnx_ratio / lrenx_ratio; Among them, && represents logical AND, Prob_thresh is a preset probability threshold, theta_thresh is a preset angle threshold, x_ratio_thresh is a preset ratio threshold, x_ratio_thresh_min is the minimum value of the preset ratio threshold, x_ratio_thresh_max is the maximum value of the preset ratio threshold, nose[0] represents the abscissa of the nose, le[0] represents the abscissa of the left eye, re[0] represents the abscissa of the right eye, l_mouth[0] represents the abscissa of the left corner of the mouth, and r_mouth[0] represents the abscissa of the right corner of the mouth.
5. The living body detection method according to claim 4, wherein The eye tilt angle theta1 satisfies the following relationship: theta1 = 180 / math.pi * math.atan2((re[1] - le[1]),(re[0] - le[0])); The mouth tilt angle theta2 satisfies the following relationship: theta2 = 180 / math.pi * math.atan2((r_mouth[1] - l_mouth[1]), (r_mouth[0] - l_mouth[0])); 6. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 5.
7. A computer-readable storage medium having executable instructions stored thereon, which when executed by a processor cause the processor to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
A method for living body detection based on an infrared camera
CN109190522A