Living body detection method, device, electronic device and computer-readable storage medium

By combining RGB and HSV color information for texture detection, and combining blink and optical flow field detection, the problems of low accuracy and poor robustness of live organism detection in the prior art are solved, and effective recognition of high-definition photos, videos and 3D masks are achieved, reducing detection costs.

CN114529958BActive Publication Date: 2025-06-17ASIAINFO TECH CHINA INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202011190472.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-30
Publication Date
2025-06-17
Estimated Expiration
2040-10-30

AI Technical Summary

Technical Problem

The existing live detection methods have problems such as low accuracy, limited adaptation scenarios, and poor robustness, especially the inability to effectively identify high-definition photos, videos or 3D masks.

Method used

By acquiring the facial image of the detected object, texture detection is performed using a pre-trained texture detection model, combining RGB color mode and HSV color space information. If the detection result meets the preset conditions, further blink detection and optical flow field detection are performed to determine whether the detected object is a living body.

Benefits of technology

While reducing the cost of live detection, it can more significantly capture the texture differences of facial images, identify common photos, videos or 3D masks with attack marks, improving the accuracy and robustness of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529958B_ABST
    Figure CN114529958B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method, apparatus, electronic device, and computer-readable storage medium for live detection, which relate to the field of image recognition. The method includes: obtaining a facial image of a detected object; performing texture detection on the facial image to obtain a texture detection result; wherein, the performing texture detection on the facial image to obtain a texture detection result includes: obtaining color information of the facial image, and inputting the color information of the facial image into a pre-trained texture detection model to obtain the texture detection result output by the texture detection model. Embodiments of the present application can identify common photos, videos, or 3D masks made of hard materials such as plastics with attack marks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and more specifically, to a living body detection method, device, electronic device, and computer-readable storage medium. Background Art

[0002] Liveness detection is a method of determining the real physiological characteristics of an object in some identity authentication scenarios. It is used to verify whether the user is actually a living person. It can effectively resist common counterfeiting methods such as photos, face swaps, masks, occlusions, and screen shots, thereby helping users identify fraud and protect their interests.

[0003] Existing liveness detection methods are usually based on the texture difference between the real face image and the secondary image. For example, the traditional method is to extract texture features from the entire facial image using LBP (Local Binary Patterns), HOG (Histogram of Oriented Gradient), Haar wavelet transform and gray-level co-occurrence matrix to obtain a distribution histogram, and then use classifiers such as SVM and random forest to discriminate liveness. This liveness detection method has the disadvantages of low accuracy, limited adaptability to scenarios, and poor robustness. The deep learning method treats liveness detection as a binary classification problem, with the liveness category recorded as 1 and the non-liveness category recorded as 0. A convolutional neural network is used to extract features from the visible light image and then classify it. This method can only identify attack images with obvious traces, and cannot distinguish high-definition photos, videos or 3D masks. In order to solve this problem, some systems propose to add near-infrared cameras and depth cameras, input visible light images, near-infrared images and depth images into the neural network for cascading, and conduct comprehensive liveness detection. Although this method is relatively accurate, it greatly increases the detection cost and cannot meet the needs of some projects. Summary of the invention

[0004] Embodiments of the present invention provide a living body detection method, device, electronic device, and computer-readable storage medium that overcome the above problems or at least partially solve the above problems.

[0005] In a first aspect, a living body detection method is provided, the method comprising:

[0006] Acquire a facial image of a detected object;

[0007] Perform texture detection on the facial image to obtain texture detection results;

[0008] The texture detection is performed on the facial image to obtain the texture detection result, including:

[0009] Obtain the color information of the facial image, input the color information of the facial image into a pre-trained texture detection model, and obtain the texture detection result output by the texture detection model;

[0010] The color information includes RGB color mode information and HSV color space information; the texture detection model is trained with the color information of the sample facial image as the training sample and the texture detection result of the sample facial image as the training label; the texture detection result is used to characterize that the texture of the detected object is the texture of a living body or the texture of a non-living body.

[0011] Obtain the facial image of the detected object, including:

[0012] Obtain the original image, which includes the face of the detected object;

[0013] Determine at least four key points in the original image, and the at least four key points are respectively located at the two eyes and the two corners of the mouth of the detected object;

[0014] Construct squares with the at least four key points as the centers respectively, and each of the at least four constructed squares does not overlap with any other square and has the maximum area;

[0015] Stitch the areas of the at least four squares in the original image to obtain the facial image.

[0016] In another possible implementation, the texture detection model includes a first sub-network model, a second sub-network model, a feature fusion layer, and a result output layer;

[0017] Input the color information of the facial image into a pre-trained texture detection model, and obtain the texture detection result output by the texture detection model, including:

[0018] Input the color information of the facial image into the first sub-network model, and obtain the color features output after the first sub-network model performs multiple convolutions and poolings on the color information;

[0019] Input the color information of the facial image into the second sub-network model, and obtain the gradient features output by the second sub-network model. The gradient features have the same size as the color features, and the gradient features are used to characterize the gradient information of the facial image;

[0020] Input the color features and the gradient features into the feature fusion layer, and obtain the fusion features output by the feature fusion layer. The fusion features are used to characterize the result after the color features and the gradient features are fused;

[0021] Input the fusion features into the result output layer, and obtain the texture detection result output by the result output layer.

[0022] In yet another possible implementation, squares are constructed with at least four key points as centers respectively, including:

[0023] Calculate the distance from the key point at the left eye to the left border of the original image, denoted as the first distance; calculate the distance from the key point at the right eye to the right border of the original image, denoted as the second distance; calculate the distance from the key point at the left corner of the mouth to the left border of the original image, denoted as the third distance; calculate the distance from the key point at the right corner of the mouth to the right border of the original image, denoted as the fourth distance; calculate the horizontal distance between the key points at the two eyes, denoted as the fifth distance; calculate the horizontal distance between the key points at the two corners of the mouth, denoted as the sixth distance; calculate the vertical distance from the center of the key points at the two eyes to the center of the key points at the two corners of the mouth, denoted as the seventh distance;

[0024] Then, the side length of the square centered at the key point at the left eye is the minimum of the fifth distance, the seventh distance, and twice the first distance; the side length of the square centered at the key point at the right eye is the minimum of the fifth distance, the seventh distance, and twice the second distance; the side length of the square centered at the key point at the left corner of the mouth is the minimum of the sixth distance, the seventh distance, and twice the third distance; the side length of the square centered at the key point at the right corner of the mouth is the minimum of the sixth distance, the seventh distance, and twice the fourth distance.

[0025] In yet another possible implementation, it further includes:

[0026] If the texture detection result of the facial image meets the first preset condition, obtain another frame of the facial image of the detected object;

[0027] Perform texture detection on the other frame of the facial image. If the texture detection result of the other frame of the facial image meets the first preset condition, perform blink detection on the two frames of the facial image to obtain a blink detection result;

[0028] If the blink detection result meets the second preset condition, perform optical flow field detection on the two frames of the facial image, and determine whether the detected object is a live body according to the optical flow field detection result.

[0029] In yet another possible implementation, performing blink detection on two frames of facial images to obtain a blink detection result includes:

[0030] Extract the regions where the eyes are located in the two frames of facial images respectively to obtain two eye images;

[0031] For each eye image, input the eye image into a pre-constructed blink detection model to obtain the eye opening degree output by the blink detection model;

[0032] If the difference between the obtained eye opening degrees of the two eyes is greater than a preset threshold, a first blink detection result is generated; if the magnitudes of the obtained eye opening degrees of the two eyes are not greater than the preset threshold, a second blink detection result is generated.

[0033] Among them, the blink detection model is trained with sample eye images as training samples and the eye opening degrees and eye opening degree intervals of the sample eye images as training labels; the first blink detection result is used to represent that the detected object blinks; the second blink detection result is used to represent that the detected object does not blink.

[0034] In another possible implementation, the blink detection model includes a feature extraction layer, an interval output layer, and an opening degree output layer;

[0035] Inputting the eye image into a pre-constructed blink detection model to obtain the eye opening degree output by the blink detection model includes:

[0036] Inputting the eye image into the feature extraction layer to obtain the eye features output by the feature extraction layer;

[0037] Inputting the eye features into the interval output layer to obtain the eye opening degree interval of the eye image output by the interval output layer;

[0038] Inputting the eye features and the eye opening degree interval of the eye image into the opening degree output layer to obtain the eye opening and closing degree obtained by the output layer through regression of the eye features within the eye opening degree interval of the eye image.

[0039] In another possible implementation, the two frame facial images are facial images acquired at two non-adjacent moments within the same time period.

[0040] In another possible implementation, determining whether the detected object is a live body according to the optical flow field detection result includes:

[0041] If optical flow appears in the area other than the face in the two frame facial images or the included angle of the optical flow directions of more than a preset proportion of the pixel points in the facial area is less than a preset threshold, it is determined that the detected object is not a live body;

[0042] If no optical flow appears in the area other than the face in the two frame facial images or the included angle of the optical flow directions of less than a preset proportion of the pixel points in the facial area is less than a preset threshold, it is determined that the detected object is a live body.

[0043] In a second aspect, a live body detection device is provided, and the device includes:

[0044] A facial image acquisition module, configured to acquire a facial image of the detected object;

[0045] A texture detection module, configured to perform texture detection on the facial image to obtain a texture detection result;

[0046] Among them, the texture detection module includes:

[0047] An information acquisition unit for acquiring the color information of the facial image;

[0048] A detection unit for inputting the color information of the facial image into a pre-trained texture detection model to obtain a texture detection result output by the texture detection model;

[0049] The color information includes RGB color mode information and HSV color space information; the texture detection model is trained with the color information of the sample facial image as the training sample and the texture detection result of the sample facial image as the training label; the texture detection result is used to characterize that the texture of the detected object is a living texture or a non-living texture.

[0050] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method provided in the first aspect are implemented.

[0051] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in the first aspect are implemented.

[0052] The beneficial effects brought by the technical solution provided by the present application are as follows: The input of the living body detection method is only the facial image of the detected object, and no additional interaction with the detected object is required. Therefore, no special hardware is needed, which can reduce the cost of living body detection. By using both the RGB color mode and the HSV color space as color information, the texture differences in the facial image can be captured more significantly. The embodiments of the present application can identify common photos, videos, or 3D masks made of hard materials such as plastic with attack marks (blurring, water ripples, local reflection, etc.). Description of the Drawings

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0054] Figure 1 It is a schematic flowchart of a living body detection method provided by an embodiment of the present application;

[0055] Figure 2 It is a schematic diagram of a facial image obtained when face detection is performed on an original image in the prior art;

[0056] Figure 3 It is a schematic diagram of obtaining 4 squares in an embodiment of the present application;

[0057] Figure 4 For the embodiments of the present application, Figure 3 it is a schematic diagram of the facial image formed after merging 4 squares in

[0058] Figure 5 it is a schematic flowchart of inputting the gradient information and color information of the facial image into a pre-trained texture detection model to obtain the texture detection result output by the texture detection model for the embodiments of the present application;

[0059] Figure 6 it is a schematic flowchart of the live detection method for another embodiment of the present application;

[0060] Figure 7 it is a schematic flowchart of performing blink detection on two frames of facial images to obtain the blink detection result for the embodiments of the present application;

[0061] Figure 8 it is a schematic flowchart of inputting the eye image into a pre-constructed blink detection model to obtain the degree of eye opening and closing output by the blink detection model for the embodiments of the present application;

[0062] Figure 9 it is a schematic structural diagram of a live detection device provided by the embodiments of the present application;

[0063] Figure 10 it is a schematic structural diagram of an electronic device provided by the embodiments of the present application. Detailed Embodiments

[0064] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation of the present invention.

[0065] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0066] The living body detection method provided by the embodiments of the present application aims to solve the above technical problems in the prior art.

[0067] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be described in further detail below in conjunction with the accompanying drawings.

[0068] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0069] In existing living body detection means, it is usually determined based on the texture differences between real face images and re-imaged images. For example, for the entire facial image, texture features are extracted using LBP (Local Binary Patterns), HOG (Histogram of Oriented Gradient), Haar wavelet transform, gray-level co-occurrence matrix, etc. to obtain a distribution histogram, and then classifiers such as SVM and random forest are used for living body discrimination. However, existing living body detection means have defects such as low accuracy, limited applicable scenarios, and poor robustness.

[0070] The living body detection method, device, electronic device, and computer-readable storage medium provided by the present application aim to solve the above technical problems in the prior art.

[0071] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0072] Figure 1 For the embodiments of the present application, a flowchart of a living body detection method is provided, as Figure 1 shown, and the method includes the following steps:

[0073] S1. Obtain a facial image of the object to be detected.

[0074] In the embodiments of the present application, the device for collecting facial images can be a camera. After collecting the facial image, the camera sends it to the execution entity of the embodiments of the present application, so that the execution entity of the embodiments of the present application obtains the facial image of the object to be detected.

[0075] The external input required for the live detection method provided in this embodiment only includes ordinary facial images and does not require interaction with the detected object. Therefore, no special hardware is needed, which can reduce the cost of live detection. This live detection method can be deployed at the facial image acquisition end. For example, in the field of security applications, it can be deployed in systems such as access control systems and identity recognition systems based on face recognition; in the field of financial applications, it can be deployed at personal terminals, and personal terminals can include, for example, smart phones, tablet computers, personal computers, etc.

[0076] For example, the live detection method of the embodiments of the present application can also be distributively deployed at the server end (or cloud end) and the facial image acquisition end. For example, a control signal can be sent at the server end (or cloud end) and transmitted to the facial image acquisition end to control the facial image acquisition end to perform live detection; the facial image acquisition end sends out indication information according to the received control signal to obtain the facial image of the detected object, then performs texture detection according to the facial image, and finally the facial image acquisition end transmits the texture detection result to the server end (or cloud end).

[0077] For another example, indication information can be sent at the server end (or cloud end) and transmitted to the facial image acquisition end. The facial image acquisition end obtains the facial image of the detected object and transmits the facial image of the detected object to the server end (or cloud end). Then the server end (or cloud end) performs texture detection according to the received facial image of the detected object. If necessary, the server end (or cloud end) can also transmit the texture detection result to the facial image acquisition end.

[0078] S2. Perform texture detection on the facial image to obtain a texture detection result;

[0079] For example, in one example, step S2 includes the following steps:

[0080] S21. Obtain the color information of the facial image.

[0081] The color information in the embodiments of the present application includes RGB color model information and HSV color space information. The RGB color model information is a color standard in the industrial field. It obtains various colors by changing the three color channels of red (R), green (G), and blue (B) and their superposition with each other. R, G, and B represent the colors of the three channels of red, green, and blue respectively. The RGB color model assigns an intensity value within the range of 0 to 255 to the RGB components of each pixel in the image. For example: for pure red, the R value is 255, the G value is 0, and the B value is 0; for gray, the R, G, and B values are equal (except 0 and 255); for white, R, G, and B are all 255; for black, R, G, and B are all 0. In the embodiments of the present application, the RGB color model information of the facial image is determined by determining the RGB components of each pixel point in the facial image.

[0082] The HSV (Hue, Saturation, Value) color space information is a color space created according to the intuitive characteristics of colors, also known as the Hexcone Model. The color parameters in this model are: hue (H), saturation (S), and value (V). The hue H is measured in degrees, and the value range is 0° to 360°. It is calculated counterclockwise starting from red. Red is 0°, green is 120°, and blue is 240°. The saturation S has a value range of 0.0 to 1.0. The larger the value, the more saturated the color. The value V has a value range of 0 (black) to 255 (white). In the embodiments of the present application, the HSV color space information of the facial image is determined by determining the HSV parameters of each pixel point in the facial image.

[0083] When the prior art performs texture detection on a facial image, the color information usually used is the RGB color model information. However, through experimental verification, it is found in the embodiments of the present application that the HSV space is more in line with the human eye's perception. By combining the RGB color model information and the HSV color space information, it is verified that the texture differences of the facial image can be captured more significantly. In the embodiments of the present application, the combination of the RGB color model information and the HSV color space information can be in the form of direct matrix splicing. For example, both the RGB color model information and the HSV color space information are three-dimensional matrices with a size of length * width * 3. After splicing the two, a matrix of length * width * 6 is obtained, and this matrix is one of the inputs of the texture detection model.

[0084] S22. Input the color information of the facial image into a pre-trained texture detection model to obtain the texture detection result output by the texture detection model.

[0085] It should be understood that before performing step S22, a texture detection model can also be pre-trained. Specifically, the texture detection model can be trained in the following way: First, collect a certain number of sample facial images, obtain the color information of each sample facial image, and determine the texture detection result of each sample facial image. The texture detection result is used to characterize whether the texture of the detected object in the sample facial image belongs to the texture of a live body or a non-live body. Then, based on the color information of the sample facial image and the texture detection result of the sample facial image, the initial model is trained to obtain the texture detection model. Among them, the initial model can be a single neural network model or a combination of multiple neural network models.

[0086] The input of the live body detection method in the embodiments of the present application is only the facial image of the detected object, and no additional interaction with the detected object is required. Therefore, no special hardware is needed, which can reduce the cost of live body detection. By using the RGB color mode and the HSV color space together as the color information, the texture differences in the facial image can be captured more significantly. The embodiments of the present application can identify common photos, videos, or 3D masks made of hard materials such as plastic with attack marks (blurring, water ripples, local reflection, etc.).

[0087] When the image is re-imaged, texture features such as pixel blurring, local reflection, water ripples, and facial gradient flattening will be formed. These features are often more obvious near the facial features such as eyes, nose, and mouth, and less obvious in the flat areas such as the forehead and the periphery of the face. If only the texture information of the obvious feature areas is extracted and the interference of irrelevant features is discarded, the effect is better than processing the entire face. Therefore, on the basis of the above embodiments, as an alternative embodiment, obtaining the facial image of the detected object includes:

[0088] S11. Obtain the original image, which includes the face of the detected object.

[0089] It should be understood that the original image obtained in the embodiments of the present application includes not only the face of the detected object, but also information other than the face of the detected object, such as the information of the environment where the detected object is located. And the face of the detected object in the original image is complete or nearly complete. When the face of the detected object is incomplete, at least the two eyes and the mouth in the face should be detectable.

[0090] S12. Determine at least four key points in the original image, and the at least four key points are respectively located at the two eyes and the two corners of the mouth of the detected object.

[0091] Specifically, in the embodiments of the present application, the two eyes and the mouth of the detected object are first recognized from the original image, and then at least four key points are determined from the two eye corners and the two mouth corners. For example, algorithms such as MTCNN (Multi-task Cascaded Convolutional Neural Networks), YOLO (You Only Look Once) V3, faster-RCNN (region proposal-CNN), and SSD (Single Shot MultiBox Detector) can be used to perform face detection on the original image to recognize at least four key points in the two eyes and the mouth. Generally, the key points at the eyes are located at the pupils, which can be understood as the middle of the eyes, or the left and right corners of each eye, or further include the highest pixel point and the lowest pixel point of the eyes. The key points at the mouth can be the two corners of the mouth. The size of the mouth can be basically determined by the positions of the two corners, or further include the highest pixel point and the lowest pixel point of the mouth. The embodiments of the present application do not specifically limit the specific algorithm used in step S12.

[0092] It should be noted that existing algorithms such as MTCNN, YOLO V3, faster-RCNN, and SSD can all recognize a smaller range of facial images from the original image. Figure 2 FIG. is a schematic diagram of a facial image obtained when performing face detection on the original image in the prior art. As Figure 2 shown, the original image includes a half-body image of the detected object. From the half-body image, the hair, eyes, nose, mouth, and ears of the detected object can be seen. The rectangular frame 100 is the facial image obtained by performing face detection on the original image using the prior art. This facial image includes the eyes, nose, and mouth of the detected object.

[0093] S13: Construct squares with at least four key points as the centers respectively, and each of the at least four constructed squares does not overlap with any other square and has the largest area.

[0094] After obtaining the above at least 4 key points, the embodiments of the present application will construct squares with each key point as the center. Considering that the facial features of the detected object may be compact or scattered, and there are also differences in the expressions made at different times, there are restrictive conditions for the squares constructed in the embodiments of the present application: each of the four squares does not overlap with any other square and has the largest area. On the Figure 2 original image shown, 4 squares are obtained by using step S13. As Figure 3 shown. Figure 3Schematic diagram of obtaining four squares in the embodiment of the present application. Among them, square 201 is a square constructed with the key point at the left eye pupil as the center, square 202 is a square constructed with the key point at the right eye pupil as the center, square 203 is a square constructed with the left mouth corner as the center, and square 204 is a square constructed with the right mouth corner as the center. From Figure 3 it can be seen that since the distance between the eyes and the mouth is relatively far, when the square constructed with the key point at one eye reaches the situation where it just coincides with the square constructed with the key point at the other eye without overlapping areas and has the largest area, there is still a relatively large gap from the square constructed with the key point at the mouth corner, and Figure 3 the schematic diagram shown only includes the eyes and the mouth and exactly does not include the nose. In fact, the texture at the eyes and the mouth is more obvious than the texture at the nose. Therefore, the four squares obtained in the embodiment of the present application can encompass more important facial organs compared to the one rectangle constructed in the prior art, reducing the interference of irrelevant organs. It should be noted that, Figure 3 This is a special case obtained by the embodiment of the present application using step S13. In practical applications, the square area usually also includes a partial area of the nose, but generally it is smaller than the complete nose in one rectangle.

[0095] As an alternative embodiment, the step of constructing squares with at least four key points as the centers respectively in the embodiment of the present application includes:

[0096] Calculate the distance from the key point at the left eye to the left border of the original image, denoted as the first distance; calculate the distance from the key point at the right eye to the right border of the original image, denoted as the second distance; calculate the distance from the key point at the left mouth corner to the left border of the original image, denoted as the third distance; calculate the distance from the key point at the right mouth corner to the right border of the original image, denoted as the fourth distance; calculate the horizontal distance between the key points at both eyes, denoted as the fifth distance; calculate the horizontal distance between the key points at both mouth corners, denoted as the sixth distance; calculate the vertical distance from the center of the key points at both eyes to the center of the key points at both mouth corners, denoted as the seventh distance;

[0097] Then the side length of the square with the key point at the left eye as the center is the minimum value among the fifth distance, the seventh distance, and twice the first distance; the side length of the square with the key point at the right eye as the center is the minimum value among the fifth distance, the seventh distance, and twice the second distance; the side length of the square with the left mouth corner as the center is the minimum value among the sixth distance, the seventh distance, and twice the third distance; the side length of the square with the right mouth corner as the center is the minimum value among the sixth distance, the seventh distance, and twice the fourth distance.

[0098] S14. Stitch the areas of at least four squares in the original image to obtain a facial image.

[0099] Figure 4 In the embodiments of the present application, Figure 3 is a schematic diagram of the facial image formed after merging 4 squares in Figure 4 As shown, the way to splice the 4 squares is to connect the squares corresponding to the left and right eyes horizontally to form a rectangle; connect the squares corresponding to the left and right mouth corners horizontally to form another rectangle, and finally connect the two rectangles vertically to obtain the facial image. Optionally, after obtaining the two rectangles, the two rectangles can also be adjusted to the same size, so that the spliced one becomes a square.

[0100] In the existing method of using a convolutional neural network to extract features, although the network can be designed to extract as many useful texture features as possible, a shallow network layer will lead to incomplete feature extraction, and an overly deep layer will cause overfitting or performance problems. Moreover, the training of the parameters in the neural network highly depends on the training data. For the entire picture, whether useful information can be extracted is also an unknown problem.

[0101] Based on the above embodiments, as an alternative embodiment, the texture detection model of the embodiments of the present application includes a first sub-network model, a second sub-network model, a feature fusion layer, and a result output layer.

[0102] Figure 5 is a schematic flowchart of inputting the gradient information and color information of the facial image into a pre-trained texture detection model to obtain the texture detection result output by the texture detection model in the embodiments of the present application, as Figure 5 shown, and this process includes:

[0103] S21. Input the color information of the facial image into the first sub-network model to obtain the color features output by the first sub-network model after performing multiple convolutional and pooling operations on the color information.

[0104] The first sub-network model of the embodiments of the present application is used to perform convolutional and pooling operations on the color information and output color features, and the color features can be encoded color information. The first sub-network model can adopt a convolutional neural network including 10 convolutional layers and 3 pooling layers, so as to achieve 10 convolutional operations and 3 pooling operations. The input of the convolutional neural network can be a two-dimensional array, and the two-dimensional array can include multiple channels. In the embodiments of the present application, the color information of the facial image includes RGB color mode information and HSV color space information. Since both RGB and HSV images are 3-channel, they are three-dimensional matrices with a size of length * width * 3. After splicing the two, a matrix of length * width * 6 is obtained, and 6 represents the number of channels. In Figure 3Taking the facial image shown as an example, the sizes of the 4 square patterns are adjusted to be the same. For example, they are all adjusted to 32*32, then 4 patterns of 32*32*3 are obtained. By splicing the 4 patterns, a pattern of 64*64*3 can be obtained. Calculate the RGB color mode information and HSV color space information of this pattern, and combine the RGB color mode information and HSV color space information of this pattern to obtain an input image of 64*64*6. The input image is the input of the first sub-network model, and the output of the first sub-network model is a color feature of 32*32*1.

[0105] S22. Input the color information of the facial image into the second sub-network model to obtain the gradient feature output by the second sub-network model. The gradient feature has the same size as the color feature, and the gradient feature is used to characterize the gradient information of the facial image.

[0106] It should be understood that if the facial image is regarded as a two-dimensional discrete function, then the gradient information is the derivative result of this two-dimensional discrete function. Optionally, in the embodiments of the present application, the gradient information of the facial image can be calculated by a sobel operator (Sobel operator). Technically, it is a discrete first-order difference operator used to calculate the approximate value of the first-order gradient of the image brightness function. Using this operator at any point in the image can generate the corresponding gradient vector at that point, that is, the gradient information.

[0107] Optionally, the second sub-network model can first perform 1*1 convolution on the gradient information, and then change the size of the convolution result to make it the same as the size of the color feature.

[0108] S23. Input the color feature and the gradient feature into the feature fusion layer to obtain the fusion feature output by the feature fusion layer. The fusion feature is used to characterize the result after the color feature and the gradient feature are fused.

[0109] In the process of obtaining the texture detection result by using the texture detection model in the embodiments of the present application, it is equivalent to using two inputs - gradient information and color information. It can enhance the texture difference between the live image and the non-live image at the places where the texture difference is large while extracting the texture by using the texture detection model, making the texture detection result output by the texture detection model more significant, and the features at the places where the texture difference is small interfere less with the classification result.

[0110] When the sizes of the color feature and the gradient feature are the same, the feature fusion layer can fuse the two features through a simple addition operation, and the result of the addition is the result after the fusion of the color feature and the gradient feature; optionally, in the embodiments of the present application, coefficients can also be added to the color feature and / or the gradient feature, and then the color feature and / or the gradient feature after adding the coefficients are fused. For example, a coefficient of 0.8 is added to the color feature, that is, each element in the color feature is multiplied by 0.8. It can be understood that if coefficients are added to the color feature and / or the gradient feature, the specific coefficients can be adjusted according to actual applications.

[0111] S24. Input the fused feature into the result output layer to obtain the texture detection result output by the result output layer.

[0112] In the embodiments of the present application, texture detection is regarded as a binary classification problem, and the texture detection result is used to characterize that the texture of the detected object is the texture of a live body or the texture of a non-live body. It can be understood that the result output layer includes a softmax layer, and the output of the softmax layer represents the relative probabilities between different categories. Further, the result output layer also includes several convolutional layers and a fully connected layer. The fused feature first undergoes convolutional processing by several convolutional layers to obtain a convolutional result, and the convolutional result passes through the fully connected layer to obtain a classification result, that is, the classification result that the texture of the detected object is the texture of a live body and the classification result that the texture of the detected object is the texture of a non-live body. Finally, the classification result is input into the softmax layer to obtain the relative probabilities of the two classification results.

[0113] It should be noted that through the live body detection methods of the above embodiments of the present application, most attack scenarios can be recognized. However, sometimes, if hackers use means such as high-definition videos / photos with high attack costs, silicone or 3D masks made of special materials for detection, the recognition rate of the live body detection methods in the above embodiments will be relatively low. Therefore, further, in the embodiments of the present application, blink detection and optical flow detection are used as detection processes after texture detection. Blink detection can effectively filter printed attack photos and screen attack photos. For video playback, since it is a secondary imaging of the video, the facial optical flow field changes in the front and rear frames are different from those of a real human face; for a 3D mask made of silicone or special materials, since there are no micro-expressions on the mask, the directions of the optical flow in the eye and mouth corners are different from those of a real human face. Based on this, in the embodiments of the present application, the local optical flow of the human face can be used to finally judge video playback and 3D masks. Through comprehensive filtering in three ways, various attack scenarios can be effectively recognized.

[0114] Figure 6 is a schematic flowchart of a live body detection method according to another embodiment of the present application. As Figure 6 shown, it includes:

[0115] S1. Obtain a frame of facial image of the object to be detected;

[0116] S2. Perform texture detection on the facial image to obtain a texture detection result;

[0117] S3. If the texture detection result of the facial image meets the first preset condition, obtain another frame of facial image of the object to be detected;

[0118] As can be seen from the above embodiments, the texture detection result is either the texture of the object to be detected being a live texture or the texture of the object to be detected being a non-live texture. When the result is that the texture of the object to be detected is a live texture, it is determined that the texture detection result of the facial image meets the first preset condition, and then another frame of facial image of the object to be detected is obtained. It can be understood that the embodiments of the present application can continuously collect facial images of the object to be detected through a camera.

[0119] As an alternative embodiment, the two frames of facial images are facial images collected at two non-adjacent moments within the same time period. Generally, a camera can collect 25 to 30 frames per second. If continuous images are used, the change in the optical flow field will not be obvious, the change in facial features cannot be shown, and the situation where the user does not blink because the time duration of two adjacent frames of facial images is too short cannot be avoided. Therefore, by selecting facial images collected at two non-adjacent moments, the embodiments of the present application can effectively avoid the above problems.

[0120] S4. Perform texture detection on the other frame of facial image. If the texture detection result of the other frame of facial image meets the first preset condition, perform blink detection on the two frames of facial images to obtain a blink detection result.

[0121] If the result of performing texture detection on the other frame of facial image also meets the first preset condition, it is necessary to perform blink detection on the two frames of facial images obtained in steps S1 and S3 to obtain a blink detection result. It can be understood that the blink detection result is used to represent whether the object to be detected blinks or not.

[0122] S5. If the blink detection result meets the second preset condition, perform optical flow field detection on the two frames of facial images, and determine whether the object to be detected is a live body according to the optical flow field detection result.

[0123] In the embodiments of the present application, the blink of the object to be detected is used as the second preset condition. If the blink detection result of the two frames of facial images meets the second preset condition, it is considered that the object to be detected has passed the blink test, and then optical flow field detection is continued on the two frames of facial images. Since the change situation of the optical flow field of a non-live body is different from that of a live body, therefore, it is determined whether the object to be detected is a live body according to the optical flow field detection result.

[0124] The existing method for blink detection is based on facial key points. Specifically, according to the positions of the upper and lower eyelid key points, the degree of eye opening and closing is calculated, and then whether a blink occurs is determined based on the degree of eye opening and closing. However, this method relies too much on key point detection. Once there is facial movement or the eyes are not very large, it is very easy to detect errors. To overcome the above problems, the embodiments of the present application adopt a neural network method to detect the picture of the area where the eyes are located, output the degree of eye opening and closing, and if the difference in the degree of eye opening and closing between two frames of images is greater than a preset threshold, it is considered that the detected object has performed a blink action.

[0125] Figure 7 This is a schematic flowchart of blink detection for two frames of facial images in the embodiments of the present application to obtain blink detection results, as Figure 7 shown, and this process includes:

[0126] S41. Respectively extract the areas where the eyes are located in the two frames of facial images to obtain two eye images.

[0127] The embodiments of the present application can adopt the method of the above embodiments. For a frame of facial image, a square constructed with the key points at the two eyes as the center is used as an eye image respectively.

[0128] S42. For each eye image, input the eye image into a pre-constructed blink detection model to obtain the degree of eye opening and closing output by the blink detection model.

[0129] It should be understood that before performing step S42, a blink detection model can also be pre-trained. Specifically, the blink detection model can be trained in the following way. First, collect a certain number of sample eye images, and determine the degree of eye opening and closing and the interval of the degree of eye opening and closing for each sample eye image, where the interval of the degree of eye opening and closing represents the numerical interval where the degree of eye opening and closing is located. For example, the value range of the degree of eye opening and closing is [0, 0.99], and [0, 0.99] is divided into 3 intervals, representing three categories, that is, [0, 0.33] represents closed, (0.33, 0.66] represents slightly open, and (0.66, 0.99] represents open. Then, based on the sample eye images, and the degree of eye opening and closing and the interval of the degree of eye opening and closing of the sample eye images, train the initial model to obtain the blink detection model. Among them, the initial model can be a single neural network model or a combination of multiple neural network models.

[0130] S43. If the difference between the two obtained degrees of eye opening and closing is greater than the preset threshold, generate a first blink detection result; if the magnitudes of the two obtained degrees of eye opening and closing are not greater than the preset threshold, generate a second blink detection result.

[0131] It should be understood that if the difference in the opening and closing degrees of the two eyes is greater than a preset threshold, it is considered that the detected object has blinked when sampling two frames of images, that is, the first blink detection result is used to represent that the detected object has blinked. If the obtained opening and closing degrees of the two eyes are not greater than the preset threshold, it is considered that the detected object has not blinked when sampling two frames of images, that is, the second blink detection result is used to represent that the detected object has not blinked.

[0132] Based on the above embodiments, as an optional embodiment, the blink detection model includes a feature extraction layer, an interval output layer, and an opening and closing degree output layer.

[0133] Figure 8 The flowchart of the process of inputting an eye image into a pre-constructed blink detection model in the embodiments of the present application to obtain the opening and closing degree of the eyes output by the blink detection model is shown in Figure 7 As shown, the process includes:

[0134] S421. Input the eye image into the feature extraction layer to obtain the eye features output by the feature extraction layer.

[0135] The feature extraction layer in the embodiments of the present application is used to perform feature encoding on the eye image and output eye features. The eye features can be the encoded eye image. Optionally, the feature extraction layer may include several (for example, 6) convolutional layers.

[0136] S422. Input the eye features into the interval output layer to obtain the opening and closing degree interval of the eyes of the eye image output by the interval output layer.

[0137] S423. Input the eye features and the opening and closing degree interval of the eyes of the eye image into the opening and closing degree output layer to obtain the opening and closing degree of the eyes obtained by the output layer through regression on the eye features within the opening and closing degree interval of the eyes of the eye image.

[0138] In the embodiments of the present application, when obtaining the opening and closing degree of the eyes, a method of classification first and then regression is adopted. The opening and closing degree interval output in step S422 is equivalent to obtaining the classification result of the opening and closing degree of the eyes. After obtaining the classification result, step S423 can determine the opening and closing degree of the eyes within a smaller interval, that is, regression. This method of classification first and then regression in the embodiments of the present application can obtain a more accurate result of the opening and closing degree of the eyes compared with the methods of only regression or only classification.

[0139] It can be seen from the above embodiments that for the video playback method, the changes in the front and rear frame facial optical flow fields are different from those of a real human face. For a 3D mask made of silicone or special materials, the direction of the optical flow in the eye and mouth corner areas is different from that of a real human face. Based on this, the embodiments of the present application determine whether the detected object is a living body according to the optical flow field detection result, including:

[0140] If optical flow appears in the area other than the face in two frames of facial images or the included angle of the optical flow directions of more than a preset proportion of pixel points in the facial area is less than a preset threshold, it is determined that the detected object is a non-living body.

[0141] If optical flow does not appear in the area other than the face in two frames of facial images or the included angle of the optical flow directions of less than a preset proportion of pixel points in the facial area is less than a preset threshold, it is determined that the detected object is a living body.

[0142] Generally, the camera for collecting portraits is always fixed. When a real portrait moves in front of the camera, optical flow will be generated in the portrait area, while there is no optical flow in the background. For video and paper attacks, the attacker holds a photo or an electronic device and moves in front of the camera. In addition to the optical flow around the portrait, the electronic device and the photo will also generate optical flow. Therefore, using this feature, the embodiments of the present application can effectively identify video and paper attacks that have not been filtered out before by determining whether optical flow appears in the area other than the face in two frames of facial images.

[0143] Moreover, when a person is unconscious, there will be some tiny movements in the eyes and the corners of the mouth, and these movements are non-rigid, which will cause the inconsistent directions of the optical flow movements of the pixels near the eyes and the corners of the mouth. Although a 3D mask can make rigid movements such as turning the head, raising the head, and lowering the head, it cannot make expressions. Therefore, the optical flow directions of the face usually remain consistent. Using this feature, the embodiments of the present application determine whether it is a 3D mask by calculating the optical flow directions of the local area.

[0144] Figure 9 It is a schematic structural diagram of the living body detection device provided by the embodiments of the present application, as Figure 9 shown, the living body detection device includes: a facial image acquisition module 101 and a texture detection module 102. Specifically:

[0145] The facial image acquisition module 101 is configured to acquire the facial image of the detected object;

[0146] The texture detection module 102 is configured to perform texture detection on the facial image to obtain a texture detection result;

[0147] Among them, the texture detection module 102 includes:

[0148] An information acquisition unit is configured to acquire the color information of the facial image;

[0149] A detection unit is configured to input the color information of the facial image into a pre-trained texture detection model to obtain the texture detection result output by the texture detection model;

[0150] The color information includes RGB color model information and HSV color space information; the texture detection model is trained with the color information of the sample facial images as the training samples and the texture detection results of the sample facial images as the training labels; the texture detection results are used to characterize whether the texture of the detected object is a texture of a live body or a texture of a non-live body.

[0151] The live body detection device provided by the embodiments of the present application specifically executes the processes of the above method embodiments. For specific details, please refer to the content of the above live body detection method embodiments, which will not be elaborated here. The input of the live body detection device provided by the embodiments of the present application is only the facial image of the detected object, without the need for additional interaction with the detected object. Therefore, no special hardware is required, which can reduce the cost of live body detection. By using both the RGB color model and the HSV color space as color information, the texture differences in the facial image can be captured more significantly. The embodiments of the present invention can identify common photos, videos, or 3D masks made of hard materials such as plastic with attack traces (blurring, water ripples, local reflection, etc.).

[0152] In the embodiments of the present application, an electronic device is provided. The electronic device includes: a memory and a processor; at least one program is stored in the memory and, when executed by the processor, enables the processor to execute the corresponding content in the foregoing method embodiments. Compared with the prior art, it can be realized that the input of the program is only the facial image of the detected object, without the need for additional interaction with the detected object. Therefore, no special hardware is required, which can reduce the cost of live body detection. By using both the RGB color model and the HSV color space as color information, the texture differences in the facial image can be captured more significantly. The embodiments of the present invention can identify common photos, videos, or 3D masks made of hard materials such as plastic with attack traces (blurring, water ripples, local reflection, etc.).

[0153] In an optional embodiment, an electronic device is provided, as Figure 10 shown. Figure 10 As shown, the electronic device 4000 includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.

[0154] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0155] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0156] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0157] The memory 4003 is used to store the application program code for executing the solution of this application, and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the application program code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0158] An embodiment of this application provides a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments. Compared with the prior art, the input of the computer program is only the facial image of the object to be detected, and no additional interaction with the object to be detected is required. Therefore, no special hardware is needed, and the cost of liveness detection can be reduced. By using the RGB color mode and the HSV color space together as color information, the texture differences of facial images can be captured more significantly. The embodiments of the present invention can identify common photos, videos, or 3D masks made of hard materials such as plastics with attack marks (blurring, water ripples, local reflection, etc.).

[0159] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0160] The above are only some implementation manners of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A method for detecting a living body, characterized in that, Including: Obtain a facial image of the object to be detected; Perform texture detection on the facial image to obtain a texture detection result; Among them, the performing texture detection on the facial image to obtain a texture detection result includes: Obtain the color information of the facial image, input the color information of the facial image into a pre-trained texture detection model, and obtain the texture detection result output by the texture detection model; The color information includes RGB color mode information and HSV color space information; the texture detection model is trained with the color information of the sample facial image as the training sample and the texture detection result of the sample facial image as the training label; the texture detection result is used to characterize that the texture of the object to be detected is the texture of a living body or the texture of a non-living body; The obtaining a facial image of the object to be detected includes: Obtain an original image, where the original image includes the face of the object to be detected; Determine at least four key points in the original image, and the at least four key points are respectively located at the two eyes and the two corners of the mouth of the object to be detected; Construct a square with each of the at least four key points as the center, and each of the at least four constructed squares does not overlap with any other square and has the largest area; Stitch the areas of the at least four squares in the original image to obtain the facial image; The constructing a square with each of the at least four key points as the center includes: Calculate the distance from the key point at the left eye to the left border of the original image, denoted as the first distance; calculate the distance from the key point at the right eye to the right border of the original image, denoted as the second distance; calculate the distance from the key point at the left corner of the mouth to the left border of the original image, denoted as the third distance; calculate the distance from the key point at the right corner of the mouth to the right border of the original image, denoted as the fourth distance; calculate the horizontal distance between the key points at the two eyes, denoted as the fifth distance; calculate the horizontal distance between the key points at the two corners of the mouth, denoted as the sixth distance; calculate the vertical distance from the center of the key points at the two eyes to the center of the key points at the two corners of the mouth, denoted as the seventh distance; Then the side length of the square with the key point at the left eye as the center is the minimum value among the fifth distance, the seventh distance, and twice the first distance; the side length of the square with the key point at the right eye as the center is the minimum value among the fifth distance, the seventh distance, and twice the second distance; the side length of the square with the key point at the left corner of the mouth as the center is the minimum value among the sixth distance, the seventh distance, and twice the third distance; the side length of the square with the key point at the right corner of the mouth as the center is the minimum value among the sixth distance, the seventh distance, and twice the fourth distance.

2. The method for detecting a living body according to claim 1, characterized in that, The texture detection model includes a first sub-network model, a second sub-network model, a feature fusion layer, and a result output layer; The inputting the color information of the facial image into a pre-trained texture detection model to obtain the texture detection result output by the texture detection model includes: Input the color information of the facial image into the first sub-network model to obtain the color features output after the first sub-network model performs multiple convolutional and pooling operations on the color information; Input the color information of the facial image into the second sub-network model to obtain the gradient features output by the second sub-network model. The gradient features have the same size as the color features and are used to characterize the gradient information of the facial image; Input the color features and gradient features into the feature fusion layer to obtain the fusion features output by the feature fusion layer. The fusion features are used to characterize the result after the color features and gradient features are fused; Input the fusion features into the result output layer to obtain the texture detection result output by the result output layer.

3. The method for detecting a living body according to claim 1 or 2, characterized in that, It further includes: If the texture detection result of the facial image meets the first preset condition, obtain another frame of facial image of the detected object; Perform texture detection on the other frame of facial image. If the texture detection result of the other frame of facial image meets the first preset condition, perform blink detection on the two frames of facial images to obtain a blink detection result; If the blink detection result meets the second preset condition, perform optical flow field detection on the two frames of facial images, and judge whether the detected object is a live body according to the optical flow field detection result.

4. The method for detecting a living body according to claim 3, characterized in that, The performing blink detection on the two frames of facial images to obtain a blink detection result includes: Extract the regions where the eyes are located in the two frames of facial images respectively to obtain two eye images; For each eye image, input the eye image into a pre-constructed blink detection model to obtain the eye opening degree output by the blink detection model; If the difference between the two obtained eye opening degrees is greater than a preset threshold, generate a first blink detection result; if the magnitudes of the two obtained eye opening degrees are not greater than the preset threshold, generate a second blink detection result. Among them, the blink detection model is trained with sample eye images as training samples and the eye opening degree and eye opening degree interval of the sample eye images as training labels; the first blink detection result is used to characterize that the detected object blinks; the second blink detection result is used to characterize that the detected object does not blink.

5. The method for detecting a living body according to claim 4, characterized in that, The blink detection model includes a feature extraction layer, an interval output layer, and an opening degree output layer; The inputting the eye image into a pre-constructed blink detection model to obtain the eye opening degree output by the blink detection model includes: Input the eye image into the feature extraction layer to obtain the eye features output by the feature extraction layer; Input the eye features into the interval output layer to obtain the eye opening degree interval of the eye image output by the interval output layer; Input the eye features and the eye opening degree interval of the eye image into the opening degree output layer to obtain the eye opening degree obtained by the output layer through regression of the eye features within the eye opening degree interval of the eye image.

6. The living body detection method according to claim 3, wherein The two frames of facial images are facial images collected at two non-adjacent moments within the same time period.

7. The living body detection method according to claim 3, wherein The judging whether the detected object is a live body according to the optical flow field detection result includes: If optical flow appears in the area other than the face in two frames of facial images or the included angle of the optical flow directions of more than a preset proportion of pixel points in the facial area is less than a preset threshold, it is determined that the detected object is a non-living body; If optical flow does not appear in the area other than the face in two frames of facial images or the included angle of the optical flow directions of less than a preset proportion of pixel points in the facial area is less than a preset threshold, it is determined that the detected object is a living body.

8. A living body detection device, wherein Including: A facial image acquisition module for acquiring the facial image of the detected object; A texture detection module for performing texture detection on the facial image to obtain a texture detection result; Among them, the texture detection module includes: An information acquisition unit for acquiring the color information of the facial image; A detection unit for inputting the color information of the facial image into a pre-trained texture detection model to obtain the texture detection result output by the texture detection model; The color information includes RGB color mode information and HSV color space information; the texture detection model is trained with the color information of sample facial images as training samples and the texture detection results of the sample facial images as training labels; the texture detection result is used to characterize the texture of the detected object as the texture of a living body or the texture of a non-living body; The acquiring the facial image of the detected object includes: Acquiring an original image, where the original image includes the face of the detected object; Determining at least four key points in the original image, and the at least four key points are respectively located at the two eyes and the two corners of the mouth of the detected object; Constructing squares with the at least four key points as the centers respectively, and each of the at least four constructed squares does not overlap with any other square and has the maximum area; Stitching the areas of the at least four squares in the original image to obtain the facial image; The constructing squares with the at least four key points as the centers respectively includes: Calculating the distance from the key point at the left eye to the left border of the original image, denoted as the first distance; calculating the distance from the key point at the right eye to the right border of the original image, denoted as the second distance; calculating the distance from the key point at the left corner of the mouth to the left border of the original image, denoted as the third distance; calculating the distance from the key point at the right corner of the mouth to the right border of the original image, denoted as the fourth distance; calculating the horizontal distance between the key points at the two eyes, denoted as the fifth distance; calculating the horizontal distance between the key points at the two corners of the mouth, denoted as the sixth distance; calculating the vertical distance from the center of the key points at the two eyes to the center of the key points at the two corners of the mouth, denoted as the seventh distance; The side length of the square centered at the key point on the left eye is the minimum of the fifth distance, the seventh distance, and twice the first distance; the side length of the square centered at the key point on the right eye is the minimum of the fifth distance, the seventh distance, and twice the second distance; the side length of the square centered at the key point on the left corner of the mouth is the minimum of the sixth distance, the seventh distance, and twice the third distance; the side length of the square centered at the key point on the right corner of the mouth is the minimum of the sixth distance, the seventh distance, and twice the fourth distance.

9. An electronic device, wherein It includes a memory and a processor, and the processor and the memory complete mutual communication through a bus; the memory stores program instructions executable by the processor, and the processor can execute the method according to any one of claims 1 to 7 by invoking the program instructions.

10. A computer-readable storage medium having a computer program stored thereon, wherein When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Human face living body detection method, system and device

    CN108875618A

  • Method, apparatus, computer device, and storage medium for ticket purchase authentication processing

    CN109493159A

  • A human face living body detection method based on local color texture characteristics

    CN109740572A

  • Living body detection method and device

    CN111274851A

  • Human face living body detection method and device

    CN111652082A