Living body detection method, device, equipment and medium
By analyzing the distance change trend between the nose tip and the face horizontal line in the face image frame, combined with the preset change trend, the problem of low accuracy of live detection in the prior art is solved, and higher accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202311819545.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has low accuracy in facial recognition detection, especially when the user's cooperation is not high or facial expression changes, which can easily lead to misrecognition.
By obtaining the position information and identification of the face key points of the image frame of the video to be detected, the first judgment value of the face in each image frame, that is, the distance between the nose tip and the face horizontal line, is determined, and the change trend is determined based on the multiple first judgment values, and the live detection result is determined in combination with the preset change trend.
It improves the accuracy and robustness of live detection, enhances the ability to identify live organs, and reduces the rate of misidentification.
Smart Images

Figure CN120220255A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and medium for live detection. Background Art
[0002] With the popularization of Internet applications and the prevalence of mobile payment, more and more scenarios require user identity verification, which can be achieved through methods such as username-password combinations, fingerprint recognition, face recognition, iris recognition, etc. Among them, during the face recognition process, the user is usually verified as a live body rather than a static picture or 3D model by interacting with the user during verification, such as prompting the user to blink, open the mouth, nod the head, etc., thereby effectively improving the security level of the system. Summary of the Invention
[0003] Embodiments of this application provide a method, apparatus, device, and medium for live detection to improve the accuracy of live detection.
[0004] In a first aspect, this application provides a method for live detection, the method including:
[0005] Obtaining the position information and identification of face key points of an image frame of a video to be detected;
[0006] Determining a first judgment value of the face in each image frame according to the position information and identification of the face key points of each image frame, where the first judgment value is the distance between the tip of the nose and the face horizontal line;
[0007] Determining a first change trend of the first judgment value according to a plurality of the first judgment values;
[0008] Determining a live detection result of the face in the video to be detected according to the first change trend and a first preset change trend.
[0009] In a second aspect, this application also provides a live detection apparatus, the apparatus including:
[0010] An obtaining module, configured to obtain the position information and identification of face key points of an image frame of a video to be detected;
[0011] A determining module, configured to determine a first judgment value of the face in each image frame according to the position information and identification of the face key points of each image frame, where the first judgment value is the distance between the tip of the nose and the face horizontal line; and determining a first change trend of the first judgment value according to a plurality of the first judgment values;
[0012] A detection module, configured to determine a live detection result of the face in the video to be detected according to the first change trend and a first preset change trend.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, which at least includes a processor and a memory. When the processor executes a computer program stored in the memory, it implements the steps of the living body detection method described in any one of the above.
[0014] In a fourth aspect, an embodiment of the present application provides a computer storage medium, which stores a computer program executable by an electronic device. When the program runs on the electronic device, it causes the electronic device to execute the steps of the living body detection method described in any one of the above.
[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes: computer program code. When the computer program code runs on the electronic device, it causes the electronic device to execute the steps of the living body detection method described in any one of the above.
[0016] In the embodiment of the present application, during living body detection, the position information and identification of the face key points of the image frames of the video to be detected are obtained. Since the nose tip is an obvious and prominent key point on the face, and the position where the nose tip is located is not easily changed by the facial expression, therefore, according to the position information and identification of the face key points of each image frame, the first judgment value of the face in each image frame is determined, that is, the distance between the nose tip of the face in each image frame and the face horizontal line is determined. And according to a plurality of first judgment values, the first change trend of the first judgment value is determined. Thus, according to the first change trend and the first preset change trend, the living body detection result of the face in the video to be detected is determined. That is, the first judgment value is determined through the obvious face key points whose position information is not easily changed, and then the living body detection result is determined based on the first judgment value, enhancing the robustness of the living body detection, thereby improving the accuracy of the living body detection. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic diagram of a living body detection process provided by an embodiment of the present application;
[0019] Figure 2 It is a schematic diagram of face key points detected by Mediapipe provided by an embodiment of the present application;
[0020] Figure 3 It is a schematic diagram of a face horizontal line provided by an embodiment of the present application;
[0021] Figure 4a Schematic diagram of a first judgment value in a human face provided by an embodiment of the present application;
[0022] Figure 4b Schematic diagram of a first judgment value in a human face provided by an embodiment of the present application;
[0023] Figure 4c Schematic diagram of a first judgment value in a human face provided by an embodiment of the present application;
[0024] Figure 5a Schematic diagram of a human face when looking straight ahead provided by an embodiment of the present application;
[0025] Figure 5b Schematic diagram of a human face when looking down provided by an embodiment of the present application;
[0026] Figure 5c Schematic diagram of a human face when looking up provided by an embodiment of the present application;
[0027] Figure 6 Schematic diagram of a human face provided by an embodiment of the present application;
[0028] Figure 7a Schematic diagram of a normal human face provided by an embodiment of the present application;
[0029] Figure 7b Schematic diagram of an abnormal human face provided by an embodiment of the present application;
[0030] Figure 7c Schematic diagram of an abnormal human face provided by an embodiment of the present application;
[0031] Figure 7d Schematic diagram of an abnormal human face provided by an embodiment of the present application;
[0032] Figure 8 Schematic diagram of a human face provided by an embodiment of the present application;
[0033] Figure 9 Schematic diagram of a human face provided by an embodiment of the present application;
[0034] Figure 10 Schematic diagram of a human face provided by an embodiment of the present application;
[0035] Figure 11 Schematic diagram of the process of live detection provided by an embodiment of the present application;
[0036] Figure 12 Schematic diagram of the process of live detection provided by an embodiment of the present application;
[0037] Figure 13 Schematic diagram of the structure of a live detection device provided by an embodiment of the present application;
[0038] Figure 14 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0039] To make the objectives and implementation manners of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0040] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the implementation manners described hereinafter, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and general meanings.
[0041] In the present application, the terms "first", "second", "third", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0042] The terms "comprising" and "having" and any variations thereof are intended to cover but not be exclusive of inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0043] The term "module" refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code that can perform functions related to the element.
[0044] Finally, it should be noted that: the embodiments of the present application are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0045] For ease of explanation, the above description has been made in conjunction with specific embodiments. However, the exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed. According to the derivations in each embodiment, various modifications and variations can be obtained. The selection and description of the embodiments in this application are for better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and the various different modified embodiments suitable for specific use considerations.
[0046] An embodiment of this application provides a live detection method, device, equipment and medium. In this method, the position information and identification of the facial key points of the image frames of the video to be detected are obtained; according to the position information and identification of the facial key points of each image frame, a first judgment value of the face in each image frame is determined, and the first judgment value is the distance between the tip of the nose and the horizontal line of the face; according to a plurality of first judgment values, a first change trend of the first judgment values is determined; according to the first change trend and a first preset change trend, a live detection result of the face in the video to be detected is determined.
[0047] The live detection method provided by the embodiment of this application can be applied to any scenario that requires live detection, and this method can be applied to products that detect the live body in these scenarios.
[0048] To better understand the live detection method provided by the embodiment of this application, the application environment applicable to the embodiment of this application will be described below. The live detection method provided by the embodiment of this application can be executed by an electronic device such as a terminal device or a server. The terminal device can be a vehicle-mounted device, a user equipment (UE), a mobile device, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a wearable device, etc. The method can be implemented by the processor calling the computer-readable program instructions stored in the memory. The server can be an independent physical server, a server cluster composed of multiple servers, or a cloud server capable of performing cloud computing.
[0049] For the detection of the action of "nodding", there are usually two detection schemes. One is to judge whether the change trend of the ratio of the longitudinal width of the forehead to the longitudinal width of the lower jaw conforms to the logic of nodding. The other scheme is to directly detect the nodding action using a trained action recognition network, which usually needs to be implemented by a 3D convolutional neural network. Comparing the above two detection schemes, the detection using a trained action recognition network has higher robustness, but much lower efficiency. Currently, it is almost impossible to be deployed on an electronic device or to achieve real-time operation. And for the scheme of detecting through the change trend of the ratio of the longitudinal width of the forehead to the longitudinal width of the lower jaw, although it is convenient to be deployed on an electronic device and can achieve real-time operation, this detection scheme is not stable. Since there are no obvious prominent key points in the forehead area and the lower jaw area of the human face, it is very easy to have the situation that the position information of the human face key points is misrecognized, thus affecting the accuracy of the live detection. And when the cooperation degree of the user to be detected for live detection is not high, the accuracy of this detection scheme will also be greatly reduced. For example, when the user makes a "pouting" or "raising eyebrows" action during the process of "nodding", because the muscles in the forehead area and the lower jaw area of the human face are relatively flexible, then, even if the user does not perform the nodding action, as long as the face makes some special expressions, it will also cause the position information of the human face key points in the recognized forehead area and the lower jaw area to change, and there is a certain probability that the change trend of the ratio of the longitudinal width of the forehead to the longitudinal width of the lower jaw conforms to the logic of nodding, thus affecting the detection result. Therefore, in order to enhance the robustness of the live detection and improve the accuracy of the live detection in this application, the embodiments of this application provide a new live detection method, and the above live detection method will be described below in combination with each embodiment.
[0050] Figure 1 It is a schematic diagram of a live detection process provided by an embodiment of this application, and this process includes:
[0051] S101: Obtain the position information and identification of the human face key points of the image frames of the video to be detected.
[0052] A live detection method provided by an embodiment of this application is applied to an electronic device, and this electronic device can be a server, a PC, a smart terminal, etc.
[0053] In the embodiments of this application, a video to be detected can be obtained, and the position information of each human face key point included in each image frame in the video to be detected can be determined. Subsequently, based on this position information, the live detection of the human face can be completed. Among them, the video to be detected can be uploaded by the user of the electronic device, or can be a video collected in real time by the image acquisition module of the electronic device. The embodiments of this application do not limit the acquisition form of the video to be detected.
[0054] In the embodiment of the present application, when obtaining the position information of the key points of the face of any image frame in the video to be detected, the image frame can be input into the recognition model, and the recognition model detects the key points of the face in the received image frame, thereby obtaining the position information of each key point of the face and the corresponding identifier. Among them, the recognition model can be Mediapipe or other key point detection models of the face. The embodiment of the present application does not limit the recognition model, and those skilled in the art can select a suitable recognition model as needed.
[0055] The position information output by the recognition model may be the coordinates of each key point of the face in the image pixel coordinate system, which coordinates usually include at least x and y coordinates, wherein the x coordinate represents the horizontal offset in the image pixel coordinate system, and y represents the vertical offset. In an embodiment of the present application, in the image pixel coordinate system of the image frame, the origin may be the upper left corner of the image frame, so the values of x and y are non-negative values, i.e., x and y ≥ 0. It should be noted that although some recognition models can identify the three-dimensional coordinates of each key point of the face in the face, due to the relatively low accuracy of the coordinates in the depth direction, in an embodiment of the present application, the position information of each key point of the face obtained may be a two-dimensional coordinate. When the position information determined by the recognition model is a three-dimensional coordinate, the x and y coordinates in the three-dimensional coordinate may be directly selected as the position information required by the embodiment of the present application. Of course, those skilled in the art may also use the three-dimensional coordinates as the position information for subsequent liveness detection according to actual needs.
[0056] Figure 2 A schematic diagram of a facial key point detected by Mediapipe provided in an embodiment of the present application. Mediapipe can detect the following: Figure 2 The key points of each face shown in Figure 2 The position of each number in represents the key point of each face. Figure 2 It can be seen that Mediapipe can recognize a variety of facial key points in the face, and Mediapipe has the advantages of stability, fast speed and low memory consumption.
[0057] It should be noted that the above-mentioned scheme for obtaining the location information and identification of facial key points is only an example given for ease of understanding, but is not limited to the above-mentioned scheme. Technical personnel in this field can configure the scheme for obtaining the location information and identification of facial key points as needed.
[0058] S102: Determine a first judgment value of the face in each image frame according to the position information and identification of the facial key points in each image frame, where the first judgment value is the distance between the nose tip and the horizontal line of the face.
[0059] Which facial key points in the face are specifically obtained is pre-configured. After obtaining the position information and identification of each facial key point, the first judgment value in the face can be determined according to each position information and the corresponding identification. Research has found that the muscles near the tip of the nose and the bridge of the nose in the face are relatively underdeveloped. For example, the average person cannot control the height of the nose tip. When a person controls the movement of the nose, generally only the nostrils can be enlarged or reduced, and during the movement of the nose, the position of the nose tip generally does not change, and the nose tip is a relatively prominent key point in the face and is easy to identify. Therefore, in the embodiment of the present application, the first judgment value can be the distance between the tip of the nose and the horizontal line of the face. In the embodiment of the present application, the horizontal line of the face can be any horizontal line in the face. For example, the eye corners of the face are generally on a horizontal line, and the eye corners are easy to identify, and the position information of the eye corners is not easy to change. Therefore, in the embodiment of the present application, the electronic device can determine a straight line according to the facial key points corresponding to the two eye corners in the face, and the determined straight line is the horizontal line of the face.
[0060] In the embodiment of the present application, which two facial key points are used to determine the horizontal line of the face can be pre-configured by relevant personnel. That is to say, the electronic device pre-saves the identification of the facial key points for determining the horizontal line of the face, and the identification of the facial key points is generally saved in pairs. Of course, the electronic device can also pre-save multiple groups of identifications of the facial key points for determining the horizontal line of the face. When performing live detection, the electronic device can randomly select a group of identifications of the facial key points from the multiple groups of identifications of the facial key points, and then determine the horizontal line of the face according to the position information corresponding to the randomly selected target identification. That is to say, in the embodiment of the present application, when performing live detection, the selection of the horizontal line of the face can be random, and malicious users cannot predict in advance which horizontal line of the face is specifically based on to determine the first judgment value in the actual detection process. It should be noted that when configuring the identification of the facial key points for determining the horizontal line of the face, obvious and prominent facial key points should be selected as much as possible, and the positions of the selected facial key points are not easy to change with the facial expressions.
[0061] S103: Determine the first change trend of the first judgment value according to the multiple first judgment values.
[0062] After obtaining the first judgment value of the human face in each image frame, the first change trend of the first judgment value can be determined according to the first judgment value corresponding to each image frame in the video to be detected. Among them, the first change trend can be understood as the change rule of the first judgment value of the human face in the image frame of the video to be detected. In this embodiment, the first change trend of the first judgment value can be determined according to multiple first judgment values. In some embodiments, the change trend of the first judgment value of the human face in all image frames of the video to be detected can be used. For example, according to the order in which each image frame appears in the video to be detected, the first judgment values corresponding to each image frame are sorted in this order to obtain a first judgment value sequence. After obtaining this first judgment value sequence, the change rule of the size of the first judgment value in the first judgment value sequence can be statistically analyzed. For example, if the first judgment value first increases, then decreases, then increases, then decreases, then it can be determined that the first change trend of the first judgment value is first to increase, then to decrease, then to increase, then to decrease.
[0063] S104: Determine the live detection result of the human face in the video to be detected according to the first change trend and the first preset change trend.
[0064] After determining the first change trend of the first judgment value, the live detection result of the human face in the video to be detected can be determined according to this first change trend and the first preset change trend. For example, when the first change trend meets the first preset change trend, it can be determined that the human face in the video to be detected is a live human face; when the first change trend does not meet the first preset change trend, it can be determined that the human face in the video to be detected is a non-live human face. Among them, the first preset change trend is pre-configured by relevant personnel. The first preset change trend can be first to increase, then to decrease, then to increase, or it can be first to decrease, then to increase, then to decrease. The embodiments of the present application do not limit the first preset change trend. However, it should be noted that due to the different positions of the face horizontal line, the change trend of the distance between the tip of the nose and the face horizontal line is also different. Therefore, the first preset change trend should be determined by relevant technical personnel after analyzing the change trend of the distance between the actually used face horizontal line and the tip of the nose.
[0065] In the embodiment of the present application during live detection, the position information and identification of the facial key points of the image frames of the video to be detected are obtained. Since the nose tip is an obvious and prominent key point on the face, and the position where the nose tip is located is not easily changed by facial expressions, the first judgment value of the face in each image frame is determined according to the position information and identification of the facial key points of each image frame, that is, the distance between the nose tip of the face and the face horizontal line in each image frame is determined. Then, according to multiple first judgment values, the first change trend of the first judgment value is determined. Thus, according to the first change trend and the first preset change trend, the live detection result of the face in the video to be detected is determined. That is, the first judgment value is determined through obvious facial key points with unchangeable position information, and then the live detection result is determined based on the first judgment value, enhancing the robustness of the live detection and thus improving the accuracy of the live detection.
[0066] In order to improve the accuracy of live detection, on the basis of the above embodiment, in the embodiment of the present application, after determining the first judgment value of the face in each image frame and before determining the first change trend of the first judgment value according to multiple first judgment values, the method further includes:
[0067] According to the positional relationship between the nose tip and the face horizontal line in each image frame, a direction symbol is added to the first judgment value of the face in each image frame.
[0068] Since the "nodding" process includes two parts: lowering the head and raising the head. When lowering the head, the nose tip is generally located below the face horizontal line, while when raising the head, there may be a situation where the nose tip is located above the face horizontal line. In some possible cases, the distance between the nose tip and the face horizontal line when the nose tip is located below the face horizontal line may be the same as the distance between the nose tip and the face horizontal line at a certain moment when the nose tip is located above the face horizontal line. Therefore, in order to facilitate determining the swinging state of the face during the "nodding" process, in the embodiment of the present application, a direction symbol can be added to the first judgment value of the face in each image. Since the face horizontal line is determined according to a pre-configured set of facial key points, the position of the face horizontal line is known, and the position information corresponding to the nose tip is also known. Thus, it can be determined whether the nose tip is located below or above the face horizontal line in the corresponding image frame. When it is determined that the nose tip in a certain image frame is located above the face horizontal line, a first direction symbol can be added to the first judgment value of the face in this image frame; when it is determined that the nose tip in a certain image frame is located below the face horizontal line, a second direction symbol can be added to the first judgment value of the face in this image frame; when it is determined that the nose tip in a certain image frame is exactly on the face horizontal line, no direction symbol is added to the first judgment value of the face in this image frame. It should be noted that in the embodiment of the present application, as long as the first direction symbol and the second direction symbol can represent different directions, those skilled in the art can configure them as needed.
[0069] To facilitate determining the first change trend, based on the above embodiments, in the embodiments of the present application, the face horizontal line is determined according to the face key points located on the same horizontal line of the human face, and the face horizontal line divides the human face in each image frame into a first region and a second region; adding a direction symbol to the first judgment value of the human face in each image frame according to the positional relationship between the tip of the nose and the face horizontal line in each image frame includes:
[0070] For each image frame, if the tip of the nose in the image frame is located in the first region, add a first direction symbol to the first judgment value of the human face in the image frame; if the tip of the nose in the image frame is located in the second region, add a second direction symbol to the first judgment value of the human face in the image frame, where the first direction symbol and the second direction symbol are in opposite directions.
[0071] To facilitate determining the first change trend based on multiple first judgment values to further improve the accuracy of liveness detection, in the embodiments of the present application, the face horizontal line can be determined according to the face key points located on the same horizontal line of the human face. That is to say, the face horizontal line can be a straight line connecting two pre-set face key points. In the embodiments of the present application, the face key points on the same horizontal line can be pre-configured. Exemplarily, the face key points corresponding to the positions of the zygomatic bones on the left and right sides of the human face can be determined as the face key points on the same horizontal line, and then the determined face horizontal line is the connection line between the face key points corresponding to the positions of the zygomatic bones on the left and right sides of the human face. Figure 3 A schematic diagram of a face horizontal line provided for the embodiments of the present application, according to Figure 3 it can be known that the face key points labeled 234 and 454 on the left and right sides of the human face are located on the same horizontal line, and then the straight line formed by the face key points corresponding to the labels 234 and 454 is the face horizontal line. The determined face horizontal line divides the human face in the corresponding image frame into a first region and a second region. For the convenience of distinction, in the embodiments of the present application, the first region can be the region including the forehead, and the second region can be the region including the mandible.
[0072] When adding a direction symbol to the first judgment value of the human face in each image frame, for each image frame, if the tip of the nose in the image frame is located in the first region, it indicates that the human face in the image frame is in a head-up state, and a first direction symbol can be added to the first judgment value corresponding to the image frame. If the tip of the nose in the image frame is located in the second region, it indicates that the person in the image frame is in a head-down state or is in the process of changing from a head-up state to a head-down state, and a second direction symbol can be added to the first judgment value corresponding to the image frame. Among them, the first direction symbol and the second direction symbol are in opposite directions.
[0073] Specifically, assume that the first direction symbol is a negative sign, i.e., the first direction symbol is "-", and the second direction symbol is a positive sign, i.e., the second direction symbol is "+". The first judgment value corresponding to the image frame A in the video to be detected is 2, and the first judgment value corresponding to the image frame B is also 2. The tip of the nose in the image frame A is located in the first region, and the tip of the nose in the image frame B is located in the second region. Then, after adding the first direction symbol to the first judgment value 2 corresponding to the image frame A, the first judgment value corresponding to the image frame A can be obtained as -2; after adding the second direction symbol to the first judgment value 2 corresponding to the image frame B, the first judgment value corresponding to the image frame A can be obtained as +2.
[0074] The following describes the determination process of the first judgment value in combination with a specific embodiment. Figure 4a This is a schematic diagram of the first judgment value in a human face provided by an embodiment of the present application. The face horizontal line D is determined according to the face key point identifiers 234 and 454. The coordinates of the midpoint of the face horizontal line D can be expressed as center(x, y), where x = (234.x + 454.x) / 2; y = (234.y + 454.y) / 2. Here, 234.x represents the x-axis coordinate of the face key point with the identifier 234, 454.x represents the x-axis coordinate of the face key point with the identifier 454, 234.y represents the y-axis coordinate of the face key point with the identifier 234, and 454.y represents the y-axis coordinate of the face key point with the identifier 454. Assume that the identifier corresponding to the nose tip, a human face key point, is pre-configured as 1. Then, after determining center(x, y), the Euclidean distance can be used to calculate the distance between center(x, y) and the face key point with the identifier 1. At this time, this distance is always a positive number. Next, it is necessary to determine whether the nose tip is located in the first region or the second region. When the nose tip is located in the first region, this distance is negative, and when the nose tip is located in the second region, this distance is positive. During the "nodding" process, this distance, that is, the first judgment value, generally first increases and then decreases. Refer to Figure 4a As shown, when the face is facing forward, the nose tip is located in the second region, that is, the nose tip is below the face horizontal line. At this time, the first judgment value is positive. Figure 4b This is a schematic diagram of the first judgment value in a human face provided by an embodiment of the present application. Figure 4b As shown in, the face is in a downward-looking state. At this time, the nose tip is located in the second region, that is, the nose tip is below the face horizontal line. At this time, the first judgment value is positive. Figure 4c This is a schematic diagram of the first judgment value in a human face provided by an embodiment of the present application. Figure 4c As shown in, the face is in a looking-up state. At this time, the nose tip is located in the first region, that is, the nose tip is above the face horizontal line. At this time, the first judgment value is negative.
[0075] According to the analysis, the first judgment value obtained by the direction symbol adding method provided in the embodiments of the present application has the same first change trend as that of the first judgment value without the added direction symbol during the nodding process, that is, it first increases, then decreases, and then increases again.
[0076] To further improve the accuracy of live detection, based on the above embodiments, in the embodiments of the present application, determining the first change trend of the first judgment value according to the first judgment value includes:
[0077] Obtain a preset number of target image frames in the video to be detected, and determine the first change trend of the first judgment values corresponding to the preset number of target image frames.
[0078] Since the user needs some time to react after receiving the "nod your head" instruction during the live detection process, that is, the "nod your head" action will not be performed immediately. Therefore, the video to be detected may also include other redundant actions, that is, the video to be detected may also include non-"nod your head" actions. Therefore, in the embodiments of the present application, a sliding window can be set. The sliding window is used to select a preset number of target image frames in the video to be detected. When the first change trend of the target image frames selected by any sliding window in the video to be detected meets the first preset change trend, it can be considered that the face in the video to be detected is a live face. That is to say, in the embodiments of the present application, when determining the first change trend of the first judgment value according to the first judgment value corresponding to each image frame included in the video to be detected, a preset number of target image frames can be obtained in the video to be detected, and the first change trend of the first judgment values corresponding to the preset number of target image frames can be determined.
[0079] Specifically, a sliding window for selecting 10 image frames is preset in advance. Then, during the live detection, the first change trend of the first judgment values corresponding to the 1st - 10th target image frames in the video to be detected can be determined, then the first change trend of the first judgment values corresponding to the 2nd - 11th target image frames in the video to be detected can be determined, and then the first change trend of the first judgment values corresponding to the 3rd - 12th target image frames in the video to be detected can be determined, and so on.
[0080] It should be noted that the embodiments of the present application do not limit the size of the sliding window, that is, do not limit the preset number. Those skilled in the art can configure it according to needs. Exemplarily, when the frame per second (FPS) of the video is 25, the normal nodding action can be completed in 4 - 8 seconds. Therefore, the size of the sliding window can be adjusted according to the FPS. For example, the preset number is 8 * 25 = 200 frames.
[0081] To further improve the accuracy of live detection, based on the above embodiments, in the embodiments of the present application, before determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend, the method further includes:
[0082] If the swing amplitude and / or swing speed of the face in the image frame of the video to be detected meet the preset swing standard, then perform the step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend.
[0083] To further improve the accuracy of live detection, before determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend, it is also possible to determine whether the swing amplitude and / or swing speed of the face in the image frame of the video to be detected meet the preset swing standard. If the swing amplitude and / or swing speed of the face in the image frame of the video to be detected do not meet the preset swing standard, it may affect the accuracy of live detection. Therefore, a prompt message for increasing the swing amplitude and / or adjusting the swing speed can be output. If the swing amplitude and / or swing speed of the face in the image frame of the video to be detected meet the preset swing standard, then the subsequent step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend can be continued.
[0084] Specifically, when determining the swing amplitude of the face, the maximum judgment value and the minimum judgment value in the first judgment value corresponding to the image frame of the video to be detected can be obtained, and the difference between the maximum judgment value and the minimum judgment value is determined as the swing amplitude of the face in the image frame of the video to be detected. When the difference is greater than the first preset threshold, it can be determined that the swing amplitude of the face meets the preset swing standard.
[0085] Through research, it is found that the normal nodding process of a person should be a roughly uniform process without sudden changes. Therefore, when it is determined that the swing speed of the face in the image frame of the video to be detected is not uniform, it can be determined that the swing speed of the face does not meet the preset swing standard. When determining the swing speed of the face, it can be determined whether the magnitudes of the first judgment values corresponding to adjacent image frames in the video to be detected change uniformly, that is, whether the variance between adjacent image frames is close to 0 or less than the second preset threshold. When the determined variance is greater than the second threshold, it can be determined that the swing speed of the face does not meet the preset swing standard.
[0086] It should be noted that when simultaneously determining whether the swing amplitude and swing speed of the face in the image frame of the video to be detected meet the preset swing criteria, only when both the swing amplitude and swing speed of the face meet the preset swing criteria can the step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend be continued. If either the swing amplitude or the swing speed of the face does not meet the preset swing criteria, it indicates that the face in the video to be detected is an abnormal face.
[0087] To further improve the accuracy of live detection, based on the above embodiments, in the embodiments of the present application, before determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend, the method further includes:
[0088] Determine a second judgment value in the face of each image frame according to the position information and identification of the face key points of each image frame, where the second judgment value is the ratio of the longitudinal width of the forehead to the longitudinal width of the mandible;
[0089] Determine the second change trend of the second judgment value according to multiple second judgment values;
[0090] If the second change trend meets the second preset change trend, then execute the step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend.
[0091] To further improve the accuracy of live detection, before determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend, it is also possible to determine the ratio of the longitudinal width of the forehead to the longitudinal width of the mandible in the face of each image frame of the video to be detected according to the position information and identification of the face key points of each image frame. For the sake of convenience of description, this ratio of the longitudinal width of the forehead to the longitudinal width of the mandible can be referred to as the second judgment value. To determine the longitudinal width of the forehead, in the embodiments of the present application, the identifications of the face key points in two forehead regions can be pre-saved, and then the longitudinal width of the forehead can be determined according to the position information corresponding to these two face key points. Similarly, to determine the longitudinal width of the mandible, the identifications of the face key points in two mandible regions can be pre-saved, and then the longitudinal width of the mandible can be determined according to the position information corresponding to these two face key points. It should be noted that how to determine the longitudinal width of the forehead and the longitudinal width of the mandible is not limited to the above examples, and those skilled in the art can configure it according to needs.
[0092] After determining the second judgment values of the human faces in each image frame of the video to be detected, the second change trend of the second judgment values can be determined based on the multiple second judgment values. The process of determining the second change trend of the second judgment values is similar to the principle of determining the first change trend, and this will not be elaborated in this embodiment of the present application.
[0093] If the determined second change trend meets the second preset change trend, it indicates that the human face in the video to be detected is performing a nodding action. To further improve the accuracy of live detection, the step of determining the live detection result of the human face in the video to be detected according to the first change trend and the first preset change trend can be continued. Figure 5a A schematic diagram of a human face when looking straight ahead provided by an embodiment of the present application Figure 5b A schematic diagram of a human face when lowering the head provided by an embodiment of the present application Figure 5c A schematic diagram of a human face when raising the head provided by an embodiment of the present application Figure 5a 、 Figure 5b and Figure 5c can be regarded as a nodding action. During the nodding process, the longitudinal width of the forehead first becomes larger and then smaller. At the same time, the longitudinal width of the lower jaw first becomes smaller and then larger. Generally, it can be obtained that the logic of nodding should satisfy becoming larger first and then smaller. Therefore, the second preset change trend can be becoming larger first and then smaller.
[0094] Specifically, Figure 6 A schematic diagram of a human face provided by an embodiment of the present application. As Figure 6 shown, in the forehead area of the human face, facial key points A and B are marked. Then, during live detection, the distance M between facial key points A and B can be determined according to the position information corresponding to facial key points A and B. This distance M is the longitudinal width of the forehead. As Figure 6 shown, in the lower jaw area of the human face, facial key points C and D are marked. Then, during live detection, the distance N between facial key points C and D can be determined according to the position information corresponding to facial key points C and D. This distance N is the longitudinal width of the lower jaw.
[0095] In the embodiments of the present application, if the liveness detection is performed only by whether the first change trend meets the first preset change trend or only by whether the second change trend meets the second preset change trend, malicious users can easily deceive the system through abnormal means. For example, by using a notebook or other objects to tentatively shake or quickly block and move around the lower jaw or forehead, resulting in the jitter or "floating" of the facial key points in the forehead or lower jaw part of the face. This unstable jitter state has a certain probability of causing the change of facial key points to meet the logic of the nodding action, thus deceiving the system. A typical malicious scenario is that a malicious user holds a high-definition photo of the target user in front of the camera. After the system prompts the user to nod, the malicious user uses some rigid objects to block the face to fraudulently access the system, thus achieving illegal intentions.
[0096] Figure 7a A schematic diagram of a normal face provided by the embodiments of the present application Figure 7a Each dot shown in the figure may represent the facial key points in the forehead area and the lower jaw area. Figure 7b A schematic diagram of an abnormal face provided by the embodiments of the present application, such as Figure 7b As shown, the facial key points in the forehead area of the face have a position offset due to the occlusion of the mobile phone. Figure 7b Compared with the facial key points in Figure 7a the facial key points in the forehead area in Figure 7a have generally moved downward (up and down in the figure), resulting in a change in the ratio of the longitudinal width of the forehead to the longitudinal width of the lower jaw. That is, compared with the ratio in Figure 7b the ratio in Figure 7b becomes smaller. It should be noted that the key position information determined by different recognition models may be inconsistent. Therefore, Figure 7c the position information of the facial key points shown in Figure 7c is only an example and does not represent the general change of the position of facial key points. Figure 7d A schematic diagram of an abnormal face provided by the embodiments of the present application, such as Figure 7c As shown, due to the occlusion of the lower jaw area, the facial key points in the lower jaw area are generally shifted upward (up and down in the figure), and the longitudinal width of the lower jaw becomes smaller. At this time, the ratio of the longitudinal width of the forehead part to the longitudinal width of the lower jaw part will become larger. Figure 7d A schematic diagram of an abnormal face provided by the embodiments of the present application, such as Figure 7d As shown, using other objects to block the lower jaw area will also produce Figure 7c the phenomenon in
[0097] In the embodiments of the present application, before determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend, it is further determined whether the first change trend meets the second preset change trend. That is to say, in the embodiments of the present application, it is not only determined whether the change trend of the distance between the tip of the nose and the horizontal line of the face meets the preset change trend, but also determined whether the change trend of the ratio of the longitudinal width of the forehead to the longitudinal width of the mandible meets the preset change trend. As long as any change trend does not meet the preset change trend, the live detection cannot pass. If a malicious user plans to use a malicious occlusion method to pass the live detection, based on the live detection method provided in the embodiments of the present application, the malicious user needs to occlude too many areas of the face. If there are too many occluded areas on the face, it is impossible to smoothly determine the position information and identification of the face key points, and thus it is impossible to smoothly perform the live detection. Therefore, the embodiments of the present application improve the accuracy of the live detection through multi-angle judgments.
[0098] In order to further improve the accuracy of the live detection, on the basis of the above embodiments, in the embodiments of the present application, the process of determining the longitudinal width of the forehead includes:
[0099] According to the position information of N preset first key point pairs, determine N first longitudinal widths in the forehead area of the face. Any first key point pair includes a face key point on one side of the eyebrow in the forehead area and a face key point on one side of the hair in the forehead area, where N is an integer greater than 1;
[0100] Determine the first average value of the N first longitudinal widths, and determine the first average value as the longitudinal width of the forehead.
[0101] There may be inaccurate problems in obtaining the position information of some face key points in the face. Therefore, in order to further improve the accuracy of the live detection, in the embodiments of the present application, when determining the longitudinal width of the forehead, N first longitudinal widths in the forehead area of the face can be determined according to the position information and identification of each determined face key point. Wherein, N is an integer greater than 1, that is to say, the position information of multiple first key point pairs is pre-saved, and a first longitudinal width can be determined according to each first key point pair. Any first key point pair includes a face key point on one side of the eyebrow in the forehead area and a face key point on one side of the hair in the forehead area, and the determined first longitudinal width is the pixel distance between the two face key points in the corresponding first key point pair.
[0102] After determining the N first longitudinal widths, in order to obtain a more robust longitudinal width of the forehead, a method of calculating and averaging multiple sets of values can be used, that is, determine the first average value of the N first longitudinal widths, and determine the first average value as the longitudinal width of the forehead.
[0103] Specifically, Figure 8 a face schematic provided by an embodiment of the present application is shown as Figure 8 shown. In the forehead area of the face, there are 3 line segments, namely R1, R2, and R3. Each line segment is determined according to two face key points. The lengths of these 3 line segments are the 3 first longitudinal widths. According to Figure 8 it can be known that the line segment R1 is determined according to the face key points labeled 55 and 109, and the length of this line segment R1 can be expressed as R1; the line segment R2 is determined according to the face key points labeled 8 and 10, and the length of this line segment R2 can be expressed as R2; the line segment R3 is determined according to the face key points labeled 285 and 338, and the length of this line segment R3 can be expressed as R3. Then, according to the above 3 first longitudinal widths, the Figure 8 forehead longitudinal width H of the face in H =(R1 + R2 + R3) / 3.
[0104] In order to further improve the accuracy of liveness detection, on the basis of the above embodiments, in an embodiment of the present application, the process of determining the mandibular longitudinal width includes:
[0105] According to the position information of M preset second key point pairs, M second longitudinal widths in the mandibular area of the face are determined. The second key point pair includes a face key point on one side of the lips in the mandibular area and a face key point on one side of the face edge in the mandibular area, and M is an integer greater than 1;
[0106] Determine the second average value of the M second heights, and determine the second average value as the mandibular longitudinal width.
[0107] Similarly, in order to further improve the accuracy of liveness detection, in an embodiment of the present application, when determining the mandibular longitudinal width, M second longitudinal widths in the mandibular area of the face can be determined according to the position information of M preset second key point pairs. Among them, M is an integer greater than 1, and the specific value of M can be the same as or different from the specific value of N. That is to say, multiple second key point pairs are pre - stored, and a second longitudinal width can be determined according to each second key point pair. Any second key point pair includes a face key point on one side of the lips in the mandibular area and a face key point on one side of the face edge in the mandibular area, and the determined second longitudinal width is the pixel distance between the two face key points in the corresponding second key point pair.
[0108] After determining the M first longitudinal widths, in order to obtain a more robust mandibular longitudinal width, the average value can be calculated using multiple sets of values, that is, the second average value of the M second longitudinal widths is determined, and this second average value is determined as the mandibular longitudinal width.
[0109] Specifically, as Figure 8 shown, there are 3 line segments in the mandibular region of the human face, namely R4, R5, and R6. Each line segment is determined based on two face key points. The lengths of these 3 line segments are the 3 second longitudinal widths. According to Figure 8 it can be known that the line segment R4 is determined based on the face key points labeled 148 and 83. The length of this line segment R4 can be expressed as R4; the line segment R5 is determined based on the face key points labeled 18 and 152. The length of this line segment R5 can be expressed as R5; the line segment R6 is determined based on the face key points labeled 313 and 377. The length of this line segment R6 can be expressed as R6. Then, based on the above 3 second longitudinal widths, the Figure 8 mandibular longitudinal width H of the human face in C =(R4 + R5 + R6) / 3.
[0110] To further improve the accuracy of live detection, based on the above embodiments, the live detection method provided by the embodiments of the present application can not only perform live detection based on the longitudinal indicators of the human face, but also add the judgment of the transverse indicators of the human face during the live detection process. When the changes in the transverse indicators corresponding to all image frames of the video to be detected conform to the nodding logic, it is determined that the live detection passes. Specifically, in the embodiments of the present application, before determining the live detection result of the human face in the video to be detected according to the first change trend and the first preset change trend, the method further includes:
[0111] If the length changes of the face horizontal lines in adjacent image frames in the video to be detected are all less than the second preset threshold, then perform the step of determining the live detection result of the human face in the video to be detected according to the first change trend and the first preset change trend.
[0112] Since a person's normal nodding process basically does not cause a change in the length of the face horizontal line, in order to further improve the accuracy of live detection, before determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend, it is possible to determine whether the length of the face horizontal line of the face in adjacent image frames in the video to be detected changes. If the length changes, it indicates that the current detected nodding behavior may be an abnormal nodding action. Since there may be slight errors when obtaining the position information of each face key point, in the embodiments of the present application, it is possible to determine whether the length change of the face horizontal line corresponding to adjacent image frames in the video to be detected is less than a second preset threshold. If so, the step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend can be executed. That is to say, a slight change in the length of the face horizontal line between adjacent image frames is allowed. Among them, the value of the second preset threshold can be a value such as 0.5, 0.4, 0.1, or 0.05. The embodiments of the present application do not limit the value of the second preset threshold, and those skilled in the art can set the second preset threshold according to actual experimental conditions.
[0113] If the face is maliciously blocked during the "nodding" process, the length of the face horizontal line corresponding to adjacent image frames will change greatly, so as to identify that the face in the video to be detected may be a non-live face.
[0114] Specifically, assume that image frame 1, image frame 2, image frame 3, and image frame 4 in the video to be detected are consecutive image frames. Among them, the length of the face horizontal line in image frame 1 is 15.3, the length of the face horizontal line in image frame 2 is 15.3, the length of the face horizontal line in image frame 3 is 15.8, the length of the face horizontal line in image frame 4 is 15.8, and the second preset threshold is 0.1. Since the length change of the face horizontal line between image frame 2 and image frame 3 is greater than the second preset threshold 0.1, it can be determined that the nodding action of the face in the video to be detected is abnormal, and an abnormal prompt message can be output.
[0115] In order to further improve the accuracy of live detection, on the basis of the above embodiments, in the embodiments of the present application, before determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend, the method further includes:
[0116] According to the position information and identification of the face key points of each image frame, determine the horizontal width of the forehead and / or the horizontal width of the mandible of the face in each image frame;
[0117] If the change in the horizontal width of the forehead and / or the horizontal width of the lower jaw of the face in adjacent image frames of the video to be detected is less than a third preset threshold, then perform the step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend.
[0118] Since when a person nods normally, it basically does not cause a change in the length of the face in the horizontal direction. Therefore, in order to further improve the accuracy of live detection, before determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend, the horizontal width of the forehead and / or the horizontal width of the lower jaw of the face in each image frame can be determined according to the position information and identification of each face key point corresponding to the image frame of the video to be detected. If the horizontal width of the forehead and / or the horizontal width of the lower jaw changes, it indicates that the current detected nodding behavior may be an abnormal nodding action.
[0119] Since there may be slight errors when obtaining the position information of each face key point, in the embodiments of the present application, it can be determined whether the change in the horizontal width of the forehead and / or the horizontal width of the lower jaw corresponding to adjacent image frames in the video to be detected is less than a third preset threshold. If so, the step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend can be performed. That is to say, a slight change in the horizontal width of the forehead and / or the horizontal width of the lower jaw between adjacent image frames is allowed. Among them, the value of the third preset threshold can be a value such as 0.5, 0.1, or 0.05, etc. The embodiments of the present application do not limit the value of the third preset threshold, and those skilled in the art can set the third preset threshold according to actual experimental situations.
[0120] If the face is maliciously blocked during the "nodding" process, the horizontal width of the forehead and / or the horizontal width of the lower jaw corresponding to adjacent image frames will change greatly, thereby identifying that the face in the video to be detected may be a non-live face.
[0121] Specifically, assume that image frame 1, image frame 2, image frame 3, and image frame 4 in the video to be detected are consecutive image frames. Among them, the horizontal width of the forehead in image frame 1 is 15.31, the horizontal width of the forehead in image frame 2 is 15.31, the horizontal width of the forehead in image frame 3 is 15.33, the horizontal width of the forehead in image frame 4 is 15.32, and the third preset threshold is 0.1. Since the change in the horizontal width of the forehead in adjacent image frames is not greater than the third preset threshold of 0.1, it can be determined that the nodding action of the face in the video to be detected is normal, and the step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend can be performed.
[0122] In order to further improve the accuracy of live detection, based on the above embodiments, in the embodiments of the present application, the process of determining the horizontal width of the forehead or the horizontal width of the lower jaw includes:
[0123] According to the position information of F preset third key point pairs, determine F horizontal widths in the forehead area or the lower jaw area of the human face. The third key point pair includes two human face key points located on one side of the eyebrows in the forehead area, or two human face key points located on one side of the hair in the forehead area, or two human face key points located on one side of the lips in the lower jaw area, or two human face key points located on one side of the edge of the human face in the lower jaw area. F is an integer greater than 1;
[0124] Determine the third average value of the F horizontal widths, and determine the third average value as the forehead width or the lower jaw width.
[0125] There may be inaccurate problems in obtaining the position information of some human face key points in the human face. Therefore, in order to further improve the accuracy of live detection, in the embodiments of the present application, when determining the horizontal width of the forehead or the horizontal width of the lower jaw, F horizontal widths in the forehead area or the lower jaw area of the human face can be determined according to the position information of F preset third key point pairs. Among them, F is an integer greater than 1, and the specific value of F can be the same as or different from the specific values of M and N. That is to say, multiple third key point pairs are pre-stored, and a horizontal width can be determined according to each set of human face key points.
[0126] In the embodiments of the present application, in order to facilitate the determination of the horizontal width of the forehead, F third key point pairs can be saved for the forehead area. Among these third key point pairs, two human face key points located on one side of the eyebrows in the forehead area can be used as a third key point pair, or two human face key points located on one side of the hair in the forehead area can be used as a third key point pair. After determining the F horizontal widths, in order to obtain a more robust horizontal width of the forehead, a method of calculating and averaging multiple sets of values can be used for calculation, that is, determine the third average value of the F horizontal widths, and determine the third average value as the horizontal width of the forehead.
[0127] Similarly, in order to facilitate the determination of the horizontal width of the lower jaw, F third key point pairs can be saved for the lower jaw area. Among these third key point pairs, two human face key points located on one side of the lips in the lower jaw area can be used as a third key point pair, or two human face key points located on one side of the edge of the human face in the lower jaw area can be used as a third key point pair. After determining the F horizontal widths, in order to obtain a more robust horizontal width of the lower jaw, a method of calculating and averaging multiple sets of values can be used for calculation, that is, determine the third average value of the F horizontal widths, and determine the third average value as the horizontal width of the lower jaw.
[0128] Specifically, Figure 9 A face illustration provided by an embodiment of the present application is shown in Figure 9 As shown, in the forehead area of the face, there are 4 horizontal line segments, namely G1, G2, G3, and G4. Each line segment is determined based on two face key points. The lengths of these 4 line segments are the 4 horizontal widths. According to Figure 9 it can be known that line segment G1 is determined based on the face key points labeled 109 and 10, and the length of this line segment G1 can be expressed as G1; line segment G2 is determined based on the face key points labeled 10 and 338, and the length of this line segment G2 can be expressed as G2; line segment G3 is determined based on the face key points labeled 55 and 8, and the length of this line segment G3 can be expressed as G3; line segment G4 is determined based on the face key points labeled 8 and 285, and the length of this line segment G4 can be expressed as G4. Then, based on the above 4 horizontal widths, the Figure 9 horizontal width M of the forehead of the face in H =(G1 + G2 + G3 + G4) / 4.
[0129] As Figure 9 shown, in the jaw area of the face, there are 4 horizontal line segments, namely G5, G6, G7, and G8. Each line segment is determined based on two face key points. The lengths of these 4 line segments are the 4 horizontal widths. According to Figure 9 it can be known that line segment G5 is determined based on the face key points labeled 83 and 18, and the length of this line segment G5 can be expressed as G5; line segment G6 is determined based on the face key points labeled 18 and 313, and the length of this line segment G6 can be expressed as G6; line segment G7 is determined based on the face key points labeled 148 and 152, and the length of this line segment G7 can be expressed as G7; line segment G8 is determined based on the face key points labeled 152 and 377, and the length of this line segment G8 can be expressed as G8. Then, based on the above 4 horizontal widths, the Figure 9 horizontal width M of the jaw of the face in H =(G5 + G6 + G7 + G8) / 4.
[0130] To further improve the accuracy of live detection, based on the above embodiments, in an embodiment of the present application, before determining the live detection result of the face in the video to be detected, the method further includes:
[0131] Determine the reference value of the face in each image frame according to the position information and label of the face key points of each image frame;
[0132] Normalize the first judgment value, the second judgment value, the horizontal width of the forehead, and the horizontal width of the mandible corresponding to the image frame based on the reference value of the human face in each image frame.
[0133] Since there is a problem of objects appearing larger when closer and smaller when farther away in the images captured by the image acquisition device when photographing a human face, in order to further improve the accuracy of live detection, in the embodiments of the present application, the first judgment value, the second judgment value, the horizontal width of the forehead, and the horizontal width of the mandible of the human face in each frame of image can be normalized. The role of normalization is to unify the first judgment value, the second judgment value, the horizontal width of the forehead, and the horizontal width of the mandible to a single dimension, so that these values will not jitter significantly due to changes in the distance from the camera or angular deviation. For example, the pixel distance between the left corner of the mouth and the right corner of the mouth will change with the change in the distance between the human face and the camera, specifically showing that objects appear larger when closer and smaller when farther away, that is, when the human face is closer to the camera, the pixel distance between the corners of the mouth is relatively large, and conversely, when farther away, it will appear relatively small. However, the ratio of the distance between the corners of the mouth of the same person to the length of other parts of the human face will not change with the change of the human face state. For example, the distance between the corners of the mouth is twice the width of the nose, etc. This two-fold size relationship will not change with the change of distance, that is, it will increase or decrease equally, but always satisfy this two-fold relationship.
[0134] In the embodiments of the present application, the reference value of the human face in each image frame can be determined according to the position information of the key points of the human face in each image frame, and subsequently, the first judgment value, the second judgment value, the horizontal width of the forehead, and the horizontal width of the mandible of the human face in the corresponding image frame can be normalized based on this reference value. When performing normalization, the value to be normalized in any image frame can be divided by the corresponding reference value to obtain the normalized data, where the value to be normalized can be the first judgment value, the second judgment value, the horizontal width of the forehead, and the horizontal width of the mandible.
[0135] In the embodiments of the present application, the reference value can be determined according to the position information corresponding to the pre-saved normalized human face key points. When determining the reference value, it can be calculated based on the following formula:
[0136] Base = (B1 + B2) / 2
[0137] Where Base is the reference value; B1 is the distance between any two key points of the human face in the image frame; B2 is the distance between any two key points of the human face in the image frame, where the key points of the human face when calculating B1 are different from those when calculating B2.
[0138] When calculating B1 and B2, the Euclidean distance calculation formula can be used to calculate the pixel distance between two key points of the human face:
[0139] Dist = sqrt((x1 - x2) 2 +(y1 - y2) 2 )
[0140] where Dist is the distance between two facial key points; sqrt is to calculate the arithmetic square root of the value therein, that is, to calculate the arithmetic square root of (x1 - x2) 2 +(y1 - y2) 2 ; x1 is the x coordinate corresponding to facial key point 1; x2 is the x coordinate corresponding to facial key point 2; y1 is the y coordinate corresponding to facial key point 1; y2 is the y coordinate corresponding to facial key point 2.
[0141] It should be noted that when calculating the reference value, the selected facial key points should follow the following principle: the selected facial key points should be as much as possible the focus of facial features, and this focus can be an intersection point or a corner point. For example, facial key points such as the corners of the mouth, the corners of the eyes, and the tip of the nose, which are easier to locate.
[0142] Specifically, Figure 10 is a facial schematic diagram provided by an embodiment of the present application. As Figure 10 shown, Figure 10 it includes a line segment determined according to the facial key points labeled 64 and labeled 294. For the convenience of description, this line segment can be called B1. It also includes a line segment determined according to the facial key points labeled 4 and labeled 2. For the convenience of description, this line segment can be called B2. The facial key points labeled 4, 294, 4, and 2 are finally determined after pre-screening and comparing many combinations of values. The reference value Base determined based on B1 and B2 is: Base = (LB1 + LB2) / 2, where LB1 represents the length of line segment B1, and LB2 represents the length of line segment B2.
[0143] The following describes the process of live detection in combination with a specific embodiment. Figure 11 is a schematic flowchart of live detection provided by an embodiment of the present application. As Figure 11As shown, the data input by the user of the electronic device can be a video or image frames with temporal significance. The electronic device determines the received data as the video to be detected and converts each image frame in the video to be detected into the RGB color mode. Each image frame in the video to be detected after the color mode conversion is input into Mediapipe. Mediapipe integrates the functions of face detection, preprocessing, and face key point detection, and can process the received image frames and output the position information and corresponding identifiers of 478 face key points of the face in the image frame. Among them, the position information of each face key point is represented in the form of (x, y, z), where x and y are coordinate values based on the image pixel coordinate system and are normalized by the width and height of the input image, and z represents the offset of the face key point from the face centroid.
[0144] After obtaining the position information and identifiers of each face key point, it is possible to judge nodding based on the position information of the face key points, so as to output a conclusion on whether the face included in the video to be detected is a live face.
[0145] Figure 12 The flow chart of a live detection process provided by an embodiment of the present application is shown in Figure 12 As shown, this process includes the following steps:
[0146] S1201: Obtain the video to be detected, and use Mediapipe to perform face key point detection on each image frame in the video to be detected, and obtain the position information and identifiers of the face key points of the face in each image frame output by Mediapipe.
[0147] S1202: According to the position information and identifiers of the face key points in each image frame, determine the first judgment value, the second judgment value, the length of the face horizontal line, the horizontal width of the forehead, and the horizontal width of the mandible in the face of each image frame.
[0148] S1203: According to the position information and identifiers of the face key points in each image frame, determine the reference value of the face in each image frame, and use the reference value to perform normalization processing on the first judgment value, the second judgment value, the length of the face horizontal line, the horizontal width of the forehead, and the horizontal width of the mandible corresponding to the face of each image frame.
[0149] S1204: Determine the first change trend and the second change trend according to the first judgment value and the second judgment value after the normalization processing.
[0150] S1205: Judge whether the changes in the horizontal width of the forehead, the horizontal width of the mandible, and the length of the face horizontal line after the normalization processing are all less than the preset threshold. If so, execute S1206; otherwise, execute S1208.
[0151] S1206: Determine whether the face in the video to be detected is a live face according to the first change trend, the second change trend, the first preset change trend, and the second preset change trend. If so, execute S1207; otherwise, execute S1208.
[0152] S1207: Determine that the face in the video to be detected is a live face.
[0153] S1208: Determine that the face in the video to be detected is a non-live face.
[0154] Based on the same technical concept, on the basis of the above embodiments, the present application provides a live detection device. Figure 13 As shown in the structural schematic diagram of a live detection device provided by an embodiment of the present application, Figure 13 The device includes:
[0155] An acquisition module 1301, configured to acquire the position information and identification of the face key points of the image frames of the video to be detected.
[0156] A determination module 1302, configured to determine the first judgment value of the face in each image frame according to the position information and identification of the face key points of each image frame, where the first judgment value is the distance between the tip of the nose and the face horizontal line; and determine the first change trend of the first judgment value according to a plurality of the first judgment values.
[0157] A detection module 1303, configured to determine the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend.
[0158] In a possible implementation manner, the device further includes:
[0159] An adding module 1304, configured to add a direction symbol to the first judgment value of the face in each image frame according to the position relationship between the tip of the nose and the face horizontal line in each image frame.
[0160] In a possible implementation manner, the face horizontal line is determined according to the face key points located on the same horizontal line on the face, and the face horizontal line divides the face in each image frame into a first region and a second region; the adding module 1304 is specifically configured to, for each image frame, if the tip of the nose in the image frame is located in the first region, add a first direction symbol to the first judgment value of the face in the image frame; if the tip of the nose in the image frame is located in the second region, add a second direction symbol to the first judgment value of the face in the image frame, where the first direction symbol and the second direction symbol are in opposite directions.
[0161] In a possible implementation, the determining module 1302 is specifically configured to obtain a preset number of target image frames in the video to be detected, and determine a first change trend of first judgment values corresponding to the preset number of target image frames.
[0162] In a possible implementation, the determining module 1302 is further configured to, if the swing amplitude and / or swing speed of the face in the image frames in the video to be detected meet a preset swing criterion, perform the step of determining a live detection result of the face in the video to be detected according to the first change trend and a first preset change trend.
[0163] In a possible implementation, the determining module 1302 is further configured to determine a second judgment value in the face of each image frame according to the position information and identification of the face key points of each image frame, where the second judgment value is the ratio of the longitudinal width of the forehead to the longitudinal width of the mandible; determine a second change trend of the second judgment values according to the plurality of second judgment values;
[0164] The detecting module 1303 is further configured to, if the second change trend meets a second preset change trend, perform the step of determining a live detection result of the face in the video to be detected according to the first change trend and the first preset change trend.
[0165] In a possible implementation, the determining module 1302 is further configured to determine N first longitudinal widths in the forehead area of the face according to the position information of N preset first key point pairs, where any one of the first key point pairs includes a face key point on one side of the eyebrow in the forehead area and a face key point on one side of the hair in the forehead area, and N is an integer greater than 1; determine a first average value of the N first longitudinal widths, and determine the first average value as the longitudinal width of the forehead.
[0166] In a possible implementation, the determining module 1302 is further configured to determine M second longitudinal widths in the mandible area of the face according to the position information of M preset second key point pairs, where the second key point pair includes a face key point on one side of the lip in the mandible area and a face key point on one side of the face edge in the mandible area, and M is an integer greater than 1; determine a second average value of the M second heights, and determine the second average value as the longitudinal width of the mandible.
[0167] In a possible implementation, the determining module 1302 is further configured to, if the length change of the face horizontal line in adjacent image frames in the video to be detected is less than a second preset threshold, perform the step of determining a live detection result of the face in the video to be detected according to the first change trend and the first preset change trend.
[0168] In a possible implementation manner, the determining module 1302 is further configured to determine the horizontal width of the forehead and / or the horizontal width of the mandible of the face in each image frame according to the position information and identification of the facial key points of each image frame;
[0169] The detecting module 1303 is further configured to, if the change in the horizontal width of the forehead and / or the horizontal width of the mandible of the face in adjacent image frames in the video to be detected is less than a third preset threshold, perform the step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend.
[0170] In a possible implementation manner, the determining module 1302 is further configured to determine F horizontal widths in the forehead area or the mandible area of the face according to the position information of F third key point pairs preset, where the third key point pair includes two facial key points on one side of the eyebrows in the forehead area, or two facial key points on one side of the hair in the forehead area, or two facial key points on one side of the lips in the mandible area, or two facial key points on one side of the edge of the face in the mandible area, and F is an integer greater than 1; determine the third average value of the F horizontal widths, and determine the third average value as the forehead width or the mandible width.
[0171] In a possible implementation manner, the determining module 1302 is further configured to determine a reference value of the face in each image frame according to the position information and identification of the facial key points of each image frame; perform normalization processing on the first judgment value, the second judgment value, the horizontal width of the forehead, and the horizontal width of the mandible corresponding to the image frame based on the reference value of the face in each image frame.
[0172] Based on the same technical concept, the present application further provides an electronic device, Figure 14 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application, as Figure 14 shown, including: a processor 1401, a communication interface 1402, a memory 1403, and a communication bus 1404, where the processor 1401, the communication interface 1402, and the memory 1403 complete mutual communication through the communication bus 1404;
[0173] A computer program is stored in the memory 1403, and when the program is executed by the processor 1401, the processor 1401 is caused to execute the following steps:
[0174] Obtain the position information and identification of the facial key points of the image frames of the video to be detected;
[0175] Determine the first judgment value of the face in each image frame according to the position information and identification of the face key points of each image frame, where the first judgment value is the distance between the tip of the nose and the face horizontal line;
[0176] Determine the first change trend of the first judgment value according to multiple first judgment values;
[0177] Determine the living detection result of the face in the video to be detected according to the first change trend and the first preset change trend.
[0178] In a possible implementation manner, the processor 1401 is further configured to add a direction symbol to the first judgment value of the face in each image frame according to the position relationship between the tip of the nose and the face horizontal line in each image frame.
[0179] In a possible implementation manner, the face horizontal line is determined according to the face key points on the same horizontal line on the face, and the face horizontal line divides the face in each image frame into a first area and a second area; the processor 1401 is further configured to, for each image frame, if the tip of the nose in the image frame is located in the first area, add a first direction symbol to the first judgment value of the face in the image frame; if the tip of the nose in the image frame is located in the second area, add a second direction symbol to the first judgment value of the face in the image frame, where the first direction symbol and the second direction symbol are opposite in direction.
[0180] In a possible implementation manner, the processor 1401 is further configured to obtain a preset number of target image frames in the video to be detected and determine the first change trend of the first judgment values corresponding to the preset number of target image frames.
[0181] In a possible implementation manner, the processor 1401 is further configured to, if the swing amplitude and / or swing speed of the face in the image frame in the video to be detected meet the preset swing standard, execute the step of determining the living detection result of the face in the video to be detected according to the first change trend and the first preset change trend.
[0182] In a possible implementation manner, the processor 1401 is further configured to determine a second judgment value in the face in each image frame according to the position information and identification of the face key points of each image frame, where the second judgment value is the ratio of the longitudinal width of the forehead to the longitudinal width of the mandible;
[0183] Determine the second change trend of the second judgment value according to multiple second judgment values;
[0184] If the second change trend satisfies the second preset change trend, perform the step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend.
[0185] In a possible implementation manner, the processor 1401 is further configured to determine N first longitudinal widths in the forehead region of the face according to the position information of N preset first key point pairs. The first key point pair includes a face key point on one side of the eyebrow in the forehead region and a face key point on one side of the hair in the forehead region, and N is an integer greater than 1.
[0186] Determine a first average value of the N first longitudinal widths, and determine the first average value as the forehead longitudinal width.
[0187] In a possible implementation manner, the processor 1401 is further configured to determine M second longitudinal widths in the mandibular region of the face according to the position information of M preset second key point pairs. The second key point pair includes a face key point on one side of the lip in the mandibular region and a face key point on one side of the face edge in the mandibular region, and M is an integer greater than 1.
[0188] Determine a second average value of the M second heights, and determine the second average value as the mandibular longitudinal width.
[0189] In a possible implementation manner, the processor 1401 is further configured to perform the step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend if the change in the length of the face horizontal line in adjacent image frames in the video to be detected is less than a second preset threshold.
[0190] In a possible implementation manner, the processor 1401 is further configured to determine the forehead horizontal width and / or the mandibular horizontal width of the face in each image frame according to the position information and identification of the face key points in each image frame.
[0191] If the change in the forehead horizontal width and / or the mandibular horizontal width of the face in adjacent image frames in the video to be detected is less than a third preset threshold, perform the step of determining the live detection result of the face in the video to be detected according to the first change trend and the first preset change trend.
[0192] In a possible implementation, the processor 1401 is further configured to determine F lateral widths in the forehead area or the jaw area of the human face according to the position information of F preset third key-point pairs. The third key-point pairs include two human face key points on one side of the eyebrows in the forehead area, or two human face key points on one side of the hair in the forehead area, or two human face key points on one side of the lips in the jaw area, or two human face key points on one side of the edge of the human face in the jaw area, where F is an integer greater than 1;
[0193] Determine a third average value of the F lateral widths, and determine the third average value as the forehead width or the jaw width.
[0194] In a possible implementation, the processor 1401 is further configured to determine a reference value of the human face in each image frame according to the position information and identification of the human face key points in each image frame;
[0195] Perform normalization processing on the corresponding first judgment value, second judgment value, forehead lateral width, and jaw lateral width based on the reference value of the human face in each image frame.
[0196] Since the principle of the above electronic device for solving problems is similar to that of the living body detection method, the implementation of the above electronic device can refer to the embodiments of the method, and the repeated parts will not be described again.
[0197] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface 1402 is used for communication between the above electronic device and other devices. The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0198] The above-mentioned processor may be a general-purpose processor, including a central processing unit, a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0199] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, in which a computer program executable by an electronic device is stored. When the program runs on the electronic device, it enables the electronic device to implement the living body detection method described in any of the above embodiments when executed.
[0200] The above-mentioned computer-readable storage medium may be any available medium or data storage device accessible by a processor in an electronic device, including but not limited to magnetic memories such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc., optical memories such as CDs, DVDs, BDs, HVDs, etc., and semiconductor memories such as ROMs, EPROMs, EEPROMs, NAND FLASH (non-volatile memories), solid-state drives (SSDs), etc.
[0201] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0202] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, in which a computer program executable by an electronic device is stored. When the program runs on the electronic device, it enables the electronic device to implement the living body detection method described in any of the above embodiments when executed.
[0203] The above-mentioned computer-readable storage medium may be any available medium or data storage device accessible by a processor in an electronic device, including but not limited to magnetic memories such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc., optical memories such as CDs, DVDs, BDs, HVDs, etc., and semiconductor memories such as ROMs, EPROMs, EEPROMs, NAND FLASH (non-volatile memories), solid-state drives (SSDs), etc.
[0204] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0205] Based on the same technical concept, the embodiments of the present application provide a computer program product, which includes computer program code. When the computer program code runs on an electronic device, it enables the electronic device to implement the living body detection method described in any of the above embodiments when executed.
[0206] The computer program for executing the operations of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet). In some embodiments, by using the status information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0207] The computer program product described herein can be specifically implemented in a manner of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0208] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0209] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0210] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0211] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0212] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A method for live detection, characterized in that, The method includes: Obtaining the position information and identification of the face key points of the image frames of the video to be detected; Determining a first judgment value of the face in each image frame according to the position information and identification of the face key points of each image frame, where the first judgment value is the distance between the tip of the nose and the horizontal line of the face; Determining a first change trend of the first judgment value according to a plurality of the first judgment values; Determining the live detection result of the face in the video to be detected according to the first change trend and a first preset change trend.
2. The method according to claim 1, characterized in that After determining the first judgment value of the face in each image frame and before determining the first change trend of the first judgment value according to a plurality of the first judgment values, the method further includes: Adding a direction symbol to the first judgment value of the face in each image frame according to the positional relationship between the tip of the nose and the horizontal line of the face in each image frame.
3. The method according to claim 2, characterized in that, The horizontal line of the face is determined according to the face key points on the same horizontal line on the face, and the horizontal line of the face divides the face in each image frame into a first region and a second region; The adding a direction symbol to the first judgment value of the face in each image frame according to the positional relationship between the tip of the nose and the horizontal line of the face in each image frame includes: For each image frame, if the tip of the nose in the image frame is located in the first region, adding a first direction symbol to the first judgment value of the face in the image frame; If the tip of the nose in the image frame is located in the second region, adding a second direction symbol to the first judgment value of the face in the image frame, where the first direction symbol and the second direction symbol are in opposite directions.
4. The method according to claim 1, characterized in that The determining a first change trend of the first judgment value according to a plurality of the first judgment values includes: Obtaining a preset number of target image frames in the video to be detected and determining the first change trend of the first judgment values corresponding to the preset number of target image frames.
5. The method according to claim 1, characterized in that Before determining the live detection result of the face in the video to be detected according to the first change trend and a first preset change trend, the method further includes: If the swing amplitude and / or swing speed of the face in the image frames of the video to be detected meet a preset swing standard, performing the step of determining the live detection result of the face in the video to be detected according to the first change trend and a first preset change trend.
6. The method according to claim 1, characterized in that Before determining the live detection result of the face in the video to be detected according to the first change trend and a first preset change trend, the method further includes: Determining a second judgment value in the face of each image frame according to the position information and identification of the face key points of each image frame, where the second judgment value is the ratio of the longitudinal width of the forehead to the longitudinal width of the mandible; Determining a second change trend of the second judgment value according to a plurality of the second judgment values; If the second change trend meets a second preset change trend, performing the step of determining the live detection result of the face in the video to be detected according to the first change trend and a first preset change trend.
7. The method according to claim 6, wherein The process of determining the longitudinal width of the forehead includes: Based on the position information of N preset first key point pairs, determine N first longitudinal widths in the forehead region of the face. The first key point pair includes a face key point on one side of the eyebrow and a face key point on one side of the hair in the forehead region, and N is an integer greater than 1; Determine the first average value of the N first longitudinal widths, and determine the first average value as the forehead longitudinal width.
8. The method according to claim 6, characterized in that, The process of determining the mandibular longitudinal width includes: Based on the position information of M preset second key point pairs, determine M second longitudinal widths in the mandibular region of the face. The second key point pair includes a face key point on one side of the lip and a face key point on one side of the face edge in the mandibular region, and M is an integer greater than 1; Determine the second average value of the M second heights, and determine the second average value as the mandibular longitudinal width.
9. The method according to claim 1, characterized in that, Before determining the live detection result of the face in the to-be-detected video according to the first change trend and the first preset change trend, the method further includes: If the length change of the face horizontal line in adjacent image frames in the to-be-detected video is less than a second preset threshold, then execute the step of determining the live detection result of the face in the to-be-detected video according to the first change trend and the first preset change trend.
10. The method according to claim 1, wherein Before determining the live detection result of the face in the to-be-detected video according to the first change trend and the first preset change trend, the method further includes: According to the position information and identification of the face key points in each image frame, determine the forehead horizontal width and / or the mandibular horizontal width of the face in each image frame; If the change in the forehead horizontal width and / or the mandibular horizontal width of the face in adjacent image frames in the to-be-detected video is less than a third preset threshold, then execute the step of determining the live detection result of the face in the to-be-detected video according to the first change trend and the first preset change trend.
11. The method according to claim 10, characterized in that, The process of determining the forehead horizontal width or the mandibular horizontal width includes: Based on the position information of F preset third key point pairs, determine F horizontal widths in the forehead region or the mandibular region of the face. The third key point pair includes two face key points on one side of the eyebrow in the forehead region, or two face key points on one side of the hair in the forehead region, or two face key points on one side of the lip in the mandibular region, or two face key points on one side of the face edge in the mandibular region, and F is an integer greater than 1; Determine the third average value of the F horizontal widths, and determine the third average value as the forehead width or the mandibular width.
12. The method according to claim 10, characterized in that, Before determining the live detection result of the face in the to-be-detected video, the method further includes: According to the position information and identification of the face key points in each image frame, determine the reference value of the face in each image frame; Perform normalization processing on the first judgment value, the second judgment value, the forehead horizontal width, and the mandibular horizontal width corresponding to the image frame based on the reference value of the face in each image frame.
13. A living body detection device, characterized in that, The device includes: An acquisition module, configured to acquire the position information and identification of the facial key points of the image frames of the video to be detected. A determination module, configured to determine a first judgment value of the face in each image frame according to the position information and identification of the facial key points of each image frame, where the first judgment value is the distance between the tip of the nose and the horizontal line of the face; and determine a first change trend of the first judgment value according to a plurality of the first judgment values. A detection module, configured to determine the live detection result of the face in the video to be detected according to the first change trend and a first preset change trend.
14. An electronic device, characterized in that, The electronic device at least includes a processor and a memory, and the processor is configured to implement the steps of the live detection method according to any one of claims 1-12 when executing the computer program stored in the memory.
15. A computer storage medium, characterized in that, It stores a computer program executable by the electronic device, and when the program runs on the electronic device, the electronic device is caused to execute the steps of the live detection method according to any one of claims 1-12.
16. A computer program product, characterized in that, The computer program product includes: computer program code, and when the computer program code runs on the electronic device, the electronic device is caused to execute the steps of the live detection method according to any one of claims 1-12.