A method for detecting a living body, an electronic device, a storage medium, and a program product

By generating a response map combined with the first and second live detection models, and taking into account the light sequence and reflection mode, the problem of inaccurate live detection in the prior art is solved, and accurate distinction between live and remakes and effective identification of other attacks is achieved.

CN115147936BActive Publication Date: 2025-07-18YUANLI JINZHI (CHONGQING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210528092.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-16
Publication Date
2025-07-18
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

The existing colorful live detection technology separates the light sequence inspection and live attack behavior detection, resulting in inaccurate detection of the detection results and making it difficult to effectively distinguish between live and remake attacks.

Method used

By generating a response map of the object to be detected, using the first live detection model and the second live detection model to combine, comprehensively consider the similarity and reflection mode of the light sequence and video reflected light, and combining the uneven features of the face with the smooth characteristics of the screen/print paper, the distinction between living objects and remakes is achieved.

Benefits of technology

It improves the accuracy and accuracy of live detection, effectively recognizes remake attacks, ensures that the probability of a real person being judged as live does not decrease, and at the same time improves the detection ability of other attack types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147936B_ABST
    Figure CN115147936B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a live detection method, an electronic device, a storage medium, and a program product. The method includes: obtaining a video to be detected of an object to be detected, where the video of the object to be detected is a video of the object to be detected collected during irradiating the object to be detected according to a first illumination sequence; generating a response map of the object to be detected according to the first illumination sequence and a second illumination sequence reflected by each entity position point of the object to be detected represented by the video to be detected; performing live detection on the object to be detected by processing the response map based on the response map using a first live detection model to obtain a first live detection result of the object to be detected; inputting the video to be detected into a second live detection model to obtain a second live detection result of the object to be detected; and determining a final live detection result of the object to be detected according to the first live detection result and the second live detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a liveness detection method, electronic equipment, storage medium and program product. Background Art

[0002] Liveness detection technology is becoming more and more mature. The related colorful liveness detection technology mainly includes two parts: one is the lighting sequence test, which compares the emitted light sequence with the sequence of reflected light presented by the object to be detected in the video to determine whether there is a camera hijacking; the other is the liveness detection method through the collected video or image of the object to be detected to detect whether the colorful video contains common liveness attack behaviors, such as screen copying, printed paper copying, etc.

[0003] The existing colorful liveness detection technology usually separates the above two parts for detection, that is, different models or algorithms are used to implement the above two parts of detection, and then the results of the two parts are combined to obtain the final liveness detection result. Therefore, the existing colorful light liveness detection technology needs to be improved. Summary of the invention

[0004] In view of the above problems, the embodiments of the present application provide a living body detection method, an electronic device, a storage medium and a program product to overcome the above problems or at least partially solve the above problems.

[0005] A first aspect of an embodiment of the present application provides a living body detection method, comprising:

[0006] Acquire a video of the object to be detected, wherein the video to be detected is: a video of the object to be detected collected during a period in which the object to be detected is irradiated according to a first illumination sequence;

[0007] Generate a response map of the object to be detected according to the first illumination sequence and the second illumination sequence reflected by each physical position point of the object to be detected represented by the video to be detected, wherein the response intensity of each pixel point in the response map represents: the similarity between the second illumination sequence reflected by the physical position point corresponding to the pixel point and the first illumination sequence;

[0008] Based on the response graph, a first liveness detection model is used to perform liveness detection on the object to be detected, so as to obtain a first liveness detection result of the object to be detected;

[0009] Inputting the video to be detected into a second liveness detection model to obtain a second liveness detection result of the object to be detected;

[0010] A final liveness detection result of the object to be detected is determined according to the first liveness detection result and the second liveness detection result.

[0011] Optionally, the second light sequence reflected by each entity position point of the object to be detected represented by the video to be detected is obtained according to the following steps:

[0012] Extract a plurality of video frames of the video to be detected;

[0013] Align the pixel points describing the same entity position point of the object to be detected in the plurality of video frames;

[0014] For each entity position point of the object to be detected, according to the light reflected by the pixel points describing this entity position point in each of the plurality of video frames, obtain the second light sequence reflected by this entity position point.

[0015] Optionally, aligning the pixel points describing the same entity position point of the object to be detected in the plurality of video frames includes:

[0016] For each of the plurality of video frames, perform face key point detection on the video frame to obtain the face key points included in the video frame. According to the face key points included in each of the video frames, perform alignment processing at the face key point level on the plurality of video frames, and perform pixel point level alignment processing on the plurality of video frames after the alignment processing at the face key point level;

[0017] Or,

[0018] Perform pixel point level alignment processing on the plurality of video frames.

[0019] Optionally, the pixel point level alignment processing is performed through the following process:

[0020] Calculate the dense optical flow data between each of the plurality of video frames participating in the pixel point level alignment processing except the reference video frame and the reference video frame respectively, where the reference video frame is any one of the plurality of video frames participating in the pixel point level alignment processing;

[0021] According to the dense optical flow data between each of the other video frames and the reference video frame, perform pixel point level alignment processing on each of the other video frames and the reference video frame.

[0022] Optionally, both the first light sequence and the second light sequence include color light sequences of multiple color channels;

[0023] Generating a response map of the object to be detected according to the first light sequence and the second light sequence reflected by each entity position point of the object to be detected represented by the video to be detected includes:

[0024] Separate the first illumination sequence to obtain the first color light sequence corresponding to each color channel, and separate the second illumination sequence to obtain the second color light sequence corresponding to each color channel;

[0025] For each color channel, generate a response sub - graph of the object to be detected in this color channel according to the similarity between the second color light sequence of this color channel reflected by each entity position point of the object to be detected and the first color light sequence corresponding to this color channel;

[0026] Fuse the response sub - graphs of the object to be detected in each color channel to obtain the response graph of the object to be detected.

[0027] Optionally, fusing the response sub - graphs of the object to be detected in each color channel to obtain the response graph of the object to be detected includes:

[0028] According to the response sub - graphs of the object to be detected in each color channel, obtain the average response intensity of the face area of the object to be detected in each color channel;

[0029] According to the average response intensity of the face area of the object to be detected in each color channel, perform regularization processing on the response sub - graph of the object to be detected in this color channel to obtain the regularized response sub - graph of the object to be detected in each color channel;

[0030] Fuse the regularized response sub - graphs of the object to be detected in each color channel to obtain the response graph of the object to be detected.

[0031] Optionally, before processing the response graph using the first live detection model, it further includes:

[0032] Obtain the attribute value of the response graph of the object to be detected, where the attribute value includes at least one of the following: average response intensity, quality value;

[0033] Determine the live detection result of the object to be detected according to the size relationship between the attribute value of the response graph of the object to be detected and the corresponding attribute threshold;

[0034] In the case where the attribute value of the response graph of the object to be detected is less than the corresponding attribute threshold, determine that the object to be detected is not a live body;

[0035] In the case where the attribute value of the response graph of the object to be detected is not less than the corresponding attribute threshold, perform the step of performing live detection on the object to be detected using the first live detection model based on the response graph.

[0036] Optionally, the first live detection model is a model that has learned the first image features of the live face's response map and the second image features of the live face;

[0037] Based on the response map, using the first live detection model to perform live detection on the object to be detected, and obtaining the first live detection result of the object to be detected, including:

[0038] Extract the first image features of the response map;

[0039] Extract any video frame of the video to be detected, and obtain the second image features of this video frame;

[0040] Fuse the first image features and the second image features to obtain fused image features;

[0041] Use the first live model to process the fused image features to obtain the first live detection result of the object to be detected.

[0042] In a second aspect of the embodiments of the present application, a live detection method is provided, including:

[0043] Obtain a video to be detected of an object to be detected, where the video to be detected is: a video of the object to be detected collected during the irradiation of the object to be detected according to a first light sequence;

[0044] According to the first light sequence and the second light sequence reflected by each entity position point of the object to be detected represented by the video to be detected, generate a response map of the object to be detected, where the response intensity of each pixel point in the response map represents: the similarity between the second light sequence reflected by the entity position point corresponding to the pixel point and the first light sequence;

[0045] Based on the response map, use a live detection model to perform live detection on the object to be detected to obtain the live detection result of the object to be detected.

[0046] Optionally, the second light sequence reflected by each entity position point of the object to be detected represented by the video to be detected is obtained according to the following steps:

[0047] Extract multiple video frames of the video to be detected;

[0048] Align the pixel points describing the same entity position point of the object to be detected in the multiple video frames;

[0049] For each entity position point of the object to be detected, according to the light reflected by the pixel points describing this entity position point in each video frame among the multiple video frames, obtain the second light sequence reflected by this entity position point.

[0050] Optionally, aligning the pixel points describing the same entity position point of the to-be-detected object in the multiple video frames includes:

[0051] For each video frame in the multiple video frames, perform face key point detection on the video frame to obtain the face key points included in the video frame. According to the face key points included in each video frame, perform alignment processing at the face key point level on the multiple video frames, and perform pixel-level alignment processing on the multiple video frames after the alignment processing at the face key point level;

[0052] Or,

[0053] Perform pixel-level alignment processing on the multiple video frames.

[0054] Optionally, perform the pixel-level alignment processing through the following process:

[0055] Calculate the dense optical flow data between each video frame other than the reference video frame among the multiple video frames participating in the pixel-level alignment processing and the reference video frame, where the reference video frame is any one of the multiple video frames participating in the pixel-level alignment processing;

[0056] According to the dense optical flow data between each other video frame and the reference video frame, perform pixel-level alignment processing between each other video frame and the reference video frame.

[0057] Optionally, both the first illumination sequence and the second illumination sequence include color light sequences of multiple color channels;

[0058] Generating a response map of the to-be-detected object according to the first illumination sequence and the second illumination sequence reflected by each entity position point of the to-be-detected object represented by the to-be-detected video includes:

[0059] Separate the first illumination sequence to obtain the first color light sequence corresponding to each color channel, and separate the second illumination sequence to obtain the second color light sequence corresponding to each color channel;

[0060] For each color channel, generate a response sub-map of the to-be-detected object in this color channel according to the similarity between the second color light sequence of this color channel reflected by each entity position point of the to-be-detected object and the first color light sequence corresponding to this color channel;

[0061] Fuse the response sub-maps of the to-be-detected object in each color channel to obtain the response map of the to-be-detected object.

[0062] Optionally, fusing the response sub - graphs of the object to be detected in each color channel to obtain the response graph of the object to be detected, including:

[0063] According to the response sub - graphs of the object to be detected in each color channel, obtaining the average response intensity of the face region of the object to be detected in each color channel;

[0064] According to the average response intensity of the face region of the object to be detected in each color channel, performing regularization processing on the response sub - graph of the object to be detected in this color channel to obtain the regularized response sub - graph of the object to be detected in each color channel;

[0065] Fusing the regularized response sub - graphs of the object to be detected in each color channel to obtain the response graph of the object to be detected.

[0066] Optionally, before processing the response graph using the live detection model, it further includes:

[0067] Obtaining the attribute value of the response graph of the object to be detected, where the attribute value includes at least one of the following: average response intensity, quality value;

[0068] Determining the live detection result of the object to be detected according to the magnitude relationship between the attribute value of the response graph of the object to be detected and the corresponding attribute threshold;

[0069] In the case where the attribute value of the response graph of the object to be detected is less than the corresponding attribute threshold, determining that the object to be detected is not a live body;

[0070] In the case where the attribute value of the response graph of the object to be detected is not less than the corresponding attribute threshold, performing the step of live detection on the object to be detected using the live detection model based on the response graph.

[0071] Optionally, the live detection model is a model that has learned the first image feature of the response graph of a live face and the second image feature of the live face;

[0072] Performing live detection on the object to be detected using the live detection model based on the response graph to obtain the live detection result of the object to be detected, including:

[0073] Extracting the first image feature of the response graph;

[0074] Extracting any video frame of the video to be detected and obtaining the second image feature of this video frame;

[0075] Fusing the first image feature and the second image feature to obtain a fused image feature;

[0076] The fused image features are processed using the living body model to obtain a living body detection result of the object to be detected.

[0077] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the liveness detection method as described in the first aspect; or, the processor executes the computer program to implement the liveness detection method as described in the second aspect.

[0078] In a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the liveness detection method as described in the first aspect is implemented; or, when the computer program / instruction is executed by a processor, the liveness detection method as described in the second aspect is implemented.

[0079] In a fifth aspect of an embodiment of the present application, a computer program product is provided, including a computer program / instruction, which, when executed by a processor, implements the liveness detection method as described in the first aspect; or, when executed by a processor, the computer program / instruction implements the liveness detection method as described in the second aspect.

[0080] The embodiments of the present application include the following advantages:

[0081] In this embodiment, the response map of the object to be detected is specific to the pixel level. According to the similarity between the second illumination sequence and the first illumination sequence reflected by the physical position points of the object to be detected represented by each pixel of the response map, different reflection patterns presented by each physical position area of the object to be detected can be obtained. The face of a living person is uneven, and the screen or printing paper used for the photocopy is relatively smooth and has a high reflectivity. Therefore, the light reflected by the living face and the photocopy has different patterns. Therefore, based on the response map of the object to be detected, it is possible to distinguish whether the object to be detected is a photocopy or a living face, thereby obtaining the first liveness detection result of the object to be detected. That is, in the colorful liveness detection method, the introduction of the feature of reflected light can further realize photocopy attack detection on the basis of colorful liveness detection, effectively improving the accuracy and detection capability of liveness detection;

[0082] In addition, the second liveness detection model can also be used to obtain liveness detection results of other attack types of the object to be detected. Combining the first liveness detection model with the second liveness detection model can effectively improve the ability to detect other attacks while ensuring that the probability of a real person being judged as a live person is not reduced, thereby effectively improving the accuracy of the final liveness detection result. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0084] Figure 1 is a flowchart of the steps of a living body detection method in an embodiment of the present application;

[0085] Figure 2 is a schematic flow chart of obtaining the first living body detection result in an embodiment of the present application;

[0086] Figure 3 is a schematic flow chart of living body detection in an embodiment of the present application;

[0087] Figure 4 is a flowchart of the steps of a living body detection method in an embodiment of the present application;

[0088] Figure 5 is a schematic structural diagram of a living body detection device in an embodiment of the present application;

[0089] Figure 6 is a schematic structural diagram of a living body detection device in an embodiment of the present application;

[0090] Figure 7 is a schematic diagram of an electronic device in an embodiment of the present application. Detailed implementation manners

[0091] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the drawings and specific implementation manners.

[0092] In recent years, research on computer vision, deep learning, machine learning, image processing, image recognition and other technologies based on artificial intelligence has made important progress. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies and application systems for simulating and extending human intelligence. Artificial intelligence is a comprehensive discipline involving many types of technologies such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, neural networks, etc. Computer vision, as an important branch of artificial intelligence, specifically allows machines to recognize the world. Computer vision technology usually includes face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, target detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, robot navigation and positioning and other technologies. With the research and advancement of artificial intelligence technology, this technology has been applied in many fields, such as security control, urban management, traffic management, building management, park management, facial access, facial attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart homes, wearable devices, unmanned driving, automatic driving, smart medical care, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile Internet, live streaming, beauty, makeup, medical beauty, and smart temperature measurement.

[0093] In the related technology, when performing liveness detection, it is necessary to judge whether there is camera hijacking by comparing the similarity between the emitted light sequence and the light sequence reflected by the collected video; however, judging only based on the similarity between the light sequences may cause the printed face to reflect light, and the light sequence reflected by the printed face and the emitted light sequence also have a high similarity. Therefore, the related liveness detection technology also needs to judge whether there is liveness attack behavior based on the collected video, such as judging whether it is a screen copy or a printed paper copy.

[0094] However, the related liveness detection technology separates the similarity detection and liveness attack behavior detection. The applicant proposes that the similarity detection and liveness attack behavior detection can be combined based on the information that the light reflected when shining on a real person's face and on a screen / printed paper has different reflection patterns. While considering the similarity, the difference in reflected light from a real person and a screen / printed paper is also considered. Because the human face is uneven, different areas reflect different light, and the background area behind the face can only reflect weak light or even no light because of the low intensity of light received; while the screen / printed paper is used for reshooting, the screen / printed paper is relatively smooth and flat, so the reflected light may be more uniform and regular. Therefore, the applicant thought that the information that the live face and the screen / printed paper have different reflection patterns can be used to improve the accuracy of colorful liveness detection.

[0095] Reference Figure 1 As shown, a flowchart of a method for detecting a living body according to an embodiment of the present application is shown. Figure 1 As shown, the liveness detection method can be applied to a background server, including the following steps:

[0096] Step S11: obtaining a video of the object to be detected, wherein the video to be detected is: a video of the object to be detected collected during the period when the object to be detected is irradiated according to a first illumination sequence;

[0097] Step S12: generating a response map of the object to be detected according to the first illumination sequence and the second illumination sequence reflected by each physical position point of the object to be detected represented by the video to be detected, wherein the response intensity of each pixel point in the response map represents: the similarity between the second illumination sequence reflected by the physical position point corresponding to the pixel point and the first illumination sequence;

[0098] Step S13: Based on the response graph, using a first liveness detection model, performing liveness detection on the object to be detected, to obtain a first liveness detection result of the object to be detected;

[0099] Step S14: inputting the video to be detected into a second liveness detection model to obtain a second liveness detection result of the object to be detected;

[0100] Step S15: determining a final liveness detection result of the object to be detected according to the first liveness detection result and the second liveness detection result.

[0101] In specific implementation, the first light sequence can be sent by the background server to the terminal. The terminal emits light according to the first light sequence to irradiate the object to be detected, and during the process of irradiating the object to be detected according to the first light sequence, it collects the video to be detected of the object to be detected. Among them, the object to be detected is the object captured by the camera of the terminal.

[0102] Optionally, in some specific implementation manners, the solution of the present application can also be executed by an electronic device such as a terminal. For example, when performing a live detection, the terminal generates the first light sequence by itself, and collects the video of the object to be detected during the process of irradiating the object to be detected according to the first light sequence; then the terminal itself executes the subsequent live detection process according to the first light sequence and the video to be detected. Whether the specific first light sequence is sent by the background server or generated by the terminal or other electronic devices by themselves, and whether the specific live detection process is executed by the background server or by the terminal or other electronic devices by themselves, or even some steps in the live detection process are executed by the terminal or other electronic devices and some steps are executed by the background server, etc. can all be set according to actual requirements. The embodiments of the present application do not limit this, nor will all possible implementation manners be listed one by one.

[0103] Process each video frame of the video to be detected, determine the pixel points corresponding to the physical position points of the object to be detected in each video frame, and obtain the light reflected by the corresponding pixel points in each video frame, so as to obtain the second light sequence reflected by each physical position point of the object to be detected. A detection position point of the object to be detected refers to a point where the object to be detected actually exists. For example, when the object to be detected is a human face, a physical position point of the object to be detected can be a point on the nose of the human face, and the size of this point is the size represented by a pixel point in the video.

[0104] Calculate the similarity between the first light sequence and the second light sequence reflected by each physical position point of the object to be detected, and use the similarity as the response intensity of the response map, so as to generate the response map of the object to be detected. Optionally, when the light sequence is white light, each element in the first light sequence and the second light sequence can represent the light intensity of the white light; when the light sequence is colored light, each element in the first light sequence and the second light sequence can represent the light intensity of the colored light and / or the color of the colored light. Among them, the similarity between the first light sequence and the second light sequence can be obtained through the dot product of the two sequences. Since the response map of the object to be detected is refined to the pixel point level, different reflection patterns presented by different position regions of the object to be detected can be obtained through the response map of the object to be detected.

[0105] The response graph of the object to be detected is input into the first liveness detection model. The first liveness detection model performs liveness detection on the object to be detected according to the response graph of the object to be detected, and a first liveness detection result of the object to be detected can be obtained. The first liveness detection model is a model that has learned the first image feature of the response graph of a live face through supervised training, and can distinguish the response graph of a live face from the response graph of other attacks (for example, a photocopy attack). Therefore, the first liveness detection model can obtain the first liveness detection result of the object to be detected through the response graph of the object to be detected. Among them, the supervised training performed by the first liveness detection model can be: obtaining the response graphs of multiple sample objects (including live faces and other objects), inputting the response graphs of the sample objects into the first liveness detection model to be trained, and obtaining the predicted probability that the sample objects are live; according to the predicted probability and whether the sample objects are actually live, a loss function is established, and the model parameters of the first liveness detection model to be trained are updated based on the loss function to obtain the first liveness detection model. In this way, the first liveness detection model can learn the response graph of live faces. The method for obtaining the response graph of the sample object may refer to the method for obtaining the response graph of the object to be detected.

[0106] In order to detect more types of attacks at the same time and improve the accuracy of the liveness detection results, a second liveness detection model can also be used to obtain a second liveness detection result of the object to be detected through the video to be detected. Among them, the second liveness detection model can be a commonly used model for liveness detection in the relevant technology, and the second liveness detection model can perform liveness detection based on the input video or video frame. Optionally, in the specific implementation, the second liveness detection model can be any type of liveness detection model, such as a mask attack liveness detection model, an action detection liveness model, etc.; for some liveness detection models, the user may be required to perform corresponding actions. Therefore, in order to obtain the first liveness detection result and the second liveness detection result based on the video to be detected, when recording the video, the user can be instructed to perform corresponding actions according to the requirements of the second liveness detection model. In a specific application scenario, it can be set according to the actual needs of the specific application scenario, and the embodiment of the present application is not limited to this.

[0107] Optionally, the first liveness detection model and the second liveness detection model may work in parallel to obtain the first liveness detection result and the second liveness detection result in parallel; the first liveness detection model may first obtain the first liveness detection result, and then the second liveness detection model may obtain the second liveness detection result; or the second liveness detection model may first obtain the second liveness detection result, and then the first liveness detection model may obtain the first liveness detection result.

[0108] By combining the first liveness detection result and the second liveness detection result, the final liveness detection result of the object to be detected can be obtained. Optionally, the first liveness detection result and the second liveness detection result can respectively represent the probability of the object to be detected being a live face. By comparing the smaller value of the probability represented by the first liveness detection result and the probability represented by the second liveness detection result, and the size relationship of the liveness threshold; when the smaller value is greater than the liveness threshold, the final liveness detection result of the object to be detected is determined to be: live; when the smaller value is not greater than the liveness threshold, the final liveness detection result of the object to be detected is determined to be: not live. Among them, the liveness threshold can be a pre-set and reasonable value. Optionally, different weights can be set for the first liveness detection result and the second liveness detection result, and the weighted first liveness detection result and the weighted second liveness detection result are combined to obtain the final liveness detection result of the object to be detected.

[0109] In this way, the ability to detect other attacks can be effectively improved while ensuring that the probability of a real person being judged as a living person is not reduced.

[0110] By adopting the technical solution of the embodiment of the present application, the response map of the object to be detected is specific to the pixel level. According to the similarity between the second illumination sequence and the first illumination sequence reflected by the physical position point of the object to be detected represented by each pixel of the response map, different reflection modes presented by each physical position area of the object to be detected can be obtained. The living face is uneven, and the screen or printing paper used for the remake is relatively smooth and has a high reflectivity. Therefore, the light reflected by the living face and the remake has different modes. Therefore, the first liveness detection model can distinguish whether the object to be detected is a remake or a living face based on the response map of the object to be detected, thereby obtaining the first liveness detection result of the object to be detected. In this way, the first liveness detection model utilizes the information that the remake and the living face have different reflection modes, realizes the combination of two detection methods (lighting sequence inspection and detection whether it is a remake), and the determined first liveness detection result is more accurate. In addition, the second liveness detection model can also obtain the second liveness detection result of the object to be detected. The final liveness detection result of the object to be detected determined by combining the first liveness detection result and the second liveness detection result can effectively improve the ability to detect other attacks while ensuring that the probability of a real person being judged as a live person is not reduced, thereby effectively improving the accuracy of the final liveness detection result.

[0111] Considering that there are still attacks such as wearing a mask, since the mask structure is similar to the human face structure, the reflected light and the generated response map are similar. And the first live detection model that only learns the first image feature of the response map of the live human face may have difficulty in recognizing such attacks. Therefore, when performing supervised training on the first live detection model, the first live detection model can also be made to learn the second image feature of the live human face. Optionally, the supervised training performed by the first live detection model can be: obtaining the response maps of multiple sample objects (including live human faces and other objects), and the video of each collected sample object, extracting any video frame from the video, and the first live detection model to be trained is based on the response map and the video frame, as well as the information on whether the sample object is actually a live body, to perform supervised training. When the model structure of the first live detection model only allows one input, the first image feature of the response map and the second image feature of the video frame can be extracted; the first image feature of the response map and the second image feature of the video frame are fused, and the fused feature is input into the first live detection model to be trained to obtain the first live detection result of the sample object.

[0112] Correspondingly, the first live detection model that has learned the image feature of the live human face can obtain the first live detection result of the object to be detected according to the fused image feature. Among them, the method for obtaining the fused image feature is: extracting the first image feature of the response map; extracting any video frame of the video to be detected and obtaining the second image feature of the video frame; fusing the first image feature and the second image feature to obtain the fused image feature.

[0113] The feature obtained by fusing the image feature of the response map of the object to be detected with the image feature of any video frame extracted from the video of the object to be detected.

[0114] Adopting the technical solution of the embodiment of the present application, the first live detection model that has learned both the response map of the live human face and the image feature of the live human face can avoid attacks such as wearing a mask, thereby improving the accuracy of the first live detection.

[0115] On the basis of the above technical solution, the second light sequence reflected by each entity position point of the object to be detected represented by the video to be detected is obtained according to the following steps: extracting multiple video frames of the video to be detected; aligning the pixel points describing the same entity position point of the object to be detected in the multiple video frames; for each entity position point of the object to be detected, according to the light reflected by the pixel points describing the entity position point in each video frame among the multiple video frames, obtaining the second light sequence reflected by the entity position point.

[0116] The video to be detected has multiple video frames. The position of an entity position point of the object to be detected may be different in each video frame. In order to obtain the second light sequence reflected by each entity position point of the object to be detected, it is necessary to align the multiple video frames to obtain the pixel points corresponding to any entity position point of the object to be detected in each video frame. According to the light reflected by the pixel points corresponding to any entity position point of the object to be detected in each video frame, and the time sequence between each video frame, the second light sequence reflected by each entity position point of the object to be detected can be obtained.

[0117] For example, the light is red light, and there are a total of 5 video frames. The pixel points of an entity position point of the object to be detected in each video frame are A, B, C, D, and E respectively. If points A, C, and D all reflect red light, while points B and E do not reflect red light, then the second entity sequence of the red light reflected by this entity position point of the object to be detected can be 10110. It can be understood that according to the intensity of the reflected light, the numbers in the second entity sequence can also be numbers between 0 and 1, and the numerical size is related to the intensity of the reflected light.

[0118] Adopting the technical solution of the embodiment of the present application can solve the problem that it is difficult to obtain the second light sequence reflected by each entity position point of the object to be detected due to the different positions of the object to be detected in each video frame.

[0119] The alignment processing of multiple video frames can be implemented based on facial key points. First, facial key point detection is performed on each video frame in the multiple video frames of the video to be detected to obtain the facial key points included in each video frame. The present application does not require a specific detection method for facial key point detection, and relevant facial key point detection algorithms or software can be used; other methods with shorter time consumption can also be used. For example, after obtaining the facial key points in the first video frame, five points, namely the corners of the left and right eyes, the tip of the nose, and the corners of the left and right lips, are used as anchor points, and the remaining video frames are mapped to the template through the Thin Plate Spline interpolation algorithm. Among them, the anchor points can be customized according to the input points of the facial key points, as long as it is ensured that the positions of some key points are fixed after rough alignment. The more the number of anchor points, the better the alignment effect, but 5 points are sufficient to meet the requirements. It can be understood that when no face is detected, the object to be detected can be directly determined as non-living.

[0120] According to the facial key points included in each video frame, the alignment processing at the facial key point level can be performed on each video frame. Further, for the multiple video frames after the alignment processing at the facial key point level, the alignment processing at the pixel point level is performed.

[0121] Optionally, the alignment processing at the pixel point level can also be directly performed on the multiple video frames.

[0122] Performing alignment processing at the pixel level directly can make the alignment process relatively simple; performing alignment processing at the facial key point level first and then at the pixel level can save the computing resources consumed during pixel-level alignment processing. How to specifically implement pixel-level alignment processing can be selected according to actual needs.

[0123] Optionally, fine alignment at the pixel level can be implemented through the following process: Calculate the dense optical flow data between each video frame other than the reference video frame and the reference video frame in multiple video frames participating in the pixel-level alignment processing, where the reference video frame is any one of the multiple video frames participating in the pixel-level alignment processing; Based on the dense optical flow data between each other video frame and the reference video frame, perform pixel-level alignment processing between each other video frame and the reference video frame.

[0124] The first video frame can be used as the reference video frame, and the dense optical flow data between each other video frame and the reference video frame is calculated in the luminance channel. Using the dense optical flow data as the mapping basis, perform fine alignment at the pixel level between each other video frame and the reference video frame. Among them, pixel-level alignment processing can be implemented based on the dense optical flow algorithm, such as the Gunnar Farneback algorithm (a dense optical flow algorithm).

[0125] In this way, fine alignment at the pixel level is achieved, and only then can the second light sequence reflected by each entity position point of the object to be detected be accurately obtained, and then an accurate response map can be generated to obtain an accurate first live detection result.

[0126] Figure 2 The flow diagram of obtaining the first live detection result is shown. The steps of obtaining the first live detection result may include: alignment processing at the facial key point level, alignment processing at the pixel level, obtaining the second light sequence, similarity calculation, generating a response map, obtaining video frames, inputting into the first live detection model, and obtaining the first live detection result. The above several steps can form a relatively complete process, but according to actual needs, one or more of these steps can be discarded. For example, the step of alignment processing at the facial key point level can be discarded to reduce the complexity of the entire process; the step of obtaining video frames can be discarded, and correspondingly, only the response map is input into the first live detection model. For the recognition of attacks such as wearing a mask, it can be achieved through the second live model to avoid duplicate work between the first live detection model and the second live detection model; The steps of alignment processing at the facial key point level and obtaining video frames can also be discarded simultaneously, reducing the complexity of the entire process while avoiding duplicate work between the first live detection model and the second live detection model.

[0127] Based on the above technical solution, both the first light sequence and the second light sequence can be color light sequences including multiple color channels. When calculating the similarity between the first light sequence and the second light sequence, the similarity between the first color light sequence and the second color light sequence in each color channel is calculated, and then the response subgraph of this color channel is obtained.

[0128] In order to obtain the color light sequences reflected by each color channel, the first light sequence and the second light sequence need to be separated respectively. Separating the first light sequence can obtain the first color light sequence corresponding to each color channel, and separating the second light sequence can obtain the second color light sequence corresponding to each color channel. When separating the first light sequence and the second light sequence, in order to improve the separation accuracy, each color channel can be regularized. The formula x′ t = x t -(∑x t ) / n can be used to realize the regularization of each color channel, where x′ t represents the regularized color light sequence (which can be the first color light sequence or the second color light sequence), x t represents the color light sequence before regularization, and n represents the length of the color light sequence.

[0129] The method of regularizing each color channel can also refer to the related technology, and this application does not limit it. Among them, since the first light sequence is sent from the background, the first color light sequence corresponding to each color channel can be directly obtained from the background.

[0130] For example, taking one second as the unit between the color light sequences, and the multiple color channels being the red channel, the yellow channel, and the blue channel, the emitted color light is red light in the first second, orange light composed of red light and yellow light in the second second, white light composed of yellow light and blue light in the third second, yellow light in the fourth second, and red light in the fifth second. Using 1 to represent the existence of light in the corresponding color channel and 0 to represent the non-existence of light in the corresponding color channel, the obtained first color light sequence can be: the first color light sequence of the red channel 11001, the first color light sequence of the yellow channel 01110, and the first color light sequence of the blue channel 00100. According to a similar principle, the second color light sequences corresponding to each color channel reflected by each entity position point of the object to be detected can be obtained. It can be understood that according to the intensity of the sent light, the numbers in the first color light sequence can be numbers between 0 and 1; according to the intensity of the reflected light, the numbers in the second color light sequence can be numbers between 0 and 1.

[0131] For each color channel, if the similarity between the second color light sequence of a color channel reflected by an entity position point of the object to be detected and the first color light sequence corresponding to this color channel is greater than a preset value, then the color of the pixel point corresponding to this entity position point in the response sub-map of this color channel is the color of this color channel. After obtaining the response sub-maps of each color channel, the response sub-maps of each color channel are fused to obtain the response map of the object to be detected.

[0132] For example, the color of the pixel point corresponding to an entity position point of the object to be detected in the response map of the red channel is red (similarity greater than the preset value), the color of the pixel point corresponding to it in the response map of the yellow channel is yellow (similarity greater than the preset value), and the color of the pixel point corresponding to it in the response map of the blue channel is none (similarity not greater than the preset value). Then, after fusing the response sub-maps of each color channel, the color of the corresponding pixel point in the response map of the object to be detected is orange composed of red and yellow.

[0133] In this way, for each color channel, the response sub-map of each color channel is obtained, and then the response map of the object to be detected is obtained based on the response sub-maps of each color channel. Compared with directly obtaining the response map of the object to be detected according to the color light integrating all color channels, it has the advantage of being more accurate.

[0134] Based on the above technical solution, in each color channel, in order to avoid the adverse effect on the response sub-map caused by the too high brightness of the response sub-map of this color channel, the response intensity mean value of the face area in this response sub-map can be used to perform regularization processing on the response sub-map of this color channel to obtain the regularized response sub-map of the object to be detected in this color channel. The regularized response sub-maps of the object to be detected in each color channel are fused to obtain the response map of the object to be detected.

[0135] Optionally, using the response intensity mean value of the face area in the response sub-map of each color channel to perform regularization processing on the response sub-map of this color channel can be: dividing the response intensity mean value of the face area of the response sub-map of this color channel by the response intensity of each pixel point of the response sub-map of this color channel, and using the obtained quotient as the response intensity of each pixel point of the regularized response sub-map of this color channel.

[0136] Among them, the response intensity of the regularized response map can be calculated by the following formula:

[0137]

[0138] Among them, F represents the face area, r i,j represents the response intensity of the pixel point (i, j) of the response map; n represents the number of pixel points, and N F represents the response intensity mean value of the face area.

[0139] The finally obtained response map is Resp′ = {r′ i,j}, where r′ i,j represents the response intensity of the pixel point (i, j) of the response map after regularization, and can be calculated by the formula obtained through calculation.

[0140] The fusion of the response sub-maps after regularization for each color channel can be achieved through the following formula:

[0141] ri i,j = uint8(min(max(r′ i,j * 255, 0), 255))

[0142] where ri i,j represents the response intensity of the pixel point (i, j) of the fused response map, unit8 means converting the floating-point number into an 8-bit non-negative integer, and r′ i,j represents the response intensity of the pixel point (i, j) of the response map after regularization.

[0143] In this way, the response intensities of the response sub-maps of each color channel can be made relatively balanced, avoiding the occurrence of over-bright scenes.

[0144] Optionally, based on the above technical solution, after obtaining the response map of the object to be detected and before inputting the response map of the object to be detected into the first live detection model, the live detection result of the object to be detected can be determined according to at least one attribute value of the response map of the object to be detected.

[0145] The attribute values of the response map of the object to be detected include at least one of the following: the mean response intensity, the quality value. In the case where any one of the attribute values is less than the corresponding attribute threshold, it is determined that the object to be detected is not a live body. In the case where each attribute value is not less than the corresponding attribute threshold, the response map of the object to be detected is input into the first live detection model.

[0146] Figure 3 Shows a schematic flow diagram of live detection. Before inputting the response map of the object to be detected into the first live detection model, first determine whether the attribute value of the response map is less than the corresponding attribute threshold. In the case where any one of the attribute values is less than the corresponding attribute threshold, it is determined that the object to be detected is not a live body, and the detection result of "not a live body" can be directly output. Otherwise, the response map of the object to be detected is input into the first live detection model for live detection.

[0147] If the average response intensity of the response map of the object to be detected is less than the response intensity threshold, it can be considered that the second color light sequence collected is too weak. Therefore, this response map is not sufficient as a clue for live verification. For security considerations, it can be considered that there is an attack, and it is directly determined that the object to be detected is not a live body. Therefore, it is not necessary to input the response map of the object to be detected into the first live body detection model.

[0148] If the quality value of the response map of the object to be detected is less than the quality threshold, there may be an attack, and it is directly determined that the object to be detected is not a live body. Therefore, it is not necessary to input the response map of the object to be detected into the first live body detection model. Among them, the quality value of the response map of the object to be detected can be determined according to the noise in the response map. The more noise there is, the lower the quality value.

[0149] The quality value of the response map can be calculated by the following formula:

[0150]

[0151] where quality is the quality value, r i,j represents the response intensity of the pixel point (i, j) of the response map, t represents each element in the light sequence, and y′ t represents the first light sequence after color channel regularization processing, and x′ t represents the second light sequence after color channel regularization processing.

[0152] In this way, some attacks can be avoided and the accuracy of the live body detection result can be improved.

[0153] Refer to Figure 4 As shown, the flowchart of the steps of a live body detection method in an embodiment of the present application is shown. As Figure 4 shown, this live body detection method can be applied to a background server and includes the following steps:

[0154] Step S41: Obtain the video to be detected of the object to be detected. The video to be detected is: the video of the object to be detected collected during the irradiation of the object to be detected according to the first light sequence;

[0155] Step S42: Generate a response map of the object to be detected according to the first light sequence and the second light sequence reflected by each entity position point of the object to be detected represented by the video to be detected. The response intensity of each pixel point in the response map represents: the similarity between the second light sequence reflected by the entity position point corresponding to the pixel point and the first light sequence;

[0156] Step S43: Based on the response map, use a live body detection model to perform live body detection on the object to be detected to obtain the live body detection result of the object to be detected.

[0157] The method for obtaining a video of an object to be detected and generating a response graph of the object to be detected can refer to the method for obtaining a video of an object to be detected and generating a response graph of the object to be detected described above; the training method of a liveness detection model can refer to the training method of the first liveness detection model.

[0158] By inputting the response graph of the object to be detected into the liveness detection model, the liveness detection result of the object to be detected can be obtained.

[0159] By adopting the technical solution of the embodiment of the present application, the response map of the detected object is specific to the pixel level. According to the similarity between the second illumination sequence and the first illumination sequence reflected by the physical position point of the object to be detected represented by each pixel of the response map, different reflection modes presented by each physical position area of the object to be detected can be obtained. The living face is uneven, and the screen or printing paper used for the remake is relatively smooth and has a high reflectivity. Therefore, the light reflected by the living face and the remake has different modes. Therefore, the liveness detection model can distinguish whether the object to be detected is a remake or a living face based on the response map of the object to be detected, thereby obtaining the liveness detection result of the object to be detected. In this way, the liveness detection model utilizes the information that the remake and the living face have different reflection modes, realizes the combination of two detection methods (lighting sequence inspection and detection whether it is a remake), and the determined liveness detection result is more accurate.

[0160] Optionally, the step of obtaining the second illumination sequence reflected by each physical position point of the object to be detected represented by the video to be detected is obtained according to the following steps:

[0161] Extract multiple video frames of the video to be detected; align the pixel points describing the same entity position point of the object to be detected in the multiple video frames; for each entity position point of the object to be detected, obtain a second illumination sequence reflected by the entity position point based on the illumination reflected by the pixel points describing the entity position point in each video frame in the multiple video frames.

[0162] Optionally, aligning the pixel points describing the same entity position point of the object to be detected in the multiple video frames specifically includes the following process:

[0163] For each video frame in the multiple video frames, perform facial key point detection on the video frame to obtain the facial key points contained in the video frame, perform facial key point level alignment processing on the multiple video frames according to the facial key points contained in each of the video frames, and perform pixel point level alignment processing on the multiple video frames after the facial key point level alignment processing;

[0164] or,

[0165] Perform pixel-level alignment processing on the multiple video frames.

[0166] Optionally, the pixel-level alignment processing can be performed through the following process:

[0167] Calculate the dense optical flow data between each of the multiple video frames participating in the pixel-level alignment processing, except the reference video frame, and the reference video frame, where the reference video frame is any one of the multiple video frames participating in the pixel-level alignment processing;

[0168] Perform pixel-level alignment processing on each of the other video frames and the reference video frame according to the dense optical flow data between each of the other video frames and the reference video frame.

[0169] Optionally, both the first illumination sequence and the second illumination sequence may include color light sequences of multiple color channels; generate a response map of the object to be detected according to the first illumination sequence and the second illumination sequence reflected by each entity position point of the object to be detected represented by the video to be detected, including: separating the first illumination sequence to obtain the first color light sequence corresponding to each color channel, and separating the second illumination sequence to obtain the second color light sequence corresponding to each color channel; for each color channel, generate a response sub-map of the object to be detected in this color channel according to the similarity between the second color light sequence of this color channel reflected by each entity position point of the object to be detected and the first color light sequence corresponding to this color channel; fuse the response sub-maps of the object to be detected in each color channel to obtain the response map of the object to be detected. The specific steps can refer to the foregoing.

[0170] Optionally, fusing the response sub-maps of the object to be detected in each color channel to obtain the response map of the object to be detected includes: obtaining the average response intensity of the face region of the object to be detected in each color channel according to the response sub-maps of the object to be detected in each color channel; performing regularization processing on the response sub-map of the object to be detected in this color channel according to the average response intensity of the face region of the object to be detected in each color channel to obtain the regularized response sub-map of the object to be detected in each color channel; fusing the regularized response sub-maps of the object to be detected in each color channel to obtain the response map of the object to be detected. Among them, fusing the response sub-maps of the object to be detected in each color channel to obtain the response map of the object to be detected includes:

[0171] Based on the response sub - graphs of the object to be detected in each color channel, obtain the mean response intensity of the face region of the object to be detected in each color channel; according to the mean response intensity of the face region of the object to be detected in each color channel, perform regularization processing on the response sub - graph of the object to be detected in this color channel to obtain the regularized response sub - graph of the object to be detected in each color channel; fuse the regularized response sub - graphs of the object to be detected in each color channel to obtain the response graph of the object to be detected.

[0172] Optionally, before processing the response graph using the live detection model, it further includes: obtaining the attribute value of the response graph of the object to be detected, where the attribute value includes at least one of the following: mean response intensity, quality value; determining the live detection result of the object to be detected according to the magnitude relationship between the attribute value of the response graph of the object to be detected and the corresponding attribute threshold; in the case where the attribute value of the response graph of the object to be detected is less than the corresponding attribute threshold, determine that the object to be detected is not a live body; in the case where the attribute value of the response graph of the object to be detected is not less than the corresponding attribute threshold, perform the step of performing live detection on the object to be detected using the live detection model based on the response graph. The specific steps can refer to the description in the previous text.

[0173] Optionally, the live detection model is a model that has learned the first image feature of the response graph of a live face and the second image feature of the live face; performing live detection on the object to be detected using the live detection model based on the response graph to obtain the live detection result of the object to be detected includes: extracting the first image feature of the response graph; extracting any video frame of the video to be detected and obtaining the second image feature of this video frame; fusing the first image feature and the second image feature to obtain a fused image feature; using the live model to process the fused image feature to obtain the live detection result of the object to be detected. The specific steps can refer to the description in the previous text.

[0174] Among them, the specific implementation process of each step in the live detection method provided in the embodiments of the present application can refer to the introduction in the previous method embodiments and will not be elaborated here.

[0175] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present application.

[0176] Figure 5 It is a schematic structural diagram of a living body detection device according to an embodiment of the present application. As Figure 5 shown, the living body detection device includes a video acquisition module 51, a response map generation module 52, a first living body detection result acquisition 53, a second living body detection result acquisition 54, and a final detection result determination module 55, where:

[0177] The video acquisition module 51 is configured to acquire a video to be detected of an object to be detected. The video to be detected is: a video of the object to be detected acquired during irradiating the object to be detected according to a first light sequence;

[0178] The response map generation module 52 is configured to generate a response map of the object to be detected according to the first light sequence and a second light sequence reflected by each entity position point of the object to be detected represented by the video to be detected. The response intensity of each pixel point in the response map represents: the similarity between the second light sequence reflected by the entity position point corresponding to the pixel point and the first light sequence;

[0179] The first living body detection result acquisition module 53 is configured to perform living body detection on the object to be detected based on the response map by using a first living body detection model, and obtain a first living body detection result of the object to be detected;

[0180] The second living body detection result acquisition module 54 is configured to input the video to be detected into a second living body detection model, and obtain a second living body detection result of the object to be detected;

[0181] The final living body detection result determination module 55 is configured to determine a final living body detection result of the object to be detected according to the first living body detection result and the second living body detection result.

[0182] Optionally, the second light sequence reflected by each entity position point of the object to be detected represented by the video to be detected is obtained according to the following steps:

[0183] Extract a plurality of video frames of the video to be detected;

[0184] Align the pixel points describing the same entity position point of the object to be detected in the plurality of video frames;

[0185] For each entity position point of the object to be detected, according to the light reflected by the pixel points describing the entity position point in each of the plurality of video frames, obtain the second light sequence reflected by the entity position point.

[0186] Optionally, aligning the pixel points describing the same entity position point of the object to be detected in the plurality of video frames includes:

[0187] For each of the multiple video frames, perform face key point detection on the video frame to obtain the face key points included in each video frame. According to the face key points included in each video frame, perform alignment processing at the face key point level on the multiple video frames, and perform alignment processing at the pixel point level on the multiple video frames after the alignment processing at the face key point level;

[0188] Alternatively, perform alignment processing at the pixel point level on the multiple video frames.

[0189] Optionally, perform the alignment processing at the pixel point level through the following process:

[0190] Calculate the dense optical flow data between each of the multiple video frames participating in the alignment processing at the pixel point level and the reference video frame, except for the reference video frame. The reference video frame is any one of the multiple video frames participating in the alignment processing at the pixel point level;

[0191] According to the dense optical flow data between each of the other video frames and the reference video frame, perform alignment processing at the pixel point level between each of the other video frames and the reference video frame.

[0192] Optionally, both the first illumination sequence and the second illumination sequence include color light sequences of multiple color channels; the response map generation module 52 includes:

[0193] A first separation unit for separating the first illumination sequence to obtain the first color light sequence corresponding to each color channel, and separating the second illumination sequence to obtain the second color light sequence corresponding to each color channel;

[0194] A first response sub-map generation unit for generating a response sub-map of the object to be detected in each color channel according to the similarity between the second color light sequence of each color channel reflected by each entity position point of the object to be detected and the first color light sequence corresponding to the color channel;

[0195] A first fusion unit for fusing the response sub-maps of the object to be detected in each color channel to obtain the response map of the object to be detected.

[0196] Optionally, the first fusion unit includes:

[0197] A first mean value acquisition sub-unit for acquiring the mean response intensity of the face region of the object to be detected in each color channel according to the response sub-maps of the object to be detected in each color channel;

[0198] A first regularization subunit, configured to regularize the response sub-graph of the object to be detected in each color channel according to the mean response intensity of the face region of the object to be detected in each color channel, so as to obtain a regularized response sub-graph of the object to be detected in each color channel;

[0199] A first fusion subunit, configured to fuse the regularized response sub-graphs of the object to be detected in each color channel to obtain a response graph of the object to be detected.

[0200] Optionally, before processing the response graph by using the first live detection model, the apparatus further includes:

[0201] A first attribute value acquisition module, configured to acquire an attribute value of the response graph of the object to be detected, where the attribute value includes at least one of the following: mean response intensity, quality value;

[0202] A first determination module, configured to determine a live detection result of the object to be detected according to a magnitude relationship between the attribute value of the response graph of the object to be detected and a corresponding attribute threshold;

[0203] A first determination module, configured to determine that the object to be detected is not a live body when the attribute value of the response graph of the object to be detected is less than the corresponding attribute threshold;

[0204] A first step execution module, configured to execute a step of performing live detection on the object to be detected by using the first live detection model based on the response graph when the attribute value of the response graph of the object to be detected is not less than the corresponding attribute threshold.

[0205] Optionally, the first live detection model is a model that has learned the first image feature of the response graph of a live face and the second image feature of the live face; the first live detection result acquisition module 53 includes:

[0206] A first response graph feature extraction unit, configured to extract the first image feature of the response graph;

[0207] A first image feature extraction unit, configured to extract any video frame of the video to be detected and acquire the second image feature of the video frame;

[0208] A first fused image feature extraction unit, configured to fuse the first image feature and the second image feature to obtain a fused image feature;

[0209] A first processing unit, configured to process the fused image feature by using the first live model to obtain a first live detection result of the object to be detected.

[0210] Figure 6 is a schematic structural diagram of a living body detection device according to an embodiment of the present application. As Figure 6 shown, the living body detection device includes a video acquisition module 61, a response map generation module 62, and a detection result determination module 63, where:

[0211] The video acquisition module 61 is configured to acquire a video to be detected of an object to be detected, and the video to be detected is: a video of the object to be detected collected during irradiating the object to be detected according to a first illumination sequence;

[0212] The response map generation module 62 is configured to generate a response map of the object to be detected according to the first illumination sequence and a second illumination sequence reflected by each entity position point of the object to be detected represented by the video to be detected, and the response intensity of each pixel point in the response map represents: the similarity between the second illumination sequence reflected by the entity position point corresponding to the pixel point and the first illumination sequence;

[0213] The detection result determination module 63 is configured to perform living body detection on the object to be detected based on the response map by using a living body detection model to obtain a living body detection result of the object to be detected.

[0214] Optionally, the second illumination sequence reflected by each entity position point of the object to be detected represented by the video to be detected is obtained according to the following steps:

[0215] Extract a plurality of video frames of the video to be detected;

[0216] Align the pixel points describing the same entity position point of the object to be detected in the plurality of video frames;

[0217] For each entity position point of the object to be detected, obtain the second illumination sequence reflected by the entity position point according to the illumination reflected by the pixel points describing the entity position point in each of the plurality of video frames.

[0218] Optionally, aligning the pixel points describing the same entity position point of the object to be detected in the plurality of video frames includes:

[0219] For each of the plurality of video frames, perform face key point detection on the video frame to obtain the face key points included in the video frame, and perform face key point level alignment processing on the plurality of video frames according to the face key points included in each of the video frames, and perform pixel point level alignment processing on the plurality of video frames after performing face key point level alignment processing;

[0220] Or,

[0221] Perform pixel-level alignment processing on the multiple video frames.

[0222] Optionally, perform the pixel-level alignment processing through the following process:

[0223] Calculate the dense optical flow data between each of the multiple video frames participating in the pixel-level alignment processing except the reference video frame and the reference video frame, where the reference video frame is any one of the multiple video frames participating in the pixel-level alignment processing;

[0224] Align each of the other video frames with the reference video frame at the pixel level according to the dense optical flow data between each of the other video frames and the reference video frame.

[0225] Optionally, both the first illumination sequence and the second illumination sequence include color light sequences of multiple color channels; the response map generation module 62 includes:

[0226] A second separation unit for separating the first illumination sequence to obtain the first color light sequence corresponding to each color channel, and separating the second illumination sequence to obtain the second color light sequence corresponding to each color channel;

[0227] A second response sub-map generation unit for generating a response sub-map of the object to be detected in each color channel according to the similarity between the second color light sequence of each color channel reflected by each entity position point of the object to be detected and the first color light sequence corresponding to the color channel;

[0228] A second fusion unit for fusing the response sub-maps of the object to be detected in each color channel to obtain the response map of the object to be detected.

[0229] Optionally, the second fusion unit includes:

[0230] A second mean value acquisition sub-unit for obtaining the mean response intensity of the face region of the object to be detected in each color channel according to the response sub-maps of the object to be detected in each color channel;

[0231] A second regularization sub-unit for regularizing the response sub-map of the object to be detected in each color channel according to the mean response intensity of the face region of the object to be detected in each color channel to obtain the regularized response sub-map of the object to be detected in each color channel;

[0232] A second fusion subunit, configured to fuse the regularized response subgraphs of the object to be detected in each color channel to obtain a response map of the object to be detected.

[0233] Optionally, before processing the response map using the live detection model, the apparatus further includes:

[0234] A second attribute value acquisition module, configured to acquire an attribute value of the response map of the object to be detected, where the attribute value includes at least one of the following: average response intensity, quality value;

[0235] A second determination module, configured to determine a live detection result of the object to be detected according to a magnitude relationship between the attribute value of the response map of the object to be detected and a corresponding attribute threshold;

[0236] A second determination module, configured to determine that the object to be detected is not a live body when the attribute value of the response map of the object to be detected is less than the corresponding attribute threshold;

[0237] A second step execution module, configured to execute a step of performing live detection on the object to be detected using the live detection model based on the response map when the attribute value of the response map of the object to be detected is not less than the corresponding attribute threshold.

[0238] Optionally, the live detection model is a model that has learned a first image feature of the response map of a live face and a second image feature of the live face; the detection result determination module 63 includes:

[0239] A second response map feature extraction unit, configured to extract a first image feature of the response map;

[0240] A second image feature extraction unit, configured to extract any video frame of the video to be detected and obtain a second image feature of the video frame;

[0241] A second fused image feature extraction unit, configured to fuse the first image feature and the second image feature to obtain a fused image feature;

[0242] A second processing unit, configured to process the fused image feature using the first live model to obtain a first live detection result of the object to be detected.

[0243] It should be noted that the apparatus embodiments are similar to the method embodiments, so the description is relatively simple. For related parts, please refer to the method embodiments.

[0244] An embodiment of the present application further provides an electronic device. Refer to Figure 7 , Figure 7 is a schematic diagram of the electronic device proposed in the embodiment of the present application. AsFigure 7 As shown in the figure, the electronic device 100 includes: a memory 110 and a processor 120. The memory 110 and the processor 120 are communicatively connected via a bus. A computer program is stored in the memory 110, and this computer program can run on the processor 120, thereby implementing the steps in the living body detection method disclosed in the embodiments of the present application.

[0245] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, it implements the living body detection method as disclosed in the embodiments of the present application.

[0246] The embodiments of the present application also provide a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, it implements the living body detection method as disclosed in the embodiments of the present application.

[0247] The embodiments of the present application also provide a computer program, which can implement the living body detection method disclosed in the embodiments of the present application when executed.

[0248] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.

[0249] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0250] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices, electronic devices, and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0251] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the process Figure 1 one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.

[0252] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0253] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.

[0254] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal device including the said element.

[0255] The above has introduced in detail a living body detection method, an electronic device, a storage medium and a program product provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for detecting a living body, characterized in that, Including: Obtaining a video to be detected of an object to be detected, where the video to be detected is a video of the object to be detected collected during irradiating the object to be detected according to a first light sequence; Generating a response map of the object to be detected based on the first light sequence and a second light sequence reflected by each entity position point of the object to be detected represented by the video to be detected, where the response intensity of each pixel point in the response map represents the similarity between the second light sequence reflected by the entity position point corresponding to the pixel point and the first light sequence; Performing live detection on the object to be detected based on the response map using a first live detection model to obtain a first live detection result of the object to be detected; Inputting the video to be detected into a second live detection model to obtain a second live detection result of the object to be detected; Determining a final live detection result of the object to be detected according to the first live detection result and the second live detection result.

2. The method according to claim 1, wherein The second light sequence reflected by each entity position point of the object to be detected represented by the video to be detected is obtained according to the following steps: Extracting a plurality of video frames of the video to be detected; Aligning pixel points describing the same entity position point of the object to be detected in the plurality of video frames; For each entity position point of the object to be detected, obtaining the second light sequence reflected by the entity position point according to the light reflected by the pixel points describing the entity position point in each of the plurality of video frames.

3. The method according to claim 2, wherein Aligning pixel points describing the same entity position point of the object to be detected in the plurality of video frames includes: For each of the plurality of video frames, performing face key point detection on the video frame to obtain face key points included in the video frame, and performing face key point-level alignment processing on the plurality of video frames according to the face key points included in each of the video frames, and then performing pixel point-level alignment processing on the plurality of video frames after the face key point-level alignment processing; Or, Performing pixel point-level alignment processing on the plurality of video frames.

4. The method according to claim 3, wherein The pixel point-level alignment processing is performed through the following process: Calculating dense optical flow data between each of the plurality of video frames other than a reference video frame participating in the pixel point-level alignment processing and the reference video frame respectively, where the reference video frame is any one of the plurality of video frames participating in the pixel point-level alignment processing; Performing pixel point-level alignment processing on each of the other video frames and the reference video frame according to the dense optical flow data between each of the other video frames and the reference video frame.

5. The method according to any one of claims 1-4, characterized in that, Both the first light sequence and the second light sequence include color light sequences of multiple color channels; Generating a response map of the object to be detected based on the first light sequence and a second light sequence reflected by each entity position point of the object to be detected represented by the video to be detected includes: Separate the first illumination sequence to obtain the first color light sequence corresponding to each color channel, and separate the second illumination sequence to obtain the second color light sequence corresponding to each color channel; For each color channel, generate a response sub - map of the object to be detected in this color channel according to the similarity between the second color light sequence of this color channel reflected by each entity position point of the object to be detected and the first color light sequence corresponding to this color channel; Fuse the response sub - maps of the object to be detected in each color channel to obtain the response map of the object to be detected.

6. The method according to claim 5, wherein Fusing the response sub - maps of the object to be detected in each color channel to obtain the response map of the object to be detected includes: According to the response sub - maps of the object to be detected in each color channel, obtain the average response intensity of the face region of the object to be detected in each color channel; According to the average response intensity of the face region of the object to be detected in each color channel, perform regularization processing on the response sub - map of the object to be detected in this color channel to obtain the regularized response sub - map of the object to be detected in each color channel; Fuse the regularized response sub - maps of the object to be detected in each color channel to obtain the response map of the object to be detected.

7. A method for detecting a living body, characterized in that, Includes: Obtain the video to be detected of the object to be detected, where the video to be detected is the video of the object to be detected collected during irradiating the object to be detected according to the first illumination sequence; According to the first illumination sequence and the second illumination sequence reflected by each entity position point of the object to be detected represented by the video to be detected, generate the response map of the object to be detected, and the response intensity of each pixel point in the response map represents the similarity between the second illumination sequence reflected by the entity position point corresponding to the pixel point and the first illumination sequence; Based on the response map, use a live detection model to perform live detection on the object to be detected to obtain the live detection result of the object to be detected.

8. The method according to claim 7, wherein The second illumination sequence reflected by each entity position point of the object to be detected represented by the video to be detected is obtained according to the following steps: Extract multiple video frames of the video to be detected; Align the pixel points describing the same entity position point of the object to be detected in the multiple video frames; For each entity position point of the object to be detected, according to the illumination reflected by the pixel points describing this entity position point in each of the multiple video frames, obtain the second illumination sequence reflected by this entity position point.

9. The method according to claim 8, wherein Aligning the pixel points describing the same entity position point of the object to be detected in the multiple video frames includes: For each of the multiple video frames, perform face key - point detection on the video frame to obtain the face key - points included in the video frame. According to the face key - points included in each video frame, perform face - key - point - level alignment processing on the multiple video frames, and then perform pixel - point - level alignment processing on the multiple video frames after face - key - point - level alignment processing; Or, Perform pixel-level alignment processing on the multiple video frames.

10. The method according to claim 9, wherein The pixel-level alignment processing is performed through the following process: Calculate the dense optical flow data between each of the multiple video frames participating in the pixel-level alignment processing, except the reference video frame, and the reference video frame, where the reference video frame is any one of the multiple video frames participating in the pixel-level alignment processing; According to the dense optical flow data between each of the other video frames and the reference video frame, perform pixel-level alignment processing on each of the other video frames and the reference video frame.

11. The method according to claim 7, wherein Both the first illumination sequence and the second illumination sequence include color light sequences of multiple color channels; Generate a response map of the object to be detected according to the first illumination sequence and the second illumination sequence reflected by each entity position point of the object to be detected represented by the video to be detected, including: Separate the first illumination sequence to obtain the first color light sequence corresponding to each color channel, and separate the second illumination sequence to obtain the second color light sequence corresponding to each color channel; For each color channel, generate a response sub-map of the object to be detected in this color channel according to the similarity between the second color light sequence of this color channel reflected by each entity position point of the object to be detected and the first color light sequence corresponding to this color channel; Fuse the response sub-maps of the object to be detected in each color channel to obtain the response map of the object to be detected.

12. The method according to claim 11, wherein Fuse the response sub-maps of the object to be detected in each color channel to obtain the response map of the object to be detected, including: According to the response sub-maps of the object to be detected in each color channel, obtain the average response intensity of the face region of the object to be detected in each color channel; According to the average response intensity of the face region of the object to be detected in each color channel, perform regularization processing on the response sub-map of the object to be detected in this color channel to obtain the regularized response sub-map of the object to be detected in each color channel; Fuse the regularized response sub-maps of the object to be detected in each color channel to obtain the response map of the object to be detected.

13. The method according to any one of claims 7-12, characterized in that, Before processing the response map using the live detection model, it further includes: Obtain the attribute value of the response map of the object to be detected, where the attribute value includes at least one of the following: average response intensity, quality value; Determine the live detection result of the object to be detected according to the magnitude relationship between the attribute value of the response map of the object to be detected and the corresponding attribute threshold; In the case where the attribute value of the response map of the object to be detected is less than the corresponding attribute threshold, determine that the object to be detected is not a live body; In the case where the attribute value of the response map of the object to be detected is not less than the corresponding attribute threshold, perform the step of performing live detection on the object to be detected using the live detection model based on the response map.

14. The method according to any one of claims 7-12, characterized in that, The live detection model is a model that has learned the first image feature of the live face's response map and the second image feature of the live face; Based on the response map, the live detection model is used to perform live detection on the object to be detected, and the live detection result of the object to be detected is obtained, including: Extracting the first image feature of the response map; Extracting any video frame of the video to be detected and obtaining the second image feature of the video frame; Fusing the first image feature and the second image feature to obtain a fused image feature; Using the live model to process the fused image feature to obtain the live detection result of the object to be detected.

15. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the live detection method according to any one of claims 1 to 6; or, the processor executes the computer program to implement the live detection method according to any one of claims 7 to 14.

16. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the live detection method according to any one of claims 1 to 6 is implemented; or, when the computer program / instructions are executed by the processor, the live detection method according to any one of claims 7 to 14 is implemented.

17. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the live detection method according to any one of claims 1 to 6 is implemented; or, when the computer program / instructions are executed by the processor, the live detection method according to any one of claims 7 to 14 is implemented.

Citation Information

Patent Citations

  • Living body detection method and device and storage medium

    CN107992794A

  • Coarse-to-fine image dense matching method, system and device and storage medium

    CN114429555A