Living body detection method and device, storage medium and electronic equipment

By acquiring video frames under different acquisition conditions and using a detection model to identify differences, the problem of traditional liveness detection being unable to resist attacks from AI-generated images has been solved, achieving a lightweight liveness detection effect.

CN121366450APending Publication Date: 2026-01-20ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511518799.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Traditional liveness detection methods are vulnerable to attacks that utilize high-definition images and videos generated by artificial intelligence.

Method used

Liveness detection is performed by acquiring video frames under different acquisition conditions and using a pre-trained detection model to identify differences between video frames, while liveness verification is performed using a lightweight model.

Benefits of technology

It effectively defends against attacks that render high-definition images, reduces the number of video frames required for detection, and is suitable for user terminals with limited computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121366450A_ABST
    Figure CN121366450A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a living body detection method, and the method comprises the steps: extracting target video frames collected under different collection conditions from a video collected by a user, inputting each target video frame into a detection model, and carrying out the living body detection of the user through the detection model according to the difference between the target video frames. According to the method, living body detection is carried out on the user through the difference between the target video frames collected under different collection conditions, presentation attacks carried out through a high-definition image can be effectively defended, living body detection can be carried out only by inputting a small number of target video frames, the whole video does not need to be input, and the user experience is improved. Therefore, living body detection can be realized only by adopting a lightweight detection model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of computer technology, and particularly relates to a live body detection method and device, a storage medium and an electronic device. BACKGROUND

[0002] Live body detection is a technology against presentation attacks, which is used to determine whether the object performing biometric recognition (such as face, fingerprint, iris) is a real, living person, or a fake tool for misusing identity, such as a photo, a video, a mask, a wax figure, etc. Live body detection has been widely used in various identity recognition processes based on biometric recognition, such as commonly used face recognition.

[0003] However, with the development of artificial intelligence (AI) and image processing technology, the quality of various high-definition images and images and videos generated through AI models is also getting higher and higher. Since such high-definition images and images and videos generated based on artificial intelligence generated content (AIGC) technology are already so lifelike that they can deceive people, the traditional live body detection method has been difficult to resist such presentation attacks based on AIGC technology.

[0004] Therefore, how to achieve live body detection that can still resist presentation attacks at the present time is an urgent problem. SUMMARY

[0005] Embodiments of the present specification provide a live body detection method, device, storage medium and electronic device to partially solve the problems existing in the prior art.

[0006] Embodiments of the present specification adopt the following technical solutions: The present specification provides a live body detection method, which comprises: acquiring a video collected in advance for a user; the video comprises video frames collected under at least two different collection conditions respectively; for each collection condition, extracting video frames collected under the collection condition from the video as target video frames; inputting each target video frame extracted into a pre-trained detection model; performing live body detection on the user according to differences between the target video frames through the detection model.

[0007] The present specification provides a live body detection device, which comprises: An acquisition module is configured to acquire a video collected in advance for a user, wherein the video comprises video frames collected under at least two different collection conditions respectively; An extraction module is configured to extract, for each collection condition, video frames collected under the collection condition from the video as target video frames; An input module is configured to input each target video frame extracted to a detection model trained in advance; A detection module is configured to perform, by the detection model, a live body detection on the user according to differences between the target video frames.

[0008] The specification provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the live body detection method.

[0009] The specification provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the live body detection method when executing the program.

[0010] The specification provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the live body detection method.

[0011] The above at least one technical solution adopted by the embodiments of the specification can achieve the following beneficial effects: The embodiments of the specification disclose a live body detection method, which extracts target video frames collected under different collection conditions from a video collected for a user, inputs each target video frame to a detection model, and performs, by the detection model, a live body detection on the user according to differences between the target video frames. The above method performs a live body detection on the user through differences between target video frames collected under different collection conditions, can effectively prevent a presentation attack by a high-definition image, and only needs to input a few target video frames to perform a live body detection, without inputting an entire video, so that a lightweight detection model can be used to implement a live body detection. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings explained herein are used to provide a further understanding of the specification, constitute a part of the specification, and the illustrative embodiments of the specification and the explanation thereof are used to explain the specification, and do not constitute an improper limitation on the specification. In the drawings: Figure 1 A live body detection method flowchart provided by the embodiments of the specification; Figure 2 A live body detection device schematic diagram provided by the embodiments of the specification; Figure 3 A structural schematic diagram of an electronic device provided by an embodiment of the present specification. DETAILED DESCRIPTION

[0013] For the purpose, technical solutions and advantages of the present specification to be clearer, the technical solutions of the present specification will be described clearly and completely in the following with reference to the embodiments of the present specification and corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by a person of ordinary skill in the art without any creative work, fall within the protection scope of the present specification.

[0014] The technical solutions provided by the embodiments of the present specification will be described in detail below with reference to the drawings.

[0015] Figure 1 A flowchart of a living body detection method provided by an embodiment of the present specification, specifically comprising the following steps: S100: acquiring a video collected in advance for a user; the video comprising video frames collected respectively under at least two different collection conditions.

[0016] In the embodiments of the present specification, the method shown in Figure 1 The device for detecting the living body of the user by the method can be a user terminal, such as a personal computer, a mobile phone, a tablet computer, etc. The present specification does not limit this.

[0017] The user terminal can perform the living body detection process on the user as shown in Figure 1 during the identity recognition process based on the biometric recognition of the user, such as performing the living body detection on the user during the process of recognizing the face of the user.

[0018] Among them, the above-mentioned video can be collected by the camera built-in or externally connected to the user terminal, or collected by other devices independent of the user terminal and then sent to the user terminal for subsequent living body detection. In short, as long as the video is collected for the user. Only the camera built-in in the user terminal is taken as an example to be described below.

[0019] Before the user terminal collects the video of the user, the user terminal can also query whether the user has authorized the behavior of the user terminal collecting the video, if the user has authorized the behavior, the user terminal collects the video, if the user has not authorized the behavior, the user terminal can display an authorization interface, the authorization interface contains inquiry information for inquiring whether the user authorizes the behavior of the user terminal collecting the video, if the user clicks a button for confirming authorization in the authorization interface, it is determined that the user has authorized the behavior of the user terminal collecting the video, and the video can be collected, if the user clicks a button for refusing authorization in the authorization interface, it is determined that the user refuses to authorize the behavior of the user terminal collecting the video, at this time, the user terminal refuses to collect the video, and other methods can be used to detect the living body of the user, or it is directly determined that the living body detection of the user is not passed.

[0020] In the embodiments of the present specification, when the user terminal collects the video of the user, the video can be collected under at least two different collection conditions. Specifically, the user terminal can first collect the video of the user under a first collection condition, change the first collection condition to a second collection condition in the process of collecting the video, and continue to collect the video of the user under the second collection condition.

[0021] In the present specification, the first collection condition refers to the collection condition before the change of the collection condition, and the second collection condition refers to the collection condition after the change of the collection condition, and the first collection condition and the second collection condition are not a specific collection condition.

[0022] The collection condition in the present specification includes both the collection condition irrelevant to the camera for collecting the video and the collection condition relevant to the camera for collecting the video. The collection condition irrelevant to the camera can be the collection environment in which the user is located when the video is collected. The collection condition relevant to the camera can be the parameter of the camera for collecting the video.

[0023] When the changed second collection condition is the collection environment in which the user is located, the first collection condition can be changed to the second collection condition by using a condition adjustment method that the user can perceive. Specifically, the user terminal can display preset prompt information in the process of collecting the video of the user under the first collection condition, so as to prompt the user to change the collection environment in which the user is located. For example, the user is prompted to adjust the distance between the user and the collection device (i.e. the camera), or the user is prompted to change the illumination of the collection environment in which the user is located.

[0024] When the changed second collection condition is the parameter of the camera for collecting the video, the user terminal can directly adjust the first collection condition to the second collection condition by using a preset condition adjustment method that the user cannot perceive in the process of collecting the video of the user under the first collection condition.

[0025] It should be noted that, no matter how the acquisition condition is changed by using any of the above condition adjustment manners, in the process of acquiring the video of the user, the interface displayed on the display screen of the user terminal (including the color, brightness, and content of the interface) remains unchanged except for displaying the above prompt information. That is, the manner of changing the acquisition condition by changing the color, brightness, and the like of the interface displayed on the display screen of the user terminal is not within the condition adjustment manners described in the specification.

[0026] S101: For each acquisition condition, extract a video frame acquired under the acquisition condition from the video as a target video frame.

[0027] In the embodiments of the present specification, after the user terminal acquires the above video, the user terminal can extract at least one video frame acquired under each acquisition condition from the video as a target video frame.

[0028] Specifically, the user terminal can determine a time period in which the above video is acquired under each acquisition condition, determine each video frame acquired in the time period in the video, determine the image quality of each video frame, and finally extract a target video from each video frame according to the image quality of each video frame. Among them, the video frame with the highest image quality can be taken as the target video frame acquired under the acquisition condition.

[0029] S102: Input each target video frame extracted into a pre-trained detection model.

[0030] In the embodiments of the present specification, the pre-trained detection model can be deployed on the user terminal. Then, after the target video frame acquired under each acquisition condition is obtained through step S101, the target video frame can be input into the detection model to perform live detection on the user through the detection model.

[0031] S103: Perform live detection on the user according to the difference between each target video frame through the detection model.

[0032] Specifically, in step S102, each target video frame can be input into the detection model respectively when the target video frames are input into the detection model. Then, in step S103, the user terminal can determine the image features of each target video frame through the detection model, fuse the image features of each target video frame to obtain a fused feature, recognize the difference between each target video frame according to the fused feature, and perform live detection on the user according to the difference.

[0033] In order to improve the live detection efficiency, the user terminal can also splice each target video frame into a to-be-detected image in step S102, and input the to-be-detected image into the detection model. Then in step S103, the user terminal can directly determine the image features of the to-be-detected image through the detection model, recognize the differences between the target video frames according to the image features of the to-be-detected image, and finally perform live detection on the user according to the differences.

[0034] In actual application scenarios, the attacker often only has a fake high-definition image in hand, so the attacker often only uses this fake high-definition image to perform a presentation attack. Specifically, the attacker generally uses the high-definition image to fake a video and performs live detection, for example, directly replaces an actual collected video with a fake video composed of video frames obtained by adjusting the high-definition image, or inserts the high-definition image or an image obtained by adjusting the high-definition image into the actual collected video as a fake video frame.

[0035] At this time, if the adjustment of the collection condition in the process of collecting the video is a user-perceptible condition adjustment manner, the user will adjust the only high-definition image in his hand according to the prompt information to obtain an image that meets the changed second collection condition. For example, the prompt information displayed by the user terminal is “please move away a little”, which means that the user needs to increase the distance between himself and the camera, so the attacker will process the only high-definition image in his hand to be smaller according to the prompt information to obtain an image that meets the changed second collection condition. For another example, the prompt information displayed by the user terminal is “please turn on the ambient light”, which means that the user needs to increase the light condition of the environment in which he is located, so the attacker will process the brightness of the only high-definition image in his hand to be larger according to the prompt information to obtain an image that meets the changed second collection condition.

[0036] For this case, although the video faked by the attacker contains the faked first target video frame collected under the first collection condition and the faked second target video frame collected under the second collection condition, and the difference between the two faked target video frames also matches the change from the first collection condition to the second collection condition, the two faked target video frames are derived from the same image, that is, the second target video frame is obtained by adjusting the first target video frame, or both the first target video frame and the second target video frame are obtained by adjusting another image.

[0037] If the adjustment of the collection condition during the collection of the video is a condition adjustment manner that is imperceptible to the user, the attacker is completely unaware of when the first collection condition has been changed to the second collection condition during the collection of the video, and thus will not make any adjustment to the high-definition image in hand. At this time, not only is the difference between the two forged target video frames completely unmatched with the change from the first collection condition to the second collection condition, but the two forged target video frames are still derived from the same image.

[0038] Therefore, in step S103, after determining the difference between the first target video frame collected under the first collection condition and the second target video frame collected under the second collection condition, the detection model can identify whether the difference between the first target video frame and the second target video frame matches the change from the first collection condition to the second collection condition, and also identify whether the first target video frame and the second target video frame are derived from the same image according to the difference between the first target video frame and the second target video frame. If the difference between the first target video frame and the second target video frame matches the change from the first collection condition to the second collection condition, and the first target video frame and the second target video frame are not derived from the same image, it is determined that the live body detection passes, otherwise it is determined that the live body detection fails.

[0039] Of course, in order to enable the detection model to accurately identify whether the difference between the first target video frame and the second target video frame matches the change from the first collection condition to the second collection condition, in addition to inputting the first target video frame and the second target video frame into the detection model, the first collection condition and the second collection condition can also be inputted into the detection model accordingly.

[0040] Through the above method, when performing live body detection on the user, not only is the live body detection performed according to only one video frame, but also the entire video does not need to be inputted into the detection model. Only a few target video frames, or even only two target video frames, are needed to perform live body detection on the user, so as to not only effectively resist the current mainstream presentation attack, but also not need a large model parameter size of the detection model. A lightweight detection model can be used to implement the above live body detection method, so as to directly deploy the detection model on a user terminal with limited computing power.

[0041] Further, since the above detection model needs to identify whether the difference between the first target video frame and the second target video frame matches the change from the first collection condition to the second collection condition, and also needs to identify whether the first target video frame and the second target video frame are derived from the same image, when training the detection model, the detection model to be trained and training samples can be obtained.

[0042] The training samples include positive samples and negative samples.

[0043] The positive sample is a first sample video frame pair. The first sample video frame pair includes a first sample video frame collected under a first collection condition and a second sample video frame collected under a second collection condition for a sample user (the sample user is a normal living body user). That is, the difference between the first sample video frame and the second sample video frame matches the change from the first collection condition to the second collection condition, and the first sample video frame and the second sample video frame are not derived from the same image. Thus, the positive sample corresponds to a detection result of living body detection passing.

[0044] The negative sample includes a second sample video frame pair and a third sample video frame pair.

[0045] The second sample video frame pair includes a third sample video frame and a fourth sample video frame. The difference between the third sample video frame and the fourth sample video frame does not match the change from the first collection condition to the second collection condition. The third sample video frame and the fourth sample video frame can be video frames in a video collected for a normal living body user as a sample user (the collection condition when the video is collected does not change from the first collection condition to the second collection condition), or can be video frames in a video generated by a fake high-definition image and not changed from the first collection condition to the second collection condition.

[0046] The third sample video frame pair includes a fifth sample video frame and a sixth sample video frame. The fifth sample video frame and the sixth sample video frame are derived from the same image, that is, the fifth sample video frame is obtained by adjusting the sixth sample video frame, or the sixth sample video frame is obtained by adjusting the fifth sample video frame, or the fifth sample video frame and the sixth sample video frame are both obtained by adjusting another image. The difference between the fifth sample video frame and the sixth sample video frame can match or not match the change from the first collection condition to the second collection condition.

[0047] The negative sample is labeled as a detection result of living body detection not passing.

[0048] After the training sample including the plurality of positive samples and the plurality of negative samples is obtained, the training sample can be input into the detection model to be trained, and the living body detection is performed according to the difference between the sample video frames included in the training sample by the detection model to be trained. The loss value is determined according to the difference between the detection result output by the detection model to be trained and the label of the training sample, wherein the loss value is positively correlated with the difference between the detection result output by the detection model to be trained and the label of the training sample. Finally, the model parameters of the detection model to be trained are adjusted to reduce the loss value. That is, the detection model to be trained is trained in a supervised training manner.

[0049] It should be noted that the device for training the detection model and the user terminal for performing the living body detection shown in the above embodiment can be the same device or different devices, which is not limited in the present specification. Figure 1 The user terminal for performing the living body detection shown in the above embodiment can be the same device or different devices, which is not limited in the present specification.

[0050] The above is a living body detection method provided by the embodiment of the present specification, based on the same idea, the present specification also provides a corresponding device, a storage medium and an electronic device.

[0051] Figure 2 The living body detection device provided by the embodiment of the present specification is shown in the schematic diagram, and the device includes: The acquisition module 201 is configured to acquire a video collected in advance for a user, wherein the video includes video frames collected under at least two different collection conditions. The extraction module 202 is configured to extract, for each collection condition, the video frames collected under the collection condition from the video as target video frames. The input module 203 is configured to input each target video frame extracted into a pre-trained detection model. The detection module 204 is configured to perform living body detection on the user according to the difference between the target video frames by the detection model.

[0052] Optionally, the device further includes: The collection module 205 is specifically configured to collect a video for a user under a first collection condition, change the first collection condition to a second collection condition, and continue to collect a video for the user under the second collection condition.

[0053] Optionally, the collection module 205 is specifically configured to display preset prompt information to prompt the user to change the collection environment in which the user is located by the prompt information, wherein the collection environment includes the distance between the user and the collection device, or a preset user-unaware adjustment method is used to adjust the first collection condition to the second collection condition.

[0054] Optionally, the extraction module 202 is specifically configured to determine, in the video, each video frame collected in a time period in which the video is collected under the acquisition condition; determine image quality of the each video frame; and extract a target video frame from the each video frame according to the image quality of the each video frame.

[0055] Optionally, the input module 203 is specifically configured to splice the extracted each target video frame into a to-be-detected image, and input the to-be-detected image into a pre-trained detection model.

[0056] Optionally, the detection module 204 is specifically configured to, for a first target video frame collected under a first acquisition condition and a second target video frame collected under a second acquisition condition, identify, by the detection model, whether a difference between the first target video frame and the second target video frame matches a change from the first acquisition condition to the second acquisition condition, and identify, according to the difference between the first target video frame and the second target video frame, whether the first target video frame and the second target video frame are derived from a same image; and perform, according to an identification result, a live body detection on the user.

[0057] Optionally, the apparatus further includes: a training module 206 configured to obtain a to-be-trained detection model and training samples, the training samples including positive samples and negative samples; the positive samples being a first sample video frame pair, the first sample video frame pair including a first sample video frame collected under a first acquisition condition and a second sample video frame collected under a second acquisition condition for a sample user; a label of the positive samples being a detection result of passing a live body detection; the negative samples including a second sample video frame pair and a third sample video frame pair, a difference between a third sample video frame and a fourth sample video frame in the second sample video frame pair not matching a change from the first acquisition condition to the second acquisition condition, a fifth sample video frame and a sixth sample video frame in the third sample video frame pair being derived from a same image; a label of the negative samples being a detection result of failing a live body detection; inputting the training samples into the to-be-trained detection model; performing, by the detection model, a live body detection according to a difference between each sample video frame contained in the training samples; and adjusting model parameters of the to-be-trained detection model according to a detection result output by the detection model and the label of the training samples.

[0058] The specification also provides a computer-readable storage medium, the storage medium storing a computer program, the computer program being executable by a processor to perform the live body detection method provided above.

[0059] The specification also provides a computer program product comprising a computer program which, when executed by a processor, implements the living body detection method described above.

[0060] Based on Figure 1 The living body detection method shown in the specification, the embodiments of the specification also provide Figure 3 The structural diagram of the electronic device shown in the specification. As Figure 3 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and of course can also include other hardware required by the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs to implement the living body detection method described above.

[0061] The above is only an embodiment of the specification and is not intended to limit the specification. For those skilled in the art, the specification can have various changes and variations. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the specification shall be included in the scope of the claims of the specification.

Claims

1. A method for living body detection, the method comprising: acquiring a video collected in advance for a user; the video comprising video frames collected respectively under at least two different collection conditions; for each collection condition, extracting, from the video, video frames collected under the collection condition as target video frames; inputting the extracted target video frames into a detection model trained in advance; performing living body detection on the user according to differences between the target video frames by the detection model.

2. The method of claim 1, wherein the video is collected in advance for the user, and the collection specifically comprises: collecting the video for the user under a first collection condition; changing the first collection condition to a second collection condition; continuing to collect the video for the user under the second collection condition.

3. The method of claim 2, wherein the first collection condition is changed to the second collection condition by: displaying preset prompt information to prompt the user to change a collection environment in which the user is located, the collection environment comprising a distance between the user and a collection device; or using a preset user-unaware adjustment method to adjust the first collection condition to the second collection condition.

4. The method of claim 1, wherein the video frames collected under the collection condition are extracted from the video as target video frames, and the extraction specifically comprises: determining, from the video, video frames collected in a time period in which the video is collected under the collection condition; determining image qualities of the video frames; extracting target video frames from the video frames according to the image qualities of the video frames.

5. The method of claim 1, wherein the extracted target video frames are inputted into the detection model trained in advance by: splicing the extracted target video frames into a to-be-detected image, and inputting the to-be-detected image into the detection model trained in advance.

6. The method of claim 3, wherein the living body detection on the user is performed according to differences between the target video frames by the detection model, and the detection specifically comprises: for a first target video frame collected under the first collection condition and a second target video frame collected under the second collection condition, identifying, by the detection model, whether differences between the first target video frame and the second target video frame match changes from the first collection condition to the second collection condition, and identifying whether the first target video frame and the second target video frame are derived from the same image according to the differences between the first target video frame and the second target video frame; performing the living body detection on the user according to the identification result.

7. The method of claim 6, wherein the detection model is trained in advance by: ​ acquire a detection model to be trained and training samples, the training samples including positive samples and negative samples; the positive samples are first sample video frame pairs, the first sample video frame pairs including first sample video frames collected under a first collection condition and second sample video frames collected under a second collection condition for a sample user; the positive samples are labeled as a detection result of a live body detection passing; the negative samples include second sample video frame pairs and third sample video frame pairs, a difference between a third sample video frame and a fourth sample video frame in the second sample video frame pairs does not match a change from the first collection condition to the second collection condition, a fifth sample video frame and a sixth sample video frame in the third sample video frame pairs are from the same image; the negative samples are labeled as a detection result of a live body detection failing; input the training samples into the detection model to be trained; perform a live body detection according to differences between sample video frames included in the training samples by the detection model, and adjust model parameters of the detection model to be trained according to a detection result output by the detection model and labels of the training samples. 8.A live body detection apparatus, the apparatus comprising: an acquisition module configured to acquire a video collected in advance for a user; the video including video frames collected under at least two different collection conditions respectively; an extraction module configured to extract, for each collection condition, video frames collected under the collection condition from the video as target video frames; an input module configured to input each target video frame extracted to a detection model trained in advance; a detection module configured to perform a live body detection on the user according to differences between the target video frames by the detection model. 9.A computer readable storage medium, the storage medium storing a computer program, the computer program being executed by a processor to implement the method of any one of claims 1-7. 10.An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor implementing the method of any one of claims 1-7 when executing the program. 11.A computer program product, the computer program product containing a computer program, the computer program being executed by a processor to implement the method of any one of claims 1-7.