Living body detection method and device, electronic equipment, storage medium and program product

By combining the feature extraction of video data and acoustic signals, comprehensive liveness detection is performed, which solves the problem of disguised identity features in face authentication and improves the accuracy and security of detection.

CN120708295APending Publication Date: 2025-09-26MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510065276.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In existing technologies, facial identity verification can be easily disguised by means such as photos, resulting in reduced security and accuracy, and existing liveness detection methods have low accuracy.

Method used

By acquiring the video data and acoustic wave signals of the target object, feature extraction is performed, and visual and acoustic wave detection are combined to conduct liveness detection in multiple dimensions.

Benefits of technology

The accuracy and robustness of liveness detection are improved, and the risk of identity authentication attacks is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708295A_ABST
    Figure CN120708295A_ABST
Patent Text Reader

Abstract

The invention provides a living body detection method and device, electronic equipment, a storage medium and a program product, and the method comprises the steps: obtaining video data collected for a target object, and obtaining a sound wave signal reflected by the target object; performing feature extraction on video frames in the video data to obtain a first feature of the video data, and determining a first living body detection result according to the first feature; performing feature extraction on the sound wave signal to obtain a second feature of the sound wave signal, and determining a second living body detection result according to the second feature; and determining a target living body detection result of the target object according to the first living body detection result and the second living body detection result, thereby improving the living body detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a liveness detection method, device, electronic device, storage medium, and program product. Background Art

[0002] With increasing emphasis on security, identity authentication is required in various applications. However, during face authentication, identity features may be disguised through means such as photos, which reduces security and the accuracy of authentication. Therefore, liveness detection is very important in face authentication. Summary of the Invention

[0003] The present disclosure provides a living body detection method, device, electronic device, storage medium and program product.

[0004] In a first aspect, the present disclosure provides a liveness detection method, the liveness detection method comprising:

[0005] Acquiring video data collected from a target object and acquiring an acoustic wave signal reflected by the target object;

[0006] Performing feature extraction on video frames in the video data to obtain a first feature of the video data, and determining a first liveness detection result based on the first feature;

[0007] performing feature extraction on the acoustic wave signal to obtain a second feature of the acoustic wave signal, and determining a second liveness detection result based on the second feature;

[0008] A target liveness detection result of the target object is determined according to the first liveness detection result and the second liveness detection result.

[0009] In a second aspect, the present disclosure provides a liveness detection device, the liveness detection device comprising:

[0010] An acquisition module, configured to acquire video data collected from a target object and to acquire an acoustic wave signal reflected from the target object;

[0011] a first processing module, configured to extract features from video frames in the video data to obtain a first feature of the video data, and determine a first liveness detection result based on the first feature;

[0012] a second processing module, configured to perform feature extraction on the acoustic wave signal to obtain a second feature of the acoustic wave signal, and determine a second liveness detection result based on the second feature;

[0013] A determination module is configured to determine a target liveness detection result of the target object based on the first liveness detection result and the second liveness detection result.

[0014] In a third aspect, the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and one or more of the computer programs are executed by the at least one processor to enable the at least one processor to perform the above-mentioned liveness detection method.

[0015] In a fourth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned liveness detection method when executed by a processor.

[0016] In a fifth aspect, the present disclosure provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned liveness detection method.

[0017] The liveness detection method provided in the embodiments of the present disclosure obtains liveness detection results by extracting and detecting features from video data and acoustic signals. This allows for the acquisition of features in multiple dimensions of vision and acoustic waves, thereby enhancing feature representation. Furthermore, liveness detection can be performed in combination with video and acoustic waves, thereby integrating liveness detection results from multiple dimensions and improving liveness detection accuracy and performance.

[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art by describing detailed example embodiments with reference to the accompanying drawings. In the accompanying drawings:

[0020] Figure 1 A diagram illustrating an application scenario of the liveness detection method and apparatus provided in an embodiment of the present disclosure;

[0021] Figure 2 A flowchart of a liveness detection method provided in an embodiment of the present disclosure;

[0022] Figure 3 is a logic block diagram of a liveness detection method according to an embodiment of the present disclosure;

[0023] Figure 4 A block diagram of a living body detection device provided in an embodiment of the present disclosure;

[0024] Figure 5 A block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] To enable those skilled in the art to better understand the technical solutions of the present disclosure, exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0026] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.

[0027] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0028] The terms used herein are only used to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof is not excluded. Similar words such as "connected" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.

[0029] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.

[0030] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution complies with relevant national laws and regulations (for example, the "Information Security Technology Personal Information Security Specification", etc.). For example: corresponding prescribed measures are taken to control access to personal information; the display of personal information is subject to prescribed restrictions; the purpose of using personal information does not exceed the scope of direct or reasonable connection; when using personal information, clear identity reference is eliminated to avoid precise positioning of specific individuals.

[0031] During identity verification, users may disguise their identity through means such as photos, reducing security and accuracy. Therefore, liveness detection is crucial. Related technologies primarily use two-dimensional image pixel texture analysis to detect forgeries of non-live individuals. However, this method is limited by lighting, pixel texture, and other factors, resulting in low accuracy.

[0032] Therefore, the present disclosure provides a liveness detection method that obtains video data and sound wave signals for a target object, and then extracts and detects features of the video data and sound wave signals to obtain liveness detection results. By combining vision and sound waves for liveness detection, liveness detection results that are integrated in multiple dimensions are achieved, thereby improving the accuracy and performance of liveness detection.

[0033] Figure 1 The following diagram schematically illustrates an application scenario of the living body detection method and device provided by the embodiments of the present disclosure.

[0034] like Figure 1 As shown, an application scenario of an embodiment of the present disclosure may include a terminal device 101, a network 103, and a server 102. The network 103 is used as a medium for providing a communication link between the terminal device 101 and the server 102. The network 103 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0035] The user can use the terminal device 101 to interact with the server 102 via the network 103 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0036] The terminal device 101 may be any electronic device having a display screen and supporting web browsing, including but not limited to a smart phone, a tablet computer, a laptop computer, a desktop computer, and the like.

[0037] The server 102 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the terminal device 101. The background management server may analyze and process received user requests and other data, and feed back the processing results (e.g., web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0038] It should be noted that the liveness detection method and apparatus provided in the embodiments of the present disclosure can be executed by the server 102. Accordingly, the liveness detection method and apparatus provided in the embodiments of the present disclosure can be set in the server 102. The liveness detection method and apparatus provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 102 and can communicate with the terminal device 101 and / or the server 102. Accordingly, the liveness detection method and apparatus provided in the embodiments of the present disclosure can also be set in a server or server cluster that is different from the server 102 and can communicate with the terminal device 101 and / or the server 102.

[0039] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0040] Figure 2 This is a flow chart of a method for detecting a living body provided by an embodiment of the present disclosure. Figure 2 , the method comprising:

[0041] S210: Acquire video data collected for the target object, and acquire the sound wave signal reflected by the target object.

[0042] Specifically, the present disclosure provides a possible implementation method, which includes transmitting an acoustic wave signal at a preset transmission frequency to a target object, receiving an acoustic wave signal reflected back after the acoustic wave signal reaches the target object, and collecting video data of the target object; in response to identifying that the target object performs a preset behavior, stopping video data collection and acoustic wave signal reception, and obtaining the collected acoustic wave signal and video data of the target object.

[0043] The emitted sound wave signal may be, for example, an ultrasonic signal or other audio signal, etc., which is not limited in the embodiments of the present disclosure.

[0044] For example, the target object triggers identity authentication based on an application in the terminal. After authorization by the target object, the application can prompt the target object to perform actions such as nodding and shaking the head, and can also require the target object to move with a larger amplitude and be within a certain distance from the terminal. This can ensure that there is a certain frequency shift from the emission of the sound wave signal to the reception, thereby improving the effectiveness of the sound wave signal; and turn on the speaker of the terminal to transmit an ultrasonic signal to the target object, while recording the sound wave signal and video data. When the correct action is recognized, the recording is stopped, and the video data and sound wave signal can be collected.

[0045] Among them, the sound wave signal can be received by the microphone in the terminal, and the video data can be collected by the camera in the terminal. There is no limitation in the embodiments of the present disclosure. In this way, the sound wave signal and video data required for liveness detection can be obtained based on the existing speakers, microphones and cameras in the terminal. No additional hardware equipment is required, and there is no need to rely on special hardware. An ordinary terminal can be used, which saves resources.

[0046] S220: Extract features from video frames in the video data to obtain a first feature of the video data, and determine a first liveness detection result based on the first feature.

[0047] S230: Extract features from the acoustic wave signal to obtain a second feature of the acoustic wave signal, and determine a second liveness detection result based on the second feature.

[0048] S240: Determine a target liveness detection result of the target object according to the first liveness detection result and the second liveness detection result.

[0049] For this step, the present disclosure provides possible embodiments, specifically including: 1) when the first liveness detection result and / or the second liveness detection result are used to indicate that the target object is non-live, determining the target liveness detection result is used to indicate that the target object is non-live; 2) when both the first liveness detection result and the second liveness detection result are used to indicate that the target object is alive, determining the target liveness detection result is used to indicate that the target object is alive.

[0050] In particular, in the embodiment of the present disclosure, when determining the target liveness detection result, the first liveness detection result may be judged first. If the first liveness detection result is not liveness, it can be determined that there is a possibility of an identity authentication attack. At this time, the second liveness detection result may no longer be judged or the process step of determining the second liveness detection result based on the acoustic signal may not be performed. This is not limited in the embodiment of the present disclosure. If the first liveness detection result is liveness, the second liveness detection result is then judged. If the second liveness detection result is also liveness, the target liveness detection result is determined to be liveness. In this way, the target object is determined to be live only when it is determined to be live based on both the video data and the acoustic signal. This can improve the accuracy of liveness detection of the target object. If the second liveness detection result is not liveness, the target liveness detection result is determined to be not liveness. In this way, based on the video data and the acoustic signal, at least one is judged to be not liveness, that is, the final target detection result is determined to be not liveness, which can further improve accuracy and reduce the risk of attack.

[0051] In the embodiment of the present disclosure, video data and acoustic wave signals for a target object are obtained, a first liveness detection result is determined based on a first feature extracted from the video data, and a second liveness detection result is determined based on a second feature extracted from the acoustic wave signal. Thus, a target liveness detection result of the target object is determined based on the first liveness detection result and the second liveness detection result. The detection results of the video data and the acoustic wave signal can be integrated to improve the accuracy of the final target liveness detection result and enhance the robustness and detection capability of the liveness detection.

[0052] The following describes the liveness detection method according to an embodiment of the present disclosure.

[0053] With respect to the above step S220, feature extraction is performed on the video frame in the video data to obtain a first feature of the video data, and a first liveness detection result is determined based on the first feature, including:

[0054] 1) Determine facial key points included in a video frame in the video data.

[0055] Among them, taking a real person as an example, facial key points may include at least one of the following: eyebrows, eyes, nose, mouth, facial contour, etc. Each determined facial key point can represent the coordinate position of the characteristic position of the facial key point in the video frame. The continuous sequence of facial key points determined by the video frame in the video data can express action information.

[0056] 2) Extract features of facial key points included in the video frame to obtain the first feature of the video data.

[0057] 3) Based on the first feature, predict the probability that the target object in the video data presents a first action.

[0058] In the embodiment of the present disclosure, a visual classifier can be pre-trained, so that based on the visual classifier, feature extraction is performed on the facial key points included in the video frame to obtain the first feature of the video data, and classification prediction is performed on the first feature to obtain the probability of the target object presenting each first action in the video data.

[0059] Among them, the first feature can represent the extracted facial key point feature information. For example, by identifying and analyzing the facial key points in the video frame, determining the position data of the facial key points included in the video frame in the video data, splicing the position data of the facial key points of the video frame, and inputting it into the visual classifier, the probability of the target object presenting the first action can be output.

[0060] 4) Determine a first liveness detection result based on the probability that the target object presents the first action.

[0061] In some possible embodiments, the first action includes: simulated nodding, simulated shaking head, the probability that the target object presents simulated nodding in the video data is a first probability, and the probability that the target object presents simulated shaking head in the video data is a second probability.

[0062] Determining a first liveness detection result based on a probability that the target object presents a first action includes: when the sum of the first probability and the second probability is greater than or equal to a first threshold, determining that the first liveness detection result is used to indicate that the target object is non-live; and when the sum of the first probability and the second probability is less than the first threshold, determining that the first liveness detection result is used to indicate that the target object is alive.

[0063] Among them, simulated nodding refers to the behavior of a non-living object simulating nodding, and simulated shaking head refers to the behavior of a non-living object simulating shaking head. It can also be understood that simulated nodding and simulated shaking head represent the behavior of a non-living object, which is an imitation of the behavior of a living object. For example, by shaking the terminal used for identity authentication to simulate the nodding or shaking head movement of a real user, or shaking an electronic screen, photo, etc. to simulate the nodding or shaking head movement of a real person, etc.

[0064] In addition, to further improve the classification accuracy, the first action may also include real-person nodding and real-person shaking head. In this way, based on the visual classifier, the probability that the target object may present simulated nodding, simulated shaking head, real-person nodding, and real-person shaking head can be obtained. Then, the first liveness detection result can be determined as non-live or alive based on the sum of the probabilities of simulated nodding and simulated shaking head. The first liveness detection result can also be determined as non-live or alive based on the sum of the probabilities of real-person nodding and real-person shaking head. There is no limitation on this. Of course, in the embodiment of the present disclosure, there is no limitation on the first action, and other actions may also be included.

[0065] In the embodiment of the present disclosure, by modeling and analyzing multiple first actions and training a visual classifier, the probability of the first action that the target object may present can be obtained based on the visual classifier, and then the first liveness detection result can be determined, thereby improving the accuracy and reliability of vision-based liveness detection.

[0066] With respect to the above step S230, feature extraction is performed on the acoustic wave signal to obtain a second feature of the acoustic wave signal, and a second liveness detection result is determined based on the second feature, including:

[0067] 1) Extract the target time-frequency information of the acoustic signal to obtain the second feature of the acoustic signal.

[0068] In an embodiment of the present disclosure, the target time-frequency information of the sound wave signal can be calculated based on short-time Fourier transform. For example, the target time-frequency information can be a time-frequency graph, and the second feature can represent the feature information of the extracted target time-frequency information.

[0069] 2) Based on the second feature, predicting the probability that the target object presents the second action, and based on the second feature, predicting the probability that the terminal presents the first terminal state, the video data and the sound wave signal are collected by the terminal.

[0070] In the embodiment of the present disclosure, a sound wave classifier can be pre-trained, so that based on the sound wave classifier, feature extraction is performed on the target time-frequency information of the first sound wave signal to obtain the second feature of the sound wave signal, and the second feature is classified and detected to obtain the probability that the target object may present each second action, and the probability that the terminal may present each first terminal state.

[0071] 3) Determine a second liveness detection result based on the probability that the target object presents the second action and the probability that the terminal presents the first terminal state.

[0072] In some possible embodiments, the second action includes: simulated nodding and simulated shaking of the head, the first terminal state includes a stationary state and a shaking state, the probability that the target object presents a simulated nod is the third probability, the probability that the target object presents a simulated shaking of the head is the fourth probability, the probability that the terminal presents a stationary state is the fifth probability, and the probability that the terminal presents a shaking state is the sixth probability.

[0073] Based on the probability that the target object presents the second action and the probability that the terminal presents the first terminal state, determining the second liveness detection result, including: 1) when the sum of the third probability, the fourth probability, the fifth probability and the sixth probability is greater than or equal to the second threshold, determining that the first liveness detection result is used to indicate that the target object is non-live; 2) when the sum of the third probability, the fourth probability, the fifth probability and the sixth probability is less than the second threshold, determining that the first liveness detection result is used to indicate that the target object is alive.

[0074] In the disclosed embodiment, simulated nodding and simulated shaking of the head represent simulated behaviors of non-living objects, and the static state and shaking state of the terminal can represent situations that may occur in the terminal when the non-living object performs an attack or simulated behavior. For example, the non-living attack behavior is verified by slightly shaking the terminal to simulate a breakthrough action while playing a video of the user nodding or shaking his head on the electronic screen. For another example, the terminal is stationary and plays a video of the user nodding or shaking his head on the electronic screen.

[0075] In addition, to further improve the classification accuracy, the second action may also include real-person nodding and real-person shaking head. Based on the sound wave classifier, the target time-frequency information of the sound wave signal is input, and the probabilities corresponding to real-person nodding, real-person shaking head, simulated nodding, simulated shaking head, terminal static state, and terminal shaking state can be output. Therefore, the second liveness detection result can be determined as non-live or alive based on the sum of the probabilities of simulated nodding, simulated shaking head, terminal static state, and terminal shaking state, or the second liveness detection result can be determined as non-live or alive based on the sum of the probabilities of real-person nodding and real-person shaking head.

[0076] In this way, in the embodiment of the present disclosure, by analyzing the sound wave signal, multiple second actions are modeled, and a sound wave classifier is trained to obtain the sound wave classifier. Then, based on the sound wave classifier, the probability of each second action that the target object may present and each terminal state that the terminal may present is determined, and then the second liveness detection result is determined, thereby improving the accuracy of liveness detection based on sound waves.

[0077] Based on the above embodiments, the training process of the visual classifier and the acoustic wave classifier in the embodiments of the present disclosure is briefly described below.

[0078] Taking a real person as an example, we can first analyze and model possible two-dimensional attack behaviors of the user. For example, possible attack behaviors include: 1) Shaking the terminal used for authentication to simulate a nod or shake of the user's head, or shaking a digital screen or photo to simulate a nod or shake of the user's head. 2) The terminal remains stationary while a video of the user nodding or shaking their head is played on the digital screen. 3) Playing a video of the user nodding or shaking their head on the digital screen while gently shaking the terminal to simulate a breakthrough action verification.

[0079] Furthermore, through analysis and modeling, and combining the respective characteristics and advantages of video and sound waves, the sound wave classifier based on the Doppler frequency shift principle has a high degree of differentiation between real-person nodding or shaking movements and terminal stillness or any slight shaking state, while the visual classifier based on facial key points has a high degree of differentiation between real-person nodding or shaking movements and simulated nodding or shaking movements. Therefore, corresponding actions can be set for the visual classifier and the sound wave classifier respectively, which can improve the accuracy of liveness detection. In a possible embodiment, the first action modeled for the visual classifier includes simulated nodding, simulated shaking head, real-person nodding and real-person shaking head, and the second action modeled for the sound wave classifier includes simulated nodding, simulated shaking head, real-person nodding and real-person shaking head, as well as the still state and shaking state representing the terminal state.

[0080] Based on the modeling analysis in the above embodiments, during training, training data is first obtained. Specifically, a possible implementation method is provided for obtaining training data. For identity authentication scenarios, operations are performed according to simulated nodding, simulated shaking head, real person nodding, real person shaking head, the terminal's static state, and the terminal's shaking state, and corresponding video data and sound wave signals are collected synchronously. For example, a real person user is prompted to nod, and video data and sound wave signals corresponding to the real person nodding are collected. Furthermore, in the disclosed embodiments, a visual classifier can be trained based on the video data corresponding to each action, and a sound wave classifier can be trained based on the sound wave signals corresponding to each action.

[0081] In one possible embodiment, for a visual classifier, facial key points included in the video frames in the video data are determined based on the video data corresponding to simulated nodding, simulated shaking head, real-person nodding and real-person shaking head, and a visual classifier is trained based on the facial key points of each corresponding video frame, wherein the cross-entropy loss function for training the visual classifier is the cross-entropy between the predicted action obtained based on the facial key points of the video frame and the corresponding action label. The model training method of the stochastic gradient descent of the cross-entropy loss function can be adopted, which is not limited in the embodiments of the present disclosure.

[0082] For example, for the facial key points included in the video frames in the video data, the facial key points of each video frame can be spliced, such as into an array of 2*H*W, where 2 is the number of channels, representing the two position coordinate points of each facial key point, H represents the number of facial key points in each video frame, and W represents the number of video frames in the video data. Then, based on the convolutional neural network, a visual classifier is trained to obtain the facial key point splicing array with four action labels of simulated nodding, simulated shaking head, real nodding and real shaking head.

[0083] In a possible embodiment, for the sound wave classifier, the target time-frequency information is calculated for the sound wave signals corresponding to simulated nodding, simulated shaking head, real person nodding, real person shaking head, the stationary state of the terminal, and the shaking state of the terminal, and the sound wave classifier is trained based on the target time-frequency information corresponding to the categories of each action and state.

[0084] For example, the target time-frequency information is a time-frequency graph, the horizontal axis of the time-frequency graph is time, and the vertical axis is frequency and energy. The time-frequency graph can be represented as an array of 1xFxT, where F is the frequency dimension and T is the time dimension. Each row in the time-frequency graph corresponds to the frequency energy value of a frequency, and the center frequency usually has a higher energy value. Furthermore, in order to improve accuracy and facilitate the display of the Doppler frequency shift phenomenon caused by the user's action, the center frequency area can also be removed, and the time-frequency graph with the center frequency area removed can be used as training data for the sound wave classifier. Then, the target time-frequency information of the six category labels of simulated nodding, simulated shaking head, real person nodding, real person shaking head, stationary state of the terminal, and shaking state of the terminal can be used to train the sound wave classifier. Among them, the cross-entropy loss function for training the sound wave classifier is the cross entropy between the predicted category and the corresponding category label obtained based on the target time-frequency information. The stochastic gradient descent model training method of the cross entropy loss function can be adopted, which is not limited in the embodiments of the present disclosure.

[0085] In this way, in the embodiment of the present disclosure, by analyzing various situations of plane attacks and determining the action modeling results for vision and sound waves, it is more comprehensive, and then through training to obtain visual classifiers and sound wave classifiers, it is possible to accurately identify the categories of various actions or terminal states, thereby improving the accuracy and detection capability of liveness detection.

[0086] Based on the above embodiments, the overall implementation logic of the liveness detection method in the embodiment of the present disclosure is described. Figure 3 FIG. 1 is a logic block diagram of a liveness detection method according to an embodiment of the present disclosure.

[0087] 1) If Figure 3 As shown, for example, when performing face identity authentication, the terminal can prompt to perform a preset action and turn on the terminal's speaker to transmit a sound wave signal to the target object, while recording the sound wave signal received by the microphone and reflected by the target object, as well as the video data collected by the camera.

[0088] 2) Performing time-frequency and other processing on the sound wave signal to obtain target time-frequency information of the sound wave signal, performing feature extraction on the target time-frequency information of the sound wave signal to obtain a second feature, and performing classification detection on the second feature to predict the probability of the target object presenting the second action, and predicting the probability of the terminal presenting the first terminal state.

[0089] Among them, in the embodiment of the present disclosure, feature extraction and classification detection can be performed on the target time-frequency information based on the trained sound wave classifier to obtain the probability that the target object presents the second action and the probability that the terminal presents the first terminal state.

[0090] 3) Determine the facial key points included in the video frame in the video data, perform feature extraction on the facial key points included in the video frame to obtain a first feature of the video data, and perform classification detection on the first feature to predict the probability of the target object presenting a first action in the video data, thereby determining a first liveness detection result.

[0091] In the embodiment of the present disclosure, feature extraction and classification detection can be performed on facial key points of the video frame based on a trained visual classifier to obtain the probability that the target object in the video data presents the first action.

[0092] 4) Determine a target liveness detection result of the target object based on the first liveness detection result and the second liveness detection result.

[0093] In the embodiment of the present disclosure, when performing liveness detection, video data and sound wave signals of the target object are obtained, and feature extraction and classification detection are performed on the sound wave signals to obtain a second liveness detection result. Feature extraction and classification detection are performed on the video data to obtain a first liveness detection result. The features of the video and sound waves can be integrated to make liveness detection decisions, thereby improving the accuracy of liveness detection.

[0094] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0095] In addition, the present disclosure also provides a liveness detection device, an electronic device, and a computer-readable storage medium, all of which can be used to implement any liveness detection method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method section and will not be repeated here.

[0096] Figure 4 A block diagram of a living body detection device provided in an embodiment of the present disclosure.

[0097] Reference Figure 4 The present disclosure provides a liveness detection device, which includes:

[0098] An acquisition module 41 is configured to acquire video data collected from a target object and to acquire an acoustic wave signal reflected from the target object;

[0099] a first processing module 42 configured to extract features from video frames in the video data to obtain a first feature of the video data, and determine a first liveness detection result based on the first feature;

[0100] a second processing module 43, configured to extract features from the acoustic wave signal to obtain a second feature of the acoustic wave signal, and determine a second liveness detection result based on the second feature;

[0101] The determination module 44 is configured to determine a target liveness detection result of the target object according to the first liveness detection result and the second liveness detection result.

[0102] In a possible embodiment, when extracting features from the video frames in the video data to obtain a first feature of the video data, and determining a first liveness detection result based on the first feature, the first processing module 42 is configured to:

[0103] Determining facial key points included in a video frame in the video data;

[0104] Extracting features of facial key points included in the video frame to obtain a first feature of the video data;

[0105] Predicting, based on the first feature, a probability that the target object in the video data exhibits a first action;

[0106] A first liveness detection result is determined according to the probability that the target object presents the first action.

[0107] In a possible embodiment, the first action includes: a simulated nod and a simulated head shake, wherein the probability that the target object presents the simulated nod in the video data is a first probability, and the probability that the target object presents the simulated head shake in the video data is a second probability;

[0108] When determining the first liveness detection result based on the probability that the target object presents the first action, the first processing module 42 is configured to:

[0109] When the sum of the first probability and the second probability is greater than or equal to a first threshold, determining that the first liveness detection result indicates that the target object is non-live;

[0110] When the sum of the first probability and the second probability is less than the first threshold, it is determined that the first living body detection result indicates that the target object is alive.

[0111] In a possible embodiment, when extracting features from the acoustic wave signal to obtain a second feature of the acoustic wave signal and determining a second liveness detection result based on the second feature, the second processing module 43 is configured to:

[0112] performing feature extraction on target time-frequency information of the acoustic wave signal to obtain a second feature of the acoustic wave signal;

[0113] Based on the second feature, predicting a probability that the target object will exhibit a second action, and based on the second feature, predicting a probability that the terminal will exhibit a first terminal state, the video data and the sound wave signal being collected by the terminal;

[0114] A second liveness detection result is determined based on the probability that the target object presents the second action and the probability that the terminal presents the first terminal state.

[0115] In one possible embodiment, the second action includes: a simulated nod and a simulated head shake, and the first terminal state includes a stationary state and a shaking state; the probability that the target object presents the simulated nod is a third probability, the probability that the target object presents the simulated head shake is a fourth probability, the probability that the terminal presents the stationary state is a fifth probability, and the probability that the terminal presents the shaking state is a sixth probability;

[0116] When determining the second liveness detection result based on the probability that the target object presents the second action and the probability that the terminal presents the first terminal state, the second processing module 43 is configured to:

[0117] When the sum of the third probability, the fourth probability, the fifth probability and the sixth probability is greater than or equal to a second threshold, determining that the first liveness detection result indicates that the target object is non-living;

[0118] In a case where the sum of the third probability, the fourth probability, the fifth probability and the sixth probability is less than a second threshold, it is determined that the first living body detection result indicates that the target object is alive.

[0119] In a possible embodiment, when determining the target liveness detection result of the target object based on the first liveness detection result and the second liveness detection result, the determination module 44 is configured to:

[0120] In a case where the first liveness detection result and / or the second liveness detection result is used to indicate that the target object is non-live, determining that the target liveness detection result is used to indicate that the target object is non-live;

[0121] In a case where both the first liveness detection result and the second liveness detection result are used to indicate that the target object is alive, it is determined that the target liveness detection result is used to indicate that the target object is alive.

[0122] Each module in the above-mentioned liveness detection device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0123] Figure 5 A block diagram of an electronic device provided in an embodiment of the present disclosure.

[0124] Reference Figure 5 An embodiment of the present disclosure provides an electronic device, which includes: at least one processor 501; at least one memory 502, and one or more I / O interfaces 503 connected between the processor 501 and the memory 502; wherein the memory 502 stores one or more computer programs that can be executed by the at least one processor 501, and the one or more computer programs are executed by the at least one processor 501 so that the at least one processor 501 can perform the above-mentioned liveness detection method.

[0125] Each module in the above-mentioned electronic device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0126] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-mentioned liveness detection method. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0127] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned liveness detection method.

[0128] It will be understood by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable storage medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium).

[0129] As is well known to those skilled in the art, the term computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information (such as computer-readable program instructions, data structures, program modules or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technology, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those skilled in the art, communication media typically contains computer-readable program instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0130] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0131] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0132] The computer program product described herein may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0133] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0134] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0135] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0136] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0137] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.

Claims

1. A method for detecting a living body, characterized in that: include: Acquiring video data collected from a target object and acquiring an acoustic wave signal reflected by the target object; Performing feature extraction on video frames in the video data to obtain a first feature of the video data, and determining a first liveness detection result based on the first feature; performing feature extraction on the acoustic wave signal to obtain a second feature of the acoustic wave signal, and determining a second liveness detection result based on the second feature; A target liveness detection result of the target object is determined according to the first liveness detection result and the second liveness detection result.

2. The method for liveness detection according to claim 1, wherein: The extracting features from the video frames in the video data to obtain a first feature of the video data, and determining a first liveness detection result based on the first feature, includes: Determining facial key points included in a video frame in the video data; Extracting features of facial key points included in the video frame to obtain a first feature of the video data; Predicting, based on the first feature, a probability that the target object in the video data exhibits a first action; A first liveness detection result is determined according to the probability that the target object presents the first action.

3. The method for liveness detection according to claim 2, wherein: The first action includes: simulated nodding and simulated shaking of the head, the probability that the target object presents the simulated nodding in the video data is a first probability, and the probability that the target object presents the simulated shaking of the head in the video data is a second probability; The determining a first liveness detection result according to the probability that the target object presents the first action includes: When the sum of the first probability and the second probability is greater than or equal to a first threshold, determining that the first liveness detection result indicates that the target object is non-live; When the sum of the first probability and the second probability is less than the first threshold, it is determined that the first living body detection result indicates that the target object is alive.

4. The method for detecting living body according to claim 1, wherein: The extracting a feature of the acoustic wave signal to obtain a second feature of the acoustic wave signal, and determining a second liveness detection result based on the second feature, includes: performing feature extraction on target time-frequency information of the acoustic wave signal to obtain a second feature of the acoustic wave signal; Based on the second feature, predicting a probability that the target object will exhibit a second action, and based on the second feature, predicting a probability that the terminal will exhibit a first terminal state, the video data and the sound wave signal being collected by the terminal; A second liveness detection result is determined based on the probability that the target object presents the second action and the probability that the terminal presents the first terminal state.

5. The method for liveness detection according to claim 4, wherein: The second action includes: a simulated nod and a simulated head shake; the first terminal state includes a stationary state and a shaking state; the probability that the target object presents the simulated nod is a third probability, the probability that the target object presents the simulated head shake is a fourth probability, the probability that the terminal presents the stationary state is a fifth probability, and the probability that the terminal presents the shaking state is a sixth probability; The determining a second liveness detection result based on the probability that the target object presents the second action and the probability that the terminal presents the first terminal state includes: When the sum of the third probability, the fourth probability, the fifth probability and the sixth probability is greater than or equal to a second threshold, determining that the first liveness detection result indicates that the target object is non-living; In a case where the sum of the third probability, the fourth probability, the fifth probability and the sixth probability is less than a second threshold, it is determined that the first living body detection result indicates that the target object is alive.

6. The liveness detection method according to any one of claims 1 to 5, characterized in that: The determining, based on the first liveness detection result and the second liveness detection result, a target liveness detection result of the target object includes: In a case where the first liveness detection result and / or the second liveness detection result is used to indicate that the target object is non-live, determining that the target liveness detection result is used to indicate that the target object is non-live; In a case where both the first liveness detection result and the second liveness detection result are used to indicate that the target object is alive, it is determined that the target liveness detection result is used to indicate that the target object is alive.

7. A living body detection device, characterized in that: include: An acquisition module, configured to acquire video data collected from a target object and to acquire an acoustic wave signal reflected by the target object; a first processing module, configured to extract features from video frames in the video data, obtain a first feature of the video data, and determine a first liveness detection result based on the first feature; a second processing module, configured to perform feature extraction on the acoustic wave signal to obtain a second feature of the acoustic wave signal, and determine a second liveness detection result based on the second feature; A determination module is configured to determine a target liveness detection result of the target object based on the first liveness detection result and the second liveness detection result.

8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor. The one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the living body detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the living body detection method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The invention comprises a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the living body detection method according to any one of claims 1 to 6.