Living body detection method, device, computer device and storage medium
By outputting motion indication information and sound wave signals, combining the characteristics of action video and reflected sound wave signals, double verification is achieved, which solves the problem that living body detection is easily attacked by forgery in the prior art and improves the accuracy of detection.
Patent Information
- Application Number
- CN202110064764.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-01-18
AI Technical Summary
Existing biometric recognition technologies are prone to forgery attacks in live detection, resulting in low accuracy of live detection results.
By outputting motion indication information and the first sound wave signal, the action video and reflected sound wave signals of the detection object are obtained, the action amplitude characteristics and sound wave motion characteristics are extracted, and the two are combined for live detection to achieve dual verification.
It improves the accuracy of live detection, effectively resists forged attacks, and ensures the accuracy of detection results.
Smart Images

Figure CN114821820B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method and device for live detection, a computer device, a storage medium, and a storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, in the era of high-speed informatization, it is very important to protect personal identity and information security. For example, in various scenarios such as terminal unlocking, online payment, and access control, it is necessary to verify the identity of users. Currently, some biometric recognition technologies have emerged, such as fingerprint recognition, face recognition, etc.
[0003] In the related art, by collecting the image of the detection object on-site and identifying the biometric features in the image for live detection to verify the identity of the detection object. However, this detection method only focuses on the information in the image, is easily attacked by forgery, and the accuracy of the live detection result is relatively low. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method and device for live detection, a computer device, and a storage medium that can effectively improve the accuracy of the live detection result.
[0005] A method for live detection, the method includes:
[0006] Output motion indication information and a first sound wave signal; the first sound wave signal is directed at the detection object moving according to the motion indication information;
[0007] Obtain an action video collected for the moving detection object, and locate the action interval corresponding to the detection object according to the action amplitude feature in the action video;
[0008] Obtain a second sound wave signal reflected by the first sound wave signal through the detection object, and extract a sound wave motion feature from the target motion signal in the second sound wave signal;
[0009] Cut out the sound wave motion feature corresponding to the action interval from the sound wave motion feature;
[0010] Perform live detection according to the action amplitude feature and the sound wave motion feature corresponding to the action interval to obtain the live detection result of the detection object.
[0011] A device for live detection, the device includes:
[0012] A data output module, configured to output motion indication information and a first sound wave signal; the first sound wave signal is directed at the detection object moving according to the motion indication information;
[0013] An action video processing module, configured to obtain an action video collected for the detected object in motion, and locate an action interval corresponding to the detected object according to the action amplitude feature in the action video;
[0014] A sound wave signal processing module, configured to obtain a second sound wave signal reflected by the detected object from the first sound wave signal, and extract a sound wave motion feature from the target motion signal in the second sound wave signal;
[0015] A living body detection module, configured to cut out a sound wave motion feature corresponding to the action interval from the sound wave motion feature; perform living body detection according to the action amplitude feature and the sound wave motion feature corresponding to the action interval, and obtain a living body detection result of the detected object.
[0016] In one embodiment, the action video processing module is further configured to perform action detection on the action video to obtain an action amplitude feature in the action video; determine an action start time and an action end time of the detected object according to the action amplitude feature; and locate an action interval corresponding to the detected object according to the action start time and the action end time.
[0017] In one embodiment, the action video processing module is further configured to perform key point detection on each video frame in the action video to obtain action key points and action regions corresponding to each video frame; perform action detection according to the action key points and the action regions corresponding to each video frame, and respectively obtain action features corresponding to each video frame; and obtain an action amplitude feature corresponding to the action video according to the time sequence of the action video and the action features corresponding to each video frame.
[0018] In one embodiment, the sound wave signal processing module is further configured to perform signal demodulation on the second sound wave signal to obtain a component signal of the second sound wave signal; perform interference cancellation on the component signal to obtain a target motion signal in the second sound wave signal; and perform feature extraction on the target motion signal to obtain a sound wave motion feature corresponding to the target motion signal.
[0019] In one embodiment, the sound wave signal processing module is further configured to perform dynamic interference cancellation on the component signal based on a preset interception frequency to obtain a component signal after dynamic interference cancellation; extract a static component in the component signal after dynamic interference cancellation, and perform static interference cancellation on the static component to obtain a target motion signal in the second sound wave signal.
[0020] In one embodiment, the living body detection module is further configured to synchronize and align the motion amplitude feature and the acoustic wave motion feature according to the time sequence of the motion amplitude feature and the time sequence of the acoustic wave motion feature; and cut the synchronized and aligned acoustic wave motion feature according to the motion start time and the motion end time corresponding to the motion interval, so as to obtain the acoustic wave motion feature corresponding to the motion interval.
[0021] In one embodiment, the living body detection module is further configured to perform motion detection on the motion amplitude feature to obtain a first motion category corresponding to the motion amplitude feature; perform motion detection on the acoustic wave motion feature corresponding to the motion interval to obtain a second motion category corresponding to the acoustic wave motion feature corresponding to the motion interval; and determine the living body detection result of the detection object according to the first motion category, the second motion category, and the motion indication information.
[0022] In one embodiment, the living body detection module is further configured to determine that the living body detection result of the detection object passes when the first motion category is consistent with the second motion category, and the first motion category and the second motion category are consistent with the indicated motion category in the motion indication information.
[0023] In one embodiment, the living body detection module is further configured to generate a corresponding acoustic wave time-frequency map according to the acoustic wave motion feature corresponding to the motion interval; input the acoustic wave time-frequency map into a trained target classification model, extract time-frequency map features from the acoustic wave time-frequency map through the target classification model; and perform motion classification on the acoustic wave time-frequency map according to the time-frequency map features to obtain a second motion category corresponding to the acoustic wave motion feature.
[0024] In one embodiment, the above-mentioned living body detection device further includes a model training module, configured to obtain a sample acoustic wave time-frequency map and a sample label; the sample acoustic wave time-frequency map is generated based on the sample acoustic wave signal reflected by the sample object from the collected first acoustic wave signal, and the sample label is an action annotation label for the sample object in the sample acoustic wave time-frequency map; input the sample acoustic wave time-frequency map into a classification model to be trained, extract sample time-frequency map features corresponding to the sample acoustic wave time-frequency map through the classification model to be trained; perform motion classification according to the sample time-frequency map features to obtain a predicted motion category; and adjust the parameters of the classification model based on the difference between the predicted motion category and the sample label and continue training until the training condition is met, and then end the training to obtain a target classification model.
[0025] In one embodiment, the above-mentioned living body detection device further includes a posture adjustment module, which is used to obtain the face image corresponding to the detection object; extract features from the face image to obtain face features; determine the face posture of the detection object according to the face features; when the face posture does not meet the posture condition, output posture adjustment information to indicate the detection object to adjust the face posture; the data output module is further used to output motion indication information and a first sound wave signal when the face posture meets the posture condition.
[0026] In one embodiment, the above-mentioned living body detection device further includes a face recognition module, which is used to obtain the face image corresponding to the detection object; extract the current face features of the face image; perform face recognition on the face image based on the current face features and the target face features corresponding to the detection object to obtain the face recognition result of the detection object; the above-mentioned living body detection device further includes an identity verification module, which is used to determine the identity verification result of the detection object according to the face recognition result and the living body detection result.
[0027] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0028] Output motion indication information and a first sound wave signal; the first sound wave signal points to the detection object moving according to the motion indication information;
[0029] Obtain the action video collected for the moving detection object, and locate the action interval corresponding to the detection object according to the action amplitude feature in the action video;
[0030] Obtain the second sound wave signal reflected by the detection object from the first sound wave signal, and extract the sound wave motion feature from the target motion signal in the second sound wave signal;
[0031] Cut out the sound wave motion feature corresponding to the action interval from the sound wave motion feature;
[0032] Perform living body detection according to the action amplitude feature and the sound wave motion feature corresponding to the action interval to obtain the living body detection result of the detection object.
[0033] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0034] Output motion indication information and a first sound wave signal; the first sound wave signal points to the detection object moving according to the motion indication information;
[0035] Obtain the action video collected for the detected object in motion, and locate the action interval corresponding to the detected object according to the action amplitude characteristics in the action video;
[0036] Obtain the second acoustic wave signal reflected by the detected object from the first acoustic wave signal, and extract the acoustic wave motion characteristics from the target motion signal in the second acoustic wave signal;
[0037] Cut out the acoustic wave motion characteristics corresponding to the action interval from the acoustic wave motion characteristics;
[0038] Perform live detection according to the action amplitude characteristics and the acoustic wave motion characteristics corresponding to the action interval to obtain the live detection result of the detected object.
[0039] A computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of the computer device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the following steps are implemented:
[0040] Output motion indication information and a first acoustic wave signal; the first acoustic wave signal is directed at the detected object moving according to the motion indication information;
[0041] Obtain the action video collected for the detected object in motion, and locate the action interval corresponding to the detected object according to the action amplitude characteristics in the action video;
[0042] Obtain the second acoustic wave signal reflected by the detected object from the first acoustic wave signal, and extract the acoustic wave motion characteristics from the target motion signal in the second acoustic wave signal;
[0043] Cut out the acoustic wave motion characteristics corresponding to the action interval from the acoustic wave motion characteristics;
[0044] Perform live detection according to the action amplitude characteristics and the acoustic wave motion characteristics corresponding to the action interval to obtain the live detection result of the detected object.
[0045] The above-mentioned living body detection method, device, computer device and storage medium output motion indication information and a first sound wave signal, where the first sound wave signal points to a detection object moving according to the motion indication information, and then obtain an action video collected for the moving detection object and a second sound wave signal reflected by the detection object from the first sound wave signal. Furthermore, by extracting action amplitude features and corresponding action intervals from the action video, extracting sound wave motion features from the second sound wave signal, and cutting out sound wave motion features corresponding to the action intervals from the sound wave motion features, living body detection is performed by combining the action amplitude features and the sound wave motion features corresponding to the action intervals. Thus, it is possible to detect whether the actions in the action video are synchronized with the motions in the reflected second sound wave signal and whether they are consistent with the action indication information, so as to be able to perform double verification from the image vision level and the sound wave signal level, effectively improving the accuracy of living body detection. Description of the Drawings
[0046] Figure 1 It is an application environment diagram of the living body detection method in an embodiment;
[0047] Figure 2 It is a flowchart of the living body detection method in an embodiment;
[0048] Figure 3 It is a timing curve diagram corresponding to the action amplitude features of the action video in an embodiment;
[0049] Figure 4 It is a time-frequency diagram of the sound wave motion features in an embodiment;
[0050] Figure 5 It is a flowchart of cutting out the sound wave motion features corresponding to the action intervals in an embodiment;
[0051] Figure 6 It is a time-frequency diagram corresponding to the sound wave signals reflected by multiple actions in an embodiment;
[0052] Figure 7 It is a flowchart of training a target classification force model in an embodiment;
[0053] Figure 8 It is a test interface diagram of living body detection in an embodiment;
[0054] Figure 9 It is a schematic diagram of a face collection interface in an embodiment;
[0055] Figure 10 It is a schematic diagram of a result display interface of living body detection results in an embodiment;
[0056] Figure 11 It is a flowchart of the living body detection method in another embodiment;
[0057] Figure 12 is a schematic flowchart of a living body detection method in another embodiment;
[0058] Figure 13 is a structural block diagram of a living body detection device in one embodiment;
[0059] Figure 14 is an internal structure diagram of a computer device in one embodiment;
[0060] Figure 15 is an internal structure diagram of a computer device in another embodiment. Detailed implementation manners
[0061] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0062] The living body detection method provided by the present application can be applied to a computer device. The computer device can be a terminal or a server. It can be understood that the living body detection method provided by the present application can be applied to a terminal, can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server.
[0063] The living body detection method provided by the present application can be applied to, for example, Figure 1 the application environment shown in the figure. Among them, the terminal 102 communicates with the server 104 through a network. Among them, the terminal 102 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The server 104 can be an independent physical server, can also be a server cluster or a distributed system composed of multiple physical servers, and can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 102 and the server 104 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.
[0064] Among them, cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". As a basic capability provider of cloud computing, a cloud computing resource pool (abbreviated as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to select and use.
[0065] Specifically, the server 104 outputs motion indication information and a first sound wave signal to the terminal 102. The terminal 102 includes a speaker 102a and a microphone 102b. The terminal 102 then outputs the motion indication information and outputs the first sound wave signal through the speaker 102a. The first sound wave signal is directed at a detection object that moves according to the motion indication information. The terminal 102 collects an action video corresponding to the moving detection object through a camera, and collects a second sound wave signal reflected by the first sound wave signal through the microphone 102b, and uploads the action video and the second sound wave signal to the server 104. The server 104 obtains the action video collected for the moving detection object, locates the action interval corresponding to the detection object according to the action amplitude characteristics in the action video; and obtains the second sound wave signal reflected by the first sound wave signal through the detection object, and extracts the sound wave motion characteristics from the target motion signal in the second sound wave signal; cuts out the sound wave motion characteristics corresponding to the action interval from the sound wave motion characteristics. The server 104 then performs a live detection based on the action amplitude characteristics and the sound wave motion characteristics corresponding to the action interval, and obtains the live detection result of the detection object.
[0066] It can be understood that the live detection methods in the embodiments of the present application adopt computer vision technology and machine learning technology in artificial intelligence technology, etc., and can effectively realize automatically identifying the action categories of the detection objects in the video, as well as identifying the action categories in the reflected sound wave signals for live detection. Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0067] Computer Vision (CV) is a science that studies how to make machines "see". To put it more concretely, it refers to the use of cameras and computers to replace human eyes to identify, follow and measure targets, and further perform graphic processing so that the computer processes the images into images that are more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. It can be understood that this application uses computer vision technology to detect the action category of an object from an image frame in a video.
[0068] Machine Learning (ML) is a multi-disciplinary cross-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all fields of artificial intelligence. Machine learning and deep learning generally include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning. It can be understood that the classification model used in some embodiments of the present application, as well as the domain-specific networks corresponding to each field, are obtained by training using machine learning technology based on artificial intelligence. The classification model obtained by training based on the machine learning technology can more accurately classify the action categories of the detection object in the image to perform live detection on the detection object.
[0069] In one embodiment, Figure 2 As shown, a method for detecting a living body is provided, and the method is described by taking the application of the method to a computer device as an example. The computer device may be Figure 1 It is understood that the method can also be applied to a system including a terminal and a server and implemented through the interaction between the terminal and the server. In this embodiment, the following steps are included:
[0070] S202, outputting motion indication information and a first sound wave signal; the first sound wave signal points to the detection object moving according to the motion indication information.
[0071] It can be understood that liveness detection is a way to determine the real physiological characteristics of the detected object in some identity verification scenarios. In face recognition applications, liveness detection can verify whether the user is the real live person through combined actions such as blinking, opening the mouth, shaking the head, and nodding, using technologies such as facial key point positioning and face tracking. It can effectively resist common attack methods such as photos, face swapping, masks, occlusion, and screen reshoots.
[0072] Among them, the motion indication information is information used to indicate that the detection object moves according to the indicated action category. It can be understood that the action category refers to the category of actions corresponding to the movement of the moving part of the detection object, where the moving part refers to the part of the detection object that moves, which can specifically be a local part or the whole body part. For example, the moving part includes at least one of the eyes, lips, face, head, and hands. For example, the action category can include at least one of opening the mouth, shaking the head, nodding the head, blinking, reading numbers, etc.
[0073] It can be understood that the form of the motion indication information can include at least one of text, image, voice, etc. For example, the motion indication information in text form directly shows the indicated action through text; the motion indication information in image form indicates the indicated action by showing it in the image, and the motion indication information in voice form outputs the indicated action through voice.
[0074] Among them, sound wave is a mechanical wave. The propagation of the vibration generated by the sound - emitting body in air or other substances is called a sound wave, which is the propagation form of sound. The sound wave signal is an audio signal. Specifically, the first sound wave signal is an ultrasonic wave signal. Among them, ultrasonic wave refers to a mechanical wave with a vibration frequency greater than 20000 Hz. Its vibration frequency per second is relatively high, exceeding the general upper limit of human ear hearing (20000 Hz), and people usually can't hear the propagation of ultrasonic waves. Ultrasonic wave signals have a high frequency and a short wavelength, and have good beam - shooting and directivity when propagating within a certain distance. It can be understood that the detection object is any object that needs to be subjected to live detection. Specifically, the detection object is a human body or a human face. It can be understood that in an actual scenario, the detection object may be a real human body or face, or may be an unreal human body or face. For example, the unreal human body or face includes at least one of a physical photo, an image displayed on an electronic screen, a physical three - dimensional face model with real human face features, etc. In each embodiment of the present application, it is precisely necessary to detect whether the detection object is a real object.
[0075] During the process of live detection, first, the computer device outputs the motion indication information and the first sound wave signal. Specifically, the server can output the motion indication information and the first sound wave signal to the terminal, and then the terminal with an information prompt function and an ultrasonic wave playback function outputs the motion indication information and the first sound wave signal to indicate that the detection object moves according to the motion indication information. In another embodiment, it can also directly be the terminal with an information prompt function and an ultrasonic wave playback function that outputs the motion indication information and the first sound wave signal.
[0076] Among them, the first acoustic wave signal is directed at the detection object moving according to the motion indication information, that is, the first acoustic wave signal propagates towards the detection object moving according to the motion indication information. As a result, after the first acoustic wave signal is reflected by the moving detection object, a second acoustic wave signal is formed by reflection. It can be understood that the signal type of the first acoustic wave signal is an ultrasonic wave signal, and the signal type of the second acoustic wave signal formed by reflecting the first acoustic wave signal is also an ultrasonic wave signal.
[0077] In one embodiment, the motion indication information is further used to indicate that the detection object moves within a specified area according to the motion indication information. Among them, the specified area means that the detection object is in the image acquisition area. For example, the specified area can be a preset distance range between the detection object and the terminal, or a distance range within which the terminal can collect an image of the face area of the detection object.
[0078] Specifically, the image acquisition area can also be displayed on the display screen of the terminal. The detection object can move its position or adjust its posture so that the detected part is within the image acquisition area. The image acquisition area can then display the picture corresponding to the detected part of the detection object. The picture corresponding to the detected part can specifically be an image or a video.
[0079] In one embodiment, before outputting the first acoustic wave signal, it further includes: obtaining a preset audio signal; randomizing the carrier frequency of the preset audio signal to generate the first acoustic wave signal.
[0080] Among them, the preset audio signal is a pre-configured ultrasonic wave signal.
[0081] Before the computer device outputs the first acoustic wave signal, it can also randomize the carrier frequency of the preset audio signal through a signal generator to generate and output the first acoustic wave signal. Specifically, after the computer device obtains the preset audio signal, when randomizing the carrier frequency of the preset audio signal, it can perform tone superposition on the preset audio signal. The specific expression can be as follows:
[0082]
[0083] Among them, 2A is the amplitude, f k is the carrier frequency of the signal, and N is the total number of subcarriers. We use a random number generator to generate the frequency f k。To avoid interference between adjacent frequency signals, the frequency interval Δf between any two tones can be specified. For example, the frequency interval can be at least 300 Hz. After the computer device performs tone superposition and carrier frequency following on the preset audio signal, a first sound wave signal is generated. By randomizing the carrier frequency of the generated audio, audio replay attacks can be resisted. Therefore, the attacker cannot pass the liveness detection-based authentication by replaying previously recorded audio signals, effectively ensuring the accuracy and security of the authentication.
[0084] In one embodiment, since the human ear cannot hear audio signals with frequencies higher than 18 KHz, and the audio hardware of most terminal devices is not very sensitive to sounds with frequencies higher than 21 KHz, the frequency of the ultrasonic signal can be set within the range of 18 - 21 KHz. Thus, it can effectively ensure that the output audio signal is a sound wave signal inaudible to the human ear and can also be effectively collected by the audio hardware of the terminal, thereby effectively ensuring the effectiveness of the output first sound wave signal to further effectively perform liveness detection on the detection object.
[0085] S204. Obtain the action video collected for the moving detection object, and locate the action interval corresponding to the detection object according to the action amplitude feature in the action video.
[0086] It can be understood that the action video refers to the video formed by the continuous frame pictures collected, that is, the continuous video frames corresponding to the moving detection object. Among them, the action video is collected in chronological order, so the collected action video has a time sequence.
[0087] Among them, the action interval refers to the time interval corresponding to the start time and end time of the same continuous action during the movement of the detection object in the action video. The action interval reflects the time period during which the detection object moves. The action interval of the detection object moving in the action video can be one or multiple. Among them, multiple means two or more.
[0088] During the process of performing liveness detection on the detection object, after the terminal outputs the motion indication information and the first sound wave signal, the terminal starts to collect the action video corresponding to the detection object in real time, and at the same time, it also collects the second sound wave signal reflected by the first sound wave signal through the detection object.
[0089] After the computer device outputs the motion indication information and the first acoustic wave signal, it acquires the action video collected for the detected object of the motion and the second acoustic wave signal reflected by the detected object from the first acoustic wave signal. Specifically, after the computer device acquires the action video collected for the detected object of the motion, it first extracts the action feature from the action video to extract the action amplitude feature corresponding to the action video. Then, the computer device locates the action interval corresponding to the detected object according to the time sequence of the action video and the corresponding action amplitude feature. In this way, the action amplitude feature and the corresponding action interval can be accurately and effectively detected from the action video.
[0090] S206. Acquire the second acoustic wave signal reflected by the detected object from the first acoustic wave signal, and extract the acoustic wave motion feature from the target motion signal in the second acoustic wave signal.
[0091] Among them, the second acoustic wave signal reflected by the detected object from the first acoustic wave signal refers to the second acoustic wave signal formed after the output first acoustic wave signal is reflected by the detected object of the motion. The signal types of the first acoustic wave signals are the same. It can be understood that after the first acoustic wave signal is output, multiple propagation paths will be generated during the propagation process. Therefore, the second acoustic wave signal collected after being reflected by the lip may include some interfering acoustic wave signals.
[0092] It can be understood that the acoustic wave motion feature refers to the feature that reflects the motion of the detected object in the collected second acoustic wave, and this feature is reflected through the feature of the acoustic wave signal. Taking the second acoustic wave signal as an ultrasonic wave as an example, the ultrasonic wave signal can produce the micro-Doppler effect, that is, when the target or the detected object has a radial motion relative to the radar or the signal acquisition device, there is still a small-amplitude motion component of the target or the components of the target relative to the radar, and this phenomenon is called micro-motion. Among them, the small amplitude refers to the radial distance between the target and the radar. For a single-scattering target, the micro-motion is reflected in the non-uniform motion of the target. For a multi-scattering target, the micro-motion is reflected in the non-rigidity of the target, and the non-rigidity means that there is still relative motion between the components of the target. Any structural component on the target or the target has vibrations, rotations, and accelerating motions in addition to the translational motion of the centroid, and these micro-motions will cause additional frequency modulation on the received signal and generate offset frequencies near the Doppler frequency shift generated by the movement of the target body. Due to the uniqueness of the micro-Doppler, the micro-Doppler frequency shifts are different. Therefore, by processing the second acoustic wave signal, the feature that reflects the motion of the detected object in the second acoustic wave signal can be extracted.
[0093] Specifically, after the computer device obtains the second acoustic wave signal reflected by the first acoustic wave signal passing through the detection object, it performs signal conditioning processing on the second acoustic wave signal to remove some interfering acoustic wave signals in the second acoustic wave signal, and only retains the signals related to the moving parts of the detection object, so as to extract the target motion signal from the second acoustic wave signal. Then, the computer device performs feature extraction on the extracted target motion signal to obtain the acoustic wave motion feature corresponding to the target motion signal.
[0094] S208, Cut out the acoustic wave motion feature corresponding to the action interval from the acoustic wave motion feature.
[0095] It can be understood that the collected action video is a continuous video frame with time sequence, and the collected second acoustic wave signal is also a continuous surrounding signal with time sequence. Therefore, both the action video and the second acoustic wave signal carry corresponding timestamps respectively.
[0096] Specifically, after the computer device extracts the action amplitude feature and the corresponding action interval from the action video, and extracts the acoustic wave motion feature from the second acoustic wave signal, it can be understood that the action amplitude feature corresponding to the action video is also a continuous action amplitude feature with time sequence corresponding to each video frame in the action video respectively. Similarly, the acoustic wave motion feature corresponding to the second acoustic wave signal is also a continuous acoustic wave feature with time sequence that can reflect the movement of the detection object.
[0097] The computer device then searches for the timestamps corresponding to the start time and end time of the action interval in the second acoustic wave signal according to the start time and end time of the action interval, and then cuts out the acoustic wave motion feature corresponding to the action interval from the time period corresponding to the acoustic wave motion feature in the second acoustic wave signal. The cut acoustic wave motion feature can be a signal segment corresponding to the action interval. When there are multiple action intervals, the acoustic wave motion features corresponding to the signal segments corresponding to the multiple action intervals are cut out. Thus, the acoustic wave motion feature corresponding to the action interval in the action video can be cut out from the second acoustic wave signal based on time sequence synchronization.
[0098] S210, Perform liveness detection according to the action amplitude feature and the acoustic wave motion feature corresponding to the action interval, and obtain the liveness detection result of the detection object.
[0099] Among them, the action amplitude feature can detect whether the detection object is a live body from the image vision level. The acoustic wave motion feature can detect whether the detection object is a live body from the acoustic wave signal level.
[0100] After the computer device extracts the motion amplitude features and the corresponding motion intervals from the motion video, extracts the acoustic wave motion features from the second acoustic wave signal, and cuts out the acoustic wave motion features corresponding to the motion intervals from the acoustic wave motion features, it performs liveness detection by combining the motion amplitude features and the acoustic wave motion features corresponding to the motion intervals. Thereby, it can detect whether the motion in the motion video is synchronized with the motion in the reflected second acoustic wave signal and whether it is consistent with the motion instruction information, so that double verification can be carried out from the image vision level and the acoustic wave signal level, effectively ensuring the accuracy of the liveness detection result of the detection object.
[0101] In the above liveness detection method, during liveness detection, by outputting motion instruction information and a first acoustic wave signal, the first acoustic wave signal points to the detection object moving according to the motion instruction information, and then the computer device acquires the motion video collected for the moving detection object and the second acoustic wave signal reflected by the detection object from the first acoustic wave signal. The computer device then extracts the motion amplitude features and the corresponding motion intervals from the motion video, extracts the acoustic wave motion features from the second acoustic wave signal, and cuts out the acoustic wave motion features corresponding to the motion intervals from the acoustic wave motion features. After that, it performs liveness detection by combining the motion amplitude features and the acoustic wave motion features corresponding to the motion intervals. Thereby, it can detect whether the motion in the motion video is synchronized with the motion in the reflected second acoustic wave signal and whether it is consistent with the motion instruction information, so that double verification can be carried out from the image vision level and the acoustic wave signal level, effectively improving the accuracy of liveness detection.
[0102] In one embodiment, locating the motion interval corresponding to the detection object according to the motion amplitude features in the motion video includes: performing motion detection on the motion video to obtain the motion amplitude features in the motion video; determining the motion start time and the motion end time of the detection object according to the motion amplitude features; and locating the motion interval corresponding to the detection object according to the motion start time and the motion end time.
[0103] It can be understood that since the motion video includes continuous video frames with time sequence. Among them, each video frame carries a corresponding timestamp.
[0104] After the computer device acquires the motion video collected for the detection object, it performs motion detection on the motion video. Specifically, the computer device performs motion detection on each video frame in the motion video respectively to obtain the motion amplitude features corresponding to each video frame. Thus, according to the timestamp carried by each video frame and the corresponding motion amplitude features, continuous motion amplitude features with time sequence corresponding to the motion video can be obtained.
[0105] For example, the amplitude of the movement of the detected object in each video frame can be detected, and for example, the amplitude value of the movement of the moving part can be obtained according to the proportion of the moving part in the video frame. Then, by arranging the amplitude values of the moving parts in each video frame in the order of time stamps, a time series curve graph corresponding to the amplitude feature of the action video can be obtained.
[0106] In one embodiment, since there are usually random or error components in the time series data, in order to more clearly distinguish the patterns in the time series data, after obtaining the time series curve graph corresponding to the amplitude feature of the action, the time series curve graph can also be smoothed. Specifically, a preset smoothing function can be used, such as algorithms like the neighborhood average filter and the linear sliding average method, to smooth the time series curve. Then, according to the preset amplitude threshold, when the amplitude value at which the curve crosses from low to high reaches the first amplitude threshold, the moment can be determined as the start time of the action. Similarly, when the amplitude value at which the curve crosses from high to low reaches the second amplitude threshold, the moment can be determined as the end time of the action.
[0107] As Figure 3 shown, it is a schematic diagram of the time series curve graph corresponding to the amplitude feature of the action video in one embodiment. Among them, the horizontal axis is time and the vertical axis is the amplitude value. From Figure 3 the time series curve graph, it can be seen that at the position of time 3a, the amplitude values at the positions before and after time 3a and the amplitude values at the positions before and after time 3b change significantly. For time 3a, the amplitude value at which the curve crosses from low to high in the previous and subsequent moments is relatively large, so time 3a can be determined as the start time of the action, specifically, it can be the moment of opening the mouth. For time 3b, the amplitude value at which the curve crosses from high to low in the previous and subsequent moments is relatively large, so time 3a can be determined as the start time of the action, specifically, it can be the moment of closing the mouth. The action interval corresponding to this amplitude feature is the time period corresponding to 3a - 3b.
[0108] Then, the computer device can determine the start time and end time of the action of the detected object according to the amplitude feature, that is, according to the amplitude values of each video frame in the action video, determine the video frames at which the detected object starts to move and ends to move in the action video, and then determine the time stamp corresponding to the video frame at which the movement starts as the start time of the action, and determine the time stamp corresponding to the video frame at which the movement ends as the end time of the action. The computer device can then locate the action interval corresponding to the detected object according to the start time and end time of the action.
[0109] In this embodiment, by performing action detection on each video frame in the action video, and then according to the action amplitude characteristics of each video frame, it is possible to effectively obtain continuous action amplitude characteristics with time sequence corresponding to the action video. And according to the continuous action amplitude characteristics with time sequence corresponding to the action video, it is possible to accurately identify the action start time and action end time when the detection object in the action video performs the movement, and further accurately locate the action interval corresponding to the detection object.
[0110] In one embodiment, action detection is performed on the action video to obtain the action amplitude characteristics in the action video, including: performing key point detection on each video frame in the action video respectively to obtain the action key points and action regions corresponding to each video frame; performing action detection according to the action key points and action regions corresponding to each video frame to obtain the action characteristics corresponding to each video frame respectively; and obtaining the action amplitude characteristics corresponding to the action video according to the time sequence of the action video and the action characteristics corresponding to each video frame.
[0111] Among them, key point detection refers to detecting the key points of the moving parts of the detection object in the video frame image.
[0112] After the computer device obtains the action video, it first performs key point detection on each video frame in the action video. Specifically, the computer device can first perform key point recognition on each video frame through a preset key point detection algorithm. For example, perform face key point recognition on each video frame to obtain the face key points of each video frame. Among them, the face key points can include key points of facial features and contour key points, etc. The face key points can specifically include key points corresponding to at least one part among the eye part, eyebrow part, nose part, lip part, ear part, and jaw line part, etc.
[0113] After the computer device recognizes the face key points in each video frame, it determines the action key points and action regions corresponding to each video frame according to the change of the face key points between consecutive video frames. For example, when the lip of the detection object moves, the position distribution of the lip key points in each video frame of the action video changes. Therefore, according to the changing face key points in consecutive video frames, the action key points and action regions can be determined. For example, the action key point can be the lip key point, and the corresponding action region is the lip region.
[0114] The computer device then performs action detection according to the action key points and action regions corresponding to each video frame. Specifically, the action characteristics are determined according to the action key points and action regions corresponding to each video frame. Specifically, the action characteristic can be an action amplitude value. For example, the action amplitude value can be the ratio of the action key point to the corresponding action region in each video frame, such as the aspect ratio of the width to the height of the action key point in the action region.
[0115] Then, the computer device can obtain the motion amplitude feature corresponding to the motion video according to the time sequence of the motion video and the motion features corresponding to each video frame. Thus, the continuous motion amplitude features corresponding to each video frame in the motion video with time sequence can be accurately obtained.
[0116] In another embodiment, the computer device can also perform motion detection on the motion video through a pre-trained motion detection network. Among them, the motion detection network can be a neural network model pre-trained based on deep learning algorithms. Specifically, the motion detection network can adopt models such as CNN (Convolutional Neural Network), LSTM (Long Short-Term Memory), DNN (Deep Neural Network), and RNN (Recurrent Neural Network), or a combination of multiple neural network models. This application does not limit this here.
[0117] Specifically, the computer device inputs the motion video into the trained motion detection network, and through the motion detection network, feature extraction and object detection are performed on each video frame in the motion video to identify the motion key points and the motion regions of interest corresponding to each video frame.
[0118] Then, motion detection is performed according to the motion key points and motion regions corresponding to each video frame to respectively obtain the motion features corresponding to each video frame. Furthermore, according to the time sequence of the motion video and the motion features corresponding to each video frame, the motion amplitude feature and the corresponding motion category corresponding to the motion video are identified. Specifically, the computer device can also predict the probability of the start of the motion and the probability of the end of the motion at each position in each video frame to obtain the motion start probability sequence, the motion end probability sequence, and the motion probability sequence. Then, based on the motion start probability sequence, the motion end probability sequence, and the motion probability sequence, the motion feature description corresponding to each motion is predicted to obtain the motion feature with the highest probability. Thus, the motion amplitude feature corresponding to the time sequence of the motion video and each video frame is obtained according to the motion feature with the highest probability. Further, the computer device can directly output the motion time sequence curve graph through the motion detection network, thereby realizing the time sequence motion detection of the motion video.
[0119] In one embodiment, extracting the acoustic wave motion feature from the target motion signal in the second acoustic wave signal includes: demodulating the second acoustic wave signal to obtain the component signal of the second acoustic wave signal; eliminating interference from the component signal to obtain the target motion signal in the second acoustic wave signal; and extracting features from the target motion signal to obtain the acoustic wave motion feature corresponding to the target motion signal.
[0120] Among them, the component signal is the signal component of the analog signal, and the component signal represents that the signal is split into two or more parts. Signals can be divided into in-phase components and quadrature components, DC components and AC components, even components and odd components, sine components and pulse components, etc. Among them, the in-phase component is the signal component in the same direction as the vector direction; the quadrature component is orthogonal to the vector signal, that is, perpendicular to the in-phase component. The component signals of the second acoustic wave signal can specifically include the in-phase component and the quadrature component corresponding to the second acoustic wave signal.
[0121] Among them, the target motion signal refers to a part of the second acoustic wave signal that is only related to the moving part of the detection object.
[0122] It can be understood that the second acoustic wave signal reflected by the first acoustic wave signal after passing through the lips includes acoustic wave signals propagating along multiple paths, such as the reflection path of the user's lips, the propagation path of solids (such as the user's face, etc.), the air propagation path, and the reflection path of surrounding objects, etc. There are some interfering acoustic wave signals among them. Therefore, the computer device needs to extract the target motion signal that is only related to the moving part of the detection object from the second acoustic wave signal.
[0123] During the process of live detection, the second acoustic wave signal reflected by the first acoustic wave signal after passing through the detection object includes multiple paths of propagation. After the computer device obtains the second acoustic wave signal, it can use a dry phase detector pair for down-conversion demodulation, for example, perform down-conversion demodulation to demodulate the baseband signal at a preset carrier frequency. Then the computer device eliminates multipath interference, thereby extracting the target motion signal in the second acoustic wave signal that is only related to the moving part.
[0124] Specifically, the computer device can perform micro-Doppler feature extraction on the obtained second acoustic wave signal to extract the acoustic wave motion characteristics corresponding to the target motion signal in the second acoustic wave. The acoustic wave motion characteristics are the extracted micro-Doppler characteristics. Among them, the micro-Doppler characteristics include parameters such as angular frequency, Doppler amplitude, and initial phase.
[0125] For example, assume that there are M paths in the obtained second acoustic wave signal Rec(t), and the obtained second acoustic wave signal can be described by the following formula:
[0126]
[0127] Among them, i represents the i-th path, where N represents the number of base signals, and k represents the k-th base signal. 2Ai(t) represents the amplitude of the acoustic wave signal in the i-th path, f k represents the carrier frequency, represents the phase shift caused by the propagation delay, represents the phase shift caused by the system delay.
[0128] The original first acoustic wave signal output through the speaker can be regarded as a carrier signal, and the second acoustic wave signal Rec(t) collected through the microphone can be regarded as the superposition of multiple baseband signals subjected to phase shift modulation. Since the generated ultrasonic signal is the superposition of audio signals with different frequencies, the audio played by the speaker can be regarded as the superposition of baseband signals with different frequencies. Since the collected signal is basically synchronized with the played output signal. Therefore, coherent detection can be used to demodulate the collected second acoustic wave signal, and the carrier frequency f k The in-phase component I(t) and the quadrature component Q(t) corresponding to the baseband signal of the second acoustic wave signal above.
[0129] Among them, the expression of the in-phase component I(t) can be as follows:
[0130]
[0131] The expression of the quadrature component Q(t) can be as follows:
[0132]
[0133] Among them, F low is a low-pass filter, and F down is a downsampling function. In the in-phase component I(t), the expression of the part of R k (t)×cos2πf k t is as follows:
[0134]
[0135] The computer device then removes the high-frequency terms of R low through the low-pass filter F k (t)×cos2πf k t, and then performs downsampling through F down . The computer device further performs frequency modulation calculation on the in-phase component I(t) of the baseband signal of the second acoustic wave signal. The calculation formula of the in-phase component I(t) can be as follows:
[0136]
[0137] Similarly, the calculation formula of the quadrature component Q(t) can be as follows:
[0138]
[0139] The in-phase component I(t) and the quadrature component Q(t) corresponding to the adjusted second acoustic wave signal can be calculated through the above formula. After the computer device performs interference cancellation processing based on the obtained in-phase component I(t) and quadrature component Q(t), the computer device further obtains the phase of the signal and performs STFT (Short-Time Fourier Transform) processing on the obtained phase, and then the target motion signal related only to the moving part of the detection object can be obtained. The target motion signal can be expressed as:
[0140] signal = I(t) + Q(t)
[0141] Furthermore, the phase function h D (t) of the target motion signal can be expressed as:
[0142]
[0143] where the initial phase is h D (0). When the phase changes linearly with time, the frequency is a fixed value. Static interference means that when the phase does not change with time, the frequency is 0, that is, the DC component. Static interference is effectively suppressed after passing through a notch filter. The notch filter is a type of band-stop filter with a very narrow stopband, so it is also called a point-stop filter and is often used to remove fixed-frequency components or areas with a very narrow stopband.
[0144] Then, the computer device can obtain the instantaneous frequency f D (t) corresponding to the target motion signal by taking the derivative of the phase function h D (t) of the target motion signal. The instantaneous frequency f D (t) can be expressed as:
[0145]
[0146] where the instantaneous frequency f D (t) is the angular frequency parameter, and the corresponding Doppler amplitude parameter and initial phase parameter can be obtained through the instantaneous frequency f D (t). Thus, the acoustic wave motion characteristics corresponding to the target motion signal are obtained according to the angular frequency parameter, Doppler amplitude parameter, and initial phase parameter, so that the useful information in the micro-Doppler signal can be effectively extracted.
[0147] After short-time Fourier transform, the time-frequency diagram of the target motion signal can be obtained. The time-frequency diagram of the target motion signal can reflect the acoustic wave motion characteristics corresponding to the target motion signal. The time-frequency diagram of the target motion signal is a two-dimensional spectrum and can represent the graph of the spectrum of the target motion signal changing with time. For example,Figure 4 As shown, it is a time-frequency diagram of the acoustic wave motion characteristics in an embodiment, specifically, the time-frequency diagram corresponding to the reflected second acoustic wave signal collected by performing an open-mouth motion on the detection object. Among them, Figure 4 in the time-frequency diagram, the vertical axis is the frequency and the horizontal axis is the time. In the time-frequency diagram, the lighter or brighter the color, the higher the frequency of the acoustic wave signal and the greater the spectral density. The darker or darker the color, the lower the frequency of the acoustic wave signal and the lower the spectral density. From Figure 4 it can be seen that in the time domain from 2s to 4s and around 6s in the time domain, the corresponding spectral density is relatively large, indicating that the open-mouth amplitude in this time domain interval is relatively large.
[0148] In this embodiment, by using dry phase detection to down-convert and demodulate the obtained second acoustic wave signal, the collected signal can be effectively processed to extract the acoustic signal component corresponding to the baseband signal of the second acoustic wave signal. By further eliminating interference from the acoustic signal component, the target motion signal related only to the moving part of the detection object can be accurately and effectively extracted.
[0149] In one embodiment, eliminating interference from the component signal to obtain the target motion signal in the second acoustic wave signal includes: dynamically eliminating interference from the component signal based on a preset interception frequency to obtain the component signal after dynamic interference elimination; extracting the static component in the component signal after dynamic interference elimination, and statically eliminating interference from the static component to obtain the target motion signal in the second acoustic wave signal.
[0150] Among them, interference elimination includes dynamic interference signal elimination and static interference signal elimination. For example, the dynamic interference signal refers to the signal reflected by other nearby moving objects in the living body detection environment except the verification object; the static interference signals include the signals reflected by the solid propagation path, air propagation path, and nearby stationary objects in the living body detection environment except the verification object.
[0151] For the in-phase component and quadrature component in the obtained second acoustic wave signal, in order to improve the accuracy of recognition, it is necessary to remove the interference signals from other paths to retain only the signals related to the moving part of the detection object. After the computer device demodulates the obtained second acoustic wave signal and extracts the acoustic signal component corresponding to the baseband signal of the second acoustic wave signal, by further eliminating interference from the extracted component signal, the computer device can perform dynamic interference elimination and static interference elimination on the component signal respectively.
[0152] Specifically, the computer device can set a preset interception frequency for the filter, and perform dynamic interference cancellation on the component signals based on the preset interception frequency, thereby filtering out the dynamic interference signals and obtaining the component signals after dynamic interference cancellation. Among them, the computer device can also cancel the dynamic interference while demodulating the baseband signal of the second acoustic wave signal, or perform dynamic interference cancellation after demodulating the second acoustic wave signal to obtain the corresponding component signals.
[0153] For example, since the movement of the human torso usually causes signal frequency shifts in the range of 50 - 200 Hz, the maximum frequency shift caused by the movement of the facial features in the human face usually does not exceed 40 Hz. For example, the maximum frequency shift caused by lip movement usually does not exceed 40 Hz. Therefore, according to the action type of the living body detection, the cut-off frequency of the low-pass filter F low used for coherent detection is set as the preset interception frequency. Specifically, the computer device can also set different preset interception frequencies for different action types. For example, for lip movement, the preset interception frequency can be 40 Hz. Perform dynamic interference cancellation on the component signals based on the preset interception frequency, so as to effectively filter out the dynamic interference signals in the component signals.
[0154] After dynamically eliminating the interference, the obtained component signals are the superposition of the acoustic wave signals reflected by the moving parts of the detection object and the static interference signals. The computer device further performs static interference cancellation on the component signals after dynamic interference cancellation.
[0155] Specifically, the in-phase component I(t) can be respectively expressed as the sum of a constant static component I s (t) and the signal reflected by the moving part, and the quadrature component Q(t) can be expressed as the sum of a constant static component Q s (t) and the signal reflected by the moving part. The specific expressions are as follows:
[0156]
[0157] Among them, A lip (t) is the amplitude of the lip reflection signal, d lip is the propagation delay, v is the propagation speed of sound in the air, and θ lip is the phase shift caused by the system delay. Further, the expressions of the in-phase component I(t) and the quadrature component Q(t) can be respectively abbreviated as:
[0158]
[0159] In order to eliminate the static component, the gradient I g (t) of the in-phase component I(t) and the gradient Q g (t) of the quadrature component Q(t) can be further calculated. The specific expressions are as follows:
[0160]
[0161] Among them, I g (t) represents the gradient of the in-phase component I(t), and Q g (t) represents the gradient of the quadrature component Q(t), and A lip (t) and Φ lip (t) are the differential coefficients of A lip (t) and Φ lip (t) respectively. Since the coefficient A lip (t) is inversely proportional to the square of the propagation distance. When the moving part is the face part of the detection object or the facial features part in the face, the movement of the detection object is relatively subtle. Therefore, the value of A lip (t) hardly changes, and thus the value of A lip (t) is approximately zero.
[0162] Therefore, the static component I s (t) corresponding to the in-phase component I(t), and the gradient Q g (t) of the quadrature component Q(t) can be respectively expressed as:
[0163]
[0164] Finally, the slow-changing terms of I g (t) and Q g (t) are eliminated using the least mean square error. After the processing is completed, the final target motion signal representing the moving part of the detection object can be obtained. When there is no movement in the moving part, the magnitudes of I g (t) and Q g (t) are close to zero.
[0165] In this embodiment, by respectively performing dynamic interference cancellation and static interference cancellation on the extracted acoustic signal components, the target motion signal related only to the moving part of the detection object can be accurately and effectively extracted.
[0166] In one embodiment, the acoustic wave motion features corresponding to the action interval are cut out from the acoustic wave motion features, including: synchronously aligning the action amplitude features with the acoustic wave motion features according to the time sequence of the action amplitude features and the time sequence of the acoustic wave motion features; cutting the synchronously aligned acoustic wave motion features according to the action start time and action end time corresponding to the action interval to obtain the acoustic wave motion features corresponding to the action interval.
[0167] Among them, the second acoustic wave signal can be acquired by sampling according to a preset sampling rate of an audio signal. Therefore, the second acoustic wave signal has timestamps corresponding to the sampling points. Synchronous alignment means synchronizing the action video and the second acoustic wave signal according to the acquisition time. For example, the action video and the second acoustic wave signal can be synchronously aligned according to the initial timestamp of the action video or the initial timestamp of the second acoustic wave signal according to the acquired timestamps.
[0168] After the computer device extracts the action amplitude features and corresponding action intervals from the action video, and extracts the acoustic wave motion features from the target motion signal of the second acoustic wave signal, it synchronizes the action video and the second acoustic wave signal, and then aligns the action start time and the action end time of the action interval in the second acoustic wave signal, that is, determines the signal start moment and the signal end moment corresponding to the action start time and the action end time in the second acoustic wave signal respectively.
[0169] Specifically, the computer device can multiply the action start time and the action end time of the action interval by the sampling rate of the second acoustic wave signal to obtain the sampling point positions of the audio signals corresponding to the action start time and the action end time respectively, and then determine the signal start moment and the signal end moment corresponding to the action start time and the action end time in the second acoustic wave signal according to the sampling point positions.
[0170] The computer device then cuts the synchronized acoustic wave motion features according to the signal start moment and the signal end moment corresponding to the action start time and the action end time in the second acoustic wave signal respectively, so as to obtain the acoustic wave motion features corresponding to the action interval.
[0171] For example, as Figure 5 shown, it is a schematic flow chart of cutting out the acoustic wave motion features corresponding to the action interval in an embodiment. Referring to Figure 5 , after the computer device acquires the collected action video 52 and the second acoustic wave signal 54, it first synchronizes the action video and the second acoustic wave signal, that is, synchronously aligns them according to the start acquisition timestamp. The computer device performs key point detection on each video frame in the action video to obtain the key point detection results 521 corresponding to each video frame. Then, according to the key points, the action amplitude value corresponding to each video frame is obtained, and then the action amplitude features with time series corresponding to the action video are obtained, and a corresponding action amplitude time series curve 56 is generated. According to the action amplitude features, the action start time 5a and the action end time 5b can be identified to locate the action interval 5a-5b. At the same time, the computer device demodulates the signal of the second acoustic wave signal and extracts features to extract the acoustic wave motion features from the target motion signal in the second acoustic wave signal, and generates a corresponding acoustic wave time-frequency diagram 58, whereFigure 5 The time-frequency diagram of the sound wave in [description] includes the time-frequency diagrams of the sound wave generated according to 4 different frequency bands. The computer device then synchronizes and aligns the time-frequency diagram of the sound wave with the action amplitude time-sequence curve, and then cuts out the sound wave motion features corresponding to the action interval 5a-5b from the time-frequency diagram of the sound wave. That is, the time-frequency interval in the time-frequency diagram of the sound wave corresponding to the action start time 5a and the action end time 5b, so that the sound wave motion features can be effectively aligned with the action amplitude features in the action video.
[0172] In this embodiment, by synchronizing and aligning the action video and the second sound wave signal, and cutting out the sound wave motion features corresponding to the action interval from the synchronized and aligned sound wave motion features, it is possible to detect whether the action in the action video is synchronized with the motion in the reflected second sound wave signal, thereby effectively improving the accuracy of live detection.
[0173] In one embodiment, live detection is performed according to the action amplitude features and the sound wave motion features corresponding to the action interval to obtain the live detection result of the detection object, including: performing action detection on the action amplitude features to obtain the first action category corresponding to the action amplitude features; performing action detection on the sound wave motion features corresponding to the action interval to obtain the second action category corresponding to the sound wave motion features corresponding to the action interval; and determining the live detection result of the detection object according to the first action category, the second action category, and the motion indication information.
[0174] Among them, the first action category refers to the action category recognized from the action video. The second action category refers to the action category recognized from the second sound wave signal. It can be understood that the first action category and the second action category may be the same or different. For example, when the first action category and the second action category are different, it means that the action corresponding to the action video and the action corresponding to the second sound wave signal are different, for example, a video forgery attack may have occurred. For this situation, it can be directly determined that the live detection of the detection object fails.
[0175] After the computer device performs action detection on the action video to obtain the action amplitude features and the corresponding action interval, and extracts the sound wave motion features from the target motion signal of the second sound wave signal, and divides the sound wave motion features corresponding to the motion interval from the sound wave motion features, the computer device can then perform live detection according to the action amplitude features and the sound wave motion features corresponding to the action interval to obtain the live detection result of the detection object.
[0176] The computer device can determine the living body detection result of the detection object according to the first action category corresponding to the action amplitude feature, the second action category corresponding to the acoustic wave motion feature, and the motion indication information. Specifically, after extracting the action amplitude feature from the action video, the computer device performs action detection on the action amplitude feature to obtain the first action category corresponding to the action amplitude feature.
[0177] After the computer device segments the acoustic wave motion feature corresponding to the motion interval from the acoustic wave motion feature, it performs action detection on the acoustic wave motion feature corresponding to the action interval to obtain the second action category corresponding to the acoustic wave motion feature corresponding to the action interval. Specifically, the computer device can obtain the corresponding second action category by classifying the time-frequency graph of the acoustic wave motion feature corresponding to the action interval.
[0178] The computer device then determines the living body detection result of the detection object according to the first action category, the second action category, and the indicated action category corresponding to the motion indication information.
[0179] In this embodiment, by combining the first action category, the second action category, and the indicated action category corresponding to the motion indication information to determine the living body detection result of the detection object, it is possible to detect whether the action in the action video is synchronized with the motion in the reflected second acoustic wave signal, thereby effectively improving the accuracy of living body detection.
[0180] In one embodiment, determining the living body detection result of the detection object according to the first action category, the second action category, and the motion indication information includes: when the first action category is consistent with the second action category, and the first action category and the second action category are consistent with the indicated action category in the motion indication information, it is determined that the living body detection result of the detection object passes.
[0181] It can be understood that performing living body detection based on a single visual level or based on a single acoustic wave signal level is vulnerable to attacks by video forgery or the acoustic wave signals of attackers. Therefore, in each embodiment of this application, by combining the visual level and the acoustic wave signal level for living body detection, the accuracy of the living body detection result can be effectively guaranteed.
[0182] After the computer device obtains the first action category corresponding to the action video according to the action amplitude feature and the second action category corresponding to the second acoustic wave signal according to the acoustic wave motion feature corresponding to the motion interval, it then determines the living body detection result of the detection object according to the first action category, the second action category, and the indicated action category corresponding to the motion indication information.
[0183] Specifically, when the first action category is inconsistent with the second action category, it is determined that the live detection result of the detection object fails. When the first action category is consistent with the second action category but inconsistent with the indicated action category corresponding to the motion indication information, it is still determined that the live detection result of the detection object fails.
[0184] Only when the first action category is consistent with the second action category, and the first action category and the second action category are consistent with the indicated action category in the motion indication information, it is determined that the live detection result of the detection object passes. That is, the action video needs to pass the detection of the motion indication information, and the second acoustic wave signal collected also needs to pass the detection of the motion indication information, and it also needs to pass the action consistency detection of the action video and the second acoustic wave signal. Only when the above detections all pass, it is determined that the live detection result of the detection object passes.
[0185] In this embodiment, by combining the action video detection and the acoustic wave signal detection, it is possible to effectively combine the vision of the video and the ultrasonic signal for live detection, effectively making up for the defect that a single vision live body is easily forged and attacked, thereby effectively improving the accuracy of the live detection result of the detection object.
[0186] In one embodiment, performing action detection on the acoustic wave motion characteristics corresponding to the action interval to obtain a second action category corresponding to the acoustic wave motion characteristics corresponding to the action interval includes: generating a corresponding acoustic wave time-frequency diagram according to the acoustic wave motion characteristics corresponding to the action interval; inputting the acoustic wave time-frequency diagram into a trained target classification model, and extracting features from the acoustic wave time-frequency diagram through the target classification model to obtain time-frequency diagram features; performing action classification on the acoustic wave time-frequency diagram according to the time-frequency diagram features to obtain a second action category corresponding to the acoustic wave motion characteristics.
[0187] Among them, the time-frequency diagram refers to representing the frequency and amplitude of a signal changing with time in a single graph, also known as a spectrogram. For example, an audio signal can be Fourier-transformed, and then with time as the horizontal axis and frequency as the vertical axis, and different colors are used to represent the assignment values, and the time-frequency diagram of the signal can be drawn. Specifically, the time-frequency diagram can be any one of a time-frequency energy diagram, a time-frequency power spectral density diagram, etc.
[0188] For example, a computer device can perform wavelet transform on the acoustic wave motion characteristics to obtain a time-frequency spectrum function, and then use this time-frequency spectrum function to draw and generate a corresponding acoustic wave time-frequency diagram for the acoustic wave motion characteristics.
[0189] Among them, the magnitude of the power spectral density at a certain frequency and a certain time can be represented by the depth of the color in a certain area of the time-frequency diagram, that is, it represents the magnitude of the ability of that area. For example, the darker the color, the smaller the spectral density of the corresponding area, and the lighter or brighter the color, the greater the spectral density of the corresponding area.
[0190] It can be understood that the target classification model is a machine learning model obtained through pre-training. The target classification model can adopt at least one of a CNN (Convolutional Neural Network) model, an LSTM (Long Short-Term Memory) model, a DNN (Deep Neural Network) model, an RNN (Recurrent Neural Network) model, etc., or can be a combination of multiple neural network models, such as a network structure combining a CNN and an LSTM model. This application does not make a limitation here.
[0191] Specifically, after the computer device generates a corresponding acoustic wave time-frequency diagram according to the acoustic wave motion characteristics, it inputs the acoustic wave time-frequency diagram into the trained target classification model. Then, the target classification model extracts features from the acoustic wave time-frequency diagram, that is, extracts the acoustic wave features at the image level in the acoustic wave time-frequency diagram, so as to obtain the time-frequency diagram features. The time-frequency diagram features reflect the motion characteristics corresponding to the acoustic wave signals in the acoustic wave time-frequency diagram. The amplitudes and spectral densities of different image regions in the acoustic wave time-frequency diagram can represent different categories of actions.
[0192] The computer device then classifies the acoustic wave time-frequency diagram according to the time-frequency diagram features through the target classification model, so as to obtain the second action category corresponding to the acoustic wave motion characteristics. By classifying the time-frequency diagram through the target classification model, the action category corresponding to the acoustic wave signal can be accurately identified from the distribution of amplitudes and spectral densities in the time-frequency diagram.
[0193] For example, as Figure 6 shown, it is the time-frequency diagram corresponding to the acoustic wave signals of multiple action reflections in an embodiment, which reflects the change of the acoustic wave signal spectrum over time. Figure 6 In each of them, the vertical axis of the time-frequency diagram is frequency, and the horizontal axis is time. Figure 6 It respectively shows the time-frequency diagrams of the acoustic wave motion characteristics corresponding to the acoustic wave signals of 4 kinds of reflections. In the time-frequency diagram corresponding to each reflected acoustic wave signal, the amplitudes and spectral densities of different image regions represent different categories of actions. By classifying the time-frequency diagram through the target classification model, the action category corresponding to the acoustic wave signal can be identified. Among them, Figure 6 in the time-frequency diagram (a), the acoustic wave signal represents lip movement, specifically the lips closing quickly, and the corresponding action category is opening the mouth. The acoustic wave signal in the time-frequency diagram (b) represents lip movement, specifically closing after opening the mouth for a period of time, and the corresponding action category is opening the mouth. The acoustic wave signal in the time-frequency diagram (c) represents shaking the head three times, and the corresponding action category is shaking the head. The acoustic wave signal in the time-frequency diagram (d) represents nodding the head three times, and the corresponding action category is nodding.
[0194] In another embodiment, the computer device can also directly output the liveness detection result of the detection object through the target classification model. Specifically, the computer device classifies the sound wave time-frequency diagram through the target classification model to obtain the second action category corresponding to the sound wave motion feature, and then the target classification model determines and outputs the liveness detection result according to the first action category, the second action category and the only action category of the motion indication information, thereby obtaining the liveness detection result of the detection object.
[0195] In one embodiment, a target classification model is obtained through a training step, which includes: obtaining a sample sound wave time-frequency graph and a sample label; the sample sound wave time-frequency graph is generated based on a sample sound wave signal reflected by a sample object after a collected first sound wave signal is generated, and the sample label is an action annotation label for the sample object in the sample sound wave time-frequency graph; the sample sound wave time-frequency graph is input into the classification model to be trained, and the sample time-frequency graph features corresponding to the sample sound wave time-frequency graph are extracted through the classification model to be trained; actions are classified according to the sample time-frequency graph features to obtain a predicted action category; based on the difference between the predicted action category and the sample label, the parameters of the classification model are adjusted and training is continued until the training conditions are met to terminate the training and obtain the target classification model.
[0196] Among them, the sample sound wave time-frequency diagram is the training data for training the target classification model, and the sample label is the training label for training the target classification model. Among them, the sample sound wave time-frequency diagram is generated based on the sample sound wave signal reflected by the sample object after the collected first sound wave signal. It can be understood that each sample sound wave time-frequency diagram is annotated with a corresponding sample label. The sample label is a label that is manually annotated on each sample sound wave time-frequency diagram according to the truth or falsity of the sample object after the sample sound wave time-frequency diagram is collected.
[0197] The sample sound wave signal reflected by the sample object may include at least one of a positive sample sound wave signal, a re-photographed sample sound wave signal, and a forged head sample sound wave signal. The positive sample sound wave signal refers to a sound wave signal collected from a real sample object. The re-photographed sample sound wave signal refers to a sound wave signal collected by re-photographing the sample object on the screen. The forged head sample sound wave signal refers to a sound wave signal collected from a three-dimensional head mold forged based on the sample object.
[0198] It can be understood that the positive sample is a real and accurate sound wave signal corresponding to a living body. The re-photographed sample sound wave signal and the forged head sample sound wave signal, that is, the negative sample, are both forged sound wave signals corresponding to a non-living body. By adding positive and negative samples to the training data, a classification model with higher classification accuracy can be trained.
[0199] It can be understood that the training step of the target classification model is a process of continuous iterative training. Iterative training refers to the process of repeatedly feeding back the results of each round of training and continuing the next round of training based on machine learning, in order to make the classification model to be trained continuously fit and converge to approach and reach the desired goal or result. Specifically, the training methods include but are not limited to supervised training, semi-supervised training and unsupervised training.
[0200] The training condition refers to the end condition of the model training. For example, the training condition may be reaching a preset number of iterations, or the classification performance index of the time-frequency graph of the classification model after adjusting the parameters reaches a preset index. For example, the preset index may include the classification accuracy of the action category of the sound wave signal in the time-frequency graph.
[0201] like Figure 7 As shown in Figure 1, the process diagram of training the target classification model is shown in Figure 1. Figure 7 , the computer device first collects the sample time-frequency diagram. Specifically, in the process of collecting the sample time-frequency diagram, due to the different frequency responses of the terminals for different signal outputs and collections, it is first necessary to perform frequency response self-calibration to select a frequency with a relatively suitable frequency response. For example, for a mobile phone terminal, since the distance between the detection object and the mobile phone terminal is small, it is necessary to select a frequency with a poor frequency response to reduce distance interference. Then the computer device generates a signal, for example, randomizing the carrier frequency of the preset audio to randomly generate an ultrasonic signal. The signal is transmitted through the terminal for signal output and collection to output a first sound wave signal and point to the sample object 72 that moves according to the sample action indication information. Then, the terminal receives the signal to collect the sample sound wave signal reflected by the first sound wave signal through the sample object, refer to the corresponding signal diagram 74. Then, the obtained sample sound wave signal is subjected to I / O demodulation processing, that is, the in-phase component I and the orthogonal component O in the sample sound wave signal are extracted, refer to the corresponding signal diagram 76. Then, the extracted in-phase component I and the orthogonal component O are subjected to differential / noise reduction processing, that is, differential processing and noise reduction processing are performed respectively. And perform STFT Fourier transform to obtain sample sound wave features, refer to the corresponding signal diagram 78, and generate the corresponding sample time-frequency diagram 710. The sample time-frequency diagram is annotated with the corresponding sample label. Among them, the sample sound wave signal reflected by the sample object may include a positive sample sound wave signal, a re-shot sample sound wave signal, and a forged head sample sound wave signal. Figure 7 The sample time-frequency graph includes a sample time-frequency graph corresponding to the positive sample sound wave signal (710a), a sample time-frequency graph corresponding to the re-shot sample sound wave signal (710b), and a sample time-frequency graph corresponding to the forged head sample sound wave signal (710c). Then the computer device inputs the sample time-frequency graph 710 and the corresponding sample label into the classification model 712 for training, so as to obtain a target classification model capable of classifying actions of the sound wave signal time-frequency graph.
[0202] Specifically, during the process of training the classification model, the computer device first inputs the sample acoustic wave time-frequency map into the classification model to be trained. Then, in each round of iterative training, the computer device extracts the sample time-frequency map features corresponding to the sample acoustic wave time-frequency map through the classification model to be trained. Specifically, the convolutional network in the classification model can be used to perform multiple convolutional operations on the sample acoustic wave time-frequency map to extract the features in the sample acoustic wave time-frequency map from multiple image levels and extract the final sample time-frequency map features.
[0203] The computer device then classifies the actions based on the sample time-frequency map features to obtain the predicted action category. The computer device then adjusts the parameters of the classification model based on the difference between the predicted action category and the sample label and continues the training. When the iterative stop condition is not met in this round, it enters the next round of training, takes the next round as this round, continues to extract the sample time-frequency map features of the sample acoustic wave time-frequency map through the classification model, classifies based on the sample time-frequency map features, continues to train the classification model of the predicted action category obtained in this round of classification, and continues the iterative training.
[0204] Specifically, when adjusting the parameters of the classification model, the weight parameters of the classification model can be solved and adjusted using the cross-entropy loss function and the SGD (Stochastic Gradient Descent) algorithm.
[0205] During the process of training the classification model, the label smoothing method can also be used to regularize the sample labels. Thus, the smoothed distribution of the labels is equivalent to adding noise to the true distribution, avoiding the model being too confident in the correct labels, reducing the difference in the output values of the predicted positive and negative samples, and thus effectively avoiding overfitting and improving the generalization ability of the classification model.
[0206] Furthermore, during the process of training the classification model, the Drop mechanism can also be used to train the classification model, that is, during the training process based on the deep learning network, for the neural network units, a certain probability is used to temporarily discard them from the network. That is, during model training, the weights of some hidden layer nodes in the network are randomly made not to work. Those nodes that do not work can be temporarily considered not to be part of the network structure, but their weights need to be retained and only not updated temporarily, and may need to participate in processing when the next sample is input.
[0207] When the training condition is met, the training is stopped, and thus the trained target classification model is obtained. For example, the training condition can be 30 times of iterative training.
[0208] It can be understood that the trained target classification model is a machine learning model capable of classifying the action categories corresponding to the acoustic spectrograms of acoustic signals of various action reflections, so as to accurately identify the action categories reflected in the acoustic spectrograms.
[0209] In another embodiment, after the target classification model extracts the sample spectrogram features corresponding to the sample acoustic spectrograms and identifies the corresponding predicted action categories based on the sample spectrogram features, it further uses a binary classification algorithm to perform liveness classification according to the obtained predicted action categories, such as real liveness or fake liveness, to obtain the result of liveness detection.
[0210] In this embodiment, the target classification model for acoustic spectrograms is trained through the sample acoustic spectrograms obtained based on the acoustic signals reflected by the collected sample objects and the corresponding sample labels, and the parameters of the classification model are gradually adjusted according to the differences between the predicted action categories and the sample labels. Thus, during the parameter adjustment process, the classification model can more accurately extract the spectrogram features reflecting the action categories in the acoustic spectrograms, and then a target classification model with a higher action classification accuracy for acoustic spectrograms can be trained.
[0211] In one embodiment, an application program for liveness detection by combining ultrasonic recognition and action video recognition is tested with the detection object being a human face as an example. As Figure 8 shown, it is a test interface diagram for liveness detection in an embodiment. The test interface includes controls such as "Input Comparison Source", "Flash Core Identity Verification", "Ultrasonic Recognition", and "Lip Reading Liveness", as well as corresponding configuration information such as application identification, security level, and test information. The test interface also includes setting buttons corresponding to saving images and saving requests respectively. The setting button corresponding to saving images is used to save the images collected during the test. The setting button corresponding to saving requests is used to save test requests such as detection requests triggered during the test. Among them, the test information can include, for example, the number of processors, the number of processing units, and the number of reflection objects. The number of processors can specifically be 2, the number of processing units can specifically be 120, and the number of reflection objects can specifically be 2. "Input Comparison Source" means submitting a photo as a reference, such as simulating the user's archived photo in an actual scenario, and performing a 1:1 face comparison with the face actually detected during subsequent face swiping. "Flash Core Identity Verification" represents the face recognition function. "Ultrasonic Recognition" represents the function of performing liveness detection by detecting the ultrasonic waves reflected by the movement of the detection object. "Lip Reading Liveness" refers to the function of performing liveness detection by the ultrasonic waves reflected by lip-reading passwords. The indicated action represents the action category indicating the detection object to perform movement. Since this ultrasonic wave depends on the disturbance generated in the recognition scene, if the action amplitude is too small, it is easily masked by environmental clutter, so opening and closing the mouth can be selected as the action for cooperative detection.
[0212] First, the user can select "ultrasonic recognition" in the test interface to start the live detection. Specifically, by triggering the "ultrasonic recognition" button, the ultrasonic recognition process is entered, and the face collection interface is displayed. As Figure 9 shown, it is a schematic diagram of the face collection interface in an embodiment. The face collection interface includes a face collection area 9a and a time-frequency diagram display area 9b. Among them, the face collection area 9a includes a preview frame 9a1 of the current face image, an action prompt area 9a2, and a light prompt area 9a3. The terminal can first collect the face image of the detection object. After detecting the face and stabilizing it, the terminal outputs action indication information and a first acoustic signal, and displays the collected current face image in the preview frame 9a1 of the face image in the face collection interface. It can be understood that, from the perspective of protecting the privacy of the user's real face image, the eye part in the face image in the preview frame 9a1 is blocked. When recognizing the face image, the actually collected face image includes the eye part. In the action prompt area 9a2 below the face image preview frame, motion indication information is displayed to instruct the detection object to move according to the motion indication information, that is, to perform an open / close mouth action. Among them, the motion indication information can specifically be "Please open your mouth once".
[0213] Furthermore, the current light situation can also be displayed in the light prompt area 9a3 in the face collection interface, such as whether the current light is appropriate. Opening and closing the mouth will respectively cause a perturbation to the ultrasonic signal. Therefore, the ultrasonic time-domain diagram displayed below will generate corresponding fluctuations in real time. The terminal then collects the action video corresponding to the detection object and the reflected ultrasonic signal, and after demodulating and extracting features from the reflected ultrasonic signal, displays the time-domain diagram corresponding to the currently collected ultrasonic signal in the time-frequency diagram display area 9b. Furthermore, a time area 9b1 can also be displayed in the time-frequency diagram display area 9b to display the duration information corresponding to the currently collected ultrasonic signal in the time area 9b1. For example, the duration of the ultrasonic signal displayed in the time area 9b1 is 5 seconds. By performing action detection on the action video, the start time and end time of the action are identified according to the action amplitude characteristics, and the action interval and action category are identified, that is, the start and end time stamps of the action in the action video. Then, the start and end time stamps and action category in the acoustic wave motion characteristics of the reflected ultrasonic signal are identified. The start and end time stamps of the action in the action video are compared with the start and end time stamps in the acoustic wave motion characteristics for consistency verification, and the action category is verified to verify whether the actions in the action video and the reflected ultrasonic signal are synchronized, so as to improve the accuracy of the live detection result.
[0214] After performing live detection by combining the action video and the reflected ultrasonic signal, the returned live detection result can also be displayed. As Figure 10As shown, it is a schematic diagram of the result display interface for the in-vivo detection result in an embodiment. Among them, the result display interface includes a result display box 10a, a time-frequency diagram display area 10b, and a result confirmation box 10c. The in-vivo detection result includes one of "recognition passed" and "recognition failed". When the in-vivo detection result is "recognition passed", the time-frequency diagram corresponding to the reflected ultrasonic signal can also be displayed in the time-frequency diagram display area 10b.
[0215] In one embodiment, as Figure 11 shown, another in-vivo detection method is provided, which specifically includes the following steps:
[0216] S1102, obtain the face image corresponding to the detection object.
[0217] S1104, extract features from the face image to obtain face features; determine the face pose of the detection object according to the face features.
[0218] S1106, when the face pose does not meet the pose condition, output pose adjustment information to instruct the detection object to adjust the face pose.
[0219] S1108, when the face pose meets the pose condition, output motion indication information and a first sound wave signal; the first sound wave signal points to the detection object moving according to the motion indication information.
[0220] S1110, obtain the action video collected for the moving detection object, and locate the action interval corresponding to the detection object according to the action amplitude feature in the action video.
[0221] S1112, obtain the second sound wave signal reflected by the first sound wave signal through the detection object, and extract the sound wave motion feature from the target motion signal in the second sound wave signal.
[0222] S1114, cut out the sound wave motion feature corresponding to the action interval from the sound wave motion feature.
[0223] S1116, perform in-vivo detection according to the action amplitude feature and the sound wave motion feature corresponding to the action interval to obtain the in-vivo detection result of the detection object.
[0224] Among them, the face pose refers to the posture and form of the face of the acquisition object. The face pose includes face distance information and face angle information. The face distance information represents the distance information of the face relative to the image acquisition device. The face angle information represents the angle information of the facial orientation of the face. It can be understood that the pose condition can be that the face distance information meets the distance threshold and the human angle information meets the angle threshold. Both the distance threshold and the angle threshold can be preset numerical ranges.
[0225] It can be understood that when the deviation amplitude of the face angle in the captured face image is large or the distance is far, it is necessary to correct the face pose of the detection object to capture high-quality motion videos and acoustic signals.
[0226] Before the computer device outputs the motion instruction information and the first acoustic signal, it first needs to detect whether the face of the detection object meets the pose condition. Specifically, the computer device first obtains the face image corresponding to the detection object collected, and then extracts features from the face image. Specifically, a preset face detection algorithm can be used to extract the face features in the face image, and the face features can specifically be face key points.
[0227] Next, the computer device determines the face frame corresponding to the detection object according to the face features, and determines the face distance information according to the proportion of the face frame in the face image. Then, further pose estimation is performed according to the face features to obtain the face angle information of the detection object. Specifically, it is possible to estimate the three rotation angles of the face in the face image, namely the pitch angle, the yaw angle, and the roll angle, and the face angle information of the detection object can be obtained according to these three rotation angles.
[0228] The computer device then determines whether the face of the current detection object meets the pose condition according to the face distance information and the face angle information. When the face distance information does not meet the distance threshold, or any one of the face angle information does not meet the angle threshold, it is determined that the face pose does not meet the pose condition. The computer device then outputs pose adjustment information. Among them, the pose adjustment information can specifically be information in text form or information in voice form. To indicate to the detection object to adjust the face pose through the pose adjustment information.
[0229] In another embodiment, when the detection object is a human body and the terminal for collecting images is a handheld device, the computer device can also detect the grip pose of the detection object. Among them, a motion sensor is installed in the terminal, and the grip pose of the detection object on the terminal can be detected through the motion sensor. When the grip pose does not meet the preset threshold, the computer device also outputs pose adjustment information to prompt the detection object to adjust the grip pose. For example, when the detection object uses the terminal with the head down and the terminal is close to the human chest, it will cause vibration interference, so it is necessary to remind the detection object to adjust the grip pose to correct the grip pose of the detection object on the terminal.
[0230] After the face pose of the detection object is adjusted, the computer device continues to obtain the face image of the detection object after the pose is adjusted and performs pose detection. Until the face pose meets the pose condition, the computer device outputs motion indication information and a first sound wave signal, and then obtains the action video corresponding to the detected object collected, and the second sound wave signal reflected by the first sound wave signal through the detected object. The computer device further extracts the action amplitude features in the action video and locates the action interval corresponding to the detected object. At the same time, the sound wave motion features are extracted from the target motion signal in the second sound wave signal, and then the sound wave motion features corresponding to the action interval are cut out from the sound wave motion features. Furthermore, based on the action amplitude features and the sound wave motion features corresponding to the action interval, a live detection is performed to obtain the live detection result of the detection object.
[0231] In this embodiment, by adjusting the face pose of the detection object, action videos and second sound wave signals of higher quality can be collected, and thus the detection object can be more accurately subjected to live detection, effectively improving the accuracy of the live detection result.
[0232] In one embodiment, as Figure 12 shown, another live detection method is provided, which specifically includes the following steps:
[0233] S1202, output motion indication information and a first sound wave signal; the first sound wave signal points to the detection object moving according to the motion indication information.
[0234] S1204, obtain the action video collected for the moving detection object, and locate the action interval corresponding to the detection object according to the action amplitude features in the action video.
[0235] S1206, obtain the second sound wave signal reflected by the first sound wave signal through the detection object, and extract the sound wave motion features from the target motion signal in the second sound wave signal.
[0236] S1208, cut out the sound wave motion features corresponding to the action interval from the sound wave motion features.
[0237] S1210, perform live detection according to the action amplitude features and the sound wave motion features corresponding to the action interval to obtain the live detection result of the detection object.
[0238] S1212, obtain the face image corresponding to the detection object.
[0239] S1214, extract the current face features of the face image.
[0240] S1216, perform face recognition on the face image based on the current face features and the target face features corresponding to the detection object to obtain the face recognition result of the detection object.
[0241] S1218. Determine the identity verification result of the detection object according to the face recognition result and the liveness detection result.
[0242] Among them, face recognition is a biometric identification technology that identifies a person's identity based on the person's facial feature information. A series of related technologies are used to collect images or video streams containing human faces with a camera or webcam, automatically detect and track human faces in the images, and then perform facial recognition on the detected human faces.
[0243] It can be understood that liveness detection can be used to verify the identity of the detection object. Identity verification based on liveness detection includes two parts, namely the face recognition part and the liveness detection part. Only when both face recognition detection and liveness detection pass can it be determined that the identity verification of the detection object passes.
[0244] The computer device can first perform face recognition on the face image corresponding to the detection object. After successful face recognition, further confirm the authenticity of the detection object through liveness detection to enhance the accuracy and security of identity verification. Face recognition detection and liveness detection can also be processed simultaneously, or liveness detection can be performed first and then face recognition detection. The present application does not limit the processing order of face recognition detection and liveness detection.
[0245] Specifically, during the process of face recognition, the computer device can obtain the face image of the detection object based on the identity verification instruction, extract the current facial features of the face image using a face recognition algorithm, and compare the current facial features with the target facial features corresponding to the detection object to perform face recognition on the face image. Among them, the face recognition algorithm can adopt algorithms such as face feature point-based recognition, whole face image-based face recognition, neural network model-based recognition, and illumination model-based recognition. Face recognition is a relatively mature technology and will not be elaborated here.
[0246] After the computer device performs face recognition on the face image, a face recognition result is obtained. The face recognition result includes successful face recognition and failed face recognition. In one embodiment, after successful face recognition, motion indication information and a first sound wave signal can also be output to further perform liveness detection on the detection object.
[0247] Specifically, the computer device acquires the action video corresponding to the detected object and the second acoustic wave signal reflected by the detected object from the first acoustic wave signal. The computer device then extracts the action amplitude feature in the action video and locates the action interval corresponding to the detected object. At the same time, the acoustic wave motion feature is extracted from the target motion signal in the second acoustic wave signal, and then the acoustic wave motion feature corresponding to the action interval is cut out from the acoustic wave motion feature. Then, based on the action amplitude feature and the acoustic wave motion feature corresponding to the action interval, a live detection is performed to obtain the live detection result of the detected object.
[0248] The computer device then determines the identity verification result of the detected object based on the face recognition result and the live detection result. Specifically, when any one of the face recognition result and the live detection result fails, it is determined that the identity verification of the detected object fails. When both the face recognition result and the live detection result pass, it is determined that the identity verification of the detected object passes.
[0249] In this embodiment, by performing face recognition on the detected object and simultaneously performing live detection on the detected object, the live authenticity of the detected object can be effectively detected. And in the live detection, the action amplitude feature at the visual level is extracted from the action video, and the acoustic wave motion feature at the audio signal level is extracted from the reflected second acoustic wave signal, so that the detected object can be effectively verified for multiple identities, effectively improving the accuracy and security of identity verification.
[0250] This application also provides an application scenario that applies the above-mentioned live detection method to implement an identity verification scenario for online payment. Specifically, when a user uses an application running on a terminal to shop or make a payment online, a payment request is initiated through the corresponding application. When making a payment, the user needs to undergo identity verification, and the detected object is the user. The terminal generates an identity verification instruction based on the payment request, and the terminal outputs motion indication information through a display screen and outputs a first acoustic wave signal through a speaker based on the identity verification instruction. During identity verification, the user faces the terminal with their face and moves according to the motion indication information, so that the first acoustic wave signal points to the moving user. For example, it can specifically be the movement of the face.
[0251] After the terminal outputs the motion indication information and the first acoustic signal, it collects the action video of the user performing the motion through the terminal camera, and at the same time collects the second acoustic signal reflected by the lips from the first acoustic signal through the terminal microphone. The terminal then extracts the action amplitude feature in the action video and locates the action interval corresponding to the detection object. At the same time, it extracts the acoustic wave motion feature from the target motion signal in the second acoustic signal, and then cuts out the acoustic wave motion feature corresponding to the action interval from the acoustic wave motion feature. Furthermore, it performs liveness detection based on the action amplitude feature and the acoustic wave motion feature corresponding to the action interval to obtain the liveness detection result of the detection object. During the verification process, the terminal also collects the user's face image, performs face recognition on the face image, and then determines the identity verification result based on the liveness detection result and the face recognition result. If the identity verification result is that the identity verification is passed, the terminal obtains the consumption value of the payment request and subtracts the consumption value from the numerical account of the user who currently requests payment, thus completing the payment.
[0252] This application also provides another application scenario, which applies the above liveness detection method to implement terminal unlocking. Specifically, when the user unlocks the terminal, an unlock request is triggered for the terminal. The terminal generates an identity verification instruction based on the unlock request and verifies the user's identity based on the identity verification instruction. Specifically, when unlocking, the user faces the terminal with the face, and the terminal outputs the motion indication information and the first acoustic signal.
[0253] After the terminal outputs the motion indication information and the first acoustic signal, it collects the action video of the user performing the motion through the terminal camera, and at the same time collects the second acoustic signal reflected by the lips from the first acoustic signal through the terminal microphone. The terminal then extracts the action amplitude feature in the action video and locates the action interval corresponding to the detection object. At the same time, it extracts the acoustic wave motion feature from the target motion signal in the second acoustic signal, and then cuts out the acoustic wave motion feature corresponding to the action interval from the acoustic wave motion feature. Furthermore, it performs liveness detection based on the action amplitude feature and the acoustic wave motion feature corresponding to the action interval to obtain the liveness detection result of the detection object. During the verification process, the terminal also collects the user's face image, performs face recognition on the face image, and then determines the identity verification result based on the liveness detection result and the face recognition result. If the identity verification result is that the identity verification is passed, the terminal performs the unlocking process, thus completing the terminal unlocking.
[0254] It can be understood that the above liveness detection method can also be applied to many other scenarios, which will not be elaborated here.
[0255] It should be understood that although Figure 2 、 11, the steps in the flowchart of 12 are displayed in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 , 11 , at least some of the steps in 12 may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.
[0256] In one embodiment, as Figure 13 shown, a living body detection device 1300 is provided. This device can be a software module, a hardware module, or a combination of both to become part of a computer device. Specifically, the device includes: a data output module 1302, an action video processing module 1304, an acoustic wave signal processing module 1306, and a living body detection module 1308, where:
[0257] The data output module 1302 is used to output motion indication information and a first acoustic wave signal; the first acoustic wave signal is directed at a detection object moving according to the motion indication information.
[0258] The action video processing module 1304 is used to obtain an action video collected for the moving detection object, and locate the action interval corresponding to the detection object according to the action amplitude feature in the action video.
[0259] The acoustic wave signal processing module 1306 is used to obtain a second acoustic wave signal reflected by the detection object from the first acoustic wave signal, and extract an acoustic wave motion feature from the target motion signal in the second acoustic wave signal.
[0260] The living body detection module 1308 is used to cut out the acoustic wave motion feature corresponding to the action interval from the acoustic wave motion feature; perform living body detection according to the action amplitude feature and the acoustic wave motion feature corresponding to the action interval to obtain the living body detection result of the detection object.
[0261] In one embodiment, the action video processing module 1304 is further used to perform action detection on the action video to obtain the action amplitude feature in the action video; determine the action start time and action end time of the detection object according to the action amplitude feature; and locate the action interval corresponding to the detection object according to the action start time and action end time.
[0262] In one embodiment, the action video processing module 1304 is further configured to perform key point detection on each video frame in the action video to obtain the action key points and action regions corresponding to each video frame; perform action detection based on the action key points and action regions corresponding to each video frame to respectively obtain the action features corresponding to each video frame; and obtain the action amplitude features corresponding to the action video according to the time sequence of the action video and the action features corresponding to each video frame.
[0263] In one embodiment, the acoustic wave signal processing module 1306 is further configured to demodulate the second acoustic wave signal to obtain the component signal of the second acoustic wave signal; eliminate interference from the component signal to obtain the target motion signal in the second acoustic wave signal; and extract features from the target motion signal to obtain the acoustic wave motion features corresponding to the target motion signal.
[0264] In one embodiment, the acoustic wave signal processing module 1306 is further configured to perform dynamic interference cancellation on the component signal based on a preset interception frequency to obtain the component signal after dynamic interference cancellation; extract the static component from the component signal after dynamic interference cancellation, and perform static interference cancellation on the static component to obtain the target motion signal in the second acoustic wave signal.
[0265] In one embodiment, the living body detection module 1308 is further configured to synchronously align the action amplitude features and the acoustic wave motion features according to the time sequences of the action amplitude features and the acoustic wave motion features; and cut the synchronized acoustic wave motion features according to the action start time and action end time corresponding to the action interval to obtain the acoustic wave motion features corresponding to the action interval.
[0266] In one embodiment, the living body detection module 1308 is further configured to perform action detection on the action amplitude features to obtain the first action category corresponding to the action amplitude features; perform action detection on the acoustic wave motion features corresponding to the action interval to obtain the second action category corresponding to the acoustic wave motion features corresponding to the action interval; and determine the living body detection result of the detection object according to the first action category, the second action category, and the motion indication information.
[0267] In one embodiment, the living body detection module 1308 is further configured to determine that the living body detection result of the detection object passes when the first action category is consistent with the second action category, and both the first action category and the second action category are consistent with the indicated action category in the motion indication information.
[0268] In one embodiment, the living body detection module 1308 is further configured to generate a corresponding acoustic wave time-frequency map according to the acoustic wave motion characteristics corresponding to the action interval; input the acoustic wave time-frequency map into a trained target classification model, extract features of the acoustic wave time-frequency map through the target classification model to obtain time-frequency map features; classify the action of the acoustic wave time-frequency map according to the time-frequency map features to obtain a second action category corresponding to the acoustic wave motion characteristics.
[0269] In one embodiment, the above-mentioned living body detection device further includes a model training module, which is configured to obtain a sample acoustic wave time-frequency map and a sample label; the sample acoustic wave time-frequency map is generated based on the sample acoustic wave signal reflected by the sample object from the collected first acoustic wave signal, and the sample label is the action annotation label for the sample object in the sample acoustic wave time-frequency map; input the sample acoustic wave time-frequency map into a classification model to be trained, extract the sample time-frequency map features corresponding to the sample acoustic wave time-frequency map through the classification model to be trained; classify the action according to the sample time-frequency map features to obtain a predicted action category; based on the difference between the predicted action category and the sample label, adjust the parameters of the classification model and continue training until the training condition is met and then end the training to obtain a target classification model.
[0270] In one embodiment, the above-mentioned living body detection device further includes a posture adjustment module, which is configured to obtain a face image corresponding to the detection object; extract features of the face image to obtain face features; determine the face posture of the detection object according to the face features; when the face posture does not meet the posture condition, output posture adjustment information to instruct the detection object to adjust the face posture; the data output module 1302 is further configured to output motion indication information and the first acoustic wave signal when the face posture meets the posture condition.
[0271] In one embodiment, the above-mentioned living body detection device further includes a face recognition module, which is configured to obtain a face image corresponding to the detection object; extract the current face features of the face image; perform face recognition on the face image based on the current face features and the target face features corresponding to the detection object to obtain a face recognition result of the detection object; the above-mentioned living body detection device further includes an identity verification module, which is configured to determine an identity verification result of the detection object according to the face recognition result and the living body detection result.
[0272] For the specific limitations of the living body detection device, reference can be made to the limitations of the living body detection method in the above text, which will not be elaborated here. Each module in the above-mentioned living body detection device can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0273] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structural diagram may be as shown in Figure 14 . The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as motion indication information, first acoustic wave signal data, action videos, second acoustic wave signals, and live detection results. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a live detection method is implemented.
[0274] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structural diagram may be as shown in Figure 15 . The computer device includes a processor, a memory, a communication interface, a display screen, a camera, a speaker, and a microphone connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, a live detection method is implemented. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen and can be used to output motion indication information. The camera of the computer device is used to collect at least one of the face images and action videos of the detection object. The speaker of the computer device is used to output the first acoustic wave signal. The microphone of the computer device is used to collect the second acoustic wave signal reflected by the first acoustic wave signal through the detection object.
[0275] Those skilled in the art can understand that Figure 14 and Figure 15 the structures shown in are only block diagrams of some structures related to the solution of this application and do not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0276] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.
[0277] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0278] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.
[0279] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0280] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0281] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for live detection, characterized in that, The method includes: Outputting motion indication information and a first sound wave signal; the first sound wave signal is directed at a detection object moving according to the motion indication information; the first sound wave signal is generated by randomizing the carrier frequency of a preset audio signal; Obtaining an action video collected for the moving detection object, and positioning an action interval corresponding to the detection object according to an action amplitude feature in the action video; Obtaining a second sound wave signal reflected by the detection object from the first sound wave signal, and extracting a sound wave motion feature from a target motion signal in the second sound wave signal; Synchronously aligning the action amplitude feature and the sound wave motion feature in time, and cutting out a sound wave motion feature corresponding to the action interval from the synchronously aligned sound wave motion feature; When a first action category corresponding to the action amplitude feature is consistent with a second action category corresponding to the sound wave motion feature corresponding to the action interval, and the first action category and the second action category are consistent with an indicated action category in the motion indication information, it is determined that the living body detection result of the detection object passes.
2. The method according to claim 1, wherein The positioning of the action interval corresponding to the detection object according to the action amplitude feature in the action video includes: Performing action detection on the action video to obtain an action amplitude feature in the action video; Determining an action start time and an action end time of the detection object according to the action amplitude feature; Positioning the action interval corresponding to the detection object according to the action start time and the action end time.
3. The method according to claim 2, wherein The performing action detection on the action video to obtain an action amplitude feature in the action video includes: Performing key point detection on each video frame in the action video respectively to obtain an action key point and an action area corresponding to each video frame; Performing action detection according to the action key point and the action area corresponding to each video frame respectively to obtain an action feature corresponding to each video frame; Obtaining an action amplitude feature corresponding to the action video according to the time sequence of the action video and the action feature corresponding to each video frame.
4. The method according to claim 1, wherein The extracting a sound wave motion feature from a target motion signal in the second sound wave signal includes: Performing signal demodulation on the second sound wave signal to obtain a component signal of the second sound wave signal; Performing interference cancellation on the component signal to obtain a target motion signal in the second sound wave signal; Performing feature extraction on the target motion signal to obtain a sound wave motion feature corresponding to the target motion signal.
5. The method according to claim 4, characterized in that The performing interference cancellation on the component signal to obtain a target motion signal in the second sound wave signal includes: Performing dynamic interference cancellation on the component signal based on a preset interception frequency to obtain a component signal after dynamic interference cancellation; Extracting a static component from the component signal after dynamic interference cancellation, and performing static interference cancellation on the static component to obtain a target motion signal in the second sound wave signal.
6. The method according to claim 1, wherein The synchronously aligning the action amplitude feature and the sound wave motion feature in time, and cutting out a sound wave motion feature corresponding to the action interval from the synchronously aligned sound wave motion feature includes: Synchronize and align the action amplitude feature with the acoustic wave motion feature according to the time sequence of the action amplitude feature and the time sequence of the acoustic wave motion feature; According to the action start time and action end time corresponding to the action interval, cut the synchronized and aligned acoustic wave motion feature to obtain the acoustic wave motion feature corresponding to the action interval.
7. The method according to claim 1, wherein The method further includes: Perform action detection on the action amplitude feature to obtain a first action category corresponding to the action amplitude feature; Perform action detection on the acoustic wave motion feature corresponding to the action interval to obtain a second action category corresponding to the acoustic wave motion feature corresponding to the action interval.
8. The method according to claim 7, wherein The performing action detection on the acoustic wave motion feature corresponding to the action interval to obtain a second action category corresponding to the acoustic wave motion feature corresponding to the action interval includes: Generate a corresponding acoustic wave time-frequency map according to the acoustic wave motion feature corresponding to the action interval; Input the acoustic wave time-frequency map into a trained target classification model, and extract time-frequency map features from the acoustic wave time-frequency map through the target classification model; Perform action classification on the acoustic wave time-frequency map according to the time-frequency map features to obtain a second action category corresponding to the acoustic wave motion feature.
9. The method according to claim 8, characterized in that The target classification model is obtained through a training step, and the training step includes: Obtain a sample acoustic wave time-frequency map and a sample label; the sample acoustic wave time-frequency map is generated based on a sample acoustic wave signal reflected by the first acoustic wave signal collected by a sample object, and the sample label is an action annotation label for the sample object in the sample acoustic wave time-frequency map; Input the sample acoustic wave time-frequency map into a classification model to be trained, and extract sample time-frequency map features corresponding to the sample acoustic wave time-frequency map through the classification model to be trained; Perform action classification according to the sample time-frequency map features to obtain a predicted action category; Based on the difference between the predicted action category and the sample label, adjust the parameters of the classification model and continue training until the training condition is met, and then end the training to obtain the target classification model.
10. The method according to any one of claims 1 to 9, characterized in that, Before outputting the motion indication information and the first acoustic wave signal, the method further includes: Obtain a face image corresponding to the detection object; Extract features from the face image to obtain face features; Determine the face pose of the detection object according to the face features; When the face pose does not meet the pose condition, output pose adjustment information to instruct the detection object to adjust the face pose; The outputting the motion indication information and the first acoustic wave signal includes: When the face pose meets the pose condition, output the motion indication information and the first acoustic wave signal.
11. The method according to any one of claims 1 to 9, characterized in that, The method further includes: Obtain a face image corresponding to the detection object; Extract the current face features of the face image; Perform face recognition on the face image based on the current face features and the target face features corresponding to the detection object to obtain a face recognition result of the detection object; Determine an identity verification result of the detection object according to the face recognition result and the live detection result.
12. A living body detection device, characterized in that, The device includes: A data output module for outputting motion indication information and a first acoustic wave signal; the first acoustic wave signal is directed at a detection object moving according to the motion indication information; the first acoustic wave signal is generated by randomizing the carrier frequency of a preset audio signal. An action video processing module for acquiring an action video collected for the moving detection object, and positioning an action interval corresponding to the detection object according to action amplitude features in the action video. An acoustic wave signal processing module for acquiring a second acoustic wave signal reflected by the detection object from the first acoustic wave signal, and extracting acoustic wave motion features from target motion signals in the second acoustic wave signal. A living body detection module for synchronously aligning the action amplitude features and the acoustic wave motion features in time, and cutting out acoustic wave motion features corresponding to the action interval from the synchronously aligned acoustic wave motion features; when a first action category corresponding to the action amplitude features is consistent with a second action category corresponding to the acoustic wave motion features corresponding to the action interval, and the first action category and the second action category are consistent with an indicated action category in the motion indication information, it is determined that the living body detection result of the detection object passes.
13. The in-vivo detection device according to claim 12, characterized in that, The action video processing module is further configured to perform action detection on the action video to obtain action amplitude features in the action video; determine an action start time and an action end time of the detection object according to the action amplitude features; and position an action interval corresponding to the detection object according to the action start time and the action end time.
14. The in-vivo detection device according to claim 13, characterized in that, The action video processing module is further configured to perform key point detection on each video frame in the action video to obtain action key points and action regions corresponding to each video frame; perform action detection according to the action key points and action regions corresponding to each video frame to respectively obtain action features corresponding to each video frame; and obtain action amplitude features corresponding to the action video according to the time sequence of the action video and the action features corresponding to each video frame.
15. The living body detection device according to claim 12, characterized in that, The acoustic wave signal processing module is further configured to demodulate the second acoustic wave signal to obtain component signals of the second acoustic wave signal; eliminate interference from the component signals to obtain target motion signals in the second acoustic wave signal; and extract features from the target motion signals to obtain acoustic wave motion features corresponding to the target motion signals.
16. The living body detection device according to claim 15, characterized in that, The acoustic wave signal processing module is further configured to perform dynamic interference cancellation on the component signals based on a preset interception frequency to obtain component signals after dynamic interference cancellation; extract static components from the component signals after dynamic interference cancellation, and perform static interference cancellation on the static components to obtain target motion signals in the second acoustic wave signal.
17. The living body detection device according to claim 12, wherein, The living body detection module is further configured to synchronously align the action amplitude features and the acoustic wave motion features according to the time sequences of the action amplitude features and the acoustic wave motion features; and cut the synchronously aligned acoustic wave motion features according to the action start time and the action end time corresponding to the action interval to obtain acoustic wave motion features corresponding to the action interval.
18. The living body detection device according to claim 12, wherein, The living body detection module is further configured to perform action detection on the action amplitude feature to obtain a first action category corresponding to the action amplitude feature; and perform action detection on the acoustic wave motion feature corresponding to the action interval to obtain a second action category corresponding to the acoustic wave motion feature corresponding to the action interval.
19. The living body detection device according to claim 18, characterized in that, The living body detection module is further configured to generate a corresponding acoustic wave time-frequency diagram according to the acoustic wave motion feature corresponding to the action interval; input the acoustic wave time-frequency diagram into a trained target classification model, extract features of the acoustic wave time-frequency diagram through the target classification model to obtain time-frequency diagram features; and perform action classification on the acoustic wave time-frequency diagram according to the time-frequency diagram features to obtain the second action category corresponding to the acoustic wave motion feature.
20. The living body detection device according to claim 19, wherein, The device further includes a model training module configured to obtain sample acoustic wave time-frequency diagrams and sample labels; the sample acoustic wave time-frequency diagrams are generated based on sample acoustic wave signals reflected by a sample object from the collected first acoustic wave signals, and the sample labels are action annotation labels for the sample object in the sample acoustic wave time-frequency diagrams. Input the sample acoustic wave time-frequency diagrams into a classification model to be trained, and extract sample time-frequency diagram features corresponding to the sample acoustic wave time-frequency diagrams through the classification model to be trained. Perform action classification according to the sample time-frequency diagram features to obtain a predicted action category; based on the difference between the predicted action category and the sample labels, adjust the parameters of the classification model and continue training until the training conditions are met, and then end the training to obtain a target classification model.
21. The living body detection device according to any one of claims 12 to 20, wherein The device further includes a posture adjustment module configured to obtain a face image corresponding to the detection object; extract features of the face image to obtain face features; determine the face posture of the detection object according to the face features; when the face posture does not meet the posture conditions, output posture adjustment information to instruct the detection object to adjust the face posture; and the data output module is further configured to output motion indication information and a first acoustic wave signal when the face posture meets the posture conditions.
22. The living body detection device according to any one of claims 12 to 20, characterized in that, The device further includes a face recognition module configured to obtain a face image corresponding to the detection object; extract current face features of the face image; perform face recognition on the face image based on the current face features and target face features corresponding to the detection object to obtain a face recognition result of the detection object; the device further includes an identity verification module configured to determine an identity verification result of the detection object according to the face recognition result and the living body detection result.
23. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.
24. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 11.
25. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for human face living body identification
CN106897658A
Living body detection method, device and equipment and computer readable storage medium
CN110119719A
Living body detection method, device and equipment and storage medium
CN111368811A
Identity verification method and device, computer equipment and storage medium
CN111563244A
KR20200084451A