Liveness detection method, device, equipment and system

By combining voiceprint recognition and lip recognition technology, using voice and video information for live detection, the problem of poor effectiveness in defending against high-definition screens and high-precision headmode attacks has been solved, and higher live detection accuracy and safety has been achieved.

CN114696988BActive Publication Date: 2025-06-06ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210241977.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-11
Publication Date
2025-06-06
Estimated Expiration
2042-03-11

AI Technical Summary

Technical Problem

The existing live detection system has limited effectiveness in defending against high-definition screen attacks and high-precision head-mode attacks. Single-frame image input leads to insufficient information, making it difficult to accurately judge live.

Method used

Information in both voice and video modes is adopted, voiceprint data is extracted from voice information, and lip recognition technology is used to extract verification passwords from video information, and live detection is performed by combining voiceprint recognition and lip recognition technology.

Benefits of technology

It significantly improves the accuracy of live detection, can effectively prevent high-definition screen attacks and high-precision head-mode attacks, and enhances the security of identity authentication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114696988B_ABST
    Figure CN114696988B_ABST
Patent Text Reader

Abstract

This specification provides a liveness detection method, device, equipment and system, which utilizes information in both voice and video modes to extract voiceprint data from voice information, and utilizes lip reading recognition technology to extract a verification password to be detected from a video image. At the same time, voiceprint recognition and lip reading recognition technology are used to jointly determine whether the object to be detected is alive, which can greatly improve the accuracy of liveness detection. At the same time, it can also play a good preventive role against attacks that are difficult to handle with single-frame images, such as high-definition screen attacks and high-precision head model attacks, thereby improving the accuracy of liveness detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and more particularly to a method, device, equipment and system for detecting a living body. Background Art

[0002] With the development of computer Internet technology, more and more products use biometric identity authentication technologies such as face recognition and pupil recognition. However, this biometric technology also has security risks. Among them, "live attack" is a security risk. For example, attackers use photos, screen displays, masks and other means to achieve the purpose of impersonating others.

[0003] Currently, most liveness detection systems use single-frame images (RGB images, IR images or depth images) as input for liveness assessment. Since a single-frame image contains limited information, it is usually difficult to defend against high-definition screen attacks and high-precision head model attacks. Summary of the invention

[0004] The purpose of the embodiments of this specification is to provide a method, device, equipment and system for liveness detection, which improves the accuracy of liveness detection.

[0005] In a first aspect, an embodiment of the present specification provides a method for detecting a living body, the method comprising:

[0006] Receiving video information to be detected, wherein the video information to be detected includes voice information of the object to be detected reading a random verification password displayed by the client;

[0007] Extracting the voiceprint data to be detected from the voice information, and extracting the verification password to be detected from the video information to be detected by using lip reading recognition technology;

[0008] The voiceprint data to be detected is compared with the voiceprint data of the pre-stored target object, and the verification password to be detected is compared with the random verification password. If the voiceprint data to be detected is the same as the voiceprint data of the target object, and the verification password to be detected is the same as the random verification password, it is determined that the liveness detection of the object to be detected has passed.

[0009] In a second aspect, this specification provides a method for detecting a living body, the method comprising:

[0010] After receiving a liveness detection request, a random verification password is generated and displayed;

[0011] Collecting video information of the object to be detected reading the random verification password, wherein the video information to be detected includes voice information of the object to be detected reading the random verification password displayed by the client;

[0012] The video information to be detected, the random verification password and the account information corresponding to the liveness detection request are sent to the server, so that the server uses lip reading recognition technology to extract the verification password to be detected from the video information to be detected, and performs liveness detection on the object to be detected in combination with the voiceprint data to be detected in the video information to be detected.

[0013] In a third aspect, this specification provides a living body detection device, including:

[0014] A video information receiving module, used for receiving the video information to be detected, wherein the video information to be detected includes the voice information of the object to be detected reading the random verification password displayed by the client;

[0015] A data extraction module, used to extract the voiceprint data to be detected from the voice information, and to extract the verification password to be detected from the video information to be detected by using lip reading recognition technology;

[0016] The data comparison module is used to compare the voiceprint data to be detected with the voiceprint data of the pre-stored target object, and to compare the verification password to be detected with the random verification password. If the voiceprint data to be detected is the same as the voiceprint data of the target object, and the verification password to be detected is the same as the random verification password, it is determined that the liveness detection of the object to be detected has passed.

[0017] In a fourth aspect, this specification provides a living body detection device, the device comprising:

[0018] A random password generation module is used to generate and display a random verification password after receiving a liveness detection request;

[0019] A video acquisition module, used for acquiring video information to be detected of the object to be detected reading the random verification password, wherein the video information to be detected includes voice information of the object to be detected reading the random verification password displayed by the client;

[0020] The liveness detection module is used to send the video information to be detected, the random verification password and the account information corresponding to the liveness detection request to the server, so that the server uses lip reading recognition technology to extract the verification password to be detected from the video information to be detected, and performs liveness detection on the object to be detected in combination with the voiceprint data to be detected in the video information to be detected.

[0021] In a fifth aspect, an embodiment of the present specification provides a liveness detection device, comprising at least one processor and a memory for storing processor executable instructions, wherein when the processor executes the instructions, the liveness detection method described in the first aspect or the second method is implemented.

[0022] In a sixth aspect, an embodiment of the present specification provides a liveness detection system, including: a client and a server, wherein the server includes at least one processor and a memory for storing processor executable instructions, and when the processor executes the instructions, the method described in the first aspect is implemented, for performing liveness detection based on video information to be detected collected by the client;

[0023] The client includes at least one processor and a memory for storing processor-executable instructions, and the method described in the second aspect is implemented when the processor executes the instructions.

[0024] The liveness detection method, device, equipment and system provided in this specification utilize information in both voice and video modes to extract voiceprint data from voice information, and utilize lip reading recognition technology to extract the verification password to be detected from the video image. At the same time, voiceprint recognition and lip reading recognition technology are used to jointly determine whether the object to be detected is alive, which can greatly improve the accuracy of liveness detection. At the same time, it can also play a good preventive role against attacks that are difficult to handle with single-frame images, such as high-definition screen attacks and high-precision head model attacks, thereby improving the accuracy of liveness detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0026] Figure 1 It is a flowchart of an embodiment of a liveness detection method provided in an embodiment of this specification;

[0027] Figure 2 This is a schematic diagram of the client interface of liveness detection in a scenario example of this manual;

[0028] Figure 3 It is a schematic diagram of the principle of deploying a liveness detection algorithm in one embodiment of this specification;

[0029] Figure 4 is a flowchart of a liveness detection method in other embodiments of this specification;

[0030] Figure 5 It is a schematic diagram of the module structure of an embodiment of a living body detection device provided in this specification;

[0031] Figure 6 is a schematic diagram of the module structure of another embodiment of the living body detection device provided in this specification;

[0032] Figure 7 It is a hardware structure block diagram of a liveness detection server in one embodiment of this specification. DETAILED DESCRIPTION

[0033] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0034] With the development of computer technology, people pay more and more attention to the security of Internet products. For example, more and more products use human biometrics for identity recognition. During identity recognition, some attackers may use pictures or videos to impersonate users, affecting the accuracy of identity recognition results. Therefore, it is an important task to perform liveness detection to ensure that the user is the real user.

[0035] In some embodiments of this specification, a liveness detection method may be provided, which collects video information of a user reading a random verification password, extracts voiceprint data therein, and uses lip reading recognition technology to extract information read by the user's lip reading from the video information, and combines voiceprint and lip reading technology to perform liveness detection on the user to ensure that it is the user himself, thereby improving the security of identity recognition, account recognition, etc. Among them, liveness detection can be understood as a technique for determining whether a user is a real person in biometrics, rather than a technique for attacking by printing photos, masks, head models, etc.

[0036] Generally, the silent liveness detection method is based on a single image for liveness detection without interaction. This method can effectively intercept some simple liveness attacks such as mobile phone screens, low-resolution screens, and printed photos, but has little interception effect on high-definition screens and high-precision head models. The liveness detection method based on actions such as blinking and shaking the head is based on simple interactive actions for liveness detection, but for simple actions such as blinking and shaking the head, recorded high-definition videos can be easily bypassed, so this type of method has very limited protection effect on high-definition videos.

[0037] Figure 1It is a flow chart of an embodiment of a liveness detection method provided in an embodiment of this specification. Although this specification provides method operation steps or device structures as shown in the following embodiments or drawings, more or fewer operation steps or module units may be included in the method or device based on routine or no creative labor. In the steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure shown in the embodiments or drawings of this specification. When the method or module structure described is applied in an actual device, server or terminal product, it can be executed sequentially or in parallel according to the method or module structure shown in the embodiment or drawings (for example, a parallel processor or multi-threaded processing environment, or even a distributed processing, server cluster implementation environment).

[0038] A specific implementation example is Figure 1 As shown, in one embodiment of the liveness detection method provided in this specification, the method can be applied to terminals such as servers, computers, tablet computers, servers, smart phones, smart wearable devices, vehicle-mounted devices, smart home devices, etc., and the method may include the following steps:

[0039] Step 102: Receive video information to be detected, wherein the video information to be detected includes voice information of the object to be detected reading a random verification password displayed by the client.

[0040] In the specific implementation process, in some application scenarios of identity recognition or account login security audit, it is often necessary to perform liveness detection on the user to confirm that the current user is the user himself, so as to avoid the problem of attackers stealing accounts by means of pictures, videos, etc. Figure 2 This is a schematic diagram of the client interface of a liveness detection scenario in this manual. Figure 2 As shown, in the embodiments of this specification, when a user performs operations such as identity recognition or account login, a random verification password can be randomly generated in the client and displayed to the user, prompting the user to read the currently displayed random verification password. The client can use its own video recording equipment such as a camera to record a video of the user reading the random verification password. The video is the video information to be detected, which may include the voice information of the object to be detected reading the random verification password displayed by the client, and of course, it may also include image information when the object to be detected reads the random verification password. Among them, the random verification password can be a combination of numbers and / or words, which is not specifically limited in the embodiments of this specification. For example, Figure 2 As shown, in one scenario example, the random verification password may be 4583.

[0041] It should be noted that the client can send the recorded video information to be detected to the server, and the server will perform data processing for liveness detection on the video information to be detected, or the client can perform data processing for liveness detection on the recorded video information to be detected by itself, which is not specifically limited in the embodiments of this specification.

[0042] Step 104: extract the voiceprint data to be detected from the voice information, and extract the verification password to be detected from the video information to be detected using lip reading recognition technology.

[0043] In the specific implementation process, the embodiments of this specification mainly adopt voiceprint recognition and lip reading recognition technology, among which voiceprint recognition is a kind of biometric recognition technology, also known as speaker recognition, including speaker identification and speaker confirmation. Voiceprint recognition is to convert sound signals into electrical signals, and then use computers to recognize them. Voiceprint is a sound wave spectrum that carries speech information displayed by electroacoustic instruments. The production of human language is a complex physiological and physical process between the human body's language center and the pronunciation organs. The vocal organs used by people when speaking, such as tongue, teeth, larynx, lungs, and nasal cavity, vary greatly in size and shape, so the voiceprint maps of any two people are different. Lip reading recognition can be understood as using machine vision technology to continuously identify faces from images, determine the person who is speaking, extract the continuous mouth shape change characteristics of this person, identify the pronunciation corresponding to the speaker's mouth shape, and then calculate the most likely natural language sentence based on the recognized pronunciation.

[0044] In the embodiment of the present specification, the voiceprint data to be detected can be extracted from the voice information in the video information to be detected, and the voiceprint data to be detected can be understood as the voiceprint of the sound appearing in the video information to be detected. At the same time, in the embodiment of the present specification, the lip reading recognition technology can also be used to extract the verification password to be detected corresponding to the user's lip shape in the video from the video information to be detected.

[0045] The extraction of voiceprint data and verification passwords to be detected can utilize intelligent learning algorithms, such as pre-training voiceprint recognition models and lip reading recognition models, extracting features from the voice information and image information in the video information to be detected, and then obtaining the corresponding voiceprint data and verification passwords to be detected. Of course, other methods can also be used to extract voiceprint data and verification passwords to be detected, which are not specifically limited in the embodiments of this specification.

[0046] Step 106: compare the voiceprint data to be detected with the voiceprint data of the target object stored in advance, and compare the verification password to be detected with the random verification password. If the voiceprint data to be detected is the same as the voiceprint data of the target object, and the verification password to be detected is the same as the random verification password, it is determined that the liveness detection of the object to be detected has passed.

[0047] In the specific implementation process, after the voiceprint data and the verification password to be detected are extracted, liveness detection can be performed. The extracted voiceprint data to be detected can be compared with the voiceprint data of the target object stored in advance, where the target object can be understood as the user corresponding to the account currently performing liveness detection, such as: if the user logs in to account A, it is necessary to perform liveness detection on the logged-in user, then the user corresponding to account A is the target object. Another example: when the user uses a smart door lock to open the door, it is necessary to perform liveness detection on the user currently opening the door, then the identifier of the smart door lock can be understood as an account, and the user bound to the smart door lock is the target object. The voice information of the target object can be pre-recorded, the voiceprint data of the target object can be extracted, and saved for use in subsequent liveness detection. The voiceprint data of the target object can be stored in the client and / or the server, and the embodiments of this specification are not specifically limited. In addition, the extracted verification password to be detected can also be compared with the random verification password displayed in the client to detect whether the password read by the user in the recorded video information to be detected is consistent with the password displayed in the client. If the voiceprint data to be detected is the same as the voiceprint data of the target object, and the verification password to be detected is the same as the random verification password displayed by the client, it can be considered that the liveness detection of the target object has passed. Based on the liveness detection results, subsequent business processes can be carried out, such as identity recognition, account login, etc.

[0048] Generally, the attacker only has a static photo of the user. If the attacker uses special technology to generate a video of the user reading a number, and the attacker reads a random password himself, if lip reading recognition and language recognition technology are combined for liveness detection, the verification password extracted from the sound and lip reading can be detected, and the liveness detection can also pass, but this is not the user himself, which affects the accuracy of the liveness detection result. And, generally speaking, the probability that the attacker can get a video of the user reading with sound, and the number read is exactly the same as the randomly generated verification password is very small. In the embodiment of this specification, a random verification password is used, and voiceprint recognition and lip reading recognition technology are combined, which can not only ensure that the mouth shape of the user to be detected corresponds to the random verification password, but also ensure that it is the password read by the user himself, which improves the accuracy of liveness detection.

[0049] In addition, the comparison process of the voiceprint data to be detected and the verification password to be detected can be performed in the client or in the server, and the embodiments of this specification do not make specific limitations. For example: the client can send the recorded video information to be detected to the server, and after the server extracts the voiceprint data to be detected and the verification password to be detected, the voiceprint data to be detected and the verification password to be detected can be returned to the client, and the client compares the voiceprint data to be detected with the voiceprint data of the target object stored in the client, and compares the verification password to be detected with the random verification password displayed in the client. Alternatively, when the client sends the video information to be detected, the account information to be detected for the current liveness detection and the random verification password currently displayed can be sent to the server. After extracting the voiceprint data to be detected, the server obtains the voiceprint data of the target object based on the account information to be detected, and compares the voiceprint data to be detected with the voiceprint data of the target object stored in advance. After extracting the verification password to be detected, the verification password to be detected can be compared with the random verification password sent by the client.

[0050] The liveness detection method provided in the embodiments of this specification utilizes information from both voice and video modes to extract voiceprint data from voice information, and utilizes lip reading recognition technology to extract a verification password to be detected from video images. At the same time, voiceprint recognition and lip reading recognition technology are used to jointly determine whether the object to be detected is alive. This can greatly improve the accuracy of liveness detection, and can also play a good preventive role against attacks that are difficult to handle with single-frame images, such as high-definition screen attacks and high-precision head model attacks, thereby improving the accuracy of liveness detection.

[0051] In some embodiments of this specification, the extracting of the verification password to be detected from the video information to be detected by using lip reading recognition technology includes:

[0052] Input the silent video information in the video information to be detected into a pre-established lip reading recognition model, and use the lip reading recognition model to extract the verification password to be detected from the silent video information; wherein the lip reading recognition model is obtained by model training and construction based on lip reading training sample data, and the lip reading training sample data of the lip reading recognition model includes: real training sample data and constructed training sample data;

[0053] The method for acquiring the real training sample data includes: collecting video data of different users reading different random verification passwords;

[0054] The method for acquiring constructed training sample data includes: utilizing video data of a user reading a specified random verification password, constructing video data of different users reading different random verification passwords, and obtaining the constructed training sample data.

[0055] In the specific implementation process, a lip reading recognition model can be pre-trained and constructed, such as: collecting lip reading training sample data for model training, and constructing an intelligent learning model capable of lip reading recognition. Among them, the lip reading training sample data can be understood as a video or image of a user reading a verification password, and the features of the lip reading training sample data are extracted and learned to obtain a lip reading recognition model. The specific algorithm used by the lip reading recognition model is not specifically limited in the embodiments of this specification, such as: a neural network model, a random forest model, etc. can be used. In some embodiments of this specification, STCNN (Spatial-Temporal Convolutional Neural Network) and GRU (Gated Recurrent Unit) can be selected. When the verification password to be detected is extracted from the video information to be detected by using lip reading recognition technology, the silent video information in the video information to be detected, that is, the image information that does not include voice information, can be input into the pre-established lip reading recognition model, and the lip reading recognition model is used to extract features of the changes in the user's mouth shape in the image information, and the meaning represented by the user's mouth shape and the verification password to be detected are identified. The lip reading recognition model is used to extract features from silent video information, obtain the meaning of the lip shape representation in the image information, and then extract the verification password to be detected, providing an accurate data basis for liveness detection.

[0056] The training of lip reading recognition models generally requires a large amount of sample data. However, the real training data may not be supplemented in a short time due to the insufficient number of training data sets, and the real data or artificially collected data cannot completely and evenly cover the representation space of the random verification password, resulting in extremely uneven training samples. Some random verification passwords have no valid training samples. Based on this, the lip reading training sample data in the embodiments of this specification may include: real training sample data and generated training sample data, wherein the real training sample data can be understood as the real video data of the user reading the random verification password, and the constructed training sample data can be understood as the sample data generated and constructed based on a certain technology.

[0057] Among them, video data of different users reading different random verification passwords can be collected as real training sample data. For example, a certain number of users can be selected, and these users can read different random verification passwords respectively, and video data of users reading random verification passwords can be recorded. The video data can be lip video data or whole face video data including lips, which is not specifically limited in the embodiments of this specification. Among them, the random verification password can be randomly generated according to actual needs, such as: a verification password library can be constructed, and verification passwords can be randomly extracted from the library as random verification passwords. The random verification password can be a combination of words or a combination of numbers or a combination of words and numbers, which is not specifically limited in the embodiments of this specification.

[0058] For constructing training sample data, video data of users reading designated random verification passwords can be collected to construct and generate video data of different users reading different random verification passwords. For example, video data of users reading designated random verification passwords can be split, synthesized, etc. to generate video data of different users reading different random verification passwords.

[0059] In some embodiments of the present specification, the video data of a user reading a specified random verification password is used to construct video data of different users reading different random verification passwords, and the constructed training sample data is obtained, including:

[0060] The video data of a designated user reading a designated random verification password is used to drive silent facial images of different users, and generated video data of different users reading the designated random verification password is generated, and the generated video data is used as the constructed training sample data.

[0061] In the specific implementation process, for constructing training sample data, the video data of the user reading the specified random verification password can be used to drive the silent facial images of different users, thereby generating the generated video data of different users reading the specified random verification password. For example, user A, i.e., the specified user, can be selected to read different random verification passwords, and a video can be recorded, such as recording the video data of user A reading three sets of random verification passwords 1234, 1345, and 3579, and using the recorded video to drive the silent facial images of different users. For example, using the video data of user A reading the random verification password 1234 to drive the silent facial image of user B, the generated video data of user B reading the random verification password 1234 can be obtained. Similarly, using the video data of user A reading the random verification passwords 1345 and 3579 to drive the silent facial image of user B, the generated video data of user B reading the random verification passwords 1345 and 3579 can be obtained respectively. The video data of user A or other designated users reading different random verification passwords can be recorded, or the real training sample data obtained above can be used to drive the silent facial images of different users, that is, the generated video data of different users reading different random verification passwords can be obtained to construct the training sample data. The number of designated users is not specifically limited.

[0062] Among them, the image generation algorithm that generates video data by using video-driven silent images can be selected according to actual needs, and the embodiments of this specification do not make specific limitations.

[0063] It can be seen that the generated video data is not actually real data, but a small amount of real data can be used to obtain a large amount of sample data, and the lip reading recognition model can be trained with the real training sample data, which makes up for the problem of insufficient sample data. At the same time, it can ensure the accuracy of model training, thereby laying an accurate data foundation for subsequent liveness detection.

[0064] In some embodiments of this specification, the random verification password includes multiple sub-passwords, each of which is a single character or a single number, and the method of using video data of a designated user reading a designated random verification password to drive silent facial images of different users to generate generated video data of different users reading the designated random verification password also includes:

[0065] Using the video data of the designated user reading the sub-password in sequence to drive the silent facial images of different users, the generated video data of different users reading the sub-password is generated;

[0066] The video data generated when different users read the sub-passwords are divided into one video for each sub-password read, so as to obtain sub-video data generated when different users read different sub-passwords;

[0067] The generated sub-video data obtained by reading different sub-passwords by the same user are randomly combined and synthesized, and the obtained synthesized video data is used as the generated video data.

[0068] In the specific implementation process, referring to the records of the above embodiments, the random verification password in the embodiment of this specification may include multiple sub-passwords, each sub-password may be a single word or a single number, and the word may be Chinese or English, that is, the random verification password in the embodiment of this specification may be a paragraph of words or a string of numbers or a combination of words and numbers, which may be set according to actual needs, and the embodiment of this specification does not make specific restrictions. In the embodiment of this specification, when constructing and generating training sample data, video data of a designated user reading each sub-password in turn may also be recorded. The combination of all sub-passwords may be used as a designated random verification password, and video data of a designated user reading all sub-password combinations may be recorded. For example, if a random verification password is generated using a digital combination of 0-9, video data of a designated user reading 10 numbers 0-9 may be recorded, and this video data may be used to drive silent facial images of other users, so that video data of different users reading sub-passwords such as 10 numbers 0-9 may be obtained. The generated video data of the same user reading a sub-password is divided into a video for each sub-password read, and the generated sub-video data of the user reading different sub-passwords is obtained. Then, according to the random verification password setting rules, the sub-video data generated by the same user reading different sub-passwords are randomly combined and synthesized to obtain the synthesized video data, that is, the generated video data, and then the training sample data is obtained.

[0069] For example, if the random verification password is a combination of 4 digits, the video data of the designated user A reading the 10 digits 0-9 can be recorded, and the silent facial image of the user B can be used to drive the video data of the user B reading the sub-passwords such as the 10 digits 0-9. The video data of the user B reading the 10 digits 0-9 is segmented, and a sub-video data is segmented for each sub-password read, and 10 generated sub-video data of the user B reading the 10 digits 0-9 are obtained. The generated sub-video data of the user B reading the 10 sub-passwords are randomly combined, and every 4 generated sub-video data are combined to form a video data, which is the generated video data. Among them, when the generated sub-video data are randomly combined and synthesized, the same generated sub-video data can be reused, that is, a synthesized video data can include the same generated sub-video data, such as: the first two sub-passwords in the random verification password 1123 are the same, so the generated sub-video data of the two users B reading 1 can be synthesized with the generated sub-video data of reading 2 and 3.

[0070] In the embodiments of this specification, by specifying the video data of the user reading the sub-password in sequence to drive the silent facial images of other users, it is possible to obtain the generated video data of other users reading each sub-password in sequence, and to obtain the generated video data covering all sub-passwords. Then, by segmenting the obtained generated video data, it is possible to obtain the generated sub-video data of the user reading each sub-password, and then by randomly combining and synthesizing each generated sub-video data, it is possible to obtain the generated video data of the user reading different random verification passwords. By means of video driving, video segmentation, and video synthesis, it is possible to obtain video data that almost covers all random verification passwords, thereby improving the uniformity of the training data, and combining with the real training sample data to improve the accuracy of the lip reading recognition model training, thereby laying an accurate data foundation for subsequent liveness detection.

[0071] In some other embodiments of the present specification, the random verification password includes multiple sub-passwords, and the sub-password is a single character or a single number. The video data of a user reading a specified random verification password is used to construct video data of different users reading different random verification passwords, and the constructed training sample data is obtained, including:

[0072] Collect video data of different users reading sub-passwords in sequence, divide the video data of different users reading sub-passwords into one video according to each sub-password read, and obtain video data of different users reading different sub-passwords;

[0073] The video data of the same user reading different sub-passwords in sequence are randomly combined and synthesized, and the obtained synthesized video data is used as the constructed training sample data.

[0074] In the specific implementation process, the lip reading recognition model can be trained by combining real training sample data and synthetic video data, wherein the synthetic video data can be understood as video data synthesized after splitting the real video data. Among them, the method for obtaining the real training sample data can refer to the records of the above embodiment, which will not be repeated here. Referring to the records of the above embodiment, the random verification password in the embodiment of this specification can include multiple sub-passwords, each sub-password can be a single word or a single number, and the word can be Chinese or English or other languages, that is, the random verification password in the embodiment of this specification can be a paragraph of text or a string of numbers or a combination of words and numbers. The combination of all sub-passwords can be used as a specified random verification password. The embodiment of this specification can collect and record video data of different users reading all sub-passwords in turn, that is, record video data of different users reading the specified random verification password. Then the video data of the same user reading the sub-password in turn is segmented to obtain video data of each user reading different sub-passwords. According to the setting rules of the random verification password, the video data of the same user reading different sub-passwords is randomly combined and synthesized to obtain synthetic video data.

[0075] For example, if the random verification password is a combination of 4 digits, the video data of different users reading the 10 digits 0-9 can be recorded, and the video data of different users reading the 10 digits 0-9 is divided according to each sub-password read, and each sub-password is divided into a video data, and 10 video data of different users reading the 10 digits 0-9 are obtained. The video data of the same user reading 10 sub-passwords are randomly combined, and every 4 video data are combined into one video data, which is the composite video data. Referring to the records of the above embodiment, when the video data are randomly combined and synthesized, the same video data can be reused, that is, the same video data can be included in one composite video data.

[0076] It can be seen that synthetic video data is also a kind of generated data. Although it is not the real video data of the verification password, the video data of each sub-password is real and more accurate than the generated training sample data. In addition, the video data of the user reading each sub-password is segmented and combined, which can almost cover all random verification passwords and improve the uniformity of the samples. Combining with real training sample data can not only ensure the authenticity and accuracy of the training sample data, but also increase the number of samples, thereby improving the accuracy of the lip reading recognition model, providing a more accurate data basis for subsequent liveness detection.

[0077] Figure 3 FIG. 1 is a schematic diagram showing the principle of deploying a liveness detection algorithm in an embodiment of this specification. Figure 3As shown, the embodiments of this specification mainly use lip reading recognition and voiceprint recognition technology. For the training of the lip reading recognition model, the sample data for model training can be obtained by combining real training sample data, generated video data and synthetic video data. Among them, the generated video data and synthetic video data are the constructed training sample data in the above embodiments. Figure 3 As shown, taking the random verification password as a 4-digit combination of 0-9 as an example, the training sample data may include real training sample data, generated video data, and synthesized video data. For the generated video data, an image generation algorithm and other methods may be used to generate a reading video of ten digits 0-9 using a single silent face image. Since a single silent face image is easy to obtain, a generated video data of reading different specified random verification passwords of more different people may be generated by the generation method. In addition, in order to solve the problem that the real data and the artificially collected data cannot completely and evenly cover the representation space of the 4-digit random number, the embodiment of this specification may use a data synthesis method to generate a reading video of a 4-digit random number. For the real data, the silent frame is suppressed by the voice signal, and the single reading video is segmented to obtain the reading video of each digit. It can be seen that only a 0-9 reading video needs to be collected for each person, and then segmented using the above method to obtain a single reading video of a real user. Since the reading video of a single digit is obtained, these videos can be used to synthesize the reading video of any 4-digit number.

[0078] By using the method described in the above embodiment, the difficulty and amount of data collection can be greatly simplified, and a large amount of training data can be synthesized by combining the above-mentioned single reading video. By generating video data and synthesizing video data, the problems of insufficient training data and uneven coverage can be compensated, and by combining real training sample data, the accuracy of sample data can be improved, thereby improving the accuracy of the lip reading recognition model.

[0079] for Figure 3 The training and construction method of the voiceprint recognition model in the embodiment can select a suitable method according to actual needs, and the embodiments of this specification do not make specific limitations.

[0080] Referring to the above embodiments, the liveness detection in the embodiments of this specification can be completed in the client or in the server. In some embodiments of this specification, the client can collect the video information to be detected, and the powerful computing power of the server can be used to perform data processing and calculation for liveness detection. Figure 4 is a flow chart of a liveness detection method in other embodiments of this specification, such as Figure 4 As shown, the method can be applied in a client such as a smart phone, a smart wearable device, a tablet computer and the like, and the method may include the following steps:

[0081] Step 402: After receiving the liveness detection request, generate and display a random verification password.

[0082] In the specific implementation process, Figure 2 As shown, in some embodiments, when a user logs into an account or performs identity recognition, a liveness detection request from the client is triggered. At this time, the client can generate a random verification password. The specific form of the random verification password can refer to the description of the above embodiment and will not be repeated here. For example, a verification password database can be configured in the client. When a random verification password needs to be generated, a specified number of passwords can be randomly selected from the database according to the configuration rules of the random verification password to form a random verification password. Alternatively, a random verification password is generated using some algorithms, which are not specifically limited in the embodiments of this specification. Figure 2 As shown, the client can display the generated random verification password on the screen of the client and prompt the user to read the password displayed on the screen.

[0083] Step 404: Collect the video information of the object to be detected reading the random verification password, wherein the video information to be detected includes the voice information of the object to be detected reading the random verification password displayed by the client.

[0084] In a specific implementation process, when the user reads the random verification password displayed on the client according to the prompt on the client screen, the client can record a video of the user reading the random verification password, i.e., the video information to be detected. As described in the above embodiment, the video information to be detected includes the voice information and image information of the object to be detected, i.e., the user to be detected, reading the random verification password displayed on the client.

[0085] Step 406: Send the video information to be detected, the random verification password, and the account information corresponding to the liveness detection request to the server, so that the server uses lip reading recognition technology to extract the verification password to be detected from the video information to be detected, and performs liveness detection on the object to be detected in combination with the voiceprint data to be detected in the video information to be detected.

[0086] In the specific implementation process, after the client collects the video information of the object to be detected reading the random verification password, the collected video information can be sent to the server, and the server performs voiceprint recognition and lip reading recognition on the video information to be detected collected by the client, extracts the voiceprint data to be detected and the verification password to be detected, and then compares the voiceprint data to be detected and the verification password to be detected with the voiceprint data and random verification password of the target user to perform liveness detection on the object to be detected. The process of liveness detection by the server and the creation of the lip reading recognition model are described in the above embodiment and will not be repeated here.

[0087] The liveness detection method provided in the embodiments of this specification is a joint liveness detection method based on voiceprint recognition and lip reading recognition. It collects the random verification password displayed on the client screen read by the user, and uses voiceprint recognition and lip reading recognition technology to identify whether the user's reading is consistent with the displayed random verification password, and identify whether the voiceprint is consistent with the target voiceprint, so as to determine whether it is a real person. At the architectural level, this method uses information from both voice and video modes, and uses voiceprint recognition and lip reading recognition technology to jointly determine whether the object to be detected is alive. This can greatly improve the accuracy of the liveness detection algorithm, and at the same time, it can also play a good preventive role against attacks that are difficult to handle with single-frame images, such as high-definition screen attacks and high-precision head model attacks.

[0088] In this specification, each embodiment of the above method is described in a progressive manner, and the same or similar parts between the embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments. For relevant parts, refer to the partial description of the method embodiment.

[0089] Based on the liveness detection method described above, one or more embodiments of this specification also provide a system for liveness detection. The system may include a system (including a distributed system), software (application), module, component, server, client, etc. using the method described in the embodiment of this specification and a device in combination with necessary implementation hardware. Based on the same innovative concept, the device in one or more embodiments provided in the embodiment of this specification is as described in the following embodiments. Since the implementation scheme and method for solving the problem of the device are similar, the implementation of the specific device in the embodiment of this specification can refer to the implementation of the aforementioned method, and the repetitions will not be repeated. As used below, the term "unit" or "module" can implement a combination of software and / or hardware for predetermined functions. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.

[0090] Specifically, Figure 5 is a schematic diagram of the module structure of an embodiment of a liveness detection device provided in this specification. The device can be applied to a server, such as Figure 5 As shown, the living body detection device provided in this specification may include:

[0091] The video information receiving module 51 is used to receive the video information to be detected, wherein the video information to be detected includes the voice information of the detected object reading the random verification password displayed by the client;

[0092] A data extraction module 52 is used to extract the voiceprint data to be detected from the voice information, and to extract the verification password to be detected from the video information to be detected by using lip reading recognition technology;

[0093] The data comparison module 53 is used to compare the voiceprint data to be detected with the voiceprint data of the target object stored in advance, and to compare the verification password to be detected with the random verification password. If the voiceprint data to be detected is the same as the voiceprint data of the target object, and the verification password to be detected is the same as the random verification password, it is determined that the liveness detection of the object to be detected has passed.

[0094] In some embodiments of this specification, the data extraction module is specifically used to:

[0095] Input the silent video information in the video information to be detected into a pre-established lip reading recognition model, and use the lip reading recognition model to extract the verification password to be detected from the silent video information; wherein the lip reading recognition model is obtained by model training and construction based on lip reading training sample data, and the lip reading training sample data of the lip reading recognition model includes: real training sample data and constructed training sample data;

[0096] The device also includes a training sample data generating module for:

[0097] Collect video data of different users reading different random verification passwords to obtain the real training sample data;

[0098] The video data of a user reading a specified random verification password is used to construct video data of different users reading different random verification passwords, thereby obtaining the constructed training sample data.

[0099] In some embodiments of this specification, the training sample data generation module is specifically used to:

[0100] The video data of a designated user reading a designated random verification password is used to drive silent facial images of different users, and generated video data of different users reading the designated random verification password is constructed, and the generated video data is used as the constructed training sample data.

[0101] In some embodiments of this specification, the random verification password includes multiple sub-passwords, each of which is a single character or a single number, and the training sample data generation module is further used to:

[0102] Using the video data of the designated user reading the sub-password to drive the silent facial images of different users, the generated video data of the different users reading the sub-password are generated;

[0103] The video data generated when different users read the sub-passwords are divided into one video for each sub-password read, so as to obtain the video data generated when different users read different sub-passwords;

[0104] The generated video data obtained by reading different sub-passwords by the same user are randomly combined and synthesized, and the obtained synthesized video data is used as the generated video data.

[0105] In the embodiment of this specification, the random verification password includes multiple sub-passwords, and the sub-password is a single character or a single number; the training sample data generation module is specifically used to:

[0106] The video data of different users reading sub-passwords are collected, and the video data of different users reading sub-passwords are segmented to obtain video data of different users reading different sub-passwords, and the video data of the same user reading different sub-passwords are randomly combined and synthesized to obtain the synthesized video data as the constructed training sample data.

[0107] Figure 6 FIG. 1 is a schematic diagram of a module structure of another embodiment of a living body detection device provided in this specification. Figure 6 As shown, the device can be applied to the client. The liveness detection device provided in this specification may include:

[0108] A random password generation module 61 is used to generate and display a random verification password after receiving a liveness detection request;

[0109] The video acquisition module 62 is used to acquire the video information of the object to be detected reading the random verification password, wherein the video information to be detected includes the voice information of the object to be detected reading the random verification password displayed by the client;

[0110] The liveness detection module 63 is used to send the video information to be detected, the random verification password and the account information corresponding to the liveness detection request to the server, so that the server can use lip reading recognition technology to extract the verification password to be detected from the video information to be detected, and perform liveness detection on the object to be detected in combination with the voiceprint data to be detected in the video information to be detected.

[0111] The liveness detection device provided in the embodiments of this specification is based on a joint liveness detection method of voiceprint recognition and lip reading recognition. It collects the random verification password displayed on the client screen read by the user, and uses voiceprint recognition and lip reading recognition technology to identify whether the user's reading is consistent with the displayed random verification password, and whether the voiceprint is consistent with the target voiceprint, thereby determining whether it is a real person. At the architectural level, this method uses information from both voice and video modalities, and uses voiceprint recognition and lip reading recognition technology to jointly determine whether the object to be detected is alive. This can greatly improve the accuracy of the liveness detection algorithm, and at the same time, it can also play a good preventive role against attacks that are difficult to handle with single-frame images, such as high-definition screen attacks and high-precision head model attacks.

[0112] It should be noted that the above-mentioned device may also include other implementations according to the description of the corresponding method embodiment. The specific implementation methods can refer to the description of the corresponding method embodiment above, and will not be described one by one here.

[0113] The embodiment of this specification also provides a living body detection device, including: at least one processor and a memory for storing processor executable instructions, and the processor implements the information recommendation data processing method of the above embodiment when executing the instructions, such as:

[0114] Receiving video information to be detected, wherein the video information to be detected includes voice information of the object to be detected reading a random verification password displayed by the client;

[0115] Extracting the voiceprint data to be detected from the voice information, and extracting the verification password to be detected from the video information to be detected by using lip reading recognition technology;

[0116] The voiceprint data to be detected is compared with the voiceprint data of the pre-stored target object, and the verification password to be detected is compared with the random verification password. If the voiceprint data to be detected is the same as the voiceprint data of the target object, and the verification password to be detected is the same as the random verification password, it is determined that the liveness detection of the object to be detected has passed.

[0117] Or, after receiving a liveness detection request, generate and display a random verification password;

[0118] Collecting video information of the object to be detected reading the random verification password, wherein the video information to be detected includes voice information of the object to be detected reading the random verification password displayed by the client;

[0119] The video information to be detected, the random verification password and the account information corresponding to the liveness detection request are sent to the server, so that the server uses lip reading recognition technology to extract the verification password to be detected from the video information to be detected, and performs liveness detection on the object to be detected in combination with the voiceprint data to be detected in the video information to be detected.

[0120] The embodiment of the present specification also provides a liveness detection system, including: a client and a server; wherein the server includes at least one processor and a memory for storing processor executable instructions, and the processor executes the method executed by the server, for performing liveness detection based on the video information to be detected collected by the client;

[0121] The client includes at least one processor and a memory for storing processor-executable instructions, and the processor implements the method executed by the client when executing the instructions.

[0122] It should be noted that the above-mentioned device and system may also include other implementation modes according to the description of the method embodiment. The specific implementation modes may refer to the description of the relevant method embodiment, which will not be described one by one here.

[0123] The liveness detection device provided in this specification can also be used in a variety of data analysis and processing systems. The system or server or terminal or device can be a separate server, or it can include a server cluster, system (including distributed system), software (application), actual operation device, logic gate circuit device, quantum computer, etc. that uses one or more of the methods or one or more embodiments of the system or server or terminal or device of this specification and a terminal device combined with necessary implementation hardware. The detection system for checking the difference data may include at least one processor and a memory storing computer executable instructions, and the processor implements the steps of the method described in any one or more of the above embodiments when executing the instructions.

[0124] The method embodiments provided in the embodiments of this specification can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Taking running on a server as an example, Figure 7 1 is a hardware structure block diagram of a liveness detection server in one embodiment of this specification. The computer terminal may be a liveness detection server or a liveness detection device in the above embodiment. Figure 7 The server 10 shown may include one or more (only one is shown in the figure) processors 100 (the processor 100 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a non-volatile memory 200 for storing data, and a transmission module 300 for communication functions. Those skilled in the art will understand that Figure 7 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 7 More or fewer components shown in the figure may also include other processing hardware, such as a database or multi-level cache, GPU, or other hardware with Figure 7 Different configurations are shown.

[0125] The non-volatile memory 200 can be used to store software programs and modules of application software, such as program instructions / modules corresponding to the liveness detection method in the embodiment of this specification. The processor 100 executes various functional applications and resource data updates by running the software programs and modules stored in the non-volatile memory 200. The non-volatile memory 200 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the non-volatile memory 200 may further include a memory remotely arranged relative to the processor 100, and these remote memories may be connected to a computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0126] The transmission module 300 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of a computer terminal. In one example, the transmission module 300 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 300 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0127] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0128] The methods or devices described in the above embodiments provided in this specification can implement business logic through computer programs and record them on storage media, and the storage media can be read and executed by computers to achieve the effects of the solutions described in the embodiments of this specification, such as:

[0129] Receiving video information to be detected, wherein the video information to be detected includes voice information of the object to be detected reading a random verification password displayed by the client;

[0130] Extracting the voiceprint data to be detected from the voice information, and extracting the verification password to be detected from the video information to be detected by using lip reading recognition technology;

[0131] The voiceprint data to be detected is compared with the voiceprint data of the pre-stored target object, and the verification password to be detected is compared with the random verification password. If the voiceprint data to be detected is the same as the voiceprint data of the target object, and the verification password to be detected is the same as the random verification password, it is determined that the liveness detection of the object to be detected has passed.

[0132] Or, after receiving a liveness detection request, generate and display a random verification password;

[0133] Collecting video information of the object to be detected reading the random verification password, wherein the video information to be detected includes voice information of the object to be detected reading the random verification password displayed by the client;

[0134] The video information to be detected, the random verification password and the account information corresponding to the liveness detection request are sent to the server, so that the server uses lip reading recognition technology to extract the verification password to be detected from the video information to be detected, and performs liveness detection on the object to be detected in combination with the voiceprint data to be detected in the video information to be detected.

[0135] The above-mentioned liveness detection method or device provided in the embodiments of this specification can be implemented by a processor in a computer executing corresponding program instructions, such as using the C++ language of the Windows operating system to implement it on a PC, on a Linux system, or other systems such as Android and iOS system programming languages ​​to implement it on a smart terminal, as well as based on the processing logic of a quantum computer.

[0136] It should be noted that the device, computer storage medium, and system described in the specification may also include other implementation methods according to the description of the relevant method embodiments. The specific implementation methods can refer to the description of the corresponding method embodiments, which will not be described one by one here.

[0137] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the hardware + program embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0138] The embodiments of this specification are not limited to complying with industry communication standards, standard computer resource data update and data storage rules, or the situations described in one or more embodiments of this specification. Certain industry standards or slightly modified implementation plans based on the implementation described in the custom method or embodiment can also achieve the same, equivalent or similar, or predictable implementation effects after deformation of the above-mentioned embodiments. The embodiments obtained by using these modified or deformed data acquisition, storage, judgment, processing methods, etc. can still fall within the scope of the optional implementation plans of the embodiments of this specification.

[0139] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a vehicle-mounted human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0140] Although one or more embodiments of the present specification provide method operation steps as described in the embodiments or flow charts, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps, and does not represent a unique execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, an environment of a parallel processor or multi-threaded processing, or even a distributed resource data update environment). The term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. The first, second, etc. words are used to represent the name, and do not represent any particular order.

[0141] For the convenience of description, the above devices are described in various modules according to their functions. Of course, when implementing one or more of the present specification, the functions of each module can be implemented in the same or more software and / or hardware, or the module implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0142] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as a combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable resource data update device to produce a machine, so that instructions executed by the processor of the computer or other programmable resource data update device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0143] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable resource data update device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0144] These computer program instructions may also be loaded onto a computer or other programmable resource data updating device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions executed on the computer or other programmable device for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0145] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0146] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0147] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0148] It should be understood by those skilled in the art that one or more embodiments of the present specification may be provided as a method, system or computer program product. Therefore, one or more embodiments of the present specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, one or more embodiments of the present specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0149] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0150] Each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. In the description of this specification, the description of the reference term "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representation of the above terms does not necessarily target the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.

[0151] The above description is only an example of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. For those skilled in the art, one or more embodiments of the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims.

Claims

1. A method for detecting a living body, the method comprising: include: Receiving video information to be detected, wherein the video information to be detected includes voice information of the object to be detected reading a random verification password displayed by the client; Extract the voiceprint data to be detected from the voice information, and extract the verification password to be detected from the video information to be detected by using lip reading recognition technology; wherein, the method of extracting the verification password to be detected from the video information to be detected by using lip reading recognition technology comprises: inputting the silent video information in the video information to be detected into a pre-established lip reading recognition model, and extracting the verification password to be detected from the silent video information by using the lip reading recognition model; the lip reading recognition model is obtained by model training and construction based on lip reading training sample data, and the lip reading training sample data of the lip reading recognition model includes construction training sample data; the method of obtaining the construction training sample data comprises: using the video data of a designated user reading a designated random verification password to drive the silent facial images of different users, generating the generated video data of different users reading the designated random verification password, and using the generated video data as the construction training sample data; The voiceprint data to be detected is compared with the voiceprint data of the pre-stored target object, and the verification password to be detected is compared with the random verification password. If the voiceprint data to be detected is the same as the voiceprint data of the target object, and the verification password to be detected is the same as the random verification password, it is determined that the liveness detection of the object to be detected has passed.

2. The method according to claim 1, wherein the lip reading training sample data of the lip reading recognition model further comprises real training sample data; in, The method for acquiring real training sample data includes: collecting video data of different users reading different random verification passwords.

3. The method according to claim 1, wherein the random verification password includes a plurality of sub-passwords, each of which is a single character or a single number, wherein the video data of a designated user reading a designated random verification password drives silent facial images of different users to generate generated video data of different users reading the designated random verification password, and further include: Using the video data of the designated user reading the sub-password in sequence to drive the silent facial images of different users, the generated video data of different users reading the sub-password is generated; The video data generated when different users read the sub-passwords are divided into one video for each sub-password read, so as to obtain sub-video data generated when different users read different sub-passwords; The generated sub-video data obtained by reading different sub-passwords by the same user are randomly combined and synthesized, and the obtained synthesized video data is used as the generated video data.

4. The method according to claim 1, wherein the random verification password includes a plurality of sub-passwords, each of which is a single character or a single number, and wherein the video data of a designated user reading a designated random verification password is used to construct video data of different users reading different random verification passwords, thereby obtaining the constructed training sample data. include: Collect video data of different users reading sub-passwords in sequence, divide the video data of different users reading sub-passwords in sequence into one video according to each sub-password read, and obtain video data of different users reading different sub-passwords; The video data of the same user reading different sub-passwords are randomly combined and synthesized, and the obtained synthesized video data is used as the constructed training sample data.

5. A method for detecting a living body, the method comprising: include: After receiving a liveness detection request, a random verification password is generated and displayed; Collecting video information of the object to be detected reading the random verification password, wherein the video information to be detected includes voice information of the object to be detected reading the random verification password displayed by the client; The video information to be detected, the random verification password and the account information corresponding to the liveness detection request are sent to the server, so that the server extracts the verification password to be detected from the video information to be detected by using lip reading recognition technology, and performs liveness detection on the object to be detected in combination with the voiceprint data to be detected in the video information to be detected; wherein, the method of extracting the verification password to be detected from the video information to be detected by using lip reading recognition technology includes: inputting the silent video information in the video information to be detected into a pre-established lip reading recognition model, and extracting the verification password to be detected from the silent video information by using the lip reading recognition model; the lip reading recognition model is obtained by model training based on lip reading training sample data, and the lip reading training sample data of the lip reading recognition model includes constructed training sample data; the method for obtaining the constructed training sample data includes: using the video data of a specified user reading a specified random verification password to drive silent facial images of different users, generating generated video data of different users reading the specified random verification password, and using the generated video data as the constructed training sample data.

6. A living body detection device, the device include: A video information receiving module, used for receiving the video information to be detected, wherein the video information to be detected includes the voice information of the object to be detected reading the random verification password displayed by the client; A data extraction module is used to extract the voiceprint data to be detected from the voice information, and to extract the verification password to be detected from the video information to be detected by using lip reading recognition technology; wherein, the method of extracting the verification password to be detected from the video information to be detected by using lip reading recognition technology comprises: inputting the silent video information in the video information to be detected into a pre-established lip reading recognition model, and extracting the verification password to be detected from the silent video information by using the lip reading recognition model; the lip reading recognition model is obtained by model training and construction based on lip reading training sample data, and the lip reading training sample data of the lip reading recognition model includes construction training sample data; the method of obtaining the construction training sample data comprises: using the video data of a designated user reading a designated random verification password to drive the silent facial images of different users, generating the generated video data of different users reading the designated random verification password, and using the generated video data as the construction training sample data; The data comparison module is used to compare the voiceprint data to be detected with the voiceprint data of the pre-stored target object, and to compare the verification password to be detected with the random verification password. If the voiceprint data to be detected is the same as the voiceprint data of the target object, and the verification password to be detected is the same as the random verification password, it is determined that the liveness detection of the object to be detected has passed.

7. The device according to claim 6, wherein the data extraction module is specifically used for: Inputting the silent video information in the video information to be detected into a pre-established lip reading recognition model, and extracting the verification password to be detected from the silent video information using the lip reading recognition model; in, The lip reading recognition model is obtained by model training and construction based on lip reading training sample data, and the lip reading training sample data of the lip reading recognition model includes: real training sample data and constructed training sample data; The device also includes a training sample data generating module for: Collect video data of different users reading different random verification passwords to obtain the real training sample data; The video data of a user reading a specified random verification password is used to construct video data of different users reading different random verification passwords, thereby obtaining the constructed training sample data.

8. The apparatus according to claim 7, wherein the training sample data generation module is specifically used for: The video data of a designated user reading a designated random verification password is used to drive silent facial images of different users, and generated video data of different users reading the designated random verification password is constructed, and the generated video data is used as the constructed training sample data.

9. The apparatus according to claim 8, wherein the random verification password comprises a plurality of sub-passwords, each of which is a single character or a single number, and the training sample data generation module is further configured to: Using the video data of the designated user reading the sub-password to drive the silent facial images of different users, the generated video data of the different users reading the sub-password are generated; The video data generated when different users read the sub-passwords are divided into one video for each sub-password read, so as to obtain the video data generated when different users read different sub-passwords; The generated video data obtained by reading different sub-passwords by the same user are randomly combined and synthesized, and the obtained synthesized video data is used as the generated video data.

10. The device according to claim 7, wherein the random verification password includes a plurality of sub-passwords, and the sub-passwords are single characters or single numbers; and the training sample data generation module is specifically used for: The video data of different users reading sub-passwords are collected, and the video data of different users reading sub-passwords are segmented to obtain video data of different users reading different sub-passwords, and the video data of the same user reading different sub-passwords are randomly combined and synthesized to obtain the synthesized video data as the constructed training sample data.

11. A living body detection device, the device include: A random password generation module is used to generate and display a random verification password after receiving a liveness detection request; A video acquisition module, used for acquiring video information to be detected of the object to be detected reading the random verification password, wherein the video information to be detected includes voice information of the object to be detected reading the random verification password displayed by the client; A liveness detection module is used to send the video information to be detected, the random verification password and the account information corresponding to the liveness detection request to a server, so that the server extracts the verification password to be detected from the video information to be detected by using lip reading recognition technology, and performs liveness detection on the object to be detected in combination with the voiceprint data to be detected in the video information to be detected; wherein, the extraction of the verification password to be detected from the video information to be detected by using lip reading recognition technology includes: inputting the silent video information in the video information to be detected into a pre-established lip reading recognition model, and extracting the verification password to be detected from the silent video information by using the lip reading recognition model; the lip reading recognition model is obtained by model training based on lip reading training sample data, and the lip reading training sample data of the lip reading recognition model includes constructed training sample data; the method for obtaining the constructed training sample data includes: using video data of a designated user reading a designated random verification password to drive silent facial images of different users, generating generated video data of different users reading the designated random verification password, and using the generated video data as the constructed training sample data.

12. A liveness detection device, include: At least one processor and a memory for storing processor-executable instructions, wherein when the processor executes the instructions, the method according to any one of claims 1 to 4 or claim 5 is implemented.

13. A liveness detection system, include: A client and a server, wherein the server comprises at least one processor and a memory for storing processor executable instructions, and when the processor executes the instructions, the method according to any one of claims 1 to 4 is implemented, for performing liveness detection based on the video information to be detected collected by the client; The client comprises at least one processor and a memory for storing processor-executable instructions, and the method of claim 5 is implemented when the processor executes the instructions.

Citation Information

Patent Citations

  • Living body recognition method and device and electronic equipment

    CN113011301A