Active liveness detection method and apparatus using face images
The active liveness detection method using command words and facial analysis improves security by preventing media attacks and enhancing accuracy in face recognition systems.
Patent Information
- Application Number
- JP2024514015
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-02
- Filing Date
- 2022-03-28
- Publication Date
- 2025-07-28
- Estimated Expiration
- 2042-03-28
AI Technical Summary
Existing face liveness detection technologies using RGB cameras are vulnerable to media streaming attacks and have limited action combinations, reducing their effectiveness and accessibility.
An active liveness detection method that requires users to pronounce randomly generated command words, analyzing mouth shape and facial muscle changes in face videos to verify authenticity.
Enhances security by preventing media streaming attacks and improving accuracy through real-time detection of liveness using mouth and facial muscle feature points.
Smart Images

Figure 0007713616000005 
Figure 0007713616000006 
Figure 0007713616000007
Abstract
Description
Technical Field
[0001] The disclosed technology relates to a method and apparatus for detecting liveness using a face image of a user who pronounces command words.
Background Art
[0002] With the increase in the non-face-to-face service environment and the growth of the face recognition market, various services utilizing face recognition technology have been launched. In this context, the importance of liveness detection or even anti-spoofing is increasing day by day.
[0003] Due to such importance, research on various liveness detection methods and development of products are underway. However, since the liveness detection technology incorporated in most commercially available face recognition systems is based on a method using additional sensors or special cameras, the accessibility from the perspective of actual users is reduced.
[0004] Recently, various face liveness detection prevention techniques based on RGB cameras have emerged. However, they still show results such as operating only under specific sensors and environmental conditions, punching holes in printed paper, or being vulnerable to media streaming attacks. As a result, the need for a more stable liveness detection technology for use on general-purpose cameras such as RGB cameras is increasing.
[0005] On the other hand, methods for face liveness detection based on RGB cameras can be classified into a single frame based method that uses only still images and a multi frame based method that uses one or more frames according to the video frames used, and can be classified into an active method and a passive method depending on whether the user's operation is required.
[0006] The existing multi-frame based liveness detection method provided by embodiment has followed the method of inducing one action such as turning the user's head according to the instructions on the screen, blinking the eyes, or opening the mouth, or combining two or more actions for judgment. However, since the combination of actions that can be induced by the corresponding method is limited, there is a problem that it still has a vulnerable structure against media streaming attacks.
Summary of the Invention
Problems to be Solved by the Invention
[0007] The disclosed technology aims to provide a method and device for detecting liveness by making the user pronounce randomly generated command words and acquiring the face video of the user pronouncing the command words.
Means for Solving the Problems
[0008] The first aspect of the disclosed technology to achieve the above technical problem is an active liveness detection method including the steps of: the authentication device generating a command word for user authentication; the authentication device outputting the command word to the screen and capturing the face video of the user pronouncing the output command word; and the authentication device extracting feature points for changes in the shape of the mouth from the face video and determining whether the user imitated and pronounced the command word.
[0009] The second aspect of the disclosed technology to achieve the above technical problem is an active liveness detection device including a sensor for identifying whether a user has entered within a range where face capture is possible, a storage device for storing data for a plurality of syllables to generate a command word for user authentication, an output device for outputting the command word to the screen and capturing the face video of the user pronouncing the output command word, and an arithmetic device for generating the command word consisting of a certain number of syllables using the data for the plurality of syllables, extracting feature points from the face video, and determining whether the user imitated and pronounced the command word.
Effects of the Invention
[0010] Examples of the disclosed technology can have effects including the following advantages. However, the scope of rights of the disclosed technology should not be understood to be limited thereby, as it does not mean that examples of the disclosed technology must include all of these.
[0011] An active liveliness detection method and apparatus using a face video according to an example of the disclosed technology can detect the liveliness of a user through an instruction word generated in real time.
[0012] Also, since an instruction word combined with syllables having meaningless meanings is generated each time, there is an effect of preventing problems caused by the leakage of the instruction word.
[0013] Also, there is an effect of improving the accuracy by detecting liveliness based on feature points for changes in the shape of the mouth that pronounces syllables and feature points for changes in facial muscles between syllables.
Brief Description of Drawings
[0014]
Figure 1
[0015]
Figure 2
[0016]
Figure 3
[0017]
Figure 4
[0018]
Figure 5
[0019]
Figure 6
Mode for Carrying Out the Invention
[0020] The present invention can be subjected to various modifications and can have various embodiments. Specific embodiments are illustrated in the drawings and will be described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, and it should be understood to include all modifications, equivalents, and alternatives included in the spirit and technical scope of the present invention.
[0021] Terms such as second, A, B, etc. can be used to describe various components, but the corresponding components are not limited by the terms, and are used only for the purpose of distinguishing one component from another. For example, without departing from the scope of the rights of the present invention, the first component can be named the second component, and similarly, the second component can also be named the first component. The term "and / or" includes a combination of a plurality of related described items or any one of a plurality of related described items.
[0022] It should be understood that the singular expressions of the terms used in this specification include plural expressions unless clearly interpreted differently in the context. And terms such as "including" can mean that the described features, numbers, steps, operations, components, parts, or combinations thereof exist, and do not exclude the possibility of the existence or addition of one or more other features, numbers, step operations, components, parts, or combinations thereof.
[0023] Prior to the detailed description of the drawings, it should be made clear that the classification of the components in this specification is only based on the main functions of each component. That is to say, two or more components described below may be integrated into one component, or one component may be divided into two or more components according to more refined functions.
[0024] And each of the components described below may additionally perform some or all of the functions that other components are responsible for in addition to the main function it is responsible for. Needless to say, some of the functions that each component is responsible for may be exclusively performed by other components. Therefore, the presence or absence of each component described throughout this specification should be interpreted functionally.
[0025] Figure 1 is an example of an active liveness detection system 100 using face video according to an embodiment of the disclosed technology. Referring to Figure 1, the authentication device 110 can generate a command word for user authentication and capture the user's pronunciation by imitating the command word. Then, it can analyze the changes in the mouth shape and facial muscles from the captured face video to determine whether the user pronounced by imitating the presented command word.
[0026] The authentication device 110 captures the user's face video. A camera can be installed for video capture. The authentication device 110 can determine the captureable distance according to the type of camera. When it senses that the user has entered within the captureable range, it can generate a command word for user authentication. Then, it can drive the camera to capture the user's face video. The face video captured by the authentication device 110 includes the changes in the mouth shape and facial muscles during the process of the user pronouncing by imitating the command word. The authentication device 110 can extract the feature points for the changes in the mouth shape and facial muscles and compare them with the feature points for the command word.
[0027] On the one hand, the command words referred to in the disclosed technology are for checking the liveness of the user, and mean a certain number of syllables that anyone can pronounce by sufficiently imitating as long as the user approaching the authentication device 110 is an actual person. The command words can randomly select a certain number of syllables out of a plurality of syllables to be used as command words.
[0028] On the other hand, there are cases where the mouth shapes for pronouncing some syllables have no difference. For example, in the case of "go" and "o", they are syllables in which the same vowel is respectively
Number
[0029] On the other hand, in order to distinguish the mouth shape, the authentication device 110 can store in the database 120 the combinations between some vowels whose pronunciation shapes are distinguishable among the Hangul vowels and consonants that become bilabial sounds when combined with some vowels, and the combinations between some vowels whose pronunciation shapes are distinguishable and consonants that do not become bilabial sounds when combined with some vowels. And a certain number of syllables randomly selected from these can be generated as command words.
[0030] On the one hand, in the database 120, the characteristic points of the mouth shape that match each other for each syllable can be stored together. For example, the characteristic point information for the mouth shape when pronouncing "O" can be stored together with the syllable "O". For each syllable stored in the database 120, the characteristic point information is matched, and this can be used to compare with the characteristic points extracted from the user's face image. The authentication device 110 can compare the characteristic points for the change in the mouth shape extracted from the face image with the characteristic points for the command word to determine whether the user imitated and pronounced the command word. At this time, according to the number of syllables included in the command word, it can be sequentially divided and the similarity can be compared one by one. For example, if the number of syllables included in the command word is 5, the characteristic points extracted from the user's face image can be divided into 5 parts, and the similarity can be compared sequentially. In this process, if the similarity exceeds the threshold value, it can be determined that the command word was accurately pronounced by imitation. In FIG. 1, Hangul syllables were used as an example, but syllables in English and other languages can also be applied based on the same principle. For example, in the case of English, after storing the remaining ones excluding some alphabets with indistinguishable pronunciations among the 26 alphabets as a whole in the database 120, a certain number can be randomly selected from these to generate a command word.
[0031] On the one hand, during the process of pronouncing the command word, the authentication device 110 can more precisely check liveness by utilizing the changes in facial muscles between syllables. When a person pronounces the next syllable after pronouncing a specific syllable, the changes in facial muscles between syllables can be in a similar form for anyone. That is, if it is an actual person pronouncing the syllables, no matter who imitates the command word, the changes in facial muscles between syllables can appear with similar characteristic points. Therefore, based on the syllables previously stored in the database 120, the characteristic point information for the changes in facial muscles corresponding to all cases can be stored, and in the future, the liveness can be determined by comparing the characteristic points of the changes in facial muscles in the user's face image. In one embodiment, if the number of syllables included in the command word is 5, first, the characteristic points extracted from the user's face image are divided into 5 parts, and then the end of the first divided characteristic point and the beginning of the second divided characteristic point are combined to be determined as the characteristic points for the changes in facial muscles between the first syllable and the second syllable. Then, the liveness can be more precisely determined by comparing with the characteristic points stored in the database 120.
[0032] On the other hand, if it is determined that the user accurately pronounces by imitating the command word, the authentication device 110 can transmit a control signal to the control device that controls the opening and closing of the gate so that the user can pass through the gate. The authentication device 110 may be connected to the control device that controls the gate, or the two devices may be combined into one. If they are separate devices, the authentication device 110 can transmit a control signal to the control device to control the gate, and if they are the same device, the gate can be controlled immediately. Therefore, by checking the active liveness using the command word generated in real time, the leakage of the password for gate access can be prevented and the security can be enhanced.
[0033] FIG. 2 is a flowchart for an active liveness detection method using a face image according to an embodiment of the disclosed technology. Referring to FIG. 2, the active liveness detection method 200 includes a command word generation stage 210, a face image capture stage 220, and a pronunciation determination stage 230. Each stage can be sequentially performed through the authentication device.
[0034] At stage 210, the authentication device generates command words for user authentication. The command words are a certain number of syllables randomly selected from a plurality of syllables, and the process of generating the command words is as described through FIG. 1. The authentication device can determine the number of syllables to be generated as command words according to a preset value. The time point of generating the command words at stage 210 can be the time point when the user enters the range where the face can be photographed.
[0035] At stage 220, the authentication device photographs the user's face image. Prior to photographing the face image, the authentication device can output the generated command words to the user. For example, the command words generated inside the device can be output through the screen. Then, the face image when the user pronounces by imitating the output command words is photographed.
[0036] At stage 230, the authentication device extracts feature points for the change in the shape of the mouth from the face image to determine whether the user pronounces by imitating the command words. At stage 230, the authentication device can compare whether the feature points of the shape of the mouth extracted from the user's face image are similar to the feature points of the shape of the mouth for pronouncing each of the certain number of syllables included in the command words, so as to determine whether the user pronounces the command words accurately. At this time, it is also possible to compare whether the feature points of the facial muscles extracted from the face image match the feature points for the change in the face between a certain number of syllables included in the command words, so as to detect more sophisticated liveness.
[0037] In the case of the conventional authentication technology based on voice recognition infrastructure, words or syllables previously registered by the user as passwords can be used. However, in this case, there is a problem that the previously registered information leaks. In the case of the authentication technology using a method of comparing the similarity of voices, it generates voices similar to the mechanically registered voices of the user, and there is a problem that unauthorized persons can pass the authentication. In order to solve such problems, this technology has the advantage of accurately authenticating the user himself because it can check the liveness of the user using command words randomly generated in real time.
[0038] Figure 3 is a block diagram of an active liveliness detection device using a face video according to an embodiment of the disclosed technology. The detection device 300 can be embodied in various forms such as a PC, a notebook computer, a smart device, a wearable device, etc. Referring to Figure 3, the active liveliness detection device 300 includes a sensor 310, a storage device 320, an output device 330, and an arithmetic device 340.
[0039] The sensor identifies whether a user has entered within a range where the face can be photographed. As an example, it can be an infrared sensor that senses the approach of a person at a gate. Of course, ultrasonic waves or other types of human body sensing sensors that are not infrared can be used.
[0040] The storage device 320 stores data for a plurality of syllables to generate command words for user authentication. The storage device 320 stores not only a plurality of syllables but also feature point information for the shape of the mouth when each syllable is pronounced and feature point information for the change of facial muscles between syllables, so a memory with a capacity sufficient to store all the data can be used.
[0041] The output device 330 includes a display that outputs a command word to the screen and a camera that captures a face video of the user who mimics and pronounces the output command word. And when outputting the command word through the display, voice or guidance messages can be further output so that the user can mimic and read the command word. Therefore, the output device can also be equipped with a speaker for voice output.
[0042] The arithmetic device 340 generates a command word consisting of a certain number of syllables using the data for the plurality of syllables stored in the storage device 320. Then, it extracts feature points from the face video captured by the output device 330 and determines whether the user mimicked and pronounced the command word. The arithmetic device can be a processor or a CPU of the active liveliness detection device 300.
[0043] On the one hand, the above-described active liveness detection device 300 can be embodied as a program (or application) including an executable algorithm executable by a computer. That is, it can be a program executed on a computer. The program can be provided and stored in a non-transitory computer readable medium.
[0044] The non-transitory computer readable medium means a medium that does not store data for a short moment such as a register, cache, memory, etc., but stores data semi-permanently and can be read by a device. Specifically, the various applications or programs described above can be provided and stored in a non-transitory computer readable medium such as a CD, DVD, hard disk, Blu-ray disk, USB, memory card, ROM (read-only memory), PROM (programmable read only memory), EPROM (Erasable PROM, EPROM), or EEPROM (Electrically EPROM), or flash memory.
[0045] The transitory computer readable medium means various types of RAM such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synclink DRAM (SLDRAM), and direct rambus RAM (DRRAM).
[0046] FIG. 4 is a drawing showing a selection of some distinguishable vowels among the Hangul vowels. As shown in FIG. 4, 14 vowels that can be distinguished in pronunciation are selected from among the 21 vowels that make up Hangul. Each of the 14 selected vowels is
Number
[0047] FIG. 5 is a drawing showing command words generated by a certain number of syllables. As shown in FIG. 5, command words can be generated with 5 syllables. The syllables selected as command words are each a combination of a consonant and a vowel. Here, the vowel is the one selected through FIG. 4, and the consonant can be partially selected from the total consonants as those that become bilabial sounds and those that do not become bilabial sounds when combined with the vowel. For example, like the command word in FIG. 5, the consonants that become bilabial sounds are
Number
Number
[0048] On the other hand, a guidance message can be output on the screen in the process of presenting the command word to the user. As shown in FIG. 5, a message such as "Please pronounce as follows for authentication." can be output, and the message can also be output in voice through a speaker.
[0049] FIG. 6 is a drawing showing the extraction of feature points from a face image. Referring to FIG. 6, a deep learning model can be used for feature point extraction. The deep learning model can be pre-trained for feature point extraction, and after the training is completed, it can be installed in the authentication device to extract feature points for the user's face image. At this time, in order to detect whether the command word presented to the user is pronounced correctly, it can be compared after extracting the feature points for the shape of the mouth.
[0050] In the data saved for command word generation, information for feature points indicating pronunciation can be matched and saved for each syllable together with a plurality of syllables. That is, after extracting the feature points of the mouth shape from the face image captured through the deep learning model, by matching them with the feature points of each syllable included in the command word, it is possible to determine whether the user pronounced the command word correctly. Needless to say, it is almost impossible for the two feature points to match 100%, so the success or failure of authentication can be determined by comparing the similarity between the feature points. If the similarity is above a certain level, it can be determined that there is no abnormality in the user liveness.
[0051] On the other hand, when the user pronounces syllables in a specific order and then pronounces the next-order syllable, the shape of the facial muscles can change in a certain determined form. That is, it is possible to predict the facial muscle changes that are roughly similar regardless of who pronounces them. In this technology, by focusing on such points, after pre-saving the feature point information corresponding to the facial changes between syllables, the feature points for the facial changes between syllables can be extracted from the face image and matched to more precisely detect the user's liveness. For example, the liveness can be detected by comparing the changes in the shape and position of the cheekbones, philtrum, eyebrows, etc. between syllables.
[0052] The active liveliness detection method and apparatus using face video according to an embodiment of the disclosed technology have been described with reference to the embodiments illustrated in the drawings to assist understanding, but these are merely exemplary, and those having ordinary knowledge in the art will understand that various modifications and equivalent other embodiments will be possible hereinafter. Therefore, the true technical protection scope of the disclosed technology should be determined by the appended claims.
Claims
1. A step in which an authentication device generates a command word for user authentication; A step in which the authentication device outputs the command word on a screen and captures a face image of a user who pronounces the output command word; and A step in which the authentication device extracts feature points for changes in the shape of the mouth from the face image and determines whether the user imitated and pronounced the command word; including The command word is a certain number of syllables randomly selected from a plurality of syllables, The certain number of syllables A combination between some vowels whose pronunciation forms are classified among the vowels of Hangul and consonants that become bilabial sounds when combined with the some vowels, and Randomly selected from a combination between some vowels whose pronunciation forms are classified and consonants that do not become bilabial sounds when combined with the some vowels, An active liveness detection method.
2. The generating step generates the command word if it identifies that the user has entered within a range where face capture is possible. The active liveness detection method according to claim 1.
3. The determining step determines whether the feature points of the shape of the mouth extracted from the face image are similar to the feature points of the shape of the mouth for pronouncing each of the certain number of syllables included in the command word. The active liveness detection method according to claim 1.
4. The determining step determines whether the feature points of the facial muscles extracted from the face image match the feature points for changes in the face between the certain number of syllables included in the command word. The active liveness detection method according to claim 1.
5. The authentication device is connected to a control device that controls the opening and closing of a gate and transmits a control signal to the control device according to the determination result. The active liveness detection method according to claim 1.
6. A sensor that identifies whether a user has entered within a range where face capture is possible; A storage device that stores data for a plurality of syllables to generate a command word for the user authentication; An output device that outputs the command word on a screen and captures a face image of a user who pronounces the output command word; and An arithmetic device that generates the command word consisting of a certain number of syllables using the data for the plurality of syllables, extracts feature points from the face image, and determines whether the user imitated and pronounced the command word; including The certain number of syllables A combination between some vowels whose pronunciation forms are classified among the vowels of Hangul and consonants that become bilabial sounds when combined with the some vowels, and Randomly selected from a combination of some vowels whose pronunciation forms are classified and consonants that do not become bilabial sounds when combined with the some vowels. Active liveness detection device. **Claim 7** The active liveness detection device according to claim 6, wherein the arithmetic unit determines whether feature points of the mouth shape extracted from the face video are similar to feature points of the mouth shape for pronouncing a certain number of syllables included in the command word. **Claim 8** The active liveness detection device according to claim 6, wherein the arithmetic unit determines whether feature points of facial muscles extracted from the face video are matched to feature points for facial changes between a certain number of syllables included in the command word. **Claim 9** The arithmetic unit includes a deep learning model for analyzing the face video. The active liveness detection device according to claim 6, wherein the pronunciation of the user is determined based on the result output by inputting the face video into the deep learning model. **Claim 10** The active liveness detection device according to claim 6, wherein the arithmetic unit transmits a control signal to a control device that controls the opening and closing of a gate based on the determination result.
Citation Information
Patent Citations
Method and device for speech input
JP1994012483A
Personal information authentication system, personal information authentication method, program, and recording medium
JP2008004050A