Face Liveness Detection
UHF echo-based methods with multiple oriented audio detectors and classifiers enhance facial liveness detection on mobile devices, addressing 3D mask attacks and improving user convenience.
Patent Information
- Application Number
- JP2024559228
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-04-13
- Filing Date
- 2023-01-27
- Publication Date
- 2025-11-06
- Estimated Expiration
- 2043-01-27
AI Technical Summary
Facial recognition systems on mobile devices are vulnerable to 3D mask attacks, and existing echo-based methods using audible audio signals are compromised and inconvenient for users.
Implementing ultra-high frequency (UHF) echo-based methods with multiple audio detectors oriented differently to capture UHF signal reflections, analyzing echo signals to extract features, and applying a classifier to determine live faces, using inaudible audio signals.
Enhances robustness against 3D mask attacks by improving facial depth resolution and user convenience through inaudible UHF audio signals, effectively distinguishing between live and fake faces.
Smart Images

Figure 0007764979000006 
Figure 0007764979000007 
Figure 0007764979000008
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to computer-readable media, methods, and apparatus. [Background technology]
[0002] Facial recognition systems are used to prevent unauthorized access to devices and services. The integrity of facial recognition systems has been attacked by unauthorized individuals attempting to gain access to protected devices and services. In recent years, mobile devices such as smartphones have utilized facial recognition to prevent unauthorized access to the mobile device.
[0003] As a result of the growing popularity of facial recognition systems on mobile devices, and as the use of facial authentication is an increasingly popular method for unlocking mobile devices such as smartphones, mobile devices are increasingly being targeted by face presentation attacks. Face presentation attacks include 2D printing attacks, which use images of an authorized user's face, replay attacks, which use videos of an authorized user's face, and more recently, 3D mask attacks, which use 3D printed masks of an authorized user's face. Summary of the Invention
[0004] According to a first exemplary aspect of the present disclosure, a computer-readable medium includes instructions executable by a computer that cause the computer to perform operations including emitting an ultra-high frequency (UHF) audio signal through a speaker; obtaining an echo signal by detecting reflections of the UHF audio signal from a surface using a plurality of audio detectors; extracting a plurality of features from the echo signal; and applying a classifier to the plurality of features to determine whether the surface is a live face.
[0005] According to a second exemplary aspect of the present disclosure, a method includes emitting an ultra-high frequency (UHF) audio signal via a speaker; obtaining an echo signal by detecting reflections of the UHF audio signal from a surface using a plurality of audio detectors; extracting a plurality of features from the echo signal; and applying a classifier to the plurality of features to determine whether the surface is a live face.
[0006] According to a third exemplary aspect of the present disclosure, an apparatus includes a plurality of audio detectors; a speaker; and a controller including circuitry configured to: emit an ultra-high frequency (UHF) audio signal through the speaker; obtain an echo signal by detecting reflections of the UHF audio signal from a surface using the plurality of audio detectors; extract a plurality of features from the echo signal; and apply a classifier to the plurality of features to determine whether the surface is a live face. [Brief explanation of the drawings]
[0007] Aspects of the present disclosure are best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be noted that, according to standard industry practice, various features have not been drawn to scale. In fact, the dimensions of various features may be arbitrarily increased or decreased for clarity of illustration. [Figure 1A] FIG. 1 is a top view of a schematic diagram of an apparatus for facial liveness detection, in accordance with at least some embodiments of the present disclosure. [Figure 1B] FIG. 1 is a front view of a schematic diagram of an apparatus for facial liveness detection, in accordance with at least some embodiments of the present disclosure. [Figure 1C] FIG. 1 is a bottom view of a schematic diagram of an apparatus for facial liveness detection, in accordance with at least some embodiments of the present disclosure. [Figure 2] 1 is an operational flow of facial recognition to prevent unauthorized access, in accordance with at least some embodiments of the present disclosure. [Figure 3] 1 is an operational flow of face recognition for face liveness detection, according to at least some embodiments of the present disclosure. [Figure 4] 1 is an operational flow of face recognition for acquiring echo signals, in accordance with at least some embodiments of the present disclosure. [Figure 5] 1 is an operational flow of a first feature extraction and classifier application process according to at least some embodiments of the present disclosure. [Figure 6] 1 is an operational flow for training a neural network for feature extraction, according to at least some embodiments of the present disclosure. [Figure 7] 10 is an operational flow of a second feature extraction and classifier application process according to at least some embodiments of the present disclosure. [Figure 8] FIG. 10 is a schematic diagram of a third feature extraction and classifier application process, according to at least some embodiments of the present disclosure. [Figure 9] FIG. 1 is a block diagram of a hardware configuration for face liveness detection, in accordance with at least some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0008] The following disclosure provides many different embodiments or examples for implementing various features of the provided subject matter. Specific examples of components, values, operations, materials, arrangements, etc. are described below to simplify the disclosure. It should be understood that these are merely examples and are not intended to be limiting. Other components, values, operations, arrangements, etc. are also contemplated. In addition, the disclosure may repeat reference numerals and / or letters in the various examples. This repetition is for purposes of simplicity and clarity and does not, in itself, dictate a relationship between the various embodiments and / or configurations being described.
[0009] There are several echo-based methods that successfully detect 2D printing and replay attacks. However, existing echo-based methods are still vulnerable to 3D mask attacks. For example, there are echo-based methods that use radar transmitter and receiver signal features along with visual features for face liveness detection, and methods that use echo and visual landmark features for face authentication.
[0010] These echo-based methods use audio signals in the 12kHz-20kHz range, which are audible to most users, resulting in user inconvenience. These echo-based methods typically capture the echo signal using only a single microphone, typically located at either the top or bottom of the device, resulting in lower facial depth resolution. These echo-based methods are routinely compromised by 3D mask attacks.
[0011] At least some embodiments described herein utilize ultra-high frequency (UHF) echo-based methods for passive mobile liveness detection. At least some embodiments described herein include echo-based methods that increase robustness by using features commonly found in handheld devices. At least some embodiments described herein analyze echo signals to extract more features to detect 3D mask attacks.
[0012] 1A, 1B, and 1C are schematic diagrams of an apparatus 100 for facial liveness detection, according to at least some embodiments of the present disclosure. Apparatus 100 includes microphone 110, microphone 111, speaker 113, camera 115, display 117, and input 119. FIG. 1B is a front view of the schematic diagram of apparatus 100, illustrating the components described above, according to at least some embodiments of the present disclosure. In at least some embodiments, apparatus 100 is within a handheld device, such that speaker 113 and multiple microphones 110 and 111 are included in the handheld device.
[0013] In at least some embodiments, an apparatus for facial liveness detection includes a plurality of audio detectors. In at least some embodiments, the plurality of audio detectors includes a plurality of microphones, such as microphones 110 and 111. In at least some embodiments, microphones 110 and 111 are configured to convert audio signals into electrical signals. In at least some embodiments, microphones 110 and 111 are configured to convert audio signals into signals of other forms of energy that can be further processed. In at least some embodiments, microphones 110 and 111 are transducers. In at least some embodiments, microphones 110 and 111 are compression microphones, dynamic microphones, etc., in any combination. In at least some embodiments, microphones 110 and 111 are further configured to detect audible signals for phone calls, voice recordings, etc.
[0014] In at least some embodiments, the multiple audio detectors include a first audio detector oriented in a first direction and a second audio detector oriented in a second direction. Microphones 110 and 111 are oriented to receive audio signals from different directions. Microphone 110 is located on the top side of device 100. FIG. 1A is a top view of a schematic diagram of device 100, showing microphone 110 opening upward, in accordance with at least some embodiments of the present disclosure. Microphone 111 is located on the bottom side of device 100. FIG. 1C is a bottom view of a schematic diagram of device 100, showing microphone 110 opening downward, in accordance with at least some embodiments of the present disclosure. In at least some embodiments, the audio detectors are oriented in other directions, such as to the right and left, at an oblique angle, or any combination, as long as the angles of audio reception are different. In at least some embodiments, the reflection patterns of echo signals captured by both microphones with different orientations, along with acoustic absorption and backscatter information, are utilized to detect 2D and 3D face presentation attacks. In at least some embodiments, using microphones at different orientations to capture UHF signal reflections helps separate echo signals by template matching, improving the signal-to-noise ratio. In at least some embodiments, the multiple audio detectors include three or more detectors.
[0015] Speaker 113 is disposed on the front of device 100. In at least some embodiments, speaker 113 is configured to emit a UHF signal. In at least some embodiments, speaker 113 is a loudspeaker, a piezoelectric speaker, or the like. In at least some embodiments, speaker 113 is configured to emit a UHF signal in substantially the same direction as the optical axis of camera 115, such that the UHF signal reflects off a surface being imaged by camera 115. In at least some embodiments, speaker 113 is a transducer configured to convert an electrical signal into an audio signal. In at least some embodiments, speaker 113 is further configured to emit an audible signal, such as for playing videos and music.
[0016] In at least some embodiments, the handheld device further comprises a camera 115. In at least some embodiments, the camera 115 is configured to capture an image of an object in front of the device 100, such as the face of a user holding the device 100. In at least some embodiments, the camera 115 comprises an image sensor configured to convert a visible light signal into an electrical signal or any other signal that can be further processed.
[0017] Display 117 is disposed on the front of device 100. In at least some embodiments, display 117 is configured to generate a visible image, such as a graphical user interface. In at least some embodiments, display 117 includes a liquid crystal display (LCD), a light emitting diode (LED) array, or any other display technology suitable for a handheld device. In at least some embodiments, display 117 is touch-sensitive, such as a touchscreen, and is further configured to accept tactile input. In at least some embodiments, display 117 is configured to show the image currently being captured by camera 115 to assist the user in aiming the camera at the user's face.
[0018] Input 119 is located on the front side of device 100. In at least some embodiments, input 119 is configured to accept tactile input. In at least some embodiments, input 119 is a button, a pressure sensor, a fingerprint sensor, or any other form of tactile input, including combinations thereof.
[0019] 2 is an operational flow of facial recognition to prevent unauthorized access according to at least some embodiments of the present disclosure. The operational flow provides a method of facial recognition to prevent unauthorized access. In at least some embodiments, one or more operations of the method are performed by a controller of a device that includes sections for performing particular operations, such as the controller and device shown in FIG. 9 described below.
[0020] In S220, the controller, or a section thereof, images the surface. In at least some embodiments, the controller images the surface with a camera to obtain a surface image. In at least some embodiments, the controller images the face of a user of a handheld device, such as apparatus 100 of FIG. 1 . In at least some embodiments, the controller images the surface to generate a digital image for image processing.
[0021] In S221, the controller, or a section thereof, analyzes the surface image. In at least some embodiments, the controller analyzes the surface image to determine whether the surface is a face. In at least some embodiments, the controller analyzes the surface image to detect facial features such as eyes, nose, mouth, and ears for further analysis. In at least some embodiments, the controller performs rotation, cropping, or other spatial manipulations to normalize the facial features for face recognition.
[0022] In S222, the controller or a section thereof determines whether the surface is a face. In at least some embodiments, the controller determines whether the surface is a face based on the surface image analysis in S221. If the controller determines that the surface is a face, the operational flow proceeds to liveness detection in S223. If the controller determines that the surface is not a face, the operational flow returns to surface imaging in S220.
[0023] At S223, the controller, or a section thereof, detects surface liveness. In at least some embodiments, the controller detects 2D and 3D face presentation attacks. In at least some embodiments, the controller performs the liveness detection process described below with reference to FIG. 3.
[0024] In S224, the controller or a section thereof determines whether the surface is live. In at least some embodiments, the controller determines whether the surface is a live human face based on the liveness detection in S223. If the controller determines that the surface is live, the operational flow proceeds to surface identification in S226. If the controller determines that the surface is not live, the operational flow proceeds to access denial in S229.
[0025] At S226, the controller, or a section thereof, identifies the surface. In at least some embodiments, the controller applies a facial recognition algorithm, such as comparing geometric or photometric features of the surface with features of identified faces. In at least some embodiments, the controller obtains distance measurements between deep features of the surface and deep features of each identified face and identifies the surface based on the shortest distance. In at least some embodiments, the controller determines that the surface is a face, and in response to determining that the surface is a live face, identifies the surface by analyzing the surface image.
[0026] In S227, the controller or a section thereof determines whether the identity is authorized. In at least some embodiments, the controller determines whether the user identified in S226 is authorized for access. If the controller determines that the identity is authorized, operational flow proceeds to access authorization in S228. If the controller determines that the identity is not authorized, operational flow proceeds to access denial in S229.
[0027] At S228, the controller, or a section thereof, grants access. In at least some embodiments, the controller grants access to at least one of a device or a service in response to identifying the surface as an authorized user. In at least some embodiments, the controller grants access to at least one of a device or a service in response to not identifying the surface as an unauthorized user. In at least some embodiments, the device to which access is granted is a device, such as the handheld device of FIG. 1. In at least some embodiments, the service is a program or application executed by the device.
[0028] At S229, the controller or a section thereof denies access. In at least some embodiments, the controller denies access to at least one of a device or a service in response to not identifying the surface as an authorized user. In at least some embodiments, the controller denies access to at least one of a device or a service in response to identifying the surface as an unauthorized user.
[0029] In at least some embodiments, the controller performs the operations in a different order. In at least some embodiments, the controller detects liveness before analyzing the surface image, or even before imaging the surface. In at least some embodiments, the controller detects liveness after identifying the surface and even after determining whether the identity is authorized. In at least some embodiments, the operational flow is repeated after an access denial, but only for a predetermined number of access denials, until a wait period is implemented, the device is powered down, the device self-destructs, further requested operations, etc.
[0030] 3 is an operational flow for facial liveness detection according to at least some embodiments of the present disclosure. The operational flow provides a method for facial liveness detection. In at least some embodiments, one or more operations of the method are performed by a controller of a device that includes sections for performing particular operations, such as the controller and device illustrated in FIG. 9 described below.
[0031] At S330, the emitting section emits a liveness detection audio signal. In at least some embodiments, the emitting section emits an ultra-high frequency (UHF) audio signal via a speaker. In at least some embodiments, the emitting section emits the UHF audio signal in response to detecting a face by the camera. In at least some embodiments, the emitting section emits the UHF audio signal in response to identifying a face.
[0032] In at least some embodiments, the emission section emits a UHF audio signal that is substantially inaudible. A substantially inaudible audio signal is one that most people cannot hear or are consciously unaware of. The higher the frequency of the audio signal, the greater the number of people who will be unable to hear the audio signal. In at least some embodiments, the emission section emits a UHF audio signal in the range of 18-22 kHz. The waveform of the audio signal also affects the number of people who will be unable to hear the audio signal. In at least some embodiments, the emission section emits a UHF audio signal that includes a sine wave and a sawtooth wave. In at least some embodiments, the emission section emits a UHF audio signal that is a combination of a sine wave and a sawtooth wave in the range of 18-22 kHz via a mobile phone to illuminate the user's face.
[0033] At S332, the acquisition section acquires echo signals. In at least some embodiments, the acquisition section acquires the echo signals by detecting reflections of the UHF audio signal from the surface with multiple audio detectors. In at least some embodiments, the acquisition section converts the raw recordings of the multiple audio detectors into a single signal representing one or more echoes of the audio signal emitted from the surface. In at least some embodiments, the acquisition section performs the echo signal acquisition process described below with respect to FIG. 4.
[0034] At S334, the extraction section extracts features from the echo signals. In at least some embodiments, the extraction section extracts multiple features from the echo signals. In at least some embodiments, the extraction section extracts handcrafted features from the echo signals, such as by using a formula to calculate specific characteristics. In at least some embodiments, the extraction section applies one or more neural networks to the echo signals to extract condensed feature representations. In at least some embodiments, the extraction section performs the feature extraction process described below with respect to FIG. 5, FIG. 7, or FIG. 8.
[0035] At S336, the application section applies a classifier to the features. In at least some embodiments, the application section applies the classifier to multiple features to determine whether the surface is a live face. In at least some embodiments, the application section applies a threshold to each feature to create a binary classification of the feature as consistent or inconsistent with a live human face. In at least some embodiments, the application section applies a neural network classifier to a concatenation of the features to generate a binary classification of the feature as consistent or inconsistent with a live human face. In at least some embodiments, the application section performs the classifier application process described below with respect to FIG. 5, FIG. 7, or FIG. 8.
[0036] 4 is an operational flow for acquiring echo signals according to at least some embodiments of the present disclosure. The operational flow provides a method of echo signal acquisition. In at least some embodiments, one or more operations of the method are performed by an acquisition section of an apparatus, such as the apparatus illustrated in FIG. 9 described below. In at least some embodiments, operations S440, S442, and S444 are performed sequentially for audio detections from each audio detector of the apparatus, each audio detection including an echo signal and / or a reflected audio signal captured by the respective audio detector.
[0037] In S440, the acquisition section or a subsection thereof separates out reflections with a time filter. In at least some embodiments, the acquisition section separates out reflections of the UHF audio signal with a time filter. In at least some embodiments, the acquisition section rejects, discards, or ignores data for detections outside a predetermined time frame measured from the time of audio signal emission. In at least some embodiments, the predetermined time frame is calculated to include echo reflections based on the assumption that the surface is 25-50 cm from the device, which is the typical distance of a user's face from a handheld device when the device's camera is pointed at their face.
[0038] At S442, the acquisition section or a subsection thereof compares the audio detection to the emitted audio signal. In at least some embodiments, as the iterations proceed, the acquisition section compares the detection of each audio detector of the plurality of audio detectors to the emitted UHF audio signal. In at least some embodiments, the acquisition section performs template matching to distinguish echo from noise in the detection.
[0039] At S444, the acquisition section or a subsection thereof removes noise from the audio detection. In at least some embodiments, the acquisition section removes noise from reflections of the UHF audio signal. In at least some embodiments, the acquisition section removes noise determined from echoes at S442.
[0040] In S446, the acquisition section or subsection determines whether all detections have been processed. If the acquisition section determines that unprocessed detections remain, the operational flow returns to reflection separation in S440. If the acquisition section determines that all detections have been processed, the operational flow proceeds to merging in S449.
[0041] At S449, the acquisition section, or a subsection thereof, merges the residual data for each sound detection into a single echo signal. In at least some embodiments, the acquisition section sums the residual data for each sound detection. In at least some embodiments, the acquisition section offsets the residual data for each sound detection based on its relative distance from the surface before summation. In at least some embodiments, the acquisition section applies an additional noise reduction process after merging. In at least some embodiments, the acquisition section merges the residual data for each sound detection such that the resulting signal-to-noise ratio is greater than the signal-to-noise ratio of each individual sound detection. In at least some embodiments, the acquisition section detects a time shift or delay between the sound detections to increase the resulting signal-to-noise ratio. In at least some embodiments, the acquisition section obtains cross-correlations between the sound detections to determine a time frame of maximum correlation between the sound detections, shifts the timing of each sound detection to coincide with the determined time frame of maximum correlation, and sums the sound detections.
[0042] 5 is an operational flow of a first feature extraction and classifier application process according to at least some embodiments of the present disclosure. The operational flow provides a first method of feature extraction and classifier application. In at least some embodiments, one or more operations of the method are performed by extraction and application sections of an apparatus, such as the apparatus shown in FIG. 9 described below.
[0043] At S550, the extraction section or a subsection thereof estimates a depth of the surface from the echo signal. In at least some embodiments, extracting a plurality of features from the echo signal includes estimating a depth of the surface from the echo signal. In at least some embodiments, the extraction section performs pseudo-depth estimation. In at least some embodiments, the extraction section estimates the depth of the surface according to the difference between the distance due to the first reflection and the distance due to the last reflection. In at least some embodiments, the distance is calculated as half the delay between the emission of the UHF sound signal and the detection of the reflection multiplied by the speed of sound. In at least some embodiments, the depth D is calculated according to the following formula:
number
[0044] In S551, the application section or a subsection thereof compares the estimated depth to a depth threshold. In at least some embodiments, applying the classifier includes comparing the depth to a depth threshold. In at least some embodiments, the application section determines that the estimated depth is consistent with a live human face in response to the estimated depth being greater than the depth threshold. In at least some embodiments, the application section determines that the estimated depth is inconsistent with a live human face in response to the estimated depth being less than or equal to the depth threshold. In at least some embodiments, the depth threshold is a parameter adjustable by an administrator of the face detection system. In at least some embodiments, the depth threshold is small because depth estimation is only intended to prevent 2D attacks. In at least some embodiments, if the depths of all final reflections compared to the depth of the first reflection are the same, the application section concludes that the surface is a planar 2D surface.
[0045] In S552, the extraction section or a subsection thereof determines an attenuation coefficient of the surface from the echo signal. In at least some embodiments, extracting a plurality of features from the echo signal includes determining an attenuation coefficient of the surface from the echo signal. When the emitted UHF audio signal strikes various surfaces, the signal is absorbed, reflected, and scattered. In particular, signal absorption and scattering result in signal attenuation. Different material properties result in different amounts of signal attenuation. In at least some embodiments, the extraction section determines the attenuation coefficient using the following formula:
number
[0046] At S553, the application section or a subsection thereof compares the determined attenuation coefficient with an attenuation coefficient threshold range. In at least some embodiments, applying the classifier includes comparing the attenuation coefficient with the attenuation coefficient threshold range. In at least some embodiments, the application section determines that the determined attenuation coefficient is consistent with a live human face in response to the determined attenuation coefficient being within the attenuation coefficient threshold range. In at least some embodiments, the application section determines that the determined attenuation coefficient is inconsistent with a live human face in response to the determined attenuation coefficient not being within the attenuation coefficient threshold range. In at least some embodiments, the attenuation coefficient threshold range comprises a parameter adjustable by an administrator of the face detection system. In at least some embodiments, the attenuation coefficient threshold range is small because attenuation coefficients of live human faces have little variability.
[0047] In S554, the extraction section or a subsection thereof estimates a backscattering coefficient of the surface from the echo signal. In at least some embodiments, extracting a plurality of features from the echo signal includes estimating a backscattering coefficient of the surface from the echo signal. The echo signal has backscattering characteristics that vary depending on the material from which the echo signal is reflected. In at least some embodiments, the extraction section estimates the backscattering coefficient to classify the input as being a 3D mask or a real face. The "backscattering coefficient" is a parameter that describes the effectiveness of an object in scattering ultrasound energy. In at least some embodiments, the backscattering coefficient η(w) is obtained from two measurements: the power spectrum of the backscattering signal and the power spectrum of the reflected signal from a flat reference surface previously obtained from a calibration process. The normalized backscattering signal power spectrum of the signal may be given as follows:
number
number
number
[0048] At S555, the application section or a subsection thereof compares the estimated backscatter coefficients to a backscatter coefficient threshold range. In at least some embodiments, applying the classifier includes comparing the backscatter coefficients to a backscatter coefficient threshold range. In at least some embodiments, the application section determines that the estimated backscatter coefficients are consistent with a live human face in response to the estimated backscatter coefficients being within the backscatter coefficient threshold range. In at least some embodiments, the application section determines that the estimated backscatter coefficients are consistent with a live human face in response to the estimated backscatter coefficients not being within the backscatter coefficient threshold range. In at least some embodiments, the backscatter coefficient threshold range comprises a parameter adjustable by an administrator of the face detection system. In at least some embodiments, the backscatter coefficient threshold range is small because backscatter coefficients of live human faces have little variability.
[0049] At S556, the extraction section, or a subsection thereof, applies a neural network to the echo signal to obtain a feature vector. In at least some embodiments, extracting a plurality of features from the echo signal includes applying a neural network to the echo signal to obtain the feature vector, the neural network being trained with a classification layer to classify the echo signal samples as live or non-live. In at least some embodiments, the extraction section applies a convolutional neural network to the echo signal to obtain a deep feature vector. In at least some embodiments, the neural network is trained to output the feature vector when applied to the echo signal. In at least some embodiments, the neural network undergoes a training process described below with reference to FIG. 6.
[0050] At S557, the applying section or a subsection thereof applies a classification layer to the feature vector. In at least some embodiments, applying the classifier includes applying the classification layer to the feature vector. In at least some embodiments, the applying section determines that the echo signal is consistent with a live human face in response to a first output value from the classification layer. In at least some embodiments, the applying section determines that the echo signal is consistent with a live human face in response to a second output value from the classification layer. In at least some embodiments, the classification layer is an anomaly detection classifier. In at least some embodiments, the classification layer is trained to output a binary value when applied to the feature vector, the binary value representing whether the echo signal is consistent with a live human face. In at least some embodiments, the classification layer undergoes a training process described below with reference to FIG. 6.
[0051] At S558, the application section or a subsection thereof weights the results for whether each feature is consistent with a live human face. In at least some embodiments, the application section applies a weight to each result proportional to the strength of the feature as a determining factor for whether the surface is a live face. In at least some embodiments, the weight is a parameter that is adjustable by an administrator of the facial recognition system. In at least some embodiments, the weight is a trainable parameter.
[0052] At S559, the application section compares the weighted results to a liveness threshold. In at least some embodiments, the weighted results are summed and compared to a single liveness threshold. In at least some embodiments, the weighted results undergo more complex calculations before comparison with the liveness threshold. In at least some embodiments, the feature vector obtained from the CNN (convolutional neural network) is combined with the handcrafted features to obtain a final score for determining whether the surface is a live human face. In at least some embodiments, the application section determines that the surface is a live human face in response to the sum of the weighted results being greater than the liveness threshold. In at least some embodiments, the application section determines that the surface is not a live human face in response to the sum of the weighted results being less than or equal to the liveness threshold.
[0053] 6 is an operational flow for training a neural network for feature extraction according to at least some embodiments of the present disclosure. The operational flow provides a method for training a neural network for feature extraction. In at least some embodiments, one or more operations of the method are performed by a controller of an apparatus such as the apparatus illustrated in FIG. 9 described below.
[0054] In S660, the emitting section emits a liveness detection audio signal. In at least some embodiments, the emitting section emits an ultra-high frequency (UHF) audio signal via a speaker. In at least some embodiments, the emitting section emits the liveness detection audio signal similar to S330 during the liveness detection process of FIG. 3. In at least some embodiments, the emitting section emits the liveness detection audio signal toward a surface known to be a live human face or toward a non-live surface such as a 3D mask, 2D print, or screen indicative of a replay attack.
[0055] In S661, an acquisition section acquires echo signal samples. In at least some embodiments, the acquisition section acquires echo signal samples by detecting reflections of UHF audio signals from surfaces using multiple audio detectors. In at least some embodiments, the acquisition section acquires echo signal samples similar to S332 during the liveness detection process of FIG. 3. In at least some embodiments, the captured and processed echo signals are used to train a single-class classifier to acquire CNN (convolutional neural network) features of real faces, such that any distribution other than a real face is considered anomalous and classified as not real.
[0056] In S663, the extraction section applies a neural network to obtain a feature vector. In at least some embodiments, the extraction section applies the neural network to the echo signal samples to obtain the feature vector. In the first iteration of applying the neural network in S663, the neural network is initialized as random values in at least some embodiments. As a result, the obtained feature vector may not be very relevant for determining liveness. As the iterations progress, the weights of the neural network are adjusted so that the obtained feature vector becomes more relevant for determining liveness.
[0057] At S664, an application section applies a classification layer to the feature vector to determine a classification of the surface. In at least some embodiments, the classification layer is a binary classifier that generates either a classification indicating that the feature vector is consistent with a live human face or a classification indicating that the feature vector is inconsistent with a live human face.
[0058] At S666, the application section adjusts the parameters of the neural network and the classification layer. In at least some embodiments, the application section adjusts the weights of the neural network and the classification layer according to a loss function based on whether the classification determined by the classification layer at S664 in light of known information about whether the surface is a live human face is correct. In at least some embodiments, the training includes adjusting the parameters of the neural network and the classification layer based on a comparison of the output classification and the corresponding label. In at least some embodiments, gradients of the weights are calculated from the output layer of the classification layer back through the neural network through a process of backpropagation, and the weights are updated according to the newly calculated gradients. In at least some embodiments, the parameters of the neural network are not adjusted after each iteration of the operations at S663 and S664. In at least some embodiments, as the iterations progress, the controller trains the neural network together with the classification layer using multiple echo signal samples, each of which is labeled as live or non-live.
[0059] In S668, the controller or a section thereof determines whether all echo signal samples have been processed. In at least some embodiments, the controller determines that all samples have been processed in response to the batch of echo signal samples being completely processed, or in response to some other termination condition, such as the neural network solution converging, the loss value of a loss function falling below a threshold, etc. If the controller determines that unprocessed echo signal samples remain, or if another termination condition has not yet been met, the operational flow returns to the signal emission in S660 of the next sample (S669). If the controller determines that all echo signal samples have been processed, or if another termination condition has been met, the operational flow ends.
[0060] In at least some embodiments, signal emission in S660 and echo signal acquisition in S661 are performed on a batch of samples before proceeding to repeat the operations in S663, S664, and S666.
[0061] 7 is an operational flow of a second feature extraction and classifier application process according to at least some embodiments of the present disclosure. The operational flow provides a second method of feature extraction and classifier application. In at least some embodiments, one or more operations of the method are performed by extraction and application sections of an apparatus, such as the apparatus shown in FIG. 9 described below.
[0062] In S770, the extraction section or a subsection thereof estimates the depth of the surface from the echo signals. The depth estimation in S770 is substantially similar to the depth estimation in S550 of FIG. 5, except as noted differently.
[0063] In S772, the extraction section or a subsection thereof determines the attenuation coefficient of the surface from the echo signal. The attenuation coefficient determination in S772 is substantially similar to the attenuation coefficient determination in S552 of FIG. 5, except as noted differently.
[0064] In S774, the extraction section or a subsection thereof estimates the backscattering coefficients of the surface from the echo signals. The backscattering coefficient estimation in S774 is substantially similar to the backscattering coefficient estimation in S554 of FIG. 5, except as noted differently.
[0065] In S776, the extraction section or a subsection thereof applies a neural network to the echo signal to obtain a feature vector. The application of the neural network in S776 is substantially similar to the application of the neural network in S556 of FIG. 5, except as noted differently.
[0066] In at least some embodiments, the extraction section performs operations S770, S772, S774, and S776 to extract features from the echo signals. In at least some embodiments, extracting a plurality of features from the echo signals includes estimating a depth of a surface from the echo signals, determining an attenuation coefficient of the surface from the echo signals, estimating a backscattering coefficient of the surface from the echo signals, and applying a neural network to the echo signals to obtain a feature vector, wherein the neural network is trained with a classification layer to classify the echo signal samples as live or non-live.
[0067] In S778, the extraction section or a subsection thereof merges the features. In at least some embodiments, the extraction section merges the estimated depth from S772, the determined attenuation coefficients from S774, the estimated backscattering coefficients from S774, and the feature vector from S776. In at least some embodiments, the extraction section concatenates the features into a single string, thereby increasing the number of features included in the feature vector.
[0068] In S779, the applying section or a subsection thereof applies a classifier to the merged features. In at least some embodiments, applying the classifier includes applying a classifier to the feature vector, depth, attenuation coefficients, and backscatter coefficients, where the classifier is trained to classify echo signal-extracted feature samples as live or non-live. In at least some embodiments, the classifier is applied to a concatenation of the features. In at least some embodiments, the classifier is an anomaly detection classifier. In at least some embodiments, the classifier is trained to output a binary value when applied to the merged features, where the binary value represents whether the echo signal is consistent with a live human face. In at least some embodiments, the classifier undergoes a training process similar to the training process of FIG. 6 , except that the classifier training process includes training a neural network with the classifier using multiple echo signal-extracted feature samples, where each echo signal-extracted feature sample is labeled as live or non-live, and the training step includes adjusting parameters of the neural network and the classifier based on a comparison of the output classification with the corresponding label.
[0069] 8 is a schematic diagram of a third feature extraction and classifier application process according to at least some embodiments of the present disclosure, including an echo signal 892, a depth estimation section 884A, an attenuation coefficient determination section 884B, a backscattering coefficient estimation section 884C, a convolutional neural network 894A, a classifier 894B, a depth estimation 896A, an attenuation coefficient determination 896B, a backscattering coefficient estimation 896C, a feature vector 896D, and a classification 898.
[0070] Echo signal 892 is input to depth estimation section 884A, attenuation coefficient determination section 884B, backscatter coefficient estimation section 884C, and convolutional neural network 894A. In response to input of echo signal 892, depth estimation section 884A outputs depth estimate 896A, attenuation coefficient determination section 884B outputs attenuation coefficient determination 896B, backscatter coefficient estimation section 884C outputs backscatter coefficient estimate 896C, and convolutional neural network 894A outputs feature vector 896D.
[0071] In at least some embodiments, the depth estimate 896A, the attenuation coefficient determination 896B, and the backscatter coefficient estimate 896C are actual values without normalization or comparison to a threshold. In at least some embodiments, the depth estimate 896A, the attenuation coefficient determination 896B, and the backscatter coefficient estimate 896C are normalized values. In at least some embodiments, the depth estimate 896A, the attenuation coefficient determination 896B, and the backscatter coefficient estimate 896C are binary values that represent the results of comparison to respective thresholds, such as those described with respect to FIG. 5.
[0072] The depth estimate 896A, the attenuation coefficient determination 896B, the backscattering coefficient estimate 896C, and the feature vector 896D are combined to form an input to the classifier 894B. In at least some embodiments, the depth estimate 896A, the attenuation coefficient determination 896B, the backscattering coefficient estimate 896C, and the feature vector 896D are concatenated into a single string of features for input to the classifier 894B.
[0073] Classifier 894B is trained to respond to the feature inputs and output a classification 898. Classification 898 indicates whether the surface associated with the echo signal is consistent with a live human face.
[0074] FIG. 9 is a block diagram of a hardware configuration for face liveness detection, according to at least some embodiments of the present disclosure.
[0075] An exemplary hardware configuration includes device 900 that interacts with microphones 910 / 911, speakers 913, cameras 915, and tactile input 919 and communicates with network 907. In at least some embodiments, device 900 incorporates microphones 910 / 911, speakers 913, cameras 915, and tactile input 919. In at least some embodiments, device 900 is a computer system that executes computer-readable instructions to perform operations for physical network function device access.
[0076] The device 900 comprises a controller 902, a storage unit 904, a communication interface 906, and an input / output interface 908. In at least some embodiments, the controller 902 includes a processor or programmable circuit that executes instructions to cause the processor or programmable circuit to perform operations in accordance with the instructions. In at least some embodiments, the controller 902 includes analog or digital programmable circuitry, or any combination thereof. In at least some embodiments, the controller 902 comprises physically separate storage or circuitry that interacts via communications. In at least some embodiments, the storage unit 904 includes a non-volatile computer-readable medium capable of storing executable and non-executable data for access by the controller 902 during execution of instructions. The communication interface 906 transmits and receives data to and from a network 907. The input / output interface 908 connects to and exchanges information with various input / output units, such as microphones 910 / 911, speakers 913, cameras 915, and tactile inputs 919, via parallel ports, serial ports, keyboard ports, mouse ports, monitor ports, etc.
[0077] The controller 902 comprises an emission section 980, an acquisition section 982, an extraction section 984, and an application section 986. The storage unit 904 includes detections 990, echo signals 992, extracted features 994, neural network parameters 996, and classification results 998.
[0078] The emission section 980 is circuitry or instructions of the controller 902 configured to emit a liveness detection audio signal. In at least some embodiments, the emission section 980 is configured to emit an ultra-high frequency (UHF) audio signal via a speaker. In at least some embodiments, the emission section 980 includes subsections for performing additional functions, as described in the flowcharts above. In at least some embodiments, such subsections are referred to by names associated with their corresponding functions.
[0079] Acquisition section 982 is circuitry or instructions in controller 902 configured to acquire echo signals. In at least some embodiments, acquisition section 982 is configured to acquire echo signals by detecting reflections of UHF audio signals from a surface with a plurality of audio detectors. In at least some embodiments, acquisition section 982 records information, such as detections 990 and echo signals 992, in storage unit 904. In at least some embodiments, acquisition section 982 includes subsections for performing additional functions, as described in the flowcharts above. In at least some embodiments, such subsections are referred to by names associated with their corresponding functions.
[0080] The extraction section 984 is circuitry or instructions in the controller 902 configured to extract features. In at least some embodiments, the extraction section 984 is configured to extract multiple features from the echo signals. In at least some embodiments, the extraction section 984 utilizes information from the storage unit 904, such as the echo signals 992 and neural network parameters 996, and records information, such as the extracted features 994, in the storage unit 904. In at least some embodiments, the extraction section 984 includes subsections for performing additional functions, as described in the flowcharts above. In at least some embodiments, such subsections are referred to by names associated with their corresponding functions.
[0081] The apply section 986 is circuitry or instructions in the controller 902 configured to apply a classifier to the features. In at least some embodiments, the apply section 986 is configured to apply the classifier to a plurality of features to determine whether a surface is a live face. In at least some embodiments, the apply section 986 utilizes information from the storage unit 904, such as extracted features 994 and neural network parameters 996, and records information, such as classification results 998, in the storage unit 904. In at least some embodiments, the extract section 984 includes subsections for performing additional functions, as described in the flowcharts above. In at least some embodiments, such subsections are referred to by names associated with their corresponding functions.
[0082] In at least some embodiments, the apparatus is a separate device capable of processing logical functions to perform the operations herein. In at least some embodiments, the controller and storage unit need not be entirely separate devices, but in some embodiments share circuitry or one or more computer-readable media. In at least some embodiments, the storage unit includes a hard drive that stores both computer-executable instructions and data accessed by the controller, and the controller includes a combination of a central processing unit (CPU) and RAM, where the computer-executable instructions can be copied, in whole or in part, for execution by the CPU during performance of the operations herein.
[0083] In at least some embodiments where the device is a computer, a program installed on the computer can cause the computer to function as or perform operations associated with the devices of the embodiments described herein, and in at least some embodiments, such a program can be executed by a processor to cause the computer to perform specific operations associated with some or all of the blocks in the flowcharts and block diagrams described herein.
[0084] At least some embodiments are described with reference to flowcharts and block diagrams, where the blocks represent (1) steps in a process in which an operation is performed or (2) sections of a controller responsible for performing an operation. In at least some embodiments, particular steps and sections are implemented by dedicated circuitry, programmable circuitry provided with computer-readable instructions stored on a computer-readable medium, and / or a processor provided with computer-readable instructions stored on a computer-readable medium. In at least some embodiments, dedicated circuitry includes digital and / or analog hardware circuitry, including integrated circuits (ICs) and / or discrete circuits. In at least some embodiments, programmable circuitry includes reconfigurable hardware circuitry, such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), etc., including logical AND, OR, XOR, NAND, NOR, and other logic operations, flip-flops, registers, memory elements, etc.
[0085] In at least some embodiments, a computer-readable storage medium comprises a tangible device capable of holding and storing instructions for use by an instruction-execution device. In some embodiments, a computer-readable storage medium comprises, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or ridge structures in grooves having instructions recorded thereon, and any suitable combination of the above. As used herein, computer-readable storage media should not be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over wires.
[0086] In at least some embodiments, the computer-readable program instructions described herein are downloadable to each computing / processing device from a computer-readable storage medium or via an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. In at least some embodiments, the network includes copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. In at least some embodiments, a network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within the respective computing / processing device.
[0087] In at least some embodiments, the computer-readable program instructions for performing the operations described above are either source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and traditional procedural programming languages such as the "C" programming language or similar programming languages. In at least some embodiments, the computer-readable program instructions execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In at least some embodiments, in the latter scenario, the remote computer is connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or is connected to an external computer (e.g., via the Internet using an Internet Service Provider). In at least some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) executes computer-readable program instructions by utilizing state information in the computer-readable program instructions to individualize the electronic circuitry to perform aspects of the present disclosure.
[0088] Although the embodiments of the present disclosure have been described above, the technical scope of the subject matter described in the claims is not limited to the above embodiments. Those skilled in the art will understand that various modifications and improvements to the above-described embodiments are possible. Furthermore, those skilled in the art will understand from the claims that embodiments with such modifications or improvements are also included in the technical scope of the present invention.
[0089] The operations, procedures, steps, and stages of each process performed by the devices, systems, programs, and methods described in the claims, embodiments, or drawings are not specifically stated as "before" or "prior to," and can be performed in any order unless the output of a previous process is used in a subsequent process. Even if the flow of a process is described using words such as "first" or "next" in the claims, embodiments, or drawings, such a description does not necessarily mean that the process must be performed in the order described.
[0090] In accordance with at least some embodiments of the present disclosure, face liveness is detected by emitting an ultra-high frequency (UHF) audio signal via a speaker, obtaining an echo signal by detecting reflections of the UHF audio signal from a surface with a plurality of audio detectors, extracting a plurality of features from the echo signal, and applying a classifier to the plurality of features to determine whether the surface is a live face.
[0091] Some embodiments include instructions in a computer program, a method performed by a processor executing the computer program instructions, and an apparatus for performing the method. In some embodiments, the apparatus comprises a controller including circuitry configured to perform the operations in the instructions.
[0092] The foregoing outlines features of several embodiments so that those skilled in the art may better understand aspects of the present disclosure. Those skilled in the art will readily appreciate that this disclosure may be used as a basis for designing or modifying other processes and structures which carry out the same purposes and / or achieve the same advantages as the embodiments presented herein. Those skilled in the art will also appreciate that such equivalent constructions do not depart from the spirit and scope of the present disclosure, and that various changes, substitutions, and alterations can be made herein without departing from the spirit and scope of the present disclosure.
[0093] Some or all of the above exemplary embodiments may also be described as follows, but are not limited to these.
[0094] (Appendix 1) emitting an ultra-high frequency (UHF) audio signal via a speaker; obtaining echo signals by detecting reflections of the UHF audio signal from a surface using a plurality of audio detectors; extracting a plurality of feature amounts from the echo signal; applying a classifier to the plurality of features to determine whether the surface is a live face; and A computer-readable medium comprising instructions executable by a computer to cause the computer to perform operations including:
[0095] (Appendix 2) extracting the plurality of features from the echo signals includes applying a neural network to the echo signals to obtain a feature vector, the neural network being trained with a classification layer to classify echo signal samples as live or non-live; applying the classifier includes applying the classification layer to the feature vector. 2. The computer-readable medium of claim 1.
[0096] (Appendix 3) training the neural network together with the classification layer using a plurality of echo signal samples, each of which is labeled as live or non-live; further comprising 3. The computer-readable medium of claim 2, wherein the training comprises adjusting parameters of the neural network and the classification layer based on a comparison of the output classifications and corresponding labels.
[0097] (Appendix 4) extracting the plurality of feature amounts from the echo signal, estimating the depth of the surface from the echo signals; determining an attenuation coefficient of the surface from the echo signals; estimating a backscattering coefficient of the surface from the echo signals; applying a neural network to the echo signals to obtain a feature vector, the neural network being trained with a classification layer to classify the echo signal samples as live or non-live; applying the classifier includes applying the classifier to the feature vector, the depth, the attenuation coefficient, and the backscatter coefficient, the classifier being trained to classify echo signal extracted feature samples as live or non-live. 2. The computer-readable medium of claim 1.
[0098] (Appendix 5) training the neural network in the classifier using a plurality of echo signal extracted feature samples, each labeled as live or non-live; further comprising 5. The computer-readable medium of claim 4, wherein the training includes adjusting parameters of the neural network and the classifier based on a comparison of the output classifications and corresponding labels.
[0099] (Appendix 6) extracting the plurality of features from the echo signals includes estimating a depth of the surface from the echo signals; applying the classifier includes comparing the depth to a depth threshold. 2. The computer-readable medium of claim 1.
[0100] (Appendix 7) extracting the plurality of features from the echo signals includes determining an attenuation coefficient of the surface from the echo signals; applying the classifier includes comparing the attenuation coefficient to an attenuation coefficient threshold range. 2. The computer-readable medium of claim 1.
[0101] (Appendix 8) extracting the plurality of features from the echo signals includes estimating a backscattering coefficient of the surface from the echo signals; applying the classifier includes comparing the backscatter coefficients to a backscatter coefficient threshold range. 2. The computer-readable medium of claim 1.
[0102] (Appendix 9) 10. The computer-readable medium of claim 1, wherein the plurality of audio detectors includes a first audio detector oriented in a first direction and a second audio detector oriented in a second direction.
[0103] (Appendix 10) the plurality of audio detectors include a plurality of microphones; 10. The computer-readable medium of claim 1, wherein the speaker and the plurality of microphones are included in a handheld device.
[0104] (Appendix 11) the handheld device further comprises a camera; the UHF audio signal is emitted in response to face detection by the camera; 11. The computer-readable medium of claim 10.
[0105] (Appendix 12) 2. The computer-readable medium of claim 1, wherein the UHF audio signal is 18 to 22 kHz.
[0106] (Appendix 13) 10. The computer-readable medium of claim 1, wherein the UHF audio signal is substantially inaudible.
[0107] (Appendix 14) 10. The computer-readable medium of claim 1, wherein the UHF audio signal includes a sine wave and a sawtooth wave.
[0108] (Appendix 15) imaging the surface with a camera to obtain a surface image; analyzing the surface image to determine whether the surface is a face; determining that the surface is a face and, in response to determining that the surface is a live face, identifying the surface by analyzing the surface image; granting access to at least one of a device or a service in response to identifying the surface as an authorized user; 2. The computer-readable medium of claim 1, further comprising:
[0109] (Appendix 16) acquiring the echo signals, isolating the reflections of the UHF audio signal with a time filter; removing noise from the reflections of the UHF audio signal by comparing the detection of each of the plurality of audio detectors with the emitted UHF audio signal; 2. The computer-readable medium of claim 1, comprising:
[0110] (Appendix 17) emitting an ultra-high frequency (UHF) audio signal via a speaker; obtaining echo signals by detecting reflections of the UHF audio signal from a surface using a plurality of audio detectors; extracting a plurality of feature amounts from the echo signal; applying a classifier to the plurality of features to determine whether the surface is a live face; and A method comprising:
[0111] (Appendix 18) extracting the plurality of features from the echo signals includes applying a neural network to the echo signals to obtain a feature vector, the neural network being trained with a classification layer to classify echo signal samples as live or non-live; applying the classifier includes applying the classification layer to the feature vector. The method described in Appendix 17.
[0112] (Appendix 19) a plurality of audio detectors; A speaker and emitting an ultra-high frequency (UHF) audio signal through said speaker; obtaining echo signals by detecting reflections of the UHF audio signals from a surface using a plurality of audio detectors; extracting a plurality of feature amounts from the echo signal; Applying a classifier to the plurality of features to determine whether the surface is a live face. a controller including a circuit configured as follows: An apparatus comprising:
[0113] (Appendix 20) the circuitry configured to extract the plurality of features from the echo signals includes applying a neural network to the echo signals to obtain a feature vector, the neural network being trained with a classification layer to classify echo signal samples as live or non-live; the circuitry configured to apply the classifier is further configured to apply the classification layer to the feature vector. 19. The apparatus of claim 19.
[0114] This application is based on and claims the benefit of priority from U.S. patent application Ser. No. 17 / 720,225, filed April 13, 2022. Claim No. 60 / 699,999, filed on Oct. 1, 2007, the disclosure of which is incorporated herein in its entirety.
Claims
1. emitting an ultra-high frequency (UHF) audio signal via a speaker; obtaining echo signals by detecting reflections of the UHF audio signal from a surface using a plurality of audio detectors; applying a neural network to the echo signals to obtain a plurality of feature vectors; applying a classification layer to the plurality of feature vectors to determine whether the surface is a live face; training the neural network with the classification layer using a plurality of echo signal samples, each of which is labeled as live or non-live; A program for causing a computer to execute an operation including: The training includes adjusting parameters of the neural network and the classification layer based on a comparison of the output classifications and corresponding labels.
2. Emitting an ultra-high frequency (UHF) audio signal via a speaker; obtaining echo signals by detecting reflections of the UHF audio signal from a surface using a plurality of audio detectors; extracting a plurality of feature amounts from the echo signal; applying a classifier to the plurality of features to determine whether the surface is a live face; and A program for causing a computer to execute an operation including: extracting the plurality of feature amounts from the echo signal, estimating the depth of the surface from the echo signals; determining an attenuation coefficient of the surface from the echo signals; estimating a backscattering coefficient of the surface from the echo signals; applying a neural network to the echo signals to obtain a feature vector, the neural network being trained with a classification layer to classify the echo signal samples as live or non-live; applying the classifier includes applying the classifier to the feature vector, the depth, the attenuation coefficients, and the backscatter coefficients, the classifier being trained to classify echo signal extracted feature samples as live or non-live. program.
3. training the neural network in the classifier using a plurality of echo signal extracted feature samples, each labeled as live or non-live; further comprising The program of claim 2 , wherein the training comprises adjusting parameters of the neural network and the classifier based on a comparison of output classifications and corresponding labels.
4. Emitting an ultra-high frequency (UHF) audio signal via a speaker; obtaining echo signals by detecting reflections of the UHF audio signal from a surface using a plurality of audio detectors; extracting a plurality of feature amounts from the echo signal; applying a classifier to the plurality of features to determine whether the surface is a live face; and A program for causing a computer to execute an operation including: extracting the plurality of features from the echo signals includes estimating a depth of the surface from the echo signals; applying the classifier includes comparing the depth to a depth threshold. program.
5. Emitting an ultra-high frequency (UHF) audio signal via a speaker; obtaining echo signals by detecting reflections of the UHF audio signal from a surface using a plurality of audio detectors; extracting a plurality of feature amounts from the echo signal; applying a classifier to the plurality of features to determine whether the surface is a live face; and A program for causing a computer to execute an operation including: extracting the plurality of features from the echo signals includes determining an attenuation coefficient of the surface from the echo signals; applying the classifier includes comparing the attenuation coefficient to an attenuation coefficient threshold range. program.
6. Emitting an ultra-high frequency (UHF) audio signal via a speaker; obtaining echo signals by detecting reflections of the UHF audio signal from a surface using a plurality of audio detectors; extracting a plurality of feature amounts from the echo signal; applying a classifier to the plurality of features to determine whether the surface is a live face; and A program for causing a computer to execute an operation including: extracting the plurality of features from the echo signals includes estimating a backscattering coefficient of the surface from the echo signals; applying the classifier includes comparing the backscatter coefficients to a backscatter coefficient threshold range. program.
7. emitting an ultra-high frequency (UHF) audio signal via a speaker; obtaining echo signals by detecting reflections of the UHF audio signal from a surface using a plurality of audio detectors; applying a neural network to the echo signals to obtain a plurality of feature vectors; applying a classification layer to the plurality of feature vectors to determine whether the surface is a live face; training the neural network with the classification layer using a plurality of echo signal samples, each of which is labeled as live or non-live; A method comprising: The method, wherein the training includes adjusting parameters of the neural network and the classification layer based on a comparison of output classifications and corresponding labels.
8. a plurality of audio detectors; A speaker and emitting an ultra-high frequency (UHF) audio signal through said speaker; obtaining echo signals by detecting reflections of the UHF audio signals from a surface using a plurality of audio detectors; applying a neural network to the echo signals to obtain a plurality of feature vectors; applying a classification layer to the plurality of feature vectors to determine whether the surface is a live face; training the neural network with the classification layer using a plurality of echo signal samples, each of which is labeled as live or non-live; a controller including a circuit configured as follows: An apparatus comprising: The apparatus, wherein the training includes adjusting parameters of the neural network and the classification layer based on a comparison of the output classifications and corresponding labels.
Citation Information
Patent Citations
Living body detection method and device thereof, medium and computer program product
CN113100734A
Self-hardening anti-breaking cable based on power grid power transmission
CN113506652A
Systems and Methods for Spoof Detection and Liveness Analysis
JP2018524072A
Personal authentication device, personal authentication method, and recording medium
WO2018198310A1