Method, electronic apparatus, and non-transitory computer-readable medium for face authentication Anti-spoofing using interferometry-based coherence
Patent Information
- Application Number
- TW111134343
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-09-12
- Publication Date
- 2026-07-21
- Estimated Expiration
- 2042-09-11
AI Technical Summary
Face authentication systems using cameras are vulnerable to spoofing attacks, such as those involving images or masks, which can compromise security and access control.
Utilizing ultrasonic sensors that employ interferometry to measure coherence between reflections from a face and a potential spoofing attempt, distinguishing between a real face and a simulated one by analyzing differences in ultrasonic energy reflections.
Effectively prevents unauthorized access by accurately differentiating between a genuine face and spoofing attempts, enhancing security without increasing latency.
Abstract
Description
[Previous Technology]
[0001] Facial recognition provides users with a convenient way to unlock their devices, increase the security of accessing accounts, or sign transactions, thus enhancing the user experience. Some facial recognition systems rely on cameras for facial authentication. However, it can be challenging to distinguish a camera from a display of an image of that user's face. Therefore, preventing unauthorized actors from deceiving camera-based facial recognition systems presents challenges. [Summary of the Invention]
[0002] This invention describes a technology and apparatus for implementing face authentication anti-spoofing using interferometric coherence. Specifically, a face authentication system uses ultrasound to distinguish a real face from a demo attack that uses an instrument to demonstrate a version of the face. An example instrument may include a piece of paper with a photograph of a user, a screen displaying a digital image of the user, or a mask that partially replicates the user's face. The face authentication system includes or communicates with an ultrasound sensor that can detect a demo attack and notify the face authentication system. Generally, the ultrasound sensor uses interferometry to evaluate a quantity of coherence (or similarity) between reflections observed by two or more sensors. By using interferometric coherence, the ultrasound sensor can distinguish the demo attack instrument from the face for face authentication anti-spoofing. Using these techniques, the ultrasound sensor can prevent unauthorized actors from using the demo attack to access a user's account or information.
[0003] The following description includes a method for anti-spoofing facial authentication using coherence, performed by an ultrasound sensor. The method includes transmitting an ultrasound transmission signal and receiving at least two ultrasound reception signals using at least two sensors of the ultrasound sensor. The at least two ultrasound reception signals include respective versions of the ultrasound transmission signal reflected by an object. The method also includes generating an interferogram based on the at least two ultrasound reception signals. The interferogram includes coherence information and phase information. The method further includes identifying a coherence feature based on the coherence information of the interferogram. The coherence feature represents a coherence quantity within a region of interest. The method further includes detecting a demo attack based on the coherence feature. The demo attack attempts to deceive a facial authentication system, and the object is associated with the demo attack. The method also includes preventing the facial authentication system from authenticating the demo attack.
[0004] The state described below includes a device having an ultrasonic sensor configured to perform any of the described methods.
[0005] The manner described below also includes a computer-readable medium comprising instructions that, when executed by a processor, cause an ultrasonic sensor to perform any of the described methods.
[0006] The state described below also includes a system having components for providing anti-counterfeiting facial authentication using coherence based on interferometry.
Implementation Method
[0008] Overview
[0009] Facial recognition provides users with a convenient way to unlock their devices, increase the security of accessing accounts, or sign transactions, thus enhancing the user experience. Some facial recognition systems rely on cameras for facial authentication. However, it can be challenging to distinguish a camera from a display of an image of that user's face. Therefore, preventing unauthorized actors from deceiving camera-based facial recognition systems presents challenges.
[0010] Some technologies distinguish between a real face and an image of a face by detecting activity (e.g., blinking or facial movement). However, such technologies may require additional acquisition time, thereby increasing the latency involved in performing facial authentication. Alternatively, these technologies can be overcome by using video recordings of a face that includes such movements.
[0011] To address these issues, this document describes a technology and apparatus for implementing anti-spoofing facial authentication using ultrasound. Specifically, a facial authentication system uses ultrasound to distinguish a real face from a demo attack that uses an instrument to demonstrate a version of the face. An example instrument may include a piece of paper with a photograph of a user, a screen displaying a digital image of the user, or a mask that partially replicates the user's face. The facial authentication system includes or communicates with an ultrasound sensor that can detect a demo attack and notify the facial authentication system. Generally, the ultrasound sensor uses interferometry to assess a quantity of coherence (or similarity) between reflections observed by two or more sensors. By using interferometric coherence, the ultrasound sensor can distinguish the demo attack instrument from the face for anti-spoofing facial authentication. Using these techniques, the ultrasound sensor can prevent unauthorized actors from using the demo attack to access a user's account or information. Instance environment
[0012] Figure 1 is a diagram illustrating exemplary environments 100-1 to 100-5, which may demonstrate the use of face authentication anti-counterfeiting technology based on interferometric coherence. In environments 100-1 to 100-5, a user device 102 performs face authentication 104. In environments 100-1 and 100-2, a user 106 controls the user device 102 and uses face authentication 104 to (e.g.) access an application or sign (electronically) a transaction. During face authentication 104, the user device 102 uses a camera system to capture one or more images 108. In this case, image 108 contains the face of user 106. Using face recognition technology, user device 102 identifies the user's face from the image and authenticates user 106.
[0013] In some cases, user 106 may wear an accessory 110 (e.g., a hat, a scarf, a headband, glasses, or jewelry), which may make it more challenging for user device 102 to identify user 106. In this scenario, user device 102 may use ultrasound to determine the absence of a demonstration attack and confirm the presence of a real face. Using ultrasound, user device 102 can successfully perform face authentication 104 in situations where user 106 is not wearing an accessory 110 (such as in environment 100-1) and in situations where user 106 chooses to wear an accessory 110 (such as in environment 100-2).
[0014] In other scenarios, an unauthorized actor 112 may control the user device 102. In this case, the unauthorized actor 112 may use various techniques to attempt to deceive (e.g., trick) the user device 102 into granting the unauthorized actor 112 access to the user 106's account or information. Environments 100-3, 100-4, and 100-5 provide examples of three different demonstration attacks 114-1, 114-2, and 114-3, respectively.
[0015] In environment 100-3, during face authentication 104, an unauthorized actor 112 presents a medium 116 with a photograph 118 of a user 106. The medium 116 may be a wood-based medium (e.g., paper, cardboard, or poster paper), a plastic-based medium (e.g., acrylic board), a cloth-based medium (e.g., a cotton fiber or polyester fabric), a glass plate, etc. In some cases, the medium 116 presents a relatively flat surface on which a photograph is attached, printed, or engraved. In other cases, the medium 116 may be bent or attached to an object having a curved surface. During face authentication 104, the unauthorized actor 112 orients the medium 116 toward a camera of a user device 102 to cause the user device 102 to capture an image 108 of the photograph 118 presented on the medium 116.
[0016] In environment 100-4, an unauthorized actor 112 presents a device 120 with a display 122 during face authentication 104. Device 120 may be a smartphone, a tablet, a wearable device, a television, a curved monitor, or a virtual or augmented reality headset. Display 122 may be a light-emitting diode (LED) display or a liquid crystal display (LCD). On display 122, device 120 displays a digital image 124 of user 106. During face authentication 104, unauthorized actor 112 orients the display 122 of device 120 toward the camera of user device 102 to cause user device 102 to capture an image 108 of the digital image 124 displayed on display 122.
[0017] In environment 100-5, an unauthorized actor 112 wears a mask 126 that replicates one or more features of the user 106's face in a certain way. For example, one color of the mask 126 may be approximately similar to the skin tone of the user 106, or the mask 126 may contain structural features representing the structure of the user's chin or cheekbones. The mask 126 may be a rigid plastic mask or a latex mask. During face authentication 104, the unauthorized actor 112 wears the mask 126 and faces the camera of the user device 102 to cause the user device 102 to capture an image 108 of the mask 126.
[0018] The demonstration attacks 114-1 to 114-3 shown in environments 100-3 to 100-5 can deceive some authentication systems, which can then grant access to unauthorized actors 112. However, by using ultrasound facial recognition anti-spoofing technology, user device 102 detects demonstration attacks 114-1 to 114-3 and denies access to unauthorized actors 112.
[0019] Using the described technique, user device 102 can distinguish between a face (e.g., a face of user 106) and a representation of a face demonstrated by a demonstration attack 114 (e.g., a photograph 118, a digital image 124, or a mask 126). Specifically, user device 102 uses ultrasound to detect differences in ultrasound reflection between a face and a representation of a face. For example, a face is composed of human tissue, which is less reflective than other types of materials (such as plastic). Furthermore, a face has curves and angles, which reduces the amount of ultrasound energy directly reflected back to user device 102. In contrast, demonstration attack instruments with relatively flat or uniformly curved surfaces (such as media 116, a display 122, or a rigid type of mask 126) can increase the amount of ultrasound energy reflected back to user device 102. By analyzing the characteristics of reflected ultrasound energy, user device 102 can determine whether to demonstrate a face or use an instrument in a demonstration attack 114 for face authentication 104. User device 102 is further described with reference to Figure 2-1. Example Face Authentication System
[0020] Figure 2-1 illustrates a face authentication system 202 as part of user device 102. User device 102 is illustrated to have various non-limiting exemplary devices, including a desktop computer 102-1, a tablet computer 102-2, a laptop computer 102-3, a television set 102-4, an arithmetic watch 102-5, a computing glasses 102-6, a gaming system 102-7, a microwave oven 102-8, and a vehicle 102-9. Other devices may also be used, including a home service device, a smart speaker, a smart thermostat, a security camera, a baby monitor, a Wi-Fi® router, a drone, a trackpad, a drawing tablet, a mini-notebook computer, an e-reader, a home automation and control system, a wall-mounted monitor, a virtual reality headset, and another home appliance. It should be noted that the user device 102 may be wearable, non-wearable but mobile, or relatively fixed (e.g., desktop computer and electrical appliance).
[0021] The user device 102 includes one or more computer processors 204 and one or more computer-readable media 206 including memory media and storage media. Applications and / or an operating system (not shown) embodied in computer-readable instructions on the computer-readable media 206 can be executed by the computer processor 204 to provide some or all of the functionality described herein. The computer-readable media 206 also includes an application 208 or settings launched in response to the facial recognition system 202 authenticating the user 106. Instance application 208 may include a password storage application, a banking application, a wallet application, a health application, or any application that provides user privacy.
[0022] User device 102 may also include a network interface 210 for transmitting data via a wired, wireless, or optical network. For example, network interface 210 may transmit data via a local area network (LAN), a wireless local area network (WLAN), a personal area network (PAN), a wide area network (WAN), an intranet, the Internet, a peer-to-peer network, a mesh network, and the like. User device 102 may also include a display (not shown).
[0023] The face authentication system 202 enables user 106 to access application 208, settings, or other resources of user device 102 using user 106's face or an image containing user 106's face. The face authentication system 202 includes at least one camera system 212, at least one face recognizer 214, and at least one ultrasound sensor 216. The face authentication system 202 may include another sensor 218 as needed. Although shown as part of the face authentication system 202 in Figure 2-1, the ultrasound sensor 216 and / or sensor 218 can be considered as separate entities capable of communicating with the face authentication system 202 (e.g., providing information) for face authentication anti-spoofing. Sometimes, in addition to supporting face authentication anti-spoofing, the ultrasound sensor 216 and / or sensor 218 also operate to support other features of user device 102. Various implementations of the face authentication system 202 may include a system-on-a-chip (SoC), one or more integrated circuits (ICs), a processor having embedded processor instructions or processor instructions configured to access stored in memory, hardware having embedded firmware, a printed circuit board having various hardware components, or any combination thereof.
[0024] The face authentication system 202 can be designed to operate under various environmental conditions. For example, the face authentication system 202 can support face authentication at a distance of approximately 70 centimeters (cm) or less. Distance refers to a distance between the user device 102 and the user 106. As another example, the face authentication system 202 can support face authentication at various tilt and lateral rotation angles that provide an angular view of the user device 102 at approximately -40 degrees to 40 degrees. These angles may include tilt angles between approximately -40 degrees and 20 degrees and lateral rotation angles between approximately -20 degrees and 20 degrees. With this range of angular views, the face authentication system 202 can operate in situations where the user 106 holds the user device 102 and / or where the user device 102 is on a surface and the user 106 is close to the user device 102.
[0025] Furthermore, the face authentication system 202 may be designed to make a decision regarding face authentication 104 within a predetermined time frame (such as 100 milliseconds (ms) to 200 ms). This frame may include the total time spent capturing image 108, using ultrasound to determine whether a demonstration attack 114 has occurred, and performing face recognition on image 108.
[0026] Camera system 212 captures one or more images 108 for face authentication 104. Camera system 212 includes at least one camera, such as a red-green-blue (RGB) camera. Camera system 212 may also include one or more illuminators to provide illumination, especially in dark environments. The illuminator may include an RGB light (such as an LED). In some embodiments, camera system 212 can be used for other applications, such as for selfies, capturing images for application 208, scanning documents, reading barcodes, etc.
[0027] Face recognition 214 performs face recognition to verify that one of the faces presented in image 108 corresponds to one of the authorized users 106 in user device 102. Face recognition 214 may be implemented in software, programmable hardware, or a combination thereof. In some embodiments, face recognition 214 is implemented using a machine learning module (e.g., a neural network).
[0028] Ultrasonic sensor 216 uses ultrasound to distinguish between a human face and a demonstration attack 114. Ultrasonic sensor 216 is further described with reference to Figure 2-2. Sensor 218 provides additional information to ultrasonic sensor 216, thereby enhancing the ability of ultrasonic sensor 216 to detect demonstration attack 114. Sensor 218 is further described with reference to Figure 2-3.
[0029] Figure 2-2 illustrates an example component of an ultrasound sensor 216. In the depicted configuration, the ultrasound sensor 216 includes a communication interface 220 for transmitting ultrasound sensor data to a remote device, but this is not required when the ultrasound sensor 216 is integrated into the user device 102. Generally, the ultrasound sensor data provided by the communication interface 220 is in a format that can be used by the face authentication system 202.
[0030] The ultrasonic sensor 216 also includes at least one sensor 222 that can convert electrical signals into sound waves. The sensor 222 can also detect sound waves and convert them into electrical signals. These electrical signals and sound waves can include frequencies within an ultrasonic range.
[0031] The spectrum (e.g., frequency range) used by sensor 222 to generate an ultrasonic signal may include frequencies within an ultrasonic range, which includes frequencies approximately between 20 kHz and 2 MHz. In some cases, the spectrum may be divided into sub-spectrums with similar or different bandwidths. For example, different frequency sub-spectrums may include 30 kHz to 500 kHz, 30 kHz to 70 kHz, 80 kHz to 500 kHz, 1 MHz to 2 MHz, 20 kHz to 48 kHz, 20 kHz to 24 kHz, 24 kHz to 28 kHz, 26 kHz to 29 kHz, 31 kHz to 34 kHz, 33 kHz to 36 kHz, or 31 kHz to 38 kHz.
[0032] These frequency sub-spectrums may be continuous or discontinuous, and may modulate the transmitted signal in phase and / or frequency. To achieve synchronization, multiple frequency sub-spectrums (continuous or discontinuous) with the same bandwidth can be used by sensor 222 to generate multiple ultrasonic signals transmitted simultaneously or temporally separately. In some cases, multiple continuous frequency sub-spectrums can be used to transmit a single ultrasonic signal, thereby giving the ultrasonic signal a wide bandwidth.
[0033] For face authentication anti-spoofing, the ultrasonic sensor 216 can use a frequency that provides a specific distance resolution beneficial to face authentication anti-spoofing. As an example, the ultrasonic sensor 216 can use a bandwidth that achieves a distance resolution of approximately 7 cm or less (e.g., approximately 5 cm or approximately 3 cm). An example bandwidth may be at least 7 kHz. Frequency can also be selected to support a specific detection range for face authentication (such as approximately 70 cm). Example frequencies include frequencies between approximately 31 kHz and 38 kHz.
[0034] Sometimes, in addition to facial recognition anti-counterfeiting, the ultrasonic sensor 216 also supports other features in the user device 102. These other features may include presence detection or grip strength detection. In this case, the ultrasonic sensor 216 can dynamically switch between different operating configurations optimized for active features. For example, the ultrasonic sensor 216 can use a frequency between approximately 26 kHz and 29 kHz for presence detection and a frequency between approximately 31 kHz and 38 kHz for facial recognition anti-counterfeiting.
[0035] In one exemplary embodiment, the sensor 222 of the ultrasonic sensor 216 has a monolithic topology. With this topology, the sensor 222 can convert electrical signals into sound waves and vice versa (e.g., transmit or receive ultrasonic signals). The exemplary monolithic sensor may include piezoelectric sensors, capacitive sensors, and microfabricated ultrasonic sensors (MUTs) using microelectromechanical systems (MEMS) technology.
[0036] Alternatively, sensor 222 may be implemented using a bistation topology comprising one of a plurality of sensors located at different locations on user device 102. In this case, a first sensor converts an electrical signal into a sound wave (e.g., transmits an ultrasound signal), and a second sensor converts the sound wave into an electrical signal (e.g., receives an ultrasound signal). An exemplary bistation topology may be implemented using at least one speaker and at least one microphone of user device 102. The speaker and microphone may be dedicated to the operation of ultrasound sensor 216. Alternatively, the speaker and microphone may be shared by both user device 102 and ultrasound sensor 216. Exemplary locations of the speaker and microphone are further described with reference to FIG5.
[0037] The ultrasonic sensor 216 includes at least one analog circuit 224, which includes circuitry and logic for modulating an electrical signal in an analog domain. The analog circuit 224 may include a waveform generator, analog-to-digital converter, amplifier, filter, mixer, phase shifter, and switcher for generating and modifying the electrical signal. In some embodiments, the analog circuit 224 includes other hardware circuitry associated with a speaker or microphone.
[0038] The ultrasonic sensor 216 also includes one or more system processors 226 and at least one system medium 228 (e.g., one or more computer-readable storage media). The system processor 226 processes electrical signals in a digital domain. The system medium 228 includes a spoofing detector 230. The spoofing detector 230 may be implemented using hardware, software, firmware, or a combination thereof. In this example, the system processor 226 implements the spoofing detector 230. The spoofing detector 230 processes responses (e.g., electrical signals) from the sensor 222 to detect a demo attack 114. The spoofing detector 230 may be implemented at least partially using a heuristic module or a machine learning module (e.g., a neural network).
[0039] In some embodiments, the ultrasound sensor 216 uses information provided by the sensor 218. For example, the sensor 218 provides the ultrasound sensor 216 with information about the location of an object (e.g., the face of user 106 or a demonstration attack device). This information may include range and / or angle measurements. In this way, the ultrasound sensor 216 can use the location information provided by the sensor 218 to customize (e.g., filter or normalize) the ultrasound data before determining whether a demonstration attack 114 exists. In some cases, this customization is performed before the ultrasound sensor 216 independently measures the location of the object using ultrasound technology. As needed, the ultrasound sensor 216 may utilize the location information provided by the sensor 218 to enhance the accuracy of measuring the location of the object using ultrasound technology.
[0040] As another example, the ultrasonic sensor 216 may utilize motion data provided by the sensor 218. Specifically, the ultrasonic sensor 216 may modify the ultrasonic data to compensate for the motion identified by the motion data. The exemplary sensor 218 is further described with reference to Figures 2-3.
[0041] Figure 2-3 illustrates an example sensor 218 for face authentication anti-spoofing. In the depicted configuration, sensor 218 may include a phase difference sensor 232, an RGB sensor 234, an inertial measurement unit (IMU) 236, and / or a radio frequency-based sensor 238. In some embodiments, phase difference sensor 232 is part of a front-facing camera 240 of camera system 212. Front-facing camera 240 may also be used to capture image 108 for face authentication. Phase difference sensor 232 can measure the distance to an object based on the detected phase difference. By measuring the distance, ultrasonic sensor 216 can determine a region of interest associated with the object and perform distance normalization to compensate for the distance between the object and ultrasonic sensor 216, which may change in different face authentication scenarios. Specifically, the ultrasonic sensor 216 is based on intensity or amplitude information correlated with the object after being calibrated by measuring distance. This enables the ultrasonic sensor 216 to support facial authentication anti-spoofing at various distances.
[0042] The RGB sensor 234 may also be part of the front-facing camera 240. The RGB sensor 234 can measure an angle of an object and provide this angle to the ultrasonic sensor 216. The ultrasonic sensor 216 can normalize energy or intensity information based on the measured angle. This enables the ultrasonic sensor 216 to better detect backscatter differences between an object or a face used in a demo attack 114 and supports anti-spoofing face authentication for various angles.
[0043] The inertial measurement unit 236 can measure the motion of one of the user devices 102. Using this motion information, the ultrasonic sensor 216 can perform motion compensation. Alternatively, the ultrasonic sensor 216 can activate in response to an instruction from the inertial measurement unit 236 that the user device 102 is approximately stationary. In this way, the inertial measurement unit 236 can reduce the number of motion artifacts observed by the ultrasonic sensor 216.
[0044] An example RF-based sensor 238 may include an ultra-wideband sensor, a radar sensor, or a WiFi® sensor. Using radio frequency, the RF-based sensor 238 can measure the distance and / or angle to an object presented for face authentication 104. Based on these measurements, the ultrasonic sensor 216 can determine an area of interest for detecting a demonstration attack and calibrate intensity or amplitude information. The interaction between the camera system 212, the face recognition device 214, the ultrasonic sensor 216, and the sensor 218 is further described with reference to Figure 3.
[0045] Figure 3 illustrates an example face authentication system 202 that performs face authentication anti-counterfeiting using ultrasound. In the depicted configuration, the face authentication system 202 includes a camera system 212, a face recognizer 214, and an ultrasound sensor 216. The face authentication system 202 may also include a sensor 218 if needed. The face recognizer 214 is coupled to the camera system 212 and the ultrasound sensor 216. The sensor 218 is coupled to the ultrasound sensor 216.
[0046] During operation, the face authentication system 202 receives a request 302 to perform face authentication 104. In some scenarios, request 302 is provided by application 208 or another service that requires face authentication 104. In other scenarios, request 302 is provided by sensor 218 in response to sensor 218 detecting interaction with one of the users 106. For example, inertial measurement unit 236 may send request 302 to face authentication system 202 in response to detecting that user 106 lifts user device 102.
[0047] In response to the receiving request 302, the face authentication system 202 initializes and starts the camera system 212. The camera system 212 captures at least one image 108 for face authentication 104. The camera system 212 provides the captured image 108 to the face recognizer 214.
[0048] Ultrasonic sensor 216 performs ultrasonic sensing to detect a demonstration attack 114. This can occur before, during, or after camera system 212 captures image 108. By analyzing the reflected ultrasonic energy, ultrasonic sensor 216 can determine whether the reflected ultrasonic energy originates from an object or a face associated with a demonstration attack 114. Ultrasonic sensor 216 generates a spoofing indicator 304 indicating whether demonstration attack 114 has been detected. Ultrasonic sensor 216 provides spoofing indicator 304 to face recognition device 214.
[0049] The face recognition device 214 receives the image 108 and the spoofing indicator 304. If the spoofing indicator 304 indicates that a demonstration attack 114 was not detected, the face recognition device 214 performs face recognition to determine whether the face of the user 106 exists in the image 108. The face recognition device 214 generates a report 306 provided to the application 208 or service that sent the request 302. If the face recognition device 214 recognizes the face of the user 106 and the spoofing indicator 304 indicates that a demonstration attack 114 is not present, the face recognition device 214 uses the report 306 to indicate successful face authentication. Alternatively, if the face recognition device 214 does not recognize the face of the user 106 and / or the spoofing indicator 304 indicates that a demonstration attack 114 has occurred, the face recognition device 214 uses the report 306 to indicate that face authentication has failed. This failure indication may cause application 208 and / or user device 102 to refuse access.
[0050] In some embodiments, sensor 218 provides sensor data 308 to ultrasound sensor 216. Ultrasonic sensor 216 uses sensor data 308 to customize the processing of ultrasound-based data, which is based on the measurement of reflected ultrasound energy. By taking into account the additional information provided by sensor 218, ultrasound sensor 216 can enhance its ability to detect demonstration attack 114.
[0051] In some configurations, the ultrasonic sensor 216 may additionally provide information to the face authentication system 202 or another component of the face authentication system 202 to support anti-spoofing techniques performed by that entity. For example, the ultrasonic sensor 216 may provide a measured distance of an object to the face authentication system 202. With this information, the face authentication system 202 may analyze the image 108 provided by the camera system 212 and determine whether the size of an object presented in the image 108 is consistent with the general size of a face at the measured distance provided by the ultrasonic sensor 216. Specifically, the face authentication system 202 may determine an upper or lower boundary associated with the size of a face at the measured distance. If the size of the object in the image 108 is within the boundary, the face authentication system 202 does not detect a spoofing attack 114. However, if the size of the object in the image 108 is outside the boundary, the face authentication system 202 detects the spoofing attack 114. In this way, the ultrasonic sensor 216 can also provide ultrasonic-based information that supports anti-counterfeiting technologies performed by other entities. (Face recognition anti-counterfeiting)
[0052] Figure 4-1 illustrates some differences in ultrasonic reflection between a demonstration attack device and a face. At 402, during face authentication 104, a demonstration attack device 404 (e.g., an object associated with demonstration attack 114) is presented to the user device 102, as shown in environments 100-3 to 100-5. At 406, during face authentication 104, a face 408 (e.g., the face of user 106) is presented to the user device 102, as shown in environments 100-1 and 100-2.
[0053] The demonstration attack device 404 may have a different physical structure than the face 408. For example, some demonstration attack devices 404 (such as medium 116 or device 120) may have a substantially flat or planar surface. Other demonstration attack devices 404 (such as mask 126 or other demonstration attack devices 404 with curved surfaces) may have an approximately segmented flat surface, wherein the demonstration attack device 404 consists of a set of substantially flat surfaces. In contrast, the face 408 may have significantly more defined contours and angles than the demonstration attack device 404. These contours form the forehead, nose, eye sockets, lips, cheeks, chin, and ears of the face 408. Based on these structural differences, the ultrasonic behavior of the demonstration attack device 404 may resemble a point target, while the ultrasonic behavior of the face 408 may resemble a distributed target.
[0054] Furthermore, the demonstration attack instrument 404 and the face 408 may be composed of different materials. These different materials may result in the demonstration attack instrument 404 and the face 408 having different absorption and reflection properties associated with ultrasound. For example, the face 408 may be composed of human tissue. Some demonstration attack instruments 404 may be composed of other materials (such as plastic) that have a higher reflectance coefficient than human tissue.
[0055] For both 402 and 406, the ultrasonic sensor 216 transmits at least one ultrasonic transmission signal 410. The ultrasonic transmission signal 410 propagates through space and impacts the demonstration attack instrument 404 and the face 408. Due to the differences in the physical structure and materials of the demonstration attack instrument 404 and the face 408, the ultrasonic transmission signal 410 interacts differently with the demonstration attack instrument 404 and the face 408. This interaction is illustrated by reflections 414-1 and 414-2 associated with the demonstration attack instrument 404 and the face 408, respectively.
[0056] Due to the planar surface of the demonstration attack instrument 404, a large amount of ultrasonic transmission signal 410 can be reflected back to the user device 102, as shown by the left-oriented reflection 414-1 in Figure 4-1. In contrast, the contour of the face 408 causes the ultrasonic transmission signal 410 to diffuse in many directions, as shown by the reflections 414-2 oriented in different directions. Therefore, less ultrasonic energy can be directed from the face 408 back towards the user device 102 relative to the demonstration attack instrument 404. Furthermore, if the demonstration attack instrument 404 contains a material with a reflectivity higher than that of human tissue, more ultrasonic energy can be reflected from the demonstration attack instrument 404 relative to the face 408. In other words, the reflection 414-1 associated with the demonstration attack instrument 404 can have a higher amplitude than the reflection 414-2 associated with the face 408. This amplitude difference is illustrated by the fact that reflection 414-1 has a longer length than reflection 414-2.
[0057] For both 402 and 406, the ultrasonic sensor 216 receives at least one ultrasonic receiving signal 412. The amplitude of the ultrasonic receiving signal 412 is based on the amount of reflected ultrasonic energy and the amount of ultrasonic energy guided back towards the ultrasonic sensor 216. Generally, the different physical structures and materials of the demonstration attack device 404 and the face 408 can cause the ultrasonic receiving signal 412 to have a higher amplitude at 402 than at 406. Based on the characteristics of the ultrasonic receiving signal 412, the ultrasonic sensor 216 can determine whether to present the demonstration attack device 404 or the face 408 for face authentication 104. The differences in the ultrasonic behavior of the demonstration attack device 404 and the face 408 are further described with reference to Figure 4-2.
[0058] Figure 4-2 illustrates the difference in received power between the demonstration attack device 404 and the face 408. A graph 416 has a first dimension associated with received power and a second dimension associated with range (e.g., distance or tilt range). In graph 416, a dashed line 418 represents the received power associated with the demonstration attack device 404. A solid line 420 represents the received power associated with the face 408. In this example, the demonstration attack device 404 and the face 408 are positioned at approximately the same distance from the user device 102.
[0059] As seen in graph 416, the demonstration attack device 404 generates a received power with a higher peak value than that of the face 408. This can be attributed at least in part to the higher reflectivity of the material forming the demonstration attack device 404. Furthermore, the received power associated with the demonstration attack device 404 is concentrated within a smaller range, while the received power associated with the face 408 is distributed over a larger range. This can be attributed at least in part to the point-like nature of the demonstration attack device 404 and the distributed nature of the face 408. By utilizing these inherent differences, the ultrasonic sensor 216 can employ various techniques to identify whether an object presented for face authentication 104 corresponds to a demonstration attack 114 or a face 408. These techniques are further described with reference to Figures 7 through 15.
[0060] Figure 5 illustrates an exemplary location of a speaker 502 and microphones 504-1 and 504-2 in a user device 102. Although the exemplary user device 102 in Figure 5 is shown to include one speaker 502 and two microphones 504-1 and 504-2, the ultrasonic sensor 216 can operate with one or more speakers and microphones at any given time. In a scenario where the ultrasonic sensor 216 can use multiple speakers, the ultrasonic sensor 216 can select one speaker to provide a higher signal-to-noise ratio. The signal-to-noise ratio may depend on the performance characteristics of the speaker and / or the position of the speaker relative to the user 106.
[0061] In this embodiment, the speaker 502 and microphones 504-1 and 504-2 are positioned on the same surface of the user device 102. In this case, this surface also includes a display 506. The microphone 504-1 is positioned on the side of the user device 102 opposite to the microphone 504-2. Consider a plane 508 that divides the user device 102 in two. In this case, a first half 510-1 of the user device 102 includes the speaker 502 and the microphone 504-1. A second half 510-2 of the user device 102 includes the microphone 504-2.
[0062] By placing microphones 504-1 and 504-2 separately (e.g., at a distance greater than the wavelength of the ultrasonic transmission signal 410), microphones 504-1 and 504-2 can observe different patterns or characteristics of reflections 414-1 or 414-2. For example, using multi-channel technology, ultrasonic sensor 216 can evaluate the differences in responses observed at microphones 504-1 and 504-2. Due to the point-like nature of the demonstration attack instrument 404, the responses observed at microphones 504-1 and 504-2 can be substantially similar. In contrast, due to the distributed nature of the face 408, the responses observed at microphones 504-1 and 504-2 for the face 408 can be significantly different. Therefore, ultrasonic sensor 216 can evaluate the differences in responses to determine whether a demonstration attack 114 is occurring. One operation of ultrasonic sensor 216 is further described using FIG. 6.
[0063] Figure 6 illustrates an exemplary embodiment of an ultrasonic sensor 216. In the depicted configuration, the ultrasonic sensor 216 includes a sensor 222, an analog circuit 224, and a system processor 226. The analog circuit 224 is coupled between the sensor 222 and the system processor 226. The analog circuit 224 includes a transmitter 602 and a receiver 604. The transmitter 602 includes a waveform generator 606 coupled to the system processor 226. Although not shown, the transmitter 602 may include one or more transmission channels. The receiver 604 includes a plurality of receiving channels 608-1 to 608-M, where M represents a positive integer. The receiving channels 608-1 to 608-M are coupled to the system processor 226.
[0064] Sensor 222 is implemented using a bistation topology comprising at least one speaker 502 and at least one microphone 504. In the depicted configuration, sensor 222 includes a plurality of speakers 502-1 to 502-S and a plurality of microphones 504-1 to 504-M, where S represents a positive integer. Speakers 502-1 to 502-S are coupled to transmitter 602, and microphones 504-1 to 504-M are coupled to respective receiving channels 608-1 to 608-M of receiver 604.
[0065] Although the ultrasonic sensor 216 in FIG. 6 includes multiple speakers 502 and multiple microphones 504, other embodiments of the ultrasonic sensor 216 may include a single speaker 502 and a single microphone 504, a single speaker 502 and multiple microphones 504, multiple speakers 502 and a single microphone 504, or other types of sensors capable of transmitting and / or receiving. In some embodiments, the speakers 502-1 to 502-S and the microphones 504-1 to 504-M may also operate with audible signals. For example, the user device 102 may play music through the speakers 502-1 to 502-S and use the microphones 504-1 to 504-M to detect the voice of the user 106.
[0066] During transmission, transmitter 602 transmits electrical signals to loudspeakers 502-1 to 502-S, which respectively transmit ultrasonic transmission signals 410-1 to 410-S. Specifically, waveform generator 606 generates electrical signals that may have similar waveforms (e.g., similar amplitude, phase, and frequency) or different waveforms (e.g., different amplitude, phase, and / or frequency). In some embodiments, system processor 226 transmits a configuration signal 610 to waveform generator 606. Configuration signal 610 may specify the characteristics of waveform generation (e.g., a center frequency, a bandwidth, and / or a modulation type). Using configuration signal 610, system processor 226 may customize ultrasonic transmission signals 410-1 to 410-S to support an active feature (such as facial recognition anti-spoofing or presence detection). Although not explicitly shown, waveform generator 606 may also transmit electrical signals to system processor 226 or receiver 604 for demodulation. The ultrasonic transmission signals 410-1 to 410-S may or may not be reflected by an object (e.g., a face 408 or a demonstration attack instrument 404).
[0067] During reception, microphones 504-1 to 504-M receive ultrasonic reception signals 412-1 to 412-M, respectively. The relative phase difference, frequency, and amplitude between the ultrasonic reception signals 412-1 to 412-M and the ultrasonic transmission signals 410-1 to 410-S may be attributed to changes in the interaction between the ultrasonic transmission signals 410-1 to 410-S and a nearby object (e.g., a face 408 or a demonstration attack device 404) or an external environment (e.g., path loss and noise sources). The ultrasonic reception signals 412-1 to 412-M represent versions of the ultrasonic transmission signal 410 propagated from one of the speakers 501-1 to 502-S to one of the microphones 504-1 to 504-M. The receiving channels 608-1 to 608-M process the ultrasonic receiving signals 412-1 to 412-M and generate the baseband receiving signals 612-1 to 612-M.
[0068] System processor 226 includes spoofing detector 230. System processor 226 and / or spoofing detector 230 may perform functions such as range compression, baseband processing, demodulation conversion, and / or filtering. Generally, spoofing detector 230 receives baseband received signals 612-1 to 612-M from receive channels 608-1 to 608-M and analyzes these signals to generate spoofing indicator 304, as further described with respect to FIG7.
[0069] Figure 7 illustrates an exemplary scheme for anti-spoofing facial authentication implemented using an ultrasonic sensor 216. In the depicted configuration, the spoofing detector 230 includes a preprocessor 702, a feature extractor 704, and a spoofing predictor 706. The preprocessor 702 operates on baseband received signals 612-1 to 612-M to provide data in a format usable by the feature extractor 704. The preprocessor 702 may perform filtering, Fourier transform (e.g., Fast Fourier Transform), range compression, normalization, and / or clutter cancellation. In some embodiments, filtering and / or normalization are based on sensor data 308, as further described below.
[0070] Feature extractor 704 can perform various functions to extract information useful for distinguishing between the demonstration attack instrument 404 and the face 408. In some cases, feature extractor 704 also utilizes the positional information of the object (e.g., range and / or angle) to extract appropriate information. Positional information may be determined by ultrasound sensor 216 using ultrasound technology, provided by sensor 218 via sensor data 308, or a combination thereof. The operation of feature extractor 704 may vary depending on the type of input data for the operation of ultrasound sensor 216 and whether the input data is associated with a single receiving channel 608 or multiple receiving channels 608-1 to 608-M. An exemplary feature extractor 704 is further described with reference to Figures 8-1, 9-1, 11-1, and 12-1.
[0071] The deception predictor 706 determines whether a demonstration attack 114 has occurred based on information provided by the feature extractor 704. The deception predictor 706 may include a comparator 708, a lookup table 710 (LUT 710) and / or a machine learning module 712.
[0072] During operation, the preprocessor 702 receives baseband received signals 612-1 to 612-M and generates complex data 714. The complex data 714 contains amplitude and phase information (e.g., real and imaginary numbers). Instance types of the complex data 714 include range-profile data 716 (e.g., data representing an intensity map), an interferogram 718, range-slow-time data 720 (e.g., data representing a range Doppler plot) and / or a power spectrum 722.
[0073] Range-profile data 716 contains amplitude (e.g., intensity) and phase information across a distance dimension and a time dimension (e.g., a fast time dimension). The distance dimension may be represented by a set of range bins (e.g., distance cells). The time dimension may be represented by a set of time intervals (e.g., linear frequency modulation pulse intervals or pulse intervals). Interferogram 718 may contain coherence and phase information across the distance and time dimensions. Range-slow time data 720 contains amplitude and phase information across the distance dimension and a Doppler (or slow time) dimension. The Doppler dimension may be represented by a set of Doppler bins. Power spectrum 722 contains power information across a frequency dimension and a time dimension.
[0074] In some cases, the preprocessor 702 may use sensor data 308 provided by sensor 218 to calibrate complex data 714 and / or filter complex data 714 according to a region of interest. As an example, the preprocessor 702 may calibrate amplitude-based information such as distance-profile data 716, interferogram 718, distance-slow-time data 720, and / or power spectrum 722 based on the measured position of one of the objects provided by sensor data 308. Specifically, the preprocessor 702 may normalize amplitude information based on the measured distance and / or angle of one of the objects presented for face authentication 104.
[0075] Additionally or alternatively, the preprocessor 702 may provide a subset of multiple data 714 associated with a region of interest identified by distance and / or angle measurements provided by sensor data 308. The region of interest may include the measured location of an object. In some cases, the region of interest is centered around the measured location of the object or around at least a portion of the measured location. The size of one of the regions of interest may be based on the general size of the face 408. Additionally or alternatively, the region of interest may be based on a region associated with face authentication (e.g., the region of interest may include a distance of up to 70 cm).
[0076] Feature extractor 704 extracts one or more features 724 from complex data 714. Instantaneous feature 724 may include a single-channel feature 726 and / or a multi-channel feature 728. Feature extractor 704 determines single-channel feature 726 based on one of the baseband received signals 612-1 to 612-M. In this manner, single-channel feature 726 is associated with one of the received channels 608-1 to 608-M. Alternatively or additionally, feature extractor 704 uses at least two of the baseband received signals 612-1 to 612-M to determine multi-channel feature 728. Multi-channel feature 728 is associated with two or more of the received channels 608-1 to 608-M.
[0077] In some cases, the feature extractor 704 may determine a feature 724 associated with an object based on a measured location of one of the objects. For example, the feature extractor 704 may identify a region covering a measured location of the object, which may be determined by analyzing multiple data 714 using ultrasound technology or based on sensor data 308 provided by sensor 218. Within this region, the feature extractor 704 evaluates the multiple data 714 to generate a feature 724. In this way, the ultrasound sensor 216 can ensure that the feature 724 is associated with the object presented for face authentication, rather than with another object present in the environment.
[0078] The deception predictor 706 analyzes the features 724 extracted by the feature extractor 704 to detect the demo attack 114 and generates a deception indicator 304. In some instances, the deception predictor 706 uses a comparator 708 to compare the features 724 with a threshold value. This threshold value is set so that the ultrasound sensor 216 can distinguish between the demo attack device 404 and a face 408 that may or may not be wearing an accessory 110. In other embodiments, the deception predictor 706 refers to a lookup table 710 to determine whether the features 724 correspond to the demo attack device 404 or the face 408. In yet another embodiment, the deception predictor 706 uses a machine learning module 712 to classify the features 724 as associated with the demo attack device 404 or the face 408.
[0079] The spoofing detector 230 generates a spoofing indicator 304 to control whether the face authentication system 202 can successfully authenticate the user 106. The spoofing indicator 304 may contain one or more elements indicating whether the object presented for face authentication 104 is associated with a demonstration attack device 404 or a face 408. An exemplary embodiment of the spoofing detector 230 is further described with reference to Figure 8-1. Single-channel face authentication anti-counterfeiting.
[0080] Figure 8-1 illustrates an example scheme implemented using an ultrasonic sensor 216 to generate a single-channel feature 726 for face authentication anti-spoofing. In the depicted configuration, the feature extractor 704 includes a distance-contour feature extractor 802. If needed, the preprocessor 702 may include a normalizer 804. The spoofing predictor 706 includes a comparator 708, a lookup table 710, and / or a machine learning module 712.
[0081] During operation, the preprocessor 702 generates distance-contour data 716. The distance-contour data 716 can be generated by performing a Fourier transform on one of the baseband received signals 612-1 to 612-M. For an embodiment including a normalizer 804, the preprocessor 702 can normalize amplitude information within the distance-contour data 716 based on sensor data 308, which may include a measured position (e.g., a measured distance and / or angle) of an object presented for face authentication 104. Alternatively or additionally, the normalizer 804 can adjust the amplitude information within the distance-contour data 716 based on a measured cross-coupling factor or a constant false alarm rate (CFAR) signal-to-noise ratio (SNR). The preprocessor 702 provides the normalized distance-contour data 806 to the feature extractor 704. For implementations that do not include a normalizer 804, the preprocessor 702 may alternatively provide the distance-contour data 716 to the feature extractor 704. Alternatively, these techniques can be applied similarly using distance-slow time data 720.
[0082] The distance-contour feature extractor 802 generates a single-channel feature 726 based on the distance-contour data 716 or normalized distance-contour data 806. The instanced single-channel feature 726 may include a peak amplitude feature 808, an energy distribution feature 810, and / or a phase feature 812. The peak amplitude feature 808 identifies a peak amplitude within the distance-contour data 716 or normalized distance-contour data 806. The energy distribution feature 810 identifies a peak energy or a shape of distributed energy within the distance dimension of the distance-contour data 716 or normalized distance-contour data 806. The phase feature 812 identifies phase information within the distance-contour data 716 or normalized distance-contour data 806.
[0083] In some cases, the distance-contour feature extractor 802 may determine a single-channel feature 726 associated with one of the objects based on its measured location. For example, the distance-contour feature extractor 802 may identify a region covering one of the measured locations of the object. Within this region, the distance-contour feature extractor 802 evaluates amplitude and / or phase information to generate a peak amplitude feature 808, an energy distribution feature 810, and / or a phase feature 812. In this way, the ultrasonic sensor 216 can ensure that the single-channel feature 726 is associated with the object presented for face authentication, rather than with another object present in the environment.
[0084] The spoofing predictor 706 detects whether a demo attack 114 occurs during face authentication 104 and generates a spoofing indicator 304 to convey this information to the face authentication system 202. Consider an instance where the distance-contour feature extractor 802 provides a peak amplitude feature 808 to the spoofing predictor 706. In this case, the spoofing predictor 706 can use a comparator 708 to compare a peak amplitude identified by the peak amplitude feature 808 with a threshold value 816. If the peak amplitude is greater than the threshold value 816, the spoofing predictor 706 detects the demo attack 114. Alternatively, if the peak amplitude is less than the threshold value 816, the spoofing predictor 706 does not detect the demo attack 114 (e.g., detects face 408). This example is further described with reference to Figure 8-2.
[0085] Figure 8-2 illustrates example distance-profile data 716 associated with face authentication anti-counterfeiting. Graphs 818, 820, 822, and 824 illustrate the distance-profile data 716 of different types of objects presented during face authentication 104. Specifically, graphs 818, 820, 822, and 824 depict the amplitude information of one of the ultrasonic received signals 412-1 to 412-M across the distance dimension.
[0086] Graph 818 represents the distance-contour data 716 collected when the object is a face 408. Graphs 820 and 822 represent the distance-contour data 716 collected when the objects are a piece of paper and a piece of cardboard, respectively. The paper and the cardboard may contain a photograph 118 of the user 106, as shown in environment 100-3. Graph 824 represents the distance-contour data 716 collected when the object is device 120 in environment 100-4. In this case, device 120 may display a digital image 124 of the user 106.
[0087] When the target is a human face 408, the amplitude of the distance-contour data 716 is below the threshold value 816, as shown in graph 818. However, when the target is a demonstration attack device 404, the amplitude of the distance-contour data 716 is above the threshold value 816 at at least some distance intervals, as shown in graphs 820, 822, and 824. In this way, the ultrasonic sensor 216 can distinguish between the human face 408 and the demonstration attack device 404.
[0088] In some embodiments, the ultrasonic sensor 216 can identify a distance 826 associated with the detected object. The ultrasonic sensor 216 can determine the distance 826 by ultrasonic sensing based on a combination of sensor data 308 or the like. At the identified distance 826, the ultrasonic sensor 216 can compare the amplitude of the distance-profile data 716 with the threshold value 816 to detect the demonstration attack 114.
[0089] Similar detection methods can be applied to other single-channel features 726. For example, the spoofing predictor 706 can use a comparator 708 to determine whether the extracted energy distribution feature 810 represents a face 408 or a demonstration attack device 404. Specifically, if the amount of energy at distance 826 is less than a threshold value 816, the ultrasonic sensor 216 does not detect the demonstration attack 114. However, if the amount of energy is greater than the threshold value 816, the ultrasonic sensor 216 detects the demonstration attack 114. Alternatively or additionally, the spoofing predictor 706 can use a lookup table 710 or a machine learning module 712 to determine whether the overall shape of one of the energy distribution features 810 across the distance dimension corresponds to a face 408 or a demonstration attack device 404. Similar techniques can be used to analyze phase features 812 using lookup table 710 or machine learning module 712.
[0090] Although described with respect to a single channel, the operations described in Figure 8-1 can be performed for more than one channel and, as needed, for each available receiving channel 608-1 to 608-M in series or parallel. In some embodiments, the feature extractor 704 provides the spoofing predictor 706 with a combination of multiple features, such as peak amplitude feature 808, energy distribution feature 810, and / or phase feature 812. By analyzing more than one single-channel feature 726, the spoofing predictor 706 can improve its ability to detect the demonstration attack 114. Another face authentication anti-spoofing technique is further described with respect to Figure 9-1. Multi-channel face authentication anti-spoofing is achieved using interferometric coherence.
[0091] Figure 9-1 illustrates an exemplary scheme implemented by an ultrasonic sensor 216 to generate a multi-channel feature 728 for face authentication anti-spoofing. In this example, the ultrasonic sensor 216 uses interferometry to evaluate a quantity of similarity between reflections observed by two or more sensors 222 (such as microphones 504-1 and 504-2). Due to the planar or uniform structure of the demonstration attack instrument 404, the ultrasonic received signals 412-1 to 412-M generated by the demonstration attack instrument 404 and received by microphones 504-1 to 504-M can have relatively similar characteristics. In other words, the amplitude and phase of the ultrasonic received signals 412-1 to 412-M can be relatively similar within a given time interval. However, the contouring and non-uniform structure of the face 408 can cause substantial differences in the ultrasonic received signals 412-1 to 412-M generated by the face 408 and received by microphones 504-1 to 504-M. In other words, the amplitude and phase of the ultrasonic received signals 412-1 to 412-M can change significantly within a similar time interval. By using interferometry, the ultrasonic sensor 216 can distinguish between the demonstration attack instrument 404 and the face 408 for face authentication anti-spoofing.
[0092] In the depicted configuration, the preprocessor 702 includes a common registration module 902, a filter 904, and an interferogram generator 906. Although not shown, the preprocessor 702 may also include the normalizer 804 of FIG. 8-1. The feature extractor 704 includes an interferogram feature extractor 908. The deception predictor 706 includes a comparator 708. Although not shown, other embodiments of the deception predictor 706 may use a lookup table 710 or a machine learning module 712.
[0093] Due to the different positions of microphones 504-1 to 504-M, an object presented for face authentication 104 can be located at different distances within the distance-contour data 716 associated with the corresponding receiving channels 608-1 to 608-M. To address this, a co-registration module 902 aligns across the distance dimension with the distance-contour data 716 associated with the different receiving channels 608-1 to 608-M. Various types of co-registration modules 902 are further described with reference to Figures 10-1 to 10-4.
[0094] Filter 904 filters the data provided by the co-registration module 902 based on a region of interest, which can be identified based on sensor data 308, distance-contour data 716, or a combination thereof. The region of interest may include a spatial region associated with face authentication or a region surrounding a measured location of an object. By filtering the data, filter 904 can improve the processing speed of subsequent procedures and / or remove data that is not associated with the object of interest.
[0095] Interference pattern generator 906 generates interference pattern 718. Specifically, interference pattern generator 906 performs complex coherence on distance-profile data 716 (or distance slow time data 720) associated with both receiving channels 608-1 to 608-M to generate interference pattern 718.
[0096] During operation, the co-registration module 902 compensates for positional differences between microphones 504-1 to 504-M by generating co-registered distance-profile data 910 based on distance-profile data 716 associated with different receiving channels 608-1 to 608-M. Within the co-registered distance-profile data 910, the distance associated with the object is similar across the different receiving channels 608-1 to 608-M. Filter 904 filters the co-registered distance-profile data 910 to generate filtered co-registered distance-profile data 912.
[0097] Prior to the interferogram generator 906, the range-profile data 716, the co-registered range-profile data 910, and the filtered co-registered range-profile data 912 have unique data associated with each of the receiving channels 608-1 to 608-M. The interferogram generator 906 operates on the filtered co-registered range-profile data 912 across the pairs of receiving channels 608-1 to 608-M to generate at least one interferogram 718. The interferogram 718 contains coherence information and phase information. The coherence information represents a measure of similarity (e.g., coherence or correlation) between the filtered co-registered range-profile data 912 of the two receiving channels 608-1 to 608-M.
[0098] The interferogram feature extractor 908 generates multi-channel features 728 based on the interferogram 718. The instance multi-channel feature 728 includes a cohomology feature 914 and a phase feature 916. The cohomology feature 914 represents a cohomology quantity over time for a subset of distances. Sometimes, the cohomology feature 914 represents an average cohomology.
[0099] A cohomology value of zero to one indicates that the filtered co-registered distance-profile data 912 is uncorrelated across the pair of receiving channels 608-1 to 608-M. In contrast, a cohomology value of one indicates that the filtered co-registered distance-profile data 912 is strongly correlated across the pair of receiving channels 608-1 to 608-M. A cohomology value between zero and one indicates that the filtered co-registered distance-profile data 912 is partially or weakly correlated across the pair of receiving channels 608-1 to 608-M.
[0100] Phase feature 916 represents the change in phase information of interferogram 718 for a subset of distances over time. Instantaneous phase feature 916 may represent the average phase difference across a region of interferogram 718, the maximum phase difference across that region of interferogram 718, and / or the standard deviation across that region of interferogram 718.
[0101] In some cases, the interferogram feature extractor 908 may determine the coherence features 914 and / or phase features 916 associated with one of the measured locations of the object. For example, the interferogram feature extractor 908 may identify a region covering one of the measured locations of the object. Within this region, the interferogram feature extractor 908 evaluates filtered co-registered distance-contour data 912 to extract the coherence features 914 and / or phase features 916. In this way, the ultrasound sensor 216 can ensure that the multi-channel features 728 are associated with the object presented for face authentication, rather than with another object present in the environment. Alternatively or additionally, the filter 904 may filter the co-registered distance-contour data 910 based on the measured location of the object to cause the subsequently generated multi-channel features 728 to be associated with the object of interest.
[0102] The deception predictor 706 uses a comparator 708 to compare the coherence feature 914 and / or the phase feature 916 with a corresponding threshold value. For example, the deception predictor 706 compares the coherence feature 914 with a coherence threshold value 918. In one instance, the coherence threshold value 918 is approximately equal to 0.5. If the coherence feature 914 (e.g., the amount of coherence observed for a subset of distances) is less than the coherence threshold value 918 and greater than zero, then the deception predictor 706 determines that the object presented for face authentication 104 corresponds to face 408. Alternatively, if the coherence feature 914 is greater than the coherence threshold value 918 and less than one, then the deception predictor 706 determines that the object presented for face authentication 104 corresponds to the demonstration attack device 404.
[0103] As another example, the deception predictor 706 compares the phase feature 916 with the phase threshold 920. If the phase feature 916 indicates that the phase information within the interferogram 718 is greater than the phase threshold 920, then the deception predictor 706 determines that the object presented for face authentication 104 corresponds to face 408. Alternatively, if the phase feature 916 is less than the phase threshold 920, then the deception predictor 706 determines that the object presented for face authentication 104 corresponds to the demonstration attack device 404.
[0104] Some implementations of the deception predictor 706 may analyze both the coherence feature 914 and the phase feature 916 to determine whether a demonstration attack 114 has occurred. Alternatively, the deception predictor 706 may use other techniques, such as employing a lookup table 710 or a machine learning module 712 to analyze the interferogram 718 and determine whether a demonstration attack 114 has occurred.
[0105] The operations described in FIG9-1 can be performed on the distance-profile data 716 provided by a pair of receiving channels 608-1 to 608-M. If the ultrasonic sensor 216 uses more than two receiving channels 608-1 to 608-M, these operations can be performed in series or in parallel for multiple pairs of receiving channels 608-1 to 608-M. Consider an instance where the ultrasonic sensor 216 uses one of four receiving channels (e.g., M equals four). In this case, the preprocessor 702 can generate up to six interferograms 718, which can be processed by the feature extractor 704 and the spoofing predictor 706 to detect the demonstration attack 114. If multiple interferograms 718 are available, the spoofing predictor 706 can detect the demonstration attack 114 in response to a multi-channel feature 728 associated with one of the multiple interferograms 718 indicating the presence of the demonstration attack instrument 404. To reduce the probability that the ultrasonic sensor 216 incorrectly determines that a face 408 presented for face authentication 104 represents a demonstration attack instrument 404, the deception predictor 706 can detect the demonstration attack 114 in response to a predetermined number of interferograms 718 (e.g., three or more of multiple interferograms 718) indicating the presence of the demonstration attack instrument 404. The exemplary phase information and coherence information provided by the interferograms 718 are further described with reference to Figure 9-2.
[0106] Figure 9-2 illustrates exemplary interference diagrams 928 to 932 related to facial authentication anti-counterfeiting. Interference diagram 928 represents data collected when the object is a face 408. Interference diagrams 930 and 932 represent data collected when the object is a device 120 or a piece of paper, respectively. The piece of paper may contain a photograph 118 of the user 106, as shown in environment 100-3.
[0107] Interferograms 928 to 932 each contain phase information 934 and coherence information 936. Different shading in phase information 934 represents different phase angles. Different shading in coherence information 936 represents different coherence amounts. Lighter shading indicates higher coherence amounts, and darker shading indicates lower coherence amounts. In Figure 9-2, identify the cells of interferograms 928 to 932 associated with a region of interest 938. Region of interest 938 is based on a measured distance between the object presented for face authentication 104 and the ultrasonic sensor 216.
[0108] Within the region of interest 938 of the interferogram 928, phase information 934 contains significantly different phases. Coherence information 936 contains numerous units weakly correlated with coherence values less than 0.5. The coherence within the region of interest 938 may appear random. These characteristics can be attributed at least in part to the contouring and distributed nature of the face 408.
[0109] Within the regions of interest 938 of interferograms 930 and 932, phase information 934 depicts similar phases over time. Coherence information 936 within the regions of interest 938 contains numerous elements strongly correlated with coherence values greater than 0.5. These characteristics can be at least partially attributed to the planar and point-like nature of these demonstration attack instruments 404.
[0110] Comparing the regions of interest 938 between interferograms 928 and 932, the phase information 934 within interferograms 930 and 932 exhibits less variation than the phase information 934 within interferogram 928. Furthermore, the coherence information 936 within interferograms 930 and 932 exhibits a higher degree of coherence than the coherence information 936 within interferogram 928. Additionally, there is less variation between the coherence information 936 within interferograms 930 and 932 compared to the coherence information 936 within interferogram 928. Generally, the multi-channel features 728 extracted by the feature extractor 704 and the analysis performed by the deception predictor 706 enable the ultrasonic sensor 216 to detect these characteristics and appropriately identify an object as corresponding to a face 408 or a demonstration attack instrument 404.
[0111] Interferometric techniques can also be applied to distinguish between a user 106 wearing an accessory 110 and a demonstration attack device 404. In some instances, the accessory 110 induces coherence information 936 with some strongly correlated units. However, compared to the strongly correlated units of the demonstration attack device 404, these strongly correlated units can exist across a wider area in terms of distance. By analyzing the phase information 934 and / or coherence information 936 within the interferogram 928, the ultrasonic sensor 216 can identify the characteristics that distinguish the face 408 from the demonstration attack device 404.
[0112] Figure 10-1 illustrates an exemplary scheme for performing co-registration using an ultrasonic sensor 216. In this case, the co-registration module 902 uses low-pass filtering to align the range-profile data 716 associated with the receiving channels 608-1 and 608-2. The co-registration module 902 includes two low-pass filters 1002-1 and 1002-2. During operation, low-pass filters 1002-1 and 1002-2 filter the range-profile data 716 of the receiving channels 608-1 and 608-2 across the range dimension to generate co-registered range-profile data 910. By using low-pass filters 1002-1 and 1002-2, amplitude and / or phase information across adjacent range baskets is combined to form a larger composite range basket. Although the distance frame of an object may differ between the distance-profile data 716 of receiving channels 608-1 and 608-2, the distance frame of the object may be the same within the larger composite distance frame of the co-registered distance-profile data 910. In this way, the co-registration module 902 is aligned with the distance-profile data 716 of receiving channels 608-1 and 608-2 in terms of distance.
[0113] Although implementing a co-registration module 902 with low-pass filters 1002-1 and 1002-2 can be relatively simple, the resulting co-registered range-profile data 910 has a lower range resolution compared to the range-profile data 716. The low-pass filtering also provides localized co-registration. This means that the range boxes associated with the object are substantially aligned, while other range boxes within the range-profile data 716 may be misaligned. To address the reduced range resolution, the co-registration module 902 can alternatively utilize the techniques described with reference to Figure 10-2.
[0114] Figure 10-2 illustrates another exemplary scheme for performing co-registration using an ultrasonic sensor 216. In this case, the co-registration module 902 evaluates the maximum echo to align with the distance-profile data 716 associated with the receiving channels 608-1 and 608-2. The co-registration module 902 includes two peak detectors 1004-1 and 1004-2, a comparator 1006, and a shifter 1008.
[0115] During operation, peak detectors 1004-1 and 1004-2 evaluate the distance-profile data 716 associated with receiving channels 608-1 and 608-2, respectively, and identify distance bins 1010-1 and 1010-2 associated with a peak amplitude. Comparator 1006 compares distance bins 1010-1 and 1010-2 to determine a difference between distance bin 1010-2 and distance bin 1010-1. This difference represents an offset 1012 (e.g., a distance offset) between the distance-profile data 716 of receiving channel 608-1 and the distance-profile data 716 of receiving channel 608-2. Shifter 1008 shifts the distance-profile data 716 of receiving channel 608-2 across the distance dimension by an amount identified by offset 1012. This results in shifter 1008 generating shifted distance-profile data 1014 for receiving channel 608-2. The distance-profile data 716 of receiving channel 608-1 and the shifted distance-profile data 1014 of receiving channel 608-2 are provided as the distance-profile data 910 for common registration.
[0116] Although slightly more complex than the co-registration module 902 of Figure 10-1, the co-registration module 902 of Figure 10-2 maintains the distance resolution provided by the distance-profile data 716. However, this technique still suffers from the localized co-registration problem described above with respect to Figure 10-1. To solve the localized co-registration problem, the co-registration module 902 can alternatively utilize the technique described with respect to Figure 10-3.
[0117] Figure 10-3 illustrates an additional example scheme for performing co-registration using an ultrasonic sensor 216. In this case, the co-registration module 902 evaluates the maximum echo to align the distance-profile data 716 associated with the receiving channels 608-1 and 608-2 and uses interpolation to provide global co-registration. The co-registration module 902 includes two peak detectors 1004-1 and 1004-2, a comparator 1006, a shifter 1008, a sub-basket distance offset detector 1016 (e.g., a sub-pixel distance offset detector), and an interpolator 1018.
[0118] During operation, peak detectors 1004-1 and 1004-2, comparator 1006, and shifter 1008 generate shifted distance-profile data 1014 for receiver channel 608-2, as described above with respect to Figure 10-2. Sub-basket distance offset detector 1016 compares the distance-profile data 716 of receiver channel 608-1 with the shifted distance-profile data 1014 of receiver channel 608-2 to identify another offset 1020 smaller than one of the distance basket sizes. Interpolator 1018 interpolates the shifted distance-profile data 1014 based on offset 1020 to generate interpolated and shifted distance-profile data 1022 for receiver channel 608-2. The distance-profile data 716 of receiving channel 608-1 and the interpolated and shifted distance-profile data 1022 of receiving channel 608-2 are provided as the distance-profile data 910 for common registration.
[0119] By using interpolation, the co-registration module 902 of FIG10-3 can more accurately align the distance-profile data 716 of the receiving channels 608-1 and 608-2 with respect to the co-registration modules 902 of FIG10-1 and FIG10-2. However, this technique also increases the computational burden. Another accuracy technique for co-registration is further described with reference to FIG10-4.
[0120] Figure 10-4 illustrates another exemplary scheme for performing co-registration using an ultrasonic sensor 216. In this case, the co-registration module 902 projects distance-profile data 716 onto a common grid. The co-registration module 902 includes a common grid projector 1024 and an interpolator 1026.
[0121] During operation, the public grid projector 1024 receives Euler angles 1028 and geometric information 1030. Euler angles 1028 represent an orientation of the user device 102. Geometric information 1030 provides information about the relative positions of microphones 504-1 and 504-2 associated with receiving channels 608-1 and 608-2. Geometric information 1030 can specify the distance and angle to microphones 504-1 and 504-2 based on a reference point associated with the user device 102. By analyzing Euler angles 1028 and geometric information 1030, the public grid projector 1024 can determine the propagation path differences between the object and microphones 504-1 and 504-2. Specifically, the public grid projector 1024 can determine the amplitude and / or phase differences attributable to the different positions of microphones 504-1 and 504-2 within the user device 102.
[0122] Interpolator 1026 uses information provided by common grid projector 1024 to resample or interpolate the distance-profile data 716 of receiving channel 608-2. This resampling adjusts the characteristics (e.g., amplitude and phase) of the distance-profile data 716 to correspond to the common grid. In some embodiments, the common grid is based on one of the positions of microphone 504-1. In this case, the distance-profile data 716 of receiving channel 608-1 does not need to pass through interpolator 1026 and can be provided as a portion of the commonly registered distance-profile data 910. Interpolator 1026 generates projected distance-profile data 1032 of receiving channel 608-2, which can be provided as another portion of the commonly registered distance-profile data 910. Face authentication anti-spoofing using skew.
[0123] Figure 11-1 illustrates an exemplary scheme for generating another single-channel or multi-channel feature for face authentication anti-spoofing using an ultrasonic sensor 216. In this example, the ultrasonic sensor 216 evaluates the skewness (e.g., asymmetry or adjusted Fisher-Pearson normalized moment coefficient) of an intensity-based probability distribution of at least one receiving channel 608. The different physical structures of the demonstration attack device 404 and the face 408 cause the intensity-based probability distribution associated with these objects to have different shapes. In this way, the ultrasonic sensor 216 can distinguish between the demonstration attack device 404 and the face 408 for face authentication anti-spoofing.
[0124] In the described configuration, the preprocessor 702 includes an intensity-based histogram generator 1102. The feature extractor 704 includes a skewness feature extractor 1104. The deception predictor 706 includes a comparator 708, a lookup table 710, and / or a machine learning module 712.
[0125] During operation, the intensity-based histogram generator 1102 generates at least one histogram 1106 based on the distance-profile data 716. If the ultrasonic sensor 216 utilizes multiple receiving channels 608-1 to 608-M, the intensity-based histogram generator 1102 can generate multiple histograms 1106. An example histogram 1106 includes a power distribution histogram 1108 and a constant false alarm rate (CFAR) signal-to-noise ratio (SNR) histogram 1110. The power distribution histogram 1108 represents the statistical frequency of the received power at different levels. The constant false alarm rate SNR histogram 1110 (CFAR SNR histogram 1110) represents the statistical frequency of different constant false alarm rate SNRs of the ultrasonic received signals 412-1 to 412-M, which have been normalized based on a CFAR kernel.
[0126] In some cases, the intensity-based histogram generator 1102 may filter the distance-profile data 716 based on the measured position of the object before generating the histogram 1106. In this way, the intensity-based histogram generator 1102 may cause a single-channel feature 726 or a multi-channel feature 728 to be associated with the object of interest, which may be determined later.
[0127] Skewness feature extractor 1104 generates skewness feature 1112 based on histogram 1106. Skewness feature 1112 represents the amount of asymmetry or adjusted Fisher-Pearson normalized moment coefficient. Spoofing predictor 706 analyzes skewness feature 1112 to determine whether the object presented for face authentication 104 corresponds to face 408 or demonstration attack device 404. Specifically, if skewness feature 1112 indicates that histogram 1106 is positively skewed relative to a threshold value, then spoofing predictor 706 can determine that the object is demonstration attack device 404.
[0128] If multiple histograms 1106 are available, the spoofing predictor 706 can also assess the differences in skewness characteristics 1112 between the multiple histograms 1106. Sometimes, a skewness difference greater than 0.1 dB can indicate that the target is a demonstration attack instrument 404. Alternatively or additionally, the spoofing predictor 706 can use a lookup table 710 or a machine learning module 712 to identify the presence of heavy tails in one or more of the histograms 1106. The term "heavy tail" refers to a higher statistical frequency of occurrence relative to a normal distribution toward one end of a distribution. The presence of heavy tails can indicate the presence of a demonstration attack instrument 404, as further described below. An example power distribution histogram 1108 and a constant false alarm rate signal-to-noise ratio histogram 1110 are depicted in Figures 11-2 and 11-3.
[0129] Figure 11-2 illustrates exemplary power distribution histograms 1114 to 1118 associated with face authentication anti-counterfeiting. Power distribution histogram 1116 represents data collected when the object is a face 408. Power distribution histograms 1116 and 1118 represent data collected when the object is a latex mask and a rigid plastic mask, respectively. Each of power distribution histograms 1114 to 1118 includes a first power distribution 1120 associated with receiving channel 608-1 (e.g., associated with microphone 504-1 in Figure 5) and a second power distribution 1122 associated with receiving channel 608-2 (e.g., associated with microphone 504-2 in Figure 5).
[0130] The power distributions 1120 and 1122 within the power distribution histogram 1114 have a relatively normal distribution. In this case, the skewness of power distribution 1120 can be approximately -0.7 dB, and the skewness of power distribution 1122 can be approximately -0.6 dB. Furthermore, the difference between the skewnesses of power distributions 1120 and 1122 can be relatively small (e.g., approximately 0.1 dB or less).
[0131] In contrast, at least one of the power distributions 1120 and 1122 within power distribution histograms 1114 and 1118 is skewed in positional direction relative to the power distributions 1120 and 1122 within power distribution histogram 1114. The exemplary skewness of the power distributions 1120 and 1122 within power distribution histogram 1116 can be approximately -0.2 dB and -0.9 dB, respectively. In this case, for the latex mask, the skewness of power distribution 1120 is 0.5 dB greater relative to the face 408. Power distribution 1122 within power distribution histogram 1114 also has a non-Gaussian distribution or shape. Specifically, power distribution 1122 has a heavy tail on the upper right side.
[0132] The exemplary skewness of power distributions 1120 and 1122 within the power distribution histogram 1118 can be approximately -0.7 dB and -0.7 dB, respectively. In this case, for a rigid plastic mask, the skewness of power distribution 1122 is 0.1 dB greater than that of the face 408. In this case, the ultrasonic sensor 216 can use other techniques (such as interferometry (Figure 9-1) or other single-channel techniques (Figure 8-1)) to distinguish the face 408 from the rigid plastic mask.
[0133] For some types of demonstration attack instruments 404, a difference between the skewness of power distributions 1120 and 1122 can be substantially large (e.g., greater than 0.3 dB), such as in the power distribution histogram 1116 used for latex masks. This is another characteristic that the deception predictor 706 can use to distinguish between a face 408 and the demonstration attack instrument 404.
[0134] This technique can also be applied to other types of demonstration attack instruments 404 (such as device 120 and a paper-based medium 116). For both of these demonstration attack instruments 404, the positive skewness can be even more pronounced. For example, if the target is device 120, the skewness of power distributions 1120 and 1122 can be approximately 0.4 dB and 0.3 dB, respectively. As another example, if the target is paper-based medium 116, the skewness of power distributions 1120 and 1122 can be approximately 0.1 dB and -0.2 dB, respectively.
[0135] Figure 11-3 illustrates example constant false alarm rate (CFRR) histograms 1124 and 1126 associated with face authentication anti-counterfeiting. CFRR histogram 1124 represents data collected when the target is a face 408. CFRR histogram 1126 represents data collected when the target is a device 120. Each of CFRR histograms 1124 and 1126 includes a first distribution 1128 associated with receiving channel 608-1 (e.g., associated with microphone 504-1) and a second distribution 1130 associated with receiving channel 608-2 (e.g., associated with microphone 504-2).
[0136] Distributions 1128 and 1130 within the constant false alarm rate (False Alarm Rate) SNR histogram 1124 have relatively normal distributions (e.g., Gaussian distributions). In contrast, the two distributions 1128 and 1130 within the constant false alarm rate SNR histogram 1126 have tails on the upper right. The tails in the constant false alarm rate SNR histogram 1126 are heavier than those in the constant false alarm rate SNR histogram 1124 (e.g., having a higher statistical frequency). Face authentication anti-spoofing based on power spectrum variation is used.
[0137] Figure 12-1 illustrates an exemplary scheme implemented by an ultrasonic sensor 216 to generate another single-channel or multi-channel feature for face authentication anti-counterfeiting. In this example, the ultrasonic sensor 216 uses spectral variation to evaluate a quantity of variation observed over time within at least one receiving channel 608. Due to the planar structure of some demonstration attack instrument 404, the spectral signature associated with the demonstration attack instrument 404 may have a higher variation than that of a distributed target (such as face 408). By analyzing the spectral variation of one or more receiving channels 608-1 to 608-M, the ultrasonic sensor 216 can distinguish between the demonstration attack instrument 404 and face 408 for face authentication anti-counterfeiting.
[0138] In the depicted configuration, feature extractor 704 includes a sliding window 1202 and a variance feature extractor 1204. Deception predictor 706 may include comparator 708, lookup table 710, and / or machine learning module 712. During operation, preprocessor 702 (not shown) provides a power spectrum 722 associated with one of the receiving channels 608 to sliding window 1202. To generate power spectrum 722, preprocessor 702 may perform a one-dimensional Fourier transform operation on distance-profile data 716 across the distance dimension. Alternatively, preprocessor 702 may perform a two-dimensional Fourier transform operation on distance-slow time data 720. Before generating power spectrum 722, preprocessor 702 may filter distance-profile data 716 based on the measured position of the object. In this way, preprocessor 702 may cause a single-channel feature 726 or multi-channel feature 728 to be associated with the object of interest for later determination.
[0139] A sliding window 1202 divides the power spectrum 722 into individual sub-frames 1206 along a time dimension. These sub-frames 1206 have different time intervals, which may or may not overlap. A variance feature extractor 1204 calculates one standard deviation (or variance) of the amplitude information within one or more sub-frames, and as needed, to generate a variance feature 1208. An example variance feature 1208 is further described below with reference to Figures 12-2 and 12-3.
[0140] Figures 12-2 and 12-3 illustrate standard deviation curves 1210 to 1220 related to face authentication anti-counterfeiting. Standard deviation curves 1210 and 1212 are associated with a face 408 and a face 408 wearing a hat as an accessory 110, respectively. Standard deviation curve 1214 is associated with a plastic mask. Standard deviation curves 1216, 1218, and 1220 are associated with a paper-based medium 116, a device 120, and a latex mask, respectively. Each of the standard deviation curves 1210 to 1220 depicts the variances 1222 and 1224 associated with receiving channels 608-1 and 608-2 (e.g., associated with microphones 504-1 and 504-2 in Figure 5).
[0141] To distinguish between a face 408 (with or without accessory 110) and various demonstration attack devices 404, the deception predictor 706 can determine whether variants 1222 and 1224 are within a set of expected values. If variants 1222 and 1224 are within the set of expected values, the deception predictor 706 determines that one of the objects presented for face authentication 104 corresponds to face 408. Alternatively, if variants 1222 or 1224 have values outside the set of expected values, the deception predictor 706 determines that the presented object corresponds to demonstration attack device 404. The set of expected values may typically include the values shown in the standard deviation curves 1210 and 1212.
[0142] Considering the standard deviation curves 1216 and 1218 associated with the paper-based medium 116 and device 120, the variances 1222 and 1224 are significantly greater than the variances 1222 and 1224 of the standard deviation curves 1210 and 1212. For the standard deviation curve 1220 associated with the latex mask, the variances 1222 and 1224 are significantly smaller than the variances 1222 and 1224 of the standard deviation curves 1210 and 1212.
[0143] Using spectral variation to distinguish standard deviation curves 1210 and 1212 from the standard deviation curve 1214 associated with the plastic mask can be challenging. In this case, the ultrasound sensor 216 can perform one or more of the other described techniques to distinguish the face 408 from the plastic mask. Example Method
[0144] Figures 13, 14, and 15 depict exemplary methods 1300, 1400, and 1500 for anti-spoofing facial authentication using ultrasound. Each method 1300, 1400, and 1500 is shown as a set of operations (or actions) performed and is not necessarily limited to the order or combination of such operations shown herein. Furthermore, one or more of the operations may be repeated, combined, reorganized, or linked to provide a wide range of additional and / or alternative methods. In the following discussion, reference may be made to the environments 100-1 to 100-5 of Figure 1 and the entities detailed in Figures 2 to 4-1, which are referred to by way of example only. The technology is not limited to the execution by a single entity or multiple entities operating on a user device 102.
[0145] At 1302 in Figure 13, an ultrasonic transmission signal is transmitted. For example, ultrasonic sensor 216 transmits ultrasonic transmission signal 410, as shown in Figure 4-1. Ultrasonic transmission signal 410 contains frequencies in the range of approximately 20 kHz to 2 MHz, for example, as described above with reference to Figure 2-2, and may represent a pulse signal or a continuous signal. In some cases, ultrasonic sensor 216 modulates a characteristic (including phase and / or frequency) of ultrasonic transmission signal 410. In some embodiments, ultrasonic sensor 216 transmits ultrasonic transmission signal 410-1 or 410-S in response to a face authentication system 202 receiving a request 302 or in response to ultrasonic sensor 216 self-sensor 218 receiving an alert indicating that user device 102 is approximately stationary (e.g., an alert from inertial measurement unit 236).
[0146] The ultrasound sensor 216 may use a dedicated sensor 222 to transmit the ultrasound transmission signal 410. In other embodiments, the ultrasound sensor 216 may use a shared speaker (e.g., speaker 502) of one of the user devices 102 to transmit the ultrasound transmission signal 410.
[0147] At 1304, an ultrasonic receiving signal is received. This ultrasonic receiving signal includes a version of an ultrasonic transmission signal reflected by an object. For example, ultrasonic sensor 216 receives ultrasonic receiving signal 412 using either microphone 504-1 or 504-2. Ultrasonic receiving signal 412 is a version of ultrasonic transmission signal 410 reflected by an object (e.g., a delayed version of ultrasonic transmission signal 410). In some cases, ultrasonic receiving signal 412 has an amplitude and / or is shifted in phase and / or frequency compared to ultrasonic transmission signal 410. In some embodiments, ultrasonic sensor 216 receives ultrasonic receiving signal 412 during at least a portion of the time during which ultrasonic transmission signal 410 is transmitted.
[0148] At 1306, range-profile data is generated based on the received ultrasound signal. The range-profile data includes amplitude and phase information associated with the received ultrasound signal. For example, the spoofing detector 230 of the ultrasound sensor 216 generates range-profile data 716 as one type of complex data 714. The range-profile data 716 contains amplitude and phase information across a distance dimension and a time dimension.
[0149] At 1308, a feature of distance-profile data within a region of interest associated with a location of an object is determined. For example, ultrasound sensor 216 determines a feature 724 (e.g., a single-channel feature 726) within a region of interest associated with a location of an object. An example single-channel feature 726 may include a peak amplitude feature 808, an energy distribution feature 810, and / or a phase feature 812, as shown in Figure 8-1. Generally, the region of interest identification covers a general spatial region encompassing the location of the object. In some cases, the identified spatial region may surround at least a portion of the object's location to account for an error margin when determining the object's location. An example spatial region may include the distance 826 of the object shown in Figure 8-2. In some embodiments, ultrasound sensor 216 may directly measure the object's location by analyzing the distance-profile data 716. Alternatively or additionally, ultrasound sensor 216 may refer to sensor data 308 provided by sensor 218 to determine the object's location.
[0150] At 1310, a demonstration attack is detected based on a feature. The demonstration attack attempts to deceive a facial authentication system. The object is associated with the demonstration attack. For example, the ultrasonic sensor 216 detects the demonstration attack 114 based on feature 724 (e.g., or determines the presence of the demonstration attack device 404), as described with respect to Figures 8-1 and 8-2.
[0151] Demonstration attack 114 attempts to deceive facial recognition system 202. The object is associated with demonstration attack 114 and represents demonstration attack device 404. Demonstration attack device 404 may contain a photograph 118 of one of the users 106, a digital image 124 of one of the users 106, or a mask 126 that replicates one or more features of the user 106.
[0152] At 1312, protection is provided against authentication demonstration attacks on the face authentication system. For example, the ultrasonic sensor 216 prevents authentication demonstration attacks 114 on the face authentication system 202. Specifically, the ultrasonic sensor 216 provides a spoofing indicator 304 to the face recognizer 214, which causes the face recognizer 214 to report 306 a face authentication failure.
[0153] In a variant of method 1300, block 1310 may alternatively determine, based on feature 724, that no demonstration attack 114 exists (e.g., the object reflecting the ultrasonic transmission signal 410 to generate the ultrasonic reception signal 412 is a face). In this case, at 1312, face authentication system 202 may authenticate face 408 to allow user 106 access.
[0154] By analyzing distance-contour data, the ultrasonic sensor 216 can perform face authentication anti-spoofing using as few as one receiver channel 608. This can be useful for space-constrained devices in which an ultrasonic sensor 216 with a single receiver channel 608 is implemented. If additional receiver channels 608 are available, the ultrasonic sensor 216 can evaluate features 724 associated with one or more receiver channels and, as needed, with each receiver channel 608 to improve its ability to correctly detect demonstration attacks 114.
[0155] At 1402 in Figure 14, an ultrasonic sensor is used to transmit an ultrasonic transmission signal. For example, ultrasonic sensor 216 transmits ultrasonic transmission signal 410, as described in Figure 13 with respect to 1302.
[0156] At 1404, at least two sensors of an ultrasound sensor receive at least two ultrasound received signals. The at least two ultrasound received signals include respective versions of an ultrasound transmission signal reflected by an object. For example, ultrasound sensor 216 uses at least two of microphones 504-1 to 504-M to receive at least two of ultrasound received signals 412-1 to 412-M. Ultrasound received signals 412-1 to 412-M are respective versions of an ultrasound transmission signal 410 reflected by an object (e.g., a demonstration attack device 404 or a face 408). Ultrasound received signals 412-1 to 412-M may have an amplitude different from the ultrasound transmission signal 410 and / or may be shifted in phase and / or frequency. In some embodiments, ultrasound sensor 216 receives ultrasound received signals 412-1 to 412-M during at least a portion of the time that the ultrasound transmission signal 410 is transmitted.
[0157] At 1406, an interferogram is generated based on at least two ultrasonic received signals. The interferogram includes coherence information and phase information. For example, the spoofing detector 230 of the ultrasonic sensor 216 generates an interferogram 718 based on at least two ultrasonic received signals 412-1 to 412-M, as shown in FIG9-1. Specifically, the preprocessor 702 of the spoofing detector 230 generates the interferogram 718 by combining information from distance-profile data 716 associated with both received channels 608-1 to 608-M. The interferogram 718 includes coherence information 936 and phase information 934, as shown in FIG9-2.
[0158] At 1408, a homology feature is identified based on the homology information of the interferogram. This homology feature represents a homology quantity within a region of interest. For example, the spoofing detector 230 of the ultrasonic sensor 216 identifies (e.g., extracts or generates) homology feature 914, which represents a homology quantity within a region of interest 938. Homology feature 914 may indicate a significantly large homology quantity when the object is associated with a demonstration attack instrument 404 (as shown in interferograms 930 or 932) or a substantially small homology quantity when the object is associated with a face 408 (as shown in interferogram 928). An example homology feature 914 may represent an average homology quantity within a region of interest 938. Another example homology feature 914 represents a quantity of homology variation within a region of interest 938.
[0159] At 1410, a demonstration attack is detected based on a homology feature. This demonstration attack attempts to deceive a face authentication system. The object is associated with the demonstration attack. For example, ultrasound sensor 216 detects demonstration attack 114 based on homology feature 914, as shown in Figure 9-1. Demonstration attack 114 attempts to deceive face authentication system 202. The object is associated with demonstration attack 114.
[0160] At 1412, a face authentication system demonstration attack 114 is prevented. For example, an ultrasonic sensor 216 prevents the face authentication system 202 from performing a face authentication demonstration attack 114. Specifically, the ultrasonic sensor 216 provides a spoofing indicator 304 to the face recognizer 214, which causes the face recognizer 214 to report 306 a face authentication failure.
[0161] In a variant of method 1400, block 1410 may alternatively determine, based on the coherence feature 914, that there is no demonstration attack 114 (e.g., the object reflecting the ultrasonic transmission signal 410 to generate the ultrasonic reception signal 412 is a face). In this case, at 1412, face authentication system 202 can authenticate face 408 to allow user 106 access.
[0162] At 1502 in Figure 15, an ultrasonic transmission signal is transmitted. For example, ultrasonic sensor 216 transmits ultrasonic transmission signal 410, as described in Figure 13 with respect to 1302.
[0163] At 1504, at least two sensors of an ultrasound sensor receive at least two ultrasound received signals. These at least two ultrasound received signals include respective versions of an ultrasound transmission signal reflected from an object. For example, ultrasound sensor 216 uses at least two of microphones 504-1 to 504-M to receive at least two of ultrasound received signals 412-1 to 412-M, as described in Figure 14 with respect to 1404.
[0164] At 1506, a power spectrum is generated based on at least two ultrasonic received signals. The power spectrum represents the power of at least two ultrasonic received signals within a set of frequencies and a time interval. For example, the deception detector 230 of the ultrasonic sensor 216 generates a power spectrum 722, as shown in Figure 12-1. The power spectrum 722 represents the power of at least two ultrasonic received signals within a set of frequencies and a time interval.
[0165] At 1508, the variation of power over time within the power spectrum is determined. These variations are associated with at least two ultrasonic received signals. For example, the spoofing detector 230 generates a variation feature 1208 representing a standard deviation of power over time within the power spectrum. Instance standard deviation graphs are shown in Figures 12-2 and 12-3 for different types of objects.
[0166] At 1510, a demo attack is detected based on the mutation number. This demo attack attempts to deceive a face authentication system. The object is associated with the demo attack. For example, the deception detector 230 detects demo attack 114 based on mutation feature 1208, as shown in Figure 12-1. Demo attack 114 attempts to deceive face authentication system 202. The object is associated with demo attack 114.
[0167] At 1512, protection is provided against authentication demonstration attacks on the face authentication system. For example, the ultrasonic sensor 216 prevents authentication demonstration attacks 114 on the face authentication system 202. Specifically, the ultrasonic sensor 216 provides a spoofing indicator 304 to the face recognizer 214, which causes the face recognizer 214 to report 306 a face authentication failure.
[0168] In a variant of method 1500, block 1510 may alternatively determine, based on the mutation feature 1208, that no demonstration attack 114 exists (e.g., the object reflecting the ultrasonic transmission signal 410 to generate the ultrasonic reception signal 412 is a face). In this case, at 1512, face authentication system 202 can authenticate face 408 to allow user 106 access.
[0169] The operations in Figures 13 to 15 can be combined in various ways. Generally, the ultrasonic sensor 216 can use any combination of single-channel feature 726 or multi-channel feature 728 to perform face authentication anti-spoofing. Although employing additional techniques for face authentication anti-spoofing may require additional computing resources, it can increase the accuracy of the ultrasonic sensor 216 in correctly distinguishing between a face 408 and a demonstration attack device 404. Furthermore, performing multiple techniques can better enable the ultrasonic sensor 216 to distinguish between a face 408 wearing an accessory 110 and a demonstration attack 114. (Example computing system)
[0170] Figure 16 illustrates various components of an instance computing system 1600 that can be implemented as any type of client, server and / or user device 102 as described with reference to the previous Figure 2-1 to perform face authentication anti-counterfeiting using an ultrasonic sensor 216.
[0171] The computing system 1600 includes a communication device 1602 that implements wired and / or wireless communication of device data 1604 (e.g., received data, data being received, data scheduled for broadcast, or data packets). The communication device 1602 or the computing system 1600 may include one or more ultrasonic sensors 216 and one or more sensors 218. The device data 1604 or other device content may include device configuration settings, media content stored on the device, and / or information associated with a user 106 of the device. The media content stored on the computing system 1600 may include any type of audio, video, and / or image data. The computing system 1600 includes one or more data inputs 1606 that can receive any type of data, media content and / or input (including human speech, input from the ultrasound sensor 216, user-selectable input (explicit or implicit), messages, music, television media content, recorded video content, and any other type of audio, video and / or image data received from any content and / or data source).
[0172] The computing system 1600 also includes a communication interface 1608, which may be implemented as a serial and / or parallel interface, a wireless interface, any type of network interface, one or more of a modem, and any other type of communication interface. The communication interface 1608 provides a connection and / or communication link between the computing system 1600 and a communication network (through which other electronic, computing and communication devices transmit data to the computing system 1600).
[0173] The computing system 1600 includes one or more processors 1610 (e.g., any of a microcontroller, controller, and the like) that process various computer-executable instructions to control the operation of the computing system 1600 and implement technologies for or in which ultrasonic facial authentication anti-spoofing can be embodied. Alternatively or additionally, the computing system 1600 may be implemented using any or a combination of hardware, firmware, or fixed logic circuitry systems implemented in conjunction with processing and control circuitry typically identified as 1612. Although not shown, the computing system 1600 may include a system bus or data transmission system of various components within a coupling device. A system bus may include any or a combination of different bus architectures, including a memory bus or memory controller, a peripheral bus, a general purpose serial bus, and / or a processor or local bus utilizing any of various bus architectures.
[0174] The computing system 1600 also includes a computer-readable medium 1614, comprising one or more memory devices that implement persistent and / or non-transitory data storage (i.e., compared to signal transmission only). Examples of the one or more memory devices include random access memory (RAM), non-volatile memory (e.g., one or more of a read-only memory (ROM), flash memory, EPROM, EEPROM, etc.), and a magnetic disk storage device. The magnetic disk storage device can be implemented as any type of magnetic or optical storage device, including a hard disk drive, a recordable and / or rewritable optical disc (CD), any type of digital versatile disc (DVD), and the like. The computing system 1600 may also include a large-capacity storage media device (storage media) 1616.
[0175] The computer-readable medium 1614 provides a data storage mechanism for storing device data 1604, various device applications 1618, and any other type of information and / or data related to the operating state of the computing system 1600. For example, an operating system 1620 may be maintained as a computer application having the computer-readable medium 1614 and executing on the processor 1610. The device application 1618 may include a device manager, including any form of control application, software application, signal processing and control module, code native to a particular device, a hardware abstraction layer of a particular device, etc. Using the ultrasonic sensor 216, the computing system 1600 can perform face authentication anti-spoofing using interferometric coherence. Summary
[0176] Although the technology and apparatus including the ultrasonic sensor for performing face authentication anti-counterfeiting using interferometric coherence have been described in language specific to features and / or methods, it should be understood that the subject matter of the appended claims is not necessarily limited to the specific features or methods described. In fact, the specific features and methods are disclosed as exemplary embodiments of face authentication anti-counterfeiting using interferometric coherence.
[0177] The following describes some examples.
[0178] Example 1: A method performed by an ultrasound sensor, the method comprising: transmitting an ultrasound transmission signal; receiving an ultrasound reception signal, the ultrasound reception signal including a version of the ultrasound transmission signal reflected by an object; generating distance-profile data based on the ultrasound reception signal, the distance-profile data including amplitude and phase information associated with the ultrasound reception signal; determining a feature of the distance-profile data within a region of interest associated with a location of the object; detecting a demo attack based on the feature, the demo attack attempting to deceive a face authentication system, the object being associated with the demo attack; and preventing the face authentication system from authenticating the demo attack.
[0179] Example 2: The method of Example 1, wherein: the determination of the feature includes detecting a peak amplitude in the region of interest of the distance-contour data; and the detection of the demonstration attack includes: comparing the peak amplitude with a threshold value; and detecting the demonstration attack in response to the peak amplitude being higher than the threshold value.
[0180] Example 3: The method of Example 2 further includes: transmitting another ultrasonic transmission signal; receiving another ultrasonic reception signal, the other ultrasonic reception signal including a version of the other ultrasonic transmission signal reflected by a face; generating other distance-contour data based on the other ultrasonic reception signal, the other distance-contour data including amplitude and phase information associated with the other ultrasonic reception signal; detecting another peak amplitude in another region of interest in the other distance-contour data, the other region of interest being associated with a location of the face; comparing the other peak amplitude with the threshold value; and in response to the other peak amplitude being less than the threshold value, enabling the face authentication system to authenticate the face.
[0181] Example 4: The method of any of the foregoing examples, wherein: the determination of the feature includes determining one of the energy distributions in the region of interest; and the detection of the demonstration attack includes detecting the demonstration attack in response to the energy distribution being greater than another threshold value.
[0182] Example 5: The method of any of the foregoing examples further includes: determining the location of the object, wherein the region of interest includes the location of the object.
[0183] Example 6: The method of Example 5 further includes: receiving sensor data from a sensor, wherein determining the location of the object includes determining the location of the object based on the sensor data.
[0184] Example 7: The method of Example 6, wherein the sensor includes: a phase difference sensor of a camera; or a radio frequency sensor.
[0185] Example 8: The method of any one of Examples 5 to 7 further includes: filtering the distance-contour data to extract the amplitude and phase information associated with the region of interest before determining the feature.
[0186] Example 9: The method of any of the foregoing examples, wherein the transmission of the ultrasonic transmission signal includes generating an ultrasonic transmission signal having a bandwidth that results in a distance resolution of the distance-profile data of less than approximately 7 cm.
[0187] Example 10: The method of any of the foregoing examples further includes: determining a distance to the object; and normalizing the amplitude information in the distance-profile data based on the distance to the object before detecting the demonstration attack.
[0188] Example 11: The method of any one of Examples 1 to 9 further includes: normalizing the amplitude information in the distance-profile data based on a cross-coupling factor or a constant false alarm rate signal-to-noise ratio before detecting the demonstration attack.
[0189] Example 12: The method of any of the foregoing examples further includes: receiving motion data from an inertial measurement unit, the motion data representing the motion of the ultrasonic sensor; and compensating for the motion of the ultrasonic sensor by modifying the distance-profile data based on the motion data.
[0190] Example 13: The method of any one of Examples 1 to 11 further includes: receiving motion data from an inertial measurement unit; and determining, based on the motion data, that the ultrasonic sensor is substantially stationary, wherein transmitting the ultrasonic transmission signal includes transmitting the ultrasonic transmission signal in response to determining that the ultrasonic sensor is substantially stationary.
[0191] Example 14: An apparatus comprising an ultrasonic sensor configured to perform any of the methods of Examples 1 to 13.
[0192] Example 15: A computer-readable medium including instructions that, when executed by a processor, cause an ultrasonic sensor to perform any of the methods of Examples 1 to 13.
[0193] Example 16: A method performed by an ultrasound sensor, the method comprising: transmitting an ultrasound transmission signal; receiving at least two ultrasound reception signals using at least two sensors of the ultrasound sensor, the at least two ultrasound reception signals including respective versions of the ultrasound transmission signal reflected by an object; generating an interferogram based on the at least two ultrasound reception signals, the interferogram including coherence information and phase information; identifying a coherence feature based on the coherence information of the interferogram, the coherence feature representing a coherence quantity in a region of interest; detecting a demo attack based on the coherence feature, the demo attack attempting to deceive a face authentication system, the object being associated with the demo attack; and preventing the face authentication system from authenticating the demo attack.
[0194] Example 17: The method of Example 16, wherein detecting the demonstration attack includes responding to the homology quantity being greater than a homology threshold value, and detecting the demonstration attack.
[0195] Example 18: The method of Example 17, wherein detecting the demonstrator attack includes: transmitting another ultrasonic transmission signal; receiving at least two other ultrasonic reception signals, the at least two other ultrasonic reception signals including a version of the other ultrasonic transmission signal reflected by a face; generating another interferogram based on the at least two other ultrasonic reception signals, the other interferogram including coherence and phase information; identifying another coherence feature based on the coherence information of the other interferogram, the other coherence feature representing another coherence amount in another region of interest; and enabling the face authentication system to authenticate the face in response to the other coherence amount being less than the coherence threshold.
[0196] Example 19: The method of any one of Examples 16 to 18, wherein the detection of the demonstration attack includes: identifying a phase feature based on phase information of the interferogram, the phase feature representing a quantity of phase variation in the region of interest; and detecting the demonstration attack in response to the quantity of phase variation being less than a phase threshold.
[0197] Example 20: The method of any one of Examples 16 to 19 further includes: generating range-profile data based on the at least two ultrasonic received signals; and performing co-registration on the range-profile data to generate co-registered range-profile data, wherein generating the interferogram includes generating the interferogram based on the co-registered range-profile data.
[0198] Example 21: The method of Example 20, wherein: the distance-profile data includes: first distance-profile data associated with a first ultrasound receiving signal, one of the at least two ultrasound receiving signals; and second distance-profile data associated with a second ultrasound receiving signal, one of the at least two ultrasound receiving signals; and the respective responses of the co-registration execution to align the object within the first distance-profile data and the second distance-profile data along a distance dimension.
[0199] Example 22: The method of Example 21, wherein the execution of the co-registration includes: filtering the first distance-profile data through a first low-pass filter; and filtering the second distance-profile data through a second low-pass filter.
[0200] Example 23: The method of Example 21, wherein the execution of the co-registration includes: detecting a first peak amplitude within the first distance-profile data; determining a first distance bin associated with the first peak amplitude; detecting a second peak amplitude within the second distance-profile data; determining a second distance bin associated with the second peak amplitude; determining an offset based on a difference between the first distance bin and the second distance bin; and shifting the second distance-profile data based on the offset to generate shifted second distance-profile data.
[0201] Example 24: The method of Example 23, wherein the execution of the co-registration includes: determining another offset between the first distance-profile data and the shifted second distance-profile data; and interpolating the shifted second distance-profile data based on the other offset to generate interpolated and shifted second distance-profile data.
[0202] Example 25: The method of Example 21, wherein the execution of the common registration includes: resampling the second distance-profile data to project the second distance-profile data onto a common grid associated with the first distance-profile data.
[0203] Example 26: The method of any one of Examples 20 to 25, wherein the detection of the demonstration attack includes: generating a power distribution histogram based on the distance-profile data, the equal power distribution histogram being associated with the at least two ultrasonic received signals respectively; and detecting the demonstration attack based on the skewness of the equal power distribution histogram.
[0204] Example 27: The method of any one of Examples 20 to 26, wherein the detection of the demonstration attack includes: generating a constant false alarm rate (FAR) distribution histogram based on the distance-profile data, the constant FAR distribution histograms being associated with the at least two ultrasonic received signals respectively; and detecting the demonstration attack based on the skewness of the constant FAR distribution histograms.
[0205] Example 28: The method of Example 27, wherein the shape of the histogram of the constant false alarm rate signal-to-noise ratio distribution exhibits one or more of the following: a non-Gaussian distribution; or a distribution with a tail.
[0206] Example 29: An apparatus comprising an ultrasonic sensor configured to perform any of the methods of Examples 16 to 28.
[0207] Example 30: The device as in Example 29, wherein: the device includes a smartphone; the ultrasonic sensor is integrated into the smartphone; and at least two sensors of the ultrasonic sensor include: a first microphone of the smartphone; and a second microphone of the smartphone.
[0208] Example 31: The device as in Example 30, wherein the first microphone and the second microphone are located at opposite ends of the smartphone.
[0209] Example 32: The device of any of Examples 29 to 31, wherein a distance between the at least two sensors is greater than a wavelength associated with an ultrasonic transmission signal.
[0210] Example 33: A computer-readable medium including instructions that, when executed by a processor, cause an ultrasonic sensor to perform any of the methods of Examples 16 to 28.
[0211] Example 34: A method performed by an ultrasound sensor, the method comprising: transmitting an ultrasound transmission signal; receiving at least two ultrasound reception signals using at least two sensors of the ultrasound sensor, the at least two ultrasound reception signals including respective versions of the ultrasound transmission signal reflected by an object; generating a power spectrum based on the at least two ultrasound reception signals, the power spectrum representing the power of the at least two ultrasound reception signals over a set of frequencies and a time interval; determining a power variation over time within the power spectrum, the variation being associated with the at least two ultrasound reception signals respectively; detecting a demo attack based on the variation, the demo attack attempting to deceive a face authentication system, the object being associated with the demo attack; and preventing the face authentication system from authenticating the demo attack.
[0212] Example 35: The method of Example 34, wherein generating the equal power spectrum includes: generating complex data based on the at least two ultrasonic received signals; and performing a Fourier transform to generate the equal power spectrum.
[0213] Example 36: The method of Example 34 or 35, wherein the complex data includes: distance-profile data; distance-slow time data; or an interferogram.
[0214] Example 37: The method of any one of Examples 34 to 36, wherein detecting the demonstration attack includes: determining whether the variants are in a set of values; and detecting the demonstration attack in response to at least one of the variants being outside the set of values.
[0215] Example 38: The method of Example 37, wherein detecting the demonstrator attack includes: transmitting another ultrasonic transmission signal; receiving at least two other ultrasonic reception signals, the at least two other ultrasonic reception signals including a version of the other ultrasonic transmission signal reflected by a face; generating other power spectra based on the at least two other ultrasonic reception signals, the other power spectra representing the power of the at least two other ultrasonic reception signals within a set of frequencies and another time interval; determining other variations of the power over time within the other power spectra, the other variations being associated with the at least two ultrasonic reception signals respectively; and responding to the other variations within the set of values, enabling the face authentication system to authenticate the face.
[0216] Example 39: The method of Example 38, wherein the face wears an accessory.
[0217] Example 40: A method as described in any one of Examples 34 to 39, wherein: the power spectrum includes: a first power spectrum associated with a first ultrasonic received signal, one of the at least two ultrasonic received signals; and a second power spectrum associated with a second ultrasonic received signal, one of the at least two ultrasonic received signals; and determining the variances includes: generating a first set of subframes of the first power spectrum using a sliding window, the subframes of the first set of subframes being associated with different portions of the time interval; calculating a first standard deviation across the first set of subframes; generating a second set of subframes of the second power spectrum using the sliding window, the subframes of the second set of subframes being associated with the different portions of the time interval; and calculating a second standard deviation across the second set of subframes.
[0218] Example 41: The method of any one of Examples 1 to 13, 16 to 28 or 34 to 40, wherein the demonstration attack includes: an unauthorized actor presenting a photograph of an authorized user; the unauthorized actor presenting a device displaying a digital image of the authorized user; or the unauthorized actor wearing a mask representing the authorized user.
[0219] Example 42: A device comprising an ultrasonic sensor configured to perform any of the methods of Examples 34 to 41.
[0220] Example 43: The device as in Example 42, wherein: the device includes a smartphone; the ultrasonic sensor is integrated into the smartphone; and at least two sensors of the ultrasonic sensor include: a first microphone of the smartphone; and a second microphone of the smartphone.
[0221] Example 44: The device as in Example 43, wherein the first microphone and the second microphone are located at opposite ends of the smartphone.
[0222] Example 45: The device of any of Examples 42 to 44, wherein a distance between the at least two sensors is greater than a wavelength associated with the ultrasonic transmission signal.
[0223] Example 46: A computer-readable medium including instructions that, when executed by a processor, cause an ultrasonic sensor to perform any of the methods of Examples 34 to 41. [Simplified Explanation of the Diagram]
[0007] Refer to the following diagram to describe the device and technology used for anti-counterfeiting facial authentication based on interferometric homology.Throughout the diagrams, the same numbers are used to refer to the same features and components: Figure 1 illustrates an example environment in which face authentication anti-counterfeiting based on interferometric coherence can be implemented; Figure 2-1 illustrates an example implementation of a face authentication system as part of a user device; Figure 2-2 illustrates an example component of an ultrasonic sensor used for face authentication anti-counterfeiting; Figure 2-3 illustrates an example sensor used for face authentication anti-counterfeiting; Figure 3 illustrates an example face authentication system implementing face authentication anti-counterfeiting using ultrasound; Figure 4-1 illustrates the difference in ultrasound reflection between a demonstration attack device and a face; Figure 4-2 illustrates the difference in received power between a demonstration attack device and a face; Figure 5 illustrates an example location of a speaker and microphone in a user device; Figure 6 illustrates an example implementation of an ultrasonic sensor for face authentication anti-counterfeiting; Figure 7 illustrates an example scheme for face authentication anti-counterfeiting implemented using an ultrasonic sensor; Figure 8-1 illustrates an example scheme for generating a single-channel feature for face authentication anti-counterfeiting using an ultrasonic sensor; Figure 8-2 illustrates example distance-contour data associated with face authentication anti-counterfeiting; Figure 9-1 illustrates an example scheme for generating a multi-channel feature for face authentication anti-counterfeiting using an ultrasonic sensor; Figure 9-2 illustrates an example interferogram associated with face authentication anti-counterfeiting; Figure 10-1 illustrates an example scheme for performing co-registration using an ultrasonic sensor; Figure 10-2 illustrates another example scheme for performing co-registration using an ultrasonic sensor; Figure 10-3 illustrates an additional example scheme for performing co-registration using an ultrasonic sensor; Figure 10-4 illustrates yet another example scheme for performing co-registration using an ultrasonic sensor. Figure 11-1 illustrates an example scheme for generating another single-channel or multi-channel feature for face authentication anti-counterfeiting using an ultrasonic sensor; Figure 11-2 illustrates an example power distribution histogram associated with face authentication anti-counterfeiting; Figure 11-3 illustrates an example constant false alarm rate (FAR) histogram associated with face authentication anti-counterfeiting; Figure 12-1 illustrates an example scheme for generating yet another single-channel or multi-channel feature for face authentication anti-counterfeiting using an ultrasonic sensor; Figure 12-2 illustrates a first set of standard deviation curves associated with face authentication anti-counterfeiting; Figure 12-3 illustrates a second set of standard deviation curves associated with face authentication anti-counterfeiting; Figure 13 illustrates an example method for performing face authentication anti-counterfeiting using ultrasound; Figure 14 illustrates another example method for performing face authentication anti-counterfeiting using coherence based on interferometry; Figure 15 illustrates an additional example method for implementing face authentication anti-counterfeiting using variance; and Figure 16 illustrates an example computational system embodying, or in which, the use of face authentication anti-counterfeiting based on interferometric homology can be implemented.
Claims
1. A method performed by an ultrasound sensor, the method comprising: Transmit an ultrasonic signal; The method involves using at least two sensors of an ultrasonic sensor to receive at least two ultrasonic received signals, each including a version of the ultrasonic transmitted signal reflected by an object; generating an interferogram based on the at least two ultrasonic received signals, the interferogram including coherence information and phase information; identifying a coherence feature based on the coherence information of the interferogram, the coherence feature representing a coherence quantity within a region of interest; detecting a demo attack based on the coherence feature, the demo attack attempting to deceive a facial recognition system, the object being associated with the demo attack; and preventing the facial recognition system from authenticating the demo attack.
2. The method of request item 1, wherein detecting the demonstration attack includes detecting the demonstration attack in response to the quantity of the homology being greater than a homology threshold.
3. As in request item 2, wherein detecting the demonstration attack includes: Transmit another ultrasonic transmission signal; The system receives at least two other ultrasound received signals, including a version of the other ultrasound transmitted signal reflected by a face; generates another interferogram based on the at least two other ultrasound received signals, the other interferogram including coherence and phase information; identifies another coherence feature based on the coherence information of the other interferogram, the other coherence feature representing another coherence quantity in another region of interest; and enables the face authentication system to authenticate the face in response to the other coherence quantity being less than the coherence threshold.
4. The method described in any of requests 1 to 3, wherein detecting the demonstration attack includes: Based on the phase information of the interferogram, a phase feature is identified, which represents a quantity of phase variation within the region of interest; and the demonstration attack is detected in response to the quantity of phase variation being less than a phase threshold.
5. The method of any one of claims 1 to 3 further includes: Range-profile data is generated based on at least two ultrasonic received signals; And perform co-registration on the distance-profile data to generate co-registered distance-profile data, wherein generating the interferogram includes generating the interferogram based on the co-registered distance-profile data.
6. As in request item 5, wherein: The distance-profile data includes: first distance-profile data associated with a first ultrasound receiving signal, one of the at least two ultrasound receiving signals; and second distance-profile data associated with a second ultrasound receiving signal, one of the at least two ultrasound receiving signals; and the respective responses of the co-registration execution along a distance dimension to align the object within the first distance-profile data and the second distance-profile data.
7. The method of request item 6, wherein the execution of the common registration includes: The first distance-profile data is filtered through a first low-pass filter; and the second distance-profile data is filtered through a second low-pass filter.
8. The method of request item 6, wherein the execution of the common registration includes: Detect a first peak amplitude within the first distance-profile data; determine a first distance frame associated with the first peak amplitude; Detect a second peak amplitude within the second distance-profile data; determine a second distance frame associated with the second peak amplitude; determine an offset based on a difference between the first distance frame and the second distance frame; and shift the second distance-profile data based on the offset to generate shifted second distance-profile data.
9. The method of request item 8, wherein the execution of the common registration includes: Determine another offset between the first distance-profile data and the shifted second distance-profile data; and based on the other offset, interpolate the shifted second distance-profile data to generate interpolated and shifted second distance-profile data.
10. The method of request item 6, wherein the execution of the common registration includes: The second distance-profile data is resampled to project the second distance-profile data onto one of the common grids associated with the first distance-profile data.
11. The method of request item 5, wherein detecting the demonstration attack includes: A power distribution histogram is generated based on the distance-profile data, and the equal power distribution histograms are respectively associated with the at least two ultrasonic received signals; And the demonstration attack was detected based on the skewness of the equal power distribution histogram.
12. The method of request item 5, wherein detecting the demonstration attack includes: The distance-profile data is used to generate constant false alarm rate (FAR) histograms, which are associated with the at least two ultrasonic received signals; and the skewness of the constant FAR histograms is used to detect the demonstration attack.
13. The method of claim 12, wherein the shape of the histogram of the constant false alarm rate signal-to-noise ratio distribution exhibits one or more of the following: a non-Gaussian distribution; or a distribution having a tail.
14. An apparatus comprising an ultrasonic sensor configured to perform any of the methods described in claims 1 to 13.
15. The device as requested in item 14, wherein: The device includes a smartphone; The ultrasonic sensor is integrated into the smartphone; The ultrasonic sensor includes at least two sensors, including: a first microphone of the smartphone; And one of the second microphones of the smartphone.
16. The device of claim 15, wherein the first microphone and the second microphone are located at opposite ends of the smartphone.
17. The device of any one of claims 14 to 16, wherein a distance between the at least two sensors is greater than a wavelength associated with the ultrasonic transmission signal.
18. A computer-readable medium including instructions that, when executed by a processor, cause an ultrasonic sensor to perform any of the methods described in claims 1 to 13.